إحصائية I2 كاختبار لتحيز الاختيار: حدود العتبة المحدثة
The I2 Statistic As Selection Bias Test: Updated Threshold Limits

شارك:
المجلة: Cureus، المجلد: 18، العدد: 4
DOI: https://doi.org/10.7759/cureus.107553
PMID: https://pubmed.ncbi.nlm.nih.gov/42181436
تاريخ النشر: 2026-04-22
المؤلف: Steffen Mickenautsch وآخرون
الموضوع الرئيسي: طرق إحصائية في التجارب السريرية

نظرة عامة

هدفت هذه الدراسة إلى مراجعة حدود تقدير نقطة I² للاختبار القائم على التجارب المقارنة المحاكية (SCT) المعتمد على التجارب المعدلة، لتعزيز نسبة التجارب العشوائية المضبوطة (RCTs) القابلة للاختبار من حيث التحيز في الاختيار. اختبرت البحث فرضيتين صفريتين: (H01) أن حجم تقديرات تأثير التجربة ليس له ارتباط إيجابي كبير مع مستويات التحيز في الاختيار، و(H02) أن تقديرات التأثير لا تختلف بشكل كبير بين RCTs ذات التحيز المنخفض والعالي. كانت المراجعة قائمة على المحاكاة وإعادة تقييم 332 RCTs من العالم الحقيقي، مما أدى إلى تحديد مستويات التحيز في الاختيار (B%) باستخدام الحدود الجديدة.

أشارت النتائج إلى أن جميع 332 RCTs كانت قابلة للاختبار، مما رفع نسبة التجارب القابلة للاختبار من 71% إلى 100%. وُجد ارتباط إيجابي ذو دلالة إحصائية (معامل سبيرمان = 0.268، p < 0.0001) بين قيم الفرق في المخاطر المطلقة (RD) ومستويات B%، حيث أظهرت RCTs ذات التحيز العالي في الاختيار متوسط RD مطلق قدره 0.18 (SD = 0.16) مقارنة بـ 0.10 (SD = 0.13) لتلك ذات التحيز المنخفض (t = -4.65؛ p < 0.0001؛ الفرق المتوسط 0.08؛ 95% CI: 0.05 - 0.11). أشار حجم التأثير (d لكوهين = 0.52) إلى تأثير متوسط، مما أدى إلى رفض كلا الفرضيتين الصفريتين. تؤكد النتائج أن العلاقة بين مستويات التحيز في الاختيار وتقديرات التأثير مستقلة عن تغييرات حدود I²، مما يثبت فائدة الحدود المعدلة دون تقويض تقييمات التحيز السابقة.

مقدمة

تناقش المقدمة قضية التحيز في الاختيار في التجارب العشوائية المضبوطة (RCTs)، والتي يمكن أن تنشأ من تخصيص غير عشوائي للمرضى بناءً على خصائص تتنبأ بنجاح العلاج. يمكن أن يؤدي هذا التحيز إلى تحريف النتائج، مما يجعل علاجًا واحدًا يبدو أكثر فعالية مما هو عليه في الحقيقة. طور ميكنوتش ويينغوبال اختبارًا لاكتشاف التحيز في الاختيار باستخدام بيانات خط الأساس من مجموعات العلاج، مستخدمين إحصائية I² للإشارة إلى المخاطر وتقدير نسبة الأفراد المخصصين بشكل غير عشوائي. كما أظهر كلارك وآخرون أن زيادة قيمة I² فوق 0% تشير إلى اختلالات تعزى إلى التحيز في الاختيار، بينما أنشأ هيكس وآخرون طريقة لتحليل نتائج الميتا التي تأخذ أيضًا في الاعتبار قياسات خط الأساس.

يقترح المؤلفون نهجًا مبتكرًا يتضمن التجارب المقارنة المحاكية (SCTs) جنبًا إلى جنب مع بيانات خط الأساس من RCT واحد لتعزيز اكتشاف التحيز في الاختيار. تشير نتائجهم إلى أن أحجام عينات التجارب المتنوعة تؤثر على تقديرات I² تحت مستويات مختلفة من التحيز المحاكي في الاختيار، مما يسمح بتصنيف تجريبي لمستويات التحيز. تهدف الدراسة إلى مراجعة حدود I² لتحسين تقييم RCTs من حيث التحيز في الاختيار، مع معالجة القيود التي تم مواجهتها في التطبيقات السابقة للاختبار. بالإضافة إلى ذلك، يعتزم المؤلفون اختبار فرضيتين صفريتين بشأن الارتباط بين تقديرات تأثير التجربة ومستويات التحيز في الاختيار، فضلاً عن الاختلافات في تقديرات التأثير بين RCTs ذات التحيز المنخفض مقابل العالي.

طرق

توضح قسم “المواد والطرق” تصميم التجربة والإجراءات المستخدمة في الدراسة. يتناول المواد المحددة المستخدمة، بما في ذلك أي مواد كيميائية، معدات، وعينات بيولوجية، بالإضافة إلى مصادرها وطرق تحضيرها. كما يصف القسم المنهجيات المطبقة لجمع البيانات وتحليلها، مما يضمن إمكانية تكرار النتائج.

تُبرز البروتوكولات التجريبية الرئيسية، بما في ذلك أي طرق إحصائية مستخدمة لتحليل البيانات، مثل تحليل الانحدار أو ANOVA، والمعايير للدلالة. يؤكد القسم على أهمية الضوابط والتكرارات في التحقق من النتائج، مما يضمن أن النتائج قوية وموثوقة. بشكل عام، يوفر هذا القسم نظرة شاملة على الإطار التجريبي الذي يدعم استنتاجات الدراسة.

نتائج

في هذه الدراسة، تم تقييم جميع 332 تجربة عشوائية مضبوطة (RCTs) باستخدام حدود الاختبار المعدلة للتحيز في الاختيار، دون أن تتجاوز أي من تقديرات نقطة I² المحسوبة هذه الحدود. أدت تطبيق الحدود الجديدة إلى تخصيص 12 RCT بتصنيف أعلى للتأثير في التحيز، وسبعة بتصنيف أقل، بينما حافظت 313 RCTs (94%) على تقييماتها الأصلية من التحليلات السابقة. وُجد ارتباط إيجابي كبير بين حجم تقديرات تأثير التجربة (الفرق في المخاطر المطلقة، RD) وفئات التحيز في الاختيار، حيث تم حساب معامل سبيرمان عند 0.268 (p < 0.0001). كشفت التحليلات أن RCTs المصنفة ذات التحيز العالي في الاختيار (B% 40-100) كان لديها متوسط RD قدره 0.18 (SD = 0.16)، مقارنة بـ 0.10 (SD = 0.13) لتلك ذات التحيز المنخفض (B% 0-30). كان الفرق في تقديرات التأثير ذو دلالة إحصائية، مع فرق متوسط (MD) قدره 0.08 (95% CI: 0.05-0.11؛ t = -4.65؛ p < 0.0001)، مما يشير إلى زيادة بمقدار ثماني نقاط مئوية مرتبطة بالتحيز الأعلى. كان حجم التأثير، المقاس بواسطة d لكوهين، 0.52، مما يشير إلى تأثير متوسط. وبالتالي، تم رفض كلا الفرضيتين الصفريتين (H01 و H02)، مما يبرز تأثير التحيز في الاختيار على نتائج التجارب.

مناقشة

في هذه الدراسة، سعى المؤلفون إلى مراجعة حدود تقدير نقطة I² للاختبار القائم على التحيز في الاختيار المعتمد على SCT لتعزيز نسبة التجارب العشوائية المضبوطة (RCTs) التي يمكن تقييمها من حيث التحيز في الاختيار. من خلال محاكاة 200 RCT بمستويات مختلفة من التحيز في الاختيار، أنشأ الباحثون حدودًا جديدة زادت من التجارب القابلة للاختبار من 71% إلى 100%. أكدت النتائج أن كلا الفرضيتين الصفريتين تم رفضهما، مما يشير إلى وجود ارتباط إيجابي كبير بين مستويات التحيز في الاختيار (B%) وتقديرات التأثير (الفرق في المخاطر المطلقة، RD)، فضلاً عن وجود فرق ذو دلالة إحصائية في RD بين RCTs ذات التحيز المنخفض والعالي.

تم تطوير الحدود المعدلة من محاكاة متعددة لـ RCTs، مع دمج نطاق أوسع من تباينات المتغيرات الأساسية مقارنة بالدراسات السابقة. سمح هذا التحسين المنهجي بتقدير أكثر دقة لحدود I² مع الحفاظ على سلامة الارتباط بين التحيز في الاختيار وتقديرات التأثير. من المهم أن الحدود الجديدة ألغت الاعتماد على التقديرات التقريبية لتقدير B%، والتي أدت سابقًا إلى عدم الدقة. تشير النتائج إلى أن حدود I² المعدلة أكثر صرامة ولكنها لا تلغي التقييمات السابقة التي تمت بموجب الحدود الأصلية، مما يعزز فائدة إحصائية I² في اختبار التحيز في الاختيار عبر RCTs.

القيود

تقدم الدراسة عدة قيود تؤثر على نتائجها. أحد القيود الرئيسية هو عدم وجود حساب رسمي لحجم العينة، والذي كان ضروريًا لتحديد العدد المناسب من التجارب العشوائية المضبوطة (RCTs) المحاكية لتعكس بشكل كافٍ التباينات في المتغيرات الأساسية التي تم مواجهتها في السيناريوهات الواقعية. تم تعيين حجم العينة بشكل عشوائي عند \( N = 50 \)، مضروبًا في ثلاثة مجالات متغيرة مستمدة من 332 RCTs من الحياة الواقعية، مما أدى إلى 165,000 نقطة بيانات \( I^2 \). على الرغم من أن حجم العينة هذا أثبت كفايته لتطوير حدود \( I^2 \)، فإن غياب نهج منهجي لتحديد حجم العينة قد يؤثر على قوة النتائج.

بالإضافة إلى ذلك، فإن اعتماد الدراسة على نطاقات المتغيرات الأساسية من مجموعة RCTs المختارة، التي تم تحديدها من خلال عملية مراجعة منهجية، يقدم تحيزات محتملة. على الرغم من أن التجارب المشمولة تمتد عبر مجموعة متنوعة من التخصصات الطبية وسنوات النشر (1985-2023)، تم استبعاد أي بيانات أساسية ذات صلة من التجارب خارج هذه المجموعة. تشير هذه القيود إلى أن الأبحاث المستقبلية يجب أن تركز على اختبار حدود \( I^2 \) الجديدة على RCTs إضافية غير مشمولة في هذه الدراسة، مما يسمح بمزيد من تحسين هذه الحدود. على الرغم من هذه القيود، تدعم النتائج فائدة اختبار \( I^2 \) القائم على SCT كطريقة كمية لاكتشاف التحيز في الاختيار في تخصيص المرضى ضمن RCTs واحدة، مما يتناقض مع الأساليب النوعية الحالية التي تفتقر إلى القدرة على تقييم التحيز بشكل كمي.

Journal: Cureus, Volume: 18, Issue: 4
DOI: https://doi.org/10.7759/cureus.107553
PMID: https://pubmed.ncbi.nlm.nih.gov/42181436
Publication Date: 2026-04-22
Author(s): Steffen Mickenautsch et al.
Primary Topic: Statistical Methods in Clinical Trials

Overview

This study aimed to revise the I² point estimate threshold limits of the trial-adjusted, simulated comparator trial (SCT)-based I² test to enhance the percentage of testable randomized controlled trials (RCTs) for selection bias. The research tested two null hypotheses: (H01) that the magnitude of trial effect estimates is not significantly positively correlated with selection bias levels, and (H02) that effect estimates do not differ significantly between RCTs with ‘low’ and ‘high’ selection bias. The revision was based on simulations and re-evaluation of 332 real-world RCTs, leading to a determination of selection bias levels (B%) using the new thresholds.

Results indicated that all 332 RCTs were testable, raising the proportion of testable trials from 71% to 100%. A statistically significant positive correlation (Spearman’s rho = 0.268, p < 0.0001) was found between absolute risk difference (RD) values and B% levels, with RCTs exhibiting 'high' selection bias showing a mean absolute RD of 0.18 (SD = 0.16) compared to 0.10 (SD = 0.13) for those with 'low' bias (t = -4.65; p < 0.0001; mean difference 0.08; 95% CI: 0.05 - 0.11). The effect size (Cohen's d = 0.52) indicated a medium effect, leading to the rejection of both null hypotheses. The findings confirm that the relationship between selection bias levels and effect estimates is independent of the I² threshold limit changes, thereby validating the utility of the revised thresholds without undermining previous selection bias ratings.

Introduction

The introduction discusses the issue of selection bias in randomized controlled trials (RCTs), which can arise from non-random allocation of patients based on characteristics that predict treatment success. This bias can skew outcomes, making one treatment appear more effective than it truly is. Mickenautsch and Yengopal developed a test to detect selection bias using baseline data from treatment groups, employing the I² statistic to indicate risk and estimate the percentage of subjects allocated non-randomly. Clark et al. further established that an increase in the I² value above 0% signifies imbalances attributable to selection bias, while Hicks et al. created a method for outcome meta-analyses that also considers baseline measurements.

The authors propose an innovative approach that incorporates simulated comparator trials (SCTs) alongside baseline data from a single RCT to enhance the detection of selection bias. Their findings indicate that varying trial sample sizes affect the I² estimates under different levels of simulated selection bias, allowing for empirical classification of bias levels. The study aims to revise the I² threshold limits to improve the evaluation of RCTs for selection bias, addressing limitations encountered in previous applications of the test. Additionally, the authors intend to test two null hypotheses regarding the correlation between trial effect estimates and selection bias levels, as well as differences in effect estimates between RCTs with low versus high selection bias.

Methods

The “Materials and Methods” section outlines the experimental design and procedures utilized in the study. It details the specific materials employed, including any reagents, equipment, and biological samples, as well as their sources and preparation methods. The section also describes the methodologies applied for data collection and analysis, ensuring reproducibility of the results.

Key experimental protocols are highlighted, including any statistical methods used to analyze the data, such as regression analysis or ANOVA, and the criteria for significance. The section emphasizes the importance of controls and replicates in validating the findings, ensuring that the results are robust and reliable. Overall, this section provides a comprehensive overview of the experimental framework that underpins the study’s conclusions.

Results

In this study, all 332 randomized controlled trials (RCTs) were evaluated using revised test thresholds for selection bias, with none of the computed I² point estimates exceeding these limits. The application of the new thresholds resulted in 12 RCTs being assigned a higher selection bias rating and seven a lower rating, while 313 RCTs (94%) maintained their original bias ratings from prior analyses. A significant positive correlation was found between the magnitude of trial effect estimates (absolute risk differences, RD) and the selection bias categories, with Spearman’s rho calculated at 0.268 (p < 0.0001). The analysis revealed that RCTs classified with 'high' selection bias (B% 40-100) had a mean RD of 0.18 (SD = 0.16), compared to 0.10 (SD = 0.13) for those with 'low' selection bias (B% 0-30). The difference in effect estimates was statistically significant, with a mean difference (MD) of 0.08 (95% CI: 0.05-0.11; t = -4.65; p < 0.0001), indicating an eight-percentage-point increase associated with higher bias. The effect size, measured by Cohen's d, was 0.52, suggesting a medium effect. Consequently, both null hypotheses (H01 and H02) were rejected, underscoring the impact of selection bias on trial outcomes.

Discussion

In this study, the authors aimed to revise the I² point estimate threshold limits of the trial-adjusted, SCT-based I² selection bias test to enhance the proportion of randomized controlled trials (RCTs) that can be evaluated for selection bias. By simulating 200 RCTs with varying levels of selection bias, the researchers generated new threshold limits that increased the testable RCTs from 71% to 100%. The results confirmed that both null hypotheses were rejected, indicating a significant positive correlation between selection bias levels (B%) and effect estimates (absolute risk difference, RD), as well as a statistically significant difference in RD between RCTs with low and high selection bias.

The revised thresholds were developed from multiple RCT simulations, incorporating a broader range of baseline variable variations than previous studies. This methodological enhancement allowed for a more precise estimation of I² thresholds while maintaining the integrity of the correlation between selection bias and effect estimates. Importantly, the new thresholds eliminated the reliance on approximations for B% estimation, which had previously led to inaccuracies. The findings suggest that the revised I² limits are more stringent yet do not invalidate prior assessments made under the original thresholds, thereby reinforcing the utility of the I² statistic in selection bias testing across RCTs.

Limitations

The study presents several limitations that impact its findings. A primary limitation is the lack of a formal sample size calculation, which was necessary to determine the appropriate number of simulated randomized controlled trials (RCTs) to adequately reflect variations in baseline variables encountered in real-world scenarios. The sample size was arbitrarily set at \( N = 50 \), multiplied by three variable domains derived from 332 real-life RCTs, resulting in 165,000 \( I^2 \) data points. Although this sample size proved sufficient for developing \( I^2 \) threshold limits, the absence of a systematic approach to sample size determination may affect the robustness of the results.

Additionally, the study’s reliance on baseline variable ranges from the selected cohort of 332 RCTs, identified through a systematic review process, introduces potential biases. While the trials included span a diverse range of medical specialties and publication years (1985-2023), any relevant baseline data from trials outside this cohort were excluded. This limitation suggests that future research should focus on testing the new \( I^2 \) threshold limits on additional RCTs not covered in this study, allowing for further refinement of these thresholds. Despite these limitations, the findings support the utility of the trial-adjusted, SCT-based \( I^2 \) test as a quantitative method for detecting selection bias in patient allocation within single RCTs, contrasting with existing qualitative approaches that lack the ability to quantitatively assess bias.

شارك: