DOI: https://doi.org/10.1017/rsm.2025.10055
PMID: https://pubmed.ncbi.nlm.nih.gov/41992906
تاريخ النشر: 2026-01-09
المؤلف: Péter Mátrai وآخرون
الموضوع الرئيسي: أساليب متنوعة في دراسات الرعاية الصحية والتعليم
نظرة عامة
تناقش هذه القسم المنهجية والآثار المترتبة على تحليل الميتا ذو التأثيرات العشوائية، مع التركيز بشكل خاص على فترة التنبؤ (PI) كأداة لتقييم التباين بين تأثيرات الدراسات. يؤكد المؤلفون أنه بينما تهدف PI إلى احتواء التأثير الحقيقي لدراسة جديدة باحتمالية محددة، يمكن أن يختلف أداؤها بشكل كبير بناءً على عدد الدراسات المضمنة في التحليل. من خلال محاكاة واسعة، تقيم الدراسة احتمالات التغطية لطرق PI المختلفة، كاشفة أن توزيع هذه الاحتمالات يمكن أن يكون غير متوازن، خاصة عندما يكون عدد الدراسات صغيرًا.
تشير النتائج الرئيسية إلى أن الطريقة الشائعة المستخدمة من قبل هيغينز-طومسون-سبايجلهاتر قد تفشل في الحفاظ على التغطية الاسمية في العينات الصغيرة، بينما تظهر طريقة البوتستراب المعلمية أداءً أفضل في ظل ظروف مماثلة. يحذر المؤلفون من أن الاعتماد فقط على متوسط التغطية يمكن أن يؤدي إلى تفسيرات خاطئة لموثوقية PI، خاصة في الحالات التي يكون فيها عدد الدراسات محدودًا. يدعون إلى مزيد من البحث لتحسين طرق بناء الفترات لضمان معايير تغطية أكثر صرامة تتجاوز متوسط التغطية.
مقدمة
تناقش مقدمة هذه الورقة البحثية منهجية تحليل الميتا (MA)، مع التركيز بشكل خاص على نموذج التأثيرات العشوائية (RE)، الذي يأخذ في الاعتبار التباين بين مجموعات الدراسات من خلال افتراض توزيع مشترك أساسي للتأثيرات الحقيقية. بينما تؤكد تقارير MA التقليدية على تقدير النقاط والفترات للمعامل المتوسط للتأثيرات الحقيقية، يمكن أن يؤدي هذا النهج إلى إخفاء التباينات الكبيرة في أحجام التأثير عبر الدراسات المختلفة، مما قد يؤدي إلى استنتاجات مضللة. لمعالجة هذه القضية، اقترح هيغينز وآخرون (2009) استخدام فترات التنبؤ (PIs) كطريقة أكثر معلوماتية لتلخيص التباين، والتي تم الدعوة منذ ذلك الحين إلى تضمينها في تقارير MA.
تسلط الورقة الضوء على أن PI بنسبة 95% تهدف إلى تغطية توزيع التأثيرات الحقيقية في المتوسط عبر عينات متكررة، لكن التفسيرات الشائعة قد تسيء تمثيل احتمالية تغطيتها. يهدف المؤلفون إلى توسيع فهم PIs من خلال التحقيق في توزيع احتمالات التغطية بدلاً من التركيز فقط على متوسط التغطية. يقدمون دراسة محاكاة شاملة تقيم أداء طرق PI المختلفة، بما في ذلك PI هيغينز-طومسون-سبايجلهاتر (HTS) واثنين من المقدرات الأحدث. كما تفحص الدراسة حساسية هذه المقدرات عندما ينحرف توزيع التأثيرات العشوائية عن التوزيع الطبيعي، مما يمهد الطريق لتحليل مفصل لطرق PI وآثارها في تحليل الميتا.
طرق
في هذا القسم، يحدد المؤلفون منهجيتهم لبناء فترات التنبؤ (PIs) من خلال طرق المحاكاة، مستندين إلى نهج محاكاة حجم التأثير الذي أسسه باكبيرغنولي وآخرون. يركزون على الفرق المتوسط (MD) كقياس التأثير الأساسي بين المجموعات التجريبية (E) ومجموعات التحكم (C)، مشيرين إلى انتشاره في الأبحاث الطبية مع التأكيد على أن نتائجهم يمكن تعميمها على أحجام تأثير أخرى مثل الفرق المتوسط المعياري، ونسبة المخاطر، ونسبة الأرجحية، ونسبة المخاطر. تشير الدراسات السابقة إلى أن أداء PIs يتأثر بشكل كبير بعدد الدراسات المضمنة ومستوى التباين الموجود.
قيم المؤلفون عدة طرق PI، بما في ذلك طريقة البوتستراب المعلمية، وطريقة التجميع، وستة تباينات من PIs من نوع HTS، التي تختلف بناءً على المقدرات المستخدمة لـ $\tau^2$، التباين ($Var(\mu)$)، والتوزيع المطبق. أجروا محاكاة لـ 5000 تحقيقات تحليل ميتا لمعظم الطرق، بينما تم تقييم طريقة البوتستراب المعلمية مع 1000 تحقيق بسبب كثافتها الحاسوبية، مما يتطلب عددًا أكبر من عينات البوتستراب. تم تنفيذ المحاكاة باستخدام R v.4.1.3، مستفيدين من موارد الحوسبة عالية الأداء لطريقة البوتستراب المعلمية وجهاز كمبيوتر شخصي قياسي للطرق الأخرى.
نتائج
تم توثيق نتائج المحاكاة التي أجريت في هذه الدراسة بشكل شامل في ملف HTML متاح في المواد عبر الإنترنت. يتضمن النص الرئيسي والمواد التكميلية أشكالًا تمثل سيناريوهين متميزين: سيناريو تباين منخفض (مع المعلمات $N = 100$، $\tau^2 = 0.2$، $I^2 = 33\%$، $v = 0.5$) وسيناريو تباين مرتفع (مع المعلمات $N = 100$، $\tau^2 = 1$، $I^2 = 71\%$، $v = 2.5$). توضح الأقسام 5.1 إلى 5.4 النتائج المتعلقة بتوزيع التغطية، ومتوسط التغطية والوسيط، ومتوسط طول الفترة الملاحظة بالنسبة للطول النظري.
بالإضافة إلى ذلك، توفر المواد التكميلية مزيدًا من الرؤى حول قياسين آخرين للأداء: الفرق المطلق المتوسط من تغطية 95% والخطأ المطلق المتوسط المعاير، كما هو موضح في الأشكال الملحقة 3 و4. تفترض التحليل أن التأثيرات العشوائية موزعة بشكل طبيعي عبر الأقسام 5.1 إلى 5.4، بينما تتناول القسم 5.5 آثار عدم الطبيعية. من الجدير بالذكر أن طريقة HTS-REML (t K-2) أسفرت عن نتائج متطابقة تقريبًا مع تلك الخاصة بطريقة HTS-DL (t K-2)، مما أدى إلى اتخاذ القرار بعرض نتائج الأخيرة فقط. تم تضمين الرسوم البيانية لطريقة HTS-HKSJ (t K-1) في الأشكال الملحقة 1 و2 من المواد التكميلية.
مناقشة
تتناول قسم المناقشة في الورقة البحثية نموذج تحليل الميتا ذو التأثيرات العشوائية (REMA)، الذي يأخذ في الاعتبار التباين عبر الدراسات بسبب اختلاف السكان، والسياقات الثقافية، وتصاميم الدراسات. يفترض النموذج أن التأثيرات الحقيقية من دراسات مختلفة مستقلة وموزعة بشكل متطابق، وعادة ما تتبع توزيعًا طبيعيًا. الهدف الأساسي هو تقدير القيمة المتوقعة ($\mu$) والتباين ($\tau^2$) لهذه التأثيرات الحقيقية، مع الاعتراف بأن النتائج الملاحظة تتأثر بأخطاء القياس العشوائية. تنتقد الورقة نموذج التأثير الشائع، مشيرة إلى أنه نادرًا ما يكون قابلاً للتطبيق في العلوم الاجتماعية أو الطبية بسبب التباين المتأصل بين الدراسات.
تتم مناقشة تقدير المعلمات في REMA، مع التأكيد على استخدام المتوسطات الموزونة للحصول على تقديرات غير متحيزة لـ $\mu$. يتم تعريف الأوزان تقليديًا على أنها $w_k = \frac{1}{\sigma^2_k + \tau^2}$، حيث تمثل $\sigma^2_k$ التباينات داخل الدراسة. يتم مراجعة طرق مختلفة لتقدير $\tau^2، بما في ذلك مقدر ديرسيمونيان وليرد ومقدر الاحتمالية القصوى المقيدة. يقدم القسم أيضًا مفهوم فترات التنبؤ (PIs)، التي توفر نطاقًا من التأثيرات الحقيقية المحتملة للدراسات الجديدة، خاصة عندما يوجد تباين كبير. توضح الورقة طرقًا مختلفة لبناء PIs، بما في ذلك طريقة هيغينز-طومسون-سبايجلهاتر وطريقة البوتستراب المعلمية، مشيرة إلى أدائها المتفاوت من حيث احتمالات التغطية وأطوال الفترات تحت سيناريوهات مختلفة من التباين وأحجام العينات.
بشكل عام، تؤكد النتائج على أهمية تمثيل عدم اليقين بدقة في تحليلات الميتا من خلال PIs، خاصة في المجالات التي تختلف فيها ظروف الدراسة بشكل كبير، وتبرز الحاجة إلى اعتبار دقيق للطرق المستخدمة لتقدير هذه الفترات.
DOI: https://doi.org/10.1017/rsm.2025.10055
PMID: https://pubmed.ncbi.nlm.nih.gov/41992906
Publication Date: 2026-01-09
Author(s): Péter Mátrai et al.
Primary Topic: Diverse Approaches in Healthcare and Education Studies
Overview
The section discusses the methodology and implications of random-effects meta-analysis, particularly focusing on the prediction interval (PI) as a tool for assessing heterogeneity among study effects. The authors emphasize that while the PI aims to contain the true effect of a new study with a specified probability, its performance can vary significantly based on the number of studies included in the analysis. Through extensive simulations, the study evaluates the coverage probabilities of various frequentist PI methods, revealing that the distribution of these probabilities can be asymmetric, especially when the number of studies is small.
Key findings indicate that the commonly used Higgins-Thompson-Spiegelhalter method may fail to maintain nominal coverage in small samples, whereas the parametric bootstrap method shows better performance under similar conditions. The authors caution that relying solely on mean coverage can lead to misinterpretations of the PI’s reliability, particularly in cases with limited study numbers. They advocate for further research to refine interval construction methods to ensure more stringent coverage criteria beyond the average mean coverage.
Introduction
The introduction of this research paper discusses the methodology of meta-analysis (MA), particularly focusing on the random-effects (RE) model, which accounts for heterogeneity among study populations by assuming a common underlying distribution of true effects. While traditional MA reports emphasize point and interval estimation of the mean parameter of true effects, this approach can obscure significant variations in effect sizes across different studies, potentially leading to misleading conclusions. To address this issue, Higgins et al. (2009) proposed the use of prediction intervals (PIs) as a more informative way to summarize heterogeneity, which has since been advocated for inclusion in MA reports.
The paper highlights that a 95% PI is intended to cover the true effects distribution on average across repeated samples, but common interpretations may misrepresent its coverage probability. The authors aim to extend the understanding of PIs by investigating the distribution of coverage probabilities rather than solely focusing on the mean coverage. They present a comprehensive simulation study assessing the performance of various frequentist PI methods, including the Higgins-Thompson-Spiegelhalter (HTS) PI and two newer estimators. The study also examines the sensitivity of these estimators when the random-effects distribution deviates from normality, setting the stage for a detailed analysis of PI methods and their implications in meta-analysis.
Methods
In this section, the authors outline their methodology for constructing prediction intervals (PIs) through simulation methods, drawing on the effect size simulation approach established by Bakbergenuly et al. They focus on the mean difference (MD) as the primary effect measure between experimental (E) and control (C) groups, noting its prevalence in medical research while asserting that their findings can be generalized to other effect sizes such as standardized mean difference, hazard ratio, odds ratio, and risk ratio. Previous studies indicate that the performance of PIs is significantly influenced by the number of studies included and the level of heterogeneity present.
The authors evaluated several PI methods, including the parametric bootstrap method, the ensemble method, and six variations of HTS-type PIs, which differ based on the estimators used for $\tau^2$, variance ($Var(\mu)$), and the distribution applied. They conducted simulations of 5000 meta-analysis realizations for most methods, while the parametric bootstrap method was assessed with 1000 realizations due to its computational intensity, requiring a higher number of bootstrap samples. The simulations were executed using R v.4.1.3, leveraging high-performance computing resources for the parametric bootstrap method and a standard personal computer for the other methods.
Results
The results of the simulations conducted in this study are comprehensively documented in an accessible HTML file available in the Online Material. The main text and Supplementary Material include figures representing two distinct scenarios: a low heterogeneity scenario (with parameters $N = 100$, $\tau^2 = 0.2$, $I^2 = 33\%$, $v = 0.5$) and a high heterogeneity scenario (with parameters $N = 100$, $\tau^2 = 1$, $I^2 = 71\%$, $v = 2.5$). Sections 5.1 to 5.4 detail the results concerning coverage distribution, mean and median coverage, and the mean observed interval length in relation to the theoretical length.
Additionally, the Supplementary Material provides further insights into two other performance measures: the mean absolute difference from 95% coverage and the normalized mean absolute error, as illustrated in Appendix Figures 3 and 4. The analysis assumes normally distributed random effects throughout Sections 5.1 to 5.4, while Section 5.5 addresses the implications of non-normality. Notably, the HTS-REML (t K-2) method yielded results nearly identical to those of the HTS-DL (t K-2) method, leading to the decision to present only the latter’s findings. Histograms for the HTS-HKSJ (t K-1) method are included in Appendix Figures 1 and 2 of the Supplementary Material.
Discussion
The discussion section of the research paper elaborates on the random-effects meta-analysis (REMA) model, which accounts for variability across studies due to differing populations, cultural contexts, and study designs. The model assumes that the true effects from various studies are independent and identically distributed, typically following a normal distribution. The primary objective is to estimate the expected value ($\mu$) and variance ($\tau^2$) of these true effects, while acknowledging that observed outcomes are influenced by random measurement errors. The paper critiques the common effect model, highlighting that it is rarely applicable in social or medical sciences due to inherent heterogeneity among studies.
Parameter estimation in REMA is discussed, emphasizing the use of weighted means to derive unbiased estimates of $\mu$. The weights are conventionally defined as $w_k = \frac{1}{\sigma^2_k + \tau^2}$, where $\sigma^2_k$ represents within-study variances. Various methods for estimating $\tau^2 are reviewed, including the DerSimonian and Laird estimator and the restricted maximum likelihood estimator. The section also introduces the concept of prediction intervals (PIs), which provide a range of plausible true effects for new studies, particularly when significant heterogeneity exists. The paper outlines different methods for constructing PIs, including the Higgins-Thompson-Spiegelhalter method and the parametric bootstrap method, noting their varying performance in terms of coverage probabilities and interval lengths under different scenarios of heterogeneity and sample sizes.
Overall, the findings underscore the importance of accurately representing uncertainty in meta-analyses through PIs, especially in fields where study conditions vary widely, and they highlight the need for careful consideration of the methods used to estimate these intervals.
