النواة GRADE 3: تقييم يقين الأدلة – تقييم التناقض
Core GRADE 3: rating certainty of evidence—assessing inconsistency

شارك:
المجلة: BMJ، المجلد: 389
DOI: https://doi.org/10.1136/bmj-2024-081905
PMID: https://pubmed.ncbi.nlm.nih.gov/40328467
تاريخ النشر: 2025-05-06
المؤلف: Gordon Guyatt وآخرون
الموضوع الرئيسي: تحليل البيانات الشامل والمراجعات المنهجية

نظرة عامة

يتناول هذا القسم من ورقة البحث مفهوم عدم الاتساق ضمن إطار عمل Core GRADE (تقييم التوصيات، التطوير والتقييم)، والذي هو جزء من سلسلة تهدف إلى تعزيز تقييم الأدلة والتوصيات للتدخلات السريرية. يُعرف عدم الاتساق بأنه تباين غير مفسر في النتائج عبر الدراسات، مما يمكن أن يؤدي إلى استنتاجات سريرية مختلفة بناءً على النتائج التي تعتبر دقيقة. تؤكد الورقة على أهمية فهم مقياس التأثير ذي الصلة بالنتائج التي يتم تقييمها؛ بالنسبة للنتائج الثنائية، يركز Core GRADE على التأثيرات النسبية مثل نسب المخاطر أو نسب الأرجحية، بينما بالنسبة للنتائج المستمرة، تُعطى الأولوية للتأثيرات المطلقة.

للتعامل بفعالية مع عدم الاتساق، يتم تشجيع مستخدمي Core GRADE على فحص الرسوم البيانية للغابات بصريًا لتقييم حجم الاختلافات في تقديرات النقاط، وتداخل فترات الثقة (CIs)، وعلاقة هذه التقديرات بالعتبات المختارة للتأثير الصفري أو الفرق المهم الحد الأدنى (MID). يتم تسليط الضوء على إحصائية I² كمقياس لعدم الاتساق، على الرغم من أنه يجب تفسيرها بحذر. توضح الورقة معايير لتقييم مصداقية تحليلات المجموعات الفرعية، بما في ذلك قيمة P من اختبارات التفاعل والتوافق مع الفرضيات السابقة. إذا تم اعتبار تأثير المجموعة الفرعية موثوقًا، سيتم تقديم ملخصات أدلة منفصلة، وسيتم تقييم يقين الأدلة وفقًا لذلك. على العكس، إذا لم يتم تحديد تأثير مجموعة فرعية موثوق، سيتم تقييم عدم الاتساق عبر جميع الدراسات المؤهلة.

النتائج

في قسم النتائج، يناقش المؤلفون آثار تقييمات مصداقية المجموعات الفرعية في نتائج البحث. إذا قرر مستخدمو Core GRADE أن مصداقية تأثير المجموعة الفرعية منخفضة أو منخفضة جدًا، فسوف يبلغون فقط عن النتائج لملخص الدراسات الكلي، مما يشير إلى عدم الاتساق عبر جميع السكان. على العكس، إذا تم تقييم المصداقية على أنها معتدلة أو عالية، سيتم إنشاء أسئلة PICO منفصلة ونتائج لكل مجموعة فرعية، جنبًا إلى جنب مع تقييمات متميزة لليقين عبر جميع المجالات الخمسة للتقييم. في الحالات التي يكون فيها تقييم المصداقية قريبًا من العتبة بين المنخفض والمعتدل، قد يختار المؤلفون تقديم كل من النتائج العامة ونتائج المجموعة الفرعية في جدول ملخص النتائج أو تضمين واحدة بشكل انتقائي أثناء تلخيص الأخرى في النص، مع التأكيد على الاعتراف بالطبيعة الدقيقة لتقييم المصداقية.

مثال توضيحي مقدم هو تحليل حاصرات بيتا لتقليل الوفيات في مرضى فشل القلب، حيث تعتبر مصداقية التأثير معتدلة. باستخدام نهج الانحدار الميتا المستمر، وجد المؤلفون أن التخفيضات الأكبر في معدل ضربات القلب تتوافق مع انخفاض أكبر في الوفيات. تشير هذه المصداقية المعتدلة إلى أن الأطباء والمرضى قد يشاركون في اتخاذ قرارات مشتركة بشأن استخدام حاصرات بيتا، داعين إلى جرعات تخفض معدل ضربات القلب بشكل فعال وآمن.

المناقشة

في قسم المناقشة من ورقة البحث، يتناول المؤلفون التحديات المتعلقة بتقييم عدم الاتساق في النتائج المستمرة، والتي يتم تقييمها عادةً كاختلافات متوسطة. يؤكدون أن عدم الاتساق في هذه الاختلافات المتوسطة يمكن أن يقوض يقين الأدلة، تمامًا كما يؤثر على النتائج الثنائية. يتم تقديم نهج Core GRADE كطريقة منهجية لتوقع وإدارة عدم الاتساقات المحتملة من خلال النظر في عوامل مختلفة، بما في ذلك الخطأ العشوائي والاختلافات في العناصر السكانية، التدخل، المقارنة، والنتيجة (PICO). يجادل المؤلفون بأنه بينما قد يبدو أن التباين في عناصر PICO يهدد في البداية يقين الأدلة، إلا أنه يمكن أن يعزز أيضًا قابلية تطبيق النتائج عبر سياقات سريرية متنوعة إذا ظلت التأثيرات متسقة.

يحدد المؤلفون ثلاثة سيناريوهات لمعالجة الاختلافات المحتملة في تأثيرات التدخل بين المجموعات الفرعية، خاصة فيما يتعلق بخصائص المرضى أو أنواع العلاج. يوصون ببناء أطر PICO واسعة لتعظيم القابلية للتعميم مع الاستعداد لعدم الاتساقات المحتملة. يجب صياغة فرضيات سابقة لشرح أي عدم اتساق ملحوظ، ويفضل أن تتضمن اتجاهًا محددًا بناءً على الأدلة السابقة. تناقش الورقة أيضًا معايير الحكم على مصداقية تأثيرات المجموعات الفرعية، مع التأكيد على أهمية التمييز بين المقارنات داخل الدراسة وبين الدراسات، وأهمية اختبارات التفاعل، والحاجة إلى عدد محدود من الفرضيات الاتجاهية. في النهاية، يدعو المؤلفون إلى تقييم دقيق لعدم الاتساق، باستخدام معايير بصرية من الرسوم البيانية للغابات وقياسات إحصائية، لضمان استنتاجات قوية في المراجعات المنهجية.

Journal: BMJ, Volume: 389
DOI: https://doi.org/10.1136/bmj-2024-081905
PMID: https://pubmed.ncbi.nlm.nih.gov/40328467
Publication Date: 2025-05-06
Author(s): Gordon Guyatt et al.
Primary Topic: Meta-analysis and systematic reviews

Overview

This section of the research paper discusses the concept of inconsistency within the Core GRADE (Grading of Recommendations Assessment, Development and Evaluation) framework, which is part of a series aimed at enhancing the evaluation of evidence and recommendations for clinical interventions. Inconsistency is defined as unexplained variability in results across studies, which can lead to differing clinical inferences based on which results are considered accurate. The paper emphasizes the importance of understanding the measure of effect relevant to the outcomes being assessed; for binary outcomes, Core GRADE focuses on relative effects such as risk ratios or odds ratios, while for continuous outcomes, absolute effects are prioritized.

To effectively address inconsistency, Core GRADE users are encouraged to visually inspect forest plots to evaluate the magnitude of differences in point estimates, the overlap of confidence intervals (CIs), and the relationship of these estimates to the chosen thresholds for null effect or minimal important difference (MID). The I² statistic is highlighted as a measure of inconsistency, though it should be interpreted with caution. The paper outlines criteria for assessing the credibility of subgroup analyses, including the P value from interaction tests and alignment with a priori hypotheses. If a subgroup effect is deemed credible, separate evidence summaries will be provided, and the certainty of evidence will be rated accordingly. Conversely, if no credible subgroup effect is identified, inconsistency will be assessed across all eligible studies.

Results

In the results section, the authors discuss the implications of subgroup credibility assessments in research findings. If Core GRADE users determine that the credibility of a subgroup effect is low or very low, they will only report results for the overall summary of studies, indicating inconsistency across the entire population. Conversely, if the credibility is assessed as moderate or high, separate PICO questions and results for each subgroup will be created, along with distinct ratings of certainty across all five domains of evaluation. In cases where the credibility assessment is close to the threshold between low and moderate, authors may choose to present both overall and subgroup results in the summary of findings table or selectively include one while summarizing the other in the text, ensuring to acknowledge the nuanced nature of the credibility assessment.

An illustrative example provided is the analysis of β blockers for reducing mortality in heart failure patients, where the credibility of the effect is deemed moderate. Utilizing a continuous meta-regression approach, the authors found that greater reductions in heart rate corresponded to a more significant decrease in mortality. This moderate credibility suggests that clinicians and patients might engage in shared decision-making regarding the use of β blockers, advocating for doses that effectively and safely lower heart rates.

Discussion

In the discussion section of the research paper, the authors address the challenges of assessing inconsistency in continuous outcomes, which are typically evaluated as mean differences. They emphasize that inconsistency in these mean differences can undermine the certainty of evidence, similar to how it affects binary outcomes. The Core GRADE approach is presented as a systematic method for anticipating and managing potential inconsistencies by considering various factors, including random error and differences in population, intervention, comparison, and outcome (PICO) elements. The authors argue that while variability in PICO elements can initially seem to compromise evidence certainty, it can also enhance the applicability of findings across diverse clinical contexts if effects remain consistent.

The authors outline three scenarios for addressing potential subgroup differences in intervention effects, particularly regarding patient demographics or treatment types. They recommend constructing broad PICO frameworks to maximize generalizability while preparing for possible inconsistencies. A priori hypotheses should be formulated to explain any observed inconsistencies, ideally including a specified direction based on prior evidence. The paper also discusses criteria for judging the credibility of subgroup effects, emphasizing the importance of distinguishing between within-study and between-study comparisons, the significance of interaction tests, and the need for a limited number of directional hypotheses. Ultimately, the authors advocate for a careful evaluation of inconsistency, using visual criteria from forest plots and statistical measures, to ensure robust conclusions in systematic reviews.

شارك: