DOI: https://doi.org/10.3348/kjr.2023.0965
PMID: https://pubmed.ncbi.nlm.nih.gov/38528689
تاريخ النشر: 2024-01-01
المؤلف: Kyunghwa Han وآخرون
الموضوع الرئيسي: الموثوقية والاتفاق في القياس
نظرة عامة
تؤكد هذه الفقرة على أهمية اتفاقية القراء كمعيار ناتج حاسم في البحث الإشعاعي، خاصة في سياق تفسير الصور الطبية، الذي يمكن أن يتأثر بالتقييمات الذاتية لمختلف القراء. تعتبر مقاييس موثوقة لاتفاقية القراء ضرورية لضمان صحة نتائج الأبحاث السريرية. تشمل الطرق الإحصائية الشائعة المستخدمة لتقييم هذا الاتفاق إحصائية كابا، ومعامل الارتباط داخل الفئة (ICC)، ومعامل الارتباط المتناغم (CCC).
بينما يكون تحليل اتفاقية القراء بسيطًا مع قارئين، غالبًا ما تتطلب الدراسات ثلاثة قراء أو أكثر لتحسين قابلية تعميم النتائج عبر مختلف الإعدادات السريرية. قد تشكل هذه الضرورة تحديات للباحثين الذين ليسوا على دراية كافية بالطرق الإحصائية القابلة للتطبيق على مجموعات القراء الأكبر، مما يبرز الحاجة إلى توجيه واضح في مثل هذه التحليلات.
نقاش
في قسم النقاش، يتناول المؤلفون تعقيدات قياس اتفاقية القراء عبر مقاييس مختلفة—ثنائية، ترتيبية، ومستمرّة—مبرزين قيود الطرق الإحصائية المستخدمة بشكل شائع. بالنسبة للمقاييس الثنائية أو الاسمية، فإن كابا كوهين مناسب لقارئين، بينما يُوصى بكابا فليس وكابا كونغر لثلاثة قراء أو أكثر. يوفر كابا لايت متوسط قيم كابا كوهين عبر جميع تركيبات القارئين. ومع ذلك، للتخفيف من مشكلات مثل مفارقة الانتشار، يُقترح استخدام مقاييس بديلة مثل معامل اتفاق غويت (AC1)، وBP لبرينان-بريدجر، وα لكريبندورف.
بالنسبة للمقاييس الترتيبية، يشير المؤلفون إلى نقص في الطرق الإحصائية المعتمدة، مما يؤدي إلى تطبيق غير كافٍ لكابا فليس وكابا كوهين الموزون، الذي يقتصر على مقارنات القارئين. يقترحون كابا موزون عام وAC2 لغويت كتحسينات محتملة، على الرغم من أن الخصائص الإحصائية لهذه الطرق لا تزال غير مستكشفة بشكل كافٍ. في تقييم المقاييس المستمرة، يتم التأكيد على إحصائيات الموثوقية مثل معامل الارتباط داخل الفئة (ICC) ومعامل الارتباط المتناغم (CCC)، جنبًا إلى جنب مع طرق رسومية مثل مخطط بلاند-ألتمن المعدل. يدعو المؤلفون إلى تقديم تقارير أوضح عن الطرق الإحصائية والنتائج، بما في ذلك نسب الاتفاق الملاحظة، لتعزيز الشفافية والجودة في دراسات اتفاقية القراء.
DOI: https://doi.org/10.3348/kjr.2023.0965
PMID: https://pubmed.ncbi.nlm.nih.gov/38528689
Publication Date: 2024-01-01
Author(s): Kyunghwa Han et al.
Primary Topic: Reliability and Agreement in Measurement
Overview
The section emphasizes the importance of inter-reader agreement as a critical outcome parameter in radiological research, particularly in the context of medical image interpretation, which can be influenced by the subjective assessments of different readers. Reliable measures of inter-reader agreement are essential for ensuring the validity of clinical research findings. Common statistical methods used to assess this agreement include the kappa statistic, intraclass correlation coefficient (ICC), and concordance correlation coefficient (CCC).
While analyzing inter-reader agreement is straightforward with two readers, studies often require three or more readers to improve the generalizability of the results across various clinical settings. This necessity may pose challenges for researchers who are less familiar with the statistical methods applicable to larger reader groups, highlighting the need for clear guidance in such analyses.
Discussion
In the discussion section, the authors address the complexities of measuring inter-reader agreement across different scales—binary, ordinal, and continuous—highlighting the limitations of commonly used statistical methods. For binary or nominal scales, Cohen’s kappa is suitable for two readers, while Fleiss’s kappa and Conger’s kappa are recommended for three or more readers. Light’s kappa offers an average of Cohen’s kappa values across all two-reader combinations. However, to mitigate issues like the prevalence paradox, alternative measures such as Gwet’s agreement coefficient (AC1), Brennan-Prediger’s BP, and Krippendorff’s α are suggested.
For ordinal scales, the authors note a lack of established statistical methods, leading to inadequate application of Fleiss’s kappa and Cohen’s weighted kappa, which is limited to two-reader comparisons. They propose the generalized weighted kappa and Gwet’s AC2 as potential improvements, although the statistical properties of these methods remain underexplored. In assessing continuous scales, reliability statistics like the Intraclass Correlation Coefficient (ICC) and Concordance Correlation Coefficient (CCC) are emphasized, along with graphical methods such as the modified Bland-Altman plot. The authors advocate for clearer reporting of statistical methods and results, including observed agreement proportions, to enhance transparency and quality in inter-reader agreement studies.
