التحليلات الإحصائية للنتائج الترتيبية في التجارب السريرية العشوائية: مراجعة استكشافية
Statistical analyses of ordinal outcomes in randomised controlled trials: a scoping review

المجلة: Trials، المجلد: 25، العدد: 1
DOI: https://doi.org/10.1186/s13063-024-08072-2
PMID: https://pubmed.ncbi.nlm.nih.gov/38582924
تاريخ النشر: 2024-04-06
المؤلف: Chris J Selman وآخرون
الموضوع الرئيسي: تحليل البيانات الشامل والمراجعات المنهجية

نظرة عامة

تستعرض هذه المراجعة استخدام وتحليل النتائج الترتيبية في التجارب السريرية العشوائية (RCTs) المنشورة بين عامي 2017 و2022 في أربع مجلات طبية رائدة. النتائج الترتيبية مهمة لمعالجة الأسئلة السريرية المعقدة، ويمكن أن يختلف المعامل المستهدف بناءً على سؤال البحث والافتراضات الأساسية. قامت المراجعة بتحليل 309 دراسات، شملت في النهاية 144 دراسة تضمنت نتيجة ترتيبية أولية أو ثانوية.

تكشف النتائج أن نسبة الأرجحية كانت المعامل المستهدف الأكثر تكرارًا، حيث ظهرت في 54% من الدراسات. ومن الجدير بالذكر أن 33% من الدراسات اختارت تقسيم النتائج الترتيبية للتحليل، بينما كانت نموذج النسب التناسبية هو الطريقة الإحصائية السائدة المستخدمة، حيث تم استخدامها في 44% من الدراسات. ومع ذلك، كانت هناك مشكلة مقلقة حيث لم تتحقق 60% من الدراسات بشكل كافٍ أو تبلغ عن قوة الافتراضات الأساسية لطرقها الإحصائية. وتخلص المراجعة إلى أن هناك تباينًا كبيرًا في الأساليب التحليلية للنتائج الترتيبية في التجارب السريرية العشوائية، مما يبرز الحاجة إلى تحسين الإرشادات والشفافية في الإبلاغ عن هذه التحليلات في الأبحاث المستقبلية.

مقدمة

تناقش مقدمة ورقة البحث أهمية النتائج الترتيبية في التجارب السريرية العشوائية (RCTs)، مع التأكيد على فائدتها في التقاط حالات المرضى المختلفة من خلال فئات مرتبة هرميًا. النتائج الترتيبية، التي تُعرف بفئات متبادلة الحصرية ومحددة بشكل غير غامض، تسهل تقييم تحسين أو تدهور حالة المريض بالنسبة لخط الأساس. تشمل الأمثلة الشائعة مقياس رانكين المعدل للإعاقة الناتجة عن السكتة الدماغية، ومقياس غلاسكو للنتائج الممتد للإصابة الدماغية الرضحية، ومقياس تقدم COVID-19 السريري التابع لمنظمة الصحة العالمية، الذي اكتسب شهرة في التجارب السريرية العشوائية التي تقيم شدة COVID-19.

تسلط هذه الفقرة الضوء على التعقيدات المرتبطة بتصميم وتحليل التجارب السريرية العشوائية التي تستخدم النتائج الترتيبية، وخاصة الحاجة إلى تحديد المعاملات المستهدفة، مثل نسبة الأرجحية التناسبية (OR)، وضمان صحة الافتراضات الأساسية. يتم مناقشة نموذج النسب التناسبية (PO) كنهج تحليلي شائع، والذي يفترض أن النسب في الفئات الأعلى متسقة عبر المقياس الترتيبي. ومع ذلك، يمكن أن تؤدي انتهاكات هذا الافتراض إلى نتائج مضللة، تؤثر على كل من الأخطاء من النوع الأول والثاني. كما تم ذكر بدائل مثل نموذج النسب التناسبية الجزئية (PPO)، والانحدار متعدد الحدود، واختبار ويلكوكسون، كل منها مع مزاياها وقيودها الخاصة. إن اختيار طريقة التحليل أمر حاسم، حيث يمكن أن يؤثر على تفسير آثار العلاج وحجم العينة المطلوب للحفاظ على القوة الإحصائية.

طرق

في هذا القسم، يوضح المؤلفون الطرق الإحصائية المستخدمة في تحليل النتائج الترتيبية عبر دراسات مختلفة. كانت الإحصائيات الوصفية، وخاصة التكرارات والنسب المئوية، هي المقاييس الأكثر تكرارًا، بينما ظهر نموذج النسب التناسبية (PO) كأكثر الطرق التحليلية استخدامًا، حيث تم استخدامه في 44% من الدراسات. ومع ذلك، تم إثارة قلق كبير بشأن التحقق من صحة افتراضات PO، حيث لم تتحقق 60% من الدراسات بشكل صريح من صحتها. من بين تلك التي فعلت، أفادت 44 من أصل 64 دراسة بأنها تحقق من الافتراض، مع استخدام اختبارات إحصائية متنوعة، بما في ذلك اختبارات نسبة الاحتمالات واختبار برانت. أشار المؤلفون إلى أنه عندما تم انتهاك افتراض PO، نادرًا ما تم اعتماد طرق بديلة، مثل نموذج النسب التناسبية (PPO) أو الانحدار اللوجستي.

يؤكد المؤلفون على أهمية تقييم افتراض PO، خاصة عند تقدير آثار العلاج عبر مستويات مختلفة من النتيجة الترتيبية. يجادلون بأنه بينما يُستخدم نموذج PO على نطاق واسع بسبب تطبيقه المباشر، فإن صحته ضرورية للتفسير الدقيق لآثار العلاج. كما يبرز القسم المخاطر المحتملة للاعتماد فقط على قيم p دون تحديد معامل مستهدف لآثار العلاج، مما يعقد ترجمة النتائج إلى الممارسة السريرية. بشكل عام، تؤكد النتائج على الحاجة إلى تحسين الصرامة المنهجية في تحليل النتائج الترتيبية لضمان نتائج صحيحة وقابلة للتفسير.

نتائج

يستعرض قسم النتائج تجميع وتحليل البيانات المستخرجة من Covidence، والتي تم إجراؤها باستخدام Stata. تم استخدام الإحصائيات الوصفية لتلخيص البيانات، مع الإبلاغ عن التكرارات والنسب المئوية للمتغيرات الفئوية، والوسائط مع النطاقات الربعية (IQRs) للمتغيرات المستمرة. تم تجميع البيانات النوعية سرديًا.

من البحث الأولي الذي شمل 309 دراسات، تم استبعاد 46 لدورها كدراسات غير عشوائية. خضعت 263 دراسة لمراجعة النص الكامل، مما أدى إلى استبعاد 119 دراسة—110 بسبب عدم وجود نتائج ترتيبية وتسع لدورها كدراسات غير عشوائية. في النهاية، تم اعتبار 144 دراسة مؤهلة لاستخراج البيانات، مع توفير مخطط تدفق يوضح عملية اختيار الدراسة في الشكل 1. الاستبيان المستخدم لاستخراج البيانات متاح في الملف الإضافي 2.

مناقشة

تسلط قسم المناقشة في هذه المراجعة الضوء على الدور الحاسم للنتائج الترتيبية في التجارب السريرية العشوائية (RCTs) وتؤكد على الحاجة إلى الوضوح في تحليلها وإبلاغها. تحدد المراجعة أن المعامل المستهدف الأكثر شيوعًا المستخدم في الدراسات كان نسبة الأرجحية (OR)، وغالبًا ما يتم اشتقاقها من النتائج الترتيبية المقسمة. ومع ذلك، كانت العديد من الدراسات تفتقر إلى مبررات لهذا التقسيم، وبعضها اعتبر النتائج الترتيبية مستمرة، مما يمكن أن يشوه التأثيرات السريرية المهمة. ومن الجدير بالذكر أن ما يقرب من نصف الدراسات لم تحدد بشكل صريح معاملاتها المستهدفة، على الرغم من التوصيات من أجل الوضوح في التقديرات، والتي تعتبر ضرورية لفهم آثار العلاج.

تشير النتائج إلى تباين كبير في كيفية تحليل النتائج الترتيبية والإبلاغ عنها عبر التجارب السريرية العشوائية، خاصة في مجالات الأعصاب والأمراض المعدية. يدعو المؤلفون الباحثين إلى تحديد وتقرير معاملاتهم المستهدفة بعناية لضمان استخدام الطرق الإحصائية المناسبة. هذه الشفافية ضرورية لتفسير النتائج ولتوافق التحليلات الإحصائية مع الطبيعة الفئوية للنتائج الترتيبية. في النهاية، تؤكد المراجعة على أهمية الممارسات المنهجية القوية لتعزيز صحة وقابلية تطبيق النتائج في السياقات السريرية.

القيود

تقدم الدراسة مراجعة شاملة للتجارب السريرية العشوائية (RCTs) التي تستخدم النتائج الترتيبية، مستمدة من أربع مجلات طبية مرموقة لتقييم الممارسات التحليلية الحالية. تشمل نقاط القوة في هذه المراجعة عملية فحص منهجية واستخراج بيانات، مدعومة باختبارات تجريبية واستخراج بيانات مزدوج لضمان موثوقية البيانات. كما أن الالتزام بقائمة مراجعة PRISMA-ScR يبرز الالتزام بمعايير الإبلاغ العالية.

ومع ذلك، فإن المراجعة ليست خالية من القيود. قد يحد التركيز على قاعدة بيانات PubMed وعدد مختار من المجلات عالية التصنيف من إمكانية تعميم النتائج، حيث تم اختيار هذا النهج للحفاظ على إمكانية التكرار وإدارة حجم الدراسات المضمنة. بينما من المتوقع أن تمثل المجلات المختارة أفضل الممارسات في البحث، قد يعكس ذلك أيضًا سيناريو الأفضل بسبب معاييرها الصارمة. بالإضافة إلى ذلك، قد تكون استراتيجية البحث قد أغفلت بعض المصطلحات المتعلقة بالنتائج الترتيبية، على الرغم من الجهود المبذولة لمعالجة ذلك خلال مرحلة التجريب. أخيرًا، يعني استبعاد أوراق البروتوكول من التجارب أنه لم يتم النظر في المعلومات المحتملة ذات الصلة بشأن المنهجيات الإحصائية، بما في ذلك التقييمات المخطط لها لافتراضات النسب التناسبية.

Journal: Trials, Volume: 25, Issue: 1
DOI: https://doi.org/10.1186/s13063-024-08072-2
PMID: https://pubmed.ncbi.nlm.nih.gov/38582924
Publication Date: 2024-04-06
Author(s): Chris J Selman et al.
Primary Topic: Meta-analysis and systematic reviews

Overview

This review examines the use and analysis of ordinal outcomes in randomized controlled trials (RCTs) published between 2017 and 2022 in four leading medical journals. Ordinal outcomes are significant for addressing complex clinical questions, and the target parameter of interest can vary based on the research question and underlying assumptions. The review analyzed 309 studies, ultimately including 144 that featured an ordinal primary or secondary outcome.

The findings reveal that the odds ratio was the most frequently reported target parameter, appearing in 54% of the studies. Notably, 33% of the studies opted to dichotomize the ordinal outcomes for analysis, while the proportional odds model was the predominant statistical method employed, utilized in 44% of the studies. However, a concerning 60% of the studies did not adequately check or report on the robustness of the assumptions underlying their statistical methods. The review concludes that there is considerable variability in the analytical approaches to ordinal outcomes in RCTs, highlighting the need for improved guidance and transparency in reporting these analyses in future research.

Introduction

The introduction of the research paper discusses the significance of ordinal outcomes in randomized controlled trials (RCTs), emphasizing their utility in capturing various patient states through hierarchically ordered categories. Ordinal outcomes, which are defined by mutually exclusive and unambiguously defined categories, facilitate the assessment of patient improvement or deterioration relative to a baseline. Common examples include the modified-Rankin scale for stroke disability, the Glasgow Outcome Scale-Extended for traumatic brain injury, and the WHO COVID-19 Clinical Progression Scale, which has gained prominence in RCTs evaluating COVID-19 severity.

The section highlights the complexities involved in designing and analyzing RCTs that utilize ordinal outcomes, particularly the need to define target parameters, such as the proportional odds ratio (OR), and to ensure the validity of underlying assumptions. The proportional odds (PO) model is discussed as a common analytical approach, which assumes that the odds of being in higher categories are consistent across the ordinal scale. However, violations of this assumption can lead to misleading results, affecting both type I and II errors. Alternatives such as the partial proportional odds (PPO) model, multinomial regression, and the Wilcoxon test are also mentioned, each with their respective advantages and limitations. The choice of analysis method is crucial, as it can impact the interpretation of treatment effects and the required sample size for maintaining statistical power.

Methods

In this section, the authors detail the statistical methods employed in analyzing ordinal outcomes across various studies. Descriptive statistics, primarily frequencies and percentages, were the most frequently reported measures, while the proportional odds (PO) model emerged as the predominant analytical method, utilized in 44% of studies. However, a significant concern was raised regarding the validation of the PO assumption, as 60% of studies did not explicitly check its validity. Among those that did, only 44 out of 64 studies reported verifying the assumption, with various statistical tests employed, including likelihood ratio tests and the Brant test. The authors noted that when the PO assumption was violated, alternative methods, such as the proportional odds model (PPO) or logistic regression, were rarely adopted.

The authors emphasize the importance of assessing the PO assumption, particularly when estimating treatment effects across different levels of the ordinal outcome. They argue that while the PO model is widely used due to its straightforward application, its validity is crucial for accurate interpretation of treatment effects. The section also highlights the potential pitfalls of relying solely on p-values without defining a target parameter for treatment effects, which complicates the translation of results into clinical practice. Overall, the findings underscore the need for improved methodological rigor in the analysis of ordinal outcomes to ensure valid and interpretable results.

Results

The results section outlines the synthesis and analysis of data extracted from Covidence, which was conducted using Stata. Descriptive statistics were employed to summarize the data, reporting frequencies and percentages for categorical variables, and medians with interquartile ranges (IQRs) for continuous variables. Qualitative data were synthesized narratively.

From the initial search of 309 studies, 46 were excluded for not being randomized controlled trials (RCTs). A total of 263 studies underwent full-text review, resulting in the exclusion of 119 studies—110 due to the absence of ordinal outcomes and nine for not being RCTs. Ultimately, 144 studies were deemed eligible for data extraction, with a flow diagram illustrating the study selection process provided in Figure 1. The questionnaire utilized for data extraction is available in Additional file 2.

Discussion

The discussion section of this review highlights the critical role of ordinal outcomes in randomized controlled trials (RCTs) and emphasizes the need for clarity in their analysis and reporting. The review identifies that the most common target parameter used in studies was the odds ratio (OR), often derived from dichotomized ordinal outcomes. However, many studies lacked justification for this dichotomization, and some treated ordinal outcomes as continuous, which can misrepresent clinically meaningful effects. Notably, nearly half of the studies did not explicitly define their target parameters, despite recommendations for clarity in estimands, which are essential for understanding treatment effects.

The findings suggest significant variability in how ordinal outcomes are analyzed and reported across RCTs, particularly in neurology and infectious diseases. The authors advocate for researchers to carefully define and report their target parameters to ensure appropriate statistical methods are employed. This transparency is crucial for the interpretation of results and for aligning statistical analyses with the categorical nature of ordinal outcomes. Ultimately, the review underscores the importance of robust methodological practices to enhance the validity and applicability of findings in clinical settings.

Limitations

The study presents a comprehensive review of randomized controlled trials (RCTs) utilizing ordinal outcomes, drawing from four prestigious medical journals to assess current analytical practices. The strengths of this review include a systematic screening and data extraction process, bolstered by pilot tests and double data extraction to ensure data reliability. Adherence to the PRISMA-ScR checklist further underscores the commitment to high reporting standards.

However, the review is not without limitations. The focus on the PubMed database and a select number of high-ranking journals may restrict the generalizability of the findings, as this approach was chosen to maintain reproducibility and manage the volume of studies included. While the selected journals are expected to represent best practices in research, this may also reflect a best-case scenario due to their rigorous standards. Additionally, the search strategy may have overlooked certain terminologies related to ordinal outcomes, despite efforts to address this during the piloting phase. Lastly, the exclusion of protocol papers from the trials means that potentially relevant information regarding statistical methodologies, including planned assessments of proportional odds assumptions, was not considered.