متى يجب التوقف عن المراجعة: التحقق من معايير التوقف في ASReview
When to stop reviewing: validation of stop criteria in ASReview

شارك:
المجلة: BMC Medical Research Methodology، المجلد: 26، العدد: 1
DOI: https://doi.org/10.1186/s12874-026-02866-5
PMID: https://pubmed.ncbi.nlm.nih.gov/42106619
تاريخ النشر: 2026-05-09
المؤلف: C. Kempny وآخرون
الموضوع الرئيسي: تحليل البيانات الشامل والمراجعات المنهجية

نظرة عامة

تستقصي هذه الدراسة فعالية ثلاثة معايير للتوقف لفرز العناوين والملخصات (T&A) في ASReview، وهي أداة تعلم آلي مصممة لتبسيط المراجعات المنهجية. تشمل المعايير التي تم اختبارها: (1) التوقف بعد تحديد عدد محسوب من الدراسات ذات الصلة من عينة أولية، (2) التوقف بعد عدد ثابت من الدراسات غير ذات الصلة المتتالية، و(3) التوقف بعد فرز نسبة محددة مسبقًا من مجموعة البيانات. شملت المحاكاة 35,000 فرز تلقائي عبر خمس مجموعات بيانات من مستودع SYNERGY. أظهرت النتائج تباينًا كبيرًا في النسبة المئوية لمجموعة البيانات التي كان يجب فرزها لتحديد جميع الدراسات ذات الصلة، حيث تراوحت من 2.9% إلى 76.9%. لم يضمن أي من معايير التوقف بشكل متسق تحديد جميع الدراسات ذات الصلة، حيث أثبت المعيار الأول موثوقيته في 2% فقط من المحاكاة.

تشير النتائج إلى أنه بينما يمكن لـ ASReview تسريع عملية الفرز من خلال إعطاء الأولوية للدراسات المحتملة ذات الصلة، فإن عدم وجود معيار توقف موثوق يثير المخاوف بشأن خطر فقدان الدراسات ذات الصلة. تؤكد الدراسة على أن فعالية معايير التوقف تختلف بشكل كبير بناءً على خصائص مجموعة البيانات، لا سيما في مجموعات البيانات الأصغر حيث يكون احتمال فقدان الدراسات أعلى. وبالتالي، يوصي المؤلفون بتطبيق حذر لمعايير التوقف، داعين إلى دمجها مع تدابير ضمان الجودة، مثل استشارات الخبراء أو فرز قوائم المراجع. يجب أن تركز الأبحاث المستقبلية على تطوير قواعد توقف أكثر قوة واستكشاف نماذج بديلة لتعزيز كفاءة فرز T&A في المراجعات المنهجية.

مقدمة

تؤكد مقدمة ورقة البحث على الدور الحاسم للمراجعات المنهجية في الأبحاث المستندة إلى الأدلة عبر مختلف التخصصات، لا سيما في أبحاث الصحة. تسلط الضوء على أن المراجعات المنهجية هي أساسيات للإرشادات السريرية وقرارات السياسة، مع الإشارة أيضًا إلى الطبيعة المرهقة لعملية الفرز، لا سيما فرز العناوين والملخصات (T&A). يزيد الحجم المتزايد للدراسات المنشورة من تعقيد هذه العملية، مما يؤدي إلى تأخيرات محتملة في النشر ونتائج قديمة. يدعو المؤلفون إلى استراتيجيات بحث أكثر دقة، يفضل أن يتم تطويرها مع أمناء مكتبات متخصصين، ودمج أدوات الذكاء الاصطناعي (AI) وتعلم الآلة (ML)، مثل ASReview، لتعزيز الكفاءة وتقليل عبء العمل المرتبط بالفرز.

تناقش الورقة وظيفة ASReview، التي تستخدم خوارزمية تعلم نشط لإعطاء الأولوية للدراسات بناءً على قرارات المستخدم، مما يحسن من صلة الدراسات المقدمة للمراجعة. يقترح المؤلفون معايير مختلفة لإيقاف عملية الفرز مبكرًا، بما في ذلك تقدير الدراسات ذات الصلة من عينة، وعدد الدراسات غير ذات الصلة المتتالية، ونسبة محددة مسبقًا من مجموعة البيانات التي تم مراجعتها. يؤكدون على ضرورة تقييم هذه المعايير بشكل نقدي لتجنب خطر استبعاد الدراسات ذات الصلة. تهدف الدراسة إلى سد الفجوة البحثية الحالية من خلال تحديد الحد الأقصى النظري لتوفير الجهد في فرز T&A وتقييم معايير التوقف العملية التي تقلل من خطر إغفال الدراسات ذات الصلة.

الطرق

في هذه الدراسة، قام المؤلفون بإجراء 35,000 فرز تلقائي للعناوين والملخصات (T&A) عبر خمس مجموعات بيانات، باستخدام ASReview (الإصدار 1.6.1) ونص برمجي مخصص بلغة بايثون. بدأ كل فرز بمجموعة فرعية عشوائية من الدراسات من مجموعات البيانات، والتي خدمت كعينة تدريبية للخوارزمية لتتعلم الخصائص المرتبطة بصلات الدراسة. شملت المنهجية معيار توقف محدد، حيث تم تسجيل موقع آخر دراسة ذات صلة كنسبة مئوية من الحجم الإجمالي لمجموعة البيانات. كان الهدف من هذا النهج هو تقييم تباين هذه النسبة عبر مجموعات بيانات ومحاكاة مختلفة لتحديد ما إذا كان يمكن إنشاء حد عالمي.

استند التحليل إلى مجموعة بيانات SYNERGY، التي تتكون من 169,288 ورقة أكاديمية من 26 مراجعة منهجية، مع تصنيفات الصلة المقدمة من المؤلفين الأصليين بناءً على معايير الإدراج والاستبعاد المحددة مسبقًا. تم اختيار خمس مراجعات منهجية للمحاكاة، كل منها يختلف في عدد الدراسات ونسبة الدراسات ذات الصلة. تم تلخيص تفاصيل هذه المراجعات، بما في ذلك مواضيعها وإحصائيات الصلة، في الجدول 2. تم تقديم أسئلة البحث وإطارات PICO/PECO لهذه المراجعات في الجدول التكميلي S1، مما يبرز التأثير المحتمل لتعقيد الموضوع وتجانس الدراسة على فعالية معايير التوقف.

النتائج

أظهر تقييم ASReview لفرز النص والملخصات (T&A) تباينًا كبيرًا في نسبة مجموعات البيانات التي كان يجب فرزها لتحديد جميع الدراسات ذات الصلة. بينما أظهرت جميع مجموعات البيانات تقليصًا في النسبة المئوية للسجلات التي تتطلب الفرز مقارنة بالطرق اليدوية (100%)، كان مدى هذا التقليص متنوعًا بشكل واسع. أظهرت مجموعة بيانات “Bos (2018)” أكبر انخفاض، بينما أظهرت مجموعة بيانات “Donners (2022)” تحسنًا طفيفًا، مما استلزم فرز ما يصل إلى 94.5% من سجلاتها. يشير هذا إلى أن ASReview يمكن أن يعزز الكفاءة في تحديد الدراسات ذات الصلة، لكن فعاليته تعتمد على خصائص مجموعة البيانات المحددة.

استخدمت الدراسة نهجًا منهجيًا، حيث تم تغيير حجم عينة التدريب الأولية عبر سبعة مستويات (من 1% إلى 30%) وإجراء 1,000 محاكاة مستقلة لكل مجموعة بيانات، ليصل الإجمالي إلى 35,000 محاكاة. تم استخدام النتائج لتقييم ثلاثة معايير توقف للفرز بأثر رجعي. أظهر معيار التوقف 1 موثوقية محدودة، حيث توقع بدقة عدد الدراسات ذات الصلة في 2% فقط من المحاكاة، بينما غالبًا ما كان يبالغ في تقدير أو يفوت الدراسات ذات الصلة. لم يكن معيار التوقف 2 يحتوي على حد ثابت يمكن تطبيقه عبر مجموعات البيانات، كما اعتُبر معيار التوقف 3، الذي ينطوي على فرز نسبة ثابتة، غير فعال أيضًا بسبب التباين الكبير في نسبة الدراسات التي تحتاج إلى مراجعة عبر مجموعات البيانات. وبالتالي، تستنتج الدراسة أنه لم يتم تحديد أي معيار توقف قابل للتطبيق عالميًا، مما يبرز الحاجة إلى نهج مخصص في تطبيقات ASReview.

المناقشة

قيمت الدراسة ثلاثة معايير توقف لفرز النص والملخصات (T&A) في ASReview من خلال محاكاة بأثر رجعي عبر خمس مجموعات بيانات. تشير النتائج إلى أنه بينما يمكن لـ ASReview تقليل عبء العمل في الفرز بشكل كبير، فإن فعالية معايير التوقف تختلف حسب هيكل مجموعة البيانات. بشكل خاص، أظهرت مجموعات البيانات الأصغر (على سبيل المثال، أقل من 1,000 دراسة) فوائد ضئيلة، حيث كان من المحتمل أن تتطلب حوالي 95% من الدراسات الفرز في سيناريوهات أقل ملاءمة. لم يوفر أي من معايير التوقف التي تم اختبارها وسيلة موثوقة لإيقاف الفرز دون المخاطرة باستبعاد الدراسات ذات الصلة، مما أدى إلى الشكوك حول تطبيقها العملي.

تتوافق النتائج مع الانتقادات من الأدبيات السابقة بشأن قيود معايير التوقف الثابتة، مما يبرز أن القواعد العالمية قد لا تكون كافية دون النظر في خصائص مجموعة البيانات وتعقيد الموضوع. تسلط الدراسة الضوء على ضرورة أن يجمع الباحثون بين معايير التوقف مع تدابير ضمان الجودة الإضافية، مثل استشارات الخبراء، لتقليل خطر فقدان الدراسات ذات الصلة. بينما يمكن لـ ASReview تسريع عملية الفرز، فإن عدم وجود معيار توقف موثوق يثير تساؤلات حول فائدته العامة، لا سيما في مجموعات البيانات الأصغر. يجب أن تبحث الأبحاث المستقبلية في نماذج بديلة واستراتيجيات فرز لتعزيز كفاءة وموثوقية عمليات فرز T&A.

Journal: BMC Medical Research Methodology, Volume: 26, Issue: 1
DOI: https://doi.org/10.1186/s12874-026-02866-5
PMID: https://pubmed.ncbi.nlm.nih.gov/42106619
Publication Date: 2026-05-09
Author(s): C. Kempny et al.
Primary Topic: Meta-analysis and systematic reviews

Overview

This study investigates the effectiveness of three stop criteria for title and abstract (T&A) screening in ASReview, a machine learning tool designed to streamline systematic reviews. The criteria tested include: (1) stopping after identifying a calculated number of relevant studies from an initial sample, (2) stopping after a fixed number of consecutive irrelevant studies, and (3) stopping after screening a predefined percentage of the dataset. The simulation involved 35,000 automated screenings across five datasets from the SYNERGY repository. Results indicated significant variability in the percentage of the dataset that needed to be screened to identify all relevant studies, ranging from 2.9% to 76.9%. None of the stop criteria consistently ensured the identification of all relevant studies, with the first criterion proving reliable in only 2% of simulations.

The findings suggest that while ASReview can expedite the screening process by prioritizing potentially relevant studies, the lack of a reliable stop criterion raises concerns about the risk of missing relevant studies. The study emphasizes that the effectiveness of stop criteria varies significantly based on dataset characteristics, particularly in smaller datasets where the potential for missed studies is higher. Consequently, the authors recommend a cautious application of stop criteria, advocating for their combination with quality assurance measures, such as expert consultations or reference list screenings. Future research should focus on developing more robust stopping rules and exploring alternative models to enhance the efficiency of T&A screening in systematic reviews.

Introduction

The introduction of the research paper emphasizes the critical role of systematic reviews in evidence-based research across various disciplines, particularly in health research. It highlights that systematic reviews are foundational for clinical guidelines and policy decisions, while also noting the labor-intensive nature of the screening process, particularly title and abstract (T&A) screening. The increasing volume of published studies complicates this process, leading to potential delays in publication and outdated results. The authors advocate for more precise search strategies, ideally developed with specialized librarians, and the integration of artificial intelligence (AI) and machine learning (ML) tools, such as ASReview, to enhance efficiency and reduce the workload associated with screening.

The paper discusses the functionality of ASReview, which employs an active learning algorithm to prioritize studies based on user decisions, thereby improving the relevance of studies presented for review. The authors propose various criteria for stopping the screening process early, including the estimation of relevant studies from a sample, the number of consecutive irrelevant studies, and a predefined percentage of the dataset reviewed. They underscore the necessity of critically evaluating these stopping criteria to avoid the risk of excluding pertinent studies. The study aims to fill the existing research gap by determining the theoretical maximum effort savings in T&A screening and assessing practical stop criteria that minimize the risk of omitting relevant studies.

Methods

In this study, the authors conducted 35,000 automated title and abstract (T&A) screenings across five datasets, utilizing ASReview (version 1.6.1) and a custom Python script. Each screening began with a random subset of studies from the datasets, which served as a training sample for the algorithm to learn the characteristics associated with study relevance. The methodology included a specific stop criterion, where the position of the last relevant study was recorded as a percentage of the total dataset size. This approach aimed to evaluate the variability of this percentage across different datasets and simulations to determine if a universal threshold could be established.

The analysis was based on the SYNERGY dataset, which comprises 169,288 academic papers from 26 systematic reviews, with relevance classifications provided by the original authors based on pre-defined inclusion and exclusion criteria. Five systematic reviews were selected for the simulations, each varying in the number of studies and the percentage of relevant studies. The details of these reviews, including their topics and relevance statistics, are summarized in Table 2. The research questions and PICO/PECO frameworks for these reviews are provided in Supplementary Table S1, highlighting the potential impact of topic complexity and study homogeneity on the effectiveness of the stopping criteria.

Results

The evaluation of ASReview for text and abstract (T&A) screening revealed significant variability in the proportion of datasets that needed to be screened to identify all relevant studies. While all datasets demonstrated a reduction in the percentage of records requiring screening compared to manual methods (100%), the extent of this reduction varied widely. The “Bos (2018)” dataset exhibited the most substantial decrease, whereas the “Donners (2022)” dataset showed minimal improvement, necessitating screening of up to 94.5% of its records. This suggests that ASReview can enhance efficiency in identifying relevant studies, but its effectiveness is contingent on the specific dataset characteristics.

The study employed a systematic approach, varying the initial training sample size across seven levels (1% to 30%) and conducting 1,000 independent simulations per dataset, totaling 35,000 simulations. The results were used to retrospectively assess three stop criteria for screening. Stop criterion 1 demonstrated limited reliability, accurately predicting the number of relevant studies in only 2% of simulations, while often overestimating or missing relevant studies. Stop criterion 2 lacked a consistent threshold applicable across datasets, and stop criterion 3, which involved screening a fixed percentage, was also deemed ineffective due to the significant variability in the proportion of studies needing review across datasets. Consequently, the study concludes that no universally applicable stop criterion was identified, highlighting the need for tailored approaches in ASReview applications.

Discussion

The study assessed three stop criteria for text and abstract (T&A) screening in ASReview through retrospective simulations across five datasets. Findings indicate that while ASReview can significantly reduce the screening workload, the effectiveness of stop criteria varies by dataset structure. Particularly, smaller datasets (e.g., fewer than 1,000 studies) showed minimal benefits, as nearly 95% of studies would still require screening in less favorable scenarios. None of the stop criteria tested provided a reliable means to halt screening without risking the exclusion of relevant studies, leading to skepticism about their practical application.

The results align with critiques from previous literature regarding the limitations of fixed stopping criteria, emphasizing that universal rules may not suffice without considering dataset characteristics and topic complexity. The study highlights the necessity for researchers to combine stop criteria with additional quality assurance measures, such as expert consultations, to mitigate the risk of missing relevant studies. While ASReview can expedite the screening process, the lack of a dependable stop criterion raises questions about its overall utility, particularly in smaller datasets. Future research should investigate alternative models and screening strategies to enhance the efficiency and reliability of T&A screening processes.

شارك: