DOI: https://doi.org/10.1001/jamanetworkopen.2024.29630
PMID: https://pubmed.ncbi.nlm.nih.gov/39576645
تاريخ النشر: 2024-11-22
المؤلف: Brooke Levis وآخرون
الموضوع الرئيسي: تحليل البيانات الشامل والمراجعات المنهجية
نظرة عامة
تتناول هذه القسم من ورقة البحث تأثير استخدام مجموعات بيانات صغيرة في تحديد درجات القطع المثلى لاستبيان صحة المريض-9 (PHQ-9) وتقدير دقتها. تهدف الدراسة إلى تقييم ما إذا كانت الطرق المعتمدة على البيانات تؤدي إلى درجات قطع تختلف عن الدرجة المثلى على مستوى السكان وما إذا كانت تنتج تقديرات دقة متحيزة. باستخدام بيانات مقطعية من تحليل بيانات المشاركين الفرديين (IPDMA) الذي شمل 100 دراسة أساسية و44,503 مشارك، أعادت الدراسة أخذ عينات من الدراسات بأحجام مشاركين مختلفة لتقييم درجات القطع المثلى بناءً على مؤشر يودن.
تكشف النتائج عن تباين كبير في درجات القطع المثلى المحددة عبر الدراسات المحاكاة، حيث تتراوح الدرجات من 2 إلى 21 للعينات الصغيرة ومن 5 إلى 11 للعينات الأكبر. فقط 17% من الدراسات التي شملت 100 مشارك و33% من الدراسات التي شملت 1000 مشارك حددت بشكل صحيح درجة القطع المثلى الحقيقية التي تبلغ 8 أو أعلى. علاوة على ذلك، وجدت الدراسة أن تقديرات الدقة غالبًا ما كانت مبالغ فيها، خاصة في العينات الصغيرة، حيث انخفضت مبالغة الحساسية مع زيادة أحجام العينات، بينما ظلت الخصوصية مستقرة نسبيًا. يستنتج المؤلفون أنه يجب على الباحثين إعطاء الأولوية لأحجام عينات كافية وطرق إحصائية قوية لضمان تحديد دقيق لدرجات القطع وتقديرات الدقة، مع تقديم نصيحة بالحذر في تفسير النتائج من الدراسات ذات الأعداد الصغيرة من المشاركين.
مقدمة
تسلط مقدمة هذه الورقة البحثية الضوء على التحديات المرتبطة بدقة أدوات فحص الاكتئاب، مع التركيز بشكل خاص على استبيان صحة المريض-9 (PHQ-9). تشير إلى أن العديد من الدراسات تستخدم أحجام عينات صغيرة وطرق معتمدة على البيانات لتحديد درجات القطع المثلى، مما يؤدي غالبًا إلى مبالغة في تقدير الدقة. أظهرت مراجعة لـ 172 دراسة أن الحجم الوسيط للعينة كان 194، مع حوالي 20 حالة من الاكتئاب، و76% من هذه الدراسات حددت درجات قطع انحرفت عن المعايير المعتمدة. ومن الجدير بالذكر أن 40% من المؤلفين أوصوا بدرجات القطع المستمدة من دراساتهم بدلاً من المعايير القياسية.
تشير الورقة إلى تحليل كبير لمقياس اكتئاب ما بعد الولادة في إدنبرة (EPDS) الذي شمل 13,255 مشاركًا، والذي أظهر تباينًا في درجات القطع المثلى بناءً على حجم العينة. على وجه التحديد، وجدت الدراسة أن درجة القطع التي تعظم مؤشر يودن تراوحت بشكل واسع اعتمادًا على حجم العينة، مع مبالغات متوسطة في الحساسية بلغت 7 نقاط مئوية للعينات الصغيرة. يهدف المؤلفون إلى التحقيق في كيفية مساهمة الطرق المعتمدة على البيانات لاختيار درجات قطع PHQ-9 في تقديرات الدقة المتحيزة، من خلال مقارنة هذه النتائج مع درجات القطع المثلى على مستوى السكان لتقييم مدى هذا التحيز عبر أحجام عينات مختلفة.
طرق البحث
في هذه الدراسة، صنف لجنة أخلاقيات البحث في مستشفى اليهود العام البحث الذي يتضمن تحليل بيانات المشاركين الفرديين (IPDMA) لمجموعات بيانات مجهولة الهوية مسبقًا على أنه معفى من متطلبات الموافقة الأخلاقية والموافقة المستنيرة. تم منح هذا الإعفاء على أساس أن البيانات المستخدمة قد تم الحصول عليها بالفعل بشكل أخلاقي، حيث حصلت الدراسات الأصلية على الموافقات الأخلاقية اللازمة وموافقة المشاركين قبل جمع البيانات.
اتبعت المنهجية المعايير الأخلاقية المعتمدة من خلال ضمان أن جميع مجموعات البيانات المدرجة في التحليل قد تم فحصها مسبقًا للتأكد من الامتثال للإرشادات الأخلاقية، مما يضمن حماية سرية المشاركين ونزاهتهم طوال عملية البحث.
النتائج
تشمل النتائج من قاعدة بيانات تحليل بيانات المرضى المتكامل (IPDMA) 100 دراسة أساسية تضم 44,503 مشارك، من بينهم 4,541 (10%) تم تحديدهم كحالات اكتئاب شديدة. كان الحجم الوسيط للعينة عبر هذه الدراسات 194 (نطاق الربع [IQR] 134-386)، مع وسطاء 28 (IQR 14-60) حالة من الاكتئاب الشديد لكل دراسة. تباينت درجات القطع المثلى لتحديد الاكتئاب الشديد بشكل كبير، حيث تراوحت من 3 إلى 18، مع درجة قطع وسيطة تبلغ 10.
كشفت تحليل درجات استبيان صحة المريض-9 (PHQ-9) عن توزيع طبيعي بين الحالات (المتوسط [SD] = 13 [6]؛ الوسيط [IQR] = 13 [9-18]) وتوزيع مائل إلى اليمين بين غير الحالات (المتوسط [SD] = 4 [4]؛ الوسيط [IQR] = 3 [1-6]). في إجمالي السكان الذين تم تحليلهم، وُجد أن الحساسية والخصوصية غير الموزونة لدرجة PHQ-9 تبلغ 8 أو أعلى كانت 80.4% و82.0%، على التوالي، مما يشير إلى أداء قوي لاستبيان PHQ-9 في التمييز بين حالات الاكتئاب الشديد وغير الحالات.
المناقشة
تستعرض الدراسة المقدمة في هذا القسم التحيز في تقديرات الدقة لاستبيان صحة المريض-9 (PHQ-9) بسبب اختيار درجات القطع المثلى المعتمدة على البيانات. تشير النتيجة الرئيسية إلى أن هذه الدرجات غالبًا ما تنحرف بشكل كبير عن الدرجة المثلى على مستوى السكان، مما يؤدي إلى تقديرات حساسية وخصوصية متحيزة. مع زيادة أحجام العينات من 100 إلى 1000 مشارك، انخفض التباين في درجات القطع المثلى، وتقلصت مبالغة الحساسية. على وجه التحديد، وجدت الدراسة أنه في العينات الصغيرة، تراوحت درجات القطع المثلى بشكل واسع، بينما أسفرت العينات الأكبر عن نتائج أكثر اتساقًا، مما يبرز تأثير حجم العينة على دقة التشخيص.
تشير تداعيات هذه النتائج إلى أنه يجب على الأطباء وصانعي السياسات توخي الحذر عند تفسير درجات القطع المثلى المستمدة من الدراسات الصغيرة. يجب أن تستند التوصيات لدرجات القطع بشكل مثالي إلى تحليلات ميتا واسعة النطاق وموثوقة بدلاً من الدراسات الفردية ذات أحجام العينات المحدودة. علاوة على ذلك، تؤكد الدراسة على أهمية إجراء حسابات حجم العينة مسبقًا لضمان تمثيل كافٍ لكل من الحالات وغير الحالات في أبحاث دقة التشخيص. يُشجع الباحثون على الإبلاغ عن تقديرات الدقة لجميع درجات القطع المحتملة واستخدام طرق إحصائية متقدمة لتعزيز موثوقية تحديد درجات القطع وتقدير الدقة. بشكل عام، تؤكد الدراسة على الحاجة إلى تقييم دقيق لدراسات دقة PHQ-9 لتجنب استنتاجات مضللة بشأن ممارسات فحص الاكتئاب.
القيود
تشمل نقاط القوة في الدراسة استخدام حجم عينة كبير وبيانات حقيقية للمشاركين، مما يعزز من قوة النتائج. ومع ذلك، توجد قيود ملحوظة. أولاً، لم تتضمن الدراسة مجموعات بيانات من الدراسات المنشورة مؤخرًا حول دقة استبيان صحة المريض-9 (PHQ-9). على الرغم من هذا الإغفال، يجادل المؤلفون بأن تضمين الدراسات الأحدث من غير المحتمل أن يؤثر بشكل كبير على النتائج، حيث من المحتمل أن تظهر هذه الدراسات أحجام عينات مماثلة وتنوعًا.
ثانيًا، اعتمدت الدراسة فقط على مؤشر يودن لتحديد درجات القطع المثلى. على الرغم من أن هذه الطريقة مستخدمة على نطاق واسع في دراسات دقة فحص الاكتئاب وتظهر أداءً مشابهًا لمؤشرات أخرى، مثل المسافة الإقليدية، إلا أنه يُعترف بأن مؤشر يودن يمكن أن يكون غير موثوق وقد يؤدي إلى مبالغة في النتائج. وبالتالي، قد تؤدي الطرق البديلة إلى نتائج مختلفة قليلاً، مما يشير إلى منطقة محتملة لمزيد من التحقيق.
DOI: https://doi.org/10.1001/jamanetworkopen.2024.29630
PMID: https://pubmed.ncbi.nlm.nih.gov/39576645
Publication Date: 2024-11-22
Author(s): Brooke Levis et al.
Primary Topic: Meta-analysis and systematic reviews
Overview
This section of the research paper investigates the impact of using small datasets in determining optimal cutoff scores for the Patient Health Questionnaire-9 (PHQ-9) and estimating their accuracy. The study aims to assess whether data-driven methods yield cutoff scores that diverge from the population-level optimal score and whether they produce biased accuracy estimates. Utilizing cross-sectional data from an individual participant data meta-analysis (IPDMA) involving 100 primary studies and 44,503 participants, the research resampled studies with varying participant sizes to evaluate the optimal cutoff scores based on the Youden index.
The findings reveal significant variability in the optimal cutoff scores identified across simulated studies, with scores ranging from 2 to 21 for smaller samples and from 5 to 11 for larger samples. Only 17% of studies with 100 participants and 33% with 1000 participants correctly identified the true optimal cutoff score of 8 or higher. Moreover, the study found that accuracy estimates were often exaggerated, particularly in smaller samples, where sensitivity overestimation decreased with larger sample sizes, while specificity remained relatively stable. The authors conclude that researchers should prioritize adequate sample sizes and robust statistical methods to ensure accurate identification of cutoff scores and accuracy estimates, advising caution in interpreting findings from studies with small participant numbers.
Introduction
The introduction of this research paper highlights the challenges associated with the accuracy of depression screening tools, particularly focusing on the Patient Health Questionnaire-9 (PHQ-9). It notes that many studies utilize small sample sizes and data-driven methods to establish optimal cutoff scores, often leading to overestimation of accuracy. A review of 172 studies revealed a median sample size of 194, with only about 20 cases of depression, and 76% of these studies identified cutoff scores that deviated from established standards. Notably, 40% of the authors recommended their derived cutoff scores over the standard ones.
The paper references a significant analysis of the Edinburgh Postnatal Depression Scale (EPDS) involving 13,255 participants, which demonstrated variability in optimal cutoff scores based on sample size. Specifically, the study found that the cutoff score maximizing the Youden index ranged widely depending on the sample size, with mean sensitivity overestimations of 7 percentage points for smaller samples. The authors aim to investigate how data-driven methods for selecting PHQ-9 cutoff scores may contribute to biased accuracy estimates, comparing these findings against population-level optimal cutoff scores to assess the extent of this bias across varying sample sizes.
Methods
In this study, the Jewish General Hospital Research Ethics Committee classified the research involving individual participant data meta-analysis (IPDMA) of pre-existing deidentified datasets as exempt from ethics approval and informed consent requirements. This exemption was granted on the basis that the data utilized had already been ethically sourced, with original studies having obtained the necessary ethics approvals and participant consent prior to data collection.
The methodology adhered to established ethical standards by ensuring that all datasets included in the analysis were previously vetted for compliance with ethical guidelines, thereby safeguarding participant confidentiality and integrity throughout the research process.
Results
The results from the Integrated Patient Data Meta-Analysis (IPDMA) database encompass 100 primary studies involving 44,503 participants, of which 4,541 (10%) were identified as cases of major depression. The median sample size across these studies was 194 (interquartile range [IQR] 134-386), with a median of 28 (IQR 14-60) cases of major depression per study. The optimal cutoff scores for identifying major depression varied significantly, ranging from 3 to 18, with a median cutoff of 10.
Analysis of the Patient Health Questionnaire-9 (PHQ-9) scores revealed a normal distribution among cases (mean [SD] = 13 [6]; median [IQR] = 13 [9-18]) and a right-skewed distribution among noncases (mean [SD] = 4 [4]; median [IQR] = 3 [1-6]). In the overall population analyzed, the unweighted sensitivity and specificity for a PHQ-9 score of 8 or higher were found to be 80.4% and 82.0%, respectively, indicating a robust performance of the PHQ-9 in distinguishing between cases and noncases of major depression.
Discussion
The study presented in this section investigates the bias in accuracy estimates for the Patient Health Questionnaire-9 (PHQ-9) due to the selection of data-driven optimal cutoff scores. The main finding indicates that these cutoff scores often diverge significantly from the population-level optimal score, leading to biased sensitivity and specificity estimates. As sample sizes increased from 100 to 1000 participants, the variability in optimal cutoff scores decreased, and the overestimation of sensitivity diminished. Specifically, the study found that in smaller samples, optimal cutoff scores ranged widely, while larger samples yielded more consistent results, highlighting the influence of sample size on diagnostic accuracy.
The implications of these findings suggest that clinicians and policymakers should exercise caution when interpreting optimal cutoff scores derived from small studies. Recommendations for cutoff scores should ideally be based on large, rigorously conducted meta-analyses rather than individual studies with limited sample sizes. Furthermore, the study emphasizes the importance of conducting a priori sample size calculations to ensure adequate representation of both cases and noncases in diagnostic accuracy research. Researchers are encouraged to report accuracy estimates for all potential cutoff scores and to utilize advanced statistical methods to enhance the reliability of cutoff score identification and accuracy estimation. Overall, the study underscores the need for careful evaluation of PHQ-9 accuracy studies to avoid misleading conclusions regarding depression screening practices.
Limitations
The study’s strengths include the utilization of a large sample size and real participant data, which enhance the robustness of the findings. However, notable limitations exist. Firstly, the research did not incorporate datasets from recently published studies on the accuracy of the Patient Health Questionnaire-9 (PHQ-9). Despite this omission, the authors argue that the inclusion of newer studies is unlikely to significantly impact the results, as these studies would probably exhibit similar sample sizes and heterogeneity.
Secondly, the study relied solely on the Youden index for determining optimal cutoff scores. While this method is widely used in depression screening accuracy studies and shows comparable performance to other indices, such as the Euclidean distance, it is acknowledged that the Youden index can be unreliable and may lead to overestimation of results. Consequently, alternative methods could yield slightly different outcomes, indicating a potential area for further investigation.
