DOI: https://doi.org/10.1186/s40537-023-00857-7
تاريخ النشر: 2024-01-03
المؤلف: Cynthia Yang وآخرون
الموضوع الرئيسي: تقنيات تصنيف البيانات غير المتوازنة
نظرة عامة
في هذه الدراسة، قام المؤلفون بفحص تأثيرات العينة العشوائية الزائدة والعينة العشوائية الناقصة على أداء نماذج التنبؤ السريرية، تحديدًا ضمن مجموعة من الأفراد الذين يعانون من الاكتئاب المعالج بالأدوية. باستخدام أربعة قواعد بيانات صحية كبيرة، قام الباحثون بتطوير والتحقق الخارجي من إجمالي 1,566 نموذجًا للتنبؤ باستخدام ثلاثة مصنفات: الانحدار اللوجستي لاسو، والغابة العشوائية، وXGBoost. قامت الدراسة بتغيير نسب عدم التوازن المستهدفة وقامت بتقييم أداء النموذج من خلال التمييز، الذي تم قياسه بواسطة المساحة تحت منحنى التشغيل الخاص بالمستقبل (AUROC)، والمعايرة، التي تم تقييمها عبر مخططات المعايرة.
أظهرت النتائج أن لا العينة العشوائية الزائدة ولا العينة العشوائية الناقصة حسنت بشكل كبير AUROC خلال التحقق الداخلي والخارجي. بالإضافة إلى ذلك، أظهرت النماذج مخاطر مبالغ فيها، والتي يمكن تصحيحها إلى حد كبير من خلال إعادة المعايرة بما يتماشى مع نسب عدم التوازن في مجموعة البيانات الأصلية. وبالتالي، يستنتج المؤلفون أن تطبيق هذه الطرق لعدم التوازن في الفئات لا يحسن أداء نماذج التنبؤ في قواعد البيانات الصحية الكبيرة، موصين بعدم استخدامها في مثل هذه السياقات.
مقدمة
تتناول مقدمة هذه الورقة البحثية القضية الشائعة لعدم التوازن في الفئات في نمذجة التنبؤ السريري، خاصة ضمن مجموعات البيانات المستمدة من السجلات الصحية الإلكترونية (EHR). يحدث عدم التوازن في الفئات عندما يكون توزيع فئات النتائج غير متساوٍ، مما يؤدي غالبًا إلى مصنفات متحيزة تفضل الفئة الغالبة وتخطئ في تصنيف الفئة الأقل. يشير المؤلفون إلى أنه على الرغم من اقتراح طرق مختلفة لعدم التوازن في الفئات لتحسين أداء التنبؤ، لا يزال هناك نقص في الإجماع بشأن فعاليتها، خاصة في سياق قواعد البيانات الصحية الكبيرة. أسفرت الدراسات السابقة عن نتائج مختلطة، حيث اقترح البعض تحسينات في أداء النموذج، بينما أشار آخرون، وخاصة أولئك الذين يركزون على الانحدار اللوجستي، إلى أن تقنيات مثل العينة العشوائية الزائدة وتقنية العينة الأقل الاصطناعية (SMOTE) قد لا تعزز تمييز النموذج.
تهدف الدراسة إلى تقييم تأثير العينة العشوائية الزائدة والعينة العشوائية الناقصة بشكل تجريبي على أداء التحقق الداخلي والخارجي لنماذج التنبؤ للمرضى الذين يعانون من الاكتئاب المعالج بالأدوية. باستخدام أربعة قواعد بيانات صحية كبيرة، طور المؤلفون نماذج باستخدام ثلاثة مصنفات—الانحدار اللوجستي لاسو، والغابة العشوائية، وXGBoost—بينما قاموا بتغيير نسبة عدم التوازن المستهدفة. تسعى هذه التحقيقات إلى تحديد ما إذا كانت طرق عدم التوازن في الفئات يمكن أن تعزز القابلية العامة والصلابة لنماذج التنبؤ السريرية، خاصة في سياق التحقق الخارجي، الذي يقيم أداء النموذج على مجموعات بيانات مستقلة لم تُستخدم خلال تدريب النموذج.
الطرق
في هذه الدراسة، استخدم المؤلفون إطار عمل التنبؤ على مستوى المريض (PLP) من مبادرة علوم البيانات الصحية والمعلوماتية (OHDSI) لتطوير والتحقق من نماذج التنبؤ. يستخدم إطار عمل OHDSI نموذج البيانات الشائعة لشراكة النتائج الطبية الملاحظة (OMOP CDM) لتوحيد مصادر البيانات المتنوعة، مما يسهل إنشاء نماذج تنبؤ قابلة للتشغيل المتبادل عبر قواعد بيانات صحية ملاحظة مختلفة. تم إنتاج إجمالي 1,566 نموذجًا للتنبؤ من خلال تغيير مهام التنبؤ، واستراتيجيات العينة، والمصنفات، مع التركيز بشكل خاص على 58 مهمة تنبؤ متميزة مستمدة من أربعة قواعد بيانات.
شملت المنهجية نهج العينة المتدرجة حيث تم تخصيص 75% من سكان المرضى كمجموعة تدريب، بينما خدمت النسبة المتبقية 25% كمجموعة اختبار. لتعزيز القوة الإحصائية، تم استبعاد المهام التي تحتوي على أقل من 100 حدث نتيجة في مجموعة الاختبار، مما أدى إلى استبعاد نتيجتين من التحليل. تراوحت نسبة عدم التوازن (IR) للبيانات من 8.6 إلى 245.3، مع متوسط قدره 84.0. نفذ المؤلفون استراتيجيات عينة متنوعة، بما في ذلك العينة العشوائية الزائدة والعينة الناقصة، مستهدفين نسب عدم التوازن المحددة لمعالجة عدم التوازن في الفئات. تم استخدام ثلاثة مصنفات—الانحدار اللوجستي المنظم L1، والغابة العشوائية، وXGBoost—مع إجراء ضبط للمعلمات الفائقة من خلال التحقق المتقاطع ثلاثي الأضعاف على مجموعة التدريب. تم التحقق من النماذج النهائية على مجموعة الاختبار، مما يضمن تقييمًا واقعيًا لأدائها التنبؤي.
النتائج
في هذه الدراسة، تم إجراء تقييم شامل لمختلف استراتيجيات العينة عبر 58 مهمة تنبؤ باستخدام ثلاثة مصنفات مختلفة، مما أدى إلى تطوير والتحقق الخارجي من 1,566 نموذجًا للتنبؤ. تراوحت المساحة تحت منحنى التشغيل الخاص بالمستقبل (AUROC) للنماذج الأصلية من 0.58 إلى 0.87. أظهر التحليل أنه بينما أسفرت بعض استراتيجيات العينة عن فرق إيجابي في AUROC، مما يشير إلى تحسين تمييز النموذج، لم تعزز العينة العشوائية الزائدة والعينة الناقصة عمومًا أداء AUROC. على وجه الخصوص، كان تأثير العينة العشوائية ضئيلًا بالنسبة للانحدار اللوجستي لاسو وXGBoost، مع اختلافات AUROC المطلقة القصوى أقل من 0.06. ومع ذلك، أظهرت العينة العشوائية الزائدة مع الغابة العشوائية تأثيرًا أكثر وضوحًا، حيث وصلت اختلافات AUROC إلى حوالي -0.3، خاصة عندما كان عدد أحداث النتائج منخفضًا.
علاوة على ذلك، قامت الدراسة بتقييم الحد الأقصى لدرجة F1 ووجدت أن استراتيجيات العينة لم تحسن هذه المقياس بشكل كبير أيضًا. أشارت مخططات المعايرة إلى أن جميع استراتيجيات العينة أدت إلى تدهور معايرة النموذج، خاصة بالنسبة لمعدلات الحدوث المستهدفة الأصغر (IRs)، مما أدى إلى زيادة المبالغة في التقدير. أظهرت جهود إعادة المعايرة بعض النجاح في استعادة المعايرة إلى مستويات مشابهة لنماذج البيانات الأصلية، على الرغم من أن العينة العشوائية الزائدة مع الغابة العشوائية كانت تميل إلى التقليل من المخاطر بعد إعادة المعايرة. في النهاية، لم تؤدِ اختيار أفضل استراتيجية عينة بناءً على AUROC للتحقق الداخلي إلى تحسين أداء التحقق الخارجي، حيث لوحظت نتائج متسقة عبر كلا مرحلتي التحقق، مما يؤكد أن العينة العشوائية الزائدة والعينة الناقصة عمومًا لم تعزز نتائج AUROC أو درجة F1.
المناقشة
في هذه الدراسة، قام المؤلفون بفحص تأثيرات العينة العشوائية الزائدة والعينة العشوائية الناقصة على أداء النماذج التنبؤية لمختلف النتائج في مجموعة مستهدفة من المرضى الذين يعانون من الاكتئاب المعالج بالأدوية، باستخدام بيانات من أربعة قواعد بيانات صحية كبيرة. شمل التحليل 58 مهمة تنبؤ، باستخدام مصنفات مختلفة (الانحدار اللوجستي لاسو، الغابة العشوائية، XGBoost) واستراتيجيات العينة. أشارت النتائج إلى أن لا العينة العشوائية الزائدة ولا العينة العشوائية الناقصة حسنت بشكل كبير أداء النموذج، كما تم قياسه بواسطة المساحة تحت منحنى التشغيل الخاص بالمستقبل (AUROC). بشكل خاص، أظهرت النماذج التي تستخدم العينة العشوائية الزائدة مع الغابة العشوائية انخفاضات ملحوظة في AUROC، خاصة في مجموعات البيانات التي تحتوي على عدد أقل من أحداث النتائج.
سلطت الدراسة أيضًا الضوء على قضية عدم معايرة النموذج الناتجة عن استراتيجيات العينة، حيث أدت كل من العينة الزائدة والعينة الناقصة إلى مخاطر مبالغ فيها. على الرغم من أن إعادة المعايرة نحو معدلات الحدوث الأصلية صححت جزئيًا هذه عدم المعايرة، إلا أنها أبرزت ضرورة تقييم معايرة النموذج عند تطبيق تقنيات العينة العشوائية. استنتج المؤلفون أن نتائجهم لا تدعم استخدام العينة العشوائية الزائدة أو العينة العشوائية الناقصة في تطوير نماذج التنبؤ ضمن قواعد البيانات الصحية الكبيرة، ويوصون بمزيد من الاستكشاف لطرق عدم التوازن في الفئات البديلة للبحث المستقبلي.
DOI: https://doi.org/10.1186/s40537-023-00857-7
Publication Date: 2024-01-03
Author(s): Cynthia Yang et al.
Primary Topic: Imbalanced Data Classification Techniques
Overview
In this study, the authors examined the effects of random oversampling and random undersampling on the performance of clinical prediction models, specifically within a population of individuals with pharmaceutically treated depression. Utilizing four large observational health databases, the researchers developed and externally validated a total of 1,566 prediction models employing three classifiers: lasso logistic regression, random forest, and XGBoost. The study varied the target imbalance ratios and assessed model performance through discrimination, measured by the area under the receiver operating characteristic curve (AUROC), and calibration, evaluated via calibration plots.
The findings revealed that neither random oversampling nor random undersampling significantly enhanced the AUROC during internal and external validation. Additionally, the models exhibited overestimated risks, which could be largely rectified through recalibration aligned with the original dataset’s imbalance ratios. Consequently, the authors conclude that the application of these class imbalance methods does not improve the performance of prediction models in large observational health databases, recommending against their use in such contexts.
Introduction
The introduction of this research paper addresses the prevalent issue of class imbalance in clinical prediction modeling, particularly within datasets derived from electronic health records (EHR). Class imbalance occurs when the distribution of outcome classes is unequal, often leading to biased classifiers that favor the majority class and misclassify the minority class. The authors note that while various class imbalance methods have been proposed to enhance prediction performance, there remains a lack of consensus regarding their effectiveness, especially in the context of large observational health databases. Previous studies have yielded mixed results, with some suggesting improvements in model performance, while others, particularly those focusing on logistic regression, indicate that techniques like random oversampling and Synthetic Minority Oversampling Technique (SMOTE) may not enhance model discrimination.
The study aims to empirically evaluate the impact of random oversampling and random undersampling on both internal and external validation performance of prediction models for patients with pharmaceutically treated depression. Utilizing four large observational health databases, the authors developed models employing three classifiers—lasso logistic regression, random forest, and XGBoost—while varying the target imbalance ratio. This investigation seeks to determine whether class imbalance methods can enhance the generalizability and robustness of clinical prediction models, particularly in the context of external validation, which assesses model performance on independent datasets not used during model training.
Methods
In this study, the authors employed the Patient-Level Prediction (PLP) framework from the Observational Health Data Sciences and Informatics (OHDSI) initiative to develop and validate prediction models. The OHDSI framework utilizes the Observational Medical Outcomes Partnership Common Data Model (OMOP CDM) to standardize heterogeneous data sources, facilitating the creation of interoperable prediction models across various observational health databases. A total of 1,566 prediction models were generated by varying prediction tasks, sampling strategies, and classifiers, specifically focusing on 58 distinct prediction tasks derived from four databases.
The methodology involved a stratified sampling approach where 75% of the patient population was designated as a training set, while the remaining 25% served as a test set. To enhance statistical power, tasks with fewer than 100 outcome events in the test set were excluded, leading to the omission of two outcomes from analysis. The imbalance ratio (IR) of the data ranged from 8.6 to 245.3, with a median of 84.0. The authors implemented various sampling strategies, including random oversampling and undersampling, targeting specific IRs to address class imbalance. Three classifiers—L1-regularized logistic regression, random forest, and XGBoost—were utilized, with hyperparameter tuning conducted through threefold cross-validation on the training set. The final models were validated on the test set, ensuring a realistic evaluation of their predictive performance.
Results
In this study, a comprehensive evaluation of various sampling strategies was conducted across 58 prediction tasks using three different classifiers, resulting in the development and external validation of 1,566 prediction models. The area under the receiver operating characteristic curve (AUROC) for the original models ranged from 0.58 to 0.87. The analysis revealed that while some sampling strategies yielded a positive AUROC difference, indicating improved model discrimination, random oversampling and undersampling generally did not enhance AUROC performance. Specifically, the impact of random sampling was minimal for lasso logistic regression and XGBoost, with maximum absolute AUROC differences below 0.06. However, random oversampling with random forest exhibited a more pronounced effect, with AUROC differences reaching approximately -0.3, particularly when the number of outcome events was low.
Furthermore, the study assessed the maximum F1-score and found that sampling strategies did not significantly improve this metric either. Calibration plots indicated that all sampling strategies led to deteriorated model calibration, particularly for smaller target incidence rates (IRs), resulting in increased overestimation. Recalibration efforts showed some success in restoring calibration to levels similar to the original data models, although random oversampling with random forest tended to underestimate risks post-recalibration. Ultimately, the selection of the best sampling strategy based on internal validation AUROC did not translate to improved external validation performance, as consistent results were observed across both validation phases, reaffirming that random oversampling and undersampling generally did not enhance AUROC or F1-score outcomes.
Discussion
In this study, the authors examined the effects of random oversampling and undersampling on the performance of predictive models for various outcomes in a target population of patients with pharmaceutically treated depression, utilizing data from four large observational health databases. The analysis involved 58 prediction tasks, employing different classifiers (lasso logistic regression, random forest, XGBoost) and sampling strategies. The findings indicated that neither random oversampling nor undersampling significantly improved model performance, as measured by the area under the receiver operating characteristic curve (AUROC). In particular, models utilizing random oversampling with random forest exhibited notable decreases in AUROC, especially in datasets with fewer outcome events.
The study also highlighted the issue of model miscalibration resulting from sampling strategies, with both oversampling and undersampling leading to overestimated risks. Although recalibration towards the original incidence rates partially corrected this miscalibration, it underscored the necessity of evaluating model calibration when applying random sampling techniques. The authors concluded that their results do not support the use of random oversampling or undersampling in developing predictive models within large observational health databases, and they recommend further exploration of alternative class imbalance methods for future research.
