طرق الانحدار المعاقبة مع التحقق المتقاطع المعدل وضبط البوتستراب تنتج نماذج توقع أفضل
Penalized Regression Methods With Modified Cross‐Validation and Bootstrap Tuning Produce Better Prediction Models

شارك:
المجلة: Biometrical Journal، المجلد: 66، العدد: 5
DOI: https://doi.org/10.1002/bimj.202300245
PMID: https://pubmed.ncbi.nlm.nih.gov/38922968
تاريخ النشر: 2024-06-24
المؤلف: Menelaos Pavlou وآخرون
الموضوع الرئيسي: طرق إحصائية واستدلال

نظرة عامة

في هذا القسم، يتناول المؤلفون قضية الإفراط في التكيف في نماذج توقع المخاطر التي تستخدم تقدير الاحتمالية القصوى (MLE)، مما يؤدي غالبًا إلى توقعات متطرفة ومنحدر معايرة (CS) أقل من 1. للتخفيف من ذلك، يتم استخدام طرق معاقبة مثل Ridge و Lasso، حيث تقلل من معاملات الانحدار نحو الصفر، مما ينتج عنه توقعات تتماشى بشكل أكبر مع المتوسط. ومع ذلك، يمكن أن يؤدي معامل الضبط، $\lambda$، الذي يتم اختياره من خلال التحقق المتقاطع القياسي إلى الإفراط في الانكماش والتباين، وهو ما يمثل مشكلة خاصة في أحجام العينات الصغيرة.

يقترح المؤلفون نهجًا جديدًا يسمى “الضبط المعدل”، والذي يتضمن تقدير $\lambda$ من مجموعة بيانات تطوير زائفة أكبر تم إنشاؤها من خلال إعادة أخذ العينات من مجموعة البيانات الأصلية. تضمن هذه الطريقة أن تكون مجموعات بيانات التدريب للتحقق المتقاطع قابلة للمقارنة في الحجم مع مجموعة البيانات الأصلية، مما يقلل من خطر الإفراط في تقدير $\lambda$. تُظهر تقييمات طرق الضبط المعدل وإعادة أخذ العينات لـ Ridge و Lasso على كل من مجموعات البيانات المحاكاة والحقيقية تحسينات كبيرة في اختيار $\lambda$ و CS، فضلاً عن دقة التوقعات المحسنة مقارنةً بتقدير الاحتمالية القصوى التقليدي وطرق الضبط القياسية.

مقدمة

تناقش مقدمة هذه الورقة البحثية الدور الحاسم لنماذج توقع المخاطر في اتخاذ القرارات السريرية، خاصة تلك المستمدة من الانحدار اللوجستي للنتائج الثنائية. تسلط الضوء على ضرورة التحقق من صحة هذه النماذج باستخدام مجموعات بيانات خارجية لضمان دقتها التنبؤية، والتي يتم تقييمها من خلال مقاييس مختلفة مثل منحدر المعايرة (CS)، إحصائية C، درجة Brier، والفائدة الصافية. تحدد الورقة مشكلة شائعة تتمثل في الإفراط في التكيف في نماذج الانحدار اللوجستي، خاصة عندما تكون حجم العينة صغيرًا بالنسبة لعدد المعلمات، مما يؤدي إلى توقعات متطرفة للمرضى الجدد. تم تقديم طرق الانحدار اللوجستي المعاقبة، مثل Ridge و Lasso، كحلول للتخفيف من الإفراط في التكيف من خلال تقليل معاملات الانحدار، على الرغم من أنها قد تُدخل أيضًا تباينًا في مقاييس المعايرة.

يلخص المؤلفون الدراسات السابقة التي استكشفت فعالية الطرق المعاقبة في الإعدادات منخفضة الأبعاد وأدائها عبر سيناريوهات مختلفة من انتشار الأحداث وارتباطات المتنبئين. يلاحظون أنه بينما تقلل الطرق المعاقبة عمومًا من الإفراط في التكيف، إلا أنها يمكن أن تؤدي أحيانًا إلى الإفراط في الانكماش وتظهر تباينًا عاليًا في تقدير معامل الضبط، مما قد يؤدي إلى نماذج غير متوازنة. تقترح الورقة طريقة ضبط متقاطعة معدلة، تشبه ضبط إعادة أخذ العينات، لتعزيز أداء الطرق المعاقبة في مجموعات البيانات الأكبر. ستفصل الأقسام اللاحقة تقنيات الانحدار المعاقبة القياسية، وطرق الضبط المقترحة، وتحليلًا مقارنًا للأداء التنبؤي باستخدام المحاكاة ومجموعة بيانات قلبية حقيقية. يهدف المؤلفون إلى إظهار الفائدة العملية لنهج الضبط المعدل الخاص بهم مع معالجة قيوده.

طرق

يستعرض قسم الطرق في الورقة البحثية الأساليب الإحصائية المستخدمة لتوقع المخاطر في النتائج الثنائية، باستخدام نماذج الانحدار اللوجستي بشكل خاص. يتم نمذجة النتيجة الثنائية، المشار إليها بـ \(Y_i\)، مع المتغيرات المرافقة التي تمثلها المتجهة \(\mathbf{x}_i\). يقدر إطار الانحدار اللوجستي احتمال الحدث المعني من خلال دالة اللوغيت، حيث يتم عادةً تقدير المعاملات \(\boldsymbol{\beta}\) باستخدام تقدير الاحتمالية القصوى (MLE). يتم التحقق من أداء النموذج التنبؤي باستخدام مقاييس مثل المعايرة، والتمييز (إحصائية C)، والدقة التنبؤية العامة (درجة Brier). تقيم المعايرة التوافق بين الاحتمالات المرصودة والمتوقعة، بينما تقيم إحصائية C قدرة النموذج على التمييز بين الأحداث وغير الأحداث.

لمعالجة الإفراط المحتمل في التكيف، خاصة في السيناريوهات ذات أحجام العينات الصغيرة بالنسبة لعدد المعلمات، تناقش الورقة طرق الانكماش الموحد وتقنيات الانحدار المعاقبة. ينطبق الانكماش الموحد على جميع تقديرات المعاملات بعد MLE لتقليل حجمها، وغالبًا ما يتم تقديره من خلال طرق إعادة أخذ العينات. يتضمن الانحدار المعاقب، بما في ذلك طرق Ridge (L2) و Lasso (L1)، فرض عقوبات على دالة اللوغ-احتمالية لتقليل المعاملات أثناء التقدير، مع كون معامل الضبط \(\lambda\) حاسمًا لتحقيق التوازن بين الانكماش وملاءمة النموذج. تشير الورقة أيضًا إلى طريقة Firth، التي تعالج مشاكل الفصل في الانحدار اللوجستي دون الحاجة إلى معامل ضبط. يتم عادةً تحسين ضبط الطرق المعاقبة من خلال التحقق المتقاطع، حيث يتم استخدام حزمة R “glmnet” للتنفيذ، ويتم استكشاف التباينات في طيات التحقق المتقاطع لتقييم تأثيرها على أداء النموذج.

نتائج

في هذا القسم، يقدم المؤلفون نتائج المحاكاة لنموذج يتميز بإحصائية C حقيقية تبلغ 0.7 وانتشار حقيقي يبلغ 0.5، مع تضمين 5 متنبئين حقيقيين و 7 متنبئين ضوضاء. تم اشتقاق معاملات الانحدار باستخدام معلمة أساسية قدرها $\beta_0 = 0$ وعامل قياس قدره $k = 0.93$. كانت حجم العينة الموصى بها للتحليل هو $N = 900$. تتوفر نتائج إضافية لسيناريوهات بديلة في المواد التكميلية 1، مما يوفر مزيدًا من الرؤى حول أداء النموذج تحت ظروف متغيرة.

مناقشة

في هذا القسم، يناقش المؤلفون قيود خوارزمية الضبط القياسية لانحدار Ridge و Lasso، خاصة فيما يتعلق باختيار معامل الضبط $\lambda$ باستخدام التحقق المتقاطع ذو k طيات. يبرزون مشكلتين رئيسيتين: التباين في $\lambda$ الأمثل بسبب تقسيم البيانات العشوائي والإفراط في تقدير $\lambda$ عند استخدام مجموعات تدريب تحقق متقاطع أصغر. لمعالجة هذه المخاوف، يقترح المؤلفون خوارزمية ضبط معدلة تستخدم مجموعة بيانات زائفة أكبر تم إنشاؤها من خلال أخذ عينات مع الاستبدال من مجموعة البيانات الأصلية، مما يضمن أن تكون مجموعات التدريب للتحقق المتقاطع بنفس حجم مجموعة البيانات الأصلية. تهدف هذه التعديلات إلى تقليل الانكماش المفرط وتحسين استقرار اختيار $\lambda$.

كما يقدم المؤلفون نهجًا بديلًا يسمى “ضبط إعادة أخذ العينات”، والذي يتضمن ملاءمة النماذج على عينات إعادة أخذ العينات وتقييم الأداء خارج العينة على مجموعة البيانات الأصلية. تُظهر كل من طرق الضبط المعدلة وإعادة أخذ العينات أنها تؤدي إلى معاملات ضبط أصغر وأقل تباينًا مقارنةً بالضبط القياسي، مما يؤدي إلى تحسين منحدرات المعايرة (CS) ومقاييس الأداء التنبؤي. تُظهر دراسات المحاكاة أن هذه الطرق تتفوق على تقدير الاحتمالية القصوى (MLE) من حيث المعايرة والتباين، خاصة عند أحجام العينات الموصى بها. تشير النتائج إلى أن استخدام الضبط المعدل أو ضبط إعادة أخذ العينات يمكن أن يعزز استقرار النموذج والمعايرة في الانحدار المعاقب، مما يجعلها مفضلة في الممارسة، خاصة عندما تكون أحجام العينات محدودة.

Journal: Biometrical Journal, Volume: 66, Issue: 5
DOI: https://doi.org/10.1002/bimj.202300245
PMID: https://pubmed.ncbi.nlm.nih.gov/38922968
Publication Date: 2024-06-24
Author(s): Menelaos Pavlou et al.
Primary Topic: Statistical Methods and Inference

Overview

In this section, the authors address the issue of overfitting in risk prediction models that utilize maximum likelihood estimation (MLE), which often leads to extreme predictions and a calibration slope (CS) of less than 1. To mitigate this, penalized methods like Ridge and Lasso are employed, as they shrink regression coefficients towards zero, yielding predictions that are more aligned with the average. However, the tuning parameter, $\lambda$, selected through standard cross-validation can lead to over-shrinkage and variability, particularly problematic in small sample sizes.

The authors propose a novel approach termed “modified tuning,” which involves estimating $\lambda$ from a larger pseudo-development dataset generated through bootstrapping the original dataset. This method ensures that the cross-validation training datasets are comparable in size to the original dataset, thereby reducing the risk of over-estimating $\lambda$. The evaluation of modified and bootstrap tuning methods for Ridge and Lasso on both simulated and real datasets demonstrates significant improvements in the selection of $\lambda$ and CS, as well as enhanced prediction accuracy compared to traditional MLE and standard tuning methods.

Introduction

The introduction of this research paper discusses the critical role of risk prediction models in clinical decision-making, particularly those derived from logistic regression for binary outcomes. It highlights the necessity of validating these models using external datasets to ensure their predictive accuracy, which is assessed through various metrics such as calibration slope (CS), C-statistic, Brier Score, and net benefit. The paper identifies a common issue of overfitting in logistic regression models, especially when the sample size is small relative to the number of parameters, leading to extreme predictions for new patients. Penalized logistic regression methods, such as Ridge and Lasso, are introduced as solutions to mitigate overfitting by shrinking regression coefficients, although they may also introduce variability in the calibration measures.

The authors summarize previous studies that have explored the effectiveness of penalized methods in low-dimensional settings and their performance across different scenarios of event prevalence and predictor correlations. They note that while penalized methods generally reduce overfitting, they can sometimes over-shrink and exhibit high variability in tuning parameter estimation, potentially leading to miscalibrated models. The paper proposes a modified cross-validation tuning method, akin to bootstrap tuning, to enhance the performance of penalized methods in larger datasets. Subsequent sections will detail standard penalized regression techniques, the proposed tuning methods, and a comparative analysis of predictive performance using simulations and a real cardiac dataset. The authors aim to demonstrate the practical utility of their modified tuning approach while addressing its limitations.

Methods

The methods section of the research paper outlines the statistical approaches employed for risk prediction in binary outcomes, specifically using logistic regression models. The binary outcome, denoted as \(Y_i\), is modeled with covariates represented by the vector \(\mathbf{x}_i\). The logistic regression framework estimates the probability of the event of interest through the logit function, where the coefficients \(\boldsymbol{\beta}\) are typically estimated using maximum likelihood estimation (MLE). The predictive performance of the model is validated using metrics such as calibration, discrimination (C-statistic), and overall predictive accuracy (Brier Score). Calibration assesses the alignment between observed and predicted probabilities, while the C-statistic evaluates the model’s ability to discriminate between events and non-events.

To address potential overfitting, particularly in scenarios with small sample sizes relative to the number of parameters, the paper discusses uniform shrinkage methods and penalized regression techniques. Uniform shrinkage applies a factor to all coefficient estimates post-MLE to reduce their magnitude, often estimated via bootstrap methods. Penalized regression, including Ridge (L2) and Lasso (L1) methods, incorporates penalties into the log-likelihood function to shrink coefficients during estimation, with the tuning parameter \(\lambda\) crucial for balancing shrinkage and model fit. The paper also mentions Firth’s method, which addresses separation issues in logistic regression without requiring a tuning parameter. The tuning of penalized methods is typically optimized through cross-validation, with the R package “glmnet” being utilized for implementation, and variations in cross-validation folds are explored to assess their impact on model performance.

Results

In this section, the authors report simulation results for a model characterized by a true C-statistic of 0.7 and a true prevalence of 0.5, incorporating 5 true predictors and 7 noise predictors. The regression coefficients were derived using a baseline parameter of $\beta_0 = 0$ and a scaling factor of $k = 0.93$. The recommended sample size for the analysis was $N = 900$. Additional results for alternative scenarios are available in Supplementary Material 1, providing further insights into the model’s performance under varying conditions.

Discussion

In this section, the authors discuss the limitations of the standard tuning algorithm for Ridge and Lasso regression, particularly regarding the selection of the tuning parameter $\lambda$ using k-fold cross-validation. They highlight two main issues: the variability in the optimal $\lambda$ due to random data splits and the overestimation of $\lambda$ when using smaller cross-validation training sets. To address these concerns, the authors propose a modified tuning algorithm that utilizes a larger pseudo-dataset created by sampling with replacement from the original dataset, ensuring that the cross-validation training sets are the same size as the original dataset. This modification aims to reduce excessive shrinkage and improve the stability of $\lambda$ selection.

The authors also introduce an alternative approach called “Bootstrap tuning,” which involves fitting models on bootstrap samples and assessing out-of-sample performance on the original dataset. Both the modified and bootstrap tuning methods are shown to yield smaller and less variable tuning parameters compared to standard tuning, leading to improved calibration slopes (CS) and predictive performance metrics. Simulation studies demonstrate that these methods outperform maximum likelihood estimation (MLE) in terms of calibration and variability, particularly at recommended sample sizes. The findings suggest that using modified or bootstrap tuning can enhance model stability and calibration in penalized regression, making them preferable in practice, especially when sample sizes are limited.

شارك: