DOI: https://doi.org/10.1186/s41512-026-00222-1
PMID: https://pubmed.ncbi.nlm.nih.gov/42152151
تاريخ النشر: 2026-05-18
المؤلف: Alexander Pate وآخرون
الموضوع الرئيسي: طرق إحصائية واستدلال
نظرة عامة
في هذه الدراسة، يحقق المؤلفون في تحيز عامل الانكماش الاستدلالي المقترح من قبل فان هويلينجن ولي سيسي ($S_{VH}$) مقارنةً بعامل الانكماش الأمثل ($S_{opt}$) وعامل الانكماش المستمد من البوتستراب ($S_{boot}$) في سياق نماذج الانحدار اللوجستي غير المعاقب لتوقع المخاطر. تتضمن الأبحاث دراستين محاكاة: الأولى تقيم تحيز $S_{VH}$ و$S_{boot}$ كموحدات لـ $S_{opt}$ عبر سيناريوهات مختلفة محددة بواسطة إحصائية C ($C_{pop}$)، بينما الثانية تقيم تحيز $S_{opt}$ عند استخدام أحجام عينات مختلفة تستهدف انكماشًا قدره 0.9.
تكشف النتائج أن $S_{VH}$ يميل إلى تقدير $S_{opt}$ بشكل زائد للقيم العالية لـ $C_{pop}$ وتقديره بشكل ناقص للقيم المنخفضة. على وجه التحديد، عندما يكون $0.8 \leq C_{pop} < 0.85$، تراوح التحيز من 0.005 إلى 0.387، بينما بالنسبة لـ $0.6 \leq C_{pop} < 0.65$، تراوح من -0.580 إلى -0.007. تضاءل التحيز مع زيادة أحجام العينات واقترب $S_{opt}$ من 1. بالمقابل، أظهر $S_{boot}$ تحيزًا أقل بشكل عام، مع تحيز وسطي قدره 0.007 مقارنةً بـ 0.032 لـ $S_{VH}$. تشير النتائج إلى أن $S_{VH}$ غالبًا ما يكون موحدًا غير كافٍ لعامل الانكماش العالمي الأمثل، ويوصي المؤلفون باستخدام تقدير البوتستراب للحصول على نتائج أكثر موثوقية، خاصة عندما تكون أحجام العينات أكبر. بالإضافة إلى ذلك، يُفضل حسابات حجم العينة المعتمدة على طرق المحاكاة على تلك التي تعتمد على $S_{VH}$.
مقدمة
تناقش مقدمة هذه الورقة البحثية تطوير وتقييم نماذج التنبؤ السريرية، والتي تعتبر ضرورية لتقدير مخاطر الأحداث المستقبلية أو تشخيص الحالات الموجودة. أحد التحديات الكبيرة في هذا السياق هو مشكلة الإفراط في التكيف، حيث تؤدي النماذج أداءً ضعيفًا على بيانات جديدة بسبب التنبؤات المتطرفة. لمعالجة ذلك، يتم استخدام تقنيات مثل العقوبة والانكماش، مع كون الممارسة الشائعة هي تطبيق عامل انكماش عالمي، يُشار إليه بـ $S$، لضبط القيم المتوقعة نحو المتوسط. يعتبر عامل الانكماش الأمثل، $S_{\text{opt}}$، حاسمًا لتقليل متوسط مربع خطأ التنبؤات في السكان المستهدفين، ومع ذلك يبقى غير معروف أثناء تطوير النموذج ويجب تقديره.
سلطت الدراسات السابقة الضوء على أوجه القصور في الموحدات الحالية لـ $S_{\text{opt}}$، مشيرةً إلى تباين عالٍ في أدائها، خاصةً مع أحجام عينات صغيرة. بينما تحسن موحدات الانكماش عمومًا المعايرة في البيانات الجديدة، فإنها غالبًا لا تعطي تقديرات دقيقة لـ $S_{\text{opt}}$. تهدف الورقة إلى مقارنة صارمة بين موحدتين محددتين: عامل الانكماش الاستدلالي، $S_{VH}$، وطريقة البوتستراب، $S_{\text{boot}}$، لتحديد فعاليتها كموحدات لـ $S_{\text{opt}}$. بالإضافة إلى ذلك، تسعى لتقييم آثار أي تحيزات في $S_{VH}$ على معايير حجم العينة لتطوير النموذج. من المتوقع أن تعزز النتائج موثوقية موحدات الانكماش وتطبيقها في نمذجة التنبؤ السريرية.
طرق
في هذا القسم، يحدد المؤلفون المنهجيات المستخدمة لتقدير المعامل المحدد باستخدام تقنيات S V H و S boot، كما تم تعريفها سابقًا. التركيز الأساسي هو تقييم ما إذا كان الميل المتوسط للمعايرة لنماذج التنبؤ، التي تم تطويرها من مجموعات بيانات بأحجام \(N_{\text{original}}\) و \(N_{\text{sim}}\)، يفي بالحد المستهدف البالغ 0.9. المقياس المستخدم لهذا التقييم هو المتوسط \(S_{\text{opt},k}\) عبر 250 تكرار محاكاة لكل سيناريو \(k\).
لتحديد الأداء، يبلغ المؤلفون عن نطاق النسبة المئوية للمتوسط \(S_{\text{opt},k}\) عبر جميع السيناريوهات وداخل المجموعات الفرعية المحددة بواسطة \(C_{\text{pop}}\). بالإضافة إلى ذلك، يحسبون التحيز لكل سيناريو \(k\) كمتوسط الفرق بين \(S_{\text{opt}}\) ومستوى الانكماش المطلوب، معبرًا عنه كـ \( \text{bias}_k = \frac{1}{250} \sum_{m=1}^{250} (S_{\text{opt},k,m} – 0.9) \). تشمل التحليلات الإبلاغ عن متوسط التحيز العام وحجمه، بالإضافة إلى تصنيف هذه النتائج حسب \(C_{\text{pop}}\) لتوفير فهم شامل لأداء النموذج عبر سياقات مختلفة.
نتائج
في قسم النتائج، يقدم المؤلفون نتائج من محاكاة تشمل 15,000 سيناريو، مع التركيز على أداء موحدات مختلفة تحت ظروف التباين غير الصفري بين المتنبئين. لاحظوا أن السيناريوهات المتطرفة، التي تتميز بالضوضاء العالية وعدم كفاية المتنبئين، أدت إلى تقديرات سلبية أو فشل في التقارب أثناء البوتستراب، والتي تم استبعادها لاحقًا من التحليل. كانت المقياس الرئيسي للاهتمام هو إحصائية C، حيث يتم الإبلاغ عنها على نطاق واسع في الأدبيات وتم تحديدها كمتنبئ مهم لأداء النموذج. بشكل ملحوظ، وُجد أن متوسط التحيز للموحد $S_{VH}$ هو 0.077، مما يشير إلى أنه غالبًا ما يقدر بشكل ضعيف عامل الانكماش الأمثل $S_{opt}$، خاصةً تحت ظروف متغيرة لإحصائية C السكانية ($C_{pop}$).
كشفت التحليلات أيضًا أنه مع زيادة حجم العينة ($N$)، تضاءل التحيز بين $S_{VH}$ و $S_{opt}$، ومع ذلك استمرت التنبؤات الزائدة والناقصة عبر نطاقات مختلفة من $R^2_{CS\_pop}$ و $C_{pop}$. أكد المؤلفون أنه بينما أظهر البوتستراب أداءً متوسطًا قويًا، إلا أنه أظهر عدم استقرار في الدراسات الفردية، كما يتضح من التباين في الموحدات $S_{boot}$ و $S_{opt}$. تشير النتائج إلى أن تحقيق حجم عينة يفي بالمعايير المقترحة من قبل رايلي وآخرين وبافلو وآخرين أمر حاسم، لكن يجب على الباحثين السعي للحصول على عينات أكبر لضمان أداء موثوق للنموذج. تؤكد النتائج على أهمية فهم التباين الكامن في الموحدات وضرورة أحجام عينات قوية في نمذجة التنبؤ.
مناقشة
في هذا القسم، يناقش المؤلفون تطوير نموذج انحدار لوجستي لتوقع النتائج باستخدام مجموعة بيانات تحتوي على متغير النتيجة \(Y_i\) ومتغيرات المتنبئ \(X_i = (X_{1i}, X_{2i}, \ldots, X_{Qi})\). يتضمن النموذج عامل انكماش عالمي \(S\) يُطبق على المعاملات المقدرة لتحسين المعايرة. يُعرف عامل الانكماش الأمثل \(S_{\text{opt}}\) بأنه القيمة التي تؤدي إلى ميل معايرة قدره 1 عند التحقق من صحة النموذج ضد مجموعة سكانية تمثيلية. يقترح المؤلفون طرقًا لتقدير \(S_{\text{opt}}\) عندما تكون مجموعة التحقق الكبيرة غير متاحة، بما في ذلك عامل الانكماش الاستدلالي \(S_{VH}\) وعامل الانكماش المستمد من البوتستراب \(S_{\text{boot}}\).
تكشف دراسات المحاكاة التي أجريت أن \(S_{\text{boot}}\) يوفر عمومًا تقديرًا أكثر دقة لـ \(S_{\text{opt}}\) مقارنةً بـ \(S_{VH}\)، خاصةً عندما يكون \(S_{\text{opt}}\) الحقيقي مرتفعًا (≥ 0.7). التحيز لـ \(S_{\text{boot}}\) أصغر باستمرار عبر سيناريوهات مختلفة، بينما يميل \(S_{VH}\) إلى تقدير $S_{\text{opt}}$ بشكل ناقص في أحجام العينات الصغيرة أو عندما يكون عدد المتنبئين مرتفعًا. تؤكد النتائج على أهمية استخدام حسابات حجم العينة المناسبة وطرق تقدير الانكماش للتخفيف من الإفراط في التكيف في نماذج الانحدار اللوجستي، خاصةً في سياقات التنبؤ السريري.
DOI: https://doi.org/10.1186/s41512-026-00222-1
PMID: https://pubmed.ncbi.nlm.nih.gov/42152151
Publication Date: 2026-05-18
Author(s): Alexander Pate et al.
Primary Topic: Statistical Methods and Inference
Overview
In this study, the authors investigate the bias of the heuristic shrinkage factor proposed by Van Houwelingen and Le Cessie ($S_{VH}$) in comparison to the optimal shrinkage factor ($S_{opt}$) and an alternative bootstrap-derived shrinkage factor ($S_{boot}$) in the context of unpenalized logistic regression models for risk prediction. The research comprises two simulation studies: the first assesses the bias of $S_{VH}$ and $S_{boot}$ as estimators of $S_{opt}$ across various scenarios defined by the C-statistic ($C_{pop}$), while the second evaluates the bias of $S_{opt}$ when using different sample sizes targeting a shrinkage of 0.9.
The results reveal that $S_{VH}$ tends to overestimate $S_{opt}$ for high $C_{pop}$ values and underestimate it for low values. Specifically, when $0.8 \leq C_{pop} < 0.85$, the bias ranged from 0.005 to 0.387, whereas for $0.6 \leq C_{pop} < 0.65$, it ranged from -0.580 to -0.007. The bias diminished as sample sizes increased and $S_{opt}$ approached 1. In contrast, $S_{boot}$ exhibited less bias overall, with a median bias of 0.007 compared to 0.032 for $S_{VH}$. The findings suggest that $S_{VH}$ is often an inadequate estimator of the optimal global shrinkage factor, and the authors recommend using the bootstrap estimate for more reliable results, particularly when sample sizes are larger. Additionally, sample size calculations based on simulation methods are preferred over those relying on $S_{VH}$.
Introduction
The introduction of this research paper discusses the development and evaluation of clinical prediction models, which are essential for estimating the risk of future events or diagnosing existing conditions. A significant challenge in this context is the issue of overfitting, where models perform poorly on new data due to extreme predictions. To address this, techniques such as penalization and shrinkage are employed, with a common practice being the application of a global shrinkage factor, denoted as $S$, to adjust predicted values towards the mean. The optimal shrinkage factor, $S_{\text{opt}}$, is crucial for minimizing the mean-square error of predictions in the target population, yet it remains unknown during model development and must be estimated.
Previous studies have highlighted the shortcomings of existing estimators for $S_{\text{opt}}$, indicating high variability in their performance, particularly with small sample sizes. While shrinkage estimators generally improve calibration in new data, they often do not yield accurate estimates of $S_{\text{opt}}$. The paper aims to rigorously compare two specific estimators: the heuristic shrinkage factor, $S_{VH}$, and a bootstrapping method, $S_{\text{boot}}$, to determine their effectiveness as estimators for $S_{\text{opt}}$. Additionally, it seeks to assess the implications of any biases in $S_{VH}$ on sample size criteria for model development. The findings are expected to enhance the reliability of shrinkage estimators and their application in clinical prediction modeling.
Methods
In this section, the authors outline the methodologies employed to estimate the specified estimand using the techniques S V H and S boot, as previously defined. The primary focus is on evaluating whether the mean calibration slope of predictive models, developed from datasets of sizes \(N_{\text{original}}\) and \(N_{\text{sim}}\), meets the target threshold of 0.9. The performance metric utilized for this assessment is the mean \(S_{\text{opt},k}\) across 250 simulation iterations for each scenario \(k\).
To quantify the performance, the authors report the percentile range of the mean \(S_{\text{opt},k}\) across all scenarios and within subgroups defined by \(C_{\text{pop}}\). Additionally, they calculate the bias for each scenario \(k\) as the mean difference between \(S_{\text{opt}}\) and the desired level of shrinkage, expressed as \( \text{bias}_k = \frac{1}{250} \sum_{m=1}^{250} (S_{\text{opt},k,m} – 0.9) \). The analysis includes reporting the overall mean bias and its magnitude, as well as stratifying these results by \(C_{\text{pop}}\) to provide a comprehensive understanding of the model performance across different contexts.
Results
In the results section, the authors present findings from simulations involving 15,000 scenarios, focusing on the performance of various estimators under conditions of non-zero covariance among predictors. They observed that extreme scenarios, characterized by high noise and inadequate predictors, led to negative estimates or convergence failures during bootstrapping, which were subsequently excluded from analysis. The primary metric of interest was the C-statistic, as it is widely reported in literature and was identified as a significant predictor of model performance. Notably, the mean bias of the estimator $S_{VH}$ was found to be 0.077, indicating that it often poorly estimated the optimal shrinkage factor $S_{opt}$, particularly under varying conditions of the population C-statistic ($C_{pop}$).
The analysis further revealed that as sample size ($N$) increased, the bias between $S_{VH}$ and $S_{opt}$ diminished, yet over- and under-predictions persisted across different ranges of $R^2_{CS\_pop}$ and $C_{pop}$. The authors emphasized that while bootstrapping demonstrated strong average performance, it exhibited instability in individual studies, as evidenced by variability in the estimators $S_{boot}$ and $S_{opt}$. The findings suggest that achieving a sample size that meets the criteria proposed by Riley et al. and Pavlou et al. is crucial, but researchers should aim for larger samples to ensure reliable model performance. The results underscore the importance of understanding the variability inherent in estimators and the necessity of robust sample sizes in predictive modeling.
Discussion
In this section, the authors discuss the development of a logistic regression model for predicting outcomes using a dataset with outcome variable \(Y_i\) and predictor variables \(X_i = (X_{1i}, X_{2i}, \ldots, X_{Qi})\). The model incorporates a global shrinkage factor \(S\) applied to the estimated coefficients to improve calibration. The optimal shrinkage factor \(S_{\text{opt}}\) is defined as the value that results in a calibration slope of 1 when validated against a representative population. The authors propose methods for estimating \(S_{\text{opt}}\) when a large validation cohort is unavailable, including the heuristic shrinkage factor \(S_{VH}\) and a bootstrap-derived shrinkage factor \(S_{\text{boot}}\).
The simulation studies conducted reveal that \(S_{\text{boot}}\) generally provides a more accurate estimate of \(S_{\text{opt}}\) compared to \(S_{VH}\), particularly when the true \(S_{\text{opt}}\) is high (≥ 0.7). The bias of \(S_{\text{boot}}\) is consistently smaller across various scenarios, while \(S_{VH}\) tends to underestimate \(S_{\text{opt}}\) in smaller sample sizes or when the number of predictors is high. The findings emphasize the importance of using appropriate sample size calculations and shrinkage estimation methods to mitigate overfitting in logistic regression models, particularly in clinical prediction contexts.
