مخططات المعايرة المرنة المجمعة للنتائج الثنائية باستخدام نمذجة التأثيرات العشوائية
Clustered flexible calibration plots for binary outcomes using random effects modeling

شارك:
المجلة: Research Synthesis Methods، المجلد: 17، العدد: 3
DOI: https://doi.org/10.1017/rsm.2025.10046
PMID: https://pubmed.ncbi.nlm.nih.gov/41992907
تاريخ النشر: 2025-12-29
المؤلف: Lasai Barreñada وآخرون
الموضوع الرئيسي: طرق إحصائية واستدلال بايزي

نظرة عامة

تقييم نماذج التنبؤ السريرية عبر مجموعات متعددة، مثل مراكز أو مجموعات بيانات مختلفة، أصبح شائعًا بشكل متزايد، مما يتطلب تقييمًا دقيقًا للتوافق – الاتفاق بين المخاطر المقدرة والنتائج الملاحظة. تقدم هذه الدراسة ثلاث منهجيات لأخذ التجميع في الاعتبار في تقييمات المعايرة: المعايرة الجماعية المجمعة (CG-C)، والمعايرة من خلال تحليل ميتا ذو مرحلتين (2MA-C)، والمعايرة باستخدام نموذج مختلط (MIX-C). تسهل هذه الأساليب إنشاء مخططات معايرة مرنة باستخدام نمذجة التأثيرات العشوائية وتوفر فترات الثقة (CIs) وفترات التنبؤ (PIs).

في دراسة حالة تتعلق بالتحقق الخارجي من نموذج يتنبأ بمخاطر الخباثة للأورام المبيضية عبر 2,489 مريضًا، قام المؤلفون بإجراء دراسات محاكاة وبيانات تركيبية لتقييم أداء الأساليب المقترحة. أظهرت النتائج أن MIX-C و2MA-C (مع الانحدارات) أنتجت منحنيات مقدرة تطابقت بشكل وثيق مع المنحنى العام الحقيقي في دراسة المحاكاة، بينما أنتجت MIX-C منحنيات خاصة بالتجمع كانت الأكثر دقة في دراسة البيانات التركيبية. أظهرت طريقة 2MA-C أفضل تغطية لفترة التنبؤ عبر مخطط المعايرة. يوصي المؤلفون باستخدام 2MA-C مع الانحدارات لتقدير المنحنى العام وفترة التنبؤ 95%، وMIX-C للمنحنيات الخاصة بالتجمع، خاصة عندما تكون أحجام العينات لكل تجمع محدودة. كما يقدمون كودًا سهل الوصول لبناء منحنيات معايرة مرنة ملخصة، بما في ذلك CIs وPIs، لتقييم تباين المعايرة عبر مجموعات البيانات أو المراكز.

مقدمة

تناقش مقدمة هذه الورقة البحثية نماذج التنبؤ السريرية (CPMs)، التي تعتبر أدوات أساسية لتقدير احتمالية الأحداث الصحية عند التشخيص أو التنبؤ. تعتمد موثوقية هذه النماذج على معايرتها، التي تقيس التوافق بين المخاطر المتوقعة والملاحظة. تسلط الورقة الضوء على اتجاه متزايد في استخدام بيانات متعددة المراكز (مجمعة) في دراسات CPM، كما يتضح من تحليل نماذج من Tufts PACE، حيث استخدم 64% مثل هذه البيانات. يؤكد المؤلفون على أهمية معالجة التجميع في تطوير النماذج وتقييمها، حيث يمكن أن تؤثر المعايرة الخاطئة سلبًا على اتخاذ القرارات الطبية.

تقدم الورقة مفهوم مخططات المعايرة كطريقة أساسية لتقييم أداء النموذج، موضحةً طريقتين: المعايرة المجمعة والنمذجة المرنة باستخدام تقنيات مثل الانحدار المحلي أو الانحدارات. بينما قد تفقد المعايرة المجمعة المعلومات بسبب التصنيف، يمكن أن توفر النماذج المرنة رؤى أكثر دقة ولكن تعتمد على معلمات التنعيم. يقترح المؤلفون ثلاث منهجيات لبناء منحنيات معايرة مرنة من بيانات مجمعة، بهدف توفير رؤى خاصة بالتجمع وقياس التباين من خلال فترات التنبؤ. يوضحون طرقهم من خلال دراسة حالة على بيانات سرطان المبيض ويحددون هيكل الورقة، التي تشمل دراسة محاكاة ومناقشة للنتائج. تمثل هذه الدراسة دراسة منهجية من المرحلة 2 تهدف إلى تعزيز تحليل المعايرة في سياق البيانات المجمعة.

الطرق

في هذه الدراسة، استخدم المؤلفون الانحدار اللوجستي مع تقاطعات عشوائية لتوليد بيانات تركيبية لتقييم نماذج التنبؤ المتعلقة بمخاطر خباثة الأورام المبيضية. تم تعريف هيكل النموذج بواسطة المعادلة $\text{logit}(\pi_{ij}) = \beta_0 + x_{ij} \beta_1 + u_j$، حيث $\beta_0$ هو التقاطع، و$\beta_1$ يمثل تأثير المتغير التابع $x$، و$u_j$ يدل على الانحراف الخاص بالتجمع. شمل عملية توليد البيانات تصميمًا كاملًا للعوامل التي تغيرت فيها كثافة التجميع (المقاسة بواسطة معامل الارتباط داخل الفئة، ICC) والمساحة تحت منحنى التشغيل الخاص بالمستقبل (AUC). تم محاكاة ما مجموعه 200 تجمع و2,000,000 مريض، مع تثبيت معدلات الأحداث عند 30% ولكنها تختلف عبر التجمعات.

لإنشاء مجموعات بيانات تركيبية، استخدم المؤلفون بيانات اتحاد تحليل الأورام المبيضية الدولية (IOTA)، حيث طبقوا حزمة synthpop في R للحفاظ على الهيكل الأساسي للبيانات مع إخفاء بيانات المرضى الفردية. تم تطوير نموذجين من الانحدار اللوجستي لكل مستشفى، مع تضمين الانحدارات للمتغيرات المستمرة ونهج الغابة العشوائية مع معلمات محددة. تم استخدام البيانات التركيبية، التي تمثل مليون فرد لكل مركز، بعد ذلك للتحقق من معايرة نموذج ADNEX عبر سيناريوهات مختلفة، بما في ذلك أعداد مختلفة من المراكز ومعدلات الأحداث لكل تجمع (EPC). شملت الدراسة 1,000 تكرار من أجل القوة، مع حساب متوسط خطأ المعايرة التربيعي (MSCE) لمقارنة النسب المقدرة الملاحظة الخاصة بالمركز ضد الاحتمالات الحقيقية، مما يبرز فعالية طريقة MIX-C في استعارة المعلومات عبر التجمعات.

النتائج

في قسم النتائج، تشير التحليلات البصرية إلى أن جميع أساليب المعايرة تحقق نتائج مشابهة لمثال الحالة، متوافقة مع النتائج السابقة التي تجاهلت التجميع. ومع ذلك، لوحظت اختلافات كبيرة في عدم اليقين المقدّر والتباين، حيث كانت ADNEX تقدر المخاطر بشكل منخفض بشكل ملحوظ. تكشف نتائج متوسط خطأ المعايرة التربيعي (MSCE) الوسيط، المعروضة في الجدول 2 والشكل 4، أن طريقة MIX-C تفوقت باستمرار على الآخرين عبر أحجام عينات التحقق المختلفة، بينما تفوقت طريقة 2MA-C (الانحدارات) في السيناريوهات ذات أحجام عينات التحقق الأكبر. كانت المعايرة اللوجستية المرنة القياسية ضعيفة الأداء، خاصة تحت ظروف معامل الارتباط داخل الفئة (ICC) العالية.

تشير التحليلات الإضافية إلى أن تغطية فترة التنبؤ (PI) اقتربت من المستويات الاسمية مع تغطية التنبؤ المتوقعة العالية (EPC) وICC المنخفض عند استخدام نهج 2MA-C (الانحدارات). ومع ذلك، كانت التغطية غير كافية عند أطراف التوزيع، خاصة بالنسبة للأساليب المجمعة عندما كانت EPC منخفضة. في تحليلات البيانات التركيبية، ظهرت طريقة MIX-C مرة أخرى كأفضل أداء عبر جميع السيناريوهات بناءً على الانحدار اللوجستي، بينما تفوقت LOESS مع أحجام عينات أكبر عندما كانت الحقيقة مستمدة من نموذج غابة عشوائية. تؤكد النتائج على أهمية مراعاة حجم العينة وتأثيرات التجميع، حيث اختلف الأداء بشكل كبير عبر مراكز مختلفة، مما يبرز توازن التحيز-التباين في تقدير المنحنيات الخاصة بالتجمع. الكود والبيانات لإعادة إنتاج الدراسة متاحة في مستودع OSF.

المناقشة

يركز قسم المناقشة في الورقة البحثية على معايرة نماذج التنبؤ السريرية (CPMs)، خاصة في سياق نموذج ADNEX لسرطان المبيض. يستخدم نموذج ADNEX، الذي يعتمد على نماذج الانحدار المختلط المتعدد الحدود، لتقدير احتمالية نتائج مختلفة تتعلق بالكتل المبيضية، مما يبرز أهمية التنبؤ بدقة بالخباثة قبل التدخل الجراحي. يقدم المؤلفون ثلاث طرق جديدة للحصول على مخططات معايرة مرنة تأخذ في الاعتبار التجميع في البيانات: المعايرة الجماعية المجمعة (CG-C)، والمعايرة من خلال تحليل ميتا ذو مرحلتين (2MA-C)، والمعايرة باستخدام نموذج مختلط (MIX-C). تعالج كل طريقة التحديات التي تطرحها البيانات المجمعة، حيث تمتد CG-C المعايرة المجمعة التقليدية، وتجمع 2MA-C بين المخططات الخاصة بالتجمع المرنة من خلال تحليل ميتا بتأثيرات عشوائية، وتستخدم MIX-C نموذج انحدار لوجستي مختلط خطي عام ذو مرحلة واحدة.

أجرى المؤلفون دراسة محاكاة واستخدموا بيانات تركيبية لتقييم أداء هذه الطرق المعايرة. وجدوا أن MIX-C و2MA-C (باستخدام الانحدارات) كانت فعالة بشكل خاص في تقدير مخططات المعايرة، خاصة للتجمعات ذات التأثيرات المتوسطة. ومع ذلك، أظهرت جميع الطرق احتمالات تغطية دون المستوى الأمثل لفترات التنبؤ (PIs)، حيث كانت 2MA-C (الانحدارات) الأفضل أداءً بشكل عام. تسلط الورقة الضوء على أهمية التمييز بين معايرة التأثير المتوسط والمعايرة الخاصة بالتجمع، محذرة من التعميم المفرط لأداء النموذج بناءً فقط على منحنيات المعايرة المتوسطة. كما يشير المؤلفون إلى أنه بينما توفر منهجياتهم رؤى قيمة حول تباين المعايرة، هناك حاجة إلى مزيد من البحث لتحسين دقة PIs، خاصة في السيناريوهات ذات أحجام العينات المحدودة. تم جعل المنهجيات وتنفيذاتها متاحة من خلال وظائف R في مستودع OSF وحزمة CalibrationCurves.

Journal: Research Synthesis Methods, Volume: 17, Issue: 3
DOI: https://doi.org/10.1017/rsm.2025.10046
PMID: https://pubmed.ncbi.nlm.nih.gov/41992907
Publication Date: 2025-12-29
Author(s): Lasai Barreñada et al.
Primary Topic: Statistical Methods and Bayesian Inference

Overview

The evaluation of clinical prediction models across multiple clusters, such as different centers or datasets, is increasingly prevalent, necessitating a thorough assessment of calibration—the agreement between estimated risks and observed outcomes. This study introduces three methodologies to account for clustering in calibration evaluations: clustered group calibration (CG-C), two-stage meta-analysis calibration (2MA-C), and mixed model calibration (MIX-C). These approaches facilitate the generation of flexible calibration plots using random effects modeling and provide confidence intervals (CIs) and prediction intervals (PIs).

In a case study involving the external validation of a model predicting the malignancy risk of ovarian tumors across 2,489 patients, the authors conducted both simulation and synthetic data studies to assess the performance of the proposed methods. Results indicated that MIX-C and 2MA-C (with splines) produced estimated curves that closely matched the true overall curve in the simulation study, while MIX-C yielded cluster-specific curves that were most accurate in the synthetic data study. The 2MA-C method demonstrated the best coverage of the PI across the calibration plot. The authors recommend employing 2MA-C with splines for estimating the overall curve and 95% PI, and MIX-C for cluster-specific curves, particularly when sample sizes per cluster are limited. They also provide accessible code for constructing summary flexible calibration curves, including CIs and PIs, to evaluate calibration heterogeneity across datasets or centers.

Introduction

The introduction of this research paper discusses Clinical Prediction Models (CPMs), which are essential tools for estimating the likelihood of health-related events at diagnosis or prognosis. The reliability of these models hinges on their calibration, which measures the alignment between predicted and observed risks. The paper highlights a growing trend in the use of multicenter (clustered) data in CPM studies, evidenced by an analysis of models from Tufts PACE, where 64% utilized such data. The authors emphasize the importance of addressing clustering in model development and evaluation, as miscalibration can adversely affect medical decision-making.

The paper introduces the concept of calibration plots as a primary method for assessing model performance, detailing two approaches: grouped calibration and flexible modeling using techniques like local regression or splines. While grouped calibration may lose information due to categorization, flexible models can provide more nuanced insights but depend on smoothing parameters. The authors propose three methodologies for constructing flexible calibration curves from clustered data, aiming to provide cluster-specific insights and quantify heterogeneity through prediction intervals. They illustrate their methods with a case study on ovarian cancer data and outline the structure of the paper, which includes a simulation study and a discussion of findings. This work represents a phase 2 methodological study aimed at enhancing calibration analysis in the context of clustered data.

Methods

In this study, the authors employed logistic regression with random intercepts to generate synthetic data for evaluating prediction models related to ovarian tumor malignancy risk. The model’s structure was defined by the equation $\text{logit}(\pi_{ij}) = \beta_0 + x_{ij} \beta_1 + u_j$, where $\beta_0$ is the intercept, $\beta_1$ represents the effect of covariate $x$, and $u_j$ denotes the cluster-specific deviation. The data generation process involved a full factorial design that varied clustering intensity (measured by the intraclass correlation coefficient, ICC) and the area under the receiver operating characteristic curve (AUC). A total of 200 clusters and 2,000,000 patients were simulated, with event rates fixed at 30% but varying across clusters.

To create synthetic datasets, the authors utilized the International Ovarian Tumor Analysis (IOTA) consortium data, applying the synthpop package in R to maintain the underlying data structure while masking individual patient data. Two logistic regression models were developed per hospital, incorporating splines for continuous variables and a random forest approach with specified parameters. The synthetic data, representing 1 million individuals per center, was then used to validate the calibration of the ADNEX model across various scenarios, including different numbers of centers and event per cluster (EPC) rates. The study involved 1,000 repetitions for robustness, calculating mean squared calibration error (MSCE) to compare center-specific estimated observed proportions against true probabilities, highlighting the effectiveness of the MIX-C method in borrowing information across clusters.

Results

In the results section, a visual analysis indicates that all calibration approaches yield similar outcomes for the case example, consistent with prior findings that disregarded clustering. However, significant differences in estimated uncertainty and heterogeneity were observed, with ADNEX notably underestimating risks. The median Mean Squared Calibration Error (MSCE) results, presented in Table 2 and Figure 4, reveal that the MIX-C method consistently outperformed others across various validation sample sizes, while the 2MA-C (splines) method excelled in scenarios with larger validation samples. Standard flexible logistic calibration performed poorly, particularly under high Intraclass Correlation Coefficient (ICC) conditions.

Further analysis indicates that the prediction interval (PI) coverage approached nominal levels with high Expected Prediction Coverage (EPC) and low ICC when using the 2MA-C (splines) approach. However, coverage was inadequate at the tails of the distribution, especially for grouped approaches when EPC was low. In synthetic data analyses, the MIX-C method again emerged as the best performer across all scenarios based on logistic regression, while LOESS excelled with larger sample sizes when the truth was derived from a random forest model. The findings underscore the importance of considering sample size and clustering effects, as performance varied significantly across different centers, highlighting a bias-variance trade-off in estimating cluster-specific curves. Code and data for reproducing the study are accessible in the OSF repository.

Discussion

The discussion section of the research paper focuses on the calibration of clinical prediction models (CPMs), particularly in the context of the ADNEX model for ovarian cancer. The ADNEX model, which utilizes multinomial regression mixed models, estimates the probability of various outcomes related to ovarian masses, emphasizing the importance of accurately predicting malignancy prior to surgical intervention. The authors introduce three novel methods for obtaining flexible calibration plots that account for clustering in data: the clustered group calibration (CG-C), two-stage meta-analysis calibration (2MA-C), and mixed model calibration (MIX-C). Each method addresses the challenges posed by clustered data, with the CG-C extending traditional grouped calibration, the 2MA-C combining flexible cluster-specific plots through random effects meta-analysis, and the MIX-C employing a one-stage logistic generalized linear mixed model.

The authors conducted a simulation study and utilized synthetic data to evaluate the performance of these calibration methods. They found that the MIX-C and 2MA-C (using splines) were particularly effective in estimating calibration plots, especially for clusters with average effects. However, all methods exhibited suboptimal coverage probabilities for prediction intervals (PIs), with 2MA-C (splines) performing best overall. The paper highlights the importance of distinguishing between average effect calibration and cluster-specific calibration, cautioning against overgeneralizing model performance based solely on average calibration curves. The authors also note that while their methodologies provide valuable insights into calibration heterogeneity, further research is needed to enhance the accuracy of PIs, particularly in scenarios with limited sample sizes. The methodologies and their implementations are made accessible through R functions in the OSF repository and the CalibrationCurves package.

شارك: