DOI: https://doi.org/10.1186/s12911-024-02525-z
PMID: https://pubmed.ncbi.nlm.nih.gov/38715002
تاريخ النشر: 2024-05-07
المؤلف: Naseem Asghar وآخرون
الموضوع الرئيسي: طرق إحصائية واستدلال
نظرة عامة
في هذه الدراسة، يتناول المؤلفون التحديات التي تطرحها بيانات المعلومات الحيوية عالية الأبعاد في تحليل البقاء، لا سيما في سياق بيانات الوقت حتى الحدث مثل الوقت حتى الفشل أو الموت. يقترحون نهجًا هجينًا جديدًا لاختيار الميزات يدمج بين أربع تقنيات لاختيار المتغيرات – LASSO، RSF-vs، SCAD، وCoxBoost – لتعزيز دقة توقعات البقاء. من خلال استخدام نماذج تحليل البقاء بما في ذلك نموذج مخاطر كوك (CoxPH)، وغابة البقاء العشوائية (RSF)، وشبكة الأعصاب DeepHit (DeepHit NN)، يوضح المؤلفون أن طريقتهم، التي تركز على المتغيرات التي يتم اختيارها باستمرار بواسطة تقنيات متعددة، تحسن بشكل كبير من الأداء التنبؤي.
تشير النتائج إلى أن الطريقة المقترحة تتفوق على تقنيات اختيار الميزات الحالية عبر مقاييس الأداء المختلفة، بما في ذلك درجة بريير المتكاملة (IBS)، ومؤشر التوافق (C-Index)، والخطأ المطلق المتكامل (IAE)، التي تم تقييمها على 11 مجموعة بيانات بقاء عالية الأبعاد. على وجه التحديد، حققت الطريقة المقترحة نتائج مثالية من خلال التفوق على طرق أخرى في 9 من أصل 11 مجموعة بيانات، وإظهار أداء متوسط متفوق عبر جميع المقاييس. تشير النتائج إلى أن دمج غابة البقاء العشوائية كنموذج تنبؤي يساهم في تحسين أداء نهج اختيار الميزات المقترح، مما يحسن في النهاية موثوقية توقعات البقاء في السياقات الطبية الحيوية.
مقدمة
تقدم مقدمة ورقة البحث لمحة عن مجال تحليل بيانات البقاء، الذي يركز على المدة حتى حدوث أحداث معينة، مثل الموت أو الفشل الميكانيكي. تشمل الأهداف الرئيسية لهذا التحليل تقدير احتمالات البقاء، وفهم تأثير عوامل الخطر المختلفة على البقاء، وتوقع أوقات البقاء بناءً على الخصائص ذات الصلة. تسلط الورقة الضوء على الأهمية المتزايدة لتقنيات اختيار الميزات (FS) في تحليل البقاء، لا سيما بسبب توفر مجموعات بيانات كبيرة وإمكانية ندرة الميزات ذات الصلة.
يؤكد المؤلفون أنه على الرغم من وجود العديد من خوارزميات FS، إلا أنه لا توجد حل شامل قابل للتطبيق، ولكل طريقة نقاط قوتها وضعفها. تتطلب هذه التباينات تحسينات مستمرة على الخوارزميات الحالية أو تطوير نهج جديدة لتقليل الخطأ في توقعات البقاء. لمعالجة هذه التحديات، يقترح المؤلفون خوارزمية FS هجينة مصممة خصيصًا لتحليل البقاء، والتي سيفصلونها في الأقسام التالية من الورقة.
طرق
في هذا القسم، يناقش المؤلفون طرق اختيار الميزات المختلفة المستخدمة لمعالجة التحديات التي تطرحها البيانات عالية الأبعاد، لا سيما لعنة الأبعاد، التي تنشأ عندما يتجاوز عدد الميزات عدد الملاحظات بشكل كبير. يمكن أن تؤدي هذه الحالة إلى زيادة التكاليف الحسابية، والضوضاء، والازدواجية، والتكيف المفرط، وسوء التعميم على البيانات غير المرئية. يصنف المؤلفون طرق اختيار الميزات إلى ثلاثة أنواع رئيسية: طرق الفلترة، وطرق التغليف، والطرق المدمجة.
تقوم طرق الفلترة بترتيب الميزات بناءً على مقاييس إحصائية قبل تدريب النموذج، مما يجعلها فعالة من حيث التكلفة وكخطوة معالجة مسبقة. من ناحية أخرى، تتضمن طرق التغليف إنشاء مجموعات فرعية من الميزات، وتدريب النماذج على كل مجموعة فرعية، وتقييم أدائها لتحديد أفضل تركيبة. تدمج الطرق المدمجة اختيار الميزات داخل خوارزمية التعلم نفسها، مما يوفر توازنًا بين الكفاءة والدقة. تستخدم الدراسة طرقًا مدمجة لاختيار المتغيرات، والتي تم تسليط الضوء عليها باعتبارها فعالة من حيث الوقت وفعالة مقارنة بالطرق الأخرى. كما يسرد القسم مجموعات بيانات البقاء عالية الأبعاد المستخدمة في الدراسة، موضحًا عدد العينات والميزات لكل مجموعة بيانات.
نتائج
في قسم النتائج، يتم تقديم تحليل لـ 11 مجموعة بيانات بقاء عالية الأبعاد من خلال أربع جداول، مع التركيز على أداء ثلاثة نماذج لتوقع البقاء: نموذج مخاطر كوك، وغابة البقاء العشوائية، وDeepHit. تم تقييم كل نموذج باستخدام خمس طرق لاختيار الميزات، بما في ذلك الطريقة المقترحة، وLASSO، وRSF-vs، وSCAD، وCoxBoost. تم حساب مقاييس الأداء – درجة بريير المتكاملة (IBS)، ومؤشر C، والخطأ المطلق المتكامل (IAE) – لكل مجموعة بيانات وطريقة اختيار ميزات. من الجدير بالذكر أن الطريقة المقترحة أظهرت أداءً متفوقًا عبر جميع المقاييس الثلاثة، متفوقة على البدائل في 8 من 11 مجموعة بيانات عند تقييمها من خلال معيار التصويت.
تسلط الجدول 2 الضوء على أن LASSO حقق أداءً تنبؤيًا أفضل باستمرار لمجموعات بيانات مثل “الثدي”، و”WPBC”، و”VDV”، بينما أظهرت مجموعة بيانات “Heart FD” نتائج مختلطة تفضل LASSO لـ IBS ومؤشر C ولكن CoxBoost لـ IAE. في الجدول 3، تفوقت الطريقة المقترحة مرة أخرى، متفوقة على جميع الطرق الأخرى عبر جميع المقاييس في 9 من 11 مجموعة بيانات. يؤكد الجدول 4 هذه النتائج، مشيرًا إلى أنه عند تطبيق الطريقة المقترحة مع DeepHit-NN، تحقق قيم IBS وIAE أقل في 8 مجموعات بيانات وقيم C-index أعلى في 9 مجموعات بيانات. أخيرًا، يجمع الجدول 5 مقاييس الأداء المتوسطة، مما يعزز تفوق الطريقة المقترحة عبر جميع النماذج والمقاييس، مدعومًا بتمثيلات رسومية في الأشكال التالية.
مناقشة
في قسم المناقشة، تستعرض الورقة منهجيات مختلفة لاختيار المتغيرات في تحليل البقاء، مع التركيز بشكل خاص على تطبيقها على مجموعات بيانات عالية الأبعاد. تسلط الضوء على تطور تقنيات مثل LASSO، التي تتضمن عقوبة L1-norm لتقليص المعاملات الأقل أهمية إلى الصفر، مما يبسط النماذج. يحسن Adaptive LASSO من ذلك من خلال تعديل الأوزان بناءً على أحجام المعاملات، بينما يجمع Elastic Net بين عقوبات L1 وL2 لمعالجة قيود LASSO، لا سيما في وجود متنبئين مترابطين. تشمل الطرق الملحوظة الأخرى استخدام الانحدارات الملساء وتعزيز التدرج لاختيار المتغيرات، بالإضافة إلى غابات البقاء العشوائية (RSF) وشبكات الأعصاب DeepHit، التي تستفيد من تقنيات التعلم الجماعي والتعلم العميق، على التوالي، لتعزيز دقة التنبؤ في تحليل البقاء.
تقترح الورقة أيضًا طريقة جديدة لاختيار الميزات الهجينة التي تدمج بين أربع تقنيات راسخة – LASSO، وRSF، وSCAD، وCoxBoost. تهدف هذه الطريقة إلى تحديد الميزات الأكثر معلوماتية من خلال اختيار تلك التي يتم اختيارها باستمرار عبر خوارزميات متعددة، مما يزيد من احتمال الاحتفاظ بالمتغيرات ذات الصلة وغير المكررة. يتم هيكلة النهج المقترح في ثلاث خطوات: تطبيق كل طريقة لاختيار الميزات بشكل فردي، ومقارنة النتائج لتشكيل مجموعة توافق من الميزات، وأخيرًا تصفية للاحتفاظ فقط بتلك الميزات التي تم اختيارها من قبل ثلاث طرق على الأقل. من المتوقع أن يحسن هذا النهج المنهجي الأداء التنبؤي في نماذج البقاء، لا سيما في السياقات عالية الأبعاد.
DOI: https://doi.org/10.1186/s12911-024-02525-z
PMID: https://pubmed.ncbi.nlm.nih.gov/38715002
Publication Date: 2024-05-07
Author(s): Naseem Asghar et al.
Primary Topic: Statistical Methods and Inference
Overview
In this study, the authors address the challenges posed by high-dimensional bioinformatics data in survival analysis, particularly in the context of time-to-event data such as time to failure or death. They propose a novel hybrid feature selection approach that integrates four variable selection techniques—LASSO, RSF-vs, SCAD, and CoxBoost—to enhance survival prediction accuracy. By employing survival analysis models including Cox Proportional Hazards (CoxPH), Random Survival Forest (RSF), and DeepHit Neural Network (DeepHit NN), the authors demonstrate that their method, which focuses on variables consistently selected by multiple techniques, significantly improves predictive performance.
The results indicate that the proposed method outperforms existing feature selection techniques across various performance metrics, including Integrated Brier Score (IBS), Concordance Index (C-Index), and Integrated Absolute Error (IAE), evaluated on 11 high-dimensional survival datasets. Specifically, the proposed method achieved ideal results by outperforming other methods in 9 out of 11 datasets and demonstrating superior average performance across all metrics. The findings suggest that the integration of Random Survival Forest as a predictive model contributes to the enhanced performance of the proposed feature selection approach, ultimately improving the reliability of survival predictions in biomedical contexts.
Introduction
The introduction to the research paper outlines the field of survival data analysis, which focuses on the duration until specific events occur, such as death or mechanical failure. The primary goals of this analysis include estimating survival probabilities, understanding the impact of various risk factors on survival, and predicting survival times based on relevant characteristics. The paper highlights the increasing importance of feature selection (FS) techniques in survival analysis, particularly due to the availability of large datasets and the potential sparsity of relevant features.
The authors emphasize that while numerous FS algorithms exist, there is no universally applicable solution, and each method has its own strengths and weaknesses. This variability necessitates ongoing improvements to existing algorithms or the development of new approaches to minimize error in survival predictions. To address these challenges, the authors propose a hybrid FS algorithm tailored for survival analysis, which they will detail in the subsequent sections of the paper.
Methods
In this section, the authors discuss various feature selection methods employed to address the challenges posed by high-dimensional data, particularly the curse of dimensionality, which arises when the number of features significantly exceeds the number of observations. This situation can lead to increased computational costs, noise, redundancy, overfitting, and poor generalization on unseen data. The authors categorize feature selection methods into three main types: filter, wrapper, and embedded methods.
Filter methods rank features based on statistical measures prior to model training, making them efficient and cost-effective as a preprocessing step. Wrapper methods, on the other hand, involve creating subsets of features, training models on each subset, and evaluating their performance to identify the best combination. Embedded methods integrate feature selection within the learning algorithm itself, offering a balance of efficiency and accuracy. The study utilizes embedded methods for variable selection, which are highlighted as both time-efficient and effective in comparison to the other methods. The section also lists benchmark high-dimensional survival datasets used in the study, detailing the number of samples and features for each dataset.
Results
In the Results section, the analysis of 11 high-dimensional survival datasets is presented through four tables, focusing on the performance of three survival prediction models: the Cox Proportional Hazards model, Random Survival Forest, and DeepHit. Each model was evaluated using five feature selection methods, including the proposed method, LASSO, RSF-vs, SCAD, and CoxBoost. Performance metrics—Integrated Brier Score (IBS), C-index, and Integrated Absolute Error (IAE)—were computed for each dataset and feature selection method. Notably, the proposed method demonstrated superior performance across all three metrics, outperforming the alternatives in 8 out of 11 datasets when evaluated through a voting criterion.
Table 2 highlights that LASSO consistently yielded better predictive performance for datasets such as ‘Breast’, ‘WPBC’, and ‘VDV’, while the ‘Heart FD’ dataset showed mixed results favoring LASSO for IBS and C-index but CoxBoost for IAE. In Table 3, the proposed method again excelled, outperforming all other methods across all metrics in 9 out of 11 datasets. Table 4 further corroborates these findings, indicating that when the proposed method is applied with DeepHit-NN, it achieves lower IBS and IAE values in 8 datasets and higher C-index values in 9 datasets. Finally, Table 5 consolidates the average performance metrics, reinforcing the proposed method’s superiority across all models and metrics, supported by graphical representations in subsequent figures.
Discussion
In the discussion section, the paper reviews various methodologies for variable selection in survival analysis, particularly focusing on their application to high-dimensional datasets. It highlights the evolution of techniques such as LASSO, which incorporates an L1-norm penalty to shrink less significant coefficients to zero, thereby simplifying models. The Adaptive LASSO improves upon this by adjusting weights based on coefficient sizes, while Elastic Net combines L1 and L2 penalties to address LASSO’s limitations, particularly in the presence of correlated predictors. Other notable methods include the use of smoothing splines and gradient boosting for variable selection, as well as Random Survival Forests (RSF) and DeepHit neural networks, which leverage ensemble learning and deep learning techniques, respectively, to enhance predictive accuracy in survival analysis.
The paper also proposes a novel hybrid feature selection method that integrates four established techniques—LASSO, RSF, SCAD, and CoxBoost. This method aims to identify the most informative features by selecting those that are consistently chosen across multiple algorithms, thereby increasing the likelihood of retaining relevant and non-redundant variables. The proposed approach is structured in three steps: applying each feature selection method individually, comparing the results to form a consensus set of features, and finally filtering to retain only those features selected by at least three of the methods. This systematic approach is expected to improve predictive performance in survival models, particularly in high-dimensional contexts.
