تأثير التعدد الخطي على أداء خوارزميات التعلم الآلي: أمثلة حالة في نمذجة الرعاية الصحية
The Multicollinearity Effect on the Performance of Machine Learning Algorithms: Case Examples in Healthcare Modelling

شارك:
المجلة: Academic Platform Journal of Engineering and Smart Systems، المجلد: 12، العدد: 3
DOI: https://doi.org/10.21541/apjess.1371070
تاريخ النشر: 2024-09-25
المؤلف: Hasan Yıldırım
الموضوع الرئيسي: الذكاء الاصطناعي في الرعاية الصحية

نظرة عامة

تتناول هذه الدراسة قضية التعدد الخطي في مجموعات البيانات، وخاصة في سياق بيانات الرعاية الصحية، التي شهدت نمواً هائلاً بسبب التقدم التكنولوجي. تؤثر هذه المشكلة سلباً على أداء كل من الخوارزميات الإحصائية وخوارزميات التعلم الآلي. تقوم الدراسة بإجراء مقارنة شاملة بين نماذج إحصائية مختلفة—مثل الانحدار Ridge وLiu وLasso وElastic Net—وثمانية خوارزميات بارزة في التعلم الآلي، بما في ذلك CART وKNN وMLP وMARS وCubist وSVM وBagging وXGBoost. يتم إجراء التقييم باستخدام مجموعتي بيانات صحية (Body Fat وCancer) تعاني من التعدد الخطي، حيث يتم تقييم أداء النموذج من خلال طرق التحقق المتبادل بناءً على متوسط الخطأ التربيعي الجذري (RMSE) ومتوسط الخطأ المطلق (MAE) ومعايير R-squared.

تشير النتائج إلى أن النماذج الإحصائية، وخاصة الانحدار Liu، تتفوق بشكل كبير على نماذج التعلم الآلي في كل من سيناريوهات التدريب والاختبار، حيث تحقق تحسينات في الأداء تتراوح من 1.55% إلى 46.08% عبر مجموعات البيانات. على الرغم من كونها غير ممثلة بشكل كافٍ في الأدبيات المتعلقة بالتعلم الآلي، إلا أن الانحدار Liu يظهر فعاليته كأداة قوية للدراسات المعتمدة على الانحدار التي تتضمن التعدد الخطي. تسلط الدراسة الضوء على محدوديتين: ضرورة تحديد درجة التعدد الخطي بدقة وأهمية ضبط معلمات النموذج الإحصائي بشكل مثالي. يجب أن تتناول الأبحاث المستقبلية هذه المحدودات لتعزيز أداء النموذج بشكل أكبر، مما يعزز الاستنتاج بأن النماذج الإحصائية هي حلول فعالة لتحديات التعدد الخطي في التطبيقات الواقعية.

مقدمة

تسلط مقدمة هذه الورقة البحثية الضوء على أهمية النمذجة الرياضية في التنبؤ والاستدلال، خاصة في سياق مجموعات البيانات الكبيرة التي تحتوي على العديد من المتغيرات التفسيرية. مع تقدم تقنيات جمع البيانات، أصبحت ظاهرة التعدد الخطي—حيث يمكن أن تؤدي المتغيرات المرتبطة بشكل كبير إلى تشويه النماذج الإحصائية ونماذج التعلم الآلي—مشكلة حرجة غالباً ما يتم تجاهلها. يمكن أن تؤدي هذه الظاهرة إلى استنتاجات غير دقيقة وأداء تنبؤي متدني، مما يستدعي فحصاً أعمق لتداعياتها عبر مجالات مختلفة، بما في ذلك الصحة والتسويق والمالية.

تهدف الورقة إلى سد فجوة في الأدبيات من خلال مقارنة الطرق الإحصائية التقليدية، مثل الانحدار Liu والانحدار Ridge، مع الأساليب البارزة في التعلم الآلي في معالجة التعدد الخطي. يؤكد المؤلفون أن النماذج الإحصائية يمكن أن تتفوق على النماذج المعقدة في التعلم الآلي في سياقات معينة ويشددون على أهمية الاعتراف بتأثير التعدد الخطي على أداء النموذج. تم هيكلة الدراسة لمراجعة الأدبيات ذات الصلة أولاً، ثم مناقشة مشكلة التعدد الخطي وطرق التشخيص، تليها شرح مفصل للنماذج المستخدمة، وعملية النمذجة، والإعدادات التجريبية، وأخيراً، نتائج المعايير والاستنتاجات.

مناقشة

في مناقشة التعدد الخطي، تسلط الورقة الضوء على المساهمات الأساسية في هذا المجال، بدءًا من مقدر ستاين الذي اقترحه جيمس وستاين، والذي قدم تقليصاً متساوياً للمعاملات في الانحدار العادي (OLS). ومع ذلك، كانت هذه الطريقة محدودة بسبب عدم قدرتها على التعامل مع المعاملات ذات العلامات المعاكسة والحاجة إلى مقدرات بديلة. عالج مقدر Ridge الذي اقترحه هويرل وكينارد هذه القضايا من خلال دمج مصطلح عقوبة لتقليص المعاملات نحو الصفر. قام مقدر Liu بتحسين ذلك من خلال دمج خصائص Ridge وستاين، بينما قدمت طريقة Lasso لتقليص المعاملات والاختيار المتغيرات من خلال السماح بتقليص المعاملات إلى الصفر. لمعالجة قيود Lasso في الإعدادات عالية الأبعاد، اقترح زو وهستي الشبكة المرنة، التي تدمج كل من تقنيات Ridge وLasso.

تتناول هذه الفقرة أيضًا عواقب التعدد الخطي، بما في ذلك تضخم تباينات تقديرات المعاملات والانخفاض المحتمل في توقعات النموذج. يتم مناقشة طرق تشخيصية مختلفة لاكتشاف التعدد الخطي، مثل حساب مصفوفة الارتباط، وعامل تضخم التباين (VIF)، وتحليل القيم الذاتية. توفر هذه الطرق رؤى حول العلاقات بين المتغيرات التفسيرية، مع عتبات محددة تشير إلى التعدد الخطي الشديد. تختتم الورقة بتلخيص الحلول للتعدد الخطي، بما في ذلك جمع البيانات، وتقنيات الانحدار البديلة (مثل Ridge وLiu وLasso والشبكة المرنة)، وطرق المعالجة المسبقة مثل التمركز والتقييس. يركز البحث على تقييم هذه النماذج البديلة لمعالجة التعدد الخطي بفعالية في تحليل الانحدار.

Journal: Academic Platform Journal of Engineering and Smart Systems, Volume: 12, Issue: 3
DOI: https://doi.org/10.21541/apjess.1371070
Publication Date: 2024-09-25
Author(s): Hasan Yıldırım
Primary Topic: Artificial Intelligence in Healthcare

Overview

The research addresses the issue of multicollinearity in datasets, particularly in the context of healthcare data, which has seen exponential growth due to technological advancements. This problem adversely affects the performance of both statistical and machine learning algorithms. The study conducts a comprehensive comparison of various statistical models—such as Ridge, Liu, Lasso, and Elastic Net regression—and eight prominent machine learning algorithms, including CART, KNN, MLP, MARS, Cubist, SVM, Bagging, and XGBoost. The evaluation is performed using two healthcare datasets (Body Fat and Cancer) afflicted by multicollinearity, with model performance assessed through cross-validation methods based on root mean square error (RMSE), mean absolute error (MAE), and R-squared criteria.

The findings indicate that statistical models, particularly Liu regression, significantly outperform machine learning models in both training and testing scenarios, achieving performance improvements ranging from 1.55% to 46.08% across the datasets. Despite being underrepresented in machine learning literature, Liu regression demonstrates its efficacy as a robust tool for regression-based studies involving multicollinearity. The study highlights two limitations: the necessity of accurately identifying the degree of multicollinearity and the importance of optimally tuning statistical model parameters. Future research should address these limitations to enhance model performance further, reinforcing the conclusion that statistical models are effective solutions for multicollinearity challenges in real-world applications.

Introduction

The introduction of this research paper highlights the significance of mathematical modeling in prediction and inference, particularly in the context of large datasets with numerous explanatory variables. As data collection technologies have advanced, the prevalence of multicollinearity—where highly correlated variables can distort statistical and machine learning models—has become a critical issue that is often overlooked. This phenomenon can lead to inaccurate inferences and diminished predictive performance, necessitating a deeper examination of its implications across various fields, including health, marketing, and finance.

The paper aims to fill a gap in the literature by comparing traditional statistical methods, such as the Liu regression and ridge regression, with prominent machine learning approaches in addressing multicollinearity. The authors assert that statistical models can outperform complex machine learning models in certain contexts and emphasize the importance of acknowledging multicollinearity’s impact on model performance. The study is structured to first review related literature, then discuss the multicollinearity problem and diagnostic methods, followed by a detailed explanation of the models employed, the modeling process, experimental settings, and finally, the benchmarking results and conclusions.

Discussion

In the discussion of multicollinearity, the paper highlights foundational contributions to the field, beginning with the Stein estimator proposed by James and Stein, which introduced equal shrinkage of coefficients in ordinary least squares (OLS) regression. However, this approach was limited by its inability to handle coefficients with opposite signs and the need for alternative estimators. The ridge estimator by Hoerl and Kennard addressed these issues by incorporating a penalty term to shrink coefficients towards zero. Liu’s estimator further refined this by combining ridge and Stein properties, while the Least Absolute Shrinkage and Selection Operator (Lasso) introduced variable selection capabilities by allowing coefficients to be shrunk to zero. To address the limitations of Lasso in high-dimensional settings, Zou and Hastie proposed the elastic net, which integrates both ridge and Lasso techniques.

The section also outlines the consequences of multicollinearity, including inflated variances of coefficient estimates and potential degradation of model predictions. Various diagnostic methods for detecting multicollinearity are discussed, such as calculating the correlation matrix, variance inflation factor (VIF), and eigenvalue analysis. These methods provide insights into the relationships among explanatory variables, with specific thresholds indicating severe multicollinearity. The paper concludes by summarizing solutions to multicollinearity, including data collection, alternative regression techniques (like ridge, Liu, Lasso, and elastic net), and preprocessing methods such as centering and scaling. The focus of the study is on evaluating these alternative models to effectively address multicollinearity in regression analysis.

شارك: