DOI: https://doi.org/10.1371/journal.pone.0335072
PMID: https://pubmed.ncbi.nlm.nih.gov/41134860
تاريخ النشر: 2025-10-24
المؤلف: Muhammad Luqman وآخرون
الموضوع الرئيسي: طرق ونماذج إحصائية متقدمة
نظرة عامة
تتناول البحث التحديات التي تطرحها البيانات ذات الحالة السيئة والتعدد الخطي في نمذجة الانحدار، والتي يمكن أن تؤدي إلى تقديرات غير دقيقة للمعلمات وزيادة في الأخطاء المعيارية. تعيق هذه القضايا القدرة على تقييم تأثير المتغيرات التفسيرية بدقة على المتغير المستجيب، مما يزيد من خطر الأخطاء من النوع الثاني. للتخفيف من هذه المشاكل، تم اقتراح تقنيات مختلفة، وكان الانحدار الجبهي واحدًا من أكثرها استخدامًا. تتضمن هذه الطريقة مصطلح عقوبة في دالة الهدف للنموذج الخطي العام، مما يقدم تحيزًا محكومًا لتقليل متوسط مربع الخطأ (MSE).
يقدم الدراسة عقوبتين جديدتين للانحدار الجبهي ويقيم أدائهما من خلال محاكاة مونت كارلو عبر سيناريوهات مختلفة، مع التركيز على MSE كمقياس للأداء. تشير النتائج إلى أن مقدرات الانحدار الجبهي المقترحة تتفوق على الثوابت الجبهي الموجودة، خاصة في الحالات التي تتميز بتعدد خطي شديد، وأعداد متغيرة من المتغيرات التفسيرية، وأحجام عينات، وهياكل تباين الخطأ. يتم التحقق من النتائج بشكل أكبر من خلال التطبيقات الواقعية، مما يظهر أن مقدرات الانحدار الجبهي الجديدة تعزز الأداء بشكل كبير في حالات التعدد الخطي المعتدل إلى القوي. يوصي المؤلفون بتبني هذه المقدرات المقترحة للانحدار الجبهي لمعالجة قضايا التعدد الخطي بفعالية في تحليل الانحدار.
مقدمة
في مقدمة هذه الورقة البحثية، يتناول المؤلفون التحديات التي تطرحها التعدد الخطي في تحليل الانحدار المتعدد، حيث تعقد العلاقة العالية بين المتغيرات التفسيرية تفسير تقديرات المعاملات وتزيد من الأخطاء المعيارية. يؤدي ذلك إلى توسيع فترات الثقة وزيادة خطر الأخطاء من النوع الثاني. يقدم المؤلفون نموذج الانحدار الخطي المتعدد في شكل مصفوفة على أنه \( Y = X\beta + \mu \)، مع التأكيد على الافتراض الأساسي بعدم وجود اعتماد خطي قوي بين المتغيرات التفسيرية. عندما يتم انتهاك هذا الافتراض، تصبح تقديرات المربعات الصغرى العادية (OLS) مشكلة، مما يدفع إلى تقديم الانحدار الجبهي (RR) كحل للتخفيف من الآثار السلبية للتعدد الخطي.
يقدم الانحدار الجبهي ثابت تنظيم \( k \) لمعاقبة المعاملات الكبيرة، مما يحسن استقرار تقديرات المعلمات. يتم اشتقاق تقديرات RR كـ \( \beta_{RR} = (X’X + kI_p)^{-1}X’Y \)، حيث \( I_p \) هي مصفوفة الهوية من الرتبة \( p \). يعد اختيار معلمة الجبهي \( k \) أمرًا حاسمًا، حيث يوازن بين التحيز والتباين لتقليل متوسط مربع الخطأ (MSE) للمعلمات المقدرة. يشير المؤلفون إلى أنه على الرغم من اقتراح طرق مختلفة لاختيار \( k \)، لا يزال هناك حاجة لاستراتيجيات عقوبة أكثر دقة تعتمد على البيانات تأخذ في الاعتبار شدة التعدد الخطي. تهدف الورقة إلى اقتراح مصطلحين جديدين للعقوبة للانحدار الجبهي ومقارنة أدائهما مع الطرق الحالية، مع تفاصيل الأقسام اللاحقة حول نظرة عامة على العقوبات الموجودة، ونتائج المحاكاة، والمقارنات على مجموعات بيانات حقيقية.
طرق
توضح قسم المنهجية تطبيق النموذج الخطي العام (GLM) لتقدير المعلمات في سياق الانحدار. يتم التعبير عن النموذج كـ \( Y = X\beta + \mu \)، حيث \( Y \) هو المتغير المستجيب، و\( X \) هو مصفوفة المتغيرات التفسيرية، و\( \beta \) يمثل المعاملات، و\( \mu \) هو مصطلح الخطأ. يتم تقديم الشكل القياسي للنموذج كـ \( y = Z\alpha + \mu \)، مع كون \( Z \) مصفوفة من المتغيرات التفسيرية الموحدة التي تم تحويلها بواسطة مصفوفة متعامدة \( E \). يتم اشتقاق مقدرات المعلمات \( \alpha \) باستخدام المربعات الصغرى العادية (OLS) والانحدار الجبهي (RR)، الممثلة بالمعادلات (9) و(10) على التوالي.
تتفاصيل القسم أيضًا حساب المعاملات الموحدة \( \beta_{\text{scaled}} \) وتقديرات المعلمات بوحدات أصلية \( \beta \) من خلال المعادلات (11) و(12). بالإضافة إلى ذلك، يوفر صيغًا لحساب متوسط مربع الخطأ (MSE) لكل من مقدرات OLS وRR، كما هو موضح في المعادلات (14) و(15). تتضمن حسابات MSE القيم الذاتية لمصفوفة التباين للمتغيرات التفسيرية، مما يبرز تأثير معلمة الجبهي \( k \) على استقرار ودقة التقديرات. بشكل عام، تؤسس هذه المنهجية إطارًا قويًا لتقدير المعلمات في نماذج الانحدار الخطي، مع معالجة قضايا التعدد الخطي من خلال استخدام الانحدار الجبهي.
نتائج
تظهر النتائج من محاكاة مونت كارلو فعالية مقدر الانحدار الجبهي المقترح، المسمى \( k_{\text{New2}} \)، عبر مجموعات مختلفة من مستويات التعدد الخطي (ρ)، وأحجام العينات (n)، والانحرافات المعيارية (σ). على وجه التحديد، تكشف الجداول من 1 إلى 6 أن \( k_{\text{New2}} \) يظهر باستمرار أقل متوسط مربع خطأ متوقع (EMSE) مقارنة بمقدرات الانحدار الجبهي الأخرى لتكوينات مختلفة، بما في ذلك الحالات التي يكون فيها \( p = 4 \) و\( p = 6 \). من الجدير بالذكر أنه حتى في السيناريوهات ذات التعدد الخطي العالي (مثل \( ρ = 0.85 \) و\( ρ = 0.90 \))، يتفوق \( k_{\text{New2}} \) على منافسيه، باستثناء حالات محددة حيث يظهر المقدار القائم على الوسيط \( k_{\text{KMED}} \) الذي اقترحه كيبريا (2003) قيم EMSE أقل بشكل طفيف.
تشير التحليلات إلى أن جميع مقدرات الانحدار الجبهي تتفوق على المربعات الصغرى العادية (OLS) في التعامل مع التعدد الخطي، مع انخفاض قيم EMSE كلما زادت حجم العينة. ومع ذلك، ترتبط مستويات التعدد الخطي الأقوى بزيادة EMSE، مما يشير إلى آثار سلبية محتملة. بشكل عام، تؤكد النتائج على قوة المقدّر المقترح \( k_{\text{New2}} \) في التخفيف من التحديات التي تطرحها التعدد الخطي، مما يجعله خيارًا متفوقًا في معظم السيناريوهات التي تم فحصها. توضح التمثيلات الرسومية للنتائج في الأشكال من 1 إلى 6 هذه الاتجاهات بشكل أكبر.
مناقشة
في قسم المناقشة، تستعرض الورقة مجموعة من مقدرات الانحدار الجبهي الموجودة التي تم تطويرها لمعالجة التعدد الخطي، بدءًا من العمل الأساسي لهوريل وكينارد (1970) الذين قدموا ثابت الجبهي \( k_{HK} = \frac{\sigma^2}{\alpha^2_{\text{max}}} \). وقد اقترحت المساهمات اللاحقة من باحثين مثل هوريل وكينارد وبالدين (1975)، ولوليس ووانغ (1976)، وغيرهم معلمات جبهي بديلة، كل منها يصقل النهج للتعامل بشكل أفضل مع التعدد الخطي في مجموعات البيانات. من الجدير بالذكر أن كيبريا (2003) قدم مقدرات تعتمد على وسائل إحصائية مختلفة، بينما اقترح خلف وشكر (2005) ومؤلفون لاحقون تعديلات تتضمن القيم الذاتية وتقديرات تباين الخطأ.
يقترح المؤلفون مقدرين جديدين للانحدار الجبهي، \( k_{NEW1} \) و\( k_{NEW2} \)، وهما معتمدان على البيانات ويأخذان في الاعتبار عدد الحالة للبيانات، مما يكيف قوة التنظيم مع شدة التعدد الخطي. يتم تقييم المقدرات المقترحة من خلال محاكاة مونت كارلو عبر سيناريوهات مختلفة، مما يظهر أداءً متفوقًا من حيث متوسط مربع الخطأ المقدر (EMSE) مقارنة بالطرق الحالية، خاصة في ظل ظروف التعدد الخطي القوي. تؤكد التطبيقات على البيانات الحقيقية فعالية المقدرات المقترحة، مما يشير إلى إمكانياتها للاستخدام العملي في تحليل الانحدار حيث تطرح التعدد الخطي تحديات كبيرة.
DOI: https://doi.org/10.1371/journal.pone.0335072
PMID: https://pubmed.ncbi.nlm.nih.gov/41134860
Publication Date: 2025-10-24
Author(s): Muhammad Luqman et al.
Primary Topic: Advanced Statistical Methods and Models
Overview
The research addresses the challenges posed by ill-conditioned data and multicollinearity in regression modeling, which can lead to inaccurate parameter estimates and inflated standard errors. These issues hinder the ability to accurately assess the impact of explanatory variables on the response variable, increasing the risk of type-II errors. To mitigate these problems, various techniques have been proposed, with ridge regression being one of the most widely used. This method incorporates a penalty term into the objective function of the general linear model, introducing a controlled bias to reduce mean square error (MSE).
The study introduces two novel ridge penalties and evaluates their performance through Monte Carlo simulations across different scenarios, focusing on MSE as the performance metric. Results indicate that the proposed ridge estimators outperform existing ridge constants, particularly in situations characterized by severe multicollinearity, varying numbers of explanatory variables, sample sizes, and error variance structures. The findings are further validated through real-life applications, demonstrating that the new ridge estimators significantly enhance performance in cases of moderate to strong multicollinearity. The authors recommend the adoption of these proposed ridge estimators to effectively address multicollinearity issues in regression analysis.
Introduction
In the introduction of this research paper, the authors address the challenges posed by multicollinearity in multiple regression analysis, where high correlation among explanatory variables complicates the interpretation of coefficient estimates and inflates standard errors. This leads to wider confidence intervals and an increased risk of type-II errors. The authors present the multiple linear regression model in matrix form as \( Y = X\beta + \mu \), emphasizing the fundamental assumption of no strong linear dependence among explanatory variables. When this assumption is violated, ordinary least squares (OLS) estimation becomes problematic, prompting the introduction of ridge regression (RR) as a solution to mitigate the adverse effects of multicollinearity.
Ridge regression introduces a regularization constant \( k \) to penalize large coefficients, thereby improving the stability of parameter estimates. The RR estimates are derived as \( \beta_{RR} = (X’X + kI_p)^{-1}X’Y \), where \( I_p \) is the identity matrix of order \( p \). The choice of the ridge parameter \( k \) is crucial, as it balances bias and variance to minimize the mean square error (MSE) of the estimated parameters. The authors note that while various methods for selecting \( k \) have been proposed, there remains a need for more refined, data-driven penalization strategies that account for the severity of multicollinearity. The paper aims to propose two new penalty terms for ridge regression and compare their performance against existing methods, with subsequent sections detailing the overview of existing penalties, simulation results, and comparisons on real-life datasets.
Methods
The methodology section outlines the application of the general linear model (GLM) for estimating parameters in a regression context. The model is expressed as \( Y = X\beta + \mu \), where \( Y \) is the response variable, \( X \) is the matrix of explanatory variables, \( \beta \) represents the coefficients, and \( \mu \) is the error term. The canonical form of the model is presented as \( y = Z\alpha + \mu \), with \( Z \) being a matrix of standardized explanatory variables transformed by an orthogonal matrix \( E \). The estimators for the parameters \( \alpha \) are derived using Ordinary Least Squares (OLS) and Ridge Regression (RR), represented by equations (9) and (10), respectively.
The section further details the computation of standardized coefficients \( \beta_{\text{scaled}} \) and parameter estimates in original units \( \beta \) through equations (11) and (12). Additionally, it provides formulas for calculating the mean square error (MSE) for both OLS and RR estimators, given by equations (14) and (15). The MSE calculations incorporate the eigenvalues of the covariance matrix of the explanatory variables, highlighting the impact of the ridge parameter \( k \) on the stability and accuracy of the estimates. Overall, this methodology establishes a robust framework for parameter estimation in linear regression models, addressing issues of multicollinearity through the use of ridge regression.
Results
The results from Monte Carlo simulations demonstrate the efficacy of the proposed ridge estimator, denoted as \( k_{\text{New2}} \), across various combinations of multicollinearity levels (ρ), sample sizes (n), and standard deviations (σ). Specifically, Tables 1 through 6 reveal that \( k_{\text{New2}} \) consistently exhibits the lowest expected mean squared error (EMSE) compared to other ridge estimators for different configurations, including cases with \( p = 4 \) and \( p = 6 \). Notably, even in scenarios with high multicollinearity (e.g., \( ρ = 0.85 \) and \( ρ = 0.90 \)), \( k_{\text{New2}} \) outperforms its competitors, except for specific instances where the median-based estimator \( k_{\text{KMED}} \) proposed by Kibria (2003) shows marginally lower EMSE values.
The analysis indicates that all ridge estimators surpass ordinary least squares (OLS) in handling multicollinearity, with EMSE values decreasing as sample size increases. However, stronger multicollinearity levels correlate with higher EMSE, suggesting potential adverse effects. Overall, the findings underscore the robustness of the proposed estimator \( k_{\text{New2}} \) in mitigating the challenges posed by multicollinearity, establishing it as a superior choice in most scenarios examined. Graphical representations of the results in Figures 1 through 6 further illustrate these trends.
Discussion
In the discussion section, the paper reviews various existing ridge regression estimators developed to address multicollinearity, starting from the foundational work of Hoerl and Kennard (1970) who introduced the ridge constant \( k_{HK} = \frac{\sigma^2}{\alpha^2_{\text{max}}} \). Subsequent contributions from researchers such as Hoerl, Kennard, and Baldwin (1975), Lawless and Wang (1976), and others have proposed alternative ridge parameters, each refining the approach to better handle multicollinearity in datasets. Notably, Kibria (2003) introduced estimators based on different statistical means, while Khalaf and Shukur (2005) and later authors suggested modifications that incorporate eigenvalues and error variance estimates.
The authors propose two new ridge estimators, \( k_{NEW1} \) and \( k_{NEW2} \), which are data-driven and account for the condition number of the data, thereby adapting the regularization strength to the severity of multicollinearity. The proposed estimators are evaluated through Monte Carlo simulations across various scenarios, demonstrating superior performance in terms of estimated mean square error (EMSE) compared to existing methods, particularly under conditions of strong multicollinearity. Real data applications further validate the effectiveness of the proposed estimators, indicating their potential for practical use in regression analysis where multicollinearity poses significant challenges.
