DOI: https://doi.org/10.1038/s41598-025-86418-9
PMID: https://pubmed.ncbi.nlm.nih.gov/39837951
تاريخ النشر: 2025-01-21
المؤلف: G. Avinash وآخرون
الموضوع الرئيسي: التأمين، الوفيات، الديموغرافيا، إدارة المخاطر
نظرة عامة
تؤكد ورقة البحث على الدور الحاسم لنماذج التعلم الآلي في التنبؤ بإشغال الأسرة في مرافق الصحة النفسية، لا سيما استجابةً للطلب المتزايد على خدمات الصحة النفسية في الهند. تقيم الدراسة ستة خوارزميات تعلم آلي – الانحدار بواسطة دعم المتجهات، وتعزيز التدرج المتطرف، والغابة العشوائية، وأقرب الجيران، وتعزيز التدرج، وشجرة القرار – باستخدام بيانات من 2008 إلى 2024. تم تقييم دقة النماذج التنبؤية من خلال متوسط الخطأ النسبي المطلق واختبار ديبولد-ماريانو. تشير النتائج إلى أن نماذج الغابة العشوائية وشجرة القرار قدمت أكثر التنبؤات دقة، وهو أمر ضروري لإدارة الموارد الفعالة وتخطيط السعة في رعاية الصحة النفسية.
تسلط الخاتمة الضوء على عدم كفاية الأساليب الإحصائية التقليدية بسبب الخصائص الفريدة للبيانات، مما يستلزم استخدام أساليب تعلم آلي أكثر قوة. تؤكد الدراسة على ضرورة معالجة الاتجاه المتزايد في إشغال الأسرة، المدفوع بالعوامل الاجتماعية التي تسهم في قضايا الصحة النفسية. يُوصى بإجراء أبحاث مستقبلية لتعزيز دقة التنبؤ من خلال دمج نماذج التعلم العميق والمتغيرات الخارجية مثل التغيرات الموسمية وبيانات قبول المرضى المصنفة حسب رموز ICD. تهدف هذه المقاربة إلى تحسين التنبؤات وتمكين مديري الرعاية الصحية من تخصيص استراتيجيات إدارة الموارد بشكل فعال، مما يؤدي في النهاية إلى تحسين رعاية المرضى والنتائج في قطاع الصحة النفسية.
الطرق
توضح قسم “الطرق” المنهجية المستخدمة لتطوير نماذج التعلم الآلي في الدراسة. تفصل اختيار الخوارزميات، وتقنيات معالجة البيانات، ومقاييس التقييم المستخدمة لتقييم أداء النموذج. من المحتمل أن تتضمن البحث مجموعة متنوعة من أساليب التعلم الآلي، مثل التعلم تحت الإشراف والتعلم غير المراقب، لمعالجة المشكلة المحددة المطروحة.
بالإضافة إلى ذلك، قد يصف القسم مجموعة البيانات المستخدمة، بما في ذلك حجمها، وميزاتها، وأي تحويلات ذات صلة تم تطبيقها لتعزيز دقة النموذج. تعتبر مقاييس التقييم، مثل الدقة، والدقة، والاسترجاع، أو درجة F1، ضرورية لتحديد فعالية النماذج في تحقيق أهداف البحث. بشكل عام، فإن هذا الإطار المنهجي ضروري لتكرار الدراسة والتحقق من النتائج.
النتائج
تكشف نتائج الدراسة حول بيانات إشغال الأسرة الأسبوعية من مستشفيات الصحة النفسية بين 1 يناير 2008 و31 يوليو 2024 عن رؤى مهمة حول أنماط استخدام الأسرة. من خلال تحليل 866 ملاحظة أسبوعية، أظهرت البيانات نطاقًا واسعًا من مستويات الإشغال، مع حد أدنى يبلغ 143 سريرًا وحد أقصى يبلغ 607 أسرّة. كان متوسط الإشغال 495.86، بينما كان الوسيط أعلى قليلاً عند 514.78، مما يشير إلى توزيع مائل إلى اليسار (الانحراف = -2.30) وتباين معتدل (الانحراف المعياري = 81.60، معامل التباين = 16.44%). تشير القيمة العالية للكراتوس (6.32) إلى توزيع لبتوكورتيك، مما يعكس انحرافات أكثر تطرفًا عن المتوسط مما هو متوقع في توزيع طبيعي، على الأرجح بسبب عوامل مثل قبول المرضى وخروجهم المتأثرين باضطرابات الصحة النفسية المحددة.
أكد تحليل السلاسل الزمنية على ثبات بيانات إشغال الأسرة من خلال اختبارات ديكي-فولر المعززة (ADF) واختبارات كفياتكوفسكي-فيليبس-شميت-شين (KPSS). ومع ذلك، أشارت اختبارات الطبيعية إلى انحرافات تتماشى مع الانحراف والكراتوس الملحوظين، المنسوبين إلى تعقيدات ديناميات مرضى الصحة النفسية، بما في ذلك أنماط الانتكاس ومدد العلاج. كانت الاتجاهات الموسمية واضحة، مع ذروات في إشغال الأسرة تحدث في 2012 و2015 و2019، وبشكل ملحوظ في 2020، على الأرجح بسبب تأثير جائحة COVID-19 على الصحة النفسية. أظهرت التغيرات الشهرية زيادة في الإشغال من مايو إلى أغسطس، بينما لوحظت مستويات أقل في يناير وفبراير. استخدمت الدراسة نماذج تعلم آلي متنوعة، بما في ذلك الانحدار بواسطة دعم المتجهات والغابة العشوائية، مع ضبط المعلمات لتحسين الدقة التنبؤية لهذه الأنماط المعقدة من الإشغال.
المناقشة
تسلط قسم المناقشة في ورقة البحث الضوء على فعالية نماذج التعلم الآلي المختلفة في التنبؤ بإشغال الأسرة الأسبوعية في مستشفيات الصحة النفسية. حددت عملية ضبط المعلمات الأمثل التكوينات المثلى لستة نماذج: الانحدار بواسطة دعم المتجهات (SVR)، XGBoost، الغابة العشوائية، أقرب الجيران (KNN)، تعزيز التدرج، وشجرة القرار. من بين هذه النماذج، برز نموذج الغابة العشوائية كالأكثر دقة، محققًا خطأ الجذر التربيعي المتوسط (RMSE) قدره 22.99، ومتوسط الخطأ المطلق (MAE) قدره 16.18، ومتوسط الخطأ النسبي المطلق (MAPE) قدره 3.57% على بيانات الاختبار. كما أدت نموذج شجرة القرار أداءً جيدًا، مع مقاييس خطأ أعلى قليلاً، بينما أظهرت الأساليب التجميعية مثل تعزيز التدرج وXGBoost نتائج تنافسية ولكنها كانت أقل دقة من الغابة العشوائية وشجرة القرار.
تؤكد التحليلات على أهمية اختيار نماذج قوية للتنبؤ بالسلاسل الزمنية في بيئات الرعاية الصحية، خاصة بالنظر إلى الطبيعة المعقدة وغير الخطية لبيانات إشغال الأسرة. كشفت التحليل الإحصائي للبيانات عن اتجاهات قابلة للتنبؤ وأنماط موسمية، يمكن أن تُفيد في تخصيص الموارد وقرارات التوظيف. أشار اختبار ديبولد-ماريانو إلى أن نماذج الغابة العشوائية وشجرة القرار أدت بشكل مشابه، مما يشير إلى قابليتها للتبادل في التطبيقات العملية. في المقابل، أظهرت KNN وSVR أداءً ضعيفًا بشكل ملحوظ، مما يدل على عدم ملاءمتها لهذه المهمة التنبؤية. بشكل عام، تدعو النتائج إلى استخدام نماذج الغابة العشوائية وشجرة القرار لإدارة الموارد بشكل فعال في مستشفيات الصحة النفسية، مع اعتبار XGBoost وتعزيز التدرج كخيارات تكميلية لتعزيز الدقة التنبؤية.
DOI: https://doi.org/10.1038/s41598-025-86418-9
PMID: https://pubmed.ncbi.nlm.nih.gov/39837951
Publication Date: 2025-01-21
Author(s): G. Avinash et al.
Primary Topic: Insurance, Mortality, Demography, Risk Management
Overview
The research paper emphasizes the critical role of machine learning models in forecasting bed occupancy in mental health facilities, particularly in response to the increasing demand for mental health services in India. The study evaluates six machine learning algorithms—Support Vector Regression, eXtreme Gradient Boosting, Random Forest, K-Nearest Neighbors, Gradient Boosting, and Decision Tree—using data from 2008 to 2024. The models’ predictive accuracy was assessed through the Mean Absolute Percentage Error and the Diebold-Mariano test. The findings indicate that Random Forest and Decision Tree models provided the most accurate forecasts, which are essential for effective resource management and capacity planning in mental health care.
The conclusion highlights the inadequacy of conventional statistical methods due to the data’s unique properties, necessitating the use of more robust machine learning approaches. The study underscores the urgency of addressing the rising trend in bed occupancy, driven by societal factors contributing to mental health issues. Future research is recommended to enhance forecasting accuracy by incorporating deep learning models and external variables such as seasonal changes and patient admission data categorized by ICD codes. This approach aims to refine predictions and enable healthcare administrators to tailor resource management strategies effectively, ultimately improving patient care and outcomes in the mental health sector.
Methods
The section on “Methods” outlines the methodology employed for developing machine learning models in the study. It details the selection of algorithms, data preprocessing techniques, and evaluation metrics used to assess model performance. The research likely incorporates a variety of machine learning approaches, such as supervised and unsupervised learning, to address the specific problem at hand.
Additionally, the section may describe the dataset utilized, including its size, features, and any relevant transformations applied to enhance model accuracy. The evaluation metrics, such as accuracy, precision, recall, or F1-score, are crucial for determining the effectiveness of the models in achieving the research objectives. Overall, this methodological framework is essential for replicating the study and validating the findings.
Results
The results of the study on weekly bed occupancy data from mental health hospitals between January 1, 2008, and July 31, 2024, reveal significant insights into bed utilization patterns. Analyzing 866 weekly observations, the data exhibited a wide range of occupancy levels, with a minimum of 143 beds and a maximum of 607 beds. The average occupancy was 495.86, while the median was slightly higher at 514.78, indicating a left-skewed distribution (skewness = -2.30) and moderate variability (standard deviation = 81.60, coefficient of variation = 16.44%). The high kurtosis value (6.32) suggests a leptokurtic distribution, reflecting more extreme deviations from the mean than expected in a normal distribution, likely due to factors such as patient admissions and discharges influenced by specific mental health disorders.
Time series analysis confirmed the stationarity of the bed occupancy data through the Augmented Dickey-Fuller (ADF) and Kwiatkowski-Phillips-Schmidt-Shin (KPSS) tests. However, tests for normality indicated deviations consistent with the observed skewness and kurtosis, attributed to the complexities of mental health patient dynamics, including relapse patterns and treatment durations. Seasonal trends were evident, with peaks in bed occupancy occurring in 2012, 2015, 2019, and notably in 2020, likely due to the COVID-19 pandemic’s impact on mental health. Monthly variations showed increased occupancy from May to August, while lower levels were observed in January and February. The study employed various machine learning models, including Support Vector Regression and Random Forest, with hyperparameter tuning to enhance predictive accuracy for these complex occupancy patterns.
Discussion
The discussion section of the research paper highlights the effectiveness of various machine learning models in forecasting weekly bed occupancy in mental health hospitals. The hyperparameter tuning process identified optimal configurations for six models: Support Vector Regression (SVR), XGBoost, Random Forest, K-Nearest Neighbors (KNN), Gradient Boosting, and Decision Tree. Among these, the Random Forest model emerged as the most accurate, achieving a Root Mean Squared Error (RMSE) of 22.99, a Mean Absolute Error (MAE) of 16.18, and a Mean Absolute Percentage Error (MAPE) of 3.57% on test data. The Decision Tree model also performed well, with slightly higher error metrics, while ensemble methods like Gradient Boosting and XGBoost showed competitive results but were less precise than Random Forest and Decision Tree.
The analysis underscores the importance of selecting robust models for time series forecasting in healthcare settings, particularly given the complex and non-linear nature of bed occupancy data. The statistical decomposition of the data revealed predictable trends and seasonal patterns, which can inform resource allocation and staffing decisions. The Diebold-Mariano test indicated that Random Forest and Decision Tree models performed similarly, suggesting their interchangeability for practical applications. In contrast, KNN and SVR exhibited significantly poorer performance, indicating their unsuitability for this forecasting task. Overall, the findings advocate for the use of Random Forest and Decision Tree models for effective resource management in mental health hospitals, while also considering XGBoost and Gradient Boosting as supplementary options for enhanced predictive accuracy.
