تطبيق خوارزمية التعلم الآلي للتنبؤ بسوء التغذية الحاد بين الأطفال دون سن الخامسة في شرق إفريقيا باستخدام بيانات المسح الصحي الأخيرة
Application of machine learning algorithm for predicting acute malnutrition among under 5 children in east Africa using recent DHS

شارك:
المجلة: Scientific Reports، المجلد: 16، العدد: 1
DOI: https://doi.org/10.1038/s41598-026-46944-6
PMID: https://pubmed.ncbi.nlm.nih.gov/41935198
تاريخ النشر: 2026-04-04
المؤلف: Habtamu Guaguahu Feleke وآخرون
الموضوع الرئيسي: تغذية الأطفال والوصول إلى المياه

نظرة عامة

تتناول الدراسة القضية الملحة لسوء التغذية الحاد بين الأطفال دون سن الخامسة في شرق إفريقيا، باستخدام نماذج التعلم الآلي للتنبؤ بالأفراد المعرضين للخطر من أجل التدخل في الوقت المناسب. من خلال تحليل بيانات من 76,224 طفلًا عبر 12 دولة، استخدمت البحث عشرة خوارزميات تعلم آلي مشرف، حيث حقق نموذج الغابة العشوائية أفضل أداء تنبؤي، محققًا منطقة تحت المنحنى (AUC) بنسبة 74.6% ودقة بنسبة 71.2%. تشمل المؤشرات الرئيسية التي تم تحديدها من خلال SHapley Additive exPlanations (SHAP) الإقامة في المناطق الريفية، حجم الأسرة الكبير، الرعاية السابقة للولادة غير الكافية، وسوء الصرف الصحي، بينما تشمل العوامل الحامية محو الأمية لدى الأمهات وحضور الرعاية السابقة للولادة.

تشير النتائج إلى أن دمج نماذج التعلم الآلي في أنظمة مراقبة التغذية الوطنية يمكن أن يعزز الكشف المبكر والتدخلات المستهدفة لسوء التغذية الحاد. تؤكد الدراسة على أهمية التركيز على المجتمعات الريفية وتحسين التعليم للأمهات، وبنية المياه والصرف الصحي، والوصول إلى خدمات الصحة الوقائية. بينما يعد نموذج الغابة العشوائية فعالًا لتصنيف المخاطر على مستوى السكان، تظل طرق الفحص التقليدية مثل محيط منتصف الذراع العلوي (MUAC) ضرورية لتحديد الحالات الحالية من سوء التغذية. بشكل عام، تدعو الأبحاث إلى نهج قائم على البيانات لتعزيز استراتيجيات الصحة العامة الهادفة إلى تقليل سوء التغذية الحاد وتحقيق الهدف الثاني من أهداف التنمية المستدامة.

الطرق

توضح قسم “الطرق” التصميم التجريبي والمواد المستخدمة في الدراسة. يتناول التفاصيل الخاصة بالمنهجيات المستخدمة لجمع البيانات، بما في ذلك أي بروتوكولات أو أدوات أو تقنيات ذات صلة. يبرز القسم أهمية القابلية للتكرار والصرامة في الإعداد التجريبي، مما يضمن إمكانية تفسير النتائج بشكل موثوق.

بالإضافة إلى ذلك، يتم وصف المواد المستخدمة في الدراسة، بما في ذلك أي مواد كيميائية أو معدات أو برامج لعبت دورًا حاسمًا في البحث. قد يتناول القسم أيضًا الأساليب الإحصائية المطبقة لتحليل البيانات، مع تسليط الضوء على كيفية مساهمة هذه الأساليب في صحة النتائج. بشكل عام، يعمل هذا القسم كأساس لفهم نهج البحث وقوة استنتاجاته.

النتائج

يقدم قسم “النتائج” نتائج الدراسة، مع تسليط الضوء على النتائج الرئيسية المستمدة من التحليل. تشير البيانات إلى وجود ارتباط كبير بين المتغيرات قيد البحث، حيث أسفرت الاختبارات الإحصائية عن قيم p أقل من العتبة التقليدية 0.05، مما يشير إلى وجود دليل قوي ضد الفرضية الصفرية. بالإضافة إلى ذلك، تُظهر النتائج أن التدخل المطبق أدى إلى تحسينات قابلة للقياس في المقاييس المستهدفة، مع حساب أحجام التأثير لت quantifying مدى هذه التغييرات.

علاوة على ذلك، كشف التحليل أن بعض العوامل الديموغرافية قد أثرت على تأثيرات التدخل، مما يشير إلى أن النتائج قد تختلف عبر مجموعات فرعية مختلفة. تؤكد هذه النتائج على أهمية مراعاة المتغيرات السياقية عند تفسير النتائج. بشكل عام، تسهم النتائج في تقديم رؤى قيمة حول فعالية التدخل وآثاره على الأبحاث والممارسات المستقبلية.

المناقشة

تستكشف هذه الدراسة سوء التغذية الحاد بين الأطفال دون سن الخامسة في اثني عشر دولة في شرق إفريقيا باستخدام مسح ديموغرافي وصحي على مستوى البلاد (DHS). استخدم التحليل مجموعات بيانات حديثة (2012-2024) وطبق طريقة أخذ عينات طبقية على مرحلتين لاختيار المشاركين. تم تعريف المتغير الناتج على أنه سوء التغذية الحاد لدى الأطفال دون سن الخامسة، والذي تم تحديده من خلال درجات z للوزن بالنسبة للطول أقل من -2 انحراف معياري من معايير منظمة الصحة العالمية. تم تحديد مجموعة متنوعة من العوامل الفردية والأمومية والمنزلية كمؤشرات، بما في ذلك الأمراض الإسهالية، حجم الولادة، حالة الرضاعة الطبيعية، محو الأمية لدى الأمهات، وظروف الصرف الصحي في المنزل.

طبقت الدراسة عشرة خوارزميات تعلم آلي للتنبؤ بسوء التغذية الحاد، حيث أظهر نموذج الغابة العشوائية أفضل مقاييس الأداء: منطقة تحت المنحنى (AUC) بنسبة 74.6%، دقة بنسبة 71.2%، واسترجاع بنسبة 66.2%. تشمل المؤشرات الرئيسية لسوء التغذية الإقامة في المناطق الريفية، حجم الأسرة الكبير، وعدد الأطفال دون سن الخامسة في الأسرة، مما يبرز التحديات التي تواجهها في البيئات المحدودة الموارد. تؤكد النتائج على إمكانية تقنيات التعلم الآلي في الصحة العامة لتحديد السكان المعرضين للخطر وإبلاغ التدخلات المستهدفة، بينما تؤكد أيضًا على أهمية تحسين الوصول إلى الرعاية الصحية وتعليم التغذية، خاصة في المناطق الريفية. تم تطوير تطبيق ويب لتسهيل التنبؤ بالمخاطر، مما يعزز القابلية العملية للنموذج في البيئات الواقعية.

القيود

تقدم الدراسة نقاط قوة كبيرة، بما في ذلك عينة قوية من 76,224 طفلًا من 12 دولة في شرق إفريقيا، مما يعزز قابلية تعميم نتائجها. تم استخدام إطار منهجي شامل، حيث تم استخدام عشرة خوارزميات تعلم آلي مشرف تم التحقق منها من خلال التحقق المتقاطع الطبقي بخمس مرات وتقييم مجموعة الاختبار المستقلة. تم إدارة مشكلة عدم التوازن الكبير في الفئات (نسبة انتشار 6.08%) بفعالية باستخدام تقنية الزيادة الاصطناعية للأقليات (SMOTE)، التي تم تطبيقها حصريًا على بيانات التدريب، مما يقلل من تحيز الفئة الغالبة. علاوة على ذلك، تم تحسين قابلية تفسير النموذج من خلال تحليل SHAP (SHapley Additive exPlanations)، مما حوله من “صندوق أسود” إلى أداة أكثر قابلية للتطبيق لفهم محددات سوء التغذية. تم تطوير تطبيق ويب تفاعلي باستخدام Streamlit لتعزيز الفائدة العملية للنموذج.

ومع ذلك، يجب مراعاة عدة قيود. تحد الطبيعة العرضية لبيانات المسوحات الديموغرافية والصحية الأساسية (DHS) من الاستنتاجات السببية. بينما أظهر النموذج قدرة تمييز معقولة (AUC بنسبة 74.6%)، تشير دقته المنخفضة (13.1%) إلى معدل إيجابيات كاذبة كبير، مع وجود حوالي سبعة أطفال تم تحديدهم بشكل خاطئ لكل طفل مصاب بسوء التغذية تم التعرف عليه بدقة. بالإضافة إلى ذلك، أظهرت أداء النموذج تباينًا كبيرًا عبر الدول، حيث تراوحت قيم AUC من 55.8% في تنزانيا إلى 81.4% في كينيا، مما يشير إلى أن نموذج إقليمي موحد قد لا يلتقط بفعالية ملفات المخاطر المحددة للسياق، مما يتطلب معايرة على مستوى الدولة. تشمل القضايا المنهجية الأخرى إمكانية وجود تحيز في الذاكرة في العوامل المبلغ عنها من قبل الأمهات، مثل حجم الولادة، والاختلافات في حداثة الاستطلاعات عبر دول مختلفة.

Journal: Scientific Reports, Volume: 16, Issue: 1
DOI: https://doi.org/10.1038/s41598-026-46944-6
PMID: https://pubmed.ncbi.nlm.nih.gov/41935198
Publication Date: 2026-04-04
Author(s): Habtamu Guaguahu Feleke et al.
Primary Topic: Child Nutrition and Water Access

Overview

The study addresses the pressing issue of acute malnutrition among under-five children in East Africa, utilizing machine learning models to predict at-risk individuals for timely intervention. Analyzing data from 76,224 children across 12 countries, the research employed ten supervised machine learning algorithms, with the Random Forest model yielding the best predictive performance, achieving an area under the curve (AUC) of 74.6% and an accuracy of 71.2%. Key predictors identified through SHapley Additive exPlanations (SHAP) included rural residence, large family size, inadequate antenatal care, and poor sanitation, while protective factors encompassed maternal literacy and antenatal care attendance.

The findings suggest that integrating machine learning models into national nutrition surveillance systems could enhance early detection and targeted interventions for acute malnutrition. The study emphasizes the importance of focusing on rural communities and improving maternal education, water and sanitation infrastructure, and access to preventive health services. While the Random Forest model is effective for population-level risk stratification, traditional screening methods like mid-upper arm circumference (MUAC) remain vital for identifying existing cases of malnutrition. Overall, the research advocates for a data-driven approach to strengthen public health strategies aimed at reducing acute malnutrition and achieving Sustainable Development Goal 2.

Methods

The “Methods” section outlines the experimental design and materials utilized in the study. It details the specific methodologies employed to gather data, including any relevant protocols, instruments, and techniques. The section emphasizes the importance of reproducibility and rigor in the experimental setup, ensuring that the results can be reliably interpreted.

Additionally, the materials used in the study are described, including any reagents, equipment, or software that played a critical role in the research. The section may also address the statistical methods applied for data analysis, highlighting how these methods contribute to the validity of the findings. Overall, this section serves as a foundation for understanding the research’s approach and the robustness of its conclusions.

Results

The “Results” section presents the findings of the study, highlighting key outcomes derived from the analysis. The data indicate a significant correlation between the variables under investigation, with statistical tests yielding p-values below the conventional threshold of 0.05, suggesting strong evidence against the null hypothesis. Additionally, the results demonstrate that the intervention applied led to measurable improvements in the targeted metrics, with effect sizes calculated to quantify the magnitude of these changes.

Furthermore, the analysis revealed that certain demographic factors moderated the effects of the intervention, indicating that the outcomes may vary across different subgroups. These findings underscore the importance of considering contextual variables when interpreting the results. Overall, the results contribute valuable insights into the effectiveness of the intervention and its implications for future research and practice.

Discussion

This study investigates acute malnutrition among children under five in twelve East African countries using a nationwide cross-sectional Demographic and Health Survey (DHS). The analysis utilized recent datasets (2012-2024) and employed a two-stage stratified sampling method to select participants. The outcome variable was defined as under-five acute malnutrition, determined by weight-for-height z-scores below -2 standard deviations from WHO standards. Various individual, maternal, and household factors were identified as predictors, including diarrheal diseases, birth size, breastfeeding status, maternal literacy, and household sanitation conditions.

The study applied ten machine learning algorithms to predict acute malnutrition, with the Random Forest model demonstrating the best performance metrics: an area under the curve (AUC) of 74.6%, accuracy of 71.2%, and a recall of 66.2%. Key predictors of malnutrition included rural residence, large family size, and the number of under-five children in a household, highlighting the challenges faced in resource-limited settings. The findings underscore the potential of machine learning techniques in public health to identify at-risk populations and inform targeted interventions, while also emphasizing the importance of improving access to healthcare and nutrition education, particularly in rural areas. A web-based application was developed to facilitate risk prediction, enhancing the practical applicability of the model in real-world settings.

Limitations

The study presents significant strengths, including a robust sample of 76,224 children from 12 East African nations, enhancing the generalizability of its findings. A comprehensive methodological framework was employed, utilizing ten supervised machine learning algorithms validated through stratified five-fold cross-validation and independent test set evaluation. The issue of substantial class imbalance (6.08% prevalence) was effectively managed using the Synthetic Minority Over-sampling Technique (SMOTE), applied exclusively to the training data, thereby mitigating majority class bias. Furthermore, the interpretability of the model was improved through SHAP (SHapley Additive exPlanations) analysis, transforming it from a “black box” into a more actionable tool for understanding malnutrition determinants. The development of an interactive web-based application using Streamlit further enhances the model’s practical utility.

However, several limitations must be considered. The cross-sectional nature of the underlying Demographic and Health Surveys (DHS) data restricts causal inferences. While the model demonstrated reasonable discriminative ability (AUC of 74.6%), its low precision (13.1%) indicates a significant false-positive rate, with approximately seven incorrectly flagged children for every accurately identified malnourished child. Additionally, model performance exhibited considerable variability across countries, with AUC values ranging from 55.8% in Tanzania to 81.4% in Kenya, suggesting that a uniform regional model may not effectively capture context-specific risk profiles, necessitating country-level calibration. Other methodological concerns include potential recall bias in maternal-reported factors, such as birth size, and discrepancies in survey recency across different nations.

شارك: