DOI: https://doi.org/10.1371/journal.pntd.0014156
PMID: https://pubmed.ncbi.nlm.nih.gov/41911272
تاريخ النشر: 2026-03-30
المؤلف: Nkuba Nyerere وآخرون
الموضوع الرئيسي: الطفيليات وتفاعلاتها مع المضيف
نظرة عامة
تبحث الدراسة في التحدي المستمر للأمراض المدارية المهملة (NTDs) في كينيا، حيث تعقد التوطن المشترك للديدان الطفيلية المنقولة بالتربة (STH) وداء البلهارسيا (SCH) وفيلاريا اللمفية (LF) جهود السيطرة. باستخدام نماذج التعلم الآلي، وتحديداً الغابة العشوائية (RF) وآلة تعزيز التدرج (GBM) وتعزيز التدرج المتطرف (XGBoost)، تهدف الدراسة إلى تحديد المناطق عالية المخاطر لتوطن الأمراض المدارية المهملة بناءً على مؤشرات ديموغرافية ومؤشرات المياه والصرف الصحي والنظافة (WASH). كشفت تحليل بيانات توطن الأمراض المدارية المهملة لعام 2022 أن نموذج RF حقق أعلى أداء تنبؤي (AUC = 0.70)، مما يبرز المؤشرات الرئيسية مثل تحسين الوصول إلى الصرف الصحي وكثافة السكان.
تشير النتائج إلى أن شرق وشمال شرق كينيا هما نقاط ساخنة هامة للأمراض المدارية المهملة، مما يتوافق مع انخفاض تغطية WASH (<40%). تؤكد الدراسة على أهمية دمج تحسينات WASH مع استراتيجيات إدارة الأدوية الجماعية لتعزيز فعالية التدخلات. علاوة على ذلك، تقترح تطوير لوحة معلومات في الوقت الحقيقي لرسم الخرائط الديناميكية للمخاطر، بما يتماشى مع خطة كينيا الاستراتيجية للقضاء على الأمراض المدارية المهملة 2030. بشكل عام، تظهر الدراسة إمكانيات التعلم الآلي كأداة استراتيجية لتوقع مخاطر الأمراض المدارية المهملة وتسهيل التدخلات المستهدفة في البيئات ذات الموارد المحدودة.
مقدمة
تناقش مقدمة ورقة البحث توطن الأمراض المشتركة، والذي يشير إلى التداخل الجغرافي لعدة أمراض ضمن نفس السكان، مع التركيز بشكل خاص على الأمراض المدارية المهملة (NTDs) التي تؤثر على حوالي 1.5 مليار شخص على مستوى العالم. في أفريقيا، تفرض هذه الأمراض، وخاصة فيلاريا اللمفية (LF) والديدان الطفيلية المنقولة بالتربة (STH) وداء البلهارسيا (SCH) وغيرها، عبئاً صحياً كبيراً، حيث يتركز 39% من عبء الأمراض المدارية المهملة العالمي في هذه المنطقة. على الرغم من الجهود المستمرة لإدارة الأدوية الجماعية، لا تزال كينيا متوطنة لـ LF وSCH وSTH، خاصة في المناطق الريفية ذات الصرف الصحي السيئ. تسلط الورقة الضوء على الحاجة إلى تحسين استراتيجيات المراقبة والتدخل بسبب بقاء بؤر الانتقال المستمرة.
يؤكد المؤلفون على دور العوامل البيئية والاجتماعية والاقتصادية في انتقال STH وتأثيرات التوطن المشترك مع LF وSCH، التي تشترك في عوامل خطر متداخلة. تعتمد جهود السيطرة الحالية على رسم الخرائط الجغرافية والخرائط الدورية، لكن هذه الطرق غالباً ما تفتقر إلى الدقة التنبؤية. تقترح الورقة استخدام تقنيات التعلم الآلي (ML) لتعزيز مراقبة الأمراض المدارية المهملة من خلال نمذجة التفاعلات المعقدة بين المتغيرات الديموغرافية ومؤشرات المياه والصرف الصحي والنظافة (WASH). من خلال الاستفادة من مجموعة بيانات التوطن المشترك في كينيا لعام 2022، تهدف الدراسة إلى تطوير نماذج ML (بما في ذلك الغابة العشوائية وآلة تعزيز التدرج وXGBoost) لتوقع ارتفاع انتشار STH وحالة التوطن المشترك، وتحديد المحركات الرئيسية للانتقال، وتحديد النقاط الساخنة ذات العبء العالي للتدخلات المستهدفة.
الطرق
تستخدم الدراسة سير عمل منظم، كما هو موضح في الشكل 1، لتوجيه عمليات تحليل البيانات من البداية إلى النهاية. يوضح هذا المخطط التدفق الخطوات المتسلسلة المعنية في النماذج والمنهجية المقترحة، مما يضمن ارتباطاً واضحاً بين كل مرحلة من مراحل الدراسة.
في سياق التعلم الآلي، تشمل المنهجية مراحل مختلفة، على الرغم من عدم تقديم تفاصيل محددة بشأن الخوارزميات أو التقنيات المستخدمة في هذا القسم. تهدف الطريقة المنهجية إلى تسهيل تحليل البيانات بشكل قوي وتعزيز موثوقية النتائج المستمدة من الدراسة.
النتائج
يقدم قسم “النتائج” من ورقة البحث النتائج الرئيسية المستمدة من التجارب أو التحليلات التي تم إجراؤها. عادةً ما يتضمن ذلك بيانات كمية، وتحليلات إحصائية، وأي أرقام أو جداول ذات صلة توضح النتائج. غالباً ما تتم مقارنة النتائج مع الفرضيات الأولية أو الدراسات السابقة لتسليط الضوء على الفروق أو التأكيدات الهامة.
في هذا القسم، قد يذكر المؤلفون مقاييس محددة، مثل القيم المتوسطة والانحرافات المعيارية أو قيم p، لإظهار الأهمية الإحصائية لنتائجهم. بالإضافة إلى ذلك، يتم مناقشة أي اتجاهات أو أنماط ملحوظة، مما يوفر رؤى حول تداعيات النتائج ضمن السياق الأوسع لمجال البحث. بشكل عام، يخدم هذا القسم لنقل الأدلة التجريبية التي تدعم استنتاجات الدراسة.
المناقشة
تبحث الدراسة في التوطن المشترك لفيلاريا اللمفية (LF) والديدان الطفيلية المنقولة بالتربة (STH) وداء البلهارسيا (SCH) في كينيا باستخدام بيانات ثانوية من المشروع الخاص الموسع للقضاء على الأمراض المدارية المهملة (ESPEN) لعام 2022. شمل التحليل 47 مقاطعة و290 وحدة تنفيذ، مع التركيز على ضرورة العلاج الكيميائي الوقائي عبر المناطق المتوطنة. تم إجراء تنظيف البيانات ومعالجتها باستخدام برنامج R، مع تطبيق تقدير المتوسط لمعالجة القيم المفقودة، مما يضمن سلامة مجموعة البيانات لنمذجة التعلم الآلي. كانت المتغير الرئيسي الناتج هو حالة التوطن المشترك، التي تم تصنيفها إلى فئات متميزة، مما يسهل تطبيق خوارزميات التعلم الآلي.
تم استخدام ثلاثة نماذج للتعلم الآلي – الغابة العشوائية (RF) وآلة تعزيز التدرج (GBM) وXGBoost – لتوقع توطن الأمراض بناءً على مجموعة متنوعة من المؤشرات، بما في ذلك حجم السكان وتدابير الصرف الصحي. أظهر نموذج RF أعلى أداء تنبؤي مع منطقة تحت منحنى ROC (AUC) تبلغ 0.703، مما يدل على قدرات تصنيف قوية. في المقابل، أظهرت نماذج GBM وXGBoost قيم AUC أقل تبلغ 0.620 و0.656 على التوالي. تؤكد النتائج على أهمية استراتيجيات إدارة الأمراض المتكاملة، حيث كانت الغالبية العظمى من وحدات التنفيذ إما غير متوطنة أو تواجه توطن مرض واحد، مما يبرز الحاجة إلى تدخلات مستهدفة في المناطق المتوطنة. بشكل عام، تسهم الدراسة في تقديم رؤى قيمة حول المشهد الوبائي للأمراض المدارية المهملة في كينيا وتؤكد على دور التعلم الآلي في تعزيز توقع الأمراض وجهود السيطرة عليها.
القيود
تقدم الدراسة حول تطبيقات التعلم الآلي في نمذجة الأمراض المدارية المهملة (NTDs) عدة قيود يجب الاعتراف بها. تتمثل إحدى القضايا الرئيسية في نقص البيانات الشاملة، مما يؤثر على قابلية المقارنة لانتشار الأمراض عبر مناطق مختلفة. بالإضافة إلى ذلك، كانت الدراسة مقيدة بمجموعة محدودة من المتغيرات الاجتماعية والديموغرافية التي يمكن أن تؤثر بشكل كبير على ديناميات انتقال الأمراض.
كان الأداء التنبؤي للنماذج معتدلاً، مع منطقة تحت المنحنى (AUC) تتراوح من 0.62 إلى 0.70. بينما تشير هذه القيم إلى قدرة تنبؤية تتجاوز التصنيف العشوائي، فإنها تشير أيضاً إلى احتمال حدوث تصنيف خاطئ عند تحديد المناطق عالية المخاطر. وبالتالي، يجب اعتبار مخرجات النموذج كأدوات دعم القرار بدلاً من كونها مؤشرات نهائية لانتقال الأمراض. علاوة على ذلك، قد يؤدي استخدام تقدير المتوسط إلى إدخال تحيز من خلال التقليل من التباين داخل مجموعة البيانات. يجب أن تأخذ الأبحاث المستقبلية في الاعتبار استخدام تقنيات تقدير أكثر قوة، مثل الجيران الأقرب (k-NN) أو التقدير المتعدد، لتعزيز تمثيل البيانات ودقة النموذج.
DOI: https://doi.org/10.1371/journal.pntd.0014156
PMID: https://pubmed.ncbi.nlm.nih.gov/41911272
Publication Date: 2026-03-30
Author(s): Nkuba Nyerere et al.
Primary Topic: Parasites and Host Interactions
Overview
The research investigates the persistent challenge of Neglected Tropical Diseases (NTDs) in Kenya, where co-endemicity of soil-transmitted helminths (STH), schistosomiasis (SCH), and lymphatic filariasis (LF) complicates control efforts. Utilizing machine learning models, specifically Random Forest (RF), Gradient Boosting Machine (GBM), and Extreme Gradient Boosting (XGBoost), the study aimed to identify high-risk areas for NTD co-endemicity based on demographic and Water, Sanitation, and Hygiene (WASH) indicators. The analysis of 2022 NTD co-endemicity data revealed that the RF model achieved the highest predictive performance (AUC = 0.70), highlighting key predictors such as improved sanitation access and population density.
The findings indicate that Eastern and North-Eastern Kenya are significant hotspots for NTDs, correlating with low WASH coverage (<40%). The study underscores the importance of integrating WASH improvements with mass drug administration strategies to enhance intervention efficacy. Furthermore, it proposes the development of a real-time dashboard for dynamic risk mapping, aligning with Kenya's NTD Elimination Strategic Plan 2030. Overall, the research demonstrates the potential of machine learning as a strategic tool for predicting NTD risks and facilitating targeted interventions in resource-limited settings.
Introduction
The introduction of the research paper discusses co-endemicity, which refers to the geographic overlap of multiple diseases within the same population, particularly focusing on neglected tropical diseases (NTDs) that affect approximately 1.5 billion people globally. In Africa, these diseases, particularly lymphatic filariasis (LF), soil-transmitted helminthiasis (STH), schistosomiasis (SCH), and others, impose a significant health burden, with 39% of the global NTD burden concentrated in this region. Despite ongoing mass drug administration efforts, Kenya remains endemic for LF, SCH, and STH, particularly in rural areas with poor sanitation. The paper highlights the need for refined surveillance and intervention strategies due to persistent transmission hotspots.
The authors emphasize the role of environmental and socioeconomic factors in STH transmission and the exacerbating effects of co-endemicity with LF and SCH, which share overlapping risk factors. Current control efforts rely on geospatial mapping and periodic surveys, but these methods often lack predictive precision. The paper proposes the use of machine learning (ML) techniques to enhance NTD surveillance by modeling the complex interactions between demographic and water, sanitation, and hygiene (WASH) variables. By leveraging Kenya’s 2022 co-endemicity dataset, the study aims to develop ML models (including Random Forest, Gradient Boosting Machine, and XGBoost) to predict high STH prevalence and co-endemicity status, identify key drivers of transmission, and pinpoint high-burden hotspots for targeted interventions.
Methods
The research employs a structured workflow, as depicted in Figure 1, to guide the data analysis processes from initiation to conclusion. This flowchart outlines the sequential steps involved in the proposed models and methodology, ensuring a clear linkage between each phase of the study.
In the context of machine learning, the methodology encompasses various stages, although specific details regarding the algorithms or techniques utilized are not provided in this section. The systematic approach aims to facilitate robust data analysis and enhance the reliability of the findings derived from the study.
Results
The “Results” section of the research paper presents the key findings derived from the conducted experiments or analyses. It typically includes quantitative data, statistical analyses, and any relevant figures or tables that illustrate the outcomes. The results are often compared against the initial hypotheses or previous studies to highlight significant differences or confirmations.
In this section, the authors may report on specific metrics, such as mean values, standard deviations, or p-values, to demonstrate the statistical significance of their findings. Additionally, any observed trends or patterns are discussed, providing insights into the implications of the results within the broader context of the research field. Overall, this section serves to convey the empirical evidence that supports the study’s conclusions.
Discussion
The study investigates the co-endemicity of lymphatic filariasis (LF), soil-transmitted helminths (STH), and schistosomiasis (SCH) in Kenya using secondary data from the Expanded Special Project for Elimination of Neglected Tropical Diseases (ESPEN) for the year 2022. The analysis encompassed 47 counties and 290 implementation units, focusing on the necessity for preventive chemotherapy across endemic regions. Data cleaning and preprocessing were performed using R software, with mean imputation applied to address missing values, ensuring the dataset’s integrity for machine learning modeling. The primary outcome variable was the co-endemicity status, which was categorized into distinct classes, facilitating the application of machine learning algorithms.
Three machine learning models—Random Forest (RF), Gradient Boosting Machine (GBM), and XGBoost—were employed to predict disease co-endemicity based on various predictors, including population size and water sanitation measures. The RF model demonstrated the highest predictive performance with an area under the ROC curve (AUC) of 0.703, indicating robust classification capabilities. In contrast, the GBM and XGBoost models exhibited lower AUC values of 0.620 and 0.656, respectively. The findings underscore the importance of integrated disease management strategies, as the majority of implementation units were either non-endemic or faced single disease endemicity, highlighting the need for targeted interventions in co-endemic regions. Overall, the study contributes valuable insights into the epidemiological landscape of NTDs in Kenya and emphasizes the role of machine learning in enhancing disease prediction and control efforts.
Limitations
The study on the applications of machine learning in modeling neglected tropical diseases (NTDs) presents several limitations that should be acknowledged. A primary concern is the lack of comprehensive data, which affects the comparability of prevalence across different regions. Additionally, the study was constrained by a limited set of social and demographic covariates that could significantly influence disease transmission dynamics.
The predictive performance of the models was moderate, with an area under the curve (AUC) ranging from 0.62 to 0.70. While these values indicate a predictive capability that exceeds random classification, they also imply potential misclassification when identifying high-risk areas. Consequently, the model outputs should be regarded as decision-support tools rather than definitive predictors of disease transmission. Furthermore, the use of mean imputation may introduce bias by underestimating variability within the dataset. Future research should consider employing more robust imputation techniques, such as k-nearest neighbors (k-NN) or multiple imputation, to enhance data representation and model accuracy.
