DOI: https://doi.org/10.1038/s41598-025-86698-1
PMID: https://pubmed.ncbi.nlm.nih.gov/39929873
تاريخ النشر: 2025-02-10
المؤلف: Yanwen Shao وآخرون
الموضوع الرئيسي: دراسات وبائية حول COVID-19
نظرة عامة
تدرس هذه الدراسة التأثير العالمي لـ COVID-19 من خلال تحليل البيانات من 15 دولة على مدى حوالي 900 يوم، مع التركيز على 44 ميزة مصنفة إلى معلومات وبائية، وخصائص الدول، والمناخ، وسياسات الوقاية. من خلال استخدام تقنيات اختيار الميزات، تحدد الأبحاث العوامل الرئيسية التي تؤثر على زيادة حالات COVID-19 الجديدة. تستخدم الدراسة نمذجة تنبؤية من خلال الذاكرة طويلة وقصيرة المدى (LSTM)، والانحدار باستخدام دعم المتجهات (SVR)، والشبكات التلافيفية الزمنية (TCN)، حيث أظهرت LSTM أداءً متفوقًا، خاصة في التنبؤ بالحالات الجديدة لألمانيا (درجة التباين المفسر 0.864)، وإيطاليا (0.860)، والولايات المتحدة (0.766).
تكشف النتائج أن مجموعة من عوامل المناخ ومعلومات الوباء تؤثر بشكل كبير على ديناميات الحالات، بينما تختلف استجابات السياسات في فعاليتها عبر دول مختلفة، لا سيما في إيطاليا، والهند، واليابان، والسويد. تؤكد التحليل على أهمية التدابير الصحية العامة الحالية، مما يشير إلى أن الدول يمكن أن تعزز استراتيجيات استجابتها من خلال إعطاء الأولوية للتدخلات مثل إغلاق المدارس وقيود التجمعات. تؤكد الدراسة أن نموذج LSTM يلتقط بشكل فعال الاعتماديات الزمنية في البيانات، متفوقًا على نماذج SVR وTCN، مما يوفر رؤى قيمة لإبلاغ استراتيجيات الصحة العامة وتحسين جهود الاستجابة لـ COVID-19 على مستوى العالم.
الطرق
تحدد قسم “الطرق” الأساليب التجريبية والتحليلية المستخدمة في الدراسة. استخدم الباحثون مجموعة من التقنيات الكمية والنوعية لجمع البيانات، مما يضمن تحليلًا شاملاً لسؤال البحث. تضمنت المنهجيات المحددة تجارب محكومة، واستطلاعات، ونمذجة إحصائية، تم تصميمها لاختبار الفرضيات التي تم صياغتها في بداية الدراسة.
شمل جمع البيانات عملية أخذ عينات منهجية لضمان التمثيل، تلاها تحليل إحصائي صارم باستخدام أدوات البرمجيات لتفسير النتائج. طبق الباحثون اختبارات إحصائية متنوعة، مثل اختبارات t وANOVA، لتقييم دلالة نتائجهم. بالإضافة إلى ذلك، يناقش القسم الاعتبارات الأخلاقية التي تم أخذها في الاعتبار خلال البحث، بما في ذلك الموافقة المستنيرة وتدابير خصوصية البيانات. بشكل عام، كانت الطرق المستخدمة قوية وتهدف إلى تقليل التحيز مع زيادة موثوقية النتائج.
النتائج
هدفت الدراسة إلى تحليل تأثير ميزات مختلفة على نمذجة العدد الإجمالي لحالات COVID-19 الجديدة المؤكدة خلال فترة 14 يومًا باستخدام نماذج الذاكرة طويلة وقصيرة المدى (LSTM). شمل التحليل أربع مجموعات بيانات متميزة: معلومات وبائية، معلومات أساسية عن الدول/المناطق، بيانات المناخ، وسياسات الوقاية من الوباء، كما هو موضح في الجدول S1. بالإضافة إلى ذلك، سعت الأبحاث إلى كشف أي علاقات كامنة بين استجابات السياسات وتقلبات الحالات الجديدة المؤكدة خلال الإطار الزمني المحدد.
المناقشة
في هذه الدراسة، بحث المؤلفون في العوامل التي تؤثر على حالات COVID-19 الجديدة المؤكدة عبر 15 دولة/منطقة، مع التركيز على دور ميزات مختلفة، بما في ذلك المناخ واستجابات السياسات. من خلال اختيار الميزات، حددوا 28 ميزة مهمة، حيث احتلت عوامل المناخ مثل الرطوبة ودرجة الحرارة مرتبة عالية بجانب إحصائيات الوباء المباشرة. ومن الجدير بالذكر أن استجابات السياسات المحددة – مثل إغلاق المدارس وقيود التجمعات – تم الاحتفاظ بها في مجموعة فرعية (المجموعة 4*) وأظهرت تأثيرًا إيجابيًا على دقة التنبؤ. تم استخدام نموذج الذاكرة طويلة وقصيرة المدى (LSTM) للتنبؤ، مما كشف أن مجموعة الميزات من المجموعات 1 و3 و4* حققت أفضل مقاييس الأداء، بما في ذلك درجات التباين المفسر (EVS) ودرجات R².
تفاوتت توقعات النموذج حسب الدولة، حيث أظهرت ألمانيا وإيطاليا والولايات المتحدة النتائج الأكثر دقة، بينما أظهرت هونغ كونغ والسويد والصين ملاءمة أقل. أشار المؤلفون إلى أن التدابير الصارمة لـ COVID-19 في الصين قد تشوه توقعات النموذج، مما أدى إلى استبعادها من مجموعة البيانات لمزيد من التحليل. على الرغم من هذا التعديل، ظلت الأداء العام لنموذج LSTM مستقرًا، مما يدل على قوته. بالإضافة إلى ذلك، قارنت الدراسة فعالية نموذج LSTM ضد نماذج التعلم الآلي الأخرى، مثل الانحدار باستخدام دعم المتجهات (SVR) والشبكات التلافيفية الزمنية (TCN)، ووجدت أن LSTM تفوقت باستمرار على الآخرين في دقة التنبؤ. تسلط هذه الأبحاث الضوء على أهمية دمج ميزات متنوعة، بما في ذلك البيانات البيئية والمتعلقة بالسياسات، لتعزيز فهم وتنبؤ اتجاهات حالات COVID-19.
DOI: https://doi.org/10.1038/s41598-025-86698-1
PMID: https://pubmed.ncbi.nlm.nih.gov/39929873
Publication Date: 2025-02-10
Author(s): Yanwen Shao et al.
Primary Topic: COVID-19 epidemiological studies
Overview
This study investigates the global impact of COVID-19 by analyzing data from 15 countries over approximately 900 days, focusing on 44 features categorized into pandemic information, country characteristics, climate, and prevention policies. By employing feature selection techniques, the research identifies key factors influencing the rise in new COVID-19 cases. The study utilizes predictive modeling through long short-term memory (LSTM), support vector regression (SVR), and temporal convolutional networks (TCN), with LSTM demonstrating superior performance, particularly in predicting new cases for Germany (explained variance score of 0.864), Italy (0.860), and the United States (0.766).
The findings reveal that a combination of climate factors and pandemic information significantly influences case dynamics, while policy responses vary in effectiveness across different countries, notably in Italy, India, Japan, and Sweden. The analysis underscores the importance of current public health measures, suggesting that countries could enhance their response strategies by prioritizing interventions like school closures and gathering restrictions. The study confirms that the LSTM model effectively captures temporal dependencies in the data, outperforming SVR and TCN models, thereby providing valuable insights for informing public health strategies and improving COVID-19 response efforts globally.
Methods
The “Methods” section outlines the experimental and analytical approaches employed in the study. The researchers utilized a combination of quantitative and qualitative techniques to gather data, ensuring a comprehensive analysis of the research question. Specific methodologies included controlled experiments, surveys, and statistical modeling, which were designed to test the hypotheses formulated at the outset of the study.
Data collection involved a systematic sampling process to ensure representativeness, followed by rigorous statistical analysis using software tools to interpret the results. The researchers applied various statistical tests, such as t-tests and ANOVA, to evaluate the significance of their findings. Additionally, the section discusses the ethical considerations taken into account during the research, including informed consent and data privacy measures. Overall, the methods employed were robust and aimed at minimizing bias while maximizing the reliability of the results.
Results
The study aimed to analyze the impact of various features on modeling the total number of new confirmed COVID-19 cases over a 14-day period using Long Short-Term Memory (LSTM) models. The analysis incorporated four distinct data groups: pandemic information, basic information about countries/regions, climate data, and pandemic prevention policies, as detailed in Table S1. Additionally, the research sought to uncover any latent relationships between policy responses and the fluctuations in new confirmed cases during the specified timeframe.
Discussion
In this study, the authors investigated the factors influencing new confirmed COVID-19 cases across 15 countries/regions, focusing on the role of various features, including climate and policy responses. Through feature selection, they identified 28 significant features, with climate factors like humidity and temperature ranking highly alongside direct pandemic statistics. Notably, specific policy responses—such as school closures and restrictions on gatherings—were retained in a subset (group 4*) and demonstrated a positive impact on prediction accuracy. The Long Short-Term Memory (LSTM) model was employed for forecasting, revealing that the combination of features from groups 1, 3, and 4* yielded the best performance metrics, including explained variance scores (EVS) and R² scores.
The model’s predictions varied by country, with Germany, Italy, and the United States showing the most accurate results, while Hong Kong, Sweden, and China exhibited poorer fits. The authors noted that the strict COVID-19 measures in China likely distorted the model’s predictions, leading to its exclusion from the dataset for further analysis. Despite this adjustment, the overall performance of the LSTM model remained stable, indicating its robustness. Additionally, the study compared the LSTM model’s effectiveness against other machine learning models, such as Support Vector Regression (SVR) and Temporal Convolutional Networks (TCN), finding that LSTM consistently outperformed the others in predictive accuracy. This research highlights the importance of integrating diverse features, including environmental and policy-related data, to enhance the understanding and forecasting of COVID-19 case trends.
