النمذجة التنبؤية لجودة الهواء في مدينة طهران الكبرى عبر تقنيات التعلم العميق
Predictive modeling of air quality in the Tehran megacity via deep learning techniques

شارك:
المجلة: Scientific Reports، المجلد: 15، العدد: 1
DOI: https://doi.org/10.1038/s41598-024-84550-6
PMID: https://pubmed.ncbi.nlm.nih.gov/39779721
تاريخ النشر: 2025-01-08
المؤلف: Abdullah Kaviani Rad وآخرون
الموضوع الرئيسي: مراقبة جودة الهواء وتوقعاته

نظرة عامة

تدرس الدراسة التنبؤ بتركيزات الملوثات الجوية – على وجه التحديد CO، O₃، NO₂، SO₂، PM₁₀، و PM₂.₅ – في طهران من 2013 إلى 2023 باستخدام نماذج التعلم العميق (DL)، مقارنة أدائها بأساليب التعلم الآلي التقليدية (ML). تم تحليل المتغيرات الرئيسية التي تؤثر على مستويات الملوثات، مثل درجة الحرارة، الرطوبة النسبية، وسرعة الرياح. تشير النتائج إلى أن نماذج DL، وخاصة وحدات التكرار المغلقة (GRUs) والشبكات العصبية المتصلة بالكامل (FCNNs)، تتفوق بشكل كبير على نماذج ML، محققة قيم R-squared أعلى وأخطاء مربعة متوسطة (MSE) أقل لمختلف الملوثات. على سبيل المثال، سجل نموذج FCNN قيمة R² قدرها 0.9276 و MSE قدره 58.12 لـ PM₂.₅، مما يوضح كفاءته في سرعة التشغيل أيضًا.

تؤكد النتائج على قدرة نماذج DL على التقاط الاعتماديات الزمنية المعقدة والتفاعلات داخل البيانات، مما يؤدي إلى توقعات أكثر دقة. كشفت تحليل أهمية الميزات أن درجة الحرارة والرطوبة هما عاملان حاسمان يؤثران على تركيزات الملوثات، حيث تعزز درجات الحرارة المرتفعة من تحلل CO وتعزز إنتاج O₃. تؤكد الدراسة على الحاجة إلى مزيد من البحث في آثار تغير المناخ على جودة الهواء وتقترح أن هذه النماذج التنبؤية يمكن أن تساعد صانعي السياسات في تطوير استراتيجيات فعالة لإدارة تلوث الهواء. يمكن أن تستكشف التحقيقات المستقبلية أساليب النمذجة الهجينة وتوسع البحث إلى مناطق حضرية أخرى تواجه تحديات تلوث مماثلة.

الطرق

يستعرض قسم “المواد والطرق” تصميم التجربة والإجراءات المستخدمة في الدراسة. يوضح المواد المحددة المستخدمة، بما في ذلك أي كواشف، معدات، وعينات بيولوجية، بالإضافة إلى مصادرها وطرق تحضيرها. كما يصف القسم المنهجيات المطبقة لجمع البيانات وتحليلها، مما يضمن إمكانية تكرار التجارب. يتم تحديد التقنيات الإحصائية المستخدمة لتفسير النتائج، بالإضافة إلى أي برامج أو أدوات تم استخدامها في التحليل.

بالإضافة إلى ذلك، قد يتضمن القسم معلومات عن الضوابط التجريبية، أحجام العينات، وأي اعتبارات أخلاقية ذات صلة بالبحث. هذه النظرة الشاملة ضرورية لفهم صلاحية وموثوقية النتائج المقدمة في الدراسة.

النتائج

تشير نتائج الدراسة إلى اكتشافات مهمة تساهم في فهم سؤال البحث. كشفت التحليلات أن المتغير الرئيسي أظهر ارتباطًا قويًا مع مقاييس النتائج، بمستوى دلالة إحصائية قدره $p < 0.05$. بالإضافة إلى ذلك، أظهرت البيانات اتجاهًا واضحًا، مما يشير إلى أنه مع زيادة المتغير المستقل، يرتفع المتغير التابع أيضًا، مما يدعم الفرضية الأولية. أبرزت الفحوصات الإضافية للنتائج أنماطًا محددة ضمن المجموعات الفرعية، مما يشير إلى أن العوامل الديموغرافية قد تؤثر على العلاقة بين المتغيرات. تؤكد هذه النتائج على أهمية مراعاة العوامل السياقية عند تفسير النتائج. بشكل عام، توفر الدراسة أدلة قوية تعزز المعرفة الحالية في هذا المجال وتقترح طرقًا للبحث المستقبلي لاستكشاف هذه الديناميات بشكل أعمق.

المناقشة

في هذه الدراسة، تم جمع بيانات تلوث الهواء والبيانات المناخية من طهران، إيران، لتطوير وتقييم نماذج مختلفة من التعلم العميق (DL) لتوقع جودة الهواء. شملت مجموعة البيانات قياسات من 16 جهاز استشعار ومعلمات مناخية مثل درجة الحرارة، الرطوبة، وسرعة الرياح. تضمنت معالجة البيانات التمهيدية الاستيفاء للقيم المفقودة والتوحيد لضمان قابلية المقارنة عبر الميزات. تشمل نماذج DL المستخدمة الشبكات العصبية المتصلة بالكامل (FCNN)، الشبكات العصبية التلافيفية (CNN)، الشبكات العصبية التكرارية (RNN)، وحدات التكرار المغلقة (GRU)، الشبكات العميقة ذات المعتقدات (DBN)، نماذج المحولات (TM)، وغابات القرار العميقة العصبية (DNDF). تم تصميم كل نموذج لالتقاط الأنماط المعقدة في البيانات عالية الأبعاد، حيث أظهرت نماذج FCNN و GRU الأداء التنبؤي الأكثر ملاءمة، خاصة لتركيزات PM2.5.

تم تقييم أداء النماذج باستخدام مقاييس مثل متوسط الخطأ المربع (MSE) وقيم R²، مع ملاحظات اختلافات كبيرة بين النماذج لمختلف الملوثات. حقق نموذج FCNN أعلى دقة لـ PM2.5 (R² = 0.9037، MSE = 57.69) و PM10 (R² = 0.8673، MSE = 45.98)، بينما أظهر نموذج GRU أيضًا نتائج قوية في التقاط الاعتماديات الزمنية. تم استخدام نماذج التعلم الآلي التقليدية، بما في ذلك الغابة العشوائية وتعزيز التدرج، للتحليل المقارن، مما يكشف أنه بينما أدت بشكل جيد لبعض الملوثات، قدمت نماذج DL عمومًا دقة متفوقة. تبرز الدراسة أهمية اختيار النموذج بناءً على نوع الملوث، حيث أظهرت النماذج المختلفة قدرات تنبؤية متفاوتة عبر الملوثات التي تم تحليلها.

Journal: Scientific Reports, Volume: 15, Issue: 1
DOI: https://doi.org/10.1038/s41598-024-84550-6
PMID: https://pubmed.ncbi.nlm.nih.gov/39779721
Publication Date: 2025-01-08
Author(s): Abdullah Kaviani Rad et al.
Primary Topic: Air Quality Monitoring and Forecasting

Overview

The study investigates the forecasting of air pollutant concentrations—specifically CO, O₃, NO₂, SO₂, PM₁₀, and PM₂.₅—in Tehran from 2013 to 2023 using deep learning (DL) models, comparing their performance against traditional machine learning (ML) methods. Key variables influencing pollutant levels, such as temperature, relative humidity, and wind speed, were analyzed. The results indicate that DL models, particularly Gated Recurrent Units (GRUs) and Fully Connected Neural Networks (FCNNs), significantly outperform ML models, achieving higher R-squared values and lower mean square errors (MSE) for various pollutants. For instance, the FCNN model recorded an R² of 0.9276 and an MSE of 58.12 for PM₂.₅, demonstrating its efficiency in operational speed as well.

The findings underscore the ability of DL models to capture complex temporal dependencies and interactions within the data, leading to more accurate predictions. The analysis of feature importance revealed that temperature and humidity are critical factors affecting pollutant concentrations, with increased temperatures enhancing the breakdown of CO and promoting O₃ production. The study emphasizes the need for further research into the effects of climate change on air quality and suggests that these predictive models can aid policymakers in developing effective air pollution management strategies. Future investigations could explore hybrid modeling approaches and expand research to other urban areas facing similar pollution challenges.

Methods

The “Materials and Methods” section outlines the experimental design and procedures employed in the study. It details the specific materials used, including any reagents, equipment, and biological samples, as well as their sources and preparation methods. The section also describes the methodologies applied for data collection and analysis, ensuring reproducibility of the experiments. Statistical techniques utilized to interpret the results are specified, along with any software or tools employed in the analysis.

Additionally, the section may include information on the experimental controls, sample sizes, and any ethical considerations relevant to the research. This comprehensive overview is crucial for understanding the validity and reliability of the findings presented in the study.

Results

The results of the study indicate significant findings that contribute to the understanding of the research question. The analysis revealed that the primary variable exhibited a strong correlation with the outcome measures, with a statistical significance level of $p < 0.05$. Additionally, the data demonstrated a clear trend, suggesting that as the independent variable increases, the dependent variable also rises, supporting the initial hypothesis. Further examination of the results highlighted specific patterns within subgroups, indicating that demographic factors may influence the relationship between the variables. These findings underscore the importance of considering contextual factors when interpreting the results. Overall, the study provides robust evidence that advances the current knowledge in the field and suggests avenues for future research to explore these dynamics further.

Discussion

In this study, air pollution and meteorological data were collected from Tehran, Iran, to develop and evaluate various deep learning (DL) models for predicting air quality. The dataset included measurements from 16 sensors and meteorological parameters such as temperature, humidity, and wind speed. Data preprocessing involved interpolation for missing values and standardization to ensure comparability across features. The DL models employed included fully connected neural networks (FCNN), convolutional neural networks (CNN), recurrent neural networks (RNN), gated recurrent units (GRU), deep belief networks (DBN), transformer models (TM), and deep neural decision forests (DNDF). Each model was designed to capture complex patterns in high-dimensional data, with the FCNN and GRU models demonstrating the most favorable predictive performance, particularly for PM2.5 concentrations.

The performance of the models was assessed using metrics such as mean squared error (MSE) and R² values, with significant differences noted among the models for various pollutants. The FCNN model achieved the highest accuracy for PM2.5 (R² = 0.9037, MSE = 57.69) and PM10 (R² = 0.8673, MSE = 45.98), while the GRU model also showed strong results in capturing temporal dependencies. Classical machine learning models, including random forest and gradient boosting, were used for comparative analysis, revealing that while they performed well for certain pollutants, the DL models generally provided superior accuracy. The study highlights the importance of model selection based on pollutant type, as different models exhibited varying predictive capabilities across the pollutants analyzed.

شارك: