DOI: https://doi.org/10.5194/gmd-17-7915-2024
تاريخ النشر: 2024-11-07
المؤلف: Leonardo Olivetti وآخرون
الموضوع الرئيسي: الظواهر الجوية والمحاكاة
نظرة عامة
في السنوات الأخيرة، ظهرت نماذج التنبؤ بالطقس المعتمدة على البيانات كبدائل تنافسية للنماذج العددية التقليدية المعتمدة على الفيزياء، لا سيما في قدرتها على التنبؤ بحدوث درجات حرارة قريبة من السطح وسرعات الرياح القصوى. تقيم هذه الدراسة موثوقية هذه النماذج المعتمدة على البيانات من خلال مقارنة أدائها مع نموذج ECMWF القائم على الفيزياء في سياق شبه تشغيلي. تشير النتائج إلى أن النماذج المعتمدة على البيانات تتفوق عمومًا على نموذج ECMWF من حيث متوسط الخطأ الجذري العالمي (RMSE) للتنبؤات التي تتراوح من 1 إلى 10 أيام مقدماً، وتظهر أداءً تنافسياً في التنبؤ بالأحداث الجوية القصوى عبر مناطق مختلفة.
ومع ذلك، فإن فعالية النماذج المعتمدة على البيانات ليست متساوية؛ حيث يتأثر أداؤها بعوامل مثل المنطقة الجغرافية، ونوع الحدث المتطرف المحدد، ومدة التنبؤ. تظهر النماذج أداءً قويًا بشكل خاص لدرجات الحرارة القصوى في المناطق الاستوائية وفي فترات التنبؤ القصيرة. وتخلص الدراسة إلى أنه بينما يمكن أن تكون النماذج المعتمدة على البيانات مكملات قيمة للتنبؤات المعتمدة على الفيزياء في المجالات التي تتفوق فيها، يجب معالجة العديد من التحديات قبل أن يمكن دمجها بالكامل في أنظمة التنبؤ التشغيلية.
مقدمة
تستعرض مقدمة هذه الورقة البحثية تطور وتقدم النماذج المعتمدة على التعلم العميق في التنبؤ بالطقس، مع تسليط الضوء على أدائها التنافسي مقارنة بالنماذج التقليدية المعتمدة على الفيزياء، لا سيما تلك الخاصة بالمركز الأوروبي للتنبؤات الجوية متوسطة المدى (ECMWF). منذ عام 2022، أفادت مجموعات بحثية متعددة أن نماذج التعلم العميق الخاصة بها يمكن أن تنتج تنبؤات حتمية أكثر دقة عبر متغيرات جوية مختلفة. وقد أكدت دراسات مستقلة هذه النتائج، مشددة على إمكانية تفوق النماذج المعتمدة على البيانات على الأساليب المعتمدة. ومع ذلك، تركز الأدبيات الحالية بشكل أساسي على مهارة التنبؤ المتوسطة، متجاهلة المجال الحرج للأحداث الجوية القصوى، والتي تعتبر ضرورية لإدارة الكوارث والاستقرار الاجتماعي والاقتصادي.
يجادل المؤلفون من أجل تقييم أكثر صرامة لنماذج التعلم العميق في التنبؤ بالأحداث الجوية القصوى، حيث قد لا تكون المنهجيات الحالية كافية لالتقاط أدائها في هذه السيناريوهات. يحددون التباينات المحتملة بين المهارة المتوسطة وتنبؤات الأحداث القصوى، مشيرين إلى مشكلات مثل ندرة الأحداث القصوى في مجموعات بيانات التدريب وقيود دوال الخسارة المتماثلة. تهدف هذه الورقة إلى تقييم مهارة نماذج التعلم العميق في التنبؤ بدرجات الحرارة القريبة من السطح وسرعات الرياح القصوى على مدى زمني يتراوح من 1 إلى 10 أيام، مقارنة بنظام التنبؤ المتكامل (IFS) الخاص بـ ECMWF باستخدام مقاييس ومعايير تقييم موحدة. ستفصل الأقسام اللاحقة النماذج والمنهجيات المستخدمة، وتعرض النتائج، وتناقش الآثار المترتبة على الاستخدام التشغيلي لنماذج التنبؤ المعتمدة على البيانات.
طرق
في هذا القسم، يوضح المؤلفون النماذج والمنهجية المستخدمة لمقارنة نماذج التنبؤ المعتمدة على البيانات والنماذج المعتمدة على الفيزياء بطريقة تضمن العدالة. يركز البحث على النماذج المعتمدة على البيانات من WeatherBench 2 التي يمكن أن تستخدم نفس الظروف الأولية كنظام التنبؤ المتكامل عالي الدقة (IFS HRES) الخاص بـ ECMWF. على وجه التحديد، تولد النماذج تنبؤات كل 6 ساعات لدرجات حرارة 2 م وسرعة الرياح 10 م، والتي يتم تقييمها مقابل بيانات إعادة التحليل ERA5 بدقة أفقية تبلغ 1.5°. تسلط الدراسة الضوء على الفروق المفاهيمية بين النماذج التشغيلية وتلك التي تستخدم بيانات إعادة التحليل، والتي لا تناسب التنبؤ في الوقت الحقيقي بسبب التأخيرات الزمنية الكامنة.
تم اختيار نموذجين معتمدين على البيانات، Pangu-Weather وGraph-Cast، لقدرتهم على تلبية المعايير المحددة وتمثيل أداء النماذج المعتمدة على البيانات بشكل فعال. تستخدم هذه النماذج هياكل متقدمة، بما في ذلك المحولات البصرية والشبكات العصبية الرسومية، وتظهر أداءً قابلاً للمقارنة عبر متغيرات جوية وسطحية مختلفة. لمعالجة القيود المحتملة في تنوع النماذج، يشير المؤلفون أيضًا إلى مقارنات إضافية في الملحق D، والتي تشمل إصدارات قائمة على إعادة التحليل من النماذج وFuXi، المعروفة بخطأ الجذر التربيعي المنخفض (RMSE) في التنبؤات المتوسطة إلى الطويلة المدى. للحصول على تفاصيل شاملة حول النماذج، يوجه المؤلفون القراء إلى الأدبيات ذات الصلة، بما في ذلك Rasp et al. (2024) وOlivetti وMessori (2024).
نتائج
في هذا القسم، يتم تقديم نتائج مقارنة نموذجية بين نماذج التنبؤ المعتمدة على البيانات والنماذج المعتمدة على الفيزياء، مع التركيز على أدائها عبر مناطق وظروف مختلفة. تكشف التحليلات أن كلا النموذجين المعتمدين على البيانات، لا سيما GraphCast، يتفوقان بشكل كبير على نموذج IFS HRES الخاص بـ ECMWF على مستوى العالم وفي معظم المناطق، خاصة في التنبؤ بسرعة الرياح 10 م ودرجة حرارة 2 م. ومن الجدير بالذكر أن GraphCast يظهر أداءً متفوقًا باستمرار في التنبؤات طويلة المدى، مع اختلافات تتراوح بين 5% إلى 20% لتنبؤات 10 أيام لدرجة حرارة 2 م. ومع ذلك، فإن أداء النماذج المعتمدة على البيانات، لا سيما Pangu-Weather، يتدهور بشكل أسرع من IFS HRES في فترات التنبؤ الأطول، مما يشير إلى مشكلة محتملة تُعرف بـ “التشويش”، حيث تتراجع التنبؤات نحو المناخ.
تشير التحليلات الإقليمية إلى أن النماذج المعتمدة على البيانات تؤدي بشكل أفضل في نصف الكرة الشمالي، لا سيما بالنسبة للحدود القصوى الباردة، بينما يتفوق IFS HRES في مناطق مثل AusNZ والقارة القطبية الجنوبية. تواجه النماذج المعتمدة على البيانات صعوبة في التعامل مع الحدود القصوى الباردة، خاصة في القارة القطبية الجنوبية، على الأرجح بسبب انخفاض جودة بيانات إعادة التحليل المستخدمة في التدريب. علاوة على ذلك، تظهر النماذج ميلاً للأداء الأفضل في درجات الحرارة القصوى مقارنةً بالحدود القصوى للرياح، ويعزى ذلك إلى نقص التدريب المحدد على سرعة الرياح. كما تسلط النتائج الضوء على التقديرات المنخفضة بشكل كبير للحدود القصوى الحارة وسرعة الرياح من قبل النماذج المعتمدة على البيانات، لا سيما في فترات التنبؤ الأطول، مما يعزز فكرة التشويش. بشكل عام، بينما تظهر النماذج المعتمدة على البيانات وعدًا، فإن أدائها يختلف بشكل كبير حسب المنطقة ومدة التنبؤ، مما يتطلب الحذر في استخدامها التشغيلي لتنبؤات الطقس القصوى.
مناقشة
تقيم قسم المناقشة في الورقة أداء التنبؤ لنموذج IFS HRES الخاص بـ ECMWF، وGraphCast، وPangu-Weather في التنبؤ بدرجات الحرارة القريبة من السطح وسرعات الرياح القصوى حتى 10 أيام مقدماً. تكشف التحليلات، بناءً على ثلاثة معايير—الدقة في التنبؤ بحجم الحدث المتطرف، وتحديد الحدث المتطرف عند نقاط الشبكة، ومعايرة الكميات القصوى—أن النماذج المعتمدة على البيانات تتفوق عمومًا على IFS HRES في معظم المناطق، لا سيما في المناطق الاستوائية. ومع ذلك، يحتفظ IFS HRES بميزة في المناطق ذات الكثافة السكانية العالية وفي التنبؤ بالحدود القصوى الباردة على فترات زمنية قصيرة. تظهر النماذج المعتمدة على البيانات انخفاضًا ملحوظًا في الأداء عند خطوط العرض العليا، ربما بسبب تحيزات التدريب، وتظهر ميلاً للعودة إلى المناخ في فترات التنبؤ الأطول، وهي ظاهرة تُعرف بالتشويش.
تشير النتائج إلى أنه بينما يمكن أن تتنافس النماذج المعتمدة على البيانات مع النماذج المعتمدة على الفيزياء، فإنها تواجه تحديات مثل محدودية المتغيرات المدخلة وتباين الأداء عبر المناطق وأنواع الأحداث. تدعو الورقة إلى نهج تنبؤ هجين يجمع بين نقاط القوة لكلا نوعي النماذج، لا سيما في المناطق التي تتفوق فيها النماذج المعتمدة على البيانات. يجب أن تركز الأبحاث المستقبلية على تحسين دوال الخسارة، ومعالجة قيود المتغيرات المدخلة، واستكشاف تأثير منهجيات التدريب على أداء النموذج، خاصة عند خطوط العرض العليا. بشكل عام، تسلط الدراسة الضوء على إمكانيات النماذج المعتمدة على البيانات في التنبؤ التشغيلي مع الاعتراف بالحاجة إلى مزيد من التقييم والتطوير.
DOI: https://doi.org/10.5194/gmd-17-7915-2024
Publication Date: 2024-11-07
Author(s): Leonardo Olivetti et al.
Primary Topic: Meteorological Phenomena and Simulations
Overview
In recent years, data-driven weather forecast models have emerged as competitive alternatives to traditional physics-based numerical models, particularly in their ability to predict near-surface temperature and wind speed extremes. This study evaluates the reliability of these data-driven models by comparing their performance against the ECMWF’s physics-based deterministic model in a semi-operational context. The findings indicate that data-driven models generally outperform the ECMWF model in terms of global root mean square error (RMSE) for forecasts ranging from 1 to 10 days ahead, and they show competitive performance in predicting extreme weather events across various regions.
However, the effectiveness of data-driven models is not uniform; their performance is influenced by factors such as geographical region, the specific type of extreme event, and the forecast lead time. The models demonstrate particularly strong performance for temperature extremes in tropical regions and at shorter lead times. The study concludes that while data-driven models can serve as valuable complements to physics-based forecasts in areas where they excel, several challenges must be addressed before they can be fully integrated into operational forecasting systems.
Introduction
The introduction of this research paper outlines the evolution and recent advancements in deep learning models for weather forecasting, highlighting their competitive performance against traditional physics-based models, particularly those from the European Centre for Medium-Range Weather Forecasts (ECMWF). Since 2022, multiple research groups have reported that their deep learning models can produce more accurate deterministic forecasts across various atmospheric variables. Independent studies have corroborated these findings, emphasizing the potential of data-driven models to outperform established methods. However, the existing literature primarily focuses on average forecasting skill, neglecting the critical area of extreme weather events, which are essential for disaster management and socio-economic stability.
The authors argue for a more rigorous evaluation of deep learning models in forecasting weather extremes, as current methodologies may not adequately capture their performance in these scenarios. They identify potential discrepancies between average skill and extreme event forecasting, citing issues such as the rarity of extreme events in training datasets and the limitations of symmetric loss functions. This paper aims to assess the skill of deep learning models in predicting near-surface temperature and wind speed extremes over a 1-10 day horizon, comparing them to ECMWF’s Integrated Forecasting System (IFS) using standardized metrics and evaluation methods. The subsequent sections will detail the models and methodologies used, present the results, and discuss implications for the operational use of data-driven forecasting models.
Methods
In this section, the authors detail the models and methodology used to compare data-driven and physics-based forecasting models in a manner that ensures fairness. The focus is on data-driven models from WeatherBench 2 that can utilize the same initial conditions as the Integrated Forecasting System High-Resolution (IFS HRES) of ECMWF. Specifically, the models generate 6-hourly forecasts for 2 m temperature and 10 m wind, which are evaluated against the ERA5 reanalysis data at a horizontal resolution of 1.5°. The study highlights the conceptual differences between operational models and those using reanalysis data, which are not suitable for real-time forecasting due to inherent time delays.
Two data-driven models, Pangu-Weather and Graph-Cast, are selected for their ability to meet the established criteria and represent the performance of deterministic data-driven models effectively. These models employ advanced architectures, including vision transformers and graph neural networks, and demonstrate comparable performance across various atmospheric and surface variables. To address potential limitations in model diversity, the authors also reference additional comparisons in Appendix D, which include reanalysis-based versions of the models and FuXi, noted for their low root mean square error (RMSE) in medium to long-range forecasting. For comprehensive details on the models, the authors direct readers to relevant literature, including Rasp et al. (2024) and Olivetti and Messori (2024).
Results
In this section, the results of a model comparison between data-driven and physics-based forecasting models are presented, focusing on their performance across various regions and conditions. The analysis reveals that both data-driven models, particularly GraphCast, significantly outperform ECMWF’s IFS HRES model globally and in most regions, especially in forecasting 10 m wind speed and 2 m temperature. Notably, GraphCast consistently shows superior performance in longer-range forecasts, with differences ranging from 5% to 20% for 10-day predictions of 2 m temperature. However, the performance of data-driven models, particularly Pangu-Weather, deteriorates more rapidly than IFS HRES at longer lead times, indicating a potential issue of “blurring,” where forecasts regress towards climatology.
Regional analyses indicate that data-driven models perform better in the Northern Hemisphere, particularly for cold extremes, while IFS HRES excels in regions like AusNZ and Antarctica. The data-driven models struggle with cold extremes, especially in Antarctica, likely due to lower quality reanalysis data used for training. Furthermore, the models exhibit a tendency to perform better for temperature extremes than for wind extremes, attributed to the lack of specific training on wind speed. The results also highlight significant underestimations of hot and wind speed extremes by data-driven models, particularly at longer lead times, reinforcing the notion of blurring. Overall, while data-driven models show promise, their performance varies significantly by region and forecast lead time, necessitating caution in their operational use for extreme weather forecasting.
Discussion
The discussion section of the paper evaluates the forecasting performance of ECMWF’s IFS HRES, GraphCast, and Pangu-Weather models in predicting near-surface temperature and wind speed extremes up to 10 days ahead. The analysis, based on three criteria—accuracy in extreme magnitude forecasting, grid-point extreme determination, and calibration of extreme quantiles—reveals that data-driven models generally outperform IFS HRES in most regions, particularly in the tropics. However, IFS HRES maintains an advantage in densely populated land areas and in forecasting cold extremes over short lead times. The data-driven models exhibit a pronounced decline in performance at higher latitudes, potentially due to training biases, and show a tendency to revert to climatology at longer lead times, a phenomenon referred to as blurring.
The findings suggest that while data-driven models can compete with physics-based models, they face challenges such as limited input variables and performance variability across regions and event types. The paper advocates for a hybrid forecasting approach that combines the strengths of both model types, particularly in regions where data-driven models excel. Future research should focus on refining loss functions, addressing input variable limitations, and exploring the impact of training methodologies on model performance, especially at higher latitudes. Overall, the study highlights the potential of data-driven models in operational forecasting while acknowledging the need for further evaluation and development.
