اكتشاف المعادلات المدفوعة بالبيانات لتوصيف غطاء السحاب
Data‐Driven Equation Discovery of a Cloud Cover Parameterization

المجلة: Journal of Advances in Modeling Earth Systems، المجلد: 16، العدد: 3
DOI: https://doi.org/10.1029/2023ms003763
تاريخ النشر: 2024-02-29
المؤلف: Arthur Grundner وآخرون
الموضوع الرئيسي: الظواهر الجوية والمحاكاة

نظرة عامة

تقدم الأبحاث نهجًا جديدًا لتحديد معلمات تغطية السحب في نماذج المناخ من خلال الاستفادة من تقنيات التعلم الآلي، وبشكل خاص الانحدار الرمزي، جنبًا إلى جنب مع الطرق التقليدية. يقترح المؤلفون إطار عمل نمذجة هرمي يدمج الانحدار الرمزي، واختيار الميزات التسلسلي، والقيود الفيزيائية لاشتقاق معادلات قابلة للتفسير لتغطية السحب من مخرجات نماذج العواصف العالمية ذات الدقة الخشنة. يعالج هذا النهج قيود الشبكات العصبية (NNs)، التي، على الرغم من أدائها العالي ضمن توزيعات تدريبها، غالبًا ما تفتقر إلى القابلية للتفسير وتتطلب أدوات إضافية لفهم توقعاتها.

تحقق المعادلة التحليلية الناتجة أداءً عاليًا (R² = 0.94) مع 11 معلمة قابلة للتدريب فقط، مما يظهر توازنًا بين الدقة والبساطة. إنها تتفوق على مخطط Xu-Randall عبر أنظمة السحب المختلفة، كما يتضح من مسافات هيلينجر التي تقل عن 0.09، وتتنافس مع أداء الشبكات العصبية في ظروف غنية بالمكثفات. علاوة على ذلك، عند ضبطها بدقة على بيانات إعادة التحليل ERA5، تظهر هذه المعادلة قابلية نقل متفوقة مقارنة بمخططات تغطية السحب الحالية. تؤكد النتائج على فعالية الانحدار الرمزي في توليد معادلات متسقة فيزيائيًا، وقابلة للتفسير، وغير خطية لتحديد معلمات تغطية السحب، مما يعزز موثوقية التوقعات المناخية.

مقدمة

في مقدمة هذه الورقة البحثية، يتناول المؤلفون قيود نماذج المناخ الحالية، وخاصة نموذج ICOsahedral Non-hydrostatic (ICON)، الذي يعمل عند دقة أفقية تتراوح بين 50-100 كم بسبب قيود الحوسبة. تتطلب هذه الدقة الخشنة تحديد معلمات العمليات دون الشبكة، مما يؤدي إلى انحيازات منهجية، خاصة في تحديد معلمات السحب. يعتمد مخطط تغطية السحب الحالي في ICON بشكل كبير على الرطوبة النسبية، مما قد يعيق تطوير معلمات أكثر دقة. يقترح المؤلفون الاستفادة من بيانات عالية الدقة وتقنيات التعلم الآلي المتقدمة لتحسين تحديد معلمات السحب والكونفكشن، مشيرين إلى أنه بينما تقدم طرق التعلم الآلي مثل الشبكات العصبية (NNs) مرونة، فإنها غالبًا ما تفتقر إلى القابلية للتفسير.

لمعالجة هذا التحدي، يقدم المؤلفون نهج نمذجة هرمي يجمع بين النماذج الفيزيائية التقليدية وتقنيات التعلم الآلي، وبشكل خاص الانحدار الرمزي. يسمح هذا الأسلوب بالاكتشاف التلقائي للمعادلات التحليلية لتغطية السحب التي تحافظ على الاتساق الفيزيائي بينما تنافس أداء الشبكات العصبية. تهدف الدراسة إلى استكشاف ما إذا كان من الممكن اشتقاق معادلة لتغطية السحب تكون قابلة للتفسير وفعالة، باستخدام الانحدار الرمزي لتخفيف الافتراضات السابقة حول الخطية والفصل. توضح الورقة نهجًا منظمًا لتقييم نماذج البيانات المختلفة، وأدائها، وقدرتها على التعميم عبر مجموعات بيانات مختلفة، ساعية في النهاية إلى تعزيز الفهم وتنفيذ تحديد معلمات تغطية السحب في نماذج المناخ.

النتائج

يقدم قسم “النتائج” في الورقة البحثية النتائج الرئيسية المستمدة من التجارب أو التحليلات التي تم إجراؤها. يوضح نتائج الدراسة، مع تسليط الضوء على نقاط البيانات والاتجاهات المهمة التي تم ملاحظتها. عادةً ما تكون النتائج مصحوبة بتحليلات إحصائية ذات صلة، بما في ذلك قيم p أو فترات الثقة، لدعم النتائج.

بالإضافة إلى ذلك، يتم الإشارة إلى أي تمثيلات رسومية، مثل الرسوم البيانية أو الجداول، لتوضيح البيانات بصريًا، مما يساعد في تفسير النتائج المعقدة. قد يناقش القسم أيضًا تداعيات النتائج فيما يتعلق بالفرضيات أو أسئلة البحث المطروحة سابقًا في الدراسة، مما يوفر أساسًا للمناقشات والاستنتاجات اللاحقة.

المناقشة

في هذا القسم، يوضح المؤلفون مجموعات البيانات المستخدمة لتدريب وتقييم مخططات تغطية السحب الخاصة بهم، مع التركيز بشكل خاص على محاكاة ICON لحل العواصف في مشروع DYAMOND وإعادة التحليل الجوي ERA5. توفر محاكاة DYAMOND، التي أجريت على مرحلتين (الصيف والشتاء)، بيانات عالية الدقة تم تقليلها لتتناسب مع دقة نماذج المناخ النموذجية. ينفذ المؤلفون نهجًا منهجيًا لضمان الاتساق في تمثيل تغطية السحب من خلال ضبط المخرجات بناءً على نسبة خلط مكثفات السحب. كما يتناولون عدم توازن الفئات من خلال تقليل عدد الخلايا الخالية من السحب وتقسيم البيانات إلى مجموعات تدريب وتحقق لتجنب الارتباطات.

لتقييم قابلية نقل نماذجهم، يستخدم المؤلفون بيانات إعادة التحليل ERA5، التي تقدم تحديات بسبب التباينات في توزيعات المتغيرات الرئيسية مثل ماء السحب والجليد. يقدم المؤلفون نهج نمذجة هرمي، يبدأ بنماذج قابلة للتفسير مثل الأشكال الخطية والمتعددة الحدود، قبل التقدم إلى الشبكات العصبية الأكثر تعقيدًا (NNs). يؤكدون على أهمية تطوير نماذج اقتصادية من خلال اختيار الميزات التسلسلي (SFS)، الذي يحسن أداء النموذج مع الحفاظ على القابلية للتفسير. يختتم القسم بمناقشة القيود الفيزيائية التي يجب أن تلبيها نماذج تغطية السحب، مما يضمن أن تظل المخرجات معقولة فيزيائيًا ومتوافقة مع المبادئ الجوية.

Journal: Journal of Advances in Modeling Earth Systems, Volume: 16, Issue: 3
DOI: https://doi.org/10.1029/2023ms003763
Publication Date: 2024-02-29
Author(s): Arthur Grundner et al.
Primary Topic: Meteorological Phenomena and Simulations

Overview

The research presents a novel approach to parameterizing cloud cover in climate models by leveraging machine learning techniques, specifically symbolic regression, alongside traditional methods. The authors propose a hierarchical modeling framework that integrates symbolic regression, sequential feature selection, and physical constraints to derive interpretable equations for cloud cover from coarse-grained outputs of global storm-resolving models. This approach addresses the limitations of neural networks (NNs), which, while high-performing within their training distributions, often lack interpretability and require additional tools for understanding their predictions.

The resulting analytical equation achieves a high performance (R² = 0.94) with only 11 trainable parameters, demonstrating a balance between accuracy and simplicity. It outperforms the Xu-Randall scheme across various cloud regimes, evidenced by Hellinger distances less than 0.09, and matches NN performance in condensate-rich conditions. Furthermore, when fine-tuned to the ERA5 reanalysis data, this equation exhibits superior transferability compared to existing cloud cover schemes. The findings underscore the efficacy of symbolic regression in generating physically-consistent, interpretable, and nonlinear equations for cloud cover parameterization, thus enhancing the reliability of climate projections.

Introduction

In the introduction of this research paper, the authors address the limitations of current climate models, particularly the ICOsahedral Non-hydrostatic (ICON) model, which operates at horizontal resolutions of 50-100 km due to computational constraints. This coarse resolution necessitates the parameterization of subgrid-scale processes, leading to systematic biases, especially in cloud parameterizations. The existing cloud cover scheme in ICON relies heavily on relative humidity, which may hinder the development of more accurate parameterizations. The authors propose leveraging high-fidelity data and advanced machine learning (ML) techniques to improve cloud and convection parameterizations, noting that while ML methods like neural networks (NNs) offer flexibility, they often lack interpretability.

To address this challenge, the authors introduce a hierarchical modeling approach that combines traditional physical models with ML techniques, specifically symbolic regression. This method allows for the automated discovery of analytical equations for cloud cover that maintain physical consistency while competing with the performance of NNs. The study aims to explore whether it is possible to derive a cloud cover equation that is both interpretable and effective, using symbolic regression to relax previous assumptions about linearity and separability. The paper outlines a structured approach to evaluate various data-driven models, their performance, and their ability to generalize across different datasets, ultimately seeking to enhance the understanding and implementation of cloud cover parameterizations in climate models.

Results

The “Results” section of the research paper presents the key findings derived from the conducted experiments or analyses. It details the outcomes of the study, highlighting significant data points and trends observed. The results are typically accompanied by relevant statistical analyses, including p-values or confidence intervals, to substantiate the findings.

Additionally, any graphical representations, such as charts or tables, are referenced to illustrate the data visually, aiding in the interpretation of complex results. The section may also discuss the implications of the findings in relation to the hypotheses or research questions posed earlier in the study, providing a foundation for subsequent discussions and conclusions.

Discussion

In this section, the authors detail the datasets utilized for training and benchmarking their cloud cover schemes, specifically focusing on the DYAMOND project’s storm-resolving ICON simulations and the ERA5 meteorological reanalysis. The DYAMOND simulations, conducted in two phases (Summer and Winter), provide high-resolution data that is coarse-grained to match typical climate model resolutions. The authors implement a systematic approach to ensure consistency in cloud cover representation by adjusting the output based on the cloud condensate mixing ratio. They also address class imbalance by undersampling predominantly cloud-free cells and splitting the data into training and validation sets to avoid correlations.

To evaluate the transferability of their models, the authors employ the ERA5 reanalysis data, which presents challenges due to discrepancies in the distributions of key variables like cloud water and ice. The authors introduce a hierarchical modeling approach, beginning with interpretable models such as linear and polynomial forms, before advancing to more complex neural networks (NNs). They emphasize the importance of developing parsimonious models through sequential feature selection (SFS), which optimizes model performance while maintaining interpretability. The section concludes with a discussion of the physical constraints that the cloud cover models must satisfy, ensuring that the outputs remain physically plausible and aligned with meteorological principles.