محاكي قابل للتفسير ومستنير بالفيزياء لطيف الطاقة المادية الخطي من التعلم الآلي
Interpretable and physics-informed emulator for the linear matter power spectrum from machine learning

شارك:
المجلة: Physical review. D/Physical review. D.، المجلد: 113، العدد: 6
DOI: https://doi.org/10.1103/hp65-mcrz
تاريخ النشر: 2026-02-17
المؤلف: J. Bayron Orjuela-Quintana وآخرون
الموضوع الرئيسي: الفيزياء الحاسوبية وتطبيقات بايثون

نظرة عامة

تقدم هذه الدراسة محاكيًا قابلًا للتفسير لطيف قوة المادة الخطي (MPS) ضمن النموذج الكوني القياسي، ΛCDM، باستخدام إطار عمل الانحدار الرمزي المدعوم بالفيزياء المعزز بالخوارزميات الجينية. ينتج المحاكي تقريبًا مغلق الشكل وسلس لطيف MPS الذي يحقق دقة أقل من واحد بالمئة عبر مجموعة واسعة من المقاييس ($k \in [10^{-5}, 1.5] \, h \, \text{Mpc}^{-1}$)، مع خطأ نسبي مطلق متوسط (MAPE) يبلغ حوالي 0.42%. من خلال دمج تصحيحات تذبذبية تتعلق بتذبذبات الباريون الصوتية (BAO)، يحافظ المحاكي على دقة عالية للطيف المحسوب عبر حلول بولتزمان، متفوقًا على صيغ التناسب التقليدية مثل إيزنشتاين-هو.

علاوة على ذلك، فإن الإطار قابل للتكيف مع نماذج الجاذبية المعدلة، كما يتضح من تطبيقه على نموذج جاذبية $f(R)$، حيث يلتقط تعديل MPS مع أخطاء متوسطة تتراوح بين 1.5-1.8%. إن قابلية تفسير المحاكي هي ميزة كبيرة، حيث يتم تأسيس كل مصطلح في التعبير النهائي على مبادئ فيزيائية راسخة، مما يسهل التشخيصات والتوسعات النظرية. يخطط المؤلفون لإصدار تنفيذ سهل الاستخدام لهذا المحاكي، مما سيمكن المستخدمين من حساب MPS الخطي بكفاءة مع دقة عالية بناءً على المعلمات الكونية، مما يعزز من فائدته في تحليلات الهيكل الكبير واستنتاج المعلمات.

مقدمة

تؤكد مقدمة الورقة على أهمية طيف قوة المادة (MPS)، الذي يُشار إليه بـ \( P(k) \)، في علم الكون الحديث، حيث يعمل كأداة حاسمة لفهم توزيع تقلبات كثافة المادة في الكون المبكر. يعد MPS ضروريًا للتنبؤ بمختلف الملاحظات الكونية، بما في ذلك تجمعات المجرات وتذبذبات الباريون الصوتية (BAO)، التي تعمل كمسطرة كونية قياسية لقياس تاريخ توسع الكون. تعتمد الاستطلاعات الحالية والمستقبلية على توقعات دقيقة لـ \( P(k) \) للتحقيق في القطاع المظلم والانحرافات المحتملة عن نموذج \(\Lambda\)CDM القياسي.

لمعالجة التحديات الحسابية المرتبطة بالتقييمات عالية الدقة لـ \( P(k) \) في خطوط استنتاج واسعة النطاق، يقترح المؤلفون نهجًا جديدًا باستخدام الانحدار الرمزي (SR) مع الخوارزميات الجينية (GAs). يهدف هذا الأسلوب إلى اشتقاق تعبيرات رياضية قابلة للتفسير لطيف MPS الخطي مع دمج أولويات فيزيائية لتعزيز الشفافية وتقليل الإفراط في التناسب. توضح الورقة المنهجية لبناء هذه التعبيرات، مما يسمح بنمذجة الانحرافات عن إطار \(\Lambda\)CDM، خاصة في سياق سيناريوهات الجاذبية المعدلة. ستفصل الأقسام اللاحقة من الورقة الخلفية النظرية، وتوليد البيانات، وأداء الصيغ الرمزية المقترحة مقارنة بالنماذج الحالية.

مناقشة

تتناول قسم المناقشة في الورقة تطبيق البرمجة الجينية (GP) والخوارزميات الجينية (GAs) للانحدار الرمزي (SR) في علم الكون، مع التركيز بشكل خاص على نمذجة طيف قوة المادة الخطي (MPS). تستلهم GP من التطور البيولوجي، حيث تقوم بتحسين التعبيرات الرمزية من خلال عمليات مثل الاختيار، والتقاطع، والطفرات، بهدف اشتقاق علاقات رياضية قابلة للتفسير من البيانات. يؤكد المؤلفون على أهمية دمج المعرفة الميدانية في نهجهم القائم على SR باستخدام GAs، مما يقيد مساحة البحث إلى تعبيرات ذات معنى فيزيائي، وبالتالي يعزز من قابلية التفسير ويقلل من الإفراط في التناسب.

يقدم المؤلفون محاكيًا شبه تحليلي لـ MPS، يتم بناؤه عن طريق فصل دالة نقل المادة (MTS) إلى مكونات سلسة وتذبذبية. يستخدمون مجموعة بيانات تم توليدها من محاكيات كونية لتدريب نموذجهم، محققين تعبيرًا رمزيًا مضغوطًا يلتقط الميزات الأساسية لـ MPS بدقة عالية. يظهر النموذج النهائي خطأ نسبي مطلق متوسط (MAPE) يبلغ 0.99% عبر مجموعة من المعلمات الكونية، متفوقًا على صيغ التناسب التقليدية مع الحفاظ على هيكل أبسط. يسلط هذا العمل الضوء على إمكانيات التعبيرات الرمزية المستمدة من GAs كأدوات تحليلية قوية في النمذجة الكونية، مما يوفر كل من الدقة وقابلية التفسير.

Journal: Physical review. D/Physical review. D., Volume: 113, Issue: 6
DOI: https://doi.org/10.1103/hp65-mcrz
Publication Date: 2026-02-17
Author(s): J. Bayron Orjuela-Quintana et al.
Primary Topic: Computational Physics and Python Applications

Overview

This research presents an interpretable emulator for the linear matter power spectrum (MPS) within the standard cosmological model, ΛCDM, utilizing a physics-informed symbolic regression framework enhanced by genetic algorithms. The emulator generates closed-form, smooth approximations of the MPS that achieve sub-percent accuracy across a wide range of scales ($k \in [10^{-5}, 1.5] \, h \, \text{Mpc}^{-1}$), with a mean absolute percentage error (MAPE) of approximately 0.42%. By incorporating oscillatory corrections related to baryon acoustic oscillations (BAO), the emulator maintains a high fidelity to spectra computed via Boltzmann solvers, outperforming traditional fitting formulas like Eisenstein-Hu.

Furthermore, the framework is adaptable to modified gravity models, exemplified by its application to an $f(R)$ gravity model, where it captures the modulation of the MPS with average errors of 1.5-1.8%. The emulator’s interpretability is a significant advantage, as each term in the final expression is grounded in established physical principles, facilitating diagnostics and theoretical extensions. The authors plan to release a user-friendly implementation of this emulator, which will enable users to efficiently compute the linear MPS with high accuracy based on cosmological parameters, thereby enhancing its utility in large-scale structure analyses and parameter inference.

Introduction

The introduction of the paper emphasizes the significance of the matter power spectrum (MPS), denoted as \( P(k) \), in modern cosmology, where it serves as a crucial tool for understanding the distribution of matter density fluctuations in the early universe. The MPS is integral for predicting various cosmological observables, including galaxy clustering and baryon acoustic oscillations (BAO), which act as a cosmic standard ruler for measuring the universe’s expansion history. Current and future surveys depend on accurate predictions of \( P(k) \) to investigate the dark sector and potential deviations from the standard \(\Lambda\)CDM model.

To address the computational challenges associated with high-precision evaluations of \( P(k) \) in large-scale inference pipelines, the authors propose a novel approach using symbolic regression (SR) combined with genetic algorithms (GAs). This method aims to derive interpretable mathematical expressions for the linear MPS while incorporating physical priors to enhance transparency and reduce overfitting. The paper outlines the methodology for constructing these expressions, which allows for the modeling of deviations from the \(\Lambda\)CDM framework, particularly in the context of modified gravity scenarios. The subsequent sections of the paper will detail the theoretical background, data generation, and the performance of the proposed symbolic formulations compared to existing models.

Discussion

The discussion section of the paper elaborates on the application of genetic programming (GP) and genetic algorithms (GAs) for symbolic regression (SR) in cosmology, particularly focusing on the modeling of the linear matter power spectrum (MPS). GP, inspired by biological evolution, optimizes symbolic expressions through processes such as selection, crossover, and mutation, aiming to derive interpretable mathematical relationships from data. The authors emphasize the importance of incorporating domain knowledge into their GA-based SR approach, which restricts the search space to physically meaningful expressions, thus enhancing interpretability and reducing overfitting.

The authors present a semi-analytical emulator for the MPS, which is constructed by separating the matter transfer function (MTS) into smooth and oscillatory components. They utilize a dataset generated from cosmological simulations to train their model, achieving a compact symbolic expression that captures the essential features of the MPS with high accuracy. The final model demonstrates a mean absolute percentage error (MAPE) of 0.99% across a range of cosmological parameters, outperforming traditional fitting formulas while maintaining a simpler structure. This work highlights the potential of GA-derived symbolic expressions as robust analytical tools in cosmological modeling, providing both accuracy and interpretability.

شارك: