DOI: https://doi.org/10.1038/s43246-024-00731-w
تاريخ النشر: 2025-01-11
المؤلف: Kangming Li وآخرون
الموضوع الرئيسي: تعلم الآلة في علوم المواد
نظرة عامة
في هذا القسم، يقوم المؤلفون بالتحقيق في إمكانية تعميم نماذج التعلم الآلي (ML) في سياق علوم المواد، كاشفين أن التقييمات التقليدية يمكن أن تؤدي إلى استنتاجات مضللة. يقومون بتقييم أداء نماذج مختلفة، بما في ذلك الأشجار المعززة، في مهام خارج التوزيع (OOD) التي تتضمن كيمياء أو تناظرات هيكلية غير مرئية. تشير النتائج إلى أن العديد من المهام تؤدي بشكل جيد عبر نماذج مختلفة، لكن تحليل أعمق لمساحة تمثيل المواد يكشف أن معظم بيانات الاختبار تقع ضمن مناطق تم تمثيلها بشكل جيد بواسطة بيانات التدريب.
تجد الدراسة أن المهام التي تظهر أداءً ضعيفًا هي عادة تلك التي تقع خارج نطاق التدريب. ومن المثير للاهتمام أن زيادة حجم أو مدة التدريب لا تؤدي إلى التحسينات المتوقعة لهذه المهام الصعبة، مما يتعارض مع التوقعات التقليدية بشأن توسيع الشبكات العصبية. يجادل المؤلفون بأن العديد من تقييمات OOD تعكس الاستيفاء بدلاً من الاستقراء الحقيقي، مما يؤدي إلى تقدير مفرط لكل من إمكانية التعميم وفوائد التوسع. وهذا يبرز ضرورة وجود معايير OOD أكثر صرامة وتحديًا لتقييم أداء نماذج ML بدقة في علوم المواد.
طرق
في هذا القسم، يقوم المؤلفون بالتحقيق في قدرات التعميم لنماذج التعلم الآلي (ML) في سياق مهام خارج التوزيع (OOD)، مقترحين نهجًا جديدًا لتحليل تمثيلات المواد. يميزون بين نوعين من OOD: OOD إحصائي، حيث يختلف التوزيع الإحصائي للسمات عن مجموعة التدريب، وOOD تمثيلي، حيث تقع بيانات الاختبار خارج منطقة التدريب في مساحة تمثيل عالية الأبعاد. لتسهيل هذا التحليل، يستخدمون تقريب وتوقع متعدد الأشكال الموحد (UMAP) لإسقاط التضمينات عالية الأبعاد في مساحة ثنائية الأبعاد، مما يسمح بتقييم تقديرات كثافة النواة لبيانات التدريب والاختبار.
تكشف النتائج أن جزءًا كبيرًا من بعض الهياكل الاختبارية، مثل 24% من الهياكل المحتوية على المغنيسيوم و87% من الهياكل المحتوية على الأكسجين، هي OOD تمثيلي، مما يشير إلى انخفاض في درجة R². على العكس، يظهر التحليل أن 92% من بيانات الاختبار التي تحتوي على خمسة عناصر أو أكثر تقع ضمن المجال التمثيلي لبيانات التدريب، مما يشير إلى أن الاختبارات الحالية قد لا تتحدى قدرات تعميم نماذج ML بشكل كافٍ. يؤكد المؤلفون على أهمية تحديد المجالات التمثيلية بدقة لتفسير أداء التعميم ويبرزون أنه بينما يمكن أن تخبر تقديرات الكثافة عن أخطاء التنبؤ، إلا أنها لا تتوافق تمامًا معها. يُقترح إجراء أبحاث مستقبلية لتعزيز تحديد المجالات وتحسين العلاقة بين تقديرات الكثافة وأخطاء التنبؤ OOD.
نتائج
في هذا القسم، يحدد المؤلفون إعداد تقييمهم لتقييم نماذج التعلم الآلي (ML) في مهام خارج التوزيع (OOD)، مبتعدين عن مقاييس الأداء التقليدية داخل التوزيع. يعرفون مهام OOD على أنها تلك التي تختلف فيها التوزيعات الإحصائية لواحد أو أكثر من السمات بين مجموعات التدريب والاختبار، مشملين سيناريوهات متنوعة من الأدبيات. تم وضع ستة معايير لتعريف بيانات اختبار OOD، بما في ذلك المواد المميزة بعناصر معينة، ومجموعات الفضاء، ومجموعات النقاط، وأنظمة البلورات. تستخدم الدراسة نهج ترك واحد من X للخارج لجميع القيم الممكنة لـ X، مستبعدة تلك التي تحتوي على أقل من 200 عينة اختبار.
لتقييم أداء OOD، يستخدم المؤلفون ثلاث قواعد بيانات مواد مستمدة من ab initio: المستودع الآلي المشترك لمحاكاة متنوعة (JARVIS)، مشروع المواد (MP)، وقاعدة بيانات المواد الكمومية المفتوحة (OQMD). تسهل هذه القواعد البيانات، مع توزيعات بياناتها المختلفة، تحليلًا شاملاً عبر أكثر من 700 مهمة OOD. يتم تدريب مجموعة تمثيلية من نماذج ML، بما في ذلك نموذج الغابة العشوائية (RF) ونموذج XGBoost (XGB) باستخدام أوصاف Matminer، وشبكة عصبية واحدة مع توسيع متعدد الأقطاب Gaussian (GMP)، وشبكة عصبية خطية ذرية (ALIGNN)، ونموذج لغة كبير (LLM) يعتمد على أوصاف نصية بلورية (LLM-Prop). تهدف هذه المجموعة المتنوعة من النماذج وتمثيلات المدخلات إلى ضمان استنتاجات قوية وعامة بشأن أداء OOD.
مناقشة
في هذا القسم، يناقش المؤلفون قدرات التعميم خارج التوزيع (OOD) لنماذج التعلم الآلي المطبقة على بيانات طاقة التكوين، وهو جانب حاسم في علوم المواد. يقدمون تحليلًا مقارنًا لأداء نماذج مختلفة، بما في ذلك ALIGNN وXGB، باستخدام مقاييس مثل متوسط الخطأ المطلق (MAE) ومعامل التحديد ($R^2$). من الجدير بالذكر أن نموذج ALIGNN يظهر أداءً قويًا في OOD، محققًا درجات $R^2$ فوق 0.95 لـ 85% من المهام، على الرغم من عدم وجود بيانات تدريب حول الروابط للعناصر المستبعدة. يحدد المؤلفون التحيزات النظامية، خاصة في توقع طاقات التكوين للمواد غير المعدنية مثل الهيدروجين (H) والفلور (F) والأكسجين (O)، والتي تُعزى إلى الاختلافات التركيبية والهيكلية بين مجموعات التدريب والاختبار. يقترحون استخدام SHAP (SHapley Additive exPlanations) لتحليل هذه التحيزات ويقترحون أن تحسين قابلية النقل الكيميائي في النماذج قد يتطلب تعزيز تمثيل السمات العنصرية.
تسلط المناقشة الضوء أيضًا على الأداء المتفوق لمهام الهيكل مقارنة بالمهام الكيميائية، حيث يتفوق نموذج ALIGNN باستمرار على الآخرين. يستكشف المؤلفون أيضًا آثار قوانين توسيع الشبكات العصبية، كاشفين أنه بينما يؤدي زيادة أحجام النماذج ومجموعات البيانات عمومًا إلى فوائد في الأداء داخل التوزيع، إلا أنه يمكن أن يؤثر سلبًا على تعميم OOD. يؤكدون على الحاجة إلى إعادة تقييم تعريفات مهام OOD، مجادلين بأن العديد من المهام التي تعتبر صعبة قد لا تختبر حقًا قدرات تعميم النموذج. يدعو المؤلفون إلى تطوير معايير أكثر صرامة لتقييم وتعزيز إمكانية تعميم نماذج التعلم الآلي في علوم المواد بشكل أفضل.
DOI: https://doi.org/10.1038/s43246-024-00731-w
Publication Date: 2025-01-11
Author(s): Kangming Li et al.
Primary Topic: Machine Learning in Materials Science
Overview
In this section, the authors investigate the generalizability of scientific machine learning (ML) models within the context of materials science, revealing that traditional heuristic evaluations can lead to misleading conclusions. They specifically assess the performance of various models, including boosted trees, on out-of-distribution (OOD) tasks that involve unseen chemistry or structural symmetries. The results indicate that many tasks perform well across different models, but a deeper analysis of the materials representation space reveals that most test data lie within regions that are well-represented by the training data.
The study finds that tasks that exhibit poor performance are typically those that fall outside the training domain. Interestingly, increasing the size or duration of training does not yield the expected improvements for these challenging tasks, contradicting conventional expectations regarding neural scaling. The authors argue that many OOD evaluations reflect interpolation rather than true extrapolation, leading to an overestimation of both generalizability and the benefits of scaling. This underscores the necessity for more rigorous and challenging OOD benchmarks to accurately assess ML model performance in materials science.
Methods
In this section, the authors investigate the generalization capabilities of machine learning (ML) models in the context of out-of-distribution (OOD) tasks, proposing a novel approach to analyze materials representations. They differentiate between two types of OOD: statistically OOD, where the statistical distribution of attributes differs from the training set, and representationally OOD, where test data lie outside the training region in a high-dimensional representation space. To facilitate this analysis, they employ Uniform Manifold Approximation and Projection (UMAP) to project high-dimensional embeddings into a two-dimensional space, allowing for the evaluation of kernel density estimates for training and test data.
The findings reveal that a significant portion of certain test structures, such as 24% of Mg-containing structures and 87% of O-containing structures, are representationally OOD, indicating a low R² score. Conversely, the analysis demonstrates that 92% of test data with five or more elements fall within the representational domain of the training data, suggesting that existing tests may not adequately challenge ML models’ generalization capabilities. The authors emphasize the importance of accurately identifying representational domains to interpret generalization performance and highlight that while density estimates can inform about prediction errors, they do not perfectly correlate with them. Future research is suggested to enhance domain identification and improve the correlation between density estimates and OOD prediction errors.
Results
In this section, the authors outline their evaluation setup for assessing machine learning (ML) models on out-of-distribution (OOD) tasks, diverging from traditional in-distribution performance metrics. They define OOD tasks as those where the statistical distributions of one or more attributes differ between training and test sets, encompassing various scenarios from the literature. Six criteria are established for defining OOD test data, including materials characterized by specific elements, space groups, point groups, and crystal systems. The study employs a leave-one-X-out approach for all possible values of X, excluding those with fewer than 200 test samples.
To evaluate OOD performance, the authors utilize three ab initio-derived materials databases: the Joint Automated Repository for Various Integrated Simulations (JARVIS), the Materials Project (MP), and the Open Quantum Materials Database (OQMD). These databases, with their differing data distributions, facilitate a comprehensive analysis across over 700 OOD tasks. A representative set of ML models is trained, including random forest (RF) and XGBoost (XGB) models using Matminer descriptors, a single neural network with Gaussian multipole (GMP) expansion, an atomistic line graph neural network (ALIGNN), and a large language model (LLM) based on crystal text descriptions (LLM-Prop). This diverse selection of models and input representations aims to ensure robust and generalized conclusions regarding OOD performance.
Discussion
In this section, the authors discuss the out-of-distribution (OOD) generalization capabilities of machine learning models applied to formation energy data, a critical aspect of materials science. They present a comparative analysis of the performance of various models, including ALIGNN and XGB, using metrics such as mean absolute error (MAE) and coefficient of determination ($R^2$). Notably, the ALIGNN model demonstrates robust OOD performance, achieving $R^2$ scores above 0.95 for 85% of tasks, despite lacking training data on the bonding of excluded elements. The authors identify systematic biases, particularly in predicting formation energies for nonmetals like hydrogen (H), fluorine (F), and oxygen (O), which are attributed to compositional and structural differences between training and test sets. They propose using SHAP (SHapley Additive exPlanations) to analyze these biases and suggest that improving chemical transferability in models may require enhanced representation of elemental attributes.
The discussion further highlights the superior OOD performance of structure-based tasks compared to chemistry-based tasks, with the ALIGNN model consistently outperforming others. The authors also explore the implications of neural scaling laws, revealing that while increasing model and dataset sizes generally benefits in-distribution performance, it can adversely affect OOD generalization. They emphasize the need for a reevaluation of OOD task definitions, arguing that many tasks considered challenging may not genuinely test model generalization capabilities. The authors advocate for the development of more rigorous benchmarks to better assess and enhance the generalizability of machine learning models in materials science.
