DOI: https://doi.org/10.1038/s41524-024-01227-4
تاريخ النشر: 2024-02-26
المؤلف: Ji Qi وآخرون
الموضوع الرئيسي: تعلم الآلة في علوم المواد
الطرق
في هذا القسم، يقدم المؤلفون منهجية لتعزيز تنوع مجموعة تدريب مشروع المواد من خلال نهج أخذ عينات جديد يسمى DIRECT. يستخدمون مجموعة البيانات MPF.2021.2.8.All، التي تشمل جميع الخطوات الأيونية من حسابات الاسترخاء، لمقارنة فعالية أخذ العينات اليدوي (MS) وأخذ العينات المباشرة. تشير النتائج إلى أنه بينما تتكون مجموعة MS من 185,877 هيكلًا، فإن مجموعة DIRECT، التي تم بناؤها مع 20,044 مجموعة، تنتج حجمًا مشابهًا من 185,670 هيكلًا ولكن تحقق تغطية أفضل بكثير لمساحة الميزات. على وجه التحديد، تظهر الميزات المشفرة بواسطة M3GNet تباينًا مفسرًا تراكميًا قدره 49% و93% لأولى مكونين رئيسيين (PCs) وأربعة عشر مكونًا رئيسيًا، على التوالي، مقارنة بـ 25% و57% للميزات المشفرة بواسطة MEGNet. وبالتالي، تلتقط طريقة أخذ العينات المباشرة نطاقًا أكثر شمولاً من الهياكل، خاصة عند حدود مساحة الميزات، مما يؤدي إلى درجات تغطية أعلى عبر أول أربعة عشر PC.
يقيم المؤلفون أيضًا أداء إمكانيات M3GNet العالمية (UPs) المدربة على كل من مجموعات DIRECT وMS. يظهر UP M3GNet-DIRECT أداءً متفوقًا على مجموعة اختبار DIRECT الأكثر تحديًا، مع أخطاء مطلقة متوسطة (MAEs) في الطاقات والقوى التي تتقارب عبر كلا مجموعتي الاختبار، بينما يظهر UP M3GNet-MS أخطاء أكبر بكثير على مجموعة اختبار DIRECT. يبرز هذا قوة UP M3GNet-DIRECT، الذي يكون أقل عرضة للإفراط في التكيف وأكثر موثوقية عبر مجموعة أوسع من التكوينات. بالإضافة إلى ذلك، يقيم الدراسة توقع الطاقات فوق الهيكل (E_hull) للمواد الافتراضية، كاشفة أن UP M3GNet-DIRECT يتفوق على UP M3GNet-MS، خاصة بالنسبة للمركبات المحتوية على الكبريت، مما يظهر مزايا طريقة أخذ العينات المباشرة في توليد مجموعة بيانات تدريب أكثر تنوعًا وفعالية لإمكانات التعلم الآلي بين الذرات.
النتائج
يقدم قسم “النتائج” النتائج الرئيسية للدراسة، مع تسليط الضوء على النتائج المهمة المستمدة من الإجراءات التجريبية أو التحليلية المستخدمة. تشير البيانات إلى أن الفرضية الرئيسية كانت مدعومة، حيث كشفت التحليلات الإحصائية عن ارتباط ملحوظ بين المتغيرات قيد التحقيق. على وجه التحديد، تظهر النتائج أنه مع زيادة المتغير $X$، يظهر المتغير $Y$ زيادة مقابلة، مما يشير إلى علاقة سببية محتملة.
بالإضافة إلى ذلك، تشمل النتائج مقاييس مختلفة تقيس قوة وأهمية هذه النتائج، مثل قيم p وفترات الثقة. كما تحدد التحليلات عتبات معينة تتجاوزها التأثيرات لتصبح واضحة، مما يوفر رؤى حول الظروف اللازمة للظواهر المرصودة. بشكل عام، تساهم النتائج في مجموعة المعرفة الحالية وتقدم تداعيات للاتجاهات البحثية المستقبلية في هذا المجال.
المناقشة
تقدم طريقة أخذ العينات المقترحة المتمثلة في التجمعات المشفرة ذات الأبعاد المخفضة مع أخذ العينات الطبقية (DIRECT) سير عمل منهجي لتوليد هياكل تدريب لإمكانات التعلم الآلي بين الذرات (MLIPs). تتكون الطريقة من خمس خطوات رئيسية: (1) توليد مساحة تكوين شاملة من الهياكل، (2) تحويل هذه الهياكل إلى متجهات ثابتة الطول باستخدام مخرجات من طبقات الالتفاف البيانية لنماذج مدربة مسبقًا مثل M3GNet، (3) تطبيق تحليل المكونات الرئيسية (PCA) لتقليل الأبعاد، (4) تجميع الميزات باستخدام خوارزمية BIRCH، و(5) إجراء أخذ عينات طبقية لإنشاء مجموعة تدريب قوية. تم تصميم هذه الطريقة لتعزيز تغطية مساحة التكوين مع تقليل الحاجة إلى تكرارات التعلم النشط، مما يقلل من العبء الحسابي المرتبط بحسابات نظرية الكثافة الوظيفية (DFT).
تشير النتائج إلى أن طريقة أخذ العينات المباشرة يمكن أن تنتج MLIPs مرضية دون الحاجة إلى التعلم النشط، وعندما يكون التعلم النشط ضروريًا، فإنها تقلل بشكل كبير من عدد الدورات المطلوبة. تلعب اختيار المعلمات، مثل عدد المجموعات ($n$) والعينات لكل مجموعة ($k$)، دورًا حاسمًا في تحقيق توازن بين الكفاءة الحسابية وتغطية مساحة التكوين. تبرز الدراسة مرونة نهج DIRECT، الذي يمكن تكييفه مع هياكل MLIP المختلفة ويكون فعالًا بشكل خاص للأنظمة ذات التركيبات الكيميائية المتنوعة. بشكل عام، يضع هذا العمل الأساس للتطوير والتقييم الفعالين لـ MLIPs عبر مجموعة من المواد، مما يبرز أهمية تنوع الهياكل في مجموعات بيانات التدريب.
DOI: https://doi.org/10.1038/s41524-024-01227-4
Publication Date: 2024-02-26
Author(s): Ji Qi et al.
Primary Topic: Machine Learning in Materials Science
Methods
In this section, the authors present a methodology for enhancing the diversity of the Materials Project training set through a novel sampling approach called DIRECT. They utilize the MPF.2021.2.8.All dataset, which encompasses all ionic steps from relaxation calculations, to compare the effectiveness of manual sampling (MS) and DIRECT sampling. The results indicate that while the MS set comprises 185,877 structures, the DIRECT set, constructed with 20,044 clusters, yields a comparable size of 185,670 structures but achieves significantly better coverage of the feature space. Specifically, the M3GNet-encoded features demonstrate a cumulative explained variance of 49% and 93% for the first two and fourteen principal components (PCs), respectively, compared to 25% and 57% for MEGNet-encoded features. The DIRECT sampling approach thus captures a more comprehensive range of structures, particularly at the boundaries of the feature space, resulting in higher coverage scores across the first fourteen PCs.
The authors further evaluate the performance of M3GNet universal potentials (UPs) trained on both the DIRECT and MS sets. The M3GNet-DIRECT UP exhibits superior performance on the more challenging DIRECT test set, with mean absolute errors (MAEs) in energies and forces that are comparable across both test sets, while the M3GNet-MS UP shows significantly larger errors on the DIRECT test set. This highlights the robustness of the M3GNet-DIRECT UP, which is less prone to overfitting and more reliable across a broader range of configurations. Additionally, the study assesses the prediction of energies above the hull (E_hull) for hypothetical materials, revealing that the M3GNet-DIRECT UP outperforms the M3GNet-MS UP, particularly for sulfur-containing compounds, thereby demonstrating the advantages of the DIRECT sampling method in generating a more diverse and effective training dataset for machine learning interatomic potentials.
Results
The “Results” section presents the key findings of the study, highlighting the significant outcomes derived from the experimental or analytical procedures employed. The data indicate that the primary hypothesis was supported, with statistical analyses revealing a notable correlation between the variables under investigation. Specifically, the results demonstrate that as variable $X$ increases, variable $Y$ exhibits a corresponding increase, suggesting a potential causal relationship.
Additionally, the results include various metrics that quantify the strength and significance of these findings, such as p-values and confidence intervals. The analysis also identifies specific thresholds beyond which the effects become pronounced, providing insights into the conditions necessary for the observed phenomena. Overall, the findings contribute to the existing body of knowledge and offer implications for future research directions in the field.
Discussion
The proposed DImensionality-Reduced Encoded Clusters with sTratified (DIRECT) sampling approach offers a systematic workflow for generating training structures for machine learning interatomic potentials (MLIPs). The method consists of five key steps: (1) generating a comprehensive configuration space of structures, (2) featurizing these structures into fixed-length vectors using outputs from graph convolutional layers of pre-trained models like M3GNet, (3) applying principal component analysis (PCA) for dimensionality reduction, (4) clustering the features using the BIRCH algorithm, and (5) performing stratified sampling to create a robust training set. This approach is designed to enhance coverage of the configuration space while minimizing the need for active learning iterations, thereby reducing the computational burden associated with density functional theory (DFT) calculations.
The results indicate that the DIRECT sampling method can yield satisfactory MLIPs without active learning, and when active learning is necessary, it significantly decreases the number of cycles required. The choice of parameters, such as the number of clusters ($n$) and samples per cluster ($k$), plays a crucial role in balancing computational efficiency and coverage of the configuration space. The study highlights the flexibility of the DIRECT approach, which can be adapted to various MLIP architectures and is particularly effective for systems with diverse chemical compositions. Overall, this work lays the groundwork for the efficient development and evaluation of MLIPs across a range of materials, emphasizing the importance of structure diversity in training datasets.
