التدريب المسبق لنموذج التعلم العميق القائم على الانتباه لمحاكاة الجزيئات
Pretraining of attention-based deep learning potential model for molecular simulation

شارك:
المجلة: npj Computational Materials، المجلد: 10، العدد: 1
DOI: https://doi.org/10.1038/s41524-024-01278-7
تاريخ النشر: 2024-05-07
المؤلف: Duo Zhang وآخرون
الموضوع الرئيسي: تعلم الآلة في علوم المواد

نظرة عامة

تقدم البحث DPA-1، نموذج Deep Potential جديد يتضمن آلية انتباه محكومة لتعزيز نمذجة سطح الطاقة المحتملة بين الذرات (PES) في المحاكاة الجزيئية. يستفيد هذا النموذج من بيانات هيكل إلكتروني واسعة، مما يسمح له بالتدريب المسبق على مجموعة بيانات شاملة تضم 56 عنصرًا، ومن ثم يتم ضبطه بدقة لمهام مختلفة مع جهد إضافي ضئيل. تظهر النتائج أن DPA-1 يتفوق على المعايير الحالية، مما يحسن بشكل كبير من كفاءة العينة عبر أنظمة ذرية مختلفة.

من الجدير بالذكر أن الدراسة تكشف أن معلمات تضمين النوع المتعلمة في DPA-1 تظهر تكوينًا حلزونيًا في الفضاء الكامن، والذي يرتبط بمواقع العناصر على الجدول الدوري. لا يسلط هذا الاكتشاف الضوء فقط على فعالية النموذج في التقاط التكوين والمساحات الكيميائية للأنظمة الذرية، بل يبرز أيضًا قابليته للتفسير، مما يوفر رؤى حول العلاقات الأساسية بين العناصر المختلفة.

الطرق

في هذا القسم، يصف المؤلفون المنهجية لنمذجة سطح الطاقة المحتملة (PES) لنظام يتكون من \( N \) ذرة، كل منها يتميز بأنواع عناصر \( A = \{ \alpha_1, \alpha_2, \ldots, \alpha_N \} \) وإحداثيات كارتيسية \( R = \{ r_1, r_2, \ldots, r_N \} \). يتم صياغة PES، المشار إليه بـ \( E = E(A, R) \)، كمجموع للطاقة الذرية \( e_i \)، حيث تعتمد كل طاقة ذرية \( e_i \) فقط على الذرات المجاورة ضمن نصف قطر قطع \( r_c \). يتم تعريف الجيران للذرة \( i \) بواسطة المجموعة \( N_{r_c}(i) \)، والتي تشمل المؤشرات \( j \) بحيث تكون المسافة الإقليدية \( r_{ji} < r_c \). يقدم المؤلفون DeepPot-SE، وهو تصور ثلاثي الأبعاد لـ PCA لتضمينات النوع المتعلمة من نموذج DPA-1 المدرب مسبقًا على مجموعة بيانات OC2M. يكشف هذا التصور عن ترتيب حلزوني لـ 56 عنصرًا في الفضاء الكامن، مع اتصالات بين العناصر في الفترة الرابعة موضحة بخطوط حمراء ومجموعات من العناصر ضمن نفس العائلة موضحة بخطوط زرقاء منقطة. يناقش القسم أيضًا خطأ الجذر المتوسط التربيعي (RMSE) للطاقة والقوة لأنظمة SSE باستخدام DPA-1، مع تسليط الضوء على استيفاء تضمينات النوع لليثيوم (Li)، الذي لا يوجد في OC2M. بالإضافة إلى ذلك، يتم تحليل كفاءة تدريب DPA-1 وDeepPot-SE مع زيادة عدد أنواع العناصر، مع اعتبار حد أقصى من 120 ذرة مجاورة في جميع التجارب. يتم تعريف القوة \( F_i \) المؤثرة على الذرة \( i \) على أنها التدرج السالب للطاقة الكلية بالنسبة لإحداثياتها، \( F_i = -\nabla_{r_i} E \). يؤكد المؤلفون على ضرورة أن يحافظ PES على عدم التغير تحت الترجمة، والدوران، والتبديل لأنواع الذرات المتطابقة، ويحددون بنية النموذج التي تسهل التنبؤ بالطاقة الذرية من خلال تقنيات حسابية متنوعة.

النتائج

في هذا القسم، يقدم المؤلفون نتائج تجاربهم التي تقيم أداء نموذج DPA-1، كما هو موضح في الشكل 1 ومفصل في قسم الطرق. تضمنت المرحلة الأولية تدريب DPA-1 من الصفر على أنظمة متنوعة لتقييم قدرته على نقل المعرفة عبر تركيبات مختلفة. بعد ذلك، تم اختبار النموذج على مجموعة بيانات AlMgCu لتقييم قابليته للنقل إلى الأنظمة الثلاثية بعد التدريب المسبق باستخدام بيانات أحادية العنصر وثنائية العنصر.

علاوة على ذلك، تم تدريب DPA-1 مسبقًا باستخدام مجموعة OC2M الفرعية من مجموعة بيانات OC20 وتم تطبيقه على مجموعة من المهام اللاحقة. لإظهار فعالية آلية تضمين النوع وآليات الانتباه المستخدمة في DPA-1، أجرى المؤلفون تحليلات مقارنة ضد نموذج DeepPot-SE عبر جميع التجارب. يختتم القسم بمقدمة موجزة حول مجموعات البيانات المستخدمة والمنهجيات التجريبية المنفذة.

المناقشة

في هذا القسم، يناقش المؤلفون تطوير وتقييم DPA-1، وهو نموذج Deep Potential قائم على الانتباه مصمم للتدريب المسبق على نطاق واسع على مجموعات بيانات ذرية. تم اختبار النموذج عبر أنظمة سبائك متنوعة، بما في ذلك AlMgCu، والكهارل الصلبة (SSE)، والسبائك عالية الإنتروبيا (HEA)، مما يظهر دقة تحقق متفوقة مقارنة بالنماذج الحالية مثل DeepPot-SE. من الجدير بالذكر أن DPA-1 حقق انخفاضًا ملحوظًا في خطأ الجذر المتوسط التربيعي (RMSE) للعينات الثلاثية، مما يدل على قدرته على تعلم التفاعلات الكامنة بفعالية من خلال مخطط تضمين النوع وآلية الانتباه. تم التحقق من أداء النموذج بشكل أكبر من خلال استرخاءات الهندسة ومحاكاة الديناميكا الجزيئية، مما أسفر عن نتائج تتماشى مع نظرية الكثافة الوظيفية (DFT) والبيانات التجريبية.

يؤكد المؤلفون على كفاءة عينة DPA-1، مما يظهر قدرته على تحقيق دقة عالية مع عدد أقل من عينات التدريب، خاصة عند التدريب المسبق على مجموعات بيانات متنوعة مثل OC2M. تتناقض هذه الكفاءة بشكل حاد مع DeepPot-SE، الذي يعاني من زيادة تربيعية مع عدد العناصر، مما يجعله مكلفًا حسابيًا. كان وقت تدريب نموذج DPA-1 أقل بكثير، حيث يتطلب أقل من 200 ساعة GPU، مع الحفاظ على مقاييس دقة تنافسية. بالإضافة إلى ذلك، تم تقييم قابلية تفسير تضمينات النوع المتعلمة من خلال تحليل المكونات الرئيسية (PCA)، مما يكشف عن علاقات كيميائية ذات مغزى بين العناصر. يقترح المؤلفون اتجاهات مستقبلية لتوسيع مجموعة بيانات التدريب واستكشاف التعلم متعدد المهام، وضغط النموذج، والتعلم الانتقالي، بهدف تعزيز قابلية تطبيق النموذج عبر مجموعة أوسع من المحاكاة الذرية.

Journal: npj Computational Materials, Volume: 10, Issue: 1
DOI: https://doi.org/10.1038/s41524-024-01278-7
Publication Date: 2024-05-07
Author(s): Duo Zhang et al.
Primary Topic: Machine Learning in Materials Science

Overview

The research presents DPA-1, a novel Deep Potential model that incorporates a gated attention mechanism to enhance the modeling of the inter-atomic potential energy surface (PES) in molecular simulations. This model leverages extensive electronic structure data, allowing it to be pretrained on a comprehensive dataset of 56 elements and subsequently fine-tuned for various downstream tasks with minimal additional effort. The results demonstrate that DPA-1 outperforms existing benchmarks, significantly improving sample efficiency across different atomic systems.

Notably, the study reveals that the learned type embedding parameters of DPA-1 exhibit a spiral configuration in the latent space, which correlates with the elements’ positions on the periodic table. This finding not only underscores the model’s effectiveness in capturing the conformation and chemical spaces of atomic systems but also highlights its interpretability, providing insights into the underlying relationships among different elements.

Methods

In this section, the authors describe the methodology for modeling the potential energy surface (PES) of a system comprising \( N \) atoms, each characterized by elemental types \( A = \{ \alpha_1, \alpha_2, \ldots, \alpha_N \} \) and Cartesian coordinates \( R = \{ r_1, r_2, \ldots, r_N \} \). The PES, denoted as \( E = E(A, R) \), is formulated as a summation of atomic energies \( e_i \), where each atomic energy \( e_i \) is dependent solely on the neighboring atoms within a cutoff radius \( r_c \). The neighbors of atom \( i \) are defined by the set \( N_{r_c}(i) \), which includes indices \( j \) such that the Euclidean distance \( r_{ji} < r_c \). The authors introduce DeepPot-SE, a three-dimensional PCA visualization of learned type embeddings from the DPA-1 model pretrained on the OC2M dataset. This visualization reveals a spiral arrangement of 56 elements in latent space, with connections among elements in the fourth period indicated by red lines and groupings of elements within the same family shown by blue dotted lines. The section also discusses the root mean square error (RMSE) of energy and force for SSE systems using DPA-1, highlighting the interpolation of type embeddings for lithium (Li), which is absent from OC2M. Additionally, the training efficiency of DPA-1 and DeepPot-SE is analyzed as the number of element types increases, with a maximum of 120 neighboring atoms considered in all experiments. The force \( F_i \) acting on atom \( i \) is defined as the negative gradient of the total energy with respect to its coordinates, \( F_i = -\nabla_{r_i} E \). The authors emphasize the necessity for the PES to maintain invariance under translation, rotation, and permutation of identical atom types, and outline the model architecture that facilitates the prediction of atomic energies through various computational techniques.

Results

In this section, the authors present the results of their experiments evaluating the performance of the DPA-1 model, as illustrated in Figure 1 and detailed in the Methods section. The initial phase involved training DPA-1 from scratch on various systems to assess its capacity for transferring knowledge across different compositions. Subsequently, the model was tested on an AlMgCu dataset to evaluate its transferability to ternary systems after pretraining with single-element and binary data.

Furthermore, DPA-1 was pretrained using the OC2M subset from the OC20 dataset and applied to a range of downstream tasks. To demonstrate the effectiveness of the type-embedding and attention mechanisms employed in DPA-1, the authors conducted comparative analyses against the DeepPot-SE model across all experiments. The section concludes with a brief introduction to the datasets utilized and the experimental methodologies implemented.

Discussion

In this section, the authors discuss the development and evaluation of DPA-1, an attention-based Deep Potential model designed for large-scale pretraining on atomistic datasets. The model was tested across various alloy systems, including AlMgCu, solid-state electrolytes (SSE), and high-entropy alloys (HEA), demonstrating superior validation accuracy compared to existing models like DeepPot-SE. Notably, DPA-1 achieved a remarkable reduction in root mean square error (RMSE) for ternary samples, indicating its capability to learn latent interactions effectively through a type-embedding scheme and attention mechanism. The model’s performance was further validated through geometry relaxations and molecular dynamics simulations, yielding results consistent with density functional theory (DFT) and experimental data.

The authors emphasize DPA-1’s sample efficiency, showcasing its ability to achieve high accuracy with fewer training samples, particularly when pretrained on diverse datasets like OC2M. This efficiency contrasts sharply with DeepPot-SE, which suffers from quadratic scaling with the number of elements, making it computationally prohibitive. The DPA-1 model’s training time was significantly lower, requiring less than 200 GPU hours, while maintaining competitive accuracy metrics. Additionally, the interpretability of the learned type embeddings was assessed through principal component analysis (PCA), revealing meaningful chemical relationships among elements. The authors propose future directions for extending the training dataset and exploring multi-task learning, model compression, and transfer learning, aiming to enhance the model’s applicability across a broader range of atomistic simulations.

شارك: