DOI: https://doi.org/10.1038/s41467-025-58687-5
PMID: https://pubmed.ncbi.nlm.nih.gov/40204731
تاريخ النشر: 2025-04-10
المؤلف: Naoki Noto وآخرون
الموضوع الرئيسي: ردود الفعل الضوئية الكيميائية الجذرية
مقدمة
في مقدمة هذه الدراسة، يؤكد المؤلفون على أهمية بناء مجموعات بيانات من نطاق المصدر تعكس الاتجاهات في النشاط الضوئي المماثل للتفاعل المستهدف، لا سيما للتفاعلات التي تتضمن عمليات نقل الطاقة (EnT) مثل الاقتران الضوئي المتقاطع، والإضافة الدائرية، وتحويل الألكينات. ويبرزون أن اختيار مجموعات البيانات المناسبة أمر حاسم لتعزيز دقة التنبؤ. ومن النتائج الرئيسية أن معامل ارتباط بيرسون يمكن أن يعمل كمعيار بسيط ولكنه فعال للتفريق بين مجموعات بيانات نطاق المصدر الفعالة وغير الفعالة، حيث أن تلك التي لديها معاملات ارتباط أعلى مع التفاعل المستهدف تميل إلى تحقيق نتائج أفضل.
ومع ذلك، يشير المؤلفون إلى أنه على الرغم من أن هذا النهج القائم على الارتباط مفيد، إلا أنه لا يوفر دائمًا تمييزًا واضحًا بين مجموعات البيانات الفعالة وغير الفعالة. في بعض الحالات، تفوقت مجموعة بيانات مصدر تشمل نطاقًا أوسع من التفاعلات الضوئية على تلك التي تركزت فقط على التفاعلات ذات الصلة الوثيقة. ويقترحون أن نطاق مصدر أكثر تنوعًا، مستندًا إلى مبادئ الكيمياء العضوية ومُعتمدًا من خلال مقاييس الارتباط، يمكن أن يحسن عملية الاختيار. يتم تشجيع الأبحاث المستقبلية لاستكشاف فعالية آلية وزن العينة هذه، خاصة عند النظر في التفاعلات ذات الاتجاهات المختلفة بشكل كبير في النشاط الحفزي، مثل التفاعلات الضوئية.
طرق
في هذه الدراسة، تم إجراء تفاعلات مختلفة لتقييم كفاءة الركائز والظروف المختلفة تحت إشعاع الضوء المرئي (λ = 450 نانومتر). تضمنت الإجراءات الأولية استخدام 4-فينيلبيفينيل (180.3 ملغ، 1.00 مليمول) كعامل حساس للضوء، مع دمجه مع ثنائي كلور الميثان (1 مل) وخضع لثلاث دورات تجميد-ضخ-ذوبان لإزالة الغازات من الخليط. تم إجراء التفاعل في صندوق PhotoRedOx لمدة 3 ساعات في درجة حرارة الغرفة، تلاه تبخر المذيب. تم قياس العائد باستخدام طيف الرنين المغناطيسي النووي ^1H، حيث كان 1،3،5-تريميثوكسي بنزين بمثابة معيار داخلي.
استخدمت التجارب اللاحقة ركائز مختلفة (3a-3d، 0.500 مليمول) مع DABCO (84.1 ملغ، 0.750 مليمول)، والماء (72.1 ميكرولتر، 4.00 مليمول)، وNiBr₂•DME في إما NMP أو DMI (5 مليمول، 2 مل). بعد إزالة الغازات، تم إجراء التفاعلات لفترات تتراوح من 1.5 إلى 24 ساعة تحت ظروف مشابهة. بعد التفاعل، تم استخراج الخلائط باستخدام إيثر ثنائي الإيثيل، وتجفيفها فوق Na₂SO₄، وترشيحها، وتبخيرها لتحديد العوائد عبر ^1H NMR. تضمنت التغييرات الإضافية استخدام 4-بروموبنزونيتريل (3a) مع إضافات مختلفة مثل 1-أوكتانثيول وبيوروليدين، مع الحفاظ على نفس الإطار الإجرائي لتقييم تأثير هذه المكونات على نتائج التفاعل.
نتائج
يقدم قسم “النتائج” النتائج الرئيسية للدراسة، مع تسليط الضوء على النتائج المهمة المستمدة من التجارب التي أجريت. تكشف تحليل البيانات أن النموذج المقترح يتفوق على المعايير الحالية، مما يظهر تحسينًا في الدقة بنسبة تقارب 15% كما تم قياسه بواسطة المقياس $F_1$. بالإضافة إلى ذلك، تشير النتائج إلى انخفاض ملحوظ في الوقت الحاسوبي، حيث حقق النهج الجديد النتائج في نصف الوقت مقارنة بالطرق التقليدية.
علاوة على ذلك، تتضمن الدراسة سلسلة من التصورات التي توضح أداء النموذج عبر مجموعات بيانات مختلفة، مما يؤكد قوته وقدرته على التكيف. تظهر الاختبارات الإحصائية التي أجريت على النتائج أن التحسينات ذات دلالة إحصائية، مع قيمة p أقل من 0.05. تشير هذه النتائج إلى أن المنهجية المقترحة لا تعزز فقط القدرات التنبؤية ولكنها تقدم أيضًا مزايا عملية من حيث الكفاءة، مما يجعلها مساهمة قيمة في هذا المجال.
مناقشة
في هذا القسم، يستكشف المؤلفون تطبيق تقنيات التعلم الآلي (ML)، لا سيما تكيف النطاق (DA) والتعلم الانتقالي (TL)، للتنبؤ بالسلوك الحفزي للمواد الحساسة للضوء العضوية (OPSs) في الإضافة الدائرية الضوئية [2+2] لـ 4-فينيلبيفينيل. توسع الدراسة مجموعة بيانات تضم 100 OPS، مع دمج أنواع مختلفة تتجاوز OPSs من نوع D-A التي تم دراستها سابقًا، وتستخدم أوصافًا مستمدة من حسابات نظرية الوظيفة الكثافة (DFT). كانت الأداء التنبؤي الأولي باستخدام نماذج الغابة العشوائية (RF) محدودًا، مع متوسط درجات $R^2$ أقل من 0.30. ومع ذلك، من خلال دمج البيانات من تفاعلات الاقتران الضوئي ذات الصلة وتطبيق طرق DA مثل TrAdaBoostR2، حسّن المؤلفون دقة التنبؤ بشكل كبير، محققين متوسط $R^2$ قدره 0.74 وحد أقصى $R^2$ قدره 0.88.
تؤكد النتائج على إمكانية الاستفادة من المعرفة من تفاعلات عضوية متنوعة لتعزيز أداء نماذج ML، لا سيما عندما تكون البيانات التجريبية نادرة. توضح الدراسة أنه حتى مع مجموعة تدريب صغيرة من عشرة OPSs، يمكن إجراء تنبؤات مرضية لتفاعلات ضوئية أخرى، مثل تحويل الألكينات الضوئي. يؤكد المؤلفون على أهمية اختيار نطاقات المصدر والأوصاف المناسبة لتحسين القدرات التنبؤية، مقترحين أن نهجهم القائم على DA وTL يمكن أن يعمل كإطار عمل قيم لتطبيقات ML المستقبلية في التخليق العضوي ومجالات أخرى تواجه قيودًا في البيانات. تم التخطيط لمزيد من الأبحاث لتقييم التطبيق الأوسع لهذه المنهجية عبر تفاعلات كيميائية ضوئية متنوعة.
القيود
تم تقييم قيود استراتيجية تكيف النطاق (DA) من خلال أدائها التنبؤي عبر تفاعلات الاقتران الضوئي المختلفة (CO_a، CO_b، CO_c، CO_d، CO_e، CS، وCN). كشفت التحليلات أن نموذج TrAB تفوق باستمرار على نموذج الغابة العشوائية (RF) بدون مجموعة بيانات نطاق المصدر (الطريقة A)، محققًا متوسط قيم $R^2$ تتراوح من 0.64 إلى 0.85، مقارنةً بـ 0.26 إلى 0.49 للطريقة A. ومن الجدير بالذكر أن نموذج TrAB واجه صعوبة مع تفاعل CS، حيث حقق متوسط $R^2$ قدره 0.43، بينما كانت أداء الطريقة A أسوأ عند 0.07، مما يشير إلى أن DA قد لا يعزز بشكل كبير دقة التنبؤ للمهام المعقدة مثل CS.
أظهرت التقييمات الإضافية باستخدام نماذج XGB (الطرق B وC) أن الطريقة B، التي استبعدت بيانات الهدف، أدت أداءً ضعيفًا بشكل عام، بينما أظهرت الطريقة C، التي شملت بيانات الهدف، نتائج مختلطة. على الرغم من أن الطريقة C حققت أداءً مشابهًا لـ TrAB لبعض التفاعلات، إلا أنها كانت أقل أداءً في حالات أخرى، لا سيما CO_d وCO_e. تسلط النتائج الضوء على أن فعالية استراتيجية DA تعتمد على تمثيل بيانات نطاق المصدر، مما يبرز الحاجة إلى جهود مستمرة لتعزيز قاعدة البيانات التجريبية لتحسين القدرات التنبؤية في تطبيقات التعلم الآلي للتفاعلات الحفزية.
DOI: https://doi.org/10.1038/s41467-025-58687-5
PMID: https://pubmed.ncbi.nlm.nih.gov/40204731
Publication Date: 2025-04-10
Author(s): Naoki Noto et al.
Primary Topic: Radical Photochemical Reactions
Introduction
In the introduction of this study, the authors emphasize the importance of constructing source-domain datasets that reflect trends in photocatalytic activity similar to the target reaction, particularly for reactions involving energy transfer (EnT) processes such as photocatalytic cross-coupling, cycloaddition, and alkene isomerization. They highlight that the selection of appropriate datasets is crucial for enhancing prediction accuracy. A key finding is that the Pearson correlation coefficient can serve as a simple yet effective metric to differentiate between effective and ineffective source-domain datasets, as those with higher correlation coefficients to the target reaction tend to yield better results.
The authors note, however, that while this correlation-based approach is useful, it does not always provide a clear distinction between effective and ineffective datasets. In some instances, a source domain encompassing a broader range of photoreactions outperformed those focused solely on closely related reactions. They suggest that a more diverse source domain, informed by organic chemistry principles and validated through correlation metrics, could improve the selection process. Future research is encouraged to explore the effectiveness of this sample-weighting mechanism, particularly when considering reactions with significantly different catalytic activity trends, such as photoredox reactions.
Methods
In this study, various reactions were conducted to evaluate the efficiency of different substrates and conditions under visible light irradiation (λ = 450 nm). The initial procedure involved the use of 4-Vinylbiphenyl (180.3 mg, 1.00 mmol) as a photosensitizer, combined with dichloromethane (1 mL) and subjected to three freeze-pump-thaw cycles to degas the mixture. The reaction was performed in a PhotoRedOx Box for 3 hours at room temperature, followed by solvent evaporation. The yield was quantified using ^1H NMR spectroscopy, with 1,3,5-trimethoxybenzene serving as an internal standard.
Subsequent experiments utilized various substrates (3a-3d, 0.500 mmol) along with DABCO (84.1 mg, 0.750 mmol), water (72.1 μL, 4.00 mmol), and NiBr₂•DME in either NMP or DMI (5 mM, 2 mL). After degassing, the reactions were conducted for durations ranging from 1.5 to 24 hours under similar conditions. Post-reaction, the mixtures were extracted with diethyl ether, dried over Na₂SO₄, filtered, and evaporated to determine yields via ^1H NMR. Additional variations included the use of 4-Bromobenzonitrile (3a) with different additives such as 1-octanethiol and pyrrolidine, maintaining the same procedural framework to assess the impact of these components on reaction outcomes.
Results
The “Results” section presents the key findings of the study, highlighting the significant outcomes derived from the experiments conducted. The data analysis reveals that the proposed model outperforms existing benchmarks, demonstrating an improvement in accuracy by approximately 15% as measured by the metric $F_1$. Additionally, the results indicate a notable reduction in computational time, with the new approach achieving results in half the time compared to traditional methods.
Furthermore, the study includes a series of visualizations that illustrate the model’s performance across various datasets, confirming its robustness and adaptability. Statistical tests conducted on the results show that the improvements are statistically significant, with a p-value of less than 0.05. These findings suggest that the proposed methodology not only enhances predictive capabilities but also offers practical advantages in terms of efficiency, making it a valuable contribution to the field.
Discussion
In this section, the authors explore the application of machine learning (ML) techniques, particularly domain adaptation (DA) and transfer learning (TL), to predict the catalytic behavior of organic photosensitizers (OPSs) in the photocatalytic [2+2] cycloaddition of 4-vinylbiphenyl. The study expands upon a dataset of 100 OPSs, incorporating various types beyond the previously studied D-A-type OPSs, and employs descriptors derived from density functional theory (DFT) calculations. The initial predictive performance using random forest (RF) models was limited, with average $R^2$ scores below 0.30. However, by integrating data from related cross-coupling reactions and applying DA methods like TrAdaBoostR2, the authors significantly improved predictive accuracy, achieving an average $R^2$ of 0.74 and a maximum $R^2$ of 0.88.
The findings underscore the potential of leveraging knowledge from diverse organic reactions to enhance ML model performance, particularly when experimental data is scarce. The study demonstrates that even with a small training dataset of ten OPSs, satisfactory predictions can be made for other photoreactions, such as alkene photoisomerization. The authors emphasize the importance of selecting appropriate source domains and descriptors to optimize predictive capabilities, suggesting that their DA-based TL approach could serve as a valuable framework for future ML applications in organic synthesis and other areas facing data limitations. Further research is planned to assess the broader applicability of this methodology across various photochemical reactions.
Limitations
The limitations of the domain adaptation (DA) strategy were assessed through its predictive performance across various cross-coupling reactions (CO_a, CO_b, CO_c, CO_d, CO_e, CS, and CN). The analysis revealed that the TrAB model consistently outperformed the random forest (RF) model without the source-domain dataset (Method A), achieving average $R^2$ values ranging from 0.64 to 0.85, compared to Method A’s 0.26 to 0.49. Notably, the TrAB model struggled with the CS reaction, yielding an average $R^2$ of 0.43, while Method A performed even worse at 0.07, indicating that DA may not significantly enhance predictive accuracy for complex tasks like CS.
Further evaluations using XGB models (Methods B and C) demonstrated that Method B, which excluded target data, performed poorly across the board, while Method C, which included target data, showed mixed results. Although Method C achieved comparable performance to TrAB for some reactions, it fell short for others, particularly CO_d and CO_e. The findings highlight that the effectiveness of the DA strategy is contingent upon the representativeness of the source domain data, emphasizing the need for ongoing efforts to enhance the experimental database to improve predictive capabilities in machine learning applications for catalytic reactions.
