تعزيز تصنيف أمراض النباتات باستخدام شبكة عصبية تلافيفية قائمة على الانتباه للتدريب داخل مجموعة البيانات وعبرها
Advancing plant disease classification using an attention-based CNN for intra-dataset and cross- dataset training

شارك:
المجلة: Scientific Reports، المجلد: 16، العدد: 1
DOI: https://doi.org/10.1038/s41598-026-45464-7
PMID: https://pubmed.ncbi.nlm.nih.gov/41896662
تاريخ النشر: 2026-03-27
المؤلف: Prateek Mahapatra وآخرون
الموضوع الرئيسي: الزراعة الذكية والذكاء الاصطناعي

نظرة عامة

تؤكد الورقة البحثية على الدور الحاسم لتصنيف أمراض النباتات بدقة في تعزيز الأمن الغذائي وإنتاجية الزراعة. بينما تقدمت تقنيات التعلم العميق في هذا المجال، لا يزال دمج التدريب عبر مجموعات البيانات مع التدريب داخل مجموعة البيانات غير مستكشف بشكل كافٍ. لمعالجة هذه الفجوة، يقترح المؤلفون نموذجًا جديدًا يعتمد على الشبكات العصبية التلافيفية (CNN) مصممًا لتعزيز استخراج الميزات ودقة التصنيف عبر مجموعات بيانات متعددة. تم تقييم النموذج على خمس مجموعات بيانات—ديجيباتوس، بقعة الأوراق الشمالية (NLB)، بلانت فيلاج، بلانت دوك، ومجموعة بيانات CD&S—مركّزًا على أمراض الأوراق التي تؤثر على الذرة والبطاطس.

حقق النموذج المقترح دقة تصنيف ملحوظة بلغت 99.38% لأمراض أوراق البطاطس خلال التدريب داخل مجموعة البيانات باستخدام مجموعة بيانات بلانت فيلاج. في التدريب عبر مجموعة البيانات، حقق متوسط دقة تصنيف بلغ 82.93% لأمراض أوراق الذرة عند تدريبه على صور من مجموعة بيانات CD&S مع إزالة الخلفيات. تشير هذه النتائج إلى أداء متفوق مقارنة بالتقنيات الحالية تحت ظروف تجريبية مماثلة، مما يظهر قدرة النموذج على التكيف لتطبيقات داخل وخارج مجموعة البيانات. تشير النتائج إلى أن هذا النموذج يمكن أن يساهم بشكل كبير في الزراعة الدقيقة وإدارة الأمراض في بيئات زراعية متنوعة. قد تتضمن الأعمال المستقبلية استكشاف مجموعات بيانات أوراق نباتية إضافية لتعزيز قدرات تعميم النموذج بشكل أكبر.

مقدمة

في هذا القسم، يصف المؤلفون تدريب وتقييم نموذج مصمم لتصنيف أمراض أوراق الذرة، تحديدًا GLS (بقعة الأوراق الرمادية) وNLB (بقعة الأوراق الشمالية). تم تدريب النموذج في البداية على صور تمت إزالة الخلفية منها من مجموعة بيانات CD وS، ثم تم اختباره لاحقًا على صور من مجموعتي بيانات بلانت فيلاج وبلانت دوك، وكلاهما يحتوي أيضًا على حالات من الأمراض المستهدفة.

تشير نتائج التقييم، المقدمة من خلال مصفوفات الارتباك وتقارير التصنيف، إلى أن النموذج حقق دقة بلغت 81.09% في تحديد حالات GLS بشكل صحيح، بينما تم تصنيف 18.91% من هذه الحالات بشكل خاطئ. بالنسبة لـ NLB، كان أداء النموذج أقل فعالية، حيث بلغت الدقة 46.70% فقط، مما أدى إلى معدل تصنيف خاطئ كبير بلغ 53.30% كـ GLS. تسلط هذه النتائج الضوء على نقاط القوة والضعف في تصنيف الأمراض عبر مجموعات بيانات مختلفة، مما يبرز الحاجة إلى مزيد من التحسين لتحسين الدقة، خاصة لحالات NLB.

طرق

تحدد المنهجية المقترحة نهجًا منهجيًا لمعالجة مشكلة البحث. تبدأ بمراجعة أدبية شاملة لتحديد الفجوات الموجودة وتأسيس إطار نظري. بعد ذلك، تستخدم الدراسة تصميمًا مختلطًا، يجمع بين تقنيات جمع البيانات الكمية والنوعية. يسمح هذا النهج المزدوج بتحليل قوي لأسئلة البحث، مما يسهل فهمًا أعمق للظواهر قيد التحقيق.

يتضمن جمع البيانات استخدام الاستطلاعات والمقابلات، مما يضمن تنوعًا في وجهات النظر. سيتم تحليل البيانات الكمية باستخدام طرق إحصائية، بينما ستخضع البيانات النوعية لتحليل موضوعي لاستخراج الرؤى الرئيسية. كما تؤكد المنهجية على أهمية الاعتبارات الأخلاقية، مما يضمن أن سرية المشاركين والموافقة المستنيرة تحظى بالأولوية طوال عملية البحث. بشكل عام، تهدف هذه المنهجية المنظمة إلى تحقيق نتائج موثوقة وصالحة تساهم في هذا المجال.

مناقشة

تسلط المناقشة الضوء على التقدم في الأنظمة الآلية لتصنيف أمراض أوراق النباتات باستخدام تقنيات التعلم العميق، وخاصة الشبكات العصبية التلافيفية (CNN) وآليات الانتباه. أظهرت دراسات مختلفة فعالية هذه الطرق، حيث حققت النماذج معدلات دقة عالية، مثل 97.55% باستخدام نموذج CNN يعتمد على Wild Horse Optimizer مع طبقات انتباه وBiLSTM، و98.12% مع نموذج EfficientPNet لأمراض أوراق البطاطس. ومع ذلك، لا تزال التحديات قائمة في تعميم البيانات عبر المجموعات، كما يتضح من الفجوات الكبيرة في الدقة عندما تم اختبار النماذج المدربة على مجموعة بيانات واحدة على أخرى، مثل نموذج VGG16 الذي حقق فقط 39.87% دقة على مجموعة بيانات بلانت دوك بعد تدريبه على بلانت فيلاج.

تهدف الدراسة المقترحة إلى معالجة هذه القيود من خلال تقديم نموذج CNN يعتمد على الانتباه مصممًا لتعزيز الأداء داخل وخارج مجموعة البيانات. يتضمن هذا النموذج طبقات انتباه موضوعة بشكل استراتيجي بعد الكتل التلافيفية لتحسين استخراج الميزات ودقة التصنيف. استخدمت الدراسة خمس مجموعات بيانات متاحة للجمهور، بما في ذلك بلانت فيلاج وبلانت دوك، لتدريب وتقييم النموذج، محققة نتائج ملحوظة مثل 98% دقة لتصنيف أمراض الذرة على مجموعة بيانات بلانت فيلاج. تؤكد النتائج على أهمية آليات الانتباه في تحسين قوة النموذج وقدرات التعميم، مما يساهم في استراتيجيات أكثر فعالية للكشف عن الأمراض النباتية وإدارتها في الممارسات الزراعية.

القيود

تعترف الدراسة بعدة قيود قد تؤثر على قوة وملاءمة نتائجها. من الجدير بالذكر أن التباين الموجود في البيئات الزراعية الحقيقية—مثل الاختلافات في الإضاءة، وظروف الخلفية، وجودة الصورة—قد لا يتم تمثيله بشكل كافٍ من خلال مجموعات البيانات المتاحة للجمهور المستخدمة في هذا البحث. علاوة على ذلك، قد يؤثر حجم وتوزيع هذه المجموعات على قدرة النموذج على التعميم بشكل فعال على بيانات جديدة.

بالإضافة إلى ذلك، تقتصر الأبحاث الحالية على تصنيف الأمراض بناءً على الصور، متجاهلة العوامل الخارجية مثل التقلبات الزمنية والظروف البيئية. لتعزيز قوة النموذج وملاءمته العملية، ستركز التحقيقات المستقبلية على تقييم النموذج باستخدام بيانات حقلية في الوقت الحقيقي ومجموعات بيانات أكثر تنوعًا.

Journal: Scientific Reports, Volume: 16, Issue: 1
DOI: https://doi.org/10.1038/s41598-026-45464-7
PMID: https://pubmed.ncbi.nlm.nih.gov/41896662
Publication Date: 2026-03-27
Author(s): Prateek Mahapatra et al.
Primary Topic: Smart Agriculture and AI

Overview

The research paper emphasizes the critical role of accurate plant disease classification in enhancing food security and agricultural productivity. While deep learning techniques have advanced this field, the integration of cross-dataset training with intra-dataset training remains underexplored. To address this gap, the authors propose a novel attention-based Convolutional Neural Network (CNN) designed to enhance feature extraction and classification accuracy across multiple datasets. The model was evaluated on five datasets—Digipathos, Northern Leaf Blight (NLB), PlantVillage, PlantDoc, and the CD&S dataset—focusing on leaf diseases affecting corn and potatoes.

The proposed model achieved a remarkable classification accuracy of 99.38% for potato leaf diseases during intra-dataset training using the PlantVillage dataset. In cross-dataset training, it attained an average classification accuracy of 82.93% for corn leaf diseases when trained on images from the CD&S dataset with backgrounds removed. These results indicate superior performance compared to existing techniques under similar experimental conditions, showcasing the model’s adaptability for both intra- and cross-dataset applications. The findings suggest that this model could significantly contribute to precision farming and disease management in diverse agricultural settings. Future work may involve exploring additional plant leaf datasets to further enhance the model’s generalization capabilities.

Introduction

In this section, the authors describe the training and evaluation of a model designed for the classification of corn leaf diseases, specifically GLS (Gray Leaf Spot) and NLB (Northern Leaf Blight). The model was initially trained on background-removed images from the CD and S dataset and subsequently tested on images from the PlantVillage and PlantDoc datasets, both of which also contained instances of the target diseases.

The evaluation results, presented through confusion matrices and classification reports, indicate that the model achieved an accuracy of 81.09% in correctly identifying GLS instances, while it misclassified 18.91% of these cases. For NLB, the model’s performance was less effective, with only 46.70% accuracy, leading to a significant misclassification rate of 53.30% as GLS. These findings highlight the model’s strengths and weaknesses in disease classification across different datasets, emphasizing the need for further refinement to improve accuracy, particularly for NLB cases.

Methods

The proposed methodology outlines a systematic approach to address the research problem. It begins with a comprehensive literature review to identify existing gaps and establish a theoretical framework. Following this, the study employs a mixed-methods design, integrating both quantitative and qualitative data collection techniques. This dual approach allows for a robust analysis of the research questions, facilitating a deeper understanding of the phenomena under investigation.

Data collection involves the use of surveys and interviews, ensuring a diverse range of perspectives. The quantitative data will be analyzed using statistical methods, while qualitative data will undergo thematic analysis to extract key insights. The methodology also emphasizes the importance of ethical considerations, ensuring that participant confidentiality and informed consent are prioritized throughout the research process. Overall, this structured methodology aims to yield reliable and valid findings that contribute to the field.

Discussion

The discussion highlights the advancements in automated systems for classifying plant leaf diseases using deep learning techniques, particularly convolutional neural networks (CNNs) and attention mechanisms. Various studies have demonstrated the efficacy of these methods, with models achieving high accuracy rates, such as 97.55% using a Wild Horse Optimizer-based CNN with attention and BiLSTM layers, and 98.12% with the EfficientPNet model for potato leaf diseases. However, challenges remain in cross-dataset generalization, as evidenced by significant discrepancies in accuracy when models trained on one dataset were tested on another, such as the VGG16 model achieving only 39.87% accuracy on the PlantDoc dataset after training on PlantVillage.

The proposed study aims to address these limitations by introducing an attention-based CNN model designed to enhance both intra- and cross-dataset performance. This model incorporates attention layers strategically placed after convolutional blocks to improve feature extraction and classification accuracy. The research utilized five publicly available datasets, including PlantVillage and PlantDoc, to train and evaluate the model, achieving notable results such as 98% accuracy for corn disease classification on the PlantVillage dataset. The findings underscore the importance of attention mechanisms in improving model robustness and generalization capabilities, ultimately contributing to more effective plant disease detection and management strategies in agricultural practices.

Limitations

The study acknowledges several limitations that may affect the robustness and applicability of its findings. Notably, the variability inherent in real agricultural environments—such as differences in illumination, background conditions, and image quality—may not be adequately represented by the publicly available datasets utilized in this research. Furthermore, the size and distribution of these datasets could influence the model’s ability to generalize effectively to new data.

Additionally, the current research is confined to image-based disease categorization, omitting consideration of external factors such as temporal fluctuations and environmental conditions. To enhance the model’s robustness and practical applicability, future investigations will focus on evaluating the model with real-time field data and more diverse datasets.

شارك: