DOI: https://doi.org/10.1038/s41598-025-14847-7
PMID: https://pubmed.ncbi.nlm.nih.gov/40817123
تاريخ النشر: 2025-08-15
المؤلف: Sankar Murugesan وآخرون
الموضوع الرئيسي: الزراعة الذكية والذكاء الاصطناعي
نظرة عامة
تقدم هذه الدراسة معمارية هجينة جديدة ConvNet-ViT تهدف إلى الكشف بدقة وتصنيف الأمراض في أوراق الموز والكرز والطماطم. قارن الباحثون أداء هذا النموذج الجديد مع عدة نماذج مسبقة التدريب من الطراز الأول، بما في ذلك EfficientNetV2 وConvNeXt وSwin Transformer وVi-Transformer (ViT)، باستخدام مجموعة بيانات متاحة للجمهور تحتوي على أوراق صحية وغير صحية. تم تقسيم مجموعة البيانات بشكل منهجي إلى مجموعات تدريب وتحقق واختبار، مما يضمن تقييمات غير متحيزة. حقق نموذج Hybrid ConvNet-ViT، الذي يدمج طبقات الالتفاف لاستخراج الميزات المحلية مع طبقات ViT للسياق العالمي، دقة تصنيف ملحوظة بلغت 99.29%، متفوقًا على جميع نماذج الأساس.
على الرغم من أدائه القوي، تعترف الدراسة بعدة قيود. قد تعيق اعتماد النموذج على المدخلات المعتمدة على الصور فعاليته في ظل ظروف العالم الحقيقي المتغيرة، مثل التغيرات في الإضاءة والفوضى الخلفية. بالإضافة إلى ذلك، بينما تشمل مجموعة البيانات ثلاثة أنواع رئيسية من المحاصيل، لا تزال قابليتها للتطبيق على مجموعة واسعة من النباتات والأمراض غير مختبرة. يتطلب النموذج أيضًا إعدادًا حسابيًا معتدلًا، مما يطرح تحديات للنشر على الأجهزة ذات الموارد المنخفضة. ستركز الأبحاث المستقبلية على التحقق من صحة معمارية Hybrid ConvNet-ViT عبر مجموعات بيانات وظروف متنوعة، بالإضافة إلى استكشاف تحسينات مثل قدرات الكشف عن الأمراض في الوقت الحقيقي للتطبيقات الزراعية العملية.
طرق
في هذه الدراسة، يستكشف المؤلفون الكشف عن أمراض النباتات باستخدام صور لأوراق الموز والكرز والطماطم، معتمدين على معمارية تعلم عميق متنوعة بما في ذلك EfficientNetV2 وConvNeXt وSwin Transformer وVision Transformer (ViT) ونموذج Hybrid ConvNet-ViT المقترح. تستخدم الأبحاث مجموعة بيانات متاحة للجمهور تحتوي على صور مصنفة عبر تسع فئات، تمثل كل من حالات الأوراق الصحية والمريضة. تم تقسيم مجموعة البيانات بشكل منهجي إلى مجموعات تدريب (70%)، تحقق (15%)، واختبار (15%) لضمان تقييم عادل لأداء النموذج. تشمل الفئات أمراضًا متنوعة مثل بقعة ورقة الطماطم Septoria وفيروس موزاييك الطماطم، مع توزيع متوازن للعينات لتعزيز قدرات تعميم النموذج.
يتضمن الإعداد التجريبي تدريب النماذج في بيئة Windows 11 باستخدام Python والعديد من المكتبات، مع استخدام مُحسِّن AdamW ومعدل تعلم قدره 0.001 على مدى 50 دورة. يتم تطبيق تقنيات زيادة البيانات لتحسين قوة النموذج. يتم تقييم الأداء باستخدام مقاييس متعددة، بما في ذلك الدقة (Acc) والحساسية (Se) والخصوصية (Sp) والدقة (Pr) ودرجة F1، مما يوفر تقييمًا شاملاً لقدرات تصنيف النماذج. تؤكد الدراسة على أهمية هذه المقاييس في فهم أداء النموذج، خاصة في مجموعات البيانات غير المتوازنة النموذجية لسيناريوهات الكشف عن أمراض النباتات. تم هيكلة المنهجية في ثلاث مراحل: التدريب، والتحقق من خلال التحقق المتقاطع 5-fold، والاختبار، مما يضمن تقييمًا شاملاً لفعالية النماذج في التطبيقات الواقعية.
نقاش
تسلط قسم النقاش في ورقة البحث الضوء على دراسات متنوعة استخدمت تقنيات تعلم عميق متقدمة للكشف عن وتصنيف أمراض النباتات، مع التركيز بشكل خاص على محاصيل الموز والكرز والطماطم. من الجدير بالذكر أن نيكزون خيمينيز وآخرين استخدموا نهج CRISP-DM مع ثلاثة نماذج مسبقة التدريب، محققين أعلى دقة تصنيف بلغت 88.33% مع EfficientNetB0 لأمراض أوراق الموز. أظهرت دراسات أخرى، مثل تلك التي أجراها كريستيان أ. إلينيسا وآخرون، فعالية U-Net في الكشف المبكر عن الأمراض، محققة معامل Dice قدره 96.45% ودرجة IoU قدرها 93.23% مع مجموعة بيانات تحتوي على 18,240 صورة. تم استكشاف دمج النماذج الهجينة، مثل الجمع بين الشبكات العصبية الالتفافية (ConvNets) وتقنيات استخراج الميزات مثل SIFT وHOG، لتعزيز دقة التنبؤ وقوة النموذج ضد التغيرات في اتجاه الصورة ومقياسها.
علاوة على ذلك، يناقش البحث تطوير نموذج Hybrid ConvNet-ViT، الذي يجمع بشكل تآزري بين نقاط القوة في ConvNets في التقاط الميزات المحلية وViTs في نمذجة الاعتماديات العالمية. تهدف هذه المعمارية إلى تحسين الكشف عن أعراض الأمراض الدقيقة في ظل ظروف صعبة، مثل الخلفيات المتغيرة والإضاءة. يظهر النموذج المقترح دقة عالية وقدرات تعميم عبر محاصيل متعددة دون الحاجة إلى إعادة تدريب محددة للأنواع، مما يقدم حلاً قابلاً للتوسع للتطبيقات الزراعية الواقعية. بشكل عام، تسلط النتائج الضوء على إمكانيات التعلم العميق في تعزيز الممارسات الزراعية من خلال تحسين منهجيات الكشف عن الأمراض وتصنيفها.
DOI: https://doi.org/10.1038/s41598-025-14847-7
PMID: https://pubmed.ncbi.nlm.nih.gov/40817123
Publication Date: 2025-08-15
Author(s): Sankar Murugesan et al.
Primary Topic: Smart Agriculture and AI
Overview
This study presents a novel Hybrid ConvNet-ViT architecture aimed at accurately detecting and classifying diseases in banana, cherry, and tomato leaves. The researchers compared the performance of this new model against several state-of-the-art pre-trained models, including EfficientNetV2, ConvNeXt, Swin Transformer, and Vi-Transformer (ViT), using a publicly available dataset of healthy and unhealthy leaves. The dataset was systematically divided into training, validation, and test sets, ensuring unbiased evaluations. The Hybrid ConvNet-ViT model, which integrates convolutional layers for local feature extraction with ViT layers for global context, achieved a remarkable classification accuracy of 99.29%, outperforming all baseline models.
Despite its strong performance, the study acknowledges several limitations. The model’s reliance on image-based inputs may hinder its effectiveness under varying real-world conditions, such as changes in lighting and background clutter. Additionally, while the dataset encompasses three major crop types, its applicability to a wider range of plants and diseases remains untested. The model also requires a moderate computational setup, posing challenges for deployment on low-resource devices. Future research will focus on validating the Hybrid ConvNet-ViT architecture across diverse datasets and conditions, as well as exploring enhancements such as real-time disease detection capabilities for practical agricultural applications.
Methods
In this study, the authors investigate plant disease detection using images of banana, cherry, and tomato leaves, employing various deep learning architectures including EfficientNetV2, ConvNeXt, Swin Transformer, Vision Transformer (ViT), and a proposed Hybrid ConvNet-ViT model. The research utilizes a publicly available dataset containing labeled images across nine classes, representing both healthy and diseased leaf conditions. The dataset is systematically divided into training (70%), validation (15%), and testing (15%) subsets to ensure a fair evaluation of model performance. The classes include various diseases such as Tomato Septoria Leaf Spot and Tomato Mosaic Virus, with a balanced distribution of samples to enhance the model’s generalization capabilities.
The experimental setup involves training the models in a Windows 11 environment using Python and several libraries, with the AdamW optimizer and a learning rate of 0.001 over 50 epochs. Data augmentation techniques are applied to improve model robustness. Performance is assessed using multiple metrics, including accuracy (Acc), sensitivity (Se), specificity (Sp), precision (Pr), and F1 Score, providing a comprehensive evaluation of the models’ classification capabilities. The study emphasizes the importance of these metrics in understanding model performance, particularly in imbalanced datasets typical of plant disease detection scenarios. The methodology is structured into three phases: training, validation through 5-fold cross-validation, and testing, ensuring a thorough assessment of the models’ effectiveness in real-world applications.
Discussion
The discussion section of the research paper highlights various studies that have employed advanced deep learning techniques for the detection and classification of plant diseases, particularly focusing on banana, cherry, and tomato crops. Notably, Nixon Jiménez et al. utilized the CRISP-DM approach with three pretrained models, achieving the highest classification accuracy of 88.33% with EfficientNetB0 for banana leaf diseases. Other studies, such as those by Christian A. Elinisa et al., demonstrated the effectiveness of U-Net for early disease detection, achieving a Dice Coefficient of 96.45% and an IoU score of 93.23% with a dataset of 18,240 images. The integration of hybrid models, such as combining Convolutional Neural Networks (ConvNets) with feature extraction techniques like SIFT and HOG, has also been explored to enhance prediction accuracy and robustness against variations in image orientation and scale.
Furthermore, the paper discusses the development of a hybrid ConvNet-ViT model, which synergistically combines the strengths of ConvNets in capturing local features and ViTs in modeling global dependencies. This architecture aims to improve the detection of subtle disease symptoms under challenging conditions, such as varying backgrounds and lighting. The proposed model demonstrates high accuracy and generalization capabilities across multiple crops without requiring species-specific retraining, thus offering a scalable solution for real-world agricultural applications. Overall, the findings underscore the potential of deep learning in advancing agricultural practices through improved disease detection and classification methodologies.
