شبكة CNN–ViT هجينة خفيفة الوزن جديدة لتصنيف أمراض أوراق الذرة
A novel lightweight hybrid CNN–ViT for maize leaf disease classification

شارك:
المجلة: Scientific Reports، المجلد: 16، العدد: 1
DOI: https://doi.org/10.1038/s41598-026-41190-2
PMID: https://pubmed.ncbi.nlm.nih.gov/41741594
تاريخ النشر: 2026-02-25
المؤلف: Saber Mehdipour وآخرون
الموضوع الرئيسي: الزراعة الذكية والذكاء الاصطناعي

نظرة عامة

تسلط الأبحاث الضوء على الدور الحاسم للذرة كمحصول عالمي والتحديات التي تطرحها الأمراض النباتية، مما يستلزم طرق تشخيص فعالة. تعتبر طرق الفحص اليدوي التقليدية غير فعالة ومعرضة للأخطاء، مما يدفع لاستكشاف الحلول الآلية. أظهرت التقدمات الأخيرة في رؤية الكمبيوتر والتعلم العميق، وخاصة من خلال الشبكات العصبية التلافيفية (CNNs) ومحولات الرؤية (ViTs)، إمكانيات في تشخيص الأمراض النباتية. ومع ذلك، غالبًا ما تفشل الشبكات العصبية التلافيفية في التقاط المعلومات السياقية العالمية، بينما تتطلب محولات الرؤية مجموعات بيانات واسعة وقوة حسابية كبيرة.

لمعالجة هذه القيود، يقترح المؤلفون نموذج هجين خفيف الوزن يجمع بين مكونات CNN وViT، مستفيدًا من بنية مزيج من الخبراء (MoE) للتركيز بشكل تكيفي على الميزات المحلية أو العالمية بناءً على خصائص بيانات الإدخال. تم تقييم هذا النموذج على مجموعة بيانات جديدة تتكون من صور كاملة لنباتات الذرة، محققًا دقة تصنيف مثيرة للإعجاب تبلغ 99.90%. تتجاوز هذه الأداء بشكل ملحوظ أداء النماذج المتطورة الحالية، بما في ذلك MobileViT وPiT وEdgeNeXt وDeiT. تشير النتائج إلى أن مثل هذه الهياكل الهجينة الخفيفة الوزن يمكن أن توفر تشخيصًا عالي الأداء للأمراض، مما يجعلها مناسبة للتطبيقات العملية في الزراعة.

مقدمة

تسلط المقدمة الضوء على الدور الحاسم للزراعة، وخاصة الذرة، في ضمان الأمن الغذائي العالمي والاستقرار الاقتصادي. تؤكد على ضعف إنتاج الذرة أمام الأمراض النباتية، مما يستلزم الكشف المبكر والدقيق للتخفيف من فقدان المحاصيل وتعزيز ممارسات الإدارة المستدامة. بينما تعتبر الفحوصات الخبيرة قيمة، إلا أنها غالبًا ما تكون غير فعالة ومعرضة للأخطاء، مما يخلق طلبًا على أنظمة التعرف على الأمراض الآلية. أدت التقدمات الأخيرة في التعلم العميق، وخاصة من خلال الشبكات العصبية التلافيفية (CNNs)، إلى تحسين التشخيص القائم على الصور؛ ومع ذلك، تكافح الشبكات العصبية التلافيفية في التقاط الاعتماديات بعيدة المدى الضرورية لتمييز أنماط الأمراض المتشابهة.

لمعالجة هذه القيود، ظهرت محولات الرؤية (ViTs) كبديل واعد بسبب قدرتها على نمذجة السياق العالمي من خلال آليات الانتباه الذاتي. على الرغم من مزاياها، فإن نشر نماذج تعتمد على ViT في السياقات الزراعية يعيقه متطلباتها لمجموعات بيانات كبيرة وموارد حسابية كبيرة. تهدف مقدمة نماذج CNN-ViT الهجينة إلى دمج نقاط القوة في كلا الهيكلين، مستفيدة من الشبكات العصبية التلافيفية لاستخراج الميزات المحلية ومحولات الرؤية للتفاعلات العالمية. ومع ذلك، لا يزال هناك تحدٍ رئيسي في تحقيق التوازن بين الكفاءة الحسابية والمرونة التمثيلية، خاصة في التكيف مع هياكل الإدخال المتنوعة بينما تكون مناسبة للبيئات ذات الموارد المحدودة. بالإضافة إلى ذلك، هناك فجوة ملحوظة في النماذج الحالية، التي تركز بشكل أساسي على صور الأوراق الفردية، مما يبرز الحاجة إلى طرق يمكن أن تحلل بفعالية صور النباتات الكاملة في ظروف العالم الحقيقي المتنوعة.

النتائج

في تحليل مجموعة بيانات PlantVillage، تشير النتائج إلى أن EdgeNeXt يظهر سلوك انتباه غير مثالي، يتميز بأنماط تنشيط متفرقة وتركيز ملحوظ على عناصر الخلفية. يقلل هذا عدم التوافق في الانتباه من قابلية تفسير النموذج وموثوقيته، خاصة في البيئات التي تكون فيها المشتتات البصرية شائعة. بالمقابل، لا يتفوق MXiT فقط على النماذج الحالية من حيث التركيز التنبؤي، بل يولد أيضًا خرائط انتباه أكثر وضوحًا وقابلية للتفسير.

يمكن أن يُعزى الأداء المتفوق لـ MXiT إلى تصميمه المعماري، الذي يجمع بين استخراج الميزات متعددة المقاييس وآليات الانتباه الذاتي الفعالة. يسمح ذلك لـ MXiT بالتقاط كل من السياق العالمي وأنماط الأمراض المحلية بشكل فعال، مما يعزز حساسيته المكانية وقدرته على استخدام الإشارات البصرية الخاصة بالمجال الضرورية لاكتشاف الأمراض النباتية. تؤكد النتائج على أهمية التصميم المعماري في تحديد الميزات البصرية الدقيقة بدقة، والتي تعتبر ضرورية لتطبيقات مثل مراقبة صحة المحاصيل.

المناقشة

تسلط قسم المناقشة في الورقة الضوء على مزايا هياكل مزيج من الخبراء (MoE)، خاصة في سياق تعزيز الكفاءة وقابلية التوسع في نماذج الرؤية. أظهرت متغيرات MoE النادرة، مثل V-MoE، أن تنشيط طبقات الخبراء بشكل انتقائي يمكن أن يقلل بشكل كبير من تكاليف الاستدلال مع الحفاظ على أداء عالٍ. يقترح المؤلفون بنية هجينة خفيفة الوزن جديدة تجمع بين خبير شبكة عصبية تلافيفية (CNN) وخبير محول رؤية (ViT) من خلال آلية بوابة. يسمح هذا التصميم باستخراج الميزات بشكل تكيفي، مما يلتقط بشكل فعال كل من الاعتماديات المحلية والعالمية، ويحقق دقة مثيرة للإعجاب تبلغ 99.90% على مجموعة بيانات حقيقية لأمراض الذرة، متفوقًا على العديد من النماذج المتطورة مع عدد أقل من المعلمات.

تشمل المساهمات الرئيسية للدراسة تقديم بنية CNN-ViT الهجينة، وإدماج وحدة انتباه ذات رأس واحد معدلة لتعزيز قابلية التوسع، وتقييم النموذج على مجموعة بيانات تعكس ظروف العالم الحقيقي بدقة أكبر من مجموعات البيانات التقليدية. تؤكد النتائج على إمكانية دمج الشبكات العصبية التلافيفية وViTs للاستفادة من نقاط القوة الخاصة بهما، مما يؤدي إلى تحسين التعميم والكفاءة لمهام اكتشاف الأمراض النباتية. كما توضح الورقة هيكلها، مشيرة إلى الأقسام المخصصة للأعمال ذات الصلة، ومجموعات البيانات، والمنهجية، والتجارب، والاستنتاجات، التي تهدف مجتمعة إلى معالجة الفجوات الموجودة في الأدبيات واقتراح اتجاهات البحث المستقبلية.

Journal: Scientific Reports, Volume: 16, Issue: 1
DOI: https://doi.org/10.1038/s41598-026-41190-2
PMID: https://pubmed.ncbi.nlm.nih.gov/41741594
Publication Date: 2026-02-25
Author(s): Saber Mehdipour et al.
Primary Topic: Smart Agriculture and AI

Overview

The research highlights the critical role of maize as a global crop and the challenges posed by plant diseases, which necessitate efficient diagnostic methods. Traditional manual inspection methods are inefficient and error-prone, prompting the exploration of automated solutions. Recent advancements in computer vision and deep learning, particularly through Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs), have shown potential in diagnosing plant diseases. However, CNNs often fail to capture global contextual information, while ViTs demand extensive datasets and significant computational power.

To address these limitations, the authors propose a lightweight hybrid model that combines CNN and ViT components, leveraging a mixture-of-experts (MoE) architecture to adaptively focus on local or global features based on the characteristics of the input data. This model was evaluated on a novel dataset comprising full maize plant images, achieving an impressive classification accuracy of 99.90%. This performance notably surpasses that of existing state-of-the-art models, including MobileViT, PiT, EdgeNeXt, and DeiT. The findings suggest that such lightweight hybrid architectures can provide high-performance disease diagnosis, making them suitable for practical applications in agriculture.

Introduction

The introduction highlights the critical role of agriculture, particularly maize, in ensuring global food security and economic stability. It emphasizes the vulnerability of maize production to plant diseases, which necessitates early and accurate detection to mitigate yield loss and promote sustainable management practices. While expert inspections are valuable, they are often inefficient and error-prone, creating a demand for automated disease recognition systems. Recent advancements in deep learning, particularly through Convolutional Neural Networks (CNNs), have enhanced image-based diagnosis; however, CNNs struggle with capturing long-range dependencies essential for distinguishing similar disease patterns.

To address these limitations, Vision Transformers (ViTs) have emerged as a promising alternative due to their ability to model global context through self-attention mechanisms. Despite their advantages, the deployment of ViT-based models in agricultural contexts is hindered by their requirement for large datasets and significant computational resources. The introduction of hybrid CNN-ViT models aims to combine the strengths of both architectures, leveraging CNNs for local feature extraction and Transformers for global interactions. Nonetheless, a key challenge remains in balancing computational efficiency with representational flexibility, particularly in adapting to diverse input structures while being suitable for resource-constrained environments. Additionally, there is a noted gap in existing models, which predominantly focus on single-leaf images, highlighting the need for methods that can effectively analyze full-plant images in varied real-world conditions.

Results

In the analysis of the PlantVillage Dataset, the results indicate that EdgeNeXt demonstrates suboptimal attention behavior, characterized by scattered activation patterns and a pronounced focus on background elements. This misalignment in attention diminishes the model’s interpretability and reliability, particularly in environments where visual distractions are prevalent. In contrast, MXiT not only outperforms existing models in terms of predictive focus but also generates clearer and more interpretable attention maps.

The superior performance of MXiT can be attributed to its architectural design, which combines multi-scale feature extraction with efficient self-attention mechanisms. This allows MXiT to effectively capture both global context and localized disease patterns, enhancing its spatial sensitivity and ability to utilize domain-specific visual cues critical for plant disease detection. The findings underscore the significance of architectural design in accurately identifying fine-grained visual features, which are essential for applications such as crop health monitoring.

Discussion

The discussion section of the paper highlights the advantages of Mixture-of-Experts (MoE) architectures, particularly in the context of enhancing efficiency and scalability in vision models. Sparse MoE variants, such as V-MoE, have shown that selectively activating expert layers can significantly reduce inference costs while maintaining high performance. The authors propose a novel lightweight hybrid architecture that combines a Convolutional Neural Network (CNN) expert and a Vision Transformer (ViT) expert through a gating mechanism. This design allows for adaptive feature extraction, effectively capturing both local and global dependencies, and achieves an impressive accuracy of 99.90% on a real-world maize disease dataset, outperforming several state-of-the-art models with fewer parameters.

Key contributions of the study include the introduction of the CNN-ViT hybrid architecture, the incorporation of a modified Single-Head Self-Attention module to enhance scalability, and the evaluation of the model on a dataset that reflects real-world conditions more accurately than traditional datasets. The findings underscore the potential of combining CNNs and ViTs to leverage their respective strengths, resulting in improved generalization and efficiency for plant disease detection tasks. The paper also outlines its structure, indicating sections dedicated to related works, datasets, methodology, experiments, and conclusions, which collectively aim to address existing gaps in the literature and propose future research directions.

شارك: