DOI: https://doi.org/10.1038/s41598-024-67526-4
تاريخ النشر: 2024-07-22
المؤلف: Jifei Zhao وآخرون
الموضوع الرئيسي: الزراعة الذكية والذكاء الاصطناعي
نظرة عامة
تقدم البحث خوارزمية خفيفة الوزن جديدة، YOLO-Granada، مصممة للكشف عن الرمان في الوقت الحقيقي خلال فترة نموه. تعتمد هذه الخوارزمية على نسخة محسنة من إطار عمل You Only Look Once النسخة 5 (YOLOv5) وتستخدم شبكة ShuffleNetv2 الخفيفة كعمود فقري لاستخراج الميزات. تشمل التحسينات الرئيسية دمج الالتفافات المجمعة ووحدة انتباه الكتلة الالتفافية (CBAM)، والتي تعمل معًا على تحسين دقة الكشف مع تقليل المتطلبات الحاسوبية بشكل كبير. تحقق خوارزمية YOLO-Granada دقة متوسطة تبلغ 0.922، متخلفة قليلاً عن دقة نموذج YOLOv5s الأصلي البالغة 0.929، بينما تحسن سرعة الكشف بنسبة 17.3% وتضغط معلمات النموذج والعمليات العائمة والحجم إلى 54.7% و51.3% و56.3% من الأصل، على التوالي. يمكن للخوارزمية معالجة 8.66 صورة في الثانية، مما يمكّن التطبيقات في الوقت الحقيقي على المنصات المحمولة منخفضة التكلفة.
تشمل اتجاهات البحث المستقبلية تعزيز دقة النموذج في البيئات المعقدة، وتوسيع مجموعة البيانات لتشمل أنواعًا مختلفة من الرمان، ودمج الصور المستندة إلى الطائرات بدون طيار لتحسين المراقبة الزراعية. تمتد التطبيقات المحتملة لـ YOLO-Granada إلى ما هو أبعد من الزراعة، مما يشير إلى فائدتها في مجالات مثل المراقبة والقيادة الذاتية. تسهم النتائج في تقدم الزراعة الذكية وتوفر إطارًا لتطوير الشبكات العصبية المخصصة للتطبيقات الزراعية.
مقدمة
ت outlines مقدمة ورقة البحث السياق الأساسي وأهمية الدراسة. تسلط الضوء على الفجوات الموجودة في الأدبيات وضرورة المزيد من التحقيق في الموضوع المحدد. يؤكد المؤلفون على أهمية بحثهم في معالجة هذه الفجوات، مما يسهم في النقاش الأكاديمي الأوسع. يتم الإشارة إلى النتائج الرئيسية من الدراسات السابقة لتأسيس إطار للبحث الحالي، مما يبرز أهمية الدراسة المقترحة في تعزيز المعرفة والفهم داخل هذا المجال.
طرق
في هذه الدراسة، استخدم المؤلفون بنية YOLOv5s كشبكة أساسية، مع تنفيذ عدة تحسينات تهدف إلى تحسين خفة وزن النموذج وتوافقه مع وحدة المعالجة المركزية. شملت التعديلات الرئيسية استبدال العمود الفقري الأصلي CSPDarknet-53 بعمود الفقري Attentioned-ShuffleNet ودمج وحدة انتباه الكتلة الالتفافية (CBAM).
يتم تمثيل شبكة كشف الكائنات YOLO-Granada المقترحة بصريًا في مخطط إطار شامل (الشكل 3)، والذي ي outlines أربعة مكونات أساسية: الإدخال، العمود الفقري، العنق، والتنبؤ. كل مكون هو جزء لا يتجزأ من وظيفة الشبكة. يوضح المخطط أيضًا عمليات مختلفة، بما في ذلك الالتفاف (Conv)، الالتفاف العمقي (DWConv)، تطبيع الدفعة (BatchNormal)، ودوال التنشيط (Leaky ReLU وLeaky SiLU). بالإضافة إلى ذلك، يبرز المخطط عملية دمج خريطة الميزات (Concat)، والتي تعزز أبعاد الميزات دون زيادة عدد المعلمات مقارنةً بالعملية الإضافية.
النتائج
يقدم قسم النتائج تقييمًا شاملاً لنموذج YOLO-Granada، مع مقارنة أدائه ضد 22 شبكة عصبية متطورة، بما في ذلك إصدارات مختلفة من YOLO. تحقق YOLO-Granada دقة تبلغ 0.922، واسترجاعًا قدره 0.833، ودرجة F1 تبلغ 87.84%، مع الحفاظ على أصغر حجم نموذج عند 7.74 ميجابايت—أصغر بنسبة 56.3% من YOLOv5s. على الرغم من أن متوسط الدقة (mAP@0.5) البالغ 0.922 هو أقل قليلاً من YOLOv5s (0.929)، إلا أن YOLO-Granada أسرع بحوالي 17.3%، حيث يعمل بسرعة 8.66 إطارًا في الثانية مقارنةً بـ 7.19 إطارًا في الثانية لـ YOLOv5s. تتضمن بنية النموذج تقنيات تحسين مثل عمود فقري ShuffleNet وآلية انتباه عصبية (CBAM) لتعزيز الدقة مع تقليل المتطلبات الحاسوبية، مما يجعلها مناسبة للنشر على الأجهزة ذات الموارد المحدودة.
تتمثل فعالية النموذج أيضًا من خلال تطبيقه على أجهزة Android لكل من تحليل الصور الثابتة ومعالجة تدفقات الفيديو في الوقت الحقيقي. في الاختبارات الثابتة، يظهر YOLO-Granada مستويات ثقة عالية (تصل إلى 0.98) في الكشف عن الرمان عبر مراحل نمو مختلفة، حتى في ظروف صعبة مثل الحجب والإضاءة المنخفضة. بالنسبة للتطبيقات في الوقت الحقيقي، يحقق النموذج معدلات FPS تتراوح بين 29.7 إلى 29.8، مما يعالج الإطارات المتتالية للحفاظ على دقة الكشف. تشير النتائج إلى أن YOLO-Granada مجهز جيدًا للتطبيقات العملية في المراقبة الزراعية، حيث يوازن بين الدقة والسرعة وحجم النموذج لتلبية احتياجات المستخدمين ذوي الموارد الحاسوبية المحدودة.
مناقشة
في قسم المناقشة، تؤكد الورقة على تفضيل المتزايد لخوارزميات كشف الكائنات ذات المرحلة الواحدة، وخاصة YOLO، على الطرق التقليدية ذات المرحلتين مثل RCNN، خاصة في التطبيقات الزراعية. تم التحقق من خفة وزن YOLO وأدائه المتفوق من خلال دراسات مختلفة، مع ملاحظات كبيرة في المهام مثل التعرف على براعم الشاي، وكشف التفاح في البساتين، وتقدير المحصول للطماطم. من الجدير بالذكر أن التحسينات على نماذج YOLO، مثل YOLOv3 وYOLOv5، حققت درجات دقة متوسطة مثيرة للإعجاب (mAP)، مما يوضح فعاليتها في سياقات زراعية متنوعة. يبرز القسم أيضًا التحديات في الكشف عن مراحل نمو الرمان، مما يشير إلى فجوة في الأبحاث الحالية التي تركز بشكل أساسي على تحديد الأمراض وتصنيف الجودة.
تناقش الورقة أيضًا تطوير ShuffleNetv2، الذي يحسن من سلفه من خلال تحسين الكفاءة الحاسوبية من خلال ابتكارات مثل الالتفافات المجمعة وخلط القنوات. لقد أظهر هذا النموذج زيادة في الدقة مع الحفاظ على بصمة حاسوبية منخفضة. بالإضافة إلى ذلك، يتم تقديم وحدة انتباه الكتلة الالتفافية (CBAM) كتحسين كبير في تعزيز استخراج الميزات في الشبكات العصبية الالتفافية، خاصة في معالجة الصور الزراعية. يتم تسليط الضوء على دمج CBAM مع ShuffleNetv2 في نموذج YOLO-Granada المقترح كاستراتيجية رئيسية لتحسين أداء الكشف مع الحفاظ على خفة وزن النموذج. تؤكد النتائج على أهمية الجمع بين الهياكل الفعالة وآليات الانتباه لتعزيز دقة وقابلية تطبيق نماذج كشف الكائنات في البيئات الزراعية المعقدة.
DOI: https://doi.org/10.1038/s41598-024-67526-4
Publication Date: 2024-07-22
Author(s): Jifei Zhao et al.
Primary Topic: Smart Agriculture and AI
Overview
The research presents a novel lightweight algorithm, YOLO-Granada, designed for the real-time detection of pomegranates during their growth period. This algorithm is based on an improved version of the You Only Look Once version 5 (YOLOv5) framework and utilizes a lightweight ShuffleNetv2 network as its backbone for feature extraction. Key enhancements include the incorporation of grouped convolutions and the Convolutional Block Attention Module (CBAM), which collectively optimize detection accuracy while significantly reducing computational demands. The YOLO-Granada algorithm achieves an average accuracy of 0.922, closely trailing the original YOLOv5s model’s accuracy of 0.929, while improving detection speed by 17.3% and compressing model parameters, floating-point operations, and size to 54.7%, 51.3%, and 56.3% of the original, respectively. The algorithm can process 8.66 images per second, enabling real-time applications on low-cost mobile platforms.
Future research directions include enhancing model accuracy in complex environments, expanding the dataset to encompass various pomegranate varieties, and integrating UAV-based imagery for improved agricultural monitoring. The potential applications of YOLO-Granada extend beyond agriculture, suggesting its utility in fields such as surveillance and autonomous driving. The findings contribute to the advancement of smart agriculture and provide a framework for developing neural networks tailored for agricultural applications.
Introduction
The introduction of the research paper outlines the foundational context and significance of the study. It highlights the existing gaps in the literature and the necessity for further investigation into the specified topic. The authors emphasize the relevance of their research in addressing these gaps, thereby contributing to the broader academic discourse. Key findings from prior studies are referenced to establish a framework for the current research, underscoring the importance of the proposed study in advancing knowledge and understanding within the field.
Methods
In this study, the authors employed the YOLOv5s architecture as the foundational network, implementing several enhancements aimed at improving model lightweightness and CPU compatibility. Key modifications included replacing the original CSPDarknet-53 backbone with the Attentioned-ShuffleNet backbone and integrating the Convolutional Block Attention Module (CBAM).
The proposed YOLO-Granada object detection network is visually represented in a comprehensive framework diagram (Fig. 3), which outlines four essential components: input, backbone, neck, and prediction. Each component is integral to the network’s functionality. The diagram also illustrates various operations, including convolution (Conv), depthwise convolution (DWConv), batch normalization (BatchNormal), and activation functions (Leaky ReLU and Leaky SiLU). Additionally, the diagram highlights the feature map concatenation (Concat) process, which enhances feature dimensionality without increasing the parameter count compared to the additive operation.
Results
The results section presents a comprehensive evaluation of the YOLO-Granada model, benchmarking its performance against 22 state-of-the-art neural networks, including various YOLO versions. YOLO-Granada achieves a precision of 0.922, recall of 0.833, and an F1 score of 87.84%, while maintaining the smallest model size at 7.74 MB—56.3% smaller than YOLOv5s. Although its mean Average Precision (mAP@0.5) of 0.922 is slightly lower than YOLOv5s (0.929), YOLO-Granada is approximately 17.3% faster, operating at 8.66 FPS compared to YOLOv5s’s 7.19 FPS. The model’s architecture incorporates optimization techniques such as a ShuffleNet backbone and a neural attention mechanism (CBAM) to enhance accuracy while minimizing computational demands, making it suitable for deployment on resource-constrained devices.
The model’s effectiveness is further validated through its application on Android devices for both static image analysis and real-time video stream processing. In static tests, YOLO-Granada demonstrates high confidence levels (up to 0.98) in detecting pomegranates across various growth stages, even in challenging conditions such as occlusion and low light. For real-time applications, the model achieves FPS rates of 29.7 to 29.8, effectively processing consecutive frames to maintain detection accuracy. The results indicate that YOLO-Granada is well-equipped for practical applications in agricultural monitoring, balancing accuracy, speed, and model size to meet the needs of users with limited computational resources.
Discussion
In the discussion section, the paper emphasizes the growing preference for single-stage object detection algorithms, particularly YOLO, over traditional two-stage methods like RCNN, especially in agricultural applications. YOLO’s lightweight nature and superior performance have been validated through various studies, with significant advancements noted in tasks such as tea bud recognition, apple detection in orchards, and yield estimation for tomatoes. Notably, enhancements to YOLO models, such as YOLOv3 and YOLOv5, have achieved impressive mean Average Precision (mAP) scores, demonstrating their effectiveness in diverse agricultural contexts. The section also highlights the challenges in detecting pomegranate growth stages, indicating a gap in existing research that primarily focuses on disease identification and quality grading.
The paper further discusses the development of ShuffleNetv2, which improves upon its predecessor by optimizing computational efficiency through innovations like group convolution and channel shuffling. This model has shown increased accuracy while maintaining a low computational footprint. Additionally, the convolutional block attention module (CBAM) is introduced as a significant advancement in enhancing feature extraction in convolutional neural networks, particularly in agricultural image processing. The integration of CBAM with ShuffleNetv2 in the proposed YOLO-Granada model is highlighted as a key strategy for improving detection performance while keeping the model lightweight. The findings underscore the importance of combining efficient architectures with attention mechanisms to enhance the accuracy and applicability of object detection models in complex agricultural environments.
