مراجعة لتقنية الإدراك البصري لروبوتات حصاد الفواكه الذكية
A review of visual perception technology for intelligent fruit harvesting robots

شارك:
المجلة: Frontiers in Plant Science، المجلد: 16
DOI: https://doi.org/10.3389/fpls.2025.1646871
PMID: https://pubmed.ncbi.nlm.nih.gov/40904864
تاريخ النشر: 2025-08-19
المؤلف: Yikun Huang وآخرون
الموضوع الرئيسي: الزراعة الذكية والذكاء الاصطناعي

نظرة عامة

تقدم هذه الورقة مراجعة شاملة للتقدم في تكنولوجيا الإدراك البصري لروبوتات قطف الفاكهة الذكية، مع التأكيد على دورها الحاسم في تعزيز الإنتاجية الزراعية. تبدأ بمناقشة الأنواع المختلفة من الكاميرات المستخدمة في قطف الفاكهة، مع تسليط الضوء على مزاياها وقيودها: الكاميرات الأحادية للمهام البسيطة، والكاميرات الثنائية للتعقيد المعتدل، وكاميرات الضوء الهيكلي وكاميرات زمن الرحلة (ToF) للإدراك العمقي عالي الدقة في البيئات المعقدة. كما تتناول الورقة تقنيات اكتشاف الأجسام، مقارنة بين معالجة الصور التقليدية وطرق التعلم العميق الحديثة مثل YOLO وSSD، مشيرة إلى أنه بينما توفر طرق التعلم العميق دقة متفوقة، فإنها تتطلب بيانات تدريب كبيرة وأجهزة قوية.

تتناول المراجعة أيضًا أهمية تحديد نقاط القطف بدقة من خلال إعادة البناء ثلاثي الأبعاد المعتمد على الرؤية واستشعار العمق، والتي تعتبر حيوية لتقليل تلف الفاكهة وزيادة معدلات النجاح في القطف. بالإضافة إلى ذلك، تستكشف دمج تقنيات مثل V-SLAM، وتخطيط المسار المتنقل، وجدولة المهام، والتي تعمل مجتمعة على تحسين الكفاءة التشغيلية في البساتين. يتم تسليط الضوء على الجمع بين الرؤية النشطة والتحكم البصري كوسيلة لتعزيز التكيف والدقة في البيئات الديناميكية، مما يسمح للروبوتات بتعديل زوايا رؤيتها واستراتيجيات التحكم في الوقت الحقيقي. تختتم الورقة بتحديد التحديات المستمرة، بما في ذلك التكيف البيئي ودمج الأنظمة، مع توقع التقدم المستقبلي في روبوتات قطف الفاكهة الذكية التي تستفيد من هذه التقنيات لمواجهة سيناريوهات زراعية أكثر تعقيدًا.

مقدمة

تسلط المقدمة الضوء على التحديات الملحة التي تواجه الإنتاج الزراعي بسبب تزايد السكان العالمي، بما في ذلك ارتفاع تكاليف العمالة وانخفاض الإنتاجية، مما يعيق تقدم الزراعة الحديثة (Zhang et al., 2024). استجابةً لذلك، حظيت تقنيات الأتمتة الزراعية، وخاصة الروبوتات الذكية للحصاد، باهتمام كبير كحلول قابلة للتطبيق (Chunjiang et al., 2023).

تكشف تحليل إحصائي لمقالات البحث من 2005 إلى 2024، مأخوذ من قاعدة بيانات Web of Science، عن زيادة ملحوظة في المنشورات المتعلقة بـ “حصاد الفاكهة”، حيث ارتفعت من 732 مقالة في 2005 إلى 2130 في 2024. تؤكد هذه الاتجاهات على الاهتمام المتزايد بتقنيات الزراعة الذكية، مع ظهور الإدراك البصري والروبوتات كمواضيع مركزية في جهود البحث المستمرة.

الطرق

تناقش هذه القسم التقدم في طرق اكتشاف الأجسام، خاصة في سياق التطبيقات الزراعية. لقد حسنت طرق اكتشاف الأجسام ذات المرحلتين، مثل R-CNN ومشتقاتها (Fast R-CNN وFaster R-CNN وMask R-CNN)، بشكل كبير من دقة وسرعة الاكتشاف من خلال دمج توليد اقتراحات المناطق مع استخراج الميزات العميقة. تم تطبيق هذه الطرق بشكل فعال في سيناريوهات زراعية متنوعة، بما في ذلك اكتشاف الفاكهة وتقييم النضج. ومع ذلك، أدت كثافتها الحاسوبية إلى تحول نحو طرق اكتشاف الأجسام ذات المرحلة الواحدة، ولا سيما عائلة YOLO (You Only Look Once)، التي تحول تحديد موقع الجسم إلى مشكلة انحدار، محققة سرعة وكفاءة عالية. قدمت الإصدارات الحديثة، مثل YOLOv8 وYOLOv12، ابتكارات معمارية تعزز قدرات الاكتشاف متعددة المقاييس مع الحفاظ على انخفاض زمن الاستجابة وتعقيد الحوسبة.

بالإضافة إلى ذلك، يبرز القسم ظهور طرق اكتشاف الأجسام المعتمدة على المحولات، التي تستفيد من آليات الانتباه الذاتي لنمذجة الاعتماديات العالمية، مما يوفر مزايا في البيئات المعقدة. على الرغم من إمكانياتها، تواجه هذه النماذج تحديات مثل التكاليف الحاسوبية العالية والاعتماد على مجموعات بيانات كبيرة. كما يؤكد النص على أهمية تقنيات تصنيف البيانات وتحديد المواقع في قطف الفاكهة، مشيرًا إلى التنوع في الأساليب وتأثير دقة التصنيف على كفاءة الحصاد. استكشفت الأبحاث الحديثة نماذج هجينة تجمع بين ميزات المحولات وإطارات YOLO لتعزيز المتانة والدقة في مهام الاكتشاف الزراعي. بشكل عام، يبرز القسم التطور المستمر لتقنيات اكتشاف الأجسام وتأثيراتها على الحصاد الذكي للفواكه، مع تحديد التحديات الرئيسية التي لا تزال قائمة في تحقيق الأداء والدقة في الوقت الحقيقي في البيئات الزراعية الديناميكية.

المناقشة

تسلط قسم المناقشة في ورقة البحث الضوء على تطور الحالة الحالية لروبوتات حصاد الفاكهة الذكية، متتبعة تطورها من الطرق الميكانيكية المبكرة في 1968 إلى الأنظمة الحديثة المعززة بالتقدم في الذكاء الاصطناعي وتقنيات الاستشعار. كانت تقنيات الحصاد الأولية، مثل الاهتزاز الميكانيكي والهوائي، محدودة بسبب ميلها إلى إتلاف الفاكهة. ومع ذلك، سمح دمج تكنولوجيا الرؤية الحاسوبية وتقنيات التحكم في التسعينيات بإنشاء أذرع روبوتية أكثر تطورًا قادرة على محاكاة إجراءات الحصاد اليدوي. على الرغم من التحديات المبكرة، بما في ذلك التكاليف العالية والدقة المنخفضة، فإن ظهور تقنيات الصناعة 4.0 قد حسّن بشكل كبير من قدرات هذه الروبوتات، خاصة من حيث الإدراك واتخاذ القرار المستقل.

تصنف الورقة روبوتات الحصاد إلى نوعين: الكمية والانتقائية، حيث يكون النوع الأخير أكثر ملاءمة للفواكه الجاهزة للسوق بسبب قدرته على تقليل التلف وتحديد المنتجات الناضجة بدقة باستخدام أنظمة الإدراك البصري المتقدمة. تؤكد المناقشة على الدور الحاسم لتكنولوجيا الإدراك البصري، التي تمكن الروبوتات من التكيف مع الظروف البيئية المتغيرة وتنفيذ مهام الحصاد بفعالية. يتم تحليل أنواع الكاميرات المختلفة، بما في ذلك الأحادية والثنائية وRGB-D وكاميرات الأحداث، لمزاياها وقيودها في التقاط معلومات العمق واللون الضرورية لاكتشاف الفاكهة. يختتم القسم بالتأكيد على أهمية تقنيات اكتشاف الأجسام، وخاصة أساليب التعلم العميق، في تعزيز كفاءة ودقة روبوتات حصاد الفاكهة في البيئات الزراعية المعقدة.

القيود

تسلط قسم القيود الضوء على العديد من التحديات التقنية المرتبطة بتقنيات التصوير المتقدمة المستخدمة في التعرف على الفاكهة وتحديد مواقعها في التطبيقات الروبوتية. بينما حسنت الكاميرات الأحادية والثنائية وأجهزة استشعار العمق ثلاثية الأبعاد بشكل كبير من الدقة، تواجه الكاميرات الثنائية مشكلات تتعلق بالمعايرة والتكيف. توفر أجهزة استشعار العمق الأكثر تطورًا، مثل كاميرات زمن الرحلة (ToF) وكاميرات الضوء الهيكلي، إدراكًا أفضل للعمق ولكن تأتي بتكاليف عالية ومتطلبات حوسبة كبيرة.

بالإضافة إلى ذلك، فإن تنفيذ خوارزميات التعلم العميق، مثل YOLO، قد عزز دقة اكتشاف الفاكهة؛ ومع ذلك، تتطلب هذه الطرق موارد حوسبة قوية، ومجموعات بيانات تدريب واسعة، ودمج فعال لبيانات العمق. وبالتالي، فإن تحقيق توازن مثالي بين الكفاءة الحاسوبية ودقة الاكتشاف لا يزال تحديًا حاسمًا لمدى قابلية هذه التقنيات للتطبيقات العملية.

Journal: Frontiers in Plant Science, Volume: 16
DOI: https://doi.org/10.3389/fpls.2025.1646871
PMID: https://pubmed.ncbi.nlm.nih.gov/40904864
Publication Date: 2025-08-19
Author(s): Yikun Huang et al.
Primary Topic: Smart Agriculture and AI

Overview

This paper provides a comprehensive review of advancements in visual perception technology for intelligent fruit-picking robots, emphasizing its critical role in enhancing agricultural productivity. It begins by discussing the various types of cameras utilized in fruit picking, highlighting their respective advantages and limitations: monocular cameras for simple tasks, binocular cameras for moderate complexity, and structured light and Time-of-Flight (ToF) cameras for high-precision depth perception in intricate environments. The paper further examines object detection techniques, contrasting traditional image processing with modern deep learning approaches like YOLO and SSD, noting that while deep learning methods yield superior accuracy, they necessitate substantial training data and robust hardware.

The review also addresses the significance of accurate picking point localization through vision-based 3D reconstruction and depth sensing, which are vital for minimizing fruit damage and enhancing picking success rates. Additionally, it explores the integration of technologies such as V-SLAM, mobile path planning, and task scheduling, which collectively improve operational efficiency in orchards. The combination of active vision and visual servoing is highlighted as a means to enhance adaptability and precision in dynamic settings, allowing robots to adjust their visual angles and control strategies in real-time. The paper concludes by identifying ongoing challenges, including environmental adaptability and system integration, while projecting future advancements in intelligent fruit-picking robots that leverage these technologies to tackle increasingly complex agricultural scenarios.

Introduction

The introduction highlights the pressing challenges faced by agricultural production due to the growing global population, including rising labor costs and low productivity, which hinder the advancement of modern agriculture (Zhang et al., 2024). In response, agricultural automation technologies, particularly intelligent harvesting robots, have garnered significant attention as viable solutions (Chunjiang et al., 2023).

A statistical analysis of research articles from 2005 to 2024, sourced from the Web of Science database, reveals a notable increase in publications related to “Fruit Harvesting,” escalating from 732 articles in 2005 to 2130 in 2024. This trend underscores the burgeoning interest in smart agriculture technologies, with visual perception and robotics emerging as central themes in ongoing research efforts.

Methods

The section discusses advancements in object detection methods, particularly in the context of agricultural applications. Two-stage object detection methods, such as R-CNN and its derivatives (Fast R-CNN, Faster R-CNN, and Mask R-CNN), have significantly improved detection accuracy and speed by integrating region proposal generation with deep feature extraction. These methods have been effectively applied in various agricultural scenarios, including fruit detection and maturity assessment. However, their computational intensity has led to a shift towards one-stage object detection methods, notably the YOLO (You Only Look Once) family, which transforms object localization into a regression problem, achieving high speed and efficiency. Recent iterations, such as YOLOv8 and YOLOv12, have introduced architectural innovations that enhance multi-scale detection capabilities while maintaining low latency and computational complexity.

Additionally, the section highlights the emergence of Transformer-based object detection methods, which leverage self-attention mechanisms to model global dependencies, offering advantages in complex environments. Despite their potential, these models face challenges such as high computational costs and reliance on large datasets. The text also emphasizes the importance of data labeling and localization techniques in fruit picking, noting the diversity in approaches and the impact of labeling accuracy on harvesting efficiency. Recent research has explored hybrid models that combine Transformer features with YOLO frameworks to enhance robustness and accuracy in agricultural detection tasks. Overall, the section underscores the ongoing evolution of object detection technologies and their implications for intelligent fruit harvesting, while also identifying key challenges that remain in achieving real-time performance and accuracy in dynamic agricultural settings.

Discussion

The discussion section of the research paper highlights the evolution and current state of intelligent fruit harvesting robots, tracing their development from early mechanical methods in 1968 to modern systems enhanced by advancements in artificial intelligence and sensor technologies. Initial harvesting techniques, such as mechanical and pneumatic vibration, were limited by their tendency to damage fruit. However, the integration of computer vision and control technologies in the 1990s allowed for more sophisticated robotic arms capable of simulating manual harvesting actions. Despite early challenges, including high costs and low precision, the emergence of Industry 4.0 technologies has significantly improved the capabilities of these robots, particularly in terms of perception and autonomous decision-making.

The paper categorizes harvesting robots into bulk and selective types, with the latter being more suitable for market-ready fruits due to their ability to minimize damage and accurately identify ripe produce using advanced visual perception systems. The discussion emphasizes the critical role of visual perception technology, which enables robots to adapt to varying environmental conditions and effectively execute harvesting tasks. Various camera types, including monocular, binocular, RGB-D, and event cameras, are analyzed for their respective advantages and limitations in capturing depth and color information essential for fruit detection. The section concludes by underscoring the importance of object detection technologies, particularly deep learning approaches, in enhancing the efficiency and accuracy of fruit harvesting robots in complex agricultural environments.

Limitations

The section on limitations highlights several technical challenges associated with advanced imaging technologies used for fruit recognition and localization in robotic applications. While monocular, binocular, and 3D depth sensors have significantly improved precision, binocular cameras face issues related to calibration and adaptability. More sophisticated depth sensors, such as Time-of-Flight (ToF) and structured light cameras, provide superior depth perception but come with high costs and substantial computational demands.

Additionally, the implementation of deep learning algorithms, like YOLO, has enhanced the accuracy of fruit detection; however, these methods necessitate robust computational resources, extensive training datasets, and effective depth data fusion. Consequently, achieving an optimal balance between computational efficiency and detection accuracy remains a critical challenge for the scalability of these technologies in practical applications.

شارك: