YOLO الديناميكي لاكتشاف الأجسام الصغيرة تحت الماء
Dynamic YOLO for small underwater object detection

شارك:
المجلة: Artificial Intelligence Review، المجلد: 57، العدد: 7
DOI: https://doi.org/10.1007/s10462-024-10788-1
تاريخ النشر: 2024-06-06
المؤلف: Jie Chen وآخرون
الموضوع الرئيسي: تقنيات تحسين الصور

نظرة عامة

تتناول البحث التحديات المتعلقة بالكشف عن الأجسام الصغيرة في البيئات تحت الماء، وهو جانب حاسم من استكشاف البحار. لمواجهة هذه المشكلة، يقترح المؤلفون كاشف ديناميكي خفيف الوزن يعتمد على YOLO، والذي يتضمن شبكة عظمية تعتمد على الالتفاف القابل للتشويه v3 مصممة خصيصًا لتعزيز الكشف عن الأجسام الصغيرة. يقدم البحث إطار عمل موحد لدمج الميزات يستخدم آليات انتباه على مستوى القناة، وعلى مستوى المقياس، وواعية مكانية لدمج خرائط الميزات من مقاييس مختلفة بشكل فعال، وبالتالي تحسين الكشف عن الأجسام الصغيرة.

كما يتميز النموذج المقترح برأس كشف جديد يفصل بين مهام التصنيف والتحديد، مما يقلل من الصراع الكامن بين هذين الهدفين. تظهر التجارب الواسعة أن YOLO الديناميكي يتفوق بشكل كبير على الأساليب الحديثة، محققًا تحسينات قدرها +0.8 AP و +1.8 AP S على مجموعة بيانات DUO، بينما يظهر أيضًا أداءً متفوقًا على مجموعات بيانات Pascal VOC و MS COCO. يختتم البحث بدراسات إبطال تؤكد فعالية كل مكون تصميم، مما يبرز التأثير المحتمل لهذا البحث على تقدم الكشف عن الأجسام الصغيرة تحت الماء.

مقدمة

تسلط مقدمة هذه الورقة البحثية الضوء على الدور الحاسم لاستكشاف البحار في استغلال الموارد والحفاظ على النظام البيئي، مما يبرز إمكانيات الكشف عن الأجسام تحت الماء المعتمد على الرؤية التي تسهلها التقدمات في الروبوتات البحرية. على الرغم من وعدها، تواجه هذه المجال تحديات كبيرة، خاصة في الكشف عن الأجسام الصغيرة، التي تشكل غالبية الأهداف تحت الماء، مثل الكائنات البحرية. يشير البحث إلى أن حوالي 43.9% من الأجسام صغيرة، مما يعقد الكشف بسبب نقص المعلومات البصرية وقيود الشبكات العصبية الالتفافية التقليدية (CNNs) في استخراج الميزات.

لمعالجة هذه التحديات، يقترح المؤلفون شبكة عظمية جديدة تعتمد على الشبكات الالتفافية القابلة للتشويه (DCN) مصممة خصيصًا للكشف عن الأجسام الصغيرة. يقدمون إطار عمل موحد لدمج الميزات متعددة المقاييس يتضمن آليات انتباه على مستوى القناة، وعلى مستوى المقياس، وواعية مكانية لتعزيز تجميع الميزات. بالإضافة إلى ذلك، يقدم البحث رأسًا مفصولًا موسعًا لتحسين دقة التحديد من خلال فصل مهام التصنيف والتحديد. يظهر كاشف YOLO الديناميكي أداءً متفوقًا على مجموعات البيانات المرجعية، محققًا نتائج حديثة مع تحسينات كبيرة مقارنة بالأساليب الحالية. تمهد المقدمة الطريق لاستكشاف مفصل للمنهجيات المقترحة وفعاليتها في الأقسام اللاحقة.

طرق

في هذا القسم، يقيم المؤلفون نموذج YOLO الديناميكي الخاص بهم على مجموعات بيانات Pascal VOC و MS COCO. بالنسبة لتجارب Pascal VOC، تم تدريب النموذج على مجموعات التدريب لعامي 2007 و 2012 واختباره على مجموعة الاختبار لعام 2007. تشير النتائج إلى أنه بينما تؤدي معظم الكواشف في الوقت الحقيقي بشكل مشابه لمجموعة بيانات DUO، يتفوق RTMDet بشكل كبير على YOLOX، محققًا دقة متوسطة (AP) قدرها 60.7، وهو زيادة قدرها 6.4 AP مقارنة بـ YOLOv8. من الجدير بالذكر أن RTMDet يتفوق في الكشف عن الأجسام الصغيرة بأفضل أداء قدره 29.0 AP_S. يحقق نموذج YOLO الديناميكي، مع عدد أقل من المعلمات، أداءً حديثًا جديدًا قدره 61.7 AP، على الرغم من أنه يعاني من تدهور قدره حوالي 3.0 AP_S في الكشف عن الأجسام الصغيرة بسبب عدم توازن الفئات أثناء التدريب، حيث تشكل الأجسام الصغيرة 7.4% فقط من مجموعة البيانات.

على مجموعة بيانات MS COCO، تم تدريب نموذج YOLO الديناميكي على مجموعة COCO Train 2017 وتم تقييمه على مجموعة Val 2017. تظهر النتائج، كما هو موضح في الجدول 3، أن النموذج يحقق 45.5 AP، متفوقًا على الكواشف السابقة في الوقت الحقيقي تحت ظروف مشابهة. ومع ذلك، فإنه يتخلف قليلاً عن YOLOv8 في الكشف عن الأجسام الصغيرة، مع AP_S قدره 25.4. بشكل عام، يظهر نموذج YOLO الديناميكي توازنًا ملائمًا بين المعلمات والدقة، مما يعزز ميزته التنافسية في مهام الكشف عن الأجسام.

النتائج

يقدم قسم النتائج تصورًا لأداء الكشف عن الأجسام تحت الماء على مجموعة بيانات DUO، مع تسليط الضوء على سيناريوهات مختلفة صعبة. توضح الشكل 4 أربع حالات تمثيلية: (أ) الأجسام الصغيرة، (ب) التباين المنخفض، (ج) الحجب، و (د) التجميع. في السيناريو (أ)، تتناثر الأجسام الصغيرة من الإكينوس على قاع البحر، حيث يفشل YOLOX في الكشف عنها، بينما ينجح RTMDet و YOLO الديناميكي في تحديد بعضها، كما هو موضح بالأسهم الصفراء. على العكس، في السيناريو (ب)، يتفوق YOLOX في الكشف عن الأجسام في ظروف التباين المنخفض، بينما ينتج RTMDet و YOLO الديناميكي اكتشافات خاطئة بسبب الظلال المربكة.

في السيناريو (ج)، ينجح كل من RTMDet و YOLO الديناميكي في الكشف عن إكينوس محجوب جزئيًا بواسطة صخرة، على الرغم من أن YOLO الديناميكي يتم تضليله بواسطة عقدة على حبل مهمل، مشابهًا لأداء YOLOX. يتضمن السيناريو (د) التجميع، حيث تظهر جميع الكواشف سلوكًا متسقًا، مع توليد YOLOX و YOLO الديناميكي لاكتشافات خاطئة في الزاوية العليا اليسرى. تعترف التحليل ببعض الإيجابيات الكاذبة التافهة، خاصة في السيناريوهات (ب) و (د)، والتي تعتبر غير ذات أهمية نظرًا للتحسينات الكمية العامة في الأداء، كما هو مفصل في الجدول 1. من الجدير بالذكر أن YOLO الديناميكي يظهر حساسية أكبر مقارنة بـ RTMDet.

نقاش

في هذا القسم، يناقش المؤلفون التقدم في الكشف عن الأجسام تحت الماء، مع التركيز بشكل خاص على الأجسام الصغيرة، التي تقدم تحديات فريدة بسبب عوامل مثل تدهور الصورة وقدرات التعميم المحدودة. يبرزون قيود الكواشف المعتمدة على المحولات في التقاط المعلومات المحلية، مما يؤدي إلى دمج الشبكات الالتفافية القابلة للتشويه (DCN) في أطر الكشف. يقترح المؤلفون شبكة عظمية خفيفة الوزن تعتمد على DCN v3، والتي تعزز استخراج الميزات للأجسام الصغيرة من خلال تقنيات أخذ العينات والتعديل التكيفية.

يتناول البحث أيضًا الصراع بين مهام التصنيف والتحديد داخل رؤوس الكشف، والتي يمكن أن تعيق الأداء. للتخفيف من ذلك، يقدم المؤلفون رأسًا متماشيًا مع المهام يستخدم انتباهًا واعيًا للقناة والالتفاف القابل للتشويه لتحسين توافق المهام. يقترحون كاشف YOLO الديناميكي الذي يتضمن آلية دمج ميزات متعددة المقاييس من خلال عنق ديناميكي، مما يعزز تمثيل الأجسام الصغيرة. يتم التحقق من فعالية النموذج المقترح من خلال تجارب واسعة على مجموعة بيانات DUO، مما يظهر تحسينات في أداء الكشف مقارنة بالأساليب الحديثة.

Journal: Artificial Intelligence Review, Volume: 57, Issue: 7
DOI: https://doi.org/10.1007/s10462-024-10788-1
Publication Date: 2024-06-06
Author(s): Jie Chen et al.
Primary Topic: Image Enhancement Techniques

Overview

The research addresses the challenges of detecting small objects in underwater environments, a critical aspect of marine exploration. To tackle this issue, the authors propose a lightweight dynamic YOLO detector, which incorporates a backbone network based on deformable convolution v3 specifically designed for enhanced small object detection. The paper introduces a unified feature fusion framework that employs channel-wise, scale-wise, and spatial-aware attention mechanisms to effectively merge feature maps from various scales, thereby optimizing the detection of small objects.

The proposed model also features a novel detection head that decouples and aligns the classification and localization tasks, mitigating the inherent conflict between these two objectives. Extensive experiments demonstrate that dynamic YOLO significantly outperforms state-of-the-art methods, achieving improvements of +0.8 AP and +1.8 AP S on the DUO dataset, while also showing superior performance on the Pascal VOC and MS COCO datasets. The paper concludes with ablation studies that validate the effectiveness of each design component, emphasizing the potential impact of this research on advancing small underwater object detection.

Introduction

The introduction of this research paper highlights the critical role of marine exploration in resource exploitation and ecosystem preservation, emphasizing the potential of vision-based underwater object detection facilitated by advancements in marine robotics. Despite its promise, the field faces significant challenges, particularly in detecting small objects, which constitute a majority of underwater targets, such as marine organisms. The paper notes that approximately 43.9% of objects are small, complicating detection due to insufficient visual information and the limitations of conventional convolutional neural networks (CNNs) in feature extraction.

To address these challenges, the authors propose a novel backbone network based on Deformable Convolution Networks (DCN) designed specifically for small object detection. They introduce a unified framework for multi-scale feature fusion that incorporates channel-, scale-, and spatial-aware attention mechanisms to enhance feature aggregation. Additionally, the paper presents an extended decoupled head to improve localization accuracy by disentangling classification and localization tasks. The proposed dynamic YOLO detector demonstrates superior performance on benchmark datasets, achieving state-of-the-art results with significant improvements over existing methods. The introduction sets the stage for a detailed exploration of the proposed methodologies and their effectiveness in subsequent sections.

Methods

In this section, the authors evaluate their dynamic YOLO model on the Pascal VOC and MS COCO datasets. For the Pascal VOC experiments, the model was trained on the 2007 and 2012 training sets and tested on the 2007 testing set. Results indicate that while most real-time detectors perform comparably to the DUO dataset, RTMDet significantly outperforms YOLOX, achieving an average precision (AP) of 60.7, which is a 6.4 AP increase over YOLOv8. Notably, RTMDet excels in small object detection with a best performance of 29.0 AP_S. The dynamic YOLO model, with fewer parameters, achieves a new state-of-the-art performance of 61.7 AP, although it experiences a degradation of approximately 3.0 AP_S in small object detection due to class imbalance during training, where small objects constitute only 7.4% of the dataset.

On the MS COCO dataset, the dynamic YOLO model was trained on the COCO Train 2017 set and evaluated on the Val 2017 set. The results, as presented in Table 3, show that the model achieves 45.5 AP, outperforming previous real-time detectors under comparable conditions. However, it falls slightly short of YOLOv8 in small object detection, with an AP_S of 25.4. Overall, the dynamic YOLO model demonstrates a favorable parameter-accuracy trade-off, reinforcing its competitive edge in object detection tasks.

Results

The results section presents a visualization of underwater object detection performance on the DUO dataset, highlighting various challenging scenarios. Figure 4 illustrates four representative cases: (a) small objects, (b) low contrast, (c) occlusion, and (d) clustering. In scenario (a), small echinus objects are scattered on the seabed, where YOLOX fails to detect them, while RTMDet and dynamic YOLO successfully identify some, as indicated by yellow arrows. Conversely, in scenario (b), YOLOX excels in detecting objects under low contrast conditions, whereas RTMDet and dynamic YOLO produce false detections due to confusing shadows.

In scenario (c), both RTMDet and dynamic YOLO successfully detect an echinus partially obscured by a rock, although dynamic YOLO is misled by a knot on a discarded rope, similar to YOLOX’s performance. Scenario (d) involves clustering, where all detectors exhibit consistent behavior, with YOLOX and dynamic YOLO generating false detections in the upper-left corner. The analysis acknowledges some trivial false positives, particularly in scenarios (b) and (d), which are deemed negligible given the overall quantitative performance improvements, as detailed in Table 1. Notably, dynamic YOLO demonstrates greater sensitivity compared to RTMDet.

Discussion

In this section, the authors discuss advancements in underwater object detection, particularly focusing on small objects, which present unique challenges due to factors like image degradation and limited generalization capabilities. They highlight the limitations of transformer-based detectors in capturing local information, leading to the integration of Deformable Convolutional Networks (DCN) into detection frameworks. The authors propose a lightweight backbone network based on DCN v3, which enhances feature extraction for small objects through adaptive sampling and modulation techniques.

The paper also addresses the conflict between classification and localization tasks within detection heads, which can hinder performance. To mitigate this, the authors introduce a task-aligned head that employs channel-aware attention and deformable convolution to improve task alignment. They propose a dynamic YOLO detector that incorporates a multi-scale feature fusion mechanism through a dynamic neck, enhancing the representation of small objects. The effectiveness of the proposed model is validated through extensive experiments on the DUO dataset, demonstrating improved detection performance compared to state-of-the-art methods.

شارك: