انتباه متعدد المقاييس قائم على SED-YOLO لاكتشاف الأجسام الصغيرة في الاستشعار عن بعد
SED-YOLO based multi-scale attention for small object detection in remote sensing

شارك:
المجلة: Scientific Reports، المجلد: 15، العدد: 1
DOI: https://doi.org/10.1038/s41598-025-87199-x
PMID: https://pubmed.ncbi.nlm.nih.gov/39856170
تاريخ النشر: 2025-01-24
المؤلف: Zhenyun Du وآخرون
الموضوع الرئيسي: تطبيقات الشبكات العصبية المتقدمة

نظرة عامة

تقدم ورقة البحث SED-YOLO، وهي نسخة محسنة من نموذج YOLOv5s مصممة خصيصًا لاكتشاف الأجسام الصغيرة في صور الاستشعار عن بُعد، مع معالجة التحديات التي تطرحها ضوضاء الصورة والخلفيات المزدحمة. تشمل الابتكارات الرئيسية دمج الالتفاف المتقطع القابل للتبديل (SAC) في الشبكة الأساسية لتحسين استخراج الميزات، وتنفيذ آلية الانتباه متعددة المقاييس الفعالة (EMA) لتسهيل تعلم الميزات متعددة المقاييس بشكل فعال مع تقليل التكاليف الحاسوبية. علاوة على ذلك، تم تقديم طريقة دمج متكيفة في قسم العنق لتحسين دمج الميزات بناءً على محتوى وخصائص الأجسام، مما يعزز من مرونة النموذج.

تم تحسين بنية النموذج بشكل أكبر من خلال توسيع رأس الكشف من ثلاثة إلى أربعة مقاييس، مع دمج طبقة مخصصة لاكتشاف الأجسام الصغيرة ووحدة الرأس الديناميكية (DyHead)، التي تعدل ديناميكيًا أوزان الانتباه في خرائط الميزات. تكشف النتائج التجريبية على مجموعة بيانات DOTA أن SED-YOLO تحقق دقة متوسطة تبلغ 71.6%، متجاوزة النموذج الأصلي YOLOv5s بنسبة 2.4%. لا تعزز هذه التحسينات دقة اكتشاف الأجسام الصغيرة فحسب، بل تضمن أيضًا الكفاءة الحاسوبية، مما يجعل SED-YOLO مناسبًا للتطبيقات في الوقت الحقيقي في الاستشعار عن بُعد. تضع النتائج أساسًا قويًا للتقدم المستقبلي في اكتشاف الأجسام الصغيرة، لا سيما للتطبيقات المحددة مثل تحديد خزانات التخزين وحقول الرياضة.

الطرق

تم إجراء التجارب في بيئة محكومة باستخدام نظام تشغيل Windows 10، مزودًا ببطاقة رسومات RTX 4090 تحتوي على 24 جيجابايت من الذاكرة. شملت مجموعة البرمجيات Python 3.8 وPyTorch 1.11.0 وCUDA 11.3. تم الحفاظ على معلمات ثابتة طوال مراحل التدريب والتحقق والاختبار، مع حجم دفعة يبلغ 16 ومعدل تعلم قدره 0.01. استخدم عملية التدريب مُحسِّن الانحدار العشوائي (SGD)، وتم معالجة صور مجموعة البيانات بدقة 640 × 640 بكسل. تم تدريب النموذج على مدى 350 دورة، باستخدام دالة خسارة تقاطع الاتحاد الكامل (CIOU) لتقييم الأداء.

النتائج

تظهر نتائج التجارب فعالية تحسينات النموذج المقترحة، لا سيما من خلال مقارنات الكشف المرئية باستخدام مجموعة بيانات DOTA. في سلسلة من أربعة سيناريوهات، يقدم العمود الأول تسميات الحقيقة الأرضية، بينما يظهر العمود الثاني نتائج الكشف من نموذج YOLOv5s، ويعرض العمود الثالث النتائج من النموذج المحسن. من الجدير بالذكر أنه في السيناريوهات التي تتضمن مركبات مزدحمة وأجسام صغيرة على مدرج المطار، واجه نموذج YOLOv5s صعوبة في اكتشاف الأجسام الصغيرة، بينما نجح نموذجنا في تحديد هذه الأهداف. بالإضافة إلى ذلك، أنتج نموذج YOLOv5s اكتشافات خاطئة في سيناريو المطار، بينما حافظ نموذجنا على الدقة وتجنب مثل هذه الأخطاء.

في المشهد الساحلي، فشل نموذج YOLOv5s مرة أخرى في اكتشاف المركبات ضد خلفيات معقدة، مما يبرز قيود قدراته على الكشف. في المقابل، أظهر نموذجنا، المعزز بوحدة الدمج المتكيفة ووحدة DyHead مع أربعة رؤوس كشف، تحسينات كبيرة في تقليل الإفراط في التكيف، والإيجابيات الكاذبة، والاكتشافات المفقودة. تؤكد هذه النتائج على دقة النموذج المحسنة وقوته، مما يجعله أكثر فعالية في مهام التعرف على الصور المعقدة.

المناقشة

في هذا القسم، يقدم المؤلفون SED-YOLO، وهي نسخة محسنة من نموذج YOLOv5s، مصممة خصيصًا لتحسين اكتشاف الأجسام الصغيرة في الخلفيات المعقدة، لا سيما في تطبيقات الاستشعار عن بُعد. تشمل الابتكارات الرئيسية إدخال الالتفاف المتقطع القابل للتبديل (SAConv) في الشبكة الأساسية، والذي يحل محل الالتفافات القياسية لالتقاط معلومات الأجسام متعددة المقاييس بشكل أفضل مع الحفاظ على دقة الصورة. بالإضافة إلى ذلك، يعزز دمج آلية الانتباه المتحركة الأسية (EMA) قدرة النموذج على التركيز على الميزات ذات الصلة عبر مقاييس مختلفة، مما يحسن دقة الكشف عن الأجسام الصغيرة. يتميز قسم العنق بوحدة دمج متكيفة تعدل استراتيجيات دمج الميزات ديناميكيًا بناءً على الخصائص المدخلة، مما يحسن من تكامل المعلومات الدلالية والمكانية.

علاوة على ذلك، تم تعزيز رأس الكشف بطبقة إضافية مخصصة لاكتشاف الأجسام الصغيرة وإدخال وحدة DyHead، التي تستخدم آلية انتباه متعددة الجوانب لتعزيز تمثيل الميزات. تظهر النتائج التجريبية أن SED-YOLO تحقق زيادة بنسبة 2.4% في دقة المتوسط (mAP) مقارنة بنموذج YOLOv5s الأصلي، مع الحفاظ على الكفاءة الحاسوبية. يظهر هذا النموذج وعدًا كبيرًا للتطبيقات في الوقت الحقيقي في الاستشعار عن بُعد، لا سيما لاكتشاف الأجسام الصغيرة مثل خزانات التخزين وحقول الرياضة، مما يضع أساسًا قويًا للتقدم المستقبلي في منهجيات اكتشاف الأجسام.

Journal: Scientific Reports, Volume: 15, Issue: 1
DOI: https://doi.org/10.1038/s41598-025-87199-x
PMID: https://pubmed.ncbi.nlm.nih.gov/39856170
Publication Date: 2025-01-24
Author(s): Zhenyun Du et al.
Primary Topic: Advanced Neural Network Applications

Overview

The research paper introduces SED-YOLO, an enhanced version of the YOLOv5s model specifically tailored for the detection of small objects in remote sensing images, addressing the challenges posed by image noise and cluttered backgrounds. Key innovations include the integration of Switchable Atrous Convolution (SAC) in the backbone network to improve feature extraction, and the implementation of an Efficient Multi-Scale Attention (EMA) mechanism to facilitate effective multi-scale feature learning while minimizing computational costs. Furthermore, an adaptive Concat method is introduced in the Neck section to optimize feature fusion based on the content and characteristics of the objects, thereby enhancing the model’s versatility.

The model’s architecture is further refined by expanding the detection head from three to four scales, incorporating a dedicated small object detection layer and the Dynamic Head (DyHead) module, which dynamically adjusts attention weights in feature maps. Experimental results on the DOTA dataset reveal that SED-YOLO achieves a mean Average Precision (mAP) of 71.6%, surpassing the original YOLOv5s by 2.4%. This improvement not only enhances the accuracy of small object detection but also ensures computational efficiency, making SED-YOLO suitable for real-time applications in remote sensing. The findings lay a robust foundation for future advancements in small object detection, particularly for specific applications such as identifying storage tanks and sports fields.

Methods

The experiments were conducted in a controlled environment using a Windows 10 operating system, equipped with an RTX 4090 GPU featuring 24GB of memory. The software stack included Python 3.8, PyTorch 1.11.0, and CUDA 11.3. Consistent hyperparameters were maintained throughout the training, validation, and testing phases, with a batch size of 16 and a learning rate of 0.01. The training process employed the Stochastic Gradient Descent (SGD) optimizer, and the dataset images were processed at a resolution of 640 × 640 pixels. The model was trained over 350 epochs, utilizing the Complete Intersection over Union (CIOU) loss function to evaluate performance.

Results

The results of the experiments demonstrate the effectiveness of the proposed model enhancements, particularly through visualized detection comparisons using the DOTA dataset. In a series of four scenarios, the first column presents ground truth labels, the second shows detection results from the YOLOv5s model, and the third displays results from the improved model. Notably, in scenarios involving densely packed vehicles and small objects on an airport tarmac, the YOLOv5s model struggled to detect small objects, while our model successfully identified these targets. Additionally, the YOLOv5s model produced false detections in the airport scenario, whereas our model maintained accuracy and avoided such errors.

In the coastal scene, the YOLOv5s model again failed to detect vehicles against complex backgrounds, highlighting the limitations of its detection capabilities. In contrast, our model, enhanced by the adaptive Concat module and the DyHead module with four detection heads, demonstrated significant improvements in reducing overfitting, false positives, and missed detections. These findings underscore the model’s enhanced precision and robustness, making it more effective for complex image recognition tasks.

Discussion

In this section, the authors present SED-YOLO, an enhanced version of the YOLOv5s model, specifically designed to improve the detection of small objects in complex backgrounds, particularly in remote sensing applications. Key innovations include the introduction of the Switchable Atrous Convolution (SAConv) in the backbone network, which replaces standard convolutions to better capture multi-scale object information while preserving image resolution. Additionally, the incorporation of an Exponential Moving Average (EMA) attention mechanism enhances the model’s ability to focus on relevant features across different scales, thereby improving detection accuracy for small objects. The Neck section features an adaptive Concat module that dynamically adjusts feature fusion strategies based on input characteristics, optimizing the integration of semantic and spatial information.

Furthermore, the detection head has been augmented with an additional layer dedicated to small object detection and the introduction of the DyHead module, which employs a multi-faceted attention mechanism to enhance feature representation. Experimental results demonstrate that SED-YOLO achieves a 2.4% increase in mean Average Precision (mAP) over the original YOLOv5s model, while maintaining computational efficiency. This model shows significant promise for real-time applications in remote sensing, particularly for detecting small objects such as storage tanks and sports fields, thereby laying a robust foundation for future advancements in object detection methodologies.

شارك: