LRDS-YOLO يعزز اكتشاف الأجسام الصغيرة في الصور الجوية للطائرات بدون طيار بتصميم خفيف وفعال
LRDS-YOLO enhances small object detection in UAV aerial images with a lightweight and efficient design

شارك:
المجلة: Scientific Reports، المجلد: 15، العدد: 1
DOI: https://doi.org/10.1038/s41598-025-07021-6
PMID: https://pubmed.ncbi.nlm.nih.gov/40595033
تاريخ النشر: 2025-07-01
المؤلف: Yuqi Han وآخرون
الموضوع الرئيسي: تطبيقات الشبكات العصبية المتقدمة

نظرة عامة

تقدم هذه القسم نظرة عامة على التحديات المرتبطة بالكشف عن الأجسام الصغيرة في صور الطائرات بدون طيار، مع تسليط الضوء على قضايا مثل التباين المنخفض، والخلفيات المعقدة، والموارد الحاسوبية المحدودة التي تؤدي إلى معدلات عالية من عدم الكشف ودقة تحديد الموقع الضعيفة. تعاني طرق الكشف التقليدية غالبًا من فقدان المعلومات، وعدم كفاية تفاعل الميزات عبر الطبقات، ورؤوس الكشف غير المرنة، مما يستدعي تطوير حلول أكثر فعالية.

لمعالجة هذه التحديات، يقدم المؤلفون LRDS-YOLO، وهو نموذج خفيف وفعال مصمم خصيصًا لتطبيقات الطائرات بدون طيار. تشمل الابتكارات الرئيسية وحدة خفض الوزن التكيفية الخفيفة (LAD)، التي تحافظ على الميزات الدقيقة وتقلل من فقدان المعلومات، وشبكة هرم الميزات المعاد معايرتها (Re-Calibration FPN)، التي تعزز دمج الميزات متعددة المقاييس من خلال التفاعلات ثنائية الاتجاه والانتباه الهجين المدرك للدقة. بالإضافة إلى ذلك، يركز آلية انتباه SegNext على الأهداف مع تقليل ضوضاء الخلفية، ورأس الكشف الديناميكي (DyHead) يحسن وزن الميزات لتحسين قوة الكشف. تظهر النتائج التجريبية أن LRDS-YOLO تحقق دقة متوسطة (mAP50) تبلغ 43.6% على مجموعة بيانات VisDrone2019—أعلى بنسبة 11.4% من الخط الأساسي—مع 4.17 مليون معلمة و24.1 GFLOPs، مما يشير إلى توازن قوي بين الدقة والكفاءة. علاوة على ذلك، تصل إلى 84.5% mAP50 على مجموعة بيانات HIT-UAV بالأشعة تحت الحمراء، مما يبرز قدراتها العامة القوية. بشكل عام، يقدم LRDS-YOLO حلاً واعدًا في الوقت الحقيقي للكشف عن الأجسام الصغيرة في سياقات الطائرات بدون طيار.

طرق

تم إنشاء بيئة التجربة للدراسة على نظام Ubuntu 22.04، باستخدام إصدار Python 3.8.13 وPytorch 1.13.7 مع دعم CUDA 11.7. يتم تفصيل تكوينات الأجهزة الرئيسية ومعلمات النموذج في الجداول 1 و2. تم إجراء التجارب بحجم دفعة قدره 8، تمتد على 300 دورة تدريبية، واستخدمت معدل تعلم قدره 0.01. بالإضافة إلى ذلك، تم تنفيذ حجم صورة تكيفي قدره 640 × 640 بكسل طوال التجارب.

نتائج

تشير نتائج الدراسة إلى أن نموذج LRDS-YOLO يتفوق بشكل كبير في مهام الكشف عن الأجسام ضمن البيئات الحضرية المعقدة، كما يتضح من التحليلات البصرية المقدمة في الأشكال 14 و15 و16. يظهر النموذج قوة ملحوظة وقدرة على التكيف عبر ظروف بيئية متنوعة، بما في ذلك سيناريوهات الإضاءة المختلفة وارتفاعات طيران الطائرات بدون طيار.

من الجدير بالذكر أن نموذج LRDS-YOLO يحافظ على دقة كشف عالية خلال العمليات النهارية والليلية، وكذلك في المناطق ذات الكثافة المرورية العالية. إن قدراته في الكشف عن الأجسام متعددة المقاييس فعالة بشكل خاص، مما يسمح بالتحديد الدقيق وتحديد الموقع للأجسام الصغيرة الموزعة بكثافة، بغض النظر عن ارتفاع الطائرة بدون طيار أو منظورها. تؤكد هذه النتائج على إمكانيات النموذج للتطبيقات العملية في المراقبة والرصد الحضري.

مناقشة

في هذا القسم، يقدم المؤلفون تقدمًا كبيرًا في الكشف عن الأجسام الصغيرة من خلال تقديم شبكة هرم الميزات المعاد معايرتها (Re-Calibration FPN) وآلية DyHead. تعالج شبكة Re-Calibration FPN تحديات تكرار الميزات وعدم التناسق الدلالي التي تنشأ من دمج الميزات السطحية والعميقة. من خلال تنفيذ وحدة تجميع الحدود الانتقائية (SBA)، يدمج التصميم بشكل انتقائي إشارات الحدود في التمثيلات الدلالية، مما يحسن من حدود الأجسام ويزيد من دقة تحديد الموقع. تعزز هذه الطريقة المبتكرة دمج الميزات الهرمية، مما يكون مفيدًا بشكل خاص للكشف عن الأجسام الصغيرة، التي غالبًا ما تعاني من تمثيل غير كافٍ في النماذج التقليدية.

بالإضافة إلى ذلك، تقدم آلية DyHead تعديلات ديناميكية على هيكل ورؤوس الكشف، مما يسمح بتعديلات تكيفية بناءً على محتوى خريطة الميزات المدخلة ودقتها. تعزز هذه المرونة استخراج الميزات واستراتيجيات الكشف عبر مقاييس متنوعة وخلفيات معقدة، مما يعزز في النهاية قوة النموذج ودقته. يقدم المؤلفون أيضًا آلية انتباه SegNext، التي تركز على مناطق الاهتمام لتخفيف ضوضاء الخلفية وتحسين دقة الكشف عن الأهداف الصغيرة. تسهم هذه التحسينات مجتمعة في نموذج كشف أكثر فعالية، خاصة في سيناريوهات الطائرات بدون طيار، مع الحفاظ على توازن بين حجم النموذج وكفاءة الحوسبة.

Journal: Scientific Reports, Volume: 15, Issue: 1
DOI: https://doi.org/10.1038/s41598-025-07021-6
PMID: https://pubmed.ncbi.nlm.nih.gov/40595033
Publication Date: 2025-07-01
Author(s): Yuqi Han et al.
Primary Topic: Advanced Neural Network Applications

Overview

The section presents an overview of the challenges associated with small object detection in UAV aerial images, highlighting issues such as low contrast, complex backgrounds, and limited computational resources that lead to high miss detection rates and poor localization accuracy. Traditional detection methods often suffer from information loss, inadequate cross-layer feature interaction, and inflexible detection heads, necessitating the development of more effective solutions.

To address these challenges, the authors introduce LRDS-YOLO, a lightweight and efficient model specifically designed for UAV applications. Key innovations include the Light Adaptive-weight Downsampling (LAD) module, which preserves fine-grained features and minimizes information loss, and the Re-Calibration Feature Pyramid Network (Re-Calibration FPN), which enhances multi-scale feature fusion through bidirectional interactions and resolution-aware hybrid attention. Additionally, the SegNext Attention mechanism focuses on targets while reducing background noise, and the dynamic detection head (DyHead) optimizes feature weighting for improved detection robustness. Experimental results demonstrate that LRDS-YOLO achieves a mean Average Precision (mAP50) of 43.6% on the VisDrone2019 dataset—11.4% higher than the baseline—with only 4.17 million parameters and 24.1 GFLOPs, indicating a strong balance between accuracy and efficiency. Furthermore, it reaches 84.5% mAP50 on the HIT-UAV infrared dataset, showcasing its strong generalization capabilities. Overall, LRDS-YOLO presents a promising real-time solution for small object detection in UAV contexts.

Methods

The experimental environment for the study was established on Ubuntu 22.04, utilizing Python version 3.8.13 and Pytorch 1.13.7 with CUDA 11.7 support. Key hardware configurations and model parameters are detailed in Tables 1 and 2. The experiments were conducted with a batch size of 8, spanning 300 training epochs, and employed a learning rate of 0.01. Additionally, an adaptive image size of 640 × 640 pixels was implemented throughout the experiments.

Results

The results of the study indicate that the LRDS-YOLO model significantly outperforms in object detection tasks within complex urban environments, as evidenced by visual analyses presented in Figures 14, 15, and 16. The model demonstrates remarkable robustness and adaptability across varying environmental conditions, including different lighting scenarios and UAV flight altitudes.

Notably, the LRDS-YOLO model maintains high detection accuracy during both daytime and nighttime operations, as well as in high-traffic-density areas. Its multi-scale object detection capabilities are particularly effective, allowing for precise identification and localization of densely distributed small objects, regardless of the UAV’s altitude or perspective. These findings underscore the model’s potential for practical applications in urban surveillance and monitoring.

Discussion

In this section, the authors present significant advancements in small object detection through the introduction of the Re-Calibration Feature Pyramid Network (Re-Calibration FPN) and the DyHead mechanism. The Re-Calibration FPN addresses the challenges of feature redundancy and semantic inconsistency that arise from the fusion of shallow and deep features. By implementing a Selective Boundary Aggregation (SBA) module, the architecture selectively integrates boundary cues into semantic representations, thereby refining object contours and improving localization accuracy. This innovative approach enhances hierarchical feature fusion, particularly beneficial for detecting small objects, which often suffer from inadequate representation in traditional models.

Additionally, the DyHead mechanism introduces dynamic adjustments to the detection head’s structure and parameters, allowing for adaptive modifications based on the input feature map’s content and resolution. This flexibility optimizes feature extraction and detection strategies across varying scales and complex backgrounds, ultimately enhancing model robustness and accuracy. The authors also introduce the SegNext Attention mechanism, which focuses on regions of interest to mitigate background noise and improve detection precision for small targets. Collectively, these enhancements contribute to a more effective detection model, particularly in UAV scenarios, while maintaining a balance between model size and computational efficiency.

شارك: