DOI: https://doi.org/10.1371/journal.pone.0341991
PMID: https://pubmed.ncbi.nlm.nih.gov/41616034
تاريخ النشر: 2026-01-30
المؤلف: Minyu Huang وآخرون
الموضوع الرئيسي: تطبيقات الشبكات العصبية المتقدمة
نظرة عامة
تقدم هذه الورقة البحثية نسخة محسنة من نموذج YOLOv11n، تُسمى YOLO الديناميكي متعدد المقاييس والقنوات (DMS-YOLO)، تهدف إلى تحسين اكتشاف المركبات من صور الطائرات بدون طيار (UAV). يتناول المؤلفون التحديات مثل حجم الكائنات الصغيرة، والدقة المنخفضة، والخلفيات المعقدة، وتvariations المقاييس من خلال تقديم عدة وحدات مبتكرة. ومن الجدير بالذكر أن شبكة تعزيز الحواف الديناميكية متعددة المقاييس (DMS-EdgeNet) مصممة لتعزيز قوة النموذج تجاه ميزات الهدف المحلية من خلال التجميع متعدد المقاييس ودمج الميزات التكيفية. بالإضافة إلى ذلك، تم دمج شبكة تجميع DynaScale (DySAN) لتحسين دلالات الحدود وزيادة دقة الاكتشاف للأهداف الضبابية، بينما تم تضمين طبقة الهدف الصغير P2 لتحسين اكتشاف الكائنات الصغيرة بشكل خاص.
تظهر النتائج التجريبية أن DMS-YOLO يتفوق بشكل كبير على خط الأساس YOLOv11n، محققًا تحسينات بنسبة 7.0% و2.9% في متوسط الدقة (mAP) عند عتبات 50 و50-95، على التوالي، على مجموعة بيانات صور حركة الطائرات، و5.1% و3.1% على مجموعة بيانات VisDrone-DET2019. علاوة على ذلك، يكشف إدخال مقياس mAP50-small أن DMS-YOLO يحقق درجة 75.6% لاكتشاف الكائنات الصغيرة، مما يمثل تحسينًا بنسبة 2.9% مقارنة بـ YOLOv11n. تؤكد هذه النتائج على قدرة النموذج المحسنة على اكتشاف الكائنات الصغيرة بدقة، وهو أمر حاسم للاكتشاف الفعال للمركبات المعتمدة على الطائرات بدون طيار.
مقدمة
تسلط المقدمة الضوء على التقدم في تقنيات الطائرات بدون طيار والتصوير عن بُعد، والتي أدت إلى تقليل تكاليف إنتاج الطائرات. ونتيجة لذلك، وجدت تقنيات اكتشاف الأهداف التي تستخدم الصور الجوية تطبيقات واسعة في مجالات مختلفة، بما في ذلك التخطيط الحضري، ورصد الهيدرولوجيا، والتحكم في حركة المرور، واكتشاف المشاة. على الرغم من هذه التقدمات، لا تزال هناك تحديات في اكتشاف الكائنات الصغيرة ضمن الصور الجوية. على وجه الخصوص، غالبًا ما تعاني الكائنات الصغيرة من دقة منخفضة وملمس ضبابي، مما يعقد استخراج المعلومات الدلالية المعنوية بواسطة الشبكات العصبية السطحية. بالإضافة إلى ذلك، غالبًا ما تمتلك هذه الكائنات الصغيرة خصائص فريدة تعيق جهود الاكتشاف بشكل أكبر.
طرق
في هذه الدراسة، تم إجراء التجارب على نظام Linux باستخدام إطار عمل التعلم العميق PyTorch، مع تسريع GPU مقدمة من NVIDIA GeForce RTX 3090 التي تحتوي على 24 جيجابايت من الذاكرة. تم إعداد بيئة البرمجة باستخدام Visual Studio Code، وتم تنفيذها بلغة بايثون. يتم تقديم المواصفات التفصيلية لبيئة التجربة والمعلمات الفائقة المستخدمة في الجداول 1 و2، على التوالي.
تشير هذه الفقرة إلى أن إعداد التجربة تم تصميمه بدقة لتحسين الأداء، على الرغم من عدم تفصيل النتائج والتحليلات المحددة في هذا الاقتباس. من المحتمل أن يتم مناقشة مزيد من الرؤى حول النتائج وآثارها في الأقسام اللاحقة من الورقة.
مناقشة
في هذا القسم، يناقش المؤلفون التحديات والتقدم في اكتشاف الكائنات الصغيرة ضمن الصور الجوية، مع التركيز بشكل خاص على قيود الطرق التقليدية والطرق المعتمدة على التعلم العميق. تكافح الخوارزميات التقليدية، مثل Viola-Jones وHOG+SVM، مع القوة والكفاءة بسبب اعتمادها على ميزات مصممة يدويًا وأحجام نوافذ ثابتة، والتي لا تكفي لمقاييس الكائنات المتغيرة والخلفيات المعقدة. بالمقابل، أظهرت أساليب التعلم العميق، وخاصة تلك التي تستخدم الشبكات العصبية التلافيفية (CNNs)، وعدًا من خلال تعلم تمثيلات ميزات متعددة المستويات تلقائيًا. ومع ذلك، تواجه النماذج الحالية مثل YOLOv11 صعوبات في اكتشاف الكائنات الصغيرة بدقة بسبب مشكلات فقدان الميزات أثناء دمج المقاييس المتعددة وأخطاء التوطين في الصور الضبابية.
لمعالجة هذه التحديات، يقترح المؤلفون DMS-YOLO، وهو نسخة محسنة من YOLOv11n، والتي تتضمن عدة وحدات مبتكرة. تم تصميم وحدة C3K2-DMS-EdgeNet لتعزيز ميزات الحواف ديناميكيًا عبر مقاييس متغيرة، بينما تعمل شبكة تجميع DynaScale (DySAN) على تحسين دمج الميزات الدلالية عالية ومنخفضة المستوى، وخاصة بالنسبة للكائنات الصغيرة ذات كثافة البكسل المنخفضة. بالإضافة إلى ذلك، تم تقديم طبقة هدف صغير جديدة P2 لالتقاط معلومات التردد العالي الضرورية لاكتشاف الأهداف الصغيرة بشكل أفضل. تظهر النتائج التجريبية أن DMS-YOLO يحسن بشكل كبير من أداء الاكتشاف، محققًا درجات أعلى في متوسط الدقة (mAP) مع الحفاظ على توازن بين كفاءة الاكتشاف والدقة. تؤكد النتائج على أهمية معالجة التحديات الفريدة التي تطرحها الصور الجوية في مهام اكتشاف الكائنات الصغيرة.
DOI: https://doi.org/10.1371/journal.pone.0341991
PMID: https://pubmed.ncbi.nlm.nih.gov/41616034
Publication Date: 2026-01-30
Author(s): Minyu Huang et al.
Primary Topic: Advanced Neural Network Applications
Overview
This research paper presents an enhanced version of the YOLOv11n model, termed Dynamic Multi-scale and Channel-scaled YOLO (DMS-YOLO), aimed at improving vehicle detection from unmanned aerial vehicle (UAV) images. The authors address challenges such as small object size, low resolution, complex backgrounds, and scale variations by introducing several innovative modules. Notably, the Dynamic Multi-Scale Edge Enhancement Network (DMS-EdgeNet) is designed to bolster the model’s robustness to local target features through multi-scale pooling and adaptive feature fusion. Additionally, the DynaScale Aggregation Network (DySAN) is integrated to refine boundary semantics and enhance detection accuracy for fuzzy targets, while a P2 small target layer is included to specifically improve small object detection.
Experimental results demonstrate that DMS-YOLO significantly outperforms the YOLOv11n baseline, achieving improvements of 7.0% and 2.9% in mean Average Precision (mAP) at thresholds of 50 and 50-95, respectively, on the Aerial Traffic Images dataset, and 5.1% and 3.1% on the VisDrone-DET2019 dataset. Furthermore, the introduction of the mAP50-small metric reveals that DMS-YOLO achieves a 75.6% score for small object detection, marking a 2.9% improvement over YOLOv11n. These findings underscore the model’s enhanced capability to accurately detect small objects, which is critical for effective UAV-based vehicle detection.
Introduction
The introduction highlights the advancements in domestic drone and remote sensing imaging technologies, which have led to a reduction in production costs for drones. Consequently, target detection technologies utilizing aerial imagery have found widespread applications across various domains, including urban planning, hydrological monitoring, traffic control, and pedestrian detection. Despite these advancements, challenges remain in the detection of small objects within aerial images. Specifically, small objects often suffer from low resolution and blurred textures, complicating the extraction of meaningful semantic information by shallow neural networks. Additionally, these small objects frequently possess unique characteristics that further hinder detection efforts.
Methods
In this study, the experiments were conducted on a Linux system utilizing the PyTorch deep learning framework, with GPU acceleration provided by an NVIDIA GeForce RTX 3090 featuring 24GB of memory. The programming environment was set up using Visual Studio Code, and the implementation was carried out in Python. Detailed specifications of the experimental environment and the hyperparameters employed are presented in Tables 1 and 2, respectively.
The section indicates that the experimental setup was meticulously designed to optimize performance, although specific results and analyses are not detailed in this excerpt. Further insights into the findings and their implications are likely discussed in subsequent sections of the paper.
Discussion
In this section, the authors discuss the challenges and advancements in small object detection within aerial imagery, particularly focusing on the limitations of traditional and deep learning-based methods. Traditional algorithms, such as Viola-Jones and HOG+SVM, struggle with robustness and efficiency due to their reliance on manually designed features and fixed window sizes, which are inadequate for varying object scales and complex backgrounds. In contrast, deep learning approaches, particularly those utilizing convolutional neural networks (CNNs), have shown promise by automatically learning multi-level feature representations. However, existing models like YOLOv11 face difficulties in accurately detecting small objects due to issues with feature loss during multi-scale fusion and localization errors in blurred images.
To address these challenges, the authors propose DMS-YOLO, an enhanced version of YOLOv11n, which incorporates several innovative modules. The C3K2-DMS-EdgeNet module is designed to dynamically enhance edge features across varying scales, while the DynaScale Aggregation Network (DySAN) improves the integration of high and low-level semantic features, particularly for small objects with low pixel density. Additionally, a new P2 small target layer is introduced to better capture high-frequency information essential for detecting small targets. Experimental results demonstrate that DMS-YOLO significantly improves detection performance, achieving higher mean average precision (mAP) scores while maintaining a balance between detection efficiency and accuracy. The findings underscore the importance of addressing the unique challenges posed by aerial imagery in small object detection tasks.
