الكشف عن الأجسام تحت الماء بكفاءة استنادًا إلى تعزيز الميزات ورأس الكشف عن الانتباه
Efficient underwater object detection based on feature enhancement and attention detection head

شارك:
المجلة: Scientific Reports، المجلد: 15، العدد: 1
DOI: https://doi.org/10.1038/s41598-025-89421-2
PMID: https://pubmed.ncbi.nlm.nih.gov/39966488
تاريخ النشر: 2025-02-18
المؤلف: Xingkun Li وآخرون
الموضوع الرئيسي: تطبيقات الشبكات العصبية المتقدمة

نظرة عامة

تتناول ورقة البحث التحديات المتعلقة بالكشف عن الأجسام تحت الماء، والتي تعيقها عوامل مثل الإضاءة الضعيفة والتشتت المعتمد على الطول الموجي. لتحسين أداء الكشف، يقترح المؤلفون مكونين جديدين: وحدة الترميز الدلالي الجزئي (PSEM) ورأس وزن الأبعاد المنفصل (SDWH). تعزز PSEM دمج الميزات متعددة المقاييس من خلال ترميز المعلومات المحلية للقناة عبر سلسلة من الإضافات البكسلية المتبقية، بينما تستخدم SDWH آلية انتباه لوزن المعلومات الدلالية المكانية والقنوية، مما يحسن من تحديد مواقع الأهداف تحت الماء في المقدمة.

تظهر التجارب الواسعة التي أجريت على مجموعات بيانات UTDAC2020 و RUOD أن دمج PSEM و SDWH في نماذج سلسلة YOLO يعزز بشكل كبير من متوسط الدقة (mAP). على وجه التحديد، حقق YOLOv8n مع هذه التحسينات زيادة بنسبة 2.8% في mAP على UTDAC2020 وتحسناً بنسبة 2.7% على RUOD، متفوقاً على خوارزميات الكشف تحت الماء الحديثة الأخرى بمعدلات mAP تبلغ 82.9% و 80.9%، على التوالي. تشير النتائج إلى أن PSEM و SDWH لا تعزز فقط دقة الكشف ولكنها أيضاً تحافظ على تكاليف حسابية مقبولة، مما يدل على إمكانياتها للتطبيقات العملية في الاستكشاف والحفاظ على البيئة البحرية.

طرق

في قسم “الطرق” من ورقة البحث، يتم تقديم نظرة عامة على هيكل الشبكة، تليها وصف تفصيلي لوحدة الترميز الدلالي الجزئي (PSEM) ورأس وزن الأبعاد المنفصل (SDWH). تستخدم الدراسة مجموعات بيانات متاحة للجمهور، وتحديداً مجموعات بيانات RUOD و UTDAC2020، لأغراض الاختبار والتقييم. جميع البيانات والصور المستخدمة متاحة، مع الحصول على إذن من المؤلفين الأصليين لإدراجها في المنشورات ذات الوصول المفتوح. تلتزم الدراسة بالإرشادات الأخلاقية، حيث تمت الموافقة على جميع بروتوكولات التجارب من قبل شركة شاندونغ تشونغتشينغ للتكنولوجيا الذكية المحدودة، وتم الحصول على موافقة مستنيرة من المشاركين أو أولياء أمورهم.

توضح القسم أيضاً مجموعات البيانات ومقاييس التقييم المستخدمة لتدريب واختبار النموذج، بالإضافة إلى إعدادات المعلمات الفائقة واختيارات المحسنين في قسم “تفاصيل التنفيذ”. يتم مناقشة التجارب المقارنة التي توضح الطرق المقترحة على مجموعات بيانات UTDAC2020 و RUOD، جنباً إلى جنب مع تحليل أداء النموذج في الوقت الحقيقي وتعقيده. تضمن هذه الطريقة المنظمة تقييمًا شاملاً للمنهجيات المقترحة في سياق مجموعات البيانات المختارة.

نقاش

في قسم النقاش، تستعرض الورقة التقدم في الكشف عن الأجسام تحت الماء، مع تسليط الضوء على التطور من طرق الميزات اليدوية التقليدية إلى أساليب التعلم العميق الحديثة. واجهت النماذج المبكرة، مثل تلك التي قدمها تشوانغ وآخرون وغوبتا وآخرون، تحديات بسبب عوامل بيئية مثل اضطراب الضوء وانحراف اللون، مما أعاق دقة الكشف. لقد استفادت التطورات الحديثة من تقنيات التعلم العميق، وخاصة خوارزميات YOLO (You Only Look Once)، التي أظهرت تحسينات كبيرة في قدرات الكشف في الوقت الحقيقي. أدت إدخال نماذج مثل YOLOv4 و YOLOv5 و YOLOv8 إلى تحسين الأداء من خلال تحسينات في هيكل العمود الفقري واستراتيجيات زيادة البيانات، مما يجعلها مناسبة لتعقيدات البيئات تحت الماء.

تؤكد الورقة أيضاً على مساهمات الطرق المقترحة، وحدة الترميز الدلالي الجزئي (PSEM) ورأس وزن الأبعاد المنفصل (SDWH)، التي تعزز دمج الميزات والتعرف على الأهداف في المقدمة في نماذج YOLO. تظهر النتائج التجريبية أن دمج PSEM و SDWH في هياكل YOLO يؤدي إلى تحسينات كبيرة في متوسط الدقة (mAP) عبر مجموعات بيانات مختلفة، بما في ذلك UTDAC2020 و RUOD. تشير النتائج إلى أنه بينما تزيد هذه التحسينات قليلاً من التعقيد الحسابي، فإنها تعزز بشكل كبير من دقة الكشف دون المساس بأداء الوقت الحقيقي، مما يلبي الاحتياجات الحرجة لمهام الكشف عن الأجسام تحت الماء.

Journal: Scientific Reports, Volume: 15, Issue: 1
DOI: https://doi.org/10.1038/s41598-025-89421-2
PMID: https://pubmed.ncbi.nlm.nih.gov/39966488
Publication Date: 2025-02-18
Author(s): Xingkun Li et al.
Primary Topic: Advanced Neural Network Applications

Overview

The research paper addresses the challenges of underwater object detection, which is hindered by factors such as poor illumination and wavelength-dependent scattering. To improve detection performance, the authors propose two novel components: the Partial Semantic Encoding Module (PSEM) and the Split Dimension Weighting Head (SDWH). PSEM enhances multi-scale feature fusion by encoding local channel information through a series of residual pixel-wise additions, while SDWH employs an attention mechanism to weight spatial and channel semantic information, thereby improving the localization of foreground underwater targets.

Extensive experiments conducted on the UTDAC2020 and RUOD datasets demonstrate that integrating PSEM and SDWH into YOLO series models significantly boosts mean Average Precision (mAP). Specifically, YOLOv8n with these enhancements achieves a 2.8% mAP increase on UTDAC2020 and a 2.7% improvement on RUOD, outperforming other state-of-the-art real-time underwater detection algorithms with mAPs of 82.9% and 80.9%, respectively. The findings suggest that PSEM and SDWH not only enhance detection accuracy but also maintain acceptable computational costs, indicating their potential for practical applications in underwater exploration and conservation.

Methods

In the “Methods” section of the research paper, an overview of the network architecture is provided, followed by a detailed description of the proposed Partial Semantic Encoding Module (PSEM) and the Split Dimension Weighting Head (SDWH). The study employs publicly available datasets, specifically the RUOD and UTDAC2020 datasets, for testing and evaluation purposes. All utilized data and images are accessible, with permissions granted by the original authors for inclusion in open-access publications. The research adheres to ethical guidelines, with all experimental protocols approved by Shandong Zhongqing Intelligent Technology Co., Ltd., and informed consent obtained from participants or their guardians.

The section further outlines the datasets and evaluation metrics used for model training and testing, as well as hyperparameter settings and optimizer choices in the “Implementation details” subsection. Comparative experiments demonstrating the proposed methods on the UTDAC2020 and RUOD datasets are discussed, alongside an analysis of the model’s real-time performance and complexity. This structured approach ensures a comprehensive evaluation of the proposed methodologies within the context of the selected datasets.

Discussion

In the discussion section, the paper reviews advancements in underwater object detection, highlighting the evolution from traditional handcrafted feature methods to modern deep learning approaches. Early models, such as those by Chuang et al. and Gupta et al., faced challenges due to environmental factors like light disturbance and color deviation, which hindered detection accuracy. Recent developments have leveraged deep learning techniques, particularly YOLO (You Only Look Once) algorithms, which have shown significant improvements in real-time detection capabilities. The introduction of models like YOLOv4, YOLOv5, and YOLOv8 has enhanced performance through optimizations in backbone architecture and data augmentation strategies, making them suitable for the complexities of underwater environments.

The paper further emphasizes the contributions of the proposed methods, Partial Semantic Encoding Module (PSEM) and Split Dimension Weighting Head (SDWH), which enhance feature fusion and foreground target recognition in YOLO models. Experimental results demonstrate that integrating PSEM and SDWH into YOLO architectures leads to substantial improvements in mean Average Precision (mAP) across various datasets, including UTDAC2020 and RUOD. The findings indicate that while these enhancements slightly increase computational complexity, they significantly boost detection accuracy without compromising real-time performance, thereby addressing the critical needs of underwater object detection tasks.

شارك: