شبكة دمج ميزات VGG-ResNet الهجينة لاكتشاف الأجسام في صور السونار الجانبي
A Hybrid VGG-ResNet Feature Fusion Network for Object Detection in Side-Scan Sonar Images

شارك:
المجلة: International Journal of Emerging Research in Engineering Science and Management، المجلد: 5، العدد: 1
DOI: https://doi.org/10.58482/ijeresm.v5i1.1
تاريخ النشر: 2026-02-08
المؤلف: Venkata Lakshmi Keerthi K وآخرون
الموضوع الرئيسي: البحوث في الصوتيات تحت الماء

نظرة عامة

تقدم ورقة البحث شبكة عصبية تلافيفية هجينة (CNN) مصممة لاكتشاف الأجسام في صور السونار الجانبي، والتي تعتبر تحديًا بطبيعتها بسبب مستويات الضوضاء العالية، والدقة المكانية المحدودة، والهياكل المعقدة في قاع البحر. يدمج النموذج نقاط القوة في استخراج الميزات من VGG-16 و ResNet-50 من خلال دمج الميزات، مما يعزز تمييز الأهداف. للتخفيف من المشكلات المتعلقة بندرة البيانات، يستخدم المؤلفون التعلم بالنقل من نماذج مدربة مسبقًا على ImageNet. تشير النتائج التجريبية إلى أن النموذج المقترح يحقق دقة تصنيف إجمالية تبلغ 84.2% ومتوسط دقة (mAP) يبلغ 88.35%، متجاوزًا العديد من طرق الكشف الحالية.

تؤكد النتائج فعالية النموذج في اكتشاف الهياكل المعقدة والمتقاربة في قاع البحر، مما يسهل التعرف الموثوق على الأهداف في مختلف العمليات البحرية، بما في ذلك علم الآثار البحرية، والبحث والإنقاذ، وفحص البنية التحتية. يوضح الاختبار الدقيق للشبكة العصبية التلافيفية الهجينة قدرتها على العمل في بيئات عالية الضوضاء وتحديد مواقع الأجسام بدقة مثل الطائرات والسفن وسط الفوضى الصوتية. يقترح المؤلفون أن الميزات الهيكلية للنموذج يمكن أن تعزز بشكل كبير الكفاءة التشغيلية والسلامة في الأنشطة تحت الماء. ستركز الأبحاث المستقبلية على تحسين الهيكل، وتقييم جدوى النشر في الوقت الحقيقي، وتوسيع تطبيقات النموذج لتشمل مجموعة واسعة من الأهداف والظروف تحت الماء.

مقدمة

تسلط مقدمة ورقة البحث هذه الضوء على أهمية اكتشاف الأجسام في صور السونار الجانبي (SSS) لتطبيقات تحت الماء المختلفة، مثل علم الآثار البحرية، وتدابير مكافحة الألغام، وفحص البنية التحتية. تعتبر الطرق البصرية التقليدية غير فعالة في البيئات تحت الماء بسبب امتصاص الضوء وتشتته، مما يجعل السونار بديلاً أكثر ملاءمة لرسم الخرائط للأجسام المغمورة. ومع ذلك، فإن اكتشاف الأجسام بشكل موثوق في صور SSS يعيقه تحديات مثل الضوضاء، والدقة المنخفضة، والفوضى الناتجة عن آثار قاع البحر. تشير الورقة إلى أن توفر بيانات السونار المعلّمة بشكل محدود وتنوع إعدادات السونار يزيد من تعقيد التعرف التلقائي على الأهداف.

أظهرت التطورات الأخيرة في التعلم العميق، وخاصة من خلال الشبكات العصبية التلافيفية (CNNs)، وعدًا في تحسين اكتشاف الأجسام في البيئات تحت الماء الصعبة. يمكن لهذه النماذج تعلم الميزات الهيكلية مباشرة من البيانات، مما يقلل من الاعتماد على تقنيات معالجة الصور التقليدية. تهدف الأبحاث إلى تطوير هيكل شبكة CNN هجينة تجمع بين VGG-16 و ResNet-50 لتعزيز تمييز الميزات في صور السونار، مع استخدام التعلم بالنقل لمعالجة ندرة البيانات. توضح الورقة مساهماتها، بما في ذلك تقييم مقارن للطريقة المقترحة مقابل الأساليب الحالية في التعلم العميق، وتضع الأساس لاستكشاف مفصل للمنهجية والنتائج في الأقسام التالية.

طرق

توضح هذه القسم المنهجية لاكتشاف الأجسام في صور السونار الجانبي (SSS)، مع الانتقال من الطرق التقليدية إلى نهج التعلم العميق الهجين. كانت التقنيات التقليدية تتضمن المعالجة المسبقة، والتجزئة، والتصنيف، وغالبًا ما تعتمد على طرق قائمة على الميزات ومطابقة القوالب، والتي واجهت صعوبة مع تعقيدات صور السونار، مثل الضوضاء وتغير خصائص الأجسام. دفعت هذه القيود إلى تطوير منهجية جديدة تؤطر مهمة الكشف كمشكلة تحديد جسم واحد، باستخدام نموذج شبكة عصبية تلافيفية هجينة (CNN) يدمج نقاط القوة في هيكلين مدربين مسبقًا: VGG-16 و ResNet-50.

يتكون النموذج المقترح من ثلاث مراحل رئيسية: إعداد مجموعة البيانات، وتصميم هيكل النموذج الهجين، والتدريب. يتم وضع علامات بدقة على مجموعة البيانات، وتستخدم تقنيات المعالجة المسبقة لتحسين جودة الصورة. يجمع النموذج الهجين بين قدرة VGG-16 على استخراج الميزات المكانية مع قدرة ResNet-50 على التعلم العميق من خلال الاتصالات المتبقية. تتوج عملية استخراج الميزات في تمثيل مشترك يتم تغذيته إلى فروع إخراج متوازية للتصنيف وتوقع صناديق الحدود. يستخدم النموذج دالة خسارة مركبة لتحسين كل من دقة التصنيف ودقة تحديد المواقع، بينما يتم استخدام التعلم بالنقل للتخفيف من الإفراط في التكيف، مما يضمن أداءً قويًا عبر ظروف السونار المتغيرة. يهدف هذا النهج المبتكر إلى تعزيز قدرات الكشف عن الأهداف المغمورة في بيئات السونار المعقدة.

نتائج

توضح قسم النتائج تطوير وتقييم هيكل شبكة عصبية تلافيفية هجينة (CNN) لاكتشاف الأجسام، مصممة خصيصًا لصور السونار الجانبي. استخدم التدريب مجموعة بيانات مخصصة تتكون من 2,220 صورة، كل منها موضوعة بعلامات صناديق وحدود وتصنيفات لتعزيز تحديد مواقع الأجسام. يدمج الهيكل عمودًا أساسيًا عالي السعة لاستخراج الميزات مع رأس إخراج محدد للمهام يهدف إلى توقع صناديق الحدود ودرجات الفئات. كانت ضبط المعلمات الفائقة أمرًا حاسمًا خلال التدريب لتقليل دالة خسارة مركبة، مما يمكّن النموذج من تعلم أنماط مكانية وطيفية معقدة ضرورية لتحديد المواقع والتصنيف الدقيق للصورة.

لتقييم أداء النموذج، تم استخدام مجموعة بيانات اختبار منفصلة تتكون من 1,486 صورة، مع ضمان عدم وجود تداخل مع بيانات التدريب. شملت مقاييس التقييم الدقة، والاسترجاع، ودرجة F1، ومتوسط الدقة (mAP)، والتي تقيس مجتمعة فعالية النموذج في اكتشاف وتحديد وتصنيف الأجسام عبر فئات مختلفة، مثل المركبات الجوية والبحرية، والمواد الصلبة متعددة الأضلاع، والحطام المغمور. تؤكد النتائج المبلغ عنها، المستمدة فقط من مجموعة الاختبار غير المرئية، على قوة النموذج وقدراته على التعميم للتطبيقات في العالم الحقيقي.

مناقشة

تسلط المناقشة الضوء على التقدم في اكتشاف الأجسام المعتمد على التعلم العميق المطبق على صور السونار، مع التأكيد على الانتقال من طرق التصنيف التقليدية إلى أطر اكتشاف الأجسام الأكثر تعقيدًا باستخدام الشبكات العصبية التلافيفية (CNNs). كانت الجهود الأولية تركز على التصنيف، ولكن القدرات الفريدة لـ CNNs في تحديد مواقع وتصنيف عدة أجسام أدت إلى اعتماد كل من النماذج ذات المرحلتين، مثل Faster R-CNN و Mask R-CNN، والنماذج ذات المرحلة الواحدة مثل YOLO و SSD. بينما تقدم النماذج ذات المرحلتين دقة تحديد مواقع عالية، فإن تعقيدها الحسابي وزمن الاستجابة يعيقان التطبيقات في الوقت الحقيقي. في المقابل، توفر كاشفات المرحلة الواحدة، وخاصة متغيرات YOLO، توازنًا ملائمًا بين السرعة والدقة، مما يجعلها مناسبة للسيناريوهات التشغيلية في البيئات تحت الماء.

تناقش الورقة أيضًا إمكانيات نماذج الشبكات العصبية العميقة الهجينة، التي تجمع بين نقاط القوة في هياكل مختلفة (مثل VGG، ResNet، Inception) لتعزيز تمثيل الميزات والصلابة في ظروف السونار الصعبة. يمكن أن تعالج هذه النماذج الهجينة بفعالية مشكلات مثل الضوضاء والدقة المنخفضة، والتي تتكرر في صور السونار. تظهر مقاييس التقييم المستخدمة، بما في ذلك تقاطع الاتحاد (IoU)، والدقة، والاسترجاع، ودرجة F1، الأداء التنافسي للنموذج المقترح، حيث حقق دقة إجمالية تبلغ 84.2% ومتوسط دقة (mAP) يبلغ 88.35%. تشير هذه الأداء إلى قدرة النموذج على اكتشاف وتحديد مواقع الأجسام بشكل موثوق في سيناريوهات تحت الماء المعقدة، مما يمهد الطريق للتطبيقات المستقبلية في علم الآثار البحرية، وعمليات البحث والإنقاذ، وفحص البنية التحتية.

Journal: International Journal of Emerging Research in Engineering Science and Management, Volume: 5, Issue: 1
DOI: https://doi.org/10.58482/ijeresm.v5i1.1
Publication Date: 2026-02-08
Author(s): Venkata Lakshmi Keerthi K et al.
Primary Topic: Underwater Acoustics Research

Overview

The research paper presents a hybrid convolutional neural network (CNN) designed for object detection in side-scan sonar imagery, which is inherently challenging due to high noise levels, limited spatial resolution, and complex seabed structures. The model integrates the feature-extraction strengths of VGG-16 and ResNet-50 through feature fusion, enhancing target discrimination. To mitigate issues related to data sparsity, the authors employ transfer learning from ImageNet-pretrained models. Experimental results indicate that the proposed model achieves an overall classification accuracy of 84.2% and a mean Average Precision (mAP) of 88.35%, surpassing several existing detection methods.

The findings underscore the model’s effectiveness in detecting tightly clustered and intricate seabed structures, thereby facilitating reliable target identification in various maritime operations, including marine archaeology, search and rescue, and infrastructure inspections. The rigorous testing of the hybrid CNN demonstrates its capability to operate in high-noise environments and accurately localize objects such as aircraft and ships amidst acoustic clutter. The authors suggest that the model’s hierarchical features can significantly enhance operational efficiency and safety in underwater activities. Future research will focus on refining the architecture, assessing real-time deployment feasibility, and expanding the model’s applicability to a wider array of underwater targets and conditions.

Introduction

The introduction of this research paper highlights the significance of object detection in side-scan sonar (SSS) images for various underwater applications, such as marine archaeology, mine countermeasures, and infrastructure inspection. Traditional optical methods are ineffective in underwater environments due to light absorption and scattering, making sonar a more suitable alternative for mapping submerged objects. However, reliable object detection in SSS images is hindered by challenges such as noise, low resolution, and clutter from seabed artifacts. The paper notes that the limited availability of annotated sonar data and the variability in sonar setups further complicate automated target recognition.

Recent advancements in deep learning, particularly through Convolutional Neural Networks (CNNs), have shown promise in improving object detection in challenging underwater settings. These models can learn hierarchical features directly from data, reducing reliance on traditional image-processing techniques. The research aims to develop a hybrid CNN architecture that combines VGG-16 and ResNet-50 for enhanced feature discrimination in sonar imagery, while also employing transfer learning to address data sparsity. The paper outlines its contributions, including a comparative evaluation of the proposed method against existing deep learning approaches, and sets the stage for a detailed exploration of the methodology and results in subsequent sections.

Methods

The section outlines the methodology for object detection in side-scan sonar (SSS) imagery, transitioning from traditional methods to a hybrid deep learning approach. Traditional techniques involved preprocessing, segmentation, and classification, often relying on feature-based methods and template matching, which struggled with the complexities of sonar imagery, such as noise and varying object characteristics. These limitations prompted the development of a new methodology that frames the detection task as a single-object identification problem, utilizing a hybrid convolutional neural network (CNN) model that integrates the strengths of two pre-trained architectures: VGG-16 and ResNet-50.

The proposed model consists of three main phases: dataset preparation, hybrid model architecture design, and training. The dataset is meticulously annotated, and preprocessing techniques are employed to enhance image quality. The hybrid model combines VGG-16’s ability to extract spatial features with ResNet-50’s capacity for deeper representation learning through residual connections. The feature extraction process culminates in a shared representation that feeds into parallel output branches for classification and bounding box prediction. The model employs a composite loss function to optimize both classification accuracy and localization precision, while transfer learning is utilized to mitigate overfitting, ensuring robust performance across varying sonar conditions. This innovative approach aims to enhance the detection capabilities of submerged targets in complex sonar environments.

Results

The results section outlines the development and evaluation of a hybrid Convolutional Neural Network (CNN) architecture for object detection, specifically tailored for side-scan sonar images. The training utilized a custom dataset comprising 2,220 images, each annotated with bounding boxes and classification labels to enhance object localization. The architecture integrates a high-capacity backbone for feature extraction with a task-specific output head aimed at predicting bounding boxes and class scores. Hyperparameter tuning was critical during training to minimize a composite loss function, thereby enabling the model to learn complex spatial and spectral patterns essential for accurate image localization and classification.

To assess the model’s performance, a separate test dataset of 1,486 images was employed, ensuring no overlap with the training data. The evaluation metrics included Precision, Recall, F1 score, and Mean Average Precision (mAP), which collectively measure the model’s effectiveness in detecting, locating, and classifying objects across various categories, such as aerial and marine vehicles, polygonal solids, and submerged debris. The reported results, derived solely from the unseen test set, confirm the model’s robustness and generalization capabilities for real-world applications.

Discussion

The discussion highlights the advancements in deep learning-based object detection applied to sonar imagery, emphasizing the transition from traditional classification methods to more complex object detection frameworks utilizing Convolutional Neural Networks (CNNs). Initial efforts focused on classification, but the unique capabilities of CNNs to localize and classify multiple objects have led to the adoption of both two-stage models, such as Faster R-CNN and Mask R-CNN, and one-stage models like YOLO and SSD. While two-stage models offer high localization accuracy, their computational complexity and latency hinder real-time applications. In contrast, one-stage detectors, particularly YOLO variants, provide a favorable speed-accuracy trade-off, making them suitable for operational scenarios in underwater environments.

The paper also discusses the potential of hybrid deep neural network models, which combine the strengths of various architectures (e.g., VGG, ResNet, Inception) to enhance feature representation and robustness in challenging sonar conditions. These hybrid models can effectively address issues such as noise and low resolution, which are prevalent in sonar imagery. The evaluation metrics employed, including Intersection over Union (IoU), Precision, Recall, and F1 score, demonstrate the proposed model’s competitive performance, achieving an overall accuracy of 84.2% and a mean Average Precision (mAP) of 88.35%. This performance indicates the model’s capability to reliably detect and localize objects in complex underwater scenarios, paving the way for future applications in marine archaeology, search and rescue operations, and infrastructure inspections.

شارك: