تحسين الصور تحت الماء وتصنيف المشاهد المعتمد على GAN باستخدام التعلم الانتقالي
GAN-based underwater image enhancement and scene classification using transfer learning

شارك:
المجلة: PLoS ONE، المجلد: 21، العدد: 3
DOI: https://doi.org/10.1371/journal.pone.0345593
PMID: https://pubmed.ncbi.nlm.nih.gov/41894524
تاريخ النشر: 2026-03-27
المؤلف: Amani Homoud وآخرون
الموضوع الرئيسي: تقنيات تحسين الصور

نظرة عامة

تقدم هذه الورقة البحثية تحليلًا استكشافيًا لتقنيات تحليل الفيديو تحت الماء تهدف إلى تحسين جودة الصورة وتحسين تصنيف الأنواع البحرية. تبدأ المنهجية بمعالجة التحديات التي تطرحها البيئات تحت الماء، مثل انخفاض شدة الضوء وتشويه الألوان، مما يعيق الكشف الدقيق عن الأجسام. يستخدم المؤلفون خوارزمية العالم الرمادي (GW) لتحسين الصورة الأولية، تليها تحليل مقارن لتقنية توازن الهيستوغرام (HE) وتوازن الهيستوغرام التكيفي المحدود التباين (CLAHE) لتقييم فعاليتها في تحسين جودة الصورة تحت الماء. ثم يتم استخدام كشف حواف كاني لاستخراج الميزات البارزة من الصور المحسنة، مما يسهل المهام التصنيفية اللاحقة.

تستفيد الدراسة من التعلم المنقول مع ثلاثة نماذج متقدمة من التعلم العميق: VGG16 وResNet50 وDenseNet121، لتصنيف أنواع بحرية مختلفة، بما في ذلك الأسماك والشعاب المرجانية وسلاحف البحر. ومن الجدير بالذكر أن DenseNet121 ظهرت كأكثر المصنفات دقة. تسلط النتائج الضوء على فعالية خوارزمية GW في التخفيف من تشويه الألوان، وتفوق CLAHE على HE التقليدية، والتطبيق الناجح لكشف حواف كاني لاستخراج الميزات. كما تؤكد الأبحاث على إمكانية تحسين GAN المعزز فائق الدقة (ESRGAN) لجودة الصور المزعجة، مما يساهم بشكل كبير في علوم البحار ورصد البيئة. يُقترح العمل المستقبلي لمعالجة القيود في المنهجيات الحالية، لا سيما في تحسين جودة الصورة تحت الماء وتحسين أداء التصنيف من خلال مجموعات بيانات واستراتيجيات أكثر واقعية لموازنة الفئات.

مقدمة

تسلط المقدمة الضوء على الدور الحيوي للمحيط في توفير الموارد والخدمات البيئية، مع التأكيد على أهمية الكائنات البحرية في عمليات مثل دورة الأكسجين وتنظيم المناخ. وتؤكد على ضرورة جمع بيانات أساسية حول تجمعات الأسماك، بما في ذلك الوفرة، وتوزيع الأنواع، والخصائص السلوكية، والتي تعتبر ضرورية للبحوث البحرية وتوقعات تغير المناخ. ومع ذلك، تم تحديد التحليل اليدوي لبيانات الفيديو تحت الماء من قبل علماء المحيطات كعملية تتطلب جهدًا كبيرًا وذات طابع شخصي، مما يثير الحاجة إلى حلول آلية في كشف الأسماك وتحليل الفيديو.

تناقش هذه الفقرة التحديات التي تواجه معالجة الصور تحت الماء، مثل الدقة المنخفضة والإضاءة السيئة، مما يعقد تصنيف الأنواع البحرية. بينما تعمل تقنيات التحسين التقليدية مثل توازن الهيستوغرام (HE) وCLAHE على تحسين تباين الصورة، فإنها غالبًا ما تضخم الضوضاء في ظروف الإضاءة المنخفضة. أظهرت التطورات الأخيرة في الشبكات التنافسية التوليدية (GANs) وعدًا في تحسين جودة الصورة، على الرغم من أنها قد تقدم تشوهات لونية. تختتم المقدمة بالإشارة إلى أن أساليب التعلم العميق، وخاصة الشبكات العصبية التلافيفية (CNNs)، أصبحت هي الأفضل في التعرف الآلي على الأسماك، مما يسهل التعرف على الأنواع وتتبعها، وبالتالي تقليل العمل اليدوي في تربية الأسماك ورصد البيئة. تهدف هذه الدراسة إلى تحسين جودة الصورة تحت الماء وتحسين تصنيف الأنواع البحرية.

طرق

تستعرض هذه الفقرة طرق التحسين المعتمدة على النماذج المستخدمة لتحسين الصور تحت الماء، والتي تعتبر حاسمة للتطبيقات في المركبات المستقلة، والهواتف الذكية، والمراقبة. يتم مناقشة تقنيات مثل الأولوية في القناة المظلمة (DCP)، التصوير القائم على الاستقطاب، الطرق المعتمدة على الهيستوغرام، الطرق المعتمدة على ريتنيكس، الطرق المعتمدة على الدمج، وطرق التعلم الآلي. تعمل DCP، على وجه الخصوص، على تحسين وضوح الصورة من خلال تقدير وإزالة الضباب بناءً على القيم ذات الشدة المنخفضة في القناة المظلمة للصور الطبيعية. على الرغم من فعاليتها، قد لا تقضي DCP تمامًا على الضباب، مما يؤدي غالبًا إلى تدرجات لونية. يحسن التصوير القائم على الاستقطاب الرؤية من خلال تحليل أنماط استقطاب الضوء، بينما تستفيد الطرق المعتمدة على ريتنيكس من الإدراك البصري البشري لتصحيح التشوهات اللونية. تجمع المنهجيات المعتمدة على الدمج بين صور متعددة أو تقنيات معالجة لتعزيز الرؤية ودقة الألوان.

تشمل المنهجية المقترحة نهجًا منهجيًا لتحليل الفيديو تحت الماء، بدءًا من تفكيك الفيديوهات إلى إطارات فردية، تليها التوضيح وتنظيف البيانات لنماذج الشبكات العصبية التلافيفية (CNN) مثل VGG16 وResNet50 وDenseNet121. يتم استخدام خوارزمية العالم الرمادي (GW) لتصحيح التشوه اللوني وتحسين التباين من خلال تطبيع شدة القنوات الحمراء والخضراء والزرقاء لتحقيق مظهر رمادي محايد. تحسب الخوارزمية متوسط الشدات لكل قناة وتعدلها وفقًا لذلك، على الرغم من أن فعاليتها قد تتناقص في الصور التي تهيمن عليها نغمة لون واحدة. تختتم الفقرة بنتائج تجريبية عددية توضح مقاييس الأداء من حيث الدقة، والاسترجاع، ودرجة F1 لكل نموذج من نماذج التعلم العميق، مما يبرز فعالية الطرق المقترحة في تحليل الفيديو تحت الماء.

مناقشة

في قسم المناقشة هذا، يبرز المؤلفون التقدم والتحديات في معالجة الفيديو تحت الماء، خاصة من خلال تقنيات التعلم العميق. يشيرون إلى أنه بينما تعمل الشبكات التنافسية التوليدية (GANs) مثل SRGAN وESRGAN على تحسين الجودة البصرية، قد تقدم أيضًا أنسجة صناعية تعيق التعرف في البيئات تحت الماء بسبب التباينات مع بيانات التدريب الحقيقية. تؤكد الدراسة على أهمية استخدام نماذج التعلم العميق، وخاصة الشبكات العصبية التلافيفية (CNNs) مثل DenseNet وResNet وVGG، لتصنيف الكائنات البحرية، على الرغم من مجموعات البيانات المحدودة وغير المتوازنة التي تعقد تدريب النموذج وأدائه.

يقدم المؤلفون منهجية جديدة تشمل استكشاف الفيديو، وتحسين الصورة باستخدام خوارزمية العالم الرمادي، والتحليل المقارن لتوازن الهيستوغرام (HE) وCLAHE. يجدون أن CLAHE تتفوق في الحفاظ على تفاصيل الصورة مع تقليل العيوب. بالإضافة إلى ذلك، يتم استخدام كاشف حواف كاني لتعزيز استخراج الميزات، مما يحسن بشكل كبير أداء التصنيف عبر النماذج المختبرة. تظهر النتائج أن DenseNet121 حققت أعلى دقة تدريب بنسبة 99.375%، مع متابعة قريبة من ResNet50، مما يشير إلى فعالية التعلم المنقول في تصنيف الأنواع البحرية بدقة. تختتم الورقة بدعوة لمزيد من استكشاف هذه التقنيات لتعزيز تحليل الفيديو تحت الماء ودقة التصنيف.

Journal: PLoS ONE, Volume: 21, Issue: 3
DOI: https://doi.org/10.1371/journal.pone.0345593
PMID: https://pubmed.ncbi.nlm.nih.gov/41894524
Publication Date: 2026-03-27
Author(s): Amani Homoud et al.
Primary Topic: Image Enhancement Techniques

Overview

This research paper presents an exploratory analysis of underwater video analysis techniques aimed at enhancing image quality and improving the classification of marine species. The methodology begins with addressing the challenges posed by underwater environments, such as reduced light intensity and color distortion, which hinder accurate object detection. The authors employ the Gray World (GW) algorithm for initial image enhancement, followed by a comparative analysis of Histogram Equalization (HE) and Contrast Limited Adaptive Histogram Equalization (CLAHE) to assess their effectiveness in improving underwater image quality. Canny Edge Detection is then utilized to extract prominent features from the enhanced images, facilitating subsequent classification tasks.

The study leverages transfer learning with three advanced deep learning models: VGG16, ResNet50, and DenseNet121, to classify various marine species, including fish, coral reefs, and sea turtles. Notably, DenseNet121 emerged as the most accurate classifier. The findings highlight the effectiveness of the GW algorithm in mitigating color distortion, the superiority of CLAHE over traditional HE, and the successful application of Canny edge detection for feature extraction. The research also emphasizes the potential of Enhanced Super-Resolution GAN (ESRGAN) to improve the quality of noisy images, contributing significantly to marine science and environmental monitoring. Future work is suggested to address limitations in the current methodologies, particularly in enhancing underwater image quality and improving classification performance through more realistic datasets and strategies for class imbalance.

Introduction

The introduction highlights the ocean’s vital role in providing resources and ecological services, emphasizing the importance of marine organisms in processes such as the oxygen cycle and climate regulation. It underscores the necessity of collecting basic data on fish populations, including abundance, species distribution, and behavioral traits, which are essential for marine research and climate change predictions. However, the manual analysis of underwater video data by oceanographers is identified as a labor-intensive and subjective process, prompting the need for automated solutions in fish detection and video analysis.

The section discusses the challenges faced in underwater image processing, such as low resolution and poor lighting, which complicate the classification of marine species. While traditional enhancement techniques like Histogram Equalization (HE) and Contrast Limited Adaptive Histogram Equalization (CLAHE) improve image contrast, they often amplify noise in low-light conditions. Recent advancements in Generative Adversarial Networks (GANs) have shown promise in enhancing image quality, although they may introduce color distortions. The introduction concludes by noting that deep learning approaches, particularly Convolutional Neural Networks (CNNs), have become the state-of-the-art in automated fish identification, facilitating efficient species recognition and tracking, thereby reducing manual labor in fish farming and ecological monitoring. This research aims to enhance underwater image quality and improve the classification of marine species.

Methods

The section outlines various model-based enhancement methods employed for improving underwater images, which are crucial for applications in autonomous vehicles, smartphones, and surveillance. Techniques such as Dark Channel Prior (DCP), polarization-based imaging, histogram-based methods, retinex-based methods, fusion-based methods, and machine learning approaches are discussed. DCP, in particular, enhances image clarity by estimating and removing haze based on the low-intensity values in the dark channel of natural images. Despite its effectiveness, DCP may not completely eliminate haze, often resulting in color casts. Polarization-based imaging improves visibility by analyzing light polarization patterns, while retinex-based methods leverage human visual perception to correct color distortions. Fusion-based methodologies combine multiple images or processing techniques to enhance visibility and color accuracy.

The proposed methodology involves a systematic approach to underwater video analysis, starting with the decomposition of videos into individual frames, followed by annotation and data cleaning for convolutional neural network (CNN) models like VGG16, ResNet50, and DenseNet121. The Gray World (GW) algorithm is utilized to correct color distortion and enhance contrast by normalizing the intensity of the red, green, and blue channels to achieve a neutral gray appearance. The algorithm calculates mean intensities for each channel and adjusts them accordingly, although its effectiveness may diminish in images dominated by a single color tone. The section concludes with numerical experimental results demonstrating the performance metrics of precision, recall, and F1-score for each deep learning model, highlighting the efficacy of the proposed methods in underwater video analysis.

Discussion

In this discussion section, the authors highlight the advancements and challenges in underwater video processing, particularly through deep learning techniques. They note that while Generative Adversarial Networks (GANs) like SRGAN and ESRGAN enhance visual quality, they may introduce artificial textures that hinder recognition in underwater environments due to discrepancies with real training data. The study emphasizes the importance of using deep learning models, particularly Convolutional Neural Networks (CNNs) such as DenseNet, ResNet, and VGG, for classifying underwater organisms, despite the limited and imbalanced datasets that complicate model training and performance.

The authors present a novel methodology that includes video exploration, image enhancement using the Gray World algorithm, and comparative analysis of histogram equalization (HE) and Contrast Limited Adaptive Histogram Equalization (CLAHE). They find CLAHE superior for preserving image details while reducing artifacts. Additionally, the Canny edge detector is employed to enhance feature extraction, significantly improving classification performance across the models tested. The results demonstrate that DenseNet121 achieved the highest training accuracy of 99.375%, with ResNet50 closely following, indicating the effectiveness of transfer learning in accurately classifying marine species. The paper concludes with a call for further exploration of these techniques to enhance underwater video analysis and classification accuracy.

شارك: