DOI: https://doi.org/10.1016/j.patcog.2025.111395
تاريخ النشر: 2025-01-24
المؤلف: Jiangzhong Cao وآخرون
الموضوع الرئيسي: تقنيات تحسين الصور
نظرة عامة
تقدم البحث إطار عمل جديد لتحسين الصور تحت الماء (UIE) يدمج وحدة فقدان الإدراك للغة والصورة القائمة على التعلم المتباين (CLIP) وتنظيم التعلم المتدرج لتحسين الجودة الإدراكية للصور تحت الماء. غالبًا ما تتجاهل طرق التعلم العميق التقليدية لتحسين الصور تحت الماء الإدراك البصري البشري وتفرض قيودًا غير كافية، مما يؤدي إلى نتائج تحسين دون المستوى الأمثل. لمعالجة هذه العيوب، يستفيد الأسلوب المقترح من قدرات استخراج الميزات الدلالية البصرية لنموذج CLIP لتطوير نموذج إدراك مصمم خصيصًا للصور تحت الماء. يقوم هذا النموذج بتقييم جودة الصورة من خلال رسم خرائط أزواج المحفزات، مما يعزز الجودة الإدراكية للصور الناتجة.
بالإضافة إلى ذلك، يتضمن الإطار تنظيم التعلم المتدرج، الذي يستخدم سلبيات متعددة لفرض قيود داخل الفضاء الإدراكي لنموذج CLIP. يقوم هذا العملية بتقييم صعوبة التعلم للسلبيات بمستويات جودة متباينة، مما يضمن تحسينًا متوازنًا يتجنب كل من التحسين الناقص والتحسين المفرط. تظهر النتائج التجريبية الواسعة أن طريقة UIE المقترحة تتفوق على التقنيات الحديثة في الجودة البصرية والتعميم، مما يثبت فعاليتها عبر مجموعات بيانات متنوعة، بما في ذلك تلك التي تفتقر إلى صور مرجعية وتحتوي على تشوهات معقدة.
مقدمة
في السنوات الأخيرة، كان لانتشار الصور تحت الماء تأثير كبير على مجالات مثل علم الأحياء البحرية، والاستكشاف تحت الماء، وعلم الآثار. تعتبر هذه الصور ضرورية لمهام رؤية الآلة، بما في ذلك تصنيف المسطحات المائية ومراقبة حالة المحيط، فضلاً عن توفير تجارب بصرية مثيرة للبيئات تحت الماء. ومع ذلك، غالبًا ما تعاني الصور تحت الماء من تشوهات مثل عدم دقة الألوان، وانخفاض التباين، وضبابية التفاصيل بسبب تلاشي الضوء وتشتته بواسطة الكائنات الدقيقة البحرية. وبالتالي، فإن تقنيات تحسين الصور تحت الماء (UIE) الفعالة ضرورية لتحسين كل من تطبيقات رؤية الآلة والتجارب البصرية البشرية.
يمكن تصنيف طرق UIE الحالية إلى ثلاث فئات: قائمة على الأولويات البصرية، قائمة على النماذج الفيزيائية، وقائمة على التعلم العميق. تعزز طرق الأولويات البصرية جودة الصورة من خلال ضبط قيم البكسل ولكنها تتجاهل عمليات التدهور الفيزيائية. تواجه الطرق القائمة على النماذج الفيزيائية تحديات في تقدير المعلمات البيئية بدقة، مما يؤدي إلى محدودية المتانة. على الرغم من أن طرق التعلم العميق قد أظهرت وعدًا، إلا أنها لا تزال تظهر قيودًا، مثل عدد المعلمات الكبير وعدم التوافق الكافي مع الإدراك البصري البشري. لمعالجة هذه القضايا، يقترح المؤلفون طريقة جديدة لتحسين الصور تحت الماء (UIE) تدمج نموذج إدراك CLIP، الذي يستخدم أزواج المحفزات المتضادة القابلة للتعلم لتوافق أفضل مع تقييمات جودة الصورة مع الإدراك البشري. يتم دمج هذا النموذج في شبكة التحسين كفقدان إدراك، مما يوجه عملية التحسين. بالإضافة إلى ذلك، يقدم المؤلفون تقنية تنظيم التعلم المتدرج التي تضبط ديناميكيًا صعوبة الأمثلة السلبية، مما يحسن عملية التحسين ويعزز الجودة الإدراكية للصور الناتجة.
طرق
في هذا القسم، يحدد المؤلفون الطرق التجريبية المستخدمة لتقييم تقنيات تحسين الصور تحت الماء المقترحة. يستخدمون مجموعات بيانات متعددة، بما في ذلك مجموعة بيانات UEQAB، التي تحتوي على 8000 صورة مع درجات متوسط الرأي (MOS) المرتبطة، حيث تم تخصيص 7200 صورة للتدريب و800 لاختبار نموذج إدراك CLIP. لتدريب واختبار شبكة تحسين الصور تحت الماء (UIE)، يتم استخدام مجموعات بيانات UIEB وU45 وSQUID، مع تكوينات محددة لمجموعات التدريب والاختبار كما هو موضح في النص.
تحدد تفاصيل التنفيذ أن التجارب أجريت باستخدام إطار عمل PyTorch على وحدة معالجة الرسومات NVIDIA TITAN RTX. تم تدريب نموذج إدراك CLIP باستخدام مُحسن الانحدار العشوائي (SGD) بمعدل تعلم قدره 0.002، بينما استخدمت نموذج UIE مُحسن آدم بمعدل تعلم أولي قدره 0.001 واستراتيجية تذويب جيبي. تم تطبيق تقنيات زيادة البيانات، مثل الانعكاس العشوائي، وتم استخراج الميزات الكامنة من طبقات معينة من نموذج VGG-19 المدرب مسبقًا لحساب المسافة. تضمنت مقاييس التقييم لمجموعات البيانات الاختبارية نسبة الإشارة إلى الضوضاء القصوى (PSNR)، ومؤشر التشابه الهيكلي (SSIM)، وتشابه رقعة الصورة الإدراكية المتعلمة (LPIPS) لمجموعات البيانات المرجعية، بينما تم استخدام مقاييس غير مرجعية مثل تقييم جودة صورة اللون تحت الماء (UCIQE) وقياس جودة الصورة تحت الماء (UIQM) لمجموعات البيانات التي تفتقر إلى صور مرجعية. تم أيضًا تقييم نموذج إدراك CLIP المقترح من حيث ارتباطه بالإدراك البصري البشري.
مناقشة
تناقش هذه القسم التقدم في تحسين الصور تحت الماء (UIE) من خلال تقنيات التعلم العميق، مع التركيز بشكل خاص على دمج الشبكات العصبية التلافيفية (CNNs)، والشبكات التنافسية التوليدية (GANs)، ونماذج المحولات. تشمل المساهمات البارزة نموذج WaterNet، الذي يجمع بين المدخلات المحسنة مع خرائط الثقة، وشبكة Ucolor التي تستخدم تضمين الفضاء متعدد الألوان. تستفيد الأساليب القائمة على GAN، مثل FUnIE-GAN وPUGAN، من نمذجة توزيع البيانات لتحسين توليد الصور. على الرغم من هذه التقدمات، غالبًا ما تفشل الطرق الحالية في إنتاج نتائج مرضية إدراكيًا، مما يشير إلى الحاجة إلى مزيد من البحث في التحسينات الموجهة بالإدراك وقيود الصور الزائفة.
كما يبرز البحث دور CLIP (التدريب المسبق المتباين للغة والصورة) في تعزيز UIE من خلال توفير نموذج إدراك يتماشى بشكل أقرب مع التقييم البصري البشري. من خلال استخدام تعلم المحفزات، يمكن لنموذج CLIP تقييم جودة الصورة بشكل أكثر فعالية، مما يعالج الغموض الدلالي الموجود في الطرق التقليدية. بالإضافة إلى ذلك، تهدف إدخال تنظيم التعلم المتدرج إلى تحسين عملية التدريب من خلال استخدام كل من العينات الإيجابية والسلبية بناءً على جودتها الإدراكية، مما يحسن القيود على فضاء الحلول ويقلل من خطر التحسين الناقص أو المفرط في الصور النهائية. بشكل عام، تظهر الطريقة المقترحة تحسينات كبيرة في الجودة الإدراكية واستعادة التفاصيل الهيكلية في الصور تحت الماء مقارنة بالتقنيات الحالية.
DOI: https://doi.org/10.1016/j.patcog.2025.111395
Publication Date: 2025-01-24
Author(s): Jiangzhong Cao et al.
Primary Topic: Image Enhancement Techniques
Overview
The research presents a novel Underwater Image Enhancement (UIE) framework that integrates a Contrastive Language-Image Pre-Training (CLIP) perception loss module and curriculum contrastive regularization to improve the perceptual quality of underwater images. Traditional deep learning methods for UIE often neglect human visual perception and impose insufficient constraints, leading to suboptimal enhancement results. To address these shortcomings, the proposed method leverages the CLIP model’s visual semantic feature extraction capabilities to develop a perception model tailored for underwater images. This model evaluates image quality through a prompt pair mapping, enhancing the perceptual quality of the output images.
Additionally, the framework incorporates curriculum contrastive regularization, which utilizes multiple negatives to impose constraints within the CLIP perceptual space. This process assesses the learning difficulty of negatives with varying quality levels, ensuring a balanced enhancement that avoids both under-enhancement and over-enhancement. Extensive experimental results demonstrate that the proposed UIE method surpasses state-of-the-art techniques in visual quality and generalization, proving effective across various datasets, including those without reference images and with complex distortions.
Introduction
In recent years, the proliferation of underwater images has significantly impacted fields such as marine biology, underwater exploration, and archaeology. These images are crucial for machine vision tasks, including water body classification and ocean condition monitoring, as well as for providing engaging visual experiences of underwater environments. However, underwater images often suffer from distortions like color inaccuracies, low contrast, and blurred details due to light attenuation and scattering by marine microorganisms. Consequently, effective Underwater Image Enhancement (UIE) techniques are essential for improving both machine vision applications and human visual experiences.
Current UIE methods can be classified into three categories: visual prior-based, physical model-based, and deep learning-based approaches. Visual prior methods enhance image quality by adjusting pixel values but neglect the physical degradation processes. Physical model-based methods face challenges in accurately estimating environmental parameters, leading to limited robustness. Although deep learning methods have shown promise, they still exhibit limitations, such as large parameter counts and inadequate alignment with human visual perception. To address these issues, the authors propose a novel UIE method that integrates a CLIP perception model, which utilizes learnable antonymic prompt pairs to better align image quality assessments with human perception. This model is incorporated into the enhancement network as a perception loss, guiding the enhancement process. Additionally, the authors introduce a curriculum contrastive regularization technique that dynamically adjusts the difficulty of negative examples, thereby refining the enhancement process and improving the perceptual quality of the resulting images.
Methods
In this section, the authors outline the experimental methods employed to evaluate their proposed underwater image enhancement techniques. They utilize multiple datasets, including the UEQAB dataset, which contains 8000 images with associated Mean Opinion Scores (MOS), from which 7200 images are designated for training and 800 for testing the CLIP perception model. For training and testing the Underwater Image Enhancement (UIE) network, the UIEB, U45, and SQUID datasets are utilized, with specific configurations for training and testing sets as detailed in the text.
The implementation details specify that experiments were conducted using the PyTorch framework on an NVIDIA TITAN RTX GPU. The CLIP perception model was trained using the Stochastic Gradient Descent (SGD) optimizer with a learning rate of 0.002, while the UIE model employed the Adam optimizer with an initial learning rate of 0.001 and a cosine annealing strategy. Data augmentation techniques, such as random flipping, were applied, and latent features were extracted from specific layers of a pre-trained VGG-19 model for distance computation. Evaluation metrics for the test datasets included Peak Signal-to-Noise Ratio (PSNR), Structural Similarity Index (SSIM), and Learned Perceptual Image Patch Similarity (LPIPS) for reference datasets, while non-reference metrics like Underwater Color Image Quality Evaluation (UCIQE) and Underwater Image Quality Measurement (UIQM) were used for datasets lacking reference images. The proposed CLIP perception model was also assessed for its correlation with human visual perception.
Discussion
The section discusses advancements in underwater image enhancement (UIE) through deep learning techniques, particularly focusing on the integration of Convolutional Neural Networks (CNNs), Generative Adversarial Networks (GANs), and Transformer models. Notable contributions include the WaterNet model, which combines enhanced inputs with confidence maps, and the Ucolor network that utilizes multi-color spatial embedding. GAN-based approaches, such as FUnIE-GAN and PUGAN, leverage data distribution modeling for improved image generation. Despite these advancements, existing methods often fail to produce perceptually pleasing results, indicating a need for further research into perception-guided enhancements and the limitations of pseudo-ground truth images.
The paper also highlights the role of CLIP (Contrastive Language-Image Pre-training) in enhancing UIE by providing a perception model that aligns more closely with human visual assessment. By employing prompt learning, the CLIP model can evaluate image quality more effectively, addressing the semantic ambiguities present in traditional methods. Additionally, the introduction of curriculum contrastive regularization aims to refine the training process by utilizing both positive and negative samples based on their perceptual quality, thus improving the constraints on the solution space and reducing the risk of under- or over-enhancement in the final images. Overall, the proposed method demonstrates significant improvements in perceptual quality and structural detail restoration in underwater images compared to existing techniques.
