دراسة حول خوارزمية الكشف عن الأجسام متعددة الأنماط المدمجة بين الأشعة تحت الحمراء والمرئية بناءً على عنق الزجاجة المعلوماتية بين الأنماط والتحول الحد الأدنى من التكرار
A study on infrared-visible fusion multimodal object detection algorithm based on cross-modal information bottleneck and minimum redundancy transformation

شارك:
المجلة: Scientific Reports، المجلد: 16، العدد: 1
DOI: https://doi.org/10.1038/s41598-026-35339-2
PMID: https://pubmed.ncbi.nlm.nih.gov/41807440
تاريخ النشر: 2026-03-10
المؤلف: Weiyan Tan وآخرون
الموضوع الرئيسي: تطبيقات الشبكات العصبية المتقدمة

نظرة عامة

تقدم ورقة البحث إطارًا جديدًا للكشف عن الأجسام متعددة الأنماط في الأشعة تحت الحمراء والمرئية، مع معالجة تحديات كبيرة مثل تقليل تكرار الأنماط ومحاذاة الميزات. يدمج الإطار المقترح بين اختناق المعلومات عبر الأنماط (CIB) وتحويل الحد الأدنى من التكرار (MRT). يستخدم وحدة CIB مسارًا للضغط والتفكيك وإعادة البناء لتعزيز الاتساق عبر الأنماط من خلال الحفاظ الانتقائي على الدلالات المشتركة، بينما تطبق وحدة MRT تحولات هيكلية نادرة لتقليل التكرار وتحسين الوعي بالحدود في المناطق المستهدفة. كما تم تقديم استراتيجية تدريب ذات مرحلتين لاستقرار عملية التمثيل أثناء عزل الأنماط ودمجها.

تظهر التجارب الواسعة على مجموعات بيانات KAIST وLLVIP فعالية النهج المقترح، حيث تحقق تحسينات ملحوظة في دقة الكشف. بشكل محدد، يعزز الأسلوب متوسط الدقة (mAP) في سيناريو KAIST ليلاً من 42.8% إلى 44.1% ويحقق دقة متوسطة عند 75% (AP@75) تبلغ 80.0% في ظروف الإضاءة المنخفضة على LLVIP، متجاوزًا الحالة السابقة الأفضل بنسبة 2.4%. كما يظهر الإطار أداءً قويًا تحت التداخل والاضطرابات الضوئية، مما يشير إلى قابليته العملية في البيئات المعقدة. ستستكشف الأعمال المستقبلية توسيع هذا الإطار لتحليل الفيديو الديناميكي، وتتبع الأجسام المتعددة، والكشف عبر المجالات، بالإضافة إلى دمج تقنيات استخراج المعرفة والتعلم الذاتي المعزز لتحسين التعميم في سيناريوهات النشر الخفيفة.

مقدمة

في مقدمة هذه الورقة البحثية، يبرز المؤلفون الأهمية المتزايدة لدمج الصور تحت الحمراء والمرئية في تعزيز قدرات الكشف عن الأجسام عبر تطبيقات متنوعة، بما في ذلك القيادة الذاتية والمراقبة الذكية. تتفوق الصور تحت الحمراء في ظروف الإضاءة المنخفضة بفضل قدراتها في التصوير الحراري، بينما تسهم الصور المرئية بمعلومات غنية عن القوام والألوان. يمكن أن يعزز الدمج الفعال لهذه الأنماط بشكل كبير من قوة النموذج ودقته في البيئات المعقدة. ومع ذلك، تواجه تقنيات الدمج متعددة الأنماط الحالية تحديات مثل التباينات بين الأنماط، وتكرار المعلومات، وصعوبات في محاذاة الميزات، مما يعيق أداء الكشف، خاصة في السيناريوهات التي تتضمن تداخلات وتغيرات في الإضاءة.

لمعالجة هذه التحديات، يقترح المؤلفون إطارًا جديدًا للكشف عن الأجسام متعددة الأنماط يجمع بين اختناق المعلومات عبر الأنماط (CIB) وتحويل الحد الأدنى من التكرار (MRT). يقدم هذا الإطار آلية ضغط المعلومات وإعادة البناء النادرة أثناء دمج الميزات، تهدف إلى تقليل المعلومات الزائدة، وتحسين المحاذاة الدلالية، وتعزيز القوة التمييزية لمناطق الأجسام من خلال الخرائط النادرة المدركة هيكليًا. يميز الأسلوب المقترح نفسه عن الأساليب الحالية الموجهة نحو المهام من خلال التركيز على تقليل التكرار وإعادة البناء الدلالي، مما يوفر في النهاية تمثيلًا أكثر قابلية للتفسير وقوة في البيئات متعددة الأنماط المعقدة. تشمل المساهمات الرئيسية للورقة تقديم آلية CIB، التي تهدف إلى تحسين الفعالية العامة لدمج الصور تحت الحمراء والمرئية للكشف عن الأجسام.

طرق

تستعرض هذه القسم الطرق المستخدمة في دراسة الكشف عن الأجسام متعددة الأنماط، مع التأكيد على أهمية دمج المعلومات من مصادر متنوعة، مثل الصور المرئية وتحت الحمراء، لتعزيز أداء الكشف تحت ظروف متغيرة. يصنف المؤلفون الأساليب الحالية إلى ثلاثة نماذج: الدمج المبكر من خلال تكديس الميزات المباشر، الدمج الوسيط المدفوع بالاهتمام، والدمج العميق الموجه نحو المهام. يحددون القيود الرئيسية في الأساليب الحالية، بما في ذلك المشكلات المتعلقة بالاستجابات الزائدة وعدم كفاية القابلية للتفسير الهيكلي، والتي تهدف الإطار المقترح إلى معالجتها من خلال اختناق المعلومات عبر الأنماط ووحدة تحويل الحد الأدنى من التكرار. تم تصميم هذا الإطار لفصل الميزات المشتركة والخاصة بالنمط، مما يحسن الاتساق الدلالي والصلابة.

تستخدم إعدادات التجارب بنية RT-DETR، المنفذة في PyTorch، وتتضمن عملية تدريب من مرحلتين على أزواج الصور المتزامنة RGB-IR. تظهر النتائج فعالية الأسلوب المقترح في تعزيز الاتساق عبر الأنماط وكبح الاستجابات غير ذات الصلة، كما يتضح من تصور خرائط الميزات الوسيطة. تكشف التحليلات المقارنة أن استراتيجية الدمج المقترحة تنتج ميزات أكثر تركيزًا دلاليًا وهيكلية، متفوقة على الأساليب التقليدية. بالإضافة إلى ذلك، تشير تحليل حساسية المعلمات إلى أن الأداء الأمثل يتحقق مع قيم محددة لتقليل التكرار وقيود التعامد، مما يبرز أهمية تحقيق توازن بين تصفية الضوضاء والحفاظ على الهيكل في مهام الكشف متعددة الأنماط.

نتائج

يقدم قسم النتائج في الدراسة نتائج من تجارب الإزالة المصممة لتقييم مساهمات الوحدات الفردية ضمن بنية النموذج المقترح. من خلال إزالة أو تغيير مكونات معينة بشكل منهجي، توضح التجارب الأدوار الوظيفية لهذه الوحدات وتوفر رؤى حول قابلية تفسير النموذج واستراتيجيات التحسين بناءً على التركيب الهيكلي.

بالإضافة إلى ذلك، تظهر النتائج النوعية فعالية نموذج الدمج في الكشف عن المشاة عبر سيناريوهات معقدة. تكشف المقارنات البصرية للحدود المتوقعة مقابل المواقع المستهدفة الفعلية أن الأسلوب المقترح يتفوق على عدة نماذج أساسية، بما في ذلك MSDS-RCNN، IAF-NET، DeepFusion، RT-DETR، وYOLOv5، من حيث دقة تحديد المواقع للأجسام والاتساق عبر الأنماط. يحافظ الأسلوب بشكل ملحوظ على حدود المشاة الدقيقة ويقلل من الاكتشافات الخاطئة، خاصة في ظروف الإضاءة الصعبة. تُعزى هذه التحسينات إلى الدمج الفعال للمعلومات عبر الأنماط، الذي يقلل من الضوضاء الخاصة بالنمط، وتطبيق تحويلات الحد الأدنى من التكرار التي تحسن محاذاة الميزات. بشكل عام، تؤكد النتائج على قوة واكتمال الدلالات للإطار المقترح في مهام الإدراك متعددة الأنماط.

مناقشة

تسلط قسم المناقشة في الورقة الضوء على التقدم الكبير في الكشف عن الأجسام متعددة الأنماط، مع التركيز بشكل خاص على دمج الأشعة تحت الحمراء والمرئية. يقدم المؤلفون بنية جديدة للدمج متعددة الأنماط تعالج بفعالية التحديات مثل التباينات بين الأنماط والتكرار. تشمل الابتكارات الرئيسية آلية اختناق المعلومات عبر الأنماط (CIB)، التي تضغط على الإشارات المشتركة بينما تعيد بناء الميزات الخاصة بالنمط، وتحويل الحد الأدنى من التكرار (MRT)، المصمم لتصفية المعلومات الزائدة وتعزيز التمثيل الهيكلي. تعمل هذه المكونات بشكل متكامل لتحسين القوة والدقة في الكشف، كما يتضح من التجارب الواسعة على مجموعات البيانات المرجعية مثل KAIST وLLVIP، التي تظهر أداءً متفوقًا تحت ظروف متغيرة، بما في ذلك الإضاءة المعقدة والتداخل.

كما تنتقد الورقة الأساليب الحالية للدمج المبكر، مشيرة إلى قيودها في التعامل مع العيوب والضوضاء الخاصة بالنمط، مما قد يؤدي إلى الإفراط في التكيف وتدرجات غير مستقرة. بينما أظهرت تقنيات الدمج الوسيطة المدفوعة بالاهتمام وعودًا، لا تزال تواجه صعوبات مع عدم التوازن الشديد بين الأنماط. يهدف الهيكل المقترح إلى التغلب على هذه التحديات من خلال اعتماد استراتيجية تدريب ذات مرحلتين تعزز التعلم التعاوني من خلال عزل الأنماط ودمجها، مما يؤدي في النهاية إلى تحسين المحاذاة الدلالية والقابلية للتفسير الهيكلي في البيئات المعقدة. تؤكد النتائج على ضرورة وجود آليات دمج أكثر مبدئية يمكنها فصل المكونات المشتركة والخاصة بالنمط بفعالية مع الحفاظ على توافق قوي عبر الأنماط.

Journal: Scientific Reports, Volume: 16, Issue: 1
DOI: https://doi.org/10.1038/s41598-026-35339-2
PMID: https://pubmed.ncbi.nlm.nih.gov/41807440
Publication Date: 2026-03-10
Author(s): Weiyan Tan et al.
Primary Topic: Advanced Neural Network Applications

Overview

The research paper presents a novel framework for infrared-visible multimodal object detection, addressing significant challenges such as modality redundancy suppression and feature alignment. The proposed framework integrates a Cross-modal Information Bottleneck (CIB) and a Minimum Redundancy Transformation (MRT). The CIB module utilizes a compress-decompose-reconstruct pathway to enhance cross-modal consistency by selectively preserving shared semantics, while the MRT module applies sparse structural transformations to reduce redundancy and improve boundary-awareness in target regions. A dual-phase training strategy is also introduced to stabilize the representation process during modality isolation and fusion.

Extensive experiments on the KAIST and LLVIP datasets demonstrate the effectiveness of the proposed approach, achieving notable improvements in detection accuracy. Specifically, the method enhances the mean Average Precision (mAP) in the KAIST nighttime scenario from 42.8% to 44.1% and achieves an Average Precision at 75% (AP@75) of 80.0% in low-light conditions on LLVIP, surpassing the previous state-of-the-art by 2.4%. The framework also shows robust performance under occlusion and illumination disturbances, indicating its practical applicability in complex environments. Future work will explore extending this framework to dynamic video analysis, multi-object tracking, and cross-domain detection, as well as integrating knowledge distillation and self-supervised learning to enhance generalization in lightweight deployment scenarios.

Introduction

In the introduction of this research paper, the authors highlight the growing importance of infrared-visible image fusion in enhancing object detection capabilities across various applications, including autonomous driving and intelligent surveillance. Infrared images excel in low-light conditions due to their thermal imaging capabilities, while visible images contribute rich texture and color information. The effective integration of these modalities can significantly enhance model robustness and accuracy in complex environments. However, existing multimodal fusion techniques face challenges such as modality discrepancies, information redundancy, and difficulties in feature alignment, which hinder detection performance, especially in scenarios with occlusion and illumination variations.

To address these challenges, the authors propose a novel multimodal object detection framework that combines a Cross-modal Information Bottleneck (CIB) with a Minimal Redundancy Transformation (MRT). This framework introduces an information compression and sparse reconstruction mechanism during feature fusion, aimed at reducing redundant information, improving semantic alignment, and enhancing the discriminative power of object regions through structure-aware sparse mapping. The proposed method distinguishes itself from existing task-oriented fusion approaches by focusing on redundancy suppression and semantic reconstruction, ultimately providing a more interpretable and robust representation for complex multimodal environments. The main contributions of the paper include the introduction of the CIB mechanism, which aims to improve the overall effectiveness of infrared-visible image fusion for object detection.

Methods

The section outlines the methods employed in a study on multimodal object detection, emphasizing the significance of integrating information from diverse sources, such as visible and infrared imagery, to enhance detection performance under varying conditions. The authors categorize existing methodologies into three paradigms: early fusion through direct feature stacking, attention-driven intermediate fusion, and task-oriented deep fusion. They identify key limitations in current approaches, including issues with redundant responses and insufficient structural interpretability, which their proposed framework aims to address through a cross-modal information bottleneck and a minimum redundancy transformation module. This framework is designed to disentangle shared and modality-specific features, thereby improving semantic consistency and robustness.

The experimental setup utilizes the RT-DETR architecture, implemented in PyTorch, and involves a two-stage training process on synchronized RGB-IR image pairs. The results demonstrate the effectiveness of the proposed method in enhancing cross-modal consistency and suppressing irrelevant responses, as evidenced by the visualization of intermediate feature maps. The comparative analysis reveals that the proposed fusion strategy yields more semantically concentrated and structured features, outperforming traditional methods. Additionally, a hyperparameter sensitivity analysis indicates that optimal performance is achieved with specific values for redundancy suppression and orthogonality constraints, highlighting the importance of balancing noise filtering and structural preservation in multimodal detection tasks.

Results

The results section of the study presents findings from ablation experiments designed to assess the contributions of individual modules within the proposed model architecture. By systematically removing or altering specific components, the experiments elucidate the functional roles of these modules and provide insights into model interpretability and optimization strategies based on structural composition.

Additionally, qualitative results demonstrate the effectiveness of the fusion model in pedestrian detection across complex scenarios. Visual comparisons of predicted boundaries against actual target locations reveal that the proposed method outperforms several baseline models, including MSDS-RCNN, IAF-NET, DeepFusion, RT-DETR, and YOLOv5, in terms of object localization accuracy and cross-modal consistency. The method notably preserves fine-grained pedestrian boundaries and minimizes false detections, particularly in challenging lighting conditions. These enhancements are attributed to the effective integration of cross-modal information, which mitigates modality-specific noise, and the application of minimum redundancy transformations that improve feature alignment. Overall, the findings affirm the robustness and semantic completeness of the proposed fusion framework in multimodal perception tasks.

Discussion

The discussion section of the paper highlights significant advancements in multimodal object detection, particularly focusing on infrared-visible integration. The authors introduce a novel multimodal fusion architecture that effectively addresses challenges such as modality discrepancies and redundancy. Key innovations include the Cross-modal Information Bottleneck (CIB) mechanism, which compresses shared cues while reconstructing modality-specific features, and the Minimum Redundancy Transformation (MRT), designed to filter out redundant information and enhance structural representation. These components work synergistically to improve robustness and detection accuracy, as evidenced by extensive experiments on benchmark datasets like KAIST and LLVIP, which demonstrate superior performance under varying conditions, including complex illumination and occlusion.

The paper also critiques existing early fusion methods, noting their limitations in handling modality-specific artifacts and noise, which can lead to overfitting and unstable gradients. While attention-driven intermediate fusion techniques have shown promise, they still struggle with severe modality imbalances. The proposed architecture aims to overcome these challenges by employing a dual-phase training strategy that optimizes cooperative learning through modality isolation and fusion, ultimately leading to enhanced semantic alignment and structural interpretability in complex environments. The findings underscore the necessity for more principled fusion mechanisms that can effectively disentangle shared and modality-specific components while maintaining robust cross-modal correspondence.

شارك: