XPoint: بنية قائمة على الفضاء البصري الذاتي الإشراف لتسجيل الصور متعددة الطيف
XPoint: A Self-Supervised Visual-State-Space-Based Architecture for Multispectral Image Registration

شارك:
المجلة: IEEE Access، المجلد: 14
DOI: https://doi.org/10.1109/access.2026.3668631
تاريخ النشر: 2026-01-01
المؤلف: İsmail Can Yağmur وآخرون
الموضوع الرئيسي: تقنيات استرجاع الصور وتصنيفها

نظرة عامة

تقدم ورقة البحث XPoint، وهو إطار عمل ذاتي الإشراف، ومرن لمطابقة الصور متعددة الطيف، يتناول التحديات التي تطرحها التغيرات غير الخطية في الكثافة، وتغيرات الزاوية الشديدة، والندرة المحدودة لمجموعات البيانات المعلّمة. غالبًا ما تركز الطرق التقليدية على اختلافات طيفية معينة، مثل المرئي-الأشعة تحت الحمراء، وتتطلب إشرافًا مكلفًا، مما يحد من قابليتها للتكيف. يتغلب XPoint على هذه القيود من خلال السماح بالتكيف السريع عبر أوضاع متعددة من خلال تصميمه المرن، مما يمكّن المستخدمين من تخصيص المكونات لمهام محددة. يستخدم الإطار مشفر VMamba، مدرب مسبقًا على مهام التقسيم، لاستخراج ميزات قوية ويشمل ثلاثة رؤوس فك مشتركة لاستخراج نقاط الاهتمام، واستخراج الوصف، وانحدار الهوموغرافيا، الذي يفرض قيودًا هندسية لتعزيز الأداء في مهام تسجيل الصور.

تظهر التقييمات التجريبية أن XPoint يتفوق باستمرار أو يتطابق مع الطرق الحالية الرائدة عبر خمسة مجموعات بيانات متعددة الطيف متميزة، بما في ذلك التركيب البصري-الحراري ومجموعات مختلفة من المرئي-الأشعة تحت الحمراء. تختتم الورقة بخطط لتحسينات مستقبلية، تركز على تعزيز الكفاءة الحسابية، وتحسين تصميمات فك التشفير لتحليل نقاط المفاتيح بدقة أكبر، وتحسين رأس انحدار الهوموغرافيا لدمج أفضل مع عملية المطابقة. تهدف هذه التحسينات إلى رفع أداء النموذج وقابليته للتكيف لمجموعة واسعة من التطبيقات.

مقدمة

تناقش مقدمة الورقة أهمية مطابقة الصور، لا سيما من خلال عدسة مطابقة الصور متعددة الأوضاع (MMIM)، التي تدمج البيانات من أوضاع تصوير مختلفة لتعزيز تفسير المشهد. هذه التقنية حيوية لتطبيقات مثل الاستشعار عن بعد والبناء من الحركة، حيث تسمح بربط الهياكل المماثلة عبر الصور الملتقطة تحت ظروف مختلفة. تواجه طرق مطابقة الصور التقليدية، بما في ذلك التقنيات اليدوية مثل SIFT وSURF، قيودًا في السيناريوهات متعددة الأوضاع بسبب كثافتها الحسابية وحساسيتها للتغيرات في الزاوية والملمس. على الرغم من أن أساليب التعلم العميق قد حسنت الأداء في مطابقة الطيف المرئي، إلا أنها لا تزال تواجه صعوبات مع التغيرات غير الخطية في الكثافة في السياقات متعددة الأوضاع.

لمعالجة هذه التحديات، تقدم الورقة XPoint، وهي بنية مرنة مصممة للتدريب التكيفي والتعديل في مطابقة الصور متعددة الأوضاع، خاصة عندما تكون البيانات المعلّمة نادرة. يستخدم XPoint نهجًا ذاتيًا للإشراف لتوليد نقاط مفاتيح زائفة من أزواج الصور المتراصة، مما يقلل بشكل كبير من الحاجة إلى تصنيف البيانات بشكل موسع. تعزز هذه الطريقة التعميم عبر أوضاع متنوعة وتدير بفعالية التغيرات في الزاوية والطيف. تشمل البنية مشفرًا مدربًا مسبقًا ورؤوسًا متعددة لاكتشاف نقاط المفاتيح، والوصف، وانحدار الهوموغرافيا، مع مساهمات رئيسية مثل التكيف الهوموغرافي متعدد الطيف واستخدام مشفرات VMamba لتحسين استخراج الميزات. بشكل عام، يمثل XPoint حلاً قابلاً للتوسع لمطابقة الصور متعددة الأوضاع، حيث يعالج قيود الطرق الحالية بينما يعزز القابلية للتكيف والأداء.

طرق

يتضمن الإعداد التجريبي لهذا البحث استخدام عدة مجموعات بيانات متعددة الأوضاع لتقييم أداء النموذج في السيناريوهات الواقعية. مجموعة البيانات الأساسية هي مجموعة بيانات أزواج الصور متعددة الطيف (البصري-الحراري) من Multipoint، والتي تتكون من 9,340 زوجًا من الصور للتدريب و4,391 للتجربة، تستهدف بشكل خاص بيانات الزراعة الجوية. بالإضافة إلى ذلك، تشمل مجموعة بيانات VEDAI 1,245 زوجًا من الصور القريبة من الأشعة تحت الحمراء (NIR)، مع 1,045 للتدريب و200 للتجربة، تهدف إلى الكشف عن المناطق الحضرية في الصور الجوية.

تعزز مجموعات البيانات الإضافية من ReDFeat التقييم، بما في ذلك مجموعات بيانات VIS-NIR وVIS-IR وVIS-SAR. تحتوي مجموعة بيانات VIS-NIR على 345 زوجًا للتدريب و128 للتجربة من مشاهد متنوعة، بينما تحتوي مجموعة بيانات VIS-IR على 221 زوجًا للتدريب و47 للتجربة من الصور المرئية والأشعة تحت الحمراء. تتكون مجموعة بيانات VIS-SAR من 2,011 زوجًا للتدريب و424 للتجربة من الصور المرئية وصور رادار الفتحة الاصطناعية (SAR). تتيح الطبيعة الشاملة لهذه المجموعات تقييمًا قويًا لقدرات النموذج، حيث تعمل مجموعة تدريب MultiPoint كنموذج أساسي، يتم تحسينه بعد ذلك باستخدام المجموعات الإضافية.

نقاش

في قسم النقاش من الورقة، يقدم المؤلفون عدة تقدمات رئيسية في إطار عملهم لمطابقة الصور متعددة الأوضاع، XPoint. يتضمن الإطار رأس انحدار الهوموغرافيا الذي يعزز اكتشاف نقاط المفاتيح والوصف من خلال تحسين تقدير الهوموغرافيا، مما يؤدي إلى تحسين دقة المطابقة وقابلية التكيف عبر مجموعات بيانات متنوعة. لمواجهة عدم التوازن في مجموعات البيانات، خاصة في سياقات VIS-SAR وVIS-NIR، يتم استخدام خسارة الانتروبيا المتقاطعة الموزونة، مما يعزز أداء النموذج وسط اختلافات طيفية كبيرة. يسمح التصميم المرن لـ XPoint بتحقيق نتائج رائدة عبر إعدادات متعددة الأوضاع، مما يظهر مرونته وفعاليته في تقليل متطلبات البيانات من خلال نهج ذاتي الإشراف.

تستعرض الورقة أيضًا الأعمال ذات الصلة في مطابقة الصور متعددة الأوضاع، مع تسليط الضوء على التحديات التي تطرحها التغيرات في المظهر عبر أجهزة الاستشعار المختلفة. تقارن بين الطرق المعتمدة على الكاشفات، التي لها أهمية أساسية ولكنها تواجه صعوبات في السيناريوهات المعقدة، مع التقنيات الناشئة المدفوعة بالبيانات التي تستفيد من التعلم العميق لتعلم تمثيل قوي. يشير المؤلفون إلى أنه على الرغم من أن التقدم الأخير في الطرق الخالية من الكاشفات يظهر وعدًا، إلا أنها غالبًا ما تواجه قيودًا في الكفاءة الحسابية والتعميم. استجابةً لذلك، يقترح المؤلفون نهجًا قائمًا على الكاشف ذاتي الإشراف يستخدم التكيف الهوموغرافي لتوليد نقاط مفاتيح زائفة، مما يعزز قابلية النموذج للتكيف مع أوضاع متنوعة وتغيرات الزاوية. تهدف هذه الطريقة المبتكرة إلى تحسين قوة اكتشاف نقاط المفاتيح والمطابقة في السياقات متعددة الأوضاع، مع معالجة قيود التقنيات الحالية.

Journal: IEEE Access, Volume: 14
DOI: https://doi.org/10.1109/access.2026.3668631
Publication Date: 2026-01-01
Author(s): İsmail Can Yağmur et al.
Primary Topic: Image Retrieval and Classification Techniques

Overview

The research paper presents XPoint, a self-supervised, modular framework for multispectral image matching that addresses the challenges posed by non-linear intensity variations, extreme viewpoint changes, and the limited availability of labeled datasets. Traditional methods often focus on specific spectral differences, such as visible-infrared, and require costly supervision, which limits their adaptability. XPoint overcomes these limitations by allowing for rapid adaptation across various modalities through its modular design, enabling users to customize components for specific tasks. The framework utilizes a VMamba encoder, pretrained on segmentation tasks, for robust feature extraction and incorporates three joint decoder heads for interest point extraction, descriptor extraction, and homography regression, which imposes geometric constraints to enhance performance in image registration tasks.

Experimental evaluations demonstrate that XPoint consistently outperforms or matches existing state-of-the-art methods across five distinct multispectral datasets, including Optical-Thermal and various visual-infrared combinations. The paper concludes with plans for future improvements, focusing on enhancing computational efficiency, refining decoder designs for more accurate keypoint analysis, and optimizing the homography regression head for better integration with the matching process. These enhancements aim to further elevate the model’s performance and adaptability for a wide range of applications.

Introduction

The introduction of the paper discusses the significance of image matching, particularly through the lens of multimodal image matching (MMIM), which integrates data from various imaging modalities to enhance scene interpretation. This technique is vital for applications such as remote sensing and structure from motion, as it allows for the correlation of similar structures across images captured under different conditions. Traditional image matching methods, including handcrafted techniques like SIFT and SURF, face limitations in multimodal scenarios due to their computational intensity and sensitivity to variations in viewpoint and texture. Although deep learning approaches have improved performance in visible spectrum matching, they still struggle with non-linear intensity variations in multimodal contexts.

To address these challenges, the paper introduces XPoint, a modular architecture designed for adaptive training and fine-tuning in multimodal image matching, particularly when labeled data is scarce. XPoint employs a self-supervised approach to generate pseudo-ground truth keypoints from aligned image pairs, significantly reducing the need for extensive data labeling. This method enhances generalization across diverse modalities and effectively manages viewpoint and spectral variations. The architecture includes a pre-trained encoder and multiple heads for keypoint detection, description, and homography regression, with key contributions such as multispectral homographic adaptation and the use of VMamba encoders for improved feature extraction. Overall, XPoint represents a scalable solution for multimodal image matching, addressing the limitations of existing methods while enhancing adaptability and performance.

Methods

The experimental setup for this research involves the utilization of several multimodal datasets to evaluate model performance in real-world scenarios. The primary dataset is the Multispectral (Optical-Thermal) Image Pair Dataset by Multipoint, comprising 9,340 training and 4,391 testing image pairs, specifically targeting aerial agricultural data. Additionally, the VEDAI dataset includes 1,245 near-infrared (NIR) image pairs, with 1,045 for training and 200 for testing, aimed at urban area detection in aerial imagery.

Further datasets from ReDFeat enhance the evaluation, including VIS-NIR, VIS-IR, and VIS-SAR datasets. The VIS-NIR dataset contains 345 training and 128 testing pairs from various scenes, while the VIS-IR dataset has 221 training and 47 testing pairs of visible and infrared images. The VIS-SAR dataset consists of 2,011 training and 424 testing pairs of visible and synthetic aperture radar (SAR) images. The comprehensive nature of these datasets allows for a robust assessment of the model’s capabilities, with the MultiPoint training set serving as the foundational baseline model, which is then fine-tuned using the additional datasets.

Discussion

In the discussion section of the paper, the authors present several key advancements in their multimodal image matching framework, XPoint. The framework incorporates a Homography Regression Head that enhances keypoint detection and description by refining homography estimation, which leads to improved matching accuracy and adaptability across various datasets. To tackle dataset imbalances, particularly in VIS-SAR and VIS-NIR contexts, a weighted cross-entropy loss is employed, enhancing model performance amidst significant spectral differences. The modular architecture of XPoint allows it to achieve state-of-the-art results across multiple multimodal settings, demonstrating its versatility and effectiveness in reducing data requirements through a self-supervised approach.

The paper also reviews related work in multimodal image matching, highlighting the challenges posed by variations in appearance across different sensors. It contrasts detector-based methods, which have foundational significance but struggle in complex scenarios, with emerging data-driven techniques that leverage deep learning for robust representation learning. The authors note that while recent advancements in detector-free methods show promise, they often face limitations in computational efficiency and generalization. In response, the authors propose a self-supervised detector-based approach that utilizes homographic adaptation to generate pseudo-ground truth keypoints, enhancing the model’s adaptability to various modalities and viewpoint changes. This innovative method aims to improve the robustness of keypoint detection and matching in multimodal contexts, addressing the limitations of existing techniques.

شارك: