دمج الميزات عبر الأنماط لتسجيل سحابة النقاط بشكل قوي مع هندسة غامضة
Cross-modal feature fusion for robust point cloud registration with ambiguous geometry

شارك:
المجلة: ISPRS Journal of Photogrammetry and Remote Sensing، المجلد: 227
DOI: https://doi.org/10.1016/j.isprsjprs.2025.05.012
تاريخ النشر: 2025-06-05
المؤلف: Zhaoyi Wang وآخرون
الموضوع الرئيسي: المسح ثلاثي الأبعاد والتراث الثقافي

نظرة عامة

في هذه الورقة، يقدم المؤلفون CoFF، وهي طريقة جديدة لدمج الميزات عبر الأنماط تهدف إلى تعزيز تسجيل سحب النقاط الزوجية من خلال دمج كل من البيانات الهندسية من سحب النقاط والمعلومات الإشعاعية من صور RGB. تعالج الطريقة المقترحة قيود التقنيات الحالية التي غالبًا ما تتجاهل استخدام بيانات RGB، خاصة في المناطق الهندسية الغامضة حيث تكون الميزات الهندسية وحدها غير كافية. يستخدم CoFF عملية دمج من مرحلتين، تشمل تخصيص ميزات الصورة على مستوى البكسل لسحب النقاط ثلاثية الأبعاد لإثراء ميزاتها ودمج ميزات الصورة على مستوى الباتش مع ميزات النقاط الفائقة لتحسين دقة المطابقة الخشنة.

تظهر فعالية CoFF من خلال تقييمات واسعة على أربعة مجموعات بيانات مستخدمة على نطاق واسع: 3DMatch و3DLoMatch وIndoorLRS وScanNet++. تشير النتائج إلى أن CoFF يحقق أداءً متفوقًا في التسجيل، مع استدعاءات تسجيل ملحوظة تبلغ 95.9% و81.6% على مجموعتي بيانات 3DMatch و3DLoMatch، على التوالي. تتفوق الطريقة بشكل خاص في السيناريوهات التي تحتوي على ميزات هندسية صعبة، مشروطة بتوافر صور RGB ذات نسيج كافٍ في المناطق المتداخلة. يقترح المؤلفون أن البحث المستقبلي يمكن أن يستكشف تقنيات التقدير المشترك لتحويلات الصورة-سحب النقاط وتحويلات سحب النقاط-سحب النقاط.

مقدمة

تتناول مقدمة الورقة المهمة الحرجة لتسجيل سحب النقاط الزوجية، والتي تتضمن تقدير التحويل اللازم لمحاذاة سحب نقطتين. هذه المهمة ضرورية لتطبيقات متنوعة، بما في ذلك إعادة البناء ثلاثي الأبعاد، والتعرف على الأماكن، وSLAM، والقيادة الذاتية. أدت التقدمات الأخيرة في التعلم العميق إلى تحسين تسجيل سحب النقاط، خاصة في السيناريوهات الصعبة ذات التداخل المنخفض. غالبًا ما تعتمد الطرق التقليدية على اكتشاف النقاط الرئيسية والمطابقة، لكن دقتها قد تكون غير كافية، مما يؤدي إلى جودة تسجيل دون المستوى. بالمقابل، أظهرت الطرق الخالية من الكاشفات التي تستخدم الميزات الهندسية أداءً متفوقًا من خلال الحفاظ على توزيع موحد للنقاط الفائقة، مما يعزز دقة المطابقة.

ومع ذلك، تكافح هذه الطرق المعتمدة على الهندسة مع الهندسات الغامضة، خاصة في السيناريوهات المتماثلة أو المستوية. لمعالجة ذلك، يقترح المؤلفون طريقة جديدة لدمج الميزات عبر الأنماط لتسجيل سحب النقاط، تُسمى CoFF. تتكون هذه الطريقة من عملية من مرحلتين: المرحلة الأولى تدمج ميزات الصورة على مستوى البكسل مع سحب النقاط ثلاثية الأبعاد لتعزيز المعلومات العالمية، بينما تجمع المرحلة الثانية ميزات الصورة على مستوى الباتش مع ميزات النقاط الفائقة لتخفيف الغموض الهندسي أثناء المطابقة الخشنة. تشمل الطريقة المقترحة خمسة مكونات: مستخلصات الميزات لبيانات النقاط والصورة، وحدة دمج عبر الأنماط، واستراتيجية مطابقة من الخشنة إلى الدقيقة، مما يؤدي إلى مُقدّر تسجيل محلي إلى عالمي. يبرز المؤلفون مساهماتهم، بما في ذلك تحليل نقدي للطرق الحالية وتقديم مجموعات بيانات فرعية لتقييم الأداء تحت الغموض الهندسي.

طرق

تستعرض هذه القسم منهجيات مختلفة لتسجيل سحب النقاط ثلاثية الأبعاد، مع التركيز على طرق التسجيل المعتمدة على الباتش المحلي، والطرق التلافيفية بالكامل، والنهج الهجينة. تستخدم تقنيات التسجيل المعتمدة على الباتش المحلي، مثل 3DMatch وPPFNet، موصوفات الميزات المستمدة من الأحياء المحلية لسحب النقاط، مستفيدة من هياكل مثل الشبكات التوأمية وPointNet. ومع ذلك، غالبًا ما تواجه هذه الطرق مشكلات تتعلق بالدقة والقدرة على التقاط الميزات العالمية.

لمعالجة هذه القيود، تستخدم طرق التسجيل التلافيفية بالكامل، بما في ذلك FCGF وD3Feat، الشبكات التلافيفية الهرمية لمعالجة سحب النقاط عالميًا، مما يعزز القدرة على التقاط ميزات أوسع. على الرغم من مزاياها، يمكن أن تفقد هذه الطرق التفاصيل المحلية الدقيقة بسبب تقليل العينة. تدمج النهج الهجينة، مثل RoReg وCoFiNet، مكونات معتمدة على الباتش المحلي في أطر تلافيفية بالكامل، مما يحسن الأداء من خلال دمج استخراج الميزات المحلية والعالمية. تستخدم هذه الطرق استراتيجيات مثل المطابقة من الخشنة إلى الدقيقة والتوافقات بين النقاط الفائقة لتحسين التحويلات بشكل تكراري، مما يوازن في النهاية بين الدقة والكفاءة والعمومية.

كما يقدم قسم المنهجية بيان مشكلة يتعلق بتقدير التحويلات بين سحب النقاط المتداخلة جزئيًا. يصف استخدام موصوفات الميزات التي تم تعلمها من كل من سحب النقاط وصور RGB، مع الافتراض بأنه قد تم تحقيق التسجيل المشترك أثناء المعالجة المسبقة. يتم إجراء تقدير التحويل باستخدام طرق مثل RANSAC أو مُقدّر LGR، مما يبرز دمج تقنيات تسجيل ثلاثية الأبعاد المختلفة لتعزيز عملية التسجيل العامة.

نتائج

في هذا القسم، يقدم المؤلفون نتائج تجريبية واسعة تقيم طريقتهم المقترحة، CoFF، عبر أربعة مجموعات بيانات مفتوحة المصدر. تشير النتائج إلى أن CoFF يتفوق باستمرار على طرق التسجيل الأساسية من حيث معدل مطابقة الميزات (FMR) ومعدل التسجيل (RR). بشكل ملحوظ، على مجموعة بيانات 3DMatch، تحقق CoFF تحسينات لا تقل عن 1.4% في FMR و4.4% في RR مقارنة بالطرق الأخرى. وبالمثل، على مجموعة بيانات 3DLoMatch، تتجاوز CoFF الطريقة الثانية الأفضل، Geo-Transformer، بنسبة 5.3% في FMR و7.6% في RR. على الرغم من أن مكاسب الأداء أصغر على مجموعة بيانات IndoorLRS بسبب الميزات الأكثر ثراءً من سحب النقاط المدمجة، إلا أن CoFF لا تزال تحقق أعلى FMR وتتجاوز CoFiNet بنسبة 0.4% في RR. على مجموعة بيانات ScanNet++، تتفوق CoFF على الطرق الثانية الأفضل بنسبة 2.1% في FMR و5.3% في RR.

يحلل المؤلفون أيضًا توزيع الخطأ باستخدام منحنيات دالة التوزيع التراكمي التجريبية (ECDF)، مما يظهر أن CoFF تحقق باستمرار معدلات خطأ أقل عبر جميع مجموعات البيانات. بالإضافة إلى ذلك، عند التقييم على مجموعات بيانات فرعية تقدم غموضًا هندسيًا، تظهر CoFF تحسينات كبيرة، متفوقة على الطريقة الثانية الأفضل بنسبة 2.5% و14.1% في FMR وRR، على التوالي، على مجموعة بيانات 3DMatch Planar. تسلط النتائج الضوء على متانة CoFF في التعامل مع الظروف الهندسية الصعبة، كما يتضح من النتائج النوعية التي توضح قدرتها على محاذاة سحب النقاط بفعالية حتى في السيناريوهات المعقدة.

نقاش

في قسم النقاش من ورقة البحث، يحلل المؤلفون فعالية طريقتهم المقترحة، CoFF، في تسجيل سحب النقاط، خاصة في السيناريوهات الهندسية الغامضة. يبرزون أن ميزات النقاط الفائقة في CoFF تظهر تميزًا عند حدود المشهد، وهو ما يُعزى إلى دمج باتشات الصورة المحلية. يعزز هذا التميز قدرة الطريقة على تمييز النقاط الكثيفة في المناطق المستوية، مما يحسن دقة التسجيل. يقارن المؤلفون أداء CoFF بأداء GeoTransformer، مشيرين إلى أنه بينما يعتمد الأخير فقط على الميزات الهندسية، فإن دمج CoFF لميزات الصورة على مستوى الباتش يؤدي إلى معدل أعلى بكثير من التوافقات الخشنة الموثوقة (42.2% و77.0% مقابل 6.6% و8.6% لـ GeoTransformer).

علاوة على ذلك، يستكشف المؤلفون تأثير أنواع مختلفة من ميزات الصورة واختيار الميزات متعددة المناظر على نتائج التسجيل. يقترحون أن الجمع بين الميزات على مستوى البكسل والميزات على مستوى الباتش في نهج الدمج من مرحلتين يساهم في متانة عملية التسجيل. يتناول النقاش أيضًا إمكانية دمج CoFF مع طرق تسجيل ثلاثية الأبعاد مختلفة ويقارن كفاءة وقت التشغيل الخاصة بها مع الطرق الأساسية. بشكل عام، تؤكد النتائج على مزايا CoFF في التعامل مع التحديات الهندسية المعقدة، مع الاعتراف أيضًا ببعض القيود التي تستدعي مزيدًا من التحقيق.

القيود

تستعرض ورقة البحث عدة قيود على الطريقة المقترحة. أولاً، تعتمد على توفر صور RGB، مما يحد من قابليتها للتطبيق في السيناريوهات التي تتواجد فيها سحب نقاط ملونة فقط. يسمح هذا الاعتماد على صور RGB باستخدام نماذج مدربة مسبقًا من مجموعات بيانات ثنائية الأبعاد ولكنه قد يعيق الأداء في غياب هذه البيانات. ثانيًا، تفترض الطريقة وجود تسجيل مشترك دقيق بين سحب النقاط وصور RGB، وهي حالة قد لا تتحقق دائمًا. في حالات عدم المحاذاة، قام المؤلفون بتنفيذ خطوة تحقق أثناء المعالجة المسبقة لاكتشاف أخطاء إعادة الإسقاط الكبيرة، موصين بإجراء تصحيحات من خلال معايرة الكاميرا-الماسح أو تقنيات التعلم العميق المتقدمة.

بالإضافة إلى ذلك، قد تواجه الطريقة صعوبات في تمييز الميزات عندما تكون كل من المعلومات الهندسية واللون غامضة، مما قد يؤدي إلى فشل التسجيل. هذه التحديات ليست حصرية لنهج المؤلفين ولكنها شائعة بين الطرق التي تعتمد فقط على البيانات الهندسية واللون. للتخفيف من هذه المشكلات، يمكن أن يؤدي دمج معلومات إضافية، مثل بيانات المستشعر التي تلتقط التغيرات الموضعية والاتجاهية، أو دمج المعرفة الدلالية من نماذج مدربة مسبقًا مثل CLIP، إلى تعزيز متانة الطريقة.

Journal: ISPRS Journal of Photogrammetry and Remote Sensing, Volume: 227
DOI: https://doi.org/10.1016/j.isprsjprs.2025.05.012
Publication Date: 2025-06-05
Author(s): Zhaoyi Wang et al.
Primary Topic: 3D Surveying and Cultural Heritage

Overview

In this paper, the authors introduce CoFF, a novel Cross-modal Feature Fusion method designed to enhance pairwise point cloud registration by integrating both geometric data from point clouds and radiometric information from RGB images. The proposed method addresses the limitations of existing techniques that often neglect the use of RGB data, particularly in geometrically ambiguous regions where geometric features alone are insufficient. CoFF employs a two-stage fusion process, which includes assigning pixel-wise image features to 3D point clouds to enrich their features and integrating patch-wise image features with superpoint features to improve coarse matching accuracy.

The effectiveness of CoFF is demonstrated through extensive evaluations on four widely-used datasets: 3DMatch, 3DLoMatch, IndoorLRS, and ScanNet++. The results indicate that CoFF achieves state-of-the-art registration performance, with notable registration recalls of 95.9% and 81.6% on the 3DMatch and 3DLoMatch datasets, respectively. The method excels particularly in scenarios with challenging geometric features, contingent upon the availability of RGB images with sufficient texture in overlapping regions. The authors suggest that future research could explore joint estimation techniques for image-point cloud and point cloud-point cloud transformations.

Introduction

The introduction of the paper addresses the critical task of pairwise point cloud registration, which involves estimating the transformation needed to align two point clouds. This task is essential for various applications, including 3D reconstruction, place recognition, SLAM, and autonomous driving. Recent advancements in deep learning have improved point cloud registration, particularly in challenging scenarios with low overlap. Traditional methods often rely on keypoint detection and matching, but their precision can be inadequate, leading to suboptimal registration quality. In contrast, detector-free methods that utilize geometric features have shown superior performance by maintaining a uniform distribution of superpoints, which enhances matching precision.

However, these geometry-based methods struggle with ambiguous geometries, particularly in symmetric or planar scenarios. To address this, the authors propose a novel cross-modal feature fusion method for point cloud registration, termed CoFF. This method consists of a two-stage process: the first stage integrates pixel-wise image features with 3D point clouds to enhance global information, while the second stage combines patch-wise image features with superpoint features to mitigate geometric ambiguity during coarse matching. The proposed approach includes five components: feature extractors for point and image data, a cross-modal fusion module, and a coarse-to-fine matching strategy, culminating in a Local-to-Global Registration estimator. The authors highlight their contributions, including a critical analysis of existing methods and the introduction of subset datasets to evaluate performance under geometric ambiguity.

Methods

The section outlines various methodologies for 3D point cloud registration, focusing on local patch-based methods, fully convolutional methods, and hybrid approaches. Local patch-based registration techniques, such as 3DMatch and PPFNet, utilize feature descriptors derived from local neighborhoods of point clouds, leveraging architectures like Siamese networks and PointNet. However, these methods often struggle with issues related to precision and the ability to capture global features.

To address these limitations, fully convolutional registration methods, including FCGF and D3Feat, employ hierarchical convolutional networks to process point clouds globally, enhancing the ability to capture broader features. Despite their advantages, these methods can lose fine-grained local details due to downsampling. Hybrid approaches, such as RoReg and CoFiNet, integrate local patch-based components into fully convolutional frameworks, improving performance by combining local and global feature extraction. These methods utilize strategies like coarse-to-fine matching and superpoint correspondences to refine transformations iteratively, ultimately balancing accuracy, efficiency, and generalizability.

The methodology section also presents a problem statement regarding the estimation of transformations between partially overlapping point clouds. It describes the use of feature descriptors learned from both point clouds and RGB images, with the assumption that co-registration has been achieved during preprocessing. The transformation estimation is performed using methods such as RANSAC or the LGR estimator, highlighting the integration of various 3D registration techniques to enhance the overall registration process.

Results

In this section, the authors present extensive experimental results evaluating their proposed method, CoFF, across four open-source datasets. The results indicate that CoFF consistently outperforms baseline registration methods in terms of Feature Matching Rate (FMR) and Registration Rate (RR). Notably, on the 3DMatch dataset, CoFF achieves improvements of at least 1.4% in FMR and 4.4% in RR compared to other methods. Similarly, on the 3DLoMatch dataset, CoFF surpasses the second-best method, Geo-Transformer, by 5.3% in FMR and 7.6% in RR. Although performance gains are smaller on the In-doorLRS dataset due to richer features from fused point clouds, CoFF still achieves the highest FMR and exceeds CoFiNet by 0.4% in RR. On the ScanNet++ dataset, CoFF outperforms the second-best methods by 2.1% in FMR and 5.3% in RR.

The authors further analyze the error distribution using Empirical Cumulative Distribution Function (ECDF) curves, demonstrating that CoFF consistently achieves lower error rates across all datasets. Additionally, when evaluating on subset datasets that present geometrical ambiguities, CoFF shows significant improvements, outperforming the second-best method by 2.5% and 14.1% in FMR and RR, respectively, on the 3DMatch Planar dataset. The results highlight CoFF’s robustness in handling geometrically challenging conditions, as evidenced by qualitative results that illustrate its capability to align point clouds effectively even in complex scenarios.

Discussion

In the discussion section of the research paper, the authors analyze the effectiveness of their proposed method, CoFF, in point cloud registration, particularly in geometrically ambiguous scenarios. They highlight that CoFF’s superpoint features exhibit distinctiveness at scene borders, which is attributed to the integration of local image patches. This distinctiveness enhances the method’s ability to differentiate dense points in planar regions, thereby improving registration accuracy. The authors compare CoFF’s performance with that of GeoTransformer, noting that while the latter relies solely on geometric features, CoFF’s incorporation of patch-wise image features results in a significantly higher rate of reliable coarse correspondences (42.2% and 77.0% versus GeoTransformer’s 6.6% and 8.6%).

Furthermore, the authors explore the impact of different types of image features and multi-view feature selection on registration outcomes. They suggest that the combination of pixel-wise and patch-wise features in their two-stage fusion approach contributes to the robustness of the registration process. The discussion also touches on the potential for integrating CoFF with various 3D registration methods and compares its runtime efficiency against baseline approaches. Overall, the findings underscore CoFF’s advantages in handling complex geometrical challenges, while also acknowledging certain limitations that warrant further investigation.

Limitations

The research paper outlines several limitations of the proposed approach. Firstly, it is contingent upon the availability of RGB images, which restricts its applicability in scenarios where only colorized point clouds are present. This reliance on RGB images allows the use of pre-trained models from 2D datasets but may hinder performance in the absence of such data. Secondly, the method assumes accurate co-registration between point clouds and RGB images, a condition that may not always be met. In cases of misalignment, the authors have implemented a validation step during preprocessing to detect significant reprojection errors, recommending corrections through camera-scanner calibration or advanced deep learning techniques.

Additionally, the method may encounter difficulties in distinguishing features when both geometric and color information are ambiguous, potentially leading to registration failures. This challenge is not exclusive to the authors’ approach but is common among methods that depend solely on geometry and color data. To mitigate these issues, the integration of additional information, such as sensor data that captures positional and orientational changes, or the incorporation of semantic knowledge from pre-trained models like CLIP, could enhance the robustness of the method.

شارك: