DOI: https://doi.org/10.1109/tvcg.2026.3663425
PMID: https://pubmed.ncbi.nlm.nih.gov/41666063
تاريخ النشر: 2026-02-10
المؤلف: Lang Nie وآخرون
الموضوع الرئيسي: تقنيات استرجاع الصور والفيديو المتقدمة
نظرة عامة
تقدم الورقة RopStitch، وهو إطار عمل غير خاضع للإشراف لتجميع الصور العميقة مصمم لتعزيز كل من المتانة والطبيعية في مهام تجميع الصور. يستخدم RopStitch بنية ذات فرعين تلتقط الميزات الخشنة والدقيقة بشكل منفصل، وتدمجها لتحسين القابلية للتعميم عبر مشاهد العالم الحقيقي المختلفة. تركز الفرع المدرب مسبقًا على التمثيلات غير المتغيرة دلاليًا، بينما يستخرج الفرع القابل للتعلم ميزات تمييز دقيقة، والتي يتم دمجها على مستوى الارتباط باستخدام عامل قابل للتحكم. بالإضافة إلى ذلك، يتناول الإطار الصراع الكامن بين محاذاة المحتوى والحفاظ على الهيكل من خلال تقديم مفهوم الطائرات المثلى الافتراضية. تتضمن هذه الطريقة تقدير معاملات تحليل التحويل واستخدام متنبئ تكراري جنبًا إلى جنب مع قيد تشويه دلالي أدنى لتحديد الطائرة المثلى لتشويه الصور في اتجاهين.
تظهر التجارب الواسعة أن RopStitch يتفوق بشكل كبير على طرق تجميع الصور الحالية، خاصة من حيث متانة المشهد وطبيعة المحتوى. تؤكد النتائج فعالية بنية الفرع المزدوج ومفهوم الطائرة المثلى الافتراضية، مما يظهر أداء RopStitch المتفوق عبر مجموعة متنوعة من مجموعات البيانات والمشاهد. الشيفرة الخاصة بـ RopStitch متاحة للجمهور، مما يسهل المزيد من البحث والتطبيق في مجال تجميع الصور.
الطرق
تستند منهجية RopStitch إلى استراتيجية تشويه من العالمية إلى المحلية، مشابهة لـ UDIS++. في البداية، يتم تقدير تحويل عالمي $H$ من زوج من الصور المرجعية والهدف ($I_{\text{ref}}$ و $I_{\text{tgt}}$) بدقة $H \times W$. يتم تحقيق ذلك من خلال بنية ذات فرعين تنتج $H$، مع تضمين أولوية عالمية ضمنية لتقدير قوي. يتم بعد ذلك تحليل التحويل إلى $H_{\text{ref}}$ و $H_{\text{tgt}}$ لتحديد الطائرة المثلى، تليها توقع التحولات الخاصة بـ Thin Plate Spline (TPS) ($T_{\text{ref}}$ و $T_{\text{tgt}}$) لتشويه مرن ثنائي الاتجاه.
تتكون دالة الهدف للنموذج من ثلاثة مكونات: $L = L_{\text{align}} + w_s L_{\text{shape}} + w_c L_{\text{coef}}$، حيث يتم تعيين الأوزان $w_s$ و $w_c$ إلى 4 و 10، على التوالي. تركز المصطلحات $L_{\text{align}}$ و $L_{\text{shape}}$ على محاذاة المحتوى والحفاظ على الشكل، بينما يكرس $L_{\text{coef}}$ لتحديد الطائرة المثلى. يتم تقديم المناقشات التفصيلية حول هذه المكونات في الأقسام التالية من الورقة.
المناقشة
تسلط قسم المناقشة في الورقة الضوء على تطور تقنيات تجميع الصور، مقارنة بين الطرق التقليدية والنهج الناشئة في التعلم العميق. يعتمد تجميع الصور التقليدي على ميزات مصنوعة يدويًا ويصنف إلى طرق موجهة نحو المحاذاة وطرق الحفاظ على الهيكل. بينما تركز التقنيات الموجهة نحو المحاذاة على محاذاة دقيقة للمحتوى باستخدام نماذج مثل التحويل العالمي والتشويهات القائمة على الشبكة، فإنها غالبًا ما تواجه صعوبات مع التوازي وقد تقدم تشوهات هيكلية. على العكس، تعطي الطرق التي تحافظ على الهيكل الأولوية للنزاهة الهندسية للأشياء البارزة ولكن قد تضحي بدقة المحاذاة، مما يؤدي إلى ظهور عيوب مرئية.
بالمقابل، تستفيد طرق التجميع المعتمدة على التعلم العميق من الشبكات العصبية التلافيفية (CNNs) لاستخراج ميزات دلالية عالية المستوى، مما يعزز المتانة والكفاءة. تطورت هذه الطرق من التعلم الخاضع للإشراف بالكامل إلى التعلم الخاضع للإشراف الضعيف وغير الخاضع للإشراف، مما يعالج التحديات المتعلقة بالتعميم عبر مشاهد متنوعة. تستخدم بنية الفرع المزدوج المقترحة في الدراسة شبكة سيامية لدمج الميزات القابلة للتعلم وغير المتغيرة بشكل فعال، مما يحسن المتانة عبر المجالات. يساهم دمج الطائرات المثلى للتجميع من خلال التشويه الثنائي الاتجاه في تقليل التشوهات الإسقاطية، بينما يضمن مولد المعاملات التكراري الجديد الحد الأدنى من التشويه الدلالي أثناء عملية التجميع. بشكل عام، يظهر الأسلوب المقترح أداءً متفوقًا في سيناريوهات عدم وجود بيانات، متفوقًا على الخوارزميات الحالية في كل من التقييمات الكمية والنوعية.
DOI: https://doi.org/10.1109/tvcg.2026.3663425
PMID: https://pubmed.ncbi.nlm.nih.gov/41666063
Publication Date: 2026-02-10
Author(s): Lang Nie et al.
Primary Topic: Advanced Image and Video Retrieval Techniques
Overview
The paper introduces RopStitch, an unsupervised deep image stitching framework designed to enhance both robustness and naturalness in image stitching tasks. RopStitch employs a dual-branch architecture that captures coarse and fine features separately, integrating them to improve generalizability across various real-world scenes. The pretrained branch focuses on semantically invariant representations, while the learnable branch extracts fine-grained discriminative features, which are merged at the correlation level using a controllable factor. Additionally, the framework addresses the inherent conflict between content alignment and structural preservation by introducing the concept of virtual optimal planes. This approach involves estimating homography decomposition coefficients and utilizing an iterative predictor alongside a minimal semantic distortion constraint to identify the optimal plane for warping images bidirectionally.
Extensive experiments demonstrate that RopStitch significantly outperforms existing image stitching methods, particularly in terms of scene robustness and content naturalness. The findings validate the effectiveness of the dual-branch architecture and the virtual optimal plane concept, showcasing RopStitch’s superior performance across a diverse range of datasets and scenes. The code for RopStitch is publicly available, facilitating further research and application in the field of image stitching.
Methods
The methodology of RopStitch is structured around a global-to-local warping strategy, akin to UDIS++. Initially, a global homography $H$ is estimated from a pair of reference and target images ($I_{\text{ref}}$ and $I_{\text{tgt}}$) of resolution $H \times W$. This is achieved through a dual-branch architecture that outputs $H$, incorporating an implicit universal prior for robust estimation. The homography is then decomposed into $H_{\text{ref}}$ and $H_{\text{tgt}}$ to identify the optimal plane, followed by the prediction of respective Thin Plate Spline (TPS) transformations ($T_{\text{ref}}$ and $T_{\text{tgt}}$) for bidirectional elastic warping.
The objective function for the model comprises three components: $L = L_{\text{align}} + w_s L_{\text{shape}} + w_c L_{\text{coef}}$, where the weights $w_s$ and $w_c$ are set to 4 and 10, respectively. The terms $L_{\text{align}}$ and $L_{\text{shape}}$ focus on content alignment and shape preservation, while $L_{\text{coef}}$ is dedicated to determining the optimal plane. The detailed discussions of these components are provided in subsequent sections of the paper.
Discussion
The discussion section of the paper highlights the evolution of image stitching techniques, contrasting traditional methods with emerging deep learning approaches. Traditional image stitching relies on hand-crafted features and is categorized into alignment-oriented and structure-preserving methods. While alignment-oriented techniques focus on precise content alignment using models like global homography and mesh-based warps, they often struggle with parallax and can introduce structural distortions. Conversely, structure-preserving methods prioritize the geometric integrity of salient objects but may sacrifice alignment accuracy, leading to visible artifacts.
In contrast, deep learning-based stitching methods leverage convolutional neural networks (CNNs) to extract high-level semantic features, enhancing robustness and efficiency. These methods have evolved from fully supervised to weakly-supervised and unsupervised learning paradigms, addressing challenges related to generalization across diverse scenes. The proposed dual-branch architecture in the study utilizes a Siamese network to effectively combine learnable and invariant features, improving cross-domain robustness. The integration of optimal stitching planes through bidirectional warping further minimizes projective distortions, while a novel iterative coefficient generator ensures minimal semantic distortion during the stitching process. Overall, the proposed method demonstrates superior performance in zero-shot scenarios, outperforming existing algorithms in both quantitative and qualitative evaluations.
