التقدم في تقسيم الكائنات: التقنيات، المقاييس والتطبيقات في رؤية الكمبيوتر
Advances in instance segmentation: Technologies, metrics and applications in computer vision

شارك:
المجلة: Neurocomputing، المجلد: 625
DOI: https://doi.org/10.1016/j.neucom.2025.129584
تاريخ النشر: 2025-01-27
المؤلف: José M. Molina وآخرون
الموضوع الرئيسي: تطبيقات الشبكات العصبية المتقدمة

نظرة عامة

تقدم هذه القسم نظرة عامة على التحسينات التي تم إجراؤها على المقالة بعد قبولها للنشر. تشمل هذه التحسينات إضافة صفحة غلاف، بيانات وصفية، وتعديلات في التنسيق تهدف إلى تحسين قابلية القراءة. ومع ذلك، من المهم ملاحظة أن هذه النسخة ليست النسخة النهائية من السجل، مما يشير إلى أنه قد تحدث مراجعات إضافية قبل النشر النهائي.

مقدمة

ت outlines مقدمة هذه الورقة البحثية أهمية تقسيم الكائنات في مجال رؤية الكمبيوتر، مع التأكيد على دورها في تقديم تفسيرات دقيقة ومفصلة للمحتوى المرئي في الصور ومقاطع الفيديو. يعمل تقسيم الكائنات عند تقاطع اكتشاف الكائنات وتقسيم المعاني، بهدف عدم التعرف على الكائنات فحسب، بل أيضًا لتحديد كل كائن بشكل مميز. تتضمن هذه العملية تحديد مناطق الاهتمام (ROIs) من خلال علامات مختلفة مثل الخطوط الخارجية والصناديق المحيطة، حيث يعمل اكتشاف الكائنات كعنصر أساسي لمهام تقسيم الكائنات. تسلط الورقة الضوء على التحديات التي تواجهها التقنيات الحالية، خاصة في التمييز بين الكائنات المتشابهة، وتلاحظ التقدم في تقسيم المعاني، خاصة في التطبيقات مثل القيادة الذاتية.

علاوة على ذلك، تناقش المقدمة تعقيدات تقسيم الكائنات في الفيديو، الذي يمد تقسيم الصور إلى المجال الزمني، مما يتطلب كل من تقسيم وتتبع الكائنات عبر الإطارات. تصنف الأساليب الحالية إلى تقسيم الكائنات في الفيديو التلقائي (AVOS)، تقسيم الكائنات في الفيديو شبه المراقب (SVOS)، وتقسيم الكائنات في الفيديو التفاعلي (IVOS)، كل منها بأساليب وتطبيقات مميزة. تتناول الورقة أيضًا التطبيقات المتنوعة لتقسيم الكائنات عبر مجالات مختلفة، بما في ذلك الرعاية الصحية، علم الفلك، والأنظمة الذاتية، مع الاعتراف بالتحديات الشائعة مثل الحجب والظروف المتغيرة التي تواجهها هذه التطبيقات. تم هيكلة الأقسام التالية من الورقة لتفصيل طرق تقسيم الكائنات، مقاييس التقييم، ومجموعات البيانات ذات الصلة، culminating في مناقشة النتائج والاستنتاجات.

طرق

يتناول القسم الخاص بـ “طرق” في الورقة البحثية مجموعة متنوعة من النماذج لتقسيم الكائنات، مصنفة إياها إلى طرق مراقبة، غير مراقبة، شبه مراقبة، وطرق التعلم المعزز. يتم التأكيد على التعلم المراقب بسبب نجاحه المستمر، على الرغم من التحديات التي تتطلب مجموعات بيانات موسومة كبيرة. تقدم الورقة تصنيفًا لطرق تقسيم الكائنات المراقبة، والتي تصنف إلى طرق من مرحلة واحدة، مرحلتين، وطرق متعددة المراحل. تهيمن طرق المرحلتين، التي تهيمن على هذا المجال، على نتائج اكتشاف الكائنات لإبلاغ مهام التقسيم ويمكن تقسيمها أيضًا إلى طرق من الأعلى إلى الأسفل ومن الأسفل إلى الأعلى. تبدأ طرق من الأعلى إلى الأسفل باكتشاف الكائنات لإنشاء صناديق محيطة، بينما تجمع طرق من الأسفل إلى الأعلى البكسلات لتشكيل الكائنات.

تُبرز طرق المرحلة الواحدة لكفاءتها، حيث تقوم بالتعرف والتقسيم في تمريرة واحدة. تصنف هذه الطرق إلى طرق قائمة على المراسي وطرق خالية من المراسي. تستخدم الطرق القائمة على المراسي، مثل Mask SSD و YOLACT++، صناديق مرجعية محددة مسبقًا للتنبؤات، بينما تعتمد الطرق الخالية من المراسي، مثل CenterMask و SOLOv2، على اكتشاف النقاط الرئيسية وتنبؤات القناع المباشرة، مما يوفر مرونة أكبر. تتناول الورقة أيضًا طرق متعددة المراحل التي تقوم بتحسين التنبؤات من خلال عمليات تكرارية، بما في ذلك الهياكل المتسلسلة وطرق المعتمدة على الشبكات العصبية المتكررة، التي تستفيد من الطبيعة التسلسلية للشبكات العصبية المتكررة لتحسين دقة التقسيم. بشكل عام، يوفر القسم نظرة شاملة على أحدث الطرق في تقسيم الكائنات، مع تسليط الضوء على نقاط قوتها وتحدياتها.

مناقشة

ت outlines قسم المناقشة في الورقة التقدمات في منهجيات تقسيم الكائنات، مصنفة إياها إلى طرق من مرحلة واحدة، مرحلتين، وطرق متعددة المراحل. تعزز طرق المرحلة الواحدة، التي تقوم بالتعرف والتقسيم في تمريرة واحدة، الكفاءة من خلال تعلم كلا المهمتين في وقت واحد. في المقابل، تقوم طرق متعددة المراحل بتحسين التنبؤات بشكل تدريجي، مستفيدة من المعلومات المتبادلة بين المراحل لتعزيز دقة التقسيم. يركز القسم على تطور هذه التقنيات، بدءًا من طرق المرحلة الواحدة التي تستخدم استراتيجيات قائمة على المراسي وتقنيات النافذة المنزلقة، والتقدم إلى أطر متعددة المراحل الأكثر تعقيدًا التي تدمج استراتيجيات مختلفة لاستخراج الميزات وتحسينها.

تشمل النماذج الرئيسية التي تم مناقشتها DeepMask وإصدارها المحسن، SharpMask، التي تستخدم النوافذ المنزلقة والميزات متعددة المقاييس للتقسيم الفعال. تسلط الورقة الضوء أيضًا على شبكة تجميع المسار (PANet) لتحسين تسلسل الميزات وMask R-CNN لدمج مهام الاكتشاف والتقسيم. بالإضافة إلى ذلك، تغطي الطرق من الأسفل إلى الأعلى التي تقوم بتحسين تقسيمات المعاني إلى تنبؤات على مستوى الكائن باستخدام طرق مثل الحقول العشوائية الشرطية (CRFs) والشبكة الديناميكية المثبتة (DIN). يختتم القسم بالتركيز على الطرق القائمة على التجميع التي تجمع البكسلات في كائنات بناءً على الألفة المتعلمة، مما يظهر تنوع وتعقيد تقنيات تقسيم الكائنات الحالية.

Journal: Neurocomputing, Volume: 625
DOI: https://doi.org/10.1016/j.neucom.2025.129584
Publication Date: 2025-01-27
Author(s): José M. Molina et al.
Primary Topic: Advanced Neural Network Applications

Overview

The section provides an overview of the enhancements made to the article following its acceptance for publication. These enhancements include the addition of a cover page, metadata, and formatting adjustments aimed at improving readability. However, it is important to note that this version is not the final version of record, indicating that further revisions may still occur before the definitive publication.

Introduction

The introduction of this research paper outlines the significance of instance segmentation within the domain of computer vision, emphasizing its role in providing detailed and accurate interpretations of visual content in images and videos. Instance segmentation operates at the intersection of object detection and semantic segmentation, aiming to not only identify objects but also to delineate each instance distinctly. This process involves identifying Regions of Interest (ROIs) through various markers such as contours and bounding boxes, with object detection serving as a foundational element for instance segmentation tasks. The paper highlights the challenges faced by current technologies, particularly in differentiating between similar objects, and notes the advancements in semantic segmentation, especially in applications like autonomous driving.

Furthermore, the introduction discusses the complexities of video instance segmentation, which extends image segmentation into the temporal domain, requiring both segmentation and tracking of instances across frames. It categorizes current approaches into Automatic Video Object Segmentation (AVOS), Semi-supervised Video Object Segmentation (SVOS), and Interactive Video Object Segmentation (IVOS), each with distinct methodologies and applications. The paper also touches on the diverse applications of instance segmentation across various fields, including healthcare, astronomy, and autonomous systems, while acknowledging the common challenges such as occlusions and varying conditions that these applications face. The subsequent sections of the paper are structured to detail instance segmentation methods, evaluation metrics, and relevant datasets, culminating in a discussion of findings and conclusions.

Methods

The section on “Methods” in the research paper discusses various paradigms for instance segmentation, categorizing them into supervised, unsupervised, semi-supervised, and reinforcement learning approaches. Supervised learning is emphasized due to its consistent success, despite the challenges of requiring large labeled datasets. The paper presents a taxonomy of supervised instance segmentation methods, which are classified into one-stage, two-stage, and multi-stage methods. Two-stage methods, which dominate the field, utilize object detection results to inform segmentation tasks and can be further divided into top-down and bottom-up approaches. Top-down methods begin with object detection to create bounding boxes, while bottom-up methods cluster pixels to form instances.

Single-stage methods are highlighted for their efficiency, performing detection and segmentation in a single pass. These methods are categorized into anchor-based and anchor-free approaches. Anchor-based methods, such as Mask SSD and YOLACT++, utilize predefined anchor boxes for predictions, while anchor-free methods, like CenterMask and SOLOv2, rely on keypoint detection and direct mask predictions, offering greater flexibility. The paper also details multi-stage methods that refine predictions through iterative processes, including cascade architectures and RNN-based methods, which leverage the sequential nature of RNNs to improve segmentation accuracy. Overall, the section provides a comprehensive overview of the state-of-the-art methods in instance segmentation, highlighting their strengths and challenges.

Discussion

The discussion section of the paper outlines advancements in instance segmentation methodologies, categorizing them into single-stage, two-stage, and multi-stage approaches. Single-stage methods, which perform detection and segmentation in one pass, enhance efficiency by simultaneously learning both tasks. In contrast, multi-stage methods refine predictions incrementally, leveraging mutual information between stages to boost segmentation accuracy. The section emphasizes the evolution of these techniques, starting with single-stage methods that utilize anchor-based strategies and sliding-window techniques, and progressing to more complex multi-stage frameworks that integrate various feature extraction and refinement strategies.

Key models discussed include DeepMask and its enhanced version, SharpMask, which utilize sliding windows and multi-scale features for effective segmentation. The paper also highlights the Path Aggregation Network (PANet) for improving feature hierarchy and Mask R-CNN for combining detection and segmentation tasks. Additionally, it covers bottom-up approaches that refine semantic segmentations into instance-level predictions using methods like Conditional Random Fields (CRFs) and the Dynamically Instantiated Network (DIN). The section concludes with a focus on clustering-based methods that group pixels into instances based on learned affinities, showcasing the diversity and sophistication of current instance segmentation techniques.

شارك: