التحقق المقارن من التعرف على مراحل الجراحة، وتقدير نقاط المفاتيح للأدوات، وتقسيم حالات الأدوات في التنظير: نتائج تحدي PhaKIR 2024
Comparative validation of surgical phase recognition, instrument keypoint estimation, and instrument instance segmentation in endoscopy: Results of the PhaKIR 2024 challenge

شارك:
المجلة: Medical Image Analysis، المجلد: 109
DOI: https://doi.org/10.1016/j.media.2026.103945
PMID: https://pubmed.ncbi.nlm.nih.gov/41564633
تاريخ النشر: 2026-01-14
المؤلف: Tobias Rueckert وآخرون
الموضوع الرئيسي: محاكاة وتدريب جراحي

نظرة عامة

تحدد هذه القسم أهمية التعرف الموثوق به وتحديد موقع الأدوات الجراحية في تسجيلات الفيديو التنظيرية، وهو أمر حاسم للتطبيقات في الجراحة الحد الأدنى من التدخل بمساعدة الكمبيوتر والروبوت (RAMIS). يبرز المؤلفون التحديات التي تطرحها الظروف الواقعية ويقترحون أن دمج السياق الجراحي، مثل مرحلة الإجراء، يمكن أن يعزز القوة والقدرة على التفسير. لمعالجة هذه القضايا، نظموا تحدي مرحلة الإجراء الجراحي، ونقطة المفتاح، والتعرف على الأدوات (PhaKIR) في MICCAI 2024، مقدّمين مجموعة بيانات جديدة متعددة المراكز تتكون من ثلاثة عشر فيديو كامل الطول لجراحة استئصال المرارة بالمنظار من ثلاث مؤسسات طبية. تحتوي هذه المجموعة على تعليقات موحدة لثلاث مهام مترابطة: التعرف على مرحلة الجراحة، تقدير نقطة المفتاح للأداة، وتقسيم حالة الأداة، مما يمكّن من تحقيق شامل لتحديد موقع الأداة والسياق الإجرائي.

في الختام، يؤكد المؤلفون على دور مجموعة البيانات في تقييم أداء الخوارزميات وتقييم مقاييس مختلفة. يشيرون إلى أنه بينما تم بحث التعرف على مرحلة الجراحة وتقسيم حالة الأداة بشكل جيد، لا يزال تقدير نقطة المفتاح للأداة غير مستكشف بشكل كافٍ. يبرز تطبيق النمذجة الزمنية في أساليب التعرف على المرحلة أهميتها، ومع ذلك، فإن نقص السياق الزمني في تقدير نقطة المفتاح وتقسيم الحالة يكشف عن مجال لمزيد من الاستكشاف. تم تحديد تحدٍ كبير وهو قابلية تعميم الأساليب المقدمة، حيث لم يظهر أي منها أداءً قويًا عبر مقاطع الفيديو من مراكز مختلفة. يؤكد المؤلفون أن تحدي PhaKIR ومجموعة البيانات تساهم في توفير موارد قيمة للمجال، بهدف إلهام المزيد من البحث والابتكار في تطوير طرق الذكاء الاصطناعي الموثوقة للتطبيقات الجراحية.

مقدمة

تسلط مقدمة هذه الورقة البحثية الضوء على التقدم في الجراحة الحد الأدنى من التدخل بمساعدة الكمبيوتر والروبوت (RAMIS)، لا سيما في الإجراءات التنظيرية. تؤكد على أهمية تحديد المرحلة الجراحية الحالية بدقة، وتقسيم الأدوات الجراحية في الصور التنظيرية، وتحديد مواقع نقاط المفتاح لتقدير أوضاع الأدوات. تعزز هذه القدرات سلامة غرفة العمليات من خلال التحذيرات الحساسة للسياق، وتحسن الإدارة، وتحسن توقعات وقت الإجراء. تنتقد الورقة مجموعات البيانات الحالية من تحديات رؤية التنظير المختلفة (EndoVis)، مشيرة إلى قيودها من حيث بيانات متعددة المراكز، وتفريق فئات الأدوات، وتعقيد بيئات الجراحة الواقعية.

لمعالجة هذه الفجوات، يقدم المؤلفون مجموعة بيانات جديدة تتكون من تسلسلات فيديو كاملة من ثلاثة عشر عملية استئصال مرارة بشرية عبر ثلاثة مراكز طبية. تسهل هذه المجموعة دمج المعلومات الزمنية والظروف الواقعية في تطوير الأساليب، وتشمل ثلاث مهام رئيسية: التعرف على مرحلة الجراحة، تقسيم حالة الأدوات بدقة بكسل، وتحديد مواقع نقاط المفتاح المحددة. تم تنظيم الورقة وفقًا لإرشادات تحديات تحليل الصور الطبية الحيوية (BIAS)، مع فصول منفصلة لكل مهمة وملاحق مفصلة لتنظيم التحدي، ووصف البيانات، وتعليمات المشاركين.

الطرق

في هذا القسم، يقدم المؤلفون نظرة شاملة على الفرق التي شاركت في التحدي ومنهجياتها الخاصة عبر ثلاث مهام متميزة: التعرف على مرحلة الجراحة، تقسيم حالة الأدوات، وتقدير نقطة المفتاح للأداة. ساهم ما مجموعه 14 فريقًا فريدًا، حيث ركز سبعة فرق على التعرف على مرحلة الجراحة، وتسعة على تقسيم حالة الأدوات، واثنان على تقدير نقطة المفتاح للأداة. من الجدير بالذكر أن خمسة فرق شاركت في مهام متعددة، مما يدل على اتساع الخبرة بين التخصصات.

تُلخص المنهجيات التي استخدمها الفرق في الجدول 2، الذي ينظم المشاركين أبجديًا ضمن كل مهمة. تم اشتقاق أوصاف أساليبهم من التقارير المنهجية المقدمة من الفرق، مما يبرز الاستراتيجيات المتنوعة المستخدمة لمعالجة التحديات المقدمة في كل مهمة. تؤكد هذه التنوع على تعقيد المشكلات والحلول المبتكرة المقترحة من قبل المشاركين.

النتائج

يوفر قسم النتائج في الورقة البحثية تقييمًا شاملاً لثلاث مهام تحدي: التعرف على مرحلة الجراحة، تقسيم حالة الأدوات، وتقدير نقطة المفتاح للأداة. يتم تفصيل أداء كل مهمة من خلال النتائج العامة، والتحليلات الإحصائية لاستقرار الترتيب، والتحليلات التكميلية، بما في ذلك تصورات مصفوفة الارتباك للتعرف على المرحلة والنتائج النوعية لتقسيم الحالة وتقدير نقطة المفتاح. من الجدير بالذكر أن التحليل يكشف عن تفاوت في مشاركة الفرق، حيث شارك عدد أكبر بكثير من الفرق في التعرف على مرحلة الجراحة (n = 7) وتقسيم حالة الأدوات (n = 9) مقارنةً بتقدير نقطة المفتاح للأداة (n = 2). يشير هذا إلى اهتمام أكبر في المهام السابقة، على الرغم من التحديات السريرية ذات الصلة التي تطرحها تقديرات نقطة المفتاح، مثل تحديد مواقع أجزاء الأدوات المحجوبة.

تشير النتائج أيضًا إلى أن جميع الفرق في مهمة التعرف على مرحلة الجراحة استخدمت السياق الزمني، بينما استخدمت فريقان فقط ذلك في تقسيم حالة الأدوات، ولم يستخدمه أي فريق في تقدير نقطة المفتاح، مما يبرز تكامل النمذجة الزمنية بشكل أكثر رسوخًا في التعرف على المرحلة. بالإضافة إلى ذلك، تشير الفرق الأعلى أداءً في المهام التي تستخدم نماذج قائمة على المحولات إلى تحول نحو هذه الهياكل في تحليل الفيديو الجراحي. يكشف مقارنة بين المستشفيات عن انحياز في الأداء، حيث كانت الخوارزميات تؤدي بشكل أفضل على مقاطع الفيديو من المستشفى 1، الذي كان أكثر تمثيلًا في بيانات التدريب، مما يبرز قيدًا كبيرًا في قابلية التعميم عبر بيئات سريرية مختلفة. يشير هذا إلى أن قابلية تكيف النماذج تتأثر بشدة بتوزيع بيانات التدريب، مما يشير إلى منطقة حاسمة للبحث المستقبلي.

المناقشة

كان الهدف من تحدي PhaKIR، الذي هو جزء من مؤتمر MICCAI 2024، هو تعزيز مجال الجراحة الحد الأدنى من التدخل بمساعدة الروبوت (RAMIS) من خلال توفير مجموعة بيانات شاملة من عمليات استئصال المرارة بالمنظار وتقييم خوارزميات مختلفة. تم تنظيمه من خلال تعاون بين الجامعة التقنية في ميونيخ، والجامعة التقنية العليا في أوستباير، وAK-TORmed، وشمل التحدي ثلاث مهام رئيسية: التعرف على مرحلة الجراحة، تقسيم حالة الأدوات، وتقدير نقطة المفتاح للأداة. كان مطلوبًا من المشاركين استخدام بيانات التدريب المقدمة فقط، ولم يكن بإمكانهم نشر النتائج قبل إصدار الورقة المشتركة للتحدي. تم وضع مجموعة البيانات، التي تتكون من 13 تسلسل فيديو من ثلاثة مراكز طبية، بعناية من قبل فريق من المتخصصين الطبيين لضمان الجودة والدقة العالية.

أكد تصميم التحدي على تطوير خوارزميات قوية قادرة على التعامل مع الظروف الجراحية الواقعية، مثل الحجب وضباب الحركة. تم تشجيع المشاركين على الاستفادة من المعلومات الزمنية من تسلسلات الفيديو لتعزيز أداء نماذجهم. شملت مقاييس التقييم للمهام درجة F1 والدقة المتوازنة للتعرف على المرحلة، بالإضافة إلى معامل تشابه Dice ومتوسط الدقة العامة لتقسيم الحالة. تم تقديم نتائج التحدي في المؤتمر، مع الاعتراف بجميع الفرق المشاركة في النشر المشترك، وستكون مجموعة البيانات متاحة للجمهور بعد إصدار الورقة.

Journal: Medical Image Analysis, Volume: 109
DOI: https://doi.org/10.1016/j.media.2026.103945
PMID: https://pubmed.ncbi.nlm.nih.gov/41564633
Publication Date: 2026-01-14
Author(s): Tobias Rueckert et al.
Primary Topic: Surgical Simulation and Training

Overview

The section outlines the significance of reliable recognition and localization of surgical instruments in endoscopic video recordings, which is crucial for applications in computer-and robot-assisted minimally invasive surgery (RAMIS). The authors highlight the challenges posed by real-world conditions and propose that incorporating surgical context, such as the procedural phase, can enhance robustness and interpretability. To address these issues, they organized the Surgical Procedure Phase, Keypoint, and Instrument Recognition (PhaKIR) subchallenge at MICCAI 2024, introducing a novel multi-center dataset of thirteen full-length laparoscopic cholecystectomy videos from three medical institutions. This dataset features unified annotations for three interrelated tasks: surgical phase recognition, instrument keypoint estimation, and instrument instance segmentation, enabling a comprehensive investigation of instrument localization and procedural context.

In the conclusion, the authors emphasize the dataset’s role in benchmarking algorithm performance and evaluating various metrics. They note that while surgical phase recognition and instrument instance segmentation are well-researched, instrument keypoint estimation remains underexplored. The application of temporal modeling in phase recognition approaches underscores its importance, yet the lack of temporal context in keypoint estimation and instance segmentation reveals an area for further exploration. A significant challenge identified is the generalizability of submitted approaches, as none demonstrated robust performance across videos from different centers. The authors assert that the PhaKIR challenge and dataset contribute valuable resources to the field, aiming to inspire further research and innovation in developing reliable AI methods for surgical applications.

Introduction

The introduction of this research paper highlights advancements in computer-and robot-assisted minimally invasive surgery (RAMIS), particularly in laparoscopic procedures. It emphasizes the importance of accurately determining the current surgical phase, segmenting surgical instruments in endoscopic images, and localizing keypoints to estimate instrument poses. These capabilities enhance operating room safety through context-sensitive warnings, optimize management, and improve procedure time predictions. The paper critiques existing datasets from various Endoscopic Vision Challenge (EndoVis) sub-challenges, noting their limitations in terms of multi-center data, instrument class differentiation, and the complexity of real-world surgical environments.

To address these gaps, the authors present a new dataset comprising full endoscopic video sequences from thirteen human cholecystectomies across three medical centers. This dataset facilitates the integration of temporal information and real-world conditions into method development, encompassing three key tasks: surgical phase recognition, pixel-precise instance segmentation of surgical instruments, and localization of specific keypoints. The paper is organized according to the Biomedical Image Analysis Challenges (BIAS) guidelines, with separate chapters for each task and detailed appendices for challenge organization, data description, and participant instructions.

Methods

In this section, the authors provide a comprehensive overview of the teams that participated in the challenge and their respective methodologies across three distinct tasks: surgical phase recognition, instrument instance segmentation, and instrument keypoint estimation. A total of 14 unique teams contributed, with seven teams focusing on surgical phase recognition, nine on instrument instance segmentation, and two on instrument keypoint estimation. Notably, five teams engaged in multiple tasks, indicating a breadth of interdisciplinary expertise.

The methodologies employed by the teams are summarized in Table 2, which organizes the participants alphabetically within each task. The descriptions of their approaches are derived from the methodological reports submitted by the teams, highlighting the diverse strategies utilized to address the challenges presented in each task. This diversity underscores the complexity of the problems and the innovative solutions proposed by the participants.

Results

The results section of the research paper provides a comprehensive evaluation of three challenge tasks: surgical phase recognition, instrument instance segmentation, and instrument keypoint estimation. Each task’s performance is detailed through overall results, statistical analyses of ranking stability, and supplementary analyses, including confusion matrix visualizations for phase recognition and qualitative results for segmentation and keypoint estimation. Notably, the analysis reveals a disparity in team participation, with significantly more teams engaging in surgical phase recognition (n = 7) and instrument instance segmentation (n = 9) compared to instrument keypoint estimation (n = 2). This suggests a greater interest in the former tasks, despite the clinically relevant challenges posed by keypoint estimation, such as localizing occluded instrument parts.

The findings also indicate that all teams in the surgical phase recognition task utilized temporal context, while only two teams did so in instrument instance segmentation, and none in keypoint estimation, highlighting a more established integration of temporal modeling in phase recognition. Additionally, the top-performing teams in the tasks utilizing transformer-based models indicate a shift towards these architectures in surgical video analysis. A cross-hospital comparison reveals a performance bias, with algorithms performing best on videos from hospital 1, which was more represented in the training data, underscoring a significant limitation in generalizability across different clinical environments. This suggests that the models’ adaptability is heavily influenced by the training data’s distribution, pointing to a critical area for future research.

Discussion

The PhaKIR sub-challenge, part of the MICCAI 2024 conference, aimed to advance the field of robotic-assisted minimally invasive surgery (RAMIS) by providing a comprehensive dataset of laparoscopic cholecystectomies and benchmarking various algorithms. Organized by a collaboration between the Technical University of Munich, Ostbayerische Technische Hochschule Regensburg, and AK-TORmed, the challenge involved three main tasks: surgical phase recognition, instrument instance segmentation, and instrument keypoint estimation. Participants were required to utilize only the provided training data and could not publish results prior to the joint challenge paper’s release. The dataset, consisting of 13 video sequences from three medical centers, was meticulously annotated by a team of medical professionals to ensure high quality and accuracy.

The challenge’s design emphasized the development of robust algorithms capable of handling real-world surgical conditions, such as occlusions and motion blur. Participants were encouraged to leverage temporal information from the video sequences to enhance their models’ performance. The evaluation metrics for the tasks included the F1-score and Balanced Accuracy for phase recognition, as well as the Dice Similarity Coefficient and mean Average Precision for instance segmentation. The results of the challenge were presented at the conference, with all participating teams acknowledged in the joint publication, and the dataset will be made publicly available following the paper’s release.

شارك: