مزايا المحولات وتطبيقها في تقسيم الصور الطبية: استعراض
Advantages of transformer and its application for medical image segmentation: a survey

المجلة: BioMedical Engineering OnLine، المجلد: 23، العدد: 1
DOI: https://doi.org/10.1186/s12938-024-01212-4
PMID: https://pubmed.ncbi.nlm.nih.gov/38310297
تاريخ النشر: 2024-02-03
المؤلف: Qiumei Pu وآخرون
الموضوع الرئيسي: تشخيص COVID-19 باستخدام الذكاء الاصطناعي

نظرة عامة

تقدم هذه القسم نظرة عامة على تطبيق الشبكات العصبية المعتمدة على مشغل الالتفاف وهياكل المحولات في تقسيم الصور الطبية. يبرز نجاح الشبكات على شكل حرف U مع هياكل الترميز في هذا المجال ويناقش التكامل الأخير لتقنية المحولات، التي تم تطويرها في الأصل لمعالجة اللغة الطبيعية، في مهام تقسيم الصور الطبية. تركز المراجعة على المنشورات من عام 2018 فصاعدًا، وتحلل بشكل منهجي الدراسات التي تستخدم هياكل المحولات لهذا الغرض.

تشمل النتائج الرئيسية شرحًا مفصلًا لبناء وحدات الترميز والمحولات، وملخصًا لنماذج تقسيم الصور الطبية المعتمدة على المحولات، وقائمة بالمقاييس التقييمية المستخدمة بشكل شائع لهذه المهام. بالإضافة إلى ذلك، تصف المراجعة مجموعات بيانات تقسيم طبية متنوعة. تؤكد الخاتمة أنه بينما يمكن أن تكون نماذج المحولات النقية مفيدة، فإن فعاليتها غالبًا ما تكون محدودة بحجم العينة المتاحة في تقسيم الصور الطبية. وبالتالي، يستمر العديد من الباحثين في تطوير نماذج تجمع بين المحولات ومشغلات الالتفاف لتعزيز الأداء.

مقدمة

تسلط مقدمة هذه الورقة البحثية الضوء على الدور الحاسم لتقسيم الصور الطبية في رؤية الكمبيوتر، مع التركيز على تصنيف الصور الطبية على مستوى البكسل المستمدة من تقنيات التصوير المتقدمة مثل التصوير بالرنين المغناطيسي، والأشعة المقطعية، والألتراساوند. تعتمد طرق التقسيم التقليدية على تقنيات مثل تحديد العتبات ونمو المنطقة، والتي تتطلب استخراج ميزات يدوياً ومعرفة متخصصة. في المقابل، اكتسبت أساليب التعلم العميق، وخاصة الشبكات العصبية الالتفافية (CNNs)، شهرة بسبب قدرتها على تعلم ميزات الصورة تلقائيًا، على الرغم من أنها تتطلب موارد حسابية كبيرة ووقتًا للتدريب. لقد كان نموذج U-Net، الذي يتميز بهندسته المعمارية للترميز والترميز العكسي والاتصالات المتجاوزة، فعالًا بشكل خاص في تعزيز أداء التقسيم من خلال دمج المعلومات متعددة المقاييس.

تشير الورقة إلى تحول كبير نحو النماذج المعتمدة على المحولات في تقسيم الصور الطبية، مدفوعًا بتقديم Vision Transformer (ViT) في عام 2020. وقد أدى ذلك إلى زيادة ملحوظة في إنتاج الأبحاث، مع زيادة بنسبة 400% في المنشورات التي تركز على المحولات لمهام التقسيم. على الرغم من الاهتمام المتزايد، غالبًا ما تركز المراجعات الحالية على تطبيقات طبية محددة أو أنواع أمراض، مما يترك فجوة في التحليلات الشاملة المستندة إلى هياكل الشبكات. تهدف هذه الورقة إلى معالجة هذه الفجوة من خلال استكشاف التقدمات الأخيرة في نماذج المحولات ونماذج الترميز والترميز العكسي لتقسيم الصور الطبية، وتقديم فحص مفصل لهياكل الشبكات وتحسينها للتطبيقات العملية. ستتناول الأقسام اللاحقة هياكل النماذج، وأساليب المحولات، ومقاييس التقييم، ومجموعات البيانات، لتختتم بملخص يحدد اتجاهات البحث المستقبلية.

نقاش

ت outlines قسم النقاش في الورقة البحثية تطور وتطبيق هياكل الترميز في تقسيم الصور الطبية، مع التركيز على هندسة الترميز والترميز العكسي. تم تقديم هذه الهندسة لأول مرة بواسطة هينتون في عام 2006، وقد تم تكييفها لمهام متنوعة، بما في ذلك التقسيم، مع نماذج بارزة مثل الشبكات الالتفافية الكاملة (FCN) وU-Net التي اكتسبت شهرة. يعزز U-Net، بشكل خاص، استخراج الميزات من خلال تصميمه المتماثل والاتصالات المتجاوزة، التي تسهل استعادة المعلومات المكانية. لقد أدى تقديم نماذج المحولات، وخاصة Vision Transformer (ViT) وSwin Transformer، إلى تقدم إضافي في هذا المجال من خلال دمج آليات الانتباه التي تلتقط الاعتماديات بعيدة المدى وتحسن دقة التقسيم.

تسلط الورقة الضوء على ظهور نماذج معتمدة على المحولات مثل TransUNet، التي تدمج وحدات المحولات في إطار عمل U-Net، مما يظهر أداءً متفوقًا في مهام تقسيم الصور الطبية. تم أيضًا تطوير متغيرات مثل 3D U-Net وSwin UNETR لمعالجة تحديات محددة في التقسيم ثلاثي الأبعاد. بالإضافة إلى ذلك، يناقش القسم تحسين هياكل المحولات لتقليل التعقيد الحسابي مع الحفاظ على الأداء، كما هو الحال في نماذج مثل MedT وCTC-Net. توضح دمج المحولات في الاتصالات المتجاوزة واستكشاف الهياكل النقية للمحولات الابتكار المستمر في هذا المجال، مع تحقيق نماذج مثل Swin-Unet نتائج تنافسية مقارنة بالشبكات الالتفافية التقليدية. بشكل عام، يؤكد النقاش على التأثير التحويلي لآليات الانتباه وهياكل المحولات في تعزيز منهجيات تقسيم الصور الطبية.

Journal: BioMedical Engineering OnLine, Volume: 23, Issue: 1
DOI: https://doi.org/10.1186/s12938-024-01212-4
PMID: https://pubmed.ncbi.nlm.nih.gov/38310297
Publication Date: 2024-02-03
Author(s): Qiumei Pu et al.
Primary Topic: COVID-19 diagnosis using AI

Overview

This section provides an overview of the application of convolution operator-based neural networks and transformer architectures in medical image segmentation. It highlights the success of U-shaped networks with codec structures in this domain and discusses the recent integration of transformer technology, originally developed for natural language processing, into medical image segmentation tasks. The review focuses on publications from 2018 onwards, systematically analyzing studies that utilize transformer architectures for this purpose.

Key findings include a detailed explanation of the construction of codec and transformer modules, a summary of transformer-based medical image segmentation models, and a list of commonly used assessment metrics for these tasks. Additionally, the review describes various medical segmentation datasets. The conclusion emphasizes that while pure transformer models can be beneficial, their effectiveness is often limited by the available sample size in medical image segmentation. Consequently, many researchers continue to develop models that combine transformer and convolutional operators to enhance performance.

Introduction

The introduction of this research paper highlights the critical role of medical image segmentation in computer vision, focusing on the pixel-level classification of medical images obtained from advanced imaging modalities such as MRI, CT, and ultrasonography. Traditional segmentation methods rely on techniques like thresholding and region growth, which necessitate manual feature extraction and expert knowledge. In contrast, deep learning approaches, particularly convolutional neural networks (CNNs), have gained prominence due to their ability to automatically learn image features, although they require substantial computational resources and time for training. The U-Net model, characterized by its encoder-decoder architecture and skip connections, has been particularly effective in enhancing segmentation performance by integrating multiscale information.

The paper notes a significant shift towards transformer-based models in medical image segmentation, spurred by the introduction of the Vision Transformer (ViT) in 2020. This has led to a marked increase in research output, with a 400% rise in publications focusing on transformers for segmentation tasks. Despite the growing interest, existing reviews often concentrate on specific medical applications or disease types, leaving a gap in comprehensive analyses based on network structures. This paper aims to address this gap by exploring recent advancements in transformer and encoder-decoder models for medical image segmentation, providing a detailed examination of network structures and their optimization for practical applications. Subsequent sections will cover model structures, transformer methods, evaluation metrics, and datasets, culminating in a summary that outlines future research directions.

Discussion

The discussion section of the research paper outlines the evolution and application of codec structures in medical image segmentation, emphasizing the encoder-decoder architecture. Initially introduced by Hinton in 2006, this architecture has been adapted for various tasks, including segmentation, with notable models such as Fully Convolutional Networks (FCN) and U-Net gaining prominence. U-Net, in particular, enhances feature extraction through its symmetric design and skip connections, which facilitate the recovery of spatial information. The introduction of transformer models, particularly the Vision Transformer (ViT) and Swin Transformer, has further advanced the field by incorporating attention mechanisms that capture long-distance dependencies and improve segmentation accuracy.

The paper highlights the emergence of transformer-based models like TransUNet, which integrates transformer modules into the U-Net framework, demonstrating superior performance in medical image segmentation tasks. Variants such as 3D U-Net and Swin UNETR have also been developed to address specific challenges in 3D segmentation. Additionally, the section discusses the optimization of transformer architectures to reduce computational complexity while maintaining performance, as seen in models like MedT and CTC-Net. The integration of transformers into skip connections and the exploration of pure transformer structures further illustrate the ongoing innovation in this domain, with models like Swin-Unet achieving competitive results against traditional convolutional networks. Overall, the discussion underscores the transformative impact of attention mechanisms and transformer architectures in enhancing medical image segmentation methodologies.