الدقة والكفاءة في تقسيم سرطان الجلد من خلال نموذج تعلم عميق مزدوج الترميز
Precision and efficiency in skin cancer segmentation through a dual encoder deep learning model

شارك:
المجلة: Scientific Reports، المجلد: 15، العدد: 1
DOI: https://doi.org/10.1038/s41598-025-88753-3
PMID: https://pubmed.ncbi.nlm.nih.gov/39924555
تاريخ النشر: 2025-02-09
المؤلف: Zhenyun Du وآخرون
الموضوع الرئيسي: كشف وإدارة الميلانوما الجلدية

نظرة عامة

تقدم ورقة البحث DuaSkinSeg، نموذج تعلم عميق جديد مصمم لتجزئة آفات الجلد، مع معالجة القيود التي تعاني منها الطرق الحالية التي غالبًا ما تتنازل عن الكفاءة واستخراج الميزات. يستخدم DuaSkinSeg بنية مزدوجة التشفير، حيث يدمج MobileNetV2 المدرب مسبقًا لاستخراج الميزات المحلية مع شبكة تحويل الرؤية-شبكة عصبية تلافيفية (ViT-CNN) لالتقاط الاعتماديات بعيدة المدى. يسمح هذا الجمع المبتكر للنموذج بالاحتفاظ بالسياق العالمي بفعالية مع الحفاظ على دقة تجزئة عالية.

تظهر التقييمات الشاملة على مجموعات البيانات المرجعية (ISIC 2016، ISIC 2017، وISIC 2018) أن DuaSkinSeg يتفوق على الطرق الحديثة، مثل U-Net وFAT-Net، خصوصًا في مقاييس مثل الدقة، والتميز، والاسترجاع، ودرجة F1، ومؤشر جاكارد. يركز تصميم النموذج على الكفاءة الحسابية وقابلية التكيف السريرية، مما يجعله مناسبًا للتطبيقات الجلدية في العالم الحقيقي. تهدف الأبحاث المستقبلية إلى تحسين DuaSkinSeg بشكل أكبر وتقييم أدائه عبر مجموعة أوسع من مجموعات البيانات السريرية، بهدف تعزيز الكشف المبكر عن الميلانوما وتحسين نتائج المرضى في dermatology.

مقدمة

في هذا القسم، يقدم المؤلفون DuaSkinSeg، بنية تجزئة جديدة مصممة لتعزيز تجزئة آفات الجلد من خلال دمج الشبكات العصبية التلافيفية (CNNs) وشبكات التحويل البصرية (ViTs) ضمن إطار مزدوج التشفير. يستخدم هذا النهج المبتكر اتصالات تخطي لربط التشفير المزدوج بمفكك تشفير CNN، مما يوازن بفعالية بين دقة التجزئة والكفاءة الحسابية. تتميز البنية بقدرتها على استخراج ميزات محلية مفصلة مع التقاط معلومات السياق العالمي الأساسية في الوقت نفسه، وهو أمر حاسم لدقة التجزئة.

يسلط المؤلفون الضوء على ثلاثة مساهمات رئيسية لعملهم: أولاً، تقديم إطار التشفير المزدوج الذي يجمع بين مشفرات CNN وViT، مما يمكّن من استخراج ميزات شاملة؛ ثانيًا، الالتقاط الفعال للاعتماديات بعيدة المدى من خلال بنية مشفر-مفكك تشفير 2D ViT-CNN، المستوحاة من UNETR 3D، مما يعزز قدرة النموذج على معالجة بيانات آفات الجلد؛ وثالثًا، تقييم DuaSkinSeg على ثلاث مجموعات بيانات مرجعية (ISIC 2016، ISIC 2017، ISIC 2018)، حيث أظهر أداءً متفوقًا مقارنة بالطرق الحالية، مما يثبت فعالية بنية التشفير المزدوج المقترحة لمهام تجزئة آفات الجلد.

طرق

يقدم نموذج DuaSkinSeg لتجزئة آفات الجلد منهجية جديدة تعزز استخراج الميزات من خلال بنية مزدوجة التشفير، كما هو موضح في الشكل 1. يدمج هذا النموذج فرعين: شبكة عصبية تلافيفية (CNN) وشبكة تحويل بصرية (ViT)، تعمل جنبًا إلى جنب لاستخراج كل من الميزات المحلية المفصلة ومعلومات السياق العالمي الأساسية. هذه الطريقة المزدوجة حاسمة لتحقيق دقة عالية في تجزئة آفات الجلد.

تم تصميم البنية بشكل متماثل، حيث تحتوي على اتصالات تخطي تسهل التواصل الفعال بين مكونات التشفير والمفكك. يقوم مفكك تشفير CNN بعد ذلك بمعالجة الميزات الغنية المستمدة من التشفير المزدوج، مما يؤدي إلى إنشاء خريطة تجزئة شاملة ومفصلة. يبرز هذا التصميم المبتكر إمكانيات دمج هياكل الشبكات العصبية المختلفة لتحسين الأداء في مهام تحليل الصور الطبية.

نتائج

تم تقييم نموذج DuaSkinSeg بدقة مقابل أحد عشر نموذجًا متقدمًا (SOTA) لتجزئة البيانات على مجموعة بيانات ISIC 2018، مما أظهر أداءً متفوقًا عبر مقاييس متعددة. حقق DuaSkinSeg دقة (ACC) تبلغ 97.08%، ومؤشر جاكارد (JI) يبلغ 84.63%، واسترجاع (RE) يبلغ 91.12%، وتميز (PR) يبلغ 92.44%. توضح المقارنات المرئية في الشكل 6 أن DuaSkinSeg يتفوق في السيناريوهات المعقدة، خصوصًا حيث تكافح النماذج الأخرى مثل U-Net وResUNet++ بسبب التباين المنخفض بين الآفات والجلد المحيط. يُعزى هذا التفوق إلى دمج DuaSkinSeg الفعال لمعلومات السياق المحلي والعالمي.

أكدت التقييمات الإضافية على مجموعات بيانات ISIC 2017 وISIC 2016 هيمنة DuaSkinSeg، مع درجات تبلغ 89.74% RE، 96.51% ACC، 80.68% JI، و90.10% PR على الأولى، و96.40% ACC، 86.43% JI، 93.41% RE، و92.68% PR على الثانية. تبرز هذه النتائج، المدعومة بنتائج تجزئة مقارنة، قوة النموذج في التعامل مع الحالات الصعبة ذات التباينات الغامضة والحدود المعقدة. بالإضافة إلى ذلك، تشير مقاييس التدريب والتحقق إلى تعلم فعال وتوافق، مع زيادات مستمرة في الدقة وانخفاضات في الخسارة عبر العصور، مما يؤكد قدرة النموذج على التعميم دون الإفراط في التكيف.

نقاش

في قسم النقاش من ورقة البحث، يسلط المؤلفون الضوء على تطور تقنيات تجزئة آفات الجلد، مقارنين بين الطرق التقليدية المعتمدة على الميزات وطرق الشبكات العصبية التلافيفية الحديثة (CNN)، خصوصًا بنية U-Net. بينما تعتمد الطرق التقليدية على ميزات الصورة المحددة مسبقًا والعتبات، أظهرت الشبكات العصبية التلافيفية، وخاصة U-Net ونسخها، أداءً متفوقًا من خلال تعلم الميزات تلقائيًا من الصور. كما يشير المؤلفون إلى ظهور شبكات التحويل البصرية (ViTs) في رؤية الكمبيوتر، التي تتفوق في التقاط السياق العالمي من خلال آليات الانتباه الذاتي. يجمع نموذجهم المقترح، DuaSkinSeg، بشكل مبتكر بين مزايا MobileNetV2 لاستخراج الميزات المحلية وViT للسياق العالمي، معالجًا القيود التي تعاني منها النماذج الحالية التي تكافح لدمج هذه الجوانب بفعالية.

أجرى المؤلفون تقييمات شاملة باستخدام ثلاث مجموعات بيانات متاحة للجمهور (ISIC 2016، ISIC 2017، وISIC 2018) واستخدموا مقاييس متنوعة، بما في ذلك الدقة، ومؤشر جاكارد، والتميز، والاسترجاع، ودرجة F1، لتقييم الأداء. تشير النتائج إلى أن DuaSkinSeg يتفوق بشكل كبير على النماذج المعروفة مثل U-Net وResUNet++ وTransUNet، مع تأكيد التحليلات الإحصائية على قوة هذه التحسينات. تتيح بنية التشفير المزدوج لـ DuaSkinSeg، التي تستخدم اتصالات تخطي لدمج الميزات بفعالية، تحسين دقة التجزئة والكفاءة، مما يجعلها أداة واعدة للتطبيقات السريرية في dermatology. تهدف الأعمال المستقبلية إلى تحسين النموذج بشكل أكبر وتقييم أدائه عبر مجموعة أوسع من مجموعات البيانات السريرية، مع التركيز على تأثيره المحتمل على الكشف المبكر عن الميلانوما ونتائج المرضى.

Journal: Scientific Reports, Volume: 15, Issue: 1
DOI: https://doi.org/10.1038/s41598-025-88753-3
PMID: https://pubmed.ncbi.nlm.nih.gov/39924555
Publication Date: 2025-02-09
Author(s): Zhenyun Du et al.
Primary Topic: Cutaneous Melanoma Detection and Management

Overview

The research paper presents DuaSkinSeg, a novel deep-learning model designed for skin lesion segmentation, addressing the limitations of existing methods that often compromise between efficiency and feature extraction. DuaSkinSeg employs a dual-encoder architecture, integrating a pre-trained MobileNetV2 for local feature extraction with a Vision Transformer-Convolutional Neural Network (ViT-CNN) for capturing long-range dependencies. This innovative combination allows the model to effectively retain global context while maintaining high segmentation accuracy.

Extensive evaluations on benchmark datasets (ISIC 2016, ISIC 2017, and ISIC 2018) reveal that DuaSkinSeg outperforms state-of-the-art methods, such as U-Net and FAT-Net, particularly in metrics like accuracy, precision, recall, F1-score, and Jaccard Index. The model’s design emphasizes computational efficiency and clinical adaptability, making it suitable for real-world dermatological applications. Future research will aim to refine DuaSkinSeg further and assess its performance across a wider array of clinical datasets, with the goal of enhancing early melanoma detection and improving patient outcomes in dermatology.

Introduction

In this section, the authors introduce DuaSkinSeg, a novel segmentation architecture designed to enhance skin lesion segmentation by integrating Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs) within a dual-encoder framework. This innovative approach employs skip connections to link the dual encoders to a CNN decoder, effectively balancing segmentation accuracy with computational efficiency. The architecture is characterized by its ability to extract detailed local features while simultaneously capturing essential global context information, which is critical for accurate segmentation.

The authors highlight three main contributions of their work: first, the introduction of the dual encoder framework that combines CNN and ViT encoders, enabling comprehensive feature extraction; second, the efficient capture of long-range dependencies through a 2D ViT-CNN encoder-decoder architecture, inspired by UNETR 3D, which enhances the model’s ability to process skin lesion data; and third, the evaluation of DuaSkinSeg on three benchmark datasets (ISIC 2016, ISIC 2017, ISIC 2018), where it demonstrated superior performance compared to existing methods, validating the effectiveness of the proposed dual-encoder architecture for skin lesion segmentation tasks.

Methods

The DuaSkinSeg model for skin lesion segmentation introduces a novel methodology that enhances feature extraction through a dual encoder architecture, as depicted in Figure 1. This model integrates two branches: a Convolutional Neural Network (CNN) and a Vision Transformer (ViT), which work in tandem to extract both detailed local features and essential global context information. This dual approach is critical for achieving high accuracy in skin lesion segmentation.

The architecture is designed symmetrically, featuring skip connections that facilitate effective communication between the encoder and decoder components. The CNN decoder subsequently processes the enriched features obtained from the dual encoder, resulting in the generation of a comprehensive and detailed segmentation map. This innovative design underscores the potential of combining different neural network architectures to improve performance in medical image analysis tasks.

Results

The DuaSkinSeg model was rigorously evaluated against eleven state-of-the-art (SOTA) segmentation models on the ISIC 2018 dataset, demonstrating superior performance across multiple metrics. Specifically, DuaSkinSeg achieved an accuracy (ACC) of 97.08%, a Jaccard Index (JI) of 84.63%, a Recall (RE) of 91.12%, and a Precision (PR) of 92.44%. Visual comparisons in Figure 6 illustrate that DuaSkinSeg excels in complex scenarios, particularly where other models like U-Net and ResUNet++ struggle due to low contrast between lesions and surrounding skin. This advantage is attributed to DuaSkinSeg’s effective integration of local and global contextual information.

Further evaluations on the ISIC 2017 and ISIC 2016 datasets reaffirmed DuaSkinSeg’s dominance, with scores of 89.74% RE, 96.51% ACC, 80.68% JI, and 90.10% PR on the former, and 96.40% ACC, 86.43% JI, 93.41% RE, and 92.68% PR on the latter. These results, supported by comparative segmentation outcomes, highlight the model’s robustness in handling challenging cases with ambiguous contrasts and complex boundaries. Additionally, training and validation metrics indicate effective learning and convergence, with consistent increases in accuracy and decreases in loss over epochs, confirming the model’s capability to generalize without overfitting.

Discussion

In the discussion section of the research paper, the authors highlight the evolution of skin lesion segmentation techniques, contrasting traditional feature-based methods with modern convolutional neural network (CNN) approaches, particularly the U-Net architecture. While traditional methods rely on predefined image features and thresholding, CNNs, especially U-Net and its variants, have shown superior performance by automatically learning features from images. The authors also note the emergence of Vision Transformers (ViTs) in computer vision, which excel at capturing global context through self-attention mechanisms. Their proposed model, DuaSkinSeg, innovatively combines the strengths of MobileNetV2 for local feature extraction and ViT for global context, addressing the limitations of existing models that struggle to integrate these aspects effectively.

The authors conducted extensive evaluations using three publicly available datasets (ISIC 2016, ISIC 2017, and ISIC 2018) and employed various metrics, including accuracy, Jaccard index, precision, recall, and F1-score, to assess performance. The results indicate that DuaSkinSeg significantly outperforms established models such as U-Net, ResUNet++, and TransUNet, with statistical analyses confirming the robustness of these improvements. The dual-encoder architecture of DuaSkinSeg, which utilizes skip connections for effective feature fusion, allows for enhanced segmentation accuracy and efficiency, making it a promising tool for clinical applications in dermatology. Future work aims to refine the model further and evaluate its performance across a broader range of clinical datasets, emphasizing its potential impact on early melanoma detection and patient outcomes.

شارك: