DOI: https://doi.org/10.1038/s41597-025-04562-3
PMID: https://pubmed.ncbi.nlm.nih.gov/39934113
تاريخ النشر: 2025-02-11
المؤلف: Noelia Vállez وآخرون
الموضوع الرئيسي: الذكاء الاصطناعي في اكتشاف السرطان
نظرة عامة
تتكون مجموعة البيانات من 38 مسح بالموجات فوق الصوتية للثدي، بإجمالي 683 صورة، تم الحصول عليها باستخدام نظام الموجات فوق الصوتية Siemens ACUSON S2000TM بين عامي 2022 و2023. تم تصميمها خصيصًا لتقسيم آفات الثدي، بهدف تحديد المنطقة وشكل الآفات مع تصنيفها على أنها حميدة أو خبيثة. تم تصنيف الصور إلى ثلاث مجموعات: 419 حالة طبيعية، 174 حالة حميدة، و90 حالة خبيثة.
تُقدم بيانات الحقيقة الأرضية في شكل أقنعة تقسيم RGB، حيث يمثل الأسود نسيج الثدي الطبيعي، ويشير الأخضر إلى الآفات الحميدة، ويدل الأحمر على الآفات الخبيثة. تعتبر مجموعة البيانات هذه أداة مهمة للباحثين الذين يقومون بتطوير وتقييم نماذج التعلم الآلي التي تهدف إلى التمييز بين الأورام الحميدة والخبيثة في صور الموجات فوق الصوتية الحقيقية للثدي. تضمن التعليقات التوضيحية للتقسيم، التي أنشأها أطباء الأشعة الخبراء، تدريب النموذج وتقييمه بدقة، مما يعزز من أهميتها في مجالات رؤية الكمبيوتر والصحة العامة.
مقدمة
تسلط المقدمة الضوء على المشهد الحالي لأنظمة الكشف المدعومة بالكمبيوتر (CAD) في فحص سرطان الثدي، مشيرة إلى فجوة كبيرة في تطبيق مثل هذه الأنظمة على صور الموجات فوق الصوتية للثدي (BUS) مقارنة بالأشعة السينية. بينما تستخدم أنظمة CAD للأشعة السينية تقنيات متنوعة، بما في ذلك التعلم الآلي التقليدي وطرق التعلم العميق، لا يزال تطوير أنظمة CAD لـ BUS محدودًا، مع كون نظام S-Detect™ من سامسونج استثناءً ملحوظًا. يعزز هذا النظام دقة الكشف عن سرطان الثدي من خلال تصنيف الآفات وفقًا لـ BI-RADS® ATLAS.
تتمثل إحدى التحديات الرئيسية في تطوير أنظمة CAD لـ BUS في ندرة مجموعات البيانات الكبيرة والمُعَلَّمة بشكل جيد اللازمة لتدريب نماذج التعلم الآلي القوية. تهدف مجموعة بيانات BUS-UCLM إلى معالجة هذه المشكلة من خلال توفير مجموعة شاملة من صور الموجات فوق الصوتية للثدي مع تعليقات توضيحية مفصلة من أطباء الأشعة الخبراء. تهدف هذه المجموعة إلى تسهيل البحث في تقسيم الآفات، والتصنيف، والكشف، مما يحسن في النهاية من تشخيص سرطان الثدي ويقلل من معدلات الوفيات من خلال تعزيز القدرات التشخيصية.
الطرق
توضح قسم “الطرق” في ورقة البحث التصميم التجريبي والتقنيات التحليلية المستخدمة للتحقيق في سؤال البحث. استخدمت الدراسة نهجًا كميًا، يتضمن تحليلات إحصائية لتقييم البيانات المجمعة من تجارب مختلفة. تضمنت المنهجيات المحددة تجارب مختبرية محكومة، حيث تم التلاعب بالمتغيرات بشكل منهجي لملاحظة تأثيراتها على النتائج المعنية.
شملت جمع البيانات استخدام أدوات موحدة لضمان الموثوقية والصلاحية، مع التركيز على تقليل التحيزات. تم إجراء التحليل باستخدام برامج إحصائية مناسبة، وتطبيق تقنيات مثل تحليل الانحدار وANOVA لتفسير النتائج. يبرز القسم أهمية إمكانية التكرار والشفافية في الطرق المستخدمة، موضحًا حجم العينة ومعايير الاختيار لدعم قوة النتائج. بشكل عام، أسست الإطار المنهجي قاعدة صلبة للتحليل اللاحق والاستنتاجات المستخلصة في الدراسة.
المناقشة
توضح قسم المناقشة في ورقة البحث المنهجية والاعتبارات الأخلاقية المتعلقة بالحصول على صور الموجات فوق الصوتية، وإخفاء الهوية، والتعليق عليها التي تم جمعها من عام 2022 إلى 2023 في مستشفى مدينة ريال الجامعي العام. تتضمن مجموعة البيانات، التي تشمل صورًا بدقة بكسل وأبعاد 768 × 1024 بكسل، بيانات مستمدة من دراسات سريرية حقيقية دون معايير اختيار محددة مسبقًا، مما يضمن تمثيلًا واقعيًا للسيناريوهات السريرية. تم الحصول على الموافقة الأخلاقية، وتم تأمين الموافقة المستنيرة من المشاركين، مع اتخاذ تدابير صارمة لإخفاء بيانات المرضى بما يتماشى مع لوائح GDPR وHIPAA. تم تحويل الصور من تنسيق DICOM إلى PNG، وتم إخفاء المعلومات الحساسة باستخدام شبكة عصبية تلافيفية (CNN) مدربة على YOLOv8.
تتكون مجموعة البيانات من 683 صورة من 38 مريضًا، مصنفة إلى آفات طبيعية، حميدة، وخبيثة، مع تعليقات توضيحية تم إنشاؤها بواسطة أطباء الأشعة الخبراء. تعزز هذه العملية الدقيقة للتعليق، التي شملت وضع علامات على جميع الآفات في الصور متعددة الآفات، من فائدة مجموعة البيانات لمهام التقسيم المختلفة. تم التحقق من جودة مجموعة البيانات من خلال تدريب نموذج UNet، محققًا درجة Dice تبلغ 0.68، مما يشير إلى إمكانيته في تطوير نماذج تعلم آلي موثوقة. ومع ذلك، يعترف المؤلفون بالتحيزات المحتملة، بما في ذلك التحيزات الديموغرافية والسريرية بسبب جمع البيانات المحلي والتعليقات التوضيحية على مستوى الخبراء. للتخفيف من هذه التحيزات، تم توفير بيانات وصفية مفصلة، مما يسمح للمستخدمين بتصفية الصور بناءً على خصائص معينة. يُوصى بدمج مجموعة بيانات BUS-UCLM مع مجموعات بيانات أخرى متاحة للجمهور لتعزيز عمومية النموذج وقوته، كما يتضح من الأداء المتفوق لنموذج Mask R-CNN في مهام التقسيم.
DOI: https://doi.org/10.1038/s41597-025-04562-3
PMID: https://pubmed.ncbi.nlm.nih.gov/39934113
Publication Date: 2025-02-11
Author(s): Noelia Vállez et al.
Primary Topic: AI in cancer detection
Overview
The dataset consists of 38 breast ultrasound scans, totaling 683 images, acquired using a Siemens ACUSON S2000TM Ultrasound System between 2022 and 2023. It is specifically designed for the segmentation of breast lesions, aiming to delineate the area and contour of lesions while classifying them as benign or malignant. The images are categorized into three groups: 419 normal, 174 benign, and 90 malignant cases.
Ground truth data is provided in the form of RGB segmentation masks, where black represents normal breast tissue, green indicates benign lesions, and red denotes malignant lesions. This dataset is instrumental for researchers developing and assessing machine learning models aimed at differentiating between benign and malignant tumors in real breast ultrasound images. The segmentation annotations, created by expert radiologists, ensure precise model training and evaluation, thereby enhancing its significance in the domains of computer vision and public health.
Introduction
The introduction highlights the current landscape of computer-aided detection (CAD) systems in breast cancer screening, noting a significant gap in the application of such systems to breast ultrasound images (BUS) compared to mammography. While mammography CAD systems utilize various techniques, including conventional machine learning and deep learning methods, the development of BUS CAD systems remains limited, with the S-Detect™ system from Samsung being a notable exception. This system enhances the accuracy of breast cancer detection by classifying lesions according to the BI-RADS® ATLAS.
A critical challenge in advancing BUS CAD systems is the scarcity of large, well-annotated datasets necessary for training robust machine learning models. The introduction of the BUS-UCLM dataset aims to address this issue by providing a comprehensive collection of breast ultrasound images with detailed annotations from expert radiologists. This dataset is intended to facilitate research in lesion segmentation, classification, and detection, ultimately improving breast cancer prognosis and reducing mortality rates through enhanced diagnostic capabilities.
Methods
The “Methods” section of the research paper outlines the experimental design and analytical techniques employed to investigate the research question. The study utilized a quantitative approach, incorporating statistical analyses to evaluate the data collected from various experiments. Specific methodologies included controlled laboratory experiments, where variables were systematically manipulated to observe their effects on the outcomes of interest.
Data collection involved the use of standardized instruments to ensure reliability and validity, with a focus on minimizing biases. The analysis was conducted using appropriate statistical software, applying techniques such as regression analysis and ANOVA to interpret the results. The section emphasizes the importance of replicability and transparency in the methods used, detailing the sample size and selection criteria to support the robustness of the findings. Overall, the methodological framework established a solid foundation for the subsequent analysis and conclusions drawn in the study.
Discussion
The discussion section of the research paper outlines the methodology and ethical considerations involved in the acquisition, anonymization, and annotation of ultrasound images collected from 2022 to 2023 at Ciudad Real General University Hospital. The dataset, which includes images with varying pixel resolutions and dimensions of 768 × 1024 pixels, was derived from authentic clinical studies without predefined selection criteria, ensuring a realistic representation of clinical scenarios. Ethical approval was obtained, and informed consent was secured from participants, with stringent measures taken to anonymize patient data in compliance with GDPR and HIPAA regulations. The images were converted from DICOM to PNG format, and sensitive information was masked using a YOLOv8-trained convolutional neural network (CNN).
The dataset comprises 683 images from 38 patients, categorized into normal, benign, and malignant lesions, with annotations generated by expert radiologists. This meticulous annotation process, which included labeling all lesions in multi-lesion images, enhances the dataset’s utility for various segmentation tasks. The dataset’s quality was validated through training a UNet model, achieving a Dice score of 0.68, indicating its potential for developing reliable machine learning models. However, the authors acknowledge potential biases, including demographic and clinical biases due to the localized collection and expert-level annotations. To mitigate these biases, detailed metadata is provided, allowing users to filter images based on specific characteristics. The integration of the BUS-UCLM dataset with other publicly available datasets is recommended to enhance model generalizability and robustness, as demonstrated by the superior performance of the Mask R-CNN model in segmentation tasks.
