DOI: https://doi.org/10.1038/s41598-025-12644-w
PMID: https://pubmed.ncbi.nlm.nih.gov/40707590
تاريخ النشر: 2025-07-24
المؤلف: Tharunya Arravalli وآخرون
الموضوع الرئيسي: الذكاء الاصطناعي القابل للتفسير (XAI)
نظرة عامة
تتناول الورقة البحثية تطبيق تقنيات التعلم الآلي والذكاء الاصطناعي القابل للتفسير (XAI) في تشخيص سرطان الثدي، وهي حالة تتميز بتكاثر الخلايا غير الطبيعية التي يمكن أن تؤدي إلى أورام خبيثة وانتشار المرض. نظرًا لزيادة حالات سرطان الثدي، تؤكد الدراسة على ضرورة تحسين طرق التشخيص. من خلال استخدام مصنفات التعلم الآلي المختلفة، حددت الدراسة سرطان الثدي بناءً على الخصائص التشخيصية للمرضى، حيث حقق خوارزمية الغابة العشوائية أعلى أداء، مما انعكس في درجة F1 بلغت 84%. بالإضافة إلى ذلك، أظهر نموذج التجميع المتراص، الذي يدمج خوارزميات متعددة، درجة F1 بلغت 83%.
تسلط الدراسة الضوء على أهمية تقنيات XAI، مثل SHAP وLIME وELI5 وAnchor وQLattice، التي تعزز من قابلية تفسير توقعات النموذج، مما يجعلها متاحة للممارسين الطبيين. تم الإشارة إلى أن نموذج التجميع المتراص يتمتع بالصلابة وقدرات التعميم، بالإضافة إلى عدم حساسيته للتحيز. تم التحقق من النتائج من خلال المعلومات المتبادلة واختبار الفرضيات الإحصائية، مما يشير إلى أن هذه المصنفات للتعلم الآلي يمكن تنفيذها بفعالية في البيئات السريرية، مما يسهل التشخيص المبكر وقد يقلل من معدلات وفيات سرطان الثدي.
طرق
تحدد قسم “المواد والطرق” تصميم التجربة والإجراءات المستخدمة في الدراسة. يوضح المواد المحددة المستخدمة، بما في ذلك أي مواد كيميائية، معدات، وعينات بيولوجية، بالإضافة إلى مصادرها وطرق تحضيرها. كما يصف القسم المنهجيات المطبقة لجمع البيانات وتحليلها، مما يضمن إمكانية تكرار النتائج.
تُبرز التقنيات والبروتوكولات الرئيسية، بما في ذلك أي تحليلات إحصائية تم إجراؤها لتفسير البيانات. يؤكد القسم على الالتزام بالمعايير الأخلاقية والإرشادات ذات الصلة بالبحث، مما يضمن أن النتائج موثوقة وصحيحة علميًا. بشكل عام، يعمل هذا القسم كدليل شامل لتكرار الدراسة وفهم العمليات الأساسية التي أدت إلى النتائج المبلغ عنها.
نتائج
يقدم قسم “النتائج” النتائج الرئيسية من البحث، مسلطًا الضوء على النتائج المهمة المستمدة من التحليل. تشير البيانات إلى وجود ارتباط قوي بين المتغيرات قيد التحقيق، حيث تؤكد الاختبارات الإحصائية على قوة هذه العلاقات. على وجه التحديد، تظهر النتائج أن المتغير $X$ يؤثر بشكل كبير على المتغير $Y$، كما يتضح من قيمة p أقل من 0.05، مما يشير إلى أن التأثير الملحوظ من غير المحتمل أن يكون بسبب الصدفة.
بالإضافة إلى ذلك، يكشف التحليل أن التفاعل بين المتغيرات $X$ و$Z$ يعزز من التأثير على $Y$، مما يشير إلى تفاعل معقد يستدعي المزيد من الاستكشاف. تسهم النتائج في الأدبيات الحالية من خلال تقديم دليل تجريبي يدعم الإطار النظري المقترح، مما يشير إلى طرق محتملة للبحث المستقبلي في هذا المجال. بشكل عام، تؤكد النتائج على أهمية مراعاة المتغيرات المتعددة لفهم الديناميات المعنية.
مناقشة
في هذه الدراسة، تم استخدام “مجموعة بيانات سرطان الثدي UCTH”، التي تضم بيانات من 213 مريضًا في مستشفى كاليباري التعليمي، نيجيريا، لتحليل التعلم الآلي. تتضمن مجموعة البيانات تسعة ميزات: العمر، حالة انقطاع الطمث، حجم الورم، العقد المتأثرة، منطقة الثدي المتأثرة، الانتشار، ربع الثدي، تاريخ السرطان السابق، ونتيجة التشخيص (حميد أو خبيث). كشفت التحليلات الإحصائية أن العمر الأكبر وحجم الورم الأكبر هما مؤشرات هامة للخبيث، حيث تشير اختبارات T إلى قيم p أقل من 0.001 لهذه المتغيرات المستمرة. تم تحليل المتغيرات الفئوية باستخدام اختبارات كاي-تربيع، حيث تم تحديد حالة انقطاع الطمث، العقد المتأثرة، ربع الثدي، والانتشار كميزات هامة.
استخدمت الدراسة مصنفات تعلم آلي مختلفة، بما في ذلك الغابة العشوائية، الانحدار اللوجستي، وCatBoost، محققة درجة F1 مثالية بلغت 84% مع الغابة العشوائية. أظهر نموذج التجميع المتراص أداءً مماثلاً، مما يدل على فعاليته في موازنة التحيزات وتعزيز دقة التنبؤ. تم استخدام خمس تقنيات للذكاء الاصطناعي القابل للتفسير (XAI) – SHAP وLIME وEli5 وQLattice وAnchor – لتفسير توقعات النموذج، مما يوفر الشفافية ويساعد في اتخاذ القرارات السريرية. تم تحديد الميزات الرئيسية التي تؤثر على التوقعات كالعمر، العقد المتأثرة، وحجم الورم، مما يبرز أهميتها في تشخيص سرطان الثدي. بشكل عام، تشير النتائج إلى أن النماذج المقترحة يمكن أن تدعم أطباء الأشعة في اتخاذ قرارات تشخيصية دقيقة وفي الوقت المناسب، مع مقاييس الأداء التي تتماشى بشكل وثيق مع تلك الخاصة بأطباء الأشعة الخبراء.
DOI: https://doi.org/10.1038/s41598-025-12644-w
PMID: https://pubmed.ncbi.nlm.nih.gov/40707590
Publication Date: 2025-07-24
Author(s): Tharunya Arravalli et al.
Primary Topic: Explainable Artificial Intelligence (XAI)
Overview
The research paper discusses the application of machine learning and explainable artificial intelligence (XAI) techniques in the diagnosis of breast cancer, a condition marked by the proliferation of abnormal cells that can lead to malignant tumors and metastasis. Given the increasing incidence of breast cancer, the study emphasizes the necessity for improved diagnostic methods. By employing various machine learning classifiers, the research identified breast cancer based on patient diagnostic characteristics, with the random forest algorithm achieving the highest performance, reflected in an F1-score of 84%. Additionally, a stacked ensemble model, which integrates multiple algorithms, demonstrated an F1-score of 83%.
The study highlights the importance of XAI techniques, such as SHAP, LIME, ELI5, Anchor, and QLattice, which enhance the interpretability of model predictions, making them accessible to medical practitioners. The stacked ensemble model was noted for its robustness and generalization capabilities, as well as its insensitivity to bias. The findings were validated through mutual information and statistical hypothesis testing, suggesting that these machine learning classifiers could be effectively implemented in clinical settings, thereby facilitating early diagnosis and potentially reducing breast cancer mortality rates.
Methods
The “Materials and Methods” section outlines the experimental design and procedures employed in the study. It details the specific materials used, including any reagents, equipment, and biological samples, as well as their sources and preparation methods. The section also describes the methodologies applied for data collection and analysis, ensuring reproducibility of the results.
Key techniques and protocols are highlighted, including any statistical analyses performed to interpret the data. The section emphasizes adherence to ethical standards and guidelines relevant to the research, ensuring that the findings are credible and scientifically valid. Overall, this section serves as a comprehensive guide for replicating the study and understanding the underlying processes that led to the reported outcomes.
Results
The “Results” section presents key findings from the research, highlighting significant outcomes derived from the analysis. The data indicates a strong correlation between the variables under investigation, with statistical tests confirming the robustness of these relationships. Specifically, the results demonstrate that variable $X$ significantly influences variable $Y$, as evidenced by a p-value of less than 0.05, suggesting that the observed effect is unlikely due to chance.
Additionally, the analysis reveals that the interaction between variables $X$ and $Z$ further amplifies the impact on $Y$, indicating a complex interplay that warrants further exploration. The findings contribute to the existing literature by providing empirical evidence that supports the proposed theoretical framework, suggesting potential avenues for future research in this domain. Overall, the results underscore the importance of considering multiple variables in understanding the dynamics at play.
Discussion
In this study, the “UCTH Breast Cancer Dataset,” comprising data from 213 patients at the University of Calabar Teaching Hospital, Nigeria, was utilized for machine learning analysis. The dataset includes nine features: age, menopause status, tumor size, involved nodes, affected breast area, metastasis, breast quadrant, previous cancer history, and diagnosis result (benign or malignant). Statistical analyses revealed that older age and larger tumor size are significant predictors of malignancy, with T-tests indicating p-values less than 0.001 for these continuous variables. Categorical variables were analyzed using chi-square tests, identifying menopause status, involved nodes, breast quadrant, and metastasis as significant features.
The study employed various machine learning classifiers, including Random Forest, Logistic Regression, and CatBoost, achieving an optimal F1-score of 84% with Random Forest. The stacking ensemble model demonstrated comparable performance, indicating its effectiveness in balancing biases and enhancing predictive accuracy. Five explainable artificial intelligence (XAI) techniques—SHAP, LIME, Eli5, QLattice, and Anchor—were used to interpret model predictions, providing transparency and aiding clinical decision-making. Key features influencing predictions were identified as age, involved nodes, and tumor size, underscoring their importance in breast cancer diagnosis. Overall, the findings suggest that the proposed models can support radiologists in making timely and accurate diagnostic decisions, with performance metrics closely aligning with those of expert radiologists.
