إطار المحول الفيدرالي القابل للتفسير لتصنيف اللوكيميا والتنبؤ بالمرحلة
Explainable federated transformer framework for joint leukemia classification and stage prediction

شارك:
المجلة: Scientific Reports، المجلد: 16، العدد: 1
DOI: https://doi.org/10.1038/s41598-025-34715-8
PMID: https://pubmed.ncbi.nlm.nih.gov/41486186
تاريخ النشر: 2026-01-04
المؤلف: Khadija Parwez وآخرون
الموضوع الرئيسي: التصوير الرقمي لأمراض الدم

نظرة عامة

تقدم ورقة البحث بنية متعددة الأنماط موزعة لتشخيص اللوكيميا تجمع بين الصور الدموية عالية الدقة والسرد السريري. تستخدم هذه الإطار محولات الرؤية (ViT) لتشفير الصور وClinicalBERT لتصنيف النصوص، مما يمكّن من التشخيص المشترك وتحديد مراحل اللوكيميا بطريقة لامركزية. يحدث دمج هذه الأنماط داخل طبقة دمج متعددة الأنماط، مدعومة برؤوس إخراج مزدوجة للتشخيص الثنائي وتحديد المراحل متعددة الفئات. من المهم أن تستخدم البنية بروتوكول تعلم موزع لضمان خصوصية البيانات، حيث تبقى البيانات المحلية ضمن حدود المؤسسات.

تظهر النتائج أن هذا النظام متعدد الأنماط يتفوق على النماذج التقليدية أحادية النمط والمركزية، محققًا دقة أعلى ودرجات F1 بينما يقدم تفسيرات قابلة للتفسير، محددة للمرضى من خلال طرق تعتمد على SHAP. تشير قوة البنية ضد توزيعات البيانات غير المستقلة وغير المتطابقة وقابليتها للتوسع ضمن شبكات الرعاية الصحية المحاكاة إلى إمكانياتها للتطبيق في العالم الحقيقي في علم الأورام التشخيصي. بشكل عام، تسلط هذه الدراسة الضوء على فعالية دمج تقنيات التعلم الآلي المتقدمة لتعزيز تشخيص اللوكيميا مع الحفاظ على الصلة السريرية وخصوصية البيانات.

مقدمة

تناقش مقدمة ورقة البحث التعقيدات المرتبطة بتشخيص اللوكيميا، مع التأكيد على الحاجة إلى تحليل شامل لكل من الأنماط الدموية والسجلات السريرية. تعتمد طرق التشخيص التقليدية بشكل كبير على التفسير البشري لطخات الدم وبيانات المرضى، مثل عدد كريات الدم البيضاء ونسب الخلايا النخاعية. ومع ذلك، غالبًا ما يتم تحليل هذه الأنواع من البيانات بشكل منفصل، مما يحد من القدرة على دمج المعلومات التي يمكن أن تعزز دقة التشخيص وتحديد المراحل.

تقدم التطورات الأخيرة في التعلم الآلي، وخاصة من خلال استخدام نماذج الالتفاف المعتمدة على الصور ومحولات النصوص، إمكانية أتمتة عملية التشخيص. ومع ذلك، تعمل هذه الأساليب عادةً بشكل مستقل، حيث تركز إما على المؤشرات المرئية من صور الطخات أو البيانات النصية من التقارير، دون دمج الرؤى عبر الأنماط. علاوة على ذلك، يثير الاعتماد على البيانات المركزية لتدريب هذه النماذج مخاوف بشأن خصوصية المرضى ويفشل في أخذ التباين في جمع البيانات وممارسات التشخيص عبر مؤسسات مختلفة في الاعتبار. تسلط الورقة الضوء على الحاجة الملحة لأنظمة ذكية تعطي الأولوية للخصوصية، وتلتقط الفروق السريرية بشكل فعال، وتوفر دعمًا قرارًا شفافًا في تشخيص اللوكيميا.

طرق

تحدد المنهجية المقترحة نظام تعلم موزع يهدف إلى التعلم متعدد الأنماط مع الحفاظ على الخصوصية للكشف الآلي عن اللوكيميا وتحديد مراحلها في بيئات الرعاية الصحية اللامركزية. يستخدم النظام بنية مزدوجة التدفق تعالج كل من صور الطخات الدموية والسرد السريري، مما يعزز القابلية للتفسير من خلال تفسير SHAP متعدد الأنماط. تشمل الابتكارات الرئيسية وحدة دمج متعددة الأنماط مع بوابة قابلة للتكيف لموازنة المساهمات من مشفر بصري قائم على محول الرؤية (ViT) ومشفر نصي ClinicalBERT، مما يضمن القوة ضد توزيعات البيانات غير المستقلة وغير المتطابقة. بالإضافة إلى ذلك، يتم استخدام مصنف برأسين للتنبؤ بشكل مشترك بوجود اللوكيميا والمرحلة السريرية، بما يتماشى مع الممارسات التشخيصية القياسية.

تتحقق الإعدادات التجريبية من فعالية الإطار المقترح في الكشف عن اللوكيميا وتحديد مراحلها. تم إجراء التجارب باستخدام PyTorch وHuggingFace Transformers على مجموعة حوسبة موزعة مع وحدات معالجة الرسوميات NVIDIA A100، حيث تحاكي التجارب بيئة تعلم موزعة مع 10 عملاء، يمثل كل منهم مزود رعاية صحية متميز. تم تقسيم البيانات بشكل استراتيجي لتعكس التحيزات المؤسسية، وتم تنفيذ التدريب المحلي باستخدام الانحدار العشوائي مع معلمات محددة. كانت الاتصالات بين العملاء والخادم محدودة بمعلمات النموذج للحفاظ على الخصوصية، بينما تم تطبيق تحسينات بيانات متنوعة لتعزيز تعميم النموذج. يضمن الإعداد إمكانية التكرار ويراقب تكاليف الاتصال وأوقات التزامن طوال عملية التدريب الموزع.

نتائج

تظهر نتائج الدراسة أن البنية متعددة الأنماط المقترحة تتفوق بشكل كبير على النماذج أحادية النمط والمركزية في الكشف عن اللوكيميا وتحديد مراحلها. حقق النموذج دقة تصنيف ثنائية قدرها 96.2%، متجاوزًا نموذج الصورة القائم على ViT (91.7%) ونموذج النص القائم على ClinicalBERT (89.4%). كانت المنطقة تحت منحنى خصائص التشغيل المستقبلية (AUC-ROC) للكشف عن اللوكيميا 0.983، مما يشير إلى قدرة قوية على التمييز بين الحالات اللوكيمية وغير اللوكيمية. أظهر تحليل مصفوفة الالتباس حدوثًا منخفضًا للسلبيات الكاذبة، وهو أمر حاسم في البيئات السريرية، خاصةً للكشف عن اللوكيميا.

أظهر التحليل الإضافي أن الدرجة العامة للنموذج لـ F1 لتصنيف المراحل كانت 0.911، مع حدوث أكبر أخطاء التصنيف بين المرحلة الثانية والمرحلة الثالثة، اللتين تشتركان في خصائص سريرية مشابهة. أكدت التصورات المعتمدة على SHAP أن النموذج حدد بفعالية الميزات الشكلية ذات الصلة سريريًا في الصور والمصطلحات الطبية ذات المعنى في النص، متماشية مع تعليقات الخبراء. أشار التحليل الطولي إلى أن أداء النموذج استقر حول 75 جولة اتصال، مما يشير إلى تقارب مبكر بسبب التكامل التآزري للأنماط. بالإضافة إلى ذلك، تحسنت اتساق SHAP جنبًا إلى جنب مع الأداء التنبؤي، مما يعزز قابلية تفسير النموذج ودقته. أشار دراسة الإزالة إلى أن قيم SHAP يمكن أن تعزز درجات F1 لكنها أدخلت بعض عدم الاستقرار خلال التقارب المبكر، مما يبرز قدرة النموذج على محاكاة التفكير التشخيصي الدموي مع الحفاظ على الأداء في بيئات موزعة ومقيدة بالخصوصية.

نقاش

تسلط قسم النقاش في ورقة البحث الضوء على الحاجة الملحة لأنظمة الذكاء الاصطناعي القابلة للتفسير والحساسة للخصوصية في اتخاذ القرارات الطبية، خاصة في علم الأورام. يعبر الأطباء عن تشككهم تجاه النماذج غير الشفافة التي تفتقر إلى الشفافية في تنبؤاتها، مؤكدين أن الأداء التنبؤي العالي وحده غير كافٍ للنشر في البيئات ذات المخاطر العالية. تعالج الدراسة الفجوة في الحلول الحالية من خلال اقتراح إطار تعلم موزع متعدد الأنماط لتشخيص اللوكيميا وتحديد مراحلها، والذي يدمج محولات الرؤية (ViT) لتحليل الصور وClinicalBERT لمعالجة النصوص السريرية. يضمن هذا الإطار خصوصية البيانات من خلال التدريب اللامركزي ويعزز القابلية للتفسير من خلال تقديم تفسيرات محلية تعتمد على SHAP لكل من مدخلات الصور والنصوص، مما يعزز ثقة الأطباء.

تتميز البنية المقترحة من خلال عملها تحت ظروف غير مستقلة وغير متطابقة واقعية وتقديم تصنيف برأسين لتحديد مراحل اللوكيميا، وهو أمر حاسم للتطبيق العملي في بيئات الرعاية الصحية. تسهم هذه البحث في المجال الناشئ للتعلم الموزع القابل للتفسير، بهدف سد الفجوة بين أنظمة الذكاء الاصطناعي عالية الأداء والحلول القابلة للنشر سريريًا. من خلال معالجة قضايا حوكمة البيانات، وشفافية التشخيص، والتحيز الإقليمي، يتم وضع الإطار كخطوة مهمة في دمج الأدلة متعددة الأنماط في التشخيص الطبي، مع إمكانيات تطبيق تمتد إلى مجالات أخرى مثل أمراض القلب والأشعة. بشكل عام، تمثل هذه الدراسة خطوة حاسمة نحو تطوير أنظمة الذكاء الاصطناعي التي لا تكون فعالة فحسب، بل أيضًا موثوقة ومتوافقة مع لوائح الخصوصية في البيئات السريرية.

Journal: Scientific Reports, Volume: 16, Issue: 1
DOI: https://doi.org/10.1038/s41598-025-34715-8
PMID: https://pubmed.ncbi.nlm.nih.gov/41486186
Publication Date: 2026-01-04
Author(s): Khadija Parwez et al.
Primary Topic: Digital Imaging for Blood Diseases

Overview

The research paper proposes a federated multimodal architecture for leukemia diagnosis that combines high-resolution hematological images with clinical narratives. This framework utilizes Vision Transformers (ViT) for image encoding and ClinicalBERT for text classification, enabling joint diagnosis and staging of leukemia in a decentralized manner. The integration of these modalities occurs within a cross-modal fusion layer, supported by dual output heads for binary diagnosis and multiclass staging. Importantly, the architecture employs a federated learning protocol to ensure data privacy, as local data remains within institutional boundaries.

The findings demonstrate that this multimodal system outperforms traditional unimodal and centralized models, achieving higher accuracy and F1-scores while providing interpretable, patient-specific explanations through SHAP-based methods. The architecture’s robustness against non-IID data distributions and its scalability within simulated healthcare networks suggest its potential for real-world application in diagnostic oncology. Overall, this study highlights the effectiveness of combining advanced machine learning techniques to enhance leukemia diagnosis while maintaining clinical relevance and data privacy.

Introduction

The introduction of the research paper discusses the complexities involved in diagnosing leukemia, emphasizing the need for a comprehensive analysis of both hematological patterns and clinical records. Traditional diagnostic methods rely heavily on human interpretation of blood smears and patient data, such as white blood cell counts and blast percentages. However, these data types are often analyzed in isolation, limiting the ability to synthesize information that could enhance diagnostic accuracy and staging.

Recent advancements in machine learning, particularly through the use of image-based convolutional models and text-based transformers, offer potential automation of the diagnostic process. Nonetheless, these approaches typically operate independently, focusing either on visual indicators from smear images or textual data from reports, without integrating insights across modalities. Furthermore, the reliance on centralized data for training these models raises concerns regarding patient privacy and fails to account for the variability in data collection and diagnostic practices across different institutions. The paper highlights the urgent need for intelligent systems that prioritize privacy, effectively capture clinical nuances, and provide transparent decision support in leukemia diagnosis.

Methods

The proposed methodology outlines a federated learning system aimed at privacy-preserving multimodal learning for automated leukemia detection and staging in decentralized healthcare environments. The system employs a dual-stream architecture that processes both hematological smear images and clinical narratives, enhancing interpretability through multimodal SHAP explainability. Key innovations include a cross-modal fusion module with adaptive gating to balance contributions from a Vision Transformer (ViT)-based visual encoder and a ClinicalBERT textual encoder, ensuring robustness against non-IID data distributions. Additionally, a two-head classifier is utilized to jointly predict leukemia presence and clinical stage, aligning with standard diagnostic practices.

The experimental setup validates the proposed framework’s effectiveness in detecting and staging leukemia. Conducted using PyTorch and HuggingFace Transformers on a distributed computing cluster with NVIDIA A100 GPUs, the experiments simulate a federated learning environment with 10 clients, each representing a distinct healthcare provider. Data partitioning was stratified to reflect institutional biases, and local training was executed using stochastic gradient descent with specific hyperparameters. Communication between clients and the server was limited to model parameters to maintain privacy, while various data augmentations were applied to enhance model generalization. The setup ensures reproducibility and monitors communication costs and synchronization times throughout the federated training process.

Results

The results of the study demonstrate that the proposed federated multimodal architecture significantly outperforms unimodal and centralized models in detecting and staging leukemia. The model achieved a binary classification accuracy of 96.2%, surpassing the ViT-only image model (91.7%) and ClinicalBERT-only text model (89.4%). The area under the receiver operating characteristic curve (AUC-ROC) for leukemia detection was 0.983, indicating a strong ability to differentiate between leukemic and non-leukemic cases. The confusion matrix analysis revealed a low incidence of false negatives, which is critical in clinical settings, particularly for leukemia detection.

Further analysis showed that the model’s overall F1 score for stage classification was 0.911, with the most significant classification errors occurring between Stage II and Stage III, which share similar clinical characteristics. SHAP-based visualizations confirmed that the model effectively identified clinically relevant morphological features in images and meaningful medical terms in text, aligning with expert annotations. The longitudinal analysis indicated that model performance plateaued around 75 communication rounds, suggesting early convergence due to the synergistic integration of modalities. Additionally, SHAP consistency improved alongside predictive performance, reinforcing the model’s interpretability and accuracy. An ablation study indicated that SHAP values could enhance F1 scores but introduced some instability during early convergence, highlighting the model’s capability to emulate hematological diagnostic reasoning while maintaining performance in decentralized, privacy-constrained environments.

Discussion

The discussion section of the research paper highlights the pressing need for interpretable and privacy-sensitive AI systems in medical decision-making, particularly in oncology. Clinicians express skepticism towards black-box models that lack transparency in their predictions, emphasizing that high predictive performance alone is insufficient for deployment in high-stakes environments. The study addresses the gap in existing solutions by proposing a federated multimodal learning framework for leukemia diagnosis and staging, which integrates Vision Transformers (ViT) for image analysis and ClinicalBERT for processing clinical text. This framework not only ensures data privacy through decentralized training but also enhances interpretability by providing SHAP-based local explanations for both image and text inputs, thereby fostering clinician trust.

The proposed architecture stands out by operating under realistic non-IID conditions and offering dual-head classification for leukemia staging, which is crucial for practical application in healthcare settings. The research contributes to the emerging field of explainable federated learning, aiming to bridge the gap between high-performing AI systems and clinically deployable solutions. By addressing issues of data governance, diagnostic transparency, and regional bias, the framework is positioned as a significant advancement in the integration of multimodal evidence in medical diagnostics, with potential applications extending to other areas such as cardiology and radiology. Overall, this study represents a critical step towards developing AI systems that are not only effective but also trustworthy and compliant with privacy regulations in clinical environments.

شارك: