إطار عمل قابل للتفسير للبلوكشين الفيدرالي مع تحسين الذكاء الاصطناعي الذي يحافظ على الخصوصية لتأمين بيانات الرعاية الصحية
An explainable federated blockchain framework with privacy-preserving AI optimization for securing healthcare data

شارك:
المجلة: Scientific Reports، المجلد: 15، العدد: 1
DOI: https://doi.org/10.1038/s41598-025-04083-4
PMID: https://pubmed.ncbi.nlm.nih.gov/40595873
تاريخ النشر: 2025-07-01
المؤلف: Tanisha Bhardwaj وآخرون
الموضوع الرئيسي: التقنيات التي تحافظ على الخصوصية في البيانات

نظرة عامة

تقدم الدراسة إطار PPFBXAIO، الذي يدمج التعلم الفيدرالي، وتكنولوجيا البلوكشين، والذكاء الاصطناعي القابل للتفسير لمعالجة التحديات الكبيرة في أنظمة التعلم الآلي اللامركزية، لا سيما في قطاع الرعاية الصحية. باستخدام تقنيات مثل LGOA لاختيار الميزات، وتطبيع Min-Max، وتشفير SHA-256، والعقود الذكية لتجميع النماذج، يضمن الإطار خصوصية البيانات، وشفافية النموذج، والقدرة على مقاومة هجمات التسميم. أظهرت التقييمات على مجموعة بيانات أمراض القلب ومجموعة بيانات سرطان الثدي في ويسكونسن أداءً مثيرًا للإعجاب، حيث حقق PPFBXAIO دقة بنسبة 95.07% ودرجة F1 بنسبة 95.98% لسرطان الثدي، و93.07% دقة و93.24% درجة F1 لأمراض القلب، متجاوزًا نماذج التعلم الفيدرالي الحالية مثل FedAvg وFL-MPC.

على الرغم من نتائجه الواعدة، يواجه الإطار تحديات تتعلق بزيادة الحمل الحسابي والتواصل بسبب دمج البلوكشين، مما قد يعيق قابلية التوسع والنشر في الوقت الحقيقي، خاصة مع زيادة عدد العقد. ستركز الأبحاث المستقبلية على تنفيذ آليات توافق خفيفة الوزن، واستكشاف حلول التخزين خارج السلسلة، ودمج بيانات طبية متعددة الأنماط لتحسين التشخيصات، ودمج تقنيات الذكاء الاصطناعي القابل للتفسير المتقدمة، وإجراء دراسات جدوى سريرية للتحقق من الإطار في الإعدادات العملية. تهدف هذه الجهود إلى تحسين PPFBXAIO ليصبح حلاً قابلاً للتوسع، وفعالاً، ويحافظ على الخصوصية لأنظمة الرعاية الصحية اللامركزية.

الطرق

تقدم المنهجية المقترحة في هذا البحث إطار تحسين الذكاء الاصطناعي القابل للتفسير المحفوظ الخصوصية (PPFBXAIO)، الذي يدمج تكنولوجيا البلوكشين مع التعلم الفيدرالي (FL) والذكاء الاصطناعي القابل للتفسير (XAI) لتعزيز أمان إدارة البيانات، والشفافية، والخصوصية. يستخدم الإطار خوارزمية تجميع لتصفية المعلمات النموذجية الشاذة، مما يقلل من هجمات التسميم مع الحفاظ على خصوصية البيانات. يستفيد من طرق التفسير المعتمدة على SHAP المحلية للتحقق من توقعات النموذج دون الوصول إلى البيانات المركزية، مما يسهل مشاركة المستخدمين ويحسن من قابلية تفسير قرارات نموذج الذكاء الاصطناعي. يساهم دمج خوارزمية تحسين الجينات المحلية (LGOA) في تحسين المعلمات الفائقة واختيار السمات المهمة من مجموعة البيانات، مما يعزز أداء النموذج في توزيعات البيانات غير IID.

تم تقييم فعالية إطار PPFBXAIO باستخدام مجموعات بيانات متاحة للجمهور لأمراض القلب وسرطان الثدي، حيث حقق معدلات دقة تبلغ 91.19% و95.44% على التوالي، متفوقًا على طرق أخرى مثل FedHFP+RNN وFedAvgBC+TabNet. تم اختبار الإطار أيضًا على مجموعات بيانات أكبر من العالم الحقيقي، بما في ذلك MIMIC-III EHR وNIH Chest X-ray14، مما يظهر قابلية التوسع والعمومية مع الحفاظ على الخصوصية. حقق المصنف المقترح أعلى استدعاء (95.39% لأمراض القلب) وأعلى مقياس F (93.24% لأمراض القلب) بين الطرق المقارنة. بالإضافة إلى ذلك، تم تقديم آلية توافق جديدة، إثبات الجودة (PoQ)، لتعزيز تجميع النموذج بناءً على اتساق التفسير والثقة التاريخية، مما يحسن الأداء والشفافية في بيئات التعلم الفيدرالي.

النتائج

تنفذ الدراسة طرق التعلم الفيدرالي (FL) باستخدام تكنولوجيا البلوكشين وPython 3.10.11 لتدريب النموذج وتوليد الضوضاء، باستخدام PyTorch 2.0.1. يتم تحقيق الاتصال بين العقد من خلال عميل HTTP الأصلي لـ Go الذي يتفاعل مع خادم HTTP Tornado المستند إلى Python. تُجرى التجارب على جهاز كمبيوتر مزود بمعالج Intel i5-13400، و32 جيجابايت من ذاكرة الوصول العشوائي، وNVIDIA RTX 4070 Ti GPU، مع التركيز على مجموعات البيانات المتعلقة بأمراض القلب وسرطان الثدي.

تقارن الدراسة النموذج المقترح مع مخططات FL المعروفة، بما في ذلك FedAvg، والتعلم الفيدرالي مع الحساب متعدد الأطراف (FL-MPC)، وFL-RAEC، وإطار FL المحفوظ الخصوصية والفعال مع البلوكشين (PEFL)، والتعلم الفيدرالي المدعوم بالبلوكشين المحفوظ الخصوصية (PPBEFL). تشير النتائج إلى أن النموذج المقترح يظهر مرونة متفوقة ضد نماذج الهجوم المختلفة. تعتبر مجموعات البيانات المستخدمة مناسبة للتعلم الفيدرالي بسبب تعقيدها، وميزاتها المتنوعة، وأهميتها في التنبؤات الحرجة في الرعاية الصحية، كما هو ملخص في الجدول 1، الذي يوضح إعداد التجربة لنموذج المحاكاة.

المناقشة

تستعرض قسم المناقشة في ورقة البحث مختلف التقدمات في التعلم الفيدرالي (FL) ودمجه مع تكنولوجيا البلوكشين، مع التركيز على تطبيقاتها في الرعاية الصحية وخصوصية البيانات. قدم Wang وآخرون إثبات التعلم الفيدرالي بدون منصة (PF-PoFL)، الذي يستخدم آليات توافق موفرة للطاقة لتعزيز مهام FL مع معالجة الثغرات الأمنية مثل هجمات Sybil من خلال أنواع معاملات مبتكرة وحوافز قائمة على الائتمان. يساهم دمج الخصوصية التفاضلية على مستوى المستخدم في حماية إضافية ضد تسريبات الخصوصية الضمنية أثناء تدريب النموذج. تبرز دراسات أخرى، مثل تلك التي أجراها Singh وآخرون وGupta وآخرون، إمكانيات FL في أنظمة الرعاية الصحية الذكية، مما يمكّن الوصول الآمن إلى البيانات وتدريب النموذج دون المساس بخصوصية المرضى.

بالإضافة إلى ذلك، يتم مناقشة منهجيات مختلفة لتجميع النماذج واختيار الميزات، بما في ذلك استخدام طرق التجميع وخوارزمية تحسين الجراد ليفي (LGOA) لاختيار الميزات بشكل فعال في التنبؤ بالأمراض. تؤكد الورقة على أهمية الذكاء الاصطناعي القابل للتفسير (XAI) في تعزيز شفافية النموذج والمساءلة، لا سيما في المجالات المنظمة مثل الرعاية الصحية. يُقترح دمج XAI مع التعلم الفيدرالي والبلوكشين كوسيلة لإنشاء إطار عمل للذكاء الاصطناعي آمن، وقابل للتدقيق، وقابل للتفسير، مما يعالج الفجوات الحرجة في الأبحاث الحالية المتعلقة بالخصوصية، والأمان، وحوكمة النموذج. بشكل عام، تؤكد النتائج على الإمكانات التحويلية للتعلم الفيدرالي والبلوكشين في تطوير حلول رعاية صحية قوية ومحافظة على الخصوصية.

Journal: Scientific Reports, Volume: 15, Issue: 1
DOI: https://doi.org/10.1038/s41598-025-04083-4
PMID: https://pubmed.ncbi.nlm.nih.gov/40595873
Publication Date: 2025-07-01
Author(s): Tanisha Bhardwaj et al.
Primary Topic: Privacy-Preserving Technologies in Data

Overview

The study presents the PPFBXAIO framework, which integrates federated learning, blockchain technology, and explainable AI to address significant challenges in decentralized machine learning systems, particularly in the healthcare sector. Utilizing techniques such as LGOA for feature selection, Min-Max normalization, SHA-256 encryption, and smart contracts for model aggregation, the framework ensures data privacy, model transparency, and resilience against poisoning attacks. Evaluations on the Heart Disease dataset and the Wisconsin Breast Cancer dataset demonstrated impressive performance metrics, with PPFBXAIO achieving 95.07% accuracy and 95.98% F1 score for breast cancer, and 93.07% accuracy and 93.24% F1 score for heart disease, surpassing existing federated learning models like FedAvg and FL-MPC.

Despite its promising results, the framework encounters challenges related to increased computational and communication overhead due to blockchain integration, which may hinder scalability and real-time deployment, especially as the number of nodes increases. Future research will focus on implementing lightweight consensus mechanisms, exploring off-chain storage solutions, incorporating multi-modal medical data for enhanced diagnostics, integrating advanced explainable AI techniques, and conducting clinical usability studies to validate the framework in practical settings. These efforts aim to refine PPFBXAIO into a scalable, effective, and privacy-preserving solution for decentralized healthcare systems.

Methods

The proposed methodology in this research introduces the Privacy Preserving Federated Blockchain Explainable Artificial Intelligence Optimization (PPFBXAIO) Framework, which integrates blockchain technology with federated learning (FL) and explainable artificial intelligence (XAI) to enhance data management security, transparency, and privacy. The framework employs an aggregation algorithm to filter out anomalous model parameters, thereby mitigating poisoning attacks while maintaining data privacy. It leverages local SHAP-based interpretation methods to validate model predictions without accessing centralized data, facilitating user participation and improving the interpretability of AI model decisions. The integration of a Local Genetic Optimization Algorithm (LGOA) optimizes hyperparameters and selects significant dataset attributes, enhancing model performance in non-IID data distributions.

The effectiveness of the PPFBXAIO framework was evaluated using publicly available datasets for heart disease and breast cancer, achieving precision rates of 91.19% and 95.44%, respectively, outperforming other methods such as FedHFP+RNN and FedAvgBC+TabNet. The framework was also tested on larger, real-world datasets, including MIMIC-III EHR and NIH Chest X-ray14, demonstrating scalability and generalizability while preserving privacy. The proposed classifier achieved the highest recall (95.39% for heart disease) and f-measure (93.24% for heart disease) among the compared methods. Additionally, a new consensus mechanism, Proof-of-Quality (PoQ), was introduced to enhance model aggregation based on explanation consistency and historical trust, further improving performance and transparency in federated learning environments.

Results

The research implements federated learning (FL) methods utilizing blockchain technology and Python 3.10.11 for model training and noise generation, employing PyTorch 2.0.1. Communication between nodes is achieved through Go’s native HTTP client interfacing with a Python-based Tornado HTTP server. The experiments are conducted on a PC equipped with an Intel i5-13400 CPU, 32 GB of DRAM, and an NVIDIA RTX 4070 Ti GPU, focusing on datasets related to heart disease and breast cancer.

The study compares the proposed model against established FL schemes, including FedAvg, Federated Learning with Multi-Party Computation (FL-MPC), FL-RAEC, a privacy-preserved and efficient FL framework with blockchain (PEFL), and Privacy-Preserving Blockchain Enabled Federated Learning (PPBEFL). The results indicate that the proposed model exhibits superior resilience against various attack models. The datasets utilized are deemed appropriate for federated learning due to their complexity, diverse features, and significance in critical healthcare predictions, as summarized in Table 1, which outlines the experimental setup of the simulation model.

Discussion

The discussion section of the research paper reviews various advancements in federated learning (FL) and its integration with blockchain technology, emphasizing their applications in healthcare and data privacy. Wang et al. introduced the Platform-Free Proof of Federated Learning (PF-PoFL), which utilizes energy-efficient consensus mechanisms to enhance FL tasks while addressing security vulnerabilities such as Sybil attacks through innovative transaction types and credit-based incentives. The incorporation of differential privacy at the user level further safeguards against implicit privacy leaks during model training. Other studies, such as those by Singh et al. and Gupta et al., highlight the potential of FL in smart healthcare systems, enabling secure data access and model training without compromising patient privacy.

Additionally, various methodologies for model aggregation and feature selection are discussed, including the use of ensemble methods and the Levy Grasshopper Optimization Algorithm (LGOA) for effective feature selection in disease prediction. The paper emphasizes the importance of explainable AI (XAI) in enhancing model transparency and accountability, particularly in regulated fields like healthcare. The integration of XAI with federated learning and blockchain is proposed as a means to create a secure, auditable, and interpretable AI framework, addressing critical gaps in current research related to privacy, security, and model governance. Overall, the findings underscore the transformative potential of federated learning and blockchain in developing robust, privacy-preserving healthcare solutions.

شارك: