DOI: https://doi.org/10.1038/s41598-026-49311-7
PMID: https://pubmed.ncbi.nlm.nih.gov/42000902
تاريخ النشر: 2026-04-18
المؤلف: Raja Vavekanand وآخرون
الموضوع الرئيسي: تعلم الآلة في الرعاية الصحية
نظرة عامة
يقدم هذا القسم نظرة عامة على التحديات والحلول المتعلقة بالقيم المفقودة في بيانات الرعاية الصحية متعددة الأنماط الزمنية، وخاصة في بيئات الرعاية الحرجة. تعيق طرق الاستيفاء المركزية التقليدية اللوائح المتعلقة بالخصوصية، وغالبًا ما تفتقر النماذج المؤسسية الفردية إلى القابلية للتعميم. يقدم البحث Fed-HealthImp، وهو إطار تعلم اتحادي مصمم لاستيفاء القيم المفقودة في بيانات السلاسل الزمنية السريرية متعددة المتغيرات، التي يتم أخذ عينات منها بشكل غير منتظم. يستخدم هذا الإطار نموذج استيفاء قائم على الانتباه الذاتي مع وزن عميل تكيفي لإدارة توزيعات البيانات غير المستقلة وغير المتطابقة عبر المستشفيات بشكل فعال.
تظهر تقييمات Fed-HealthImp على ثلاثة مجموعات بيانات ICU حقيقية—eICU-CRD وMIMIC-IV وHiRID—أنه يحقق جودة استيفاء ضمن 3.5% من نموذج مركزي ينتهك الخصوصية، ويتجاوز بشكل كبير أساليب التدريب المحلية فقط، ويعزز توقعات الوفيات اللاحقة، كما يتضح من زيادة في AUROC تصل إلى 3.4%. تؤكد النتائج على إمكانيات الإطار كحل عملي يحافظ على الخصوصية لتطوير نماذج استيفاء قوية من بيانات ICU العالمية المجزأة، مما يسهل ممارسات الرعاية الحرجة الأكثر تعاونًا والمبنية على البيانات.
الطرق
في هذا القسم، يحدد المؤلفون الطرق الأساسية المستخدمة للمقارنة في دراستهم، والتي تشمل مجموعة من تقنيات الاستيفاء. تشمل الطرق خوارزميات سريرية بسيطة مثل استيفاء المتوسط/الوسيط واستيفاء آخر ملاحظة تم نقلها للأمام (LOCF)، والتي تعمل كمعايير أساسية. بالإضافة إلى ذلك، يستخدمون الاستيفاء المتعدد بواسطة المعادلات المتسلسلة (MICE) كنهج كلاسيكي، إلى جانب طرق التعلم العميق المتقدمة مثل GRU-D، وهو شبكة عصبية متكررة تأخذ في الاعتبار الفقدان مع التغطية وتدهور الفجوة الزمنية، وBRITS، التي تستخدم الاستيفاء المتكرر ثنائي الاتجاه مع تنظيم الاتساق.
علاوة على ذلك، يقيم المؤلفون الأساليب المركزية والاتحادية، بما في ذلك الاستيفاء الذاتي المركزي القائم على الانتباه مع بيانات مجمعة (يعتبر طريقة أوراكل بسبب طبيعته المنتهكة للخصوصية) واستيفاء الانتباه الذاتي الاتحادي العادي (FedAvg-SAITS)، الذي يتم تدريبه باستخدام تقنيات الت averaging الاتحادي القياسية عبر العملاء. يهدف هذا المقارنة الشاملة إلى تقييم فعالية الطرق المقترحة ضد استراتيجيات الاستيفاء المعتمدة.
النتائج
في هذه الدراسة، تم محاكاة بيئة تعلم اتحادي تشمل 50 عميلًا، تتكون من 40 من مجموعة بيانات eICU-CRD، و5 من MIMIC-IV، و5 من HiRID، كل منها يمثل بيئات مستشفيات متنوعة ومجموعات مرضى. تم تقسيم البيانات لكل عميل إلى مجموعات تدريب (70%)، والتحقق (10%)، والاختبار (20%)، مع تلخيص خصائص هذه المجموعات في الجدول 2. تم إجراء عملية التدريب على مدى 100 جولة اتحادية، باستخدام نسبة عميل \( C = 0.3 \)، وحقب محلية \( E = 5 \)، وحجم دفعة 32. تم تعديل معدل التعلم باستخدام التخفيف الكوني، يتراوح من \( 10^{-3} \) إلى \( 10^{-5} \).
تم تقييم أداء الطريقة المقترحة مقابل ستة نماذج أساسية، كما هو موضح في القسم 3.6. توفر النتائج من هذه التجارب رؤى حول فعالية التعلم الاتحادي في التعامل مع مجموعات بيانات الرعاية الصحية المتنوعة مع الحفاظ على خصوصية البيانات وسلامتها عبر بيئات سريرية مختلفة.
المناقشة
تسلط قسم المناقشة في ورقة البحث الضوء على التقدم المحرز في معالجة البيانات المفقودة في السلاسل الزمنية للرعاية الصحية من خلال إطار Fed-HealthImp المقترح. يستفيد هذا الإطار من التعلم الاتحادي (FL) لتمكين استيفاء القيم المفقودة مع الحفاظ على الخصوصية في مجموعات بيانات ICU غير المتجانسة، متجاوزًا قيود الطرق الإحصائية التقليدية وتعزيز أداء نماذج التعلم العميق. من خلال استخدام نموذج استيفاء قائم على الانتباه الذاتي ونظام وزن عميل تكيفي، يحقق Fed-HealthImp دقة استيفاء ضمن 3.5% من نموذج مركزي بينما يتفوق بشكل كبير على أساليب التدريب المحلية فقط. هذا أمر ذو صلة خاصة في البيئات السريرية حيث يكون التفسير الدقيق للإشارات الفسيولوجية أمرًا حاسمًا لرعاية المرضى.
تناقش الورقة أيضًا التحديات الفريدة التي تطرحها التباين الإحصائي والزمني عبر مؤسسات الرعاية الصحية المختلفة، والتي يمكن أن تؤثر على استقرار ودقة القيم المستوفاة. يعالج آلية وزن العميل التكيفية هذه التحديات من خلال إعطاء الأولوية للتحديثات من العملاء الذين تتماشى توزيعات بياناتهم بشكل أقرب مع الهدف التعليمي العالمي، مما يحسن من قوة النموذج عبر مجموعات بيانات متنوعة. بشكل عام، تؤكد النتائج على إمكانيات أطر التعلم الاتحادي مثل Fed-HealthImp في التوفيق بين الحاجة إلى استيفاء عالي الجودة مع اللوائح الصارمة لخصوصية البيانات، مما يمهد الطريق لاتخاذ قرارات سريرية أكثر فعالية في بيئات الرعاية الصحية الواقعية.
DOI: https://doi.org/10.1038/s41598-026-49311-7
PMID: https://pubmed.ncbi.nlm.nih.gov/42000902
Publication Date: 2026-04-18
Author(s): Raja Vavekanand et al.
Primary Topic: Machine Learning in Healthcare
Overview
The section presents an overview of the challenges and solutions related to missing values in multimodal time-series healthcare data, particularly in critical care settings. Traditional centralized imputation methods are hindered by privacy regulations, and single-institution models often lack generalizability. The paper introduces Fed-HealthImp, a federated learning framework designed for deep imputation of missing values in multivariate, irregularly sampled clinical time-series data. This framework utilizes a self-attention-based imputation model with adaptive client weighting to effectively manage non-IID data distributions across hospitals.
The evaluation of Fed-HealthImp on three real-world ICU datasets—eICU-CRD, MIMIC-IV, and HiRID—demonstrates that it achieves imputation quality within 3.5% of a centralized model that violates privacy, significantly surpasses local-only training approaches, and enhances downstream mortality prediction, as indicated by an increase in AUROC by up to 3.4%. The findings underscore the framework’s potential as a practical, privacy-preserving solution for developing robust imputation models from fragmented global ICU data, thereby facilitating more collaborative and data-informed critical care practices.
Methods
In this section, the authors outline the baseline methods employed for comparison in their study, which encompass a range of imputation techniques. The methods include simple clinical heuristics such as mean/median imputation and last-observation-carried-forward (LOCF), which serve as basic benchmarks. Additionally, they utilize Multiple Imputation by Chained Equations (MICE) as a classical approach, alongside advanced deep learning methods like GRU-D, a recurrent neural network that accounts for missingness with masking and time-gap decay, and BRITS, which employs bidirectional recurrent imputation with consistency regularization.
Furthermore, the authors assess centralized and federated approaches, including Centralised Self-Attention Imputation with pooled data (considered an oracle method due to its privacy-violating nature) and Vanilla Federated Averaging Self-Attention Imputation (FedAvg-SAITS), which is trained using standard federated averaging techniques across clients. This comprehensive comparison aims to evaluate the effectiveness of their proposed methods against established imputation strategies.
Results
In this study, a federated learning environment was simulated involving 50 clients, comprising 40 from the eICU-CRD dataset, 5 from MIMIC-IV, and 5 from HiRID, each representing diverse hospital settings and patient cohorts. The data for each client was partitioned into training (70%), validation (10%), and test (20%) sets, with the characteristics of these datasets summarized in Table 2. The training process was conducted over 100 federated rounds, utilizing a client fraction \( C = 0.3 \), local epochs \( E = 5 \), and a batch size of 32. The learning rate was adjusted using cosine annealing, ranging from \( 10^{-3} \) to \( 10^{-5} \).
The performance of the proposed method was evaluated against six baseline models, as detailed in Section 3.6. The results from these experiments provide insights into the effectiveness of federated learning in handling diverse healthcare datasets while maintaining data privacy and integrity across different clinical environments.
Discussion
The discussion section of the research paper highlights the advancements made in addressing missing data in healthcare time-series through the proposed Fed-HealthImp framework. This framework leverages federated learning (FL) to enable privacy-preserving imputation of missing values in heterogeneous ICU datasets, circumventing the limitations of traditional statistical methods and enhancing the performance of deep learning models. By employing a self-attention-based imputation model and an adaptive client weighting scheme, Fed-HealthImp achieves imputation accuracy that is within 3.5% of a centralized model while significantly outperforming local-only training approaches. This is particularly relevant in clinical settings where the accurate interpretation of physiological signals is critical for patient care.
The paper also discusses the unique challenges posed by statistical and temporal heterogeneity across different healthcare institutions, which can affect the stability and accuracy of imputed values. The adaptive client weighting mechanism addresses these challenges by prioritizing updates from clients whose data distributions align more closely with the global learning objective, thus improving the robustness of the model across diverse datasets. Overall, the findings underscore the potential of federated learning frameworks like Fed-HealthImp to reconcile the need for high-quality imputation with stringent data privacy regulations, paving the way for more effective clinical decision-making in real-world healthcare environments.
