أخذ عينات العملاء المثلى في التعلم الفيدرالي مع خصوصية تفاضلية غير متجانسة على مستوى العميل
Optimal Client Sampling in Federated Learning With Client-Level Heterogeneous Differential Privacy

شارك:
المجلة: IEEE Internet of Things Journal، المجلد: 13، العدد: 9
DOI: https://doi.org/10.1109/jiot.2026.3661913
تاريخ النشر: 2026-02-06
المؤلف: Zhenyun Du وآخرون
الموضوع الرئيسي: التقنيات التي تحافظ على الخصوصية في البيانات

نظرة عامة

تتناول ورقة البحث التحديات المتعلقة بتنفيذ الخصوصية التفاضلية على مستوى العميل (DP) في التعلم الفيدرالي (FL) عندما يكون لدى العملاء متطلبات خصوصية غير متجانسة. تفرض الطرق التقليدية، مثل DP-FedAvg، مستوى خصوصية موحد عبر جميع العملاء، مما قد يؤدي إلى ضوضاء DP مفرطة وانخفاض كبير في فائدة النموذج. يقترح المؤلفون نهجًا جديدًا يسمى GDPFed، والذي يجمع العملاء بناءً على ميزانياتهم للخصوصية، مما يسمح بتطبيق DP على مستوى العميل داخل كل مجموعة. تقلل هذه الطريقة من هدر ميزانية الخصوصية وتعزز أداء النموذج، خاصةً في ظل نموذج هجوم عملي حيث يكون كل من العملاء والخادم صادقين ولكن فضوليين.

لتحسين GDPFed بشكل أكبر، يقدم المؤلفون GDPFed+، الذي يتضمن تقليل النموذج لتقليل الضوضاء غير الضرورية ويضبط نسب أخذ العينات لكل مجموعة من العملاء لتقليل خطأ التقارب. تظهر تقييماتهم التجريبية الواسعة عبر مجموعات بيانات مرجعية مختلفة أن GDPFed+ يتفوق بشكل كبير على الطرق الحالية الرائدة، محققًا مكاسب كبيرة في دقة النموذج بينما يستوعب بشكل فعال تفضيلات الخصوصية المتنوعة بين العملاء. تؤكد النتائج على إمكانيات GDPFed وGDPFed+ في تعزيز أطر التعلم الفيدرالي التي تعطي الأولوية لكل من الخصوصية وفائدة النموذج.

مقدمة

في مقدمة هذه الورقة البحثية، يبرز المؤلفون قيود أطر التعلم الآلي المركزية التقليدية، لا سيما فيما يتعلق بالخصوصية وأعباء الاتصال في أنظمة إنترنت الأشياء (IoT). لمعالجة هذه التحديات، ظهر التعلم الفيدرالي (FL) كنهج لامركزي يسمح بتدريب النماذج بشكل تعاوني دون مركزية البيانات الحساسة. تؤكد الورقة على أهمية الخصوصية في FL، خاصة في التطبيقات مثل الرعاية الصحية والاستشعار عن بعد، حيث يمكن للخصوم استغلال تحديثات النموذج لاستنتاج معلومات خاصة. لتعزيز الخصوصية، يركز المؤلفون على الخصوصية التفاضلية على مستوى العميل (DP) داخل FL، مما يوفر فائدة أفضل للنموذج مقارنةً بـ DP على مستوى السجل.

تقدم الورقة التعلم الفيدرالي القائم على الخصوصية التفاضلية الجماعية (GDPFed)، وهو نهج جديد مصمم لاستيعاب تفضيلات الخصوصية غير المتجانسة بين العملاء مع الحفاظ على فائدة عالية للنموذج. من خلال تجميع العملاء وفقًا لميزانياتهم للخصوصية وتطبيق DP على مستوى العميل على مستوى المجموعة، يقلل GDPFed من قيود الطرق الحالية التي تفرض ميزانية الخصوصية الأكثر صرامة عبر جميع العملاء. يقترح المؤلفون أيضًا GDPFed +، الذي يتضمن تقليل النموذج ويقوم بتحسين نسب أخذ العينات للعملاء لتعزيز فائدة النموذج بشكل أكبر مع الحفاظ على ضمانات الخصوصية. تعد الورقة بتقييمات شاملة ضد الطرق الرائدة، مما يظهر أن GDPFed وGDPFed + يحسنان الأداء بشكل كبير في إعدادات DP غير المتجانسة.

الطرق

في هذا القسم، يستكشف المؤلفون ديناميات نسب أخذ العينات للعملاء وأداء طرق التعلم الفيدرالي المختلفة تحت قيود الخصوصية التفاضلية. يقومون بإجراء تجارب على مجموعات بيانات FMNIST وCIFAR-10، موضحين كيف تتكيف نسب أخذ العينات المثلى للعملاء استجابةً لميزانيات الخصوصية المتغيرة. على وجه التحديد، مع زيادة ميزانية الخصوصية $\epsilon_1$، ترتفع نسبة أخذ العينات للمجموعة 1 ($q_1$) بينما تنخفض تلك للمجموعات 2 و3 ($q_2$ و$q_3$) للحفاظ على قيود عالمية. تشير النتائج إلى أن الطريقة المقترحة، GDPFed +، تتفوق باستمرار على المعايير الأخرى، محققة دقة متوسطة تبلغ 56.93% عبر جميع مجموعات البيانات، وهو تحسن بنسبة 5.75% مقارنةً بـ GDPFed وتتفوق على الطرق الرائدة مثل PFA وIDP-FedAvg.

كما يقوم المؤلفون بتحليل سلوك التقارب للطرق المختلفة، مشيرين إلى أن DP-FedAvg يعاني من تدهور الأداء بسبب حقن الضوضاء الموحد عبر العملاء، مما يؤثر سلبًا على تعلم النموذج. بالمقابل، يطبق GDPFed الخصوصية التفاضلية على مستوى المجموعة، مما يقلل من الضوضاء غير الضرورية ويؤدي إلى نتائج تقارب أفضل. تستكشف الدراسة أيضًا تأثير مستويات الضوضاء، وتفضيلات خصوصية العملاء، وتنوع البيانات على أداء النموذج، كاشفة أن GDPFed + لا يقلل فقط من إجمالي الضوضاء المضافة إلى تحديثات النموذج ولكن أيضًا يحافظ على أداء قوي عبر توزيعات العملاء المختلفة. بشكل عام، تسلط النتائج الضوء على فعالية GDPFed + في تحقيق التوازن بين الخصوصية والفائدة في أنظمة التعلم الفيدرالي.

المناقشة

في هذا القسم، يناقش المؤلفون تنفيذ الخصوصية التفاضلية على مستوى العميل (DP) ضمن إطار التعلم الفيدرالي (FL)، لا سيما تحت نموذج هجوم أكثر تحديًا حيث قد يكون كل من العملاء والخادم معاديين. تقسم الطريقة المقترحة، GDPFed، العملاء إلى مجموعات بناءً على ميزانياتهم الفردية للخصوصية، مما يسمح بحقن ضوضاء مخصصة لضمان DP مع تحسين فائدة النموذج. يبرز المؤلفون أهمية الآلية الغاوسية لإضافة الضوضاء ويشرحون الأسس النظرية التي تضمن الخصوصية لكل مجموعة، بالإضافة إلى النظام ككل.

يكشف التحليل أن الضوضاء المطلوبة للحفاظ على الخصوصية تتأثر بعدة عوامل، بما في ذلك عدد العملاء الذين تم أخذ عينات منهم وأبعاد النموذج. يقترح المؤلفون تقنية تقليل لتحسين فائدة النموذج من خلال تقليل عدد المعلمات الخاضعة للضوضاء، مما يقلل من تأثير DP على أداء النموذج. كما يقدمون نسخة محسنة، GDPFed+، التي تتضمن نسب أخذ عينات مثلى للعملاء واستراتيجيات تقليل لتحسين التوازن بين الخصوصية والفائدة. يشير تحليل التقارب إلى أنه تحت ظروف معينة، يمكن أن يحقق GDPFed أداءً مرضيًا مع الالتزام بقيود الخصوصية التي وضعتها تفضيلات العملاء غير المتجانسة.

Journal: IEEE Internet of Things Journal, Volume: 13, Issue: 9
DOI: https://doi.org/10.1109/jiot.2026.3661913
Publication Date: 2026-02-06
Author(s): Zhenyun Du et al.
Primary Topic: Privacy-Preserving Technologies in Data

Overview

The research paper addresses the challenges of implementing client-level differential privacy (DP) in federated learning (FL) when clients have heterogeneous privacy requirements. Traditional methods, such as DP-FedAvg, enforce a uniform privacy level across all clients, which can lead to excessive DP noise and a significant drop in model utility. The authors propose a novel approach called GDPFed, which groups clients based on their privacy budgets, allowing for client-level DP within each group. This method reduces privacy budget waste and enhances model performance, particularly under a practical attack model where both clients and the server are honest-but-curious.

To further optimize GDPFed, the authors introduce GDPFed+, which incorporates model sparsification to minimize unnecessary noise and fine-tunes per-group client sampling ratios to reduce convergence error. Their extensive empirical evaluations across various benchmark datasets demonstrate that GDPFed+ significantly outperforms existing state-of-the-art methods, achieving substantial gains in model accuracy while effectively accommodating diverse privacy preferences among clients. The findings underscore the potential of GDPFed and GDPFed+ in advancing federated learning frameworks that prioritize both privacy and model utility.

Introduction

In the introduction of this research paper, the authors highlight the limitations of traditional centralized Machine Learning (ML) frameworks, particularly concerning privacy and communication burdens in Internet of Things (IoT) systems. To address these challenges, Federated Learning (FL) has emerged as a decentralized approach that allows collaborative model training without centralizing sensitive data. The paper emphasizes the importance of privacy in FL, especially in applications like healthcare and remote sensing, where adversaries can exploit model updates to infer private information. To enhance privacy, the authors focus on client-level Differential Privacy (DP) within FL, which offers better model utility compared to record-level DP.

The paper introduces Group-based Differentially Private Federated Learning (GDPFed), a novel approach designed to accommodate heterogeneous privacy preferences among clients while maintaining high model utility. By grouping clients according to their privacy budgets and applying client-level DP at the group level, GDPFed mitigates the limitations of existing methods that enforce the most stringent privacy budget across all clients. The authors also propose GDPFed +, which incorporates model sparsification and optimizes client sampling ratios to further enhance model utility while preserving privacy guarantees. The paper promises extensive evaluations against state-of-the-art methods, demonstrating that GDPFed and GDPFed + significantly improve performance in heterogeneous DP settings.

Methods

In this section, the authors investigate the dynamics of client sampling ratios and the performance of various federated learning methods under differential privacy constraints. They conduct experiments on the FMNIST and CIFAR-10 datasets, demonstrating how optimal client sampling ratios adjust in response to varying privacy budgets. Specifically, as the privacy budget $\epsilon_1$ increases, the sampling ratio for Group 1 ($q_1$) rises while those for Groups 2 and 3 ($q_2$ and $q_3$) decrease to maintain a global constraint. The results indicate that the proposed method, GDPFed +, consistently outperforms other baselines, achieving an average accuracy of 56.93% across all datasets, which is a 5.75% improvement over GDPFed and surpasses state-of-the-art methods like PFA and IDP-FedAvg.

The authors also analyze the convergence behavior of different methods, noting that DP-FedAvg suffers from performance degradation due to uniform noise injection across clients, which adversely affects model learning. In contrast, GDPFed applies differential privacy at the group level, reducing unnecessary noise and leading to better convergence outcomes. The study further explores the impact of noise levels, client privacy preferences, and data heterogeneity on model performance, revealing that GDPFed + not only reduces the total noise added to model updates but also maintains robust performance across various client distributions. Overall, the findings highlight the effectiveness of GDPFed + in balancing privacy and utility in federated learning systems.

Discussion

In this section, the authors discuss the implementation of client-level differential privacy (DP) within a federated learning (FL) framework, particularly under a more challenging attack model where both clients and the server may be adversarial. The proposed method, GDPFed, partitions clients into groups based on their individual privacy budgets, allowing for tailored noise injection to ensure DP while optimizing model utility. The authors highlight the importance of the Gaussian mechanism for noise addition and detail the theoretical foundations that guarantee privacy for each group, as well as the overall system.

The analysis reveals that the noise required for maintaining privacy is influenced by several factors, including the number of clients sampled and the dimensionality of the model. The authors propose a sparsification technique to enhance model utility by reducing the number of parameters subject to noise, thereby mitigating the impact of DP on model performance. They also introduce an enhanced version, GDPFed+, which incorporates optimal client sampling ratios and sparsification strategies to further improve the balance between privacy and utility. The convergence analysis indicates that under certain conditions, GDPFed can achieve satisfactory performance while adhering to the privacy constraints set by heterogeneous client preferences.

شارك: