ProFed: معيار للتعلم الفيدرالي غير المستقل عن الهوية القائم على القرب
ProFed: A Benchmark for Proximity-Based Non-IID Federated Learning

شارك:
المجلة: Journal of Open Research Software، المجلد: 14
DOI: https://doi.org/10.5334/jors.624
تاريخ النشر: 2026-01-01
المؤلف: Davide Domini وآخرون
الموضوع الرئيسي: التقنيات التي تحافظ على الخصوصية في البيانات

نظرة عامة

يقدم هذا القسم PROFED، وهو معيار مصمم لتعزيز تقييم خوارزميات التعلم الفيدرالي (FL) تحت بيانات العملاء غير المستقلة والموزعة بشكل متطابق (non-IID)، والتي تعكس غالبًا التباينات الجغرافية. عادةً ما تقوم تقييمات FL التقليدية بتوزيع بيانات non-IID بشكل عشوائي عبر الأجهزة، متجاهلة التوزيع المكاني الذي يمكن أن يؤثر بشكل كبير على الأداء. يتناول PROFED هذا من خلال محاكاة تقسيمات البيانات مع درجات متفاوتة من الانحراف عبر مناطق مختلفة، باستخدام طرق الانحراف المعروفة على مجموعات بيانات شائعة مثل MNIST وCIFAR-10 وUTKFace. يهدف هذا الإطار إلى توفير نهج موحد للباحثين لتقييم خوارزميات FL مقابل معايير ثابتة.

يختتم المؤلفون بالتأكيد على قابلية استخدام PROFED العملية وإمكاناته في تسهيل التقييمات الواقعية لخوارزميات FL. تشمل اتجاهات البحث المستقبلية التحقيق في تأثير تقسيم المستوى الإقليمي على ديناميات التعلم مقارنةً باستراتيجيات non-IID التقليدية، بالإضافة إلى استكشاف سلوك التقارب والصلابة تجاه تباين البيانات. بالإضافة إلى ذلك، يخطط المؤلفون لتوسيع PROFED من خلال دمج المزيد من مجموعات البيانات وخوارزميات الأساس، لا سيما في مجالات التعلم الفيدرالي المخصص والمجموعة، لتعزيز مرونته. تم تصميم PROFED ليكون مستقلًا عن المنصة، ومتوافقًا مع أي نظام تشغيل يدعم Python 3.12 وما فوق.

مقدمة

تناقش مقدمة الورقة الاهتمام المتزايد في التعلم الفيدرالي (FL) كحل لمخاوف الخصوصية المرتبطة بجمع البيانات المركزية. يتيح FL تدريب نموذج عالمي مع الحفاظ على بيانات المستخدم موزعة. ومع ذلك، تشير الأبحاث الحالية إلى أن FL يعمل بشكل جيد مع البيانات الموزعة بشكل متجانس ولكنه يواجه صعوبات مع البيانات غير المستقلة والموزعة بشكل متطابق (non-IID). يتضح هذا بشكل خاص في توقع حركة المرور الحضرية، حيث تظهر البيانات من المناطق الجغرافية القريبة ارتباطات مكانية قوية، مما يؤدي إلى دقة تنبؤية أفضل داخل نفس المنطقة مقارنة بالمناطق المختلفة.

لمعالجة التحديات التي تطرحها تباينات البيانات، تم اقتراح خوارزميات مختلفة مثل Scaffold وFedProx، والتي تفترض عادةً بيانات IID داخل مجموعات العملاء. ومع ذلك، تكشف السيناريوهات الواقعية غالبًا أن البيانات من الأجهزة القريبة تشترك في المزيد من التشابهات مقارنة بتلك البعيدة. بينما حاولت بعض الدراسات تجميع العملاء بناءً على مقاييس التشابه، لا يزال نقص المعايير الموحدة لتقييم هذه الأساليب يمثل قيدًا كبيرًا. لسد هذه الفجوة، يقدم المؤلفون PROFED، وهو معيار جديد مصمم خصيصًا للتعلم الفيدرالي غير المستقل والموزع بناءً على القرب. يستخدم PROFED مجموعات بيانات معروفة في رؤية الكمبيوتر وطرق تقسيم البيانات، مما يسمح للباحثين بالتلاعب بانحراف البيانات لتجارب وتحليلات أكثر دقة. يتم التأكيد على فعاليته من خلال تطبيقه في مساهمات علمية متنوعة.

نقاش

يسلط قسم النقاش في الورقة الضوء على تطور المعايير في التعلم الفيدرالي (FL)، مشددًا على التحول من تقسيمات البيانات المتجانسة إلى سيناريوهات أكثر تعقيدًا تأخذ في الاعتبار تحولات البيانات والتوزيعات المكانية. حققت الأطر الحالية مثل FedScale وLEAF تقدمًا في تقييم جوانب مختلفة من FL، لكنها غالبًا ما تتجاهل التوزيع الجغرافي للأجهزة وارتباط البيانات بالموقع. لمعالجة هذه الفجوة، يقدم المؤلفون PROFED، وهو معيار يحاكي تقسيمات البيانات مع درجات متفاوتة من الانحراف عبر مناطق مختلفة، مما يمكّن من تقييم أكثر واقعية لخوارزميات FL.

تم تنفيذ PROFED باستخدام PyTorch وTorchVision، مما يسمح بإنشاء تقسيمات بيانات مدركة للمناطق تعكس الطبيعة غير المستقلة والموزعة للبيانات في العالم الحقيقي. يدعم الإطار طرق تقسيم مختلفة، بما في ذلك تقسيم Dirichlet وتقسيم صعب، ويوفر تحكمًا دقيقًا في توزيع البيانات بين الأجهزة. أجرى المؤلفون تجارب باستخدام خوارزميات متطورة (FedAvg وFedProx وScaffold) لتقييم فعالية PROFED، كاشفين أن الأداء يتدهور بشكل كبير تحت ظروف non-IID، خاصة مع انحراف البيانات الشديد. تؤكد النتائج على قيود الأساليب الحالية في FL في التعامل مع توزيعات البيانات غير المتجانسة وتقترح أن هناك حاجة لمزيد من البحث لتطوير حلول أكثر قوة.

Journal: Journal of Open Research Software, Volume: 14
DOI: https://doi.org/10.5334/jors.624
Publication Date: 2026-01-01
Author(s): Davide Domini et al.
Primary Topic: Privacy-Preserving Technologies in Data

Overview

The section presents PROFED, a benchmark designed to enhance the evaluation of Federated Learning (FL) algorithms under non-independent and identically distributed (non-IID) client data, which often reflects geographic variations. Traditional FL evaluations typically randomize non-IID data across devices, neglecting the spatial distribution that can significantly influence performance. PROFED addresses this by simulating data splits with varying degrees of skewness across different regions, utilizing established skewness methods on popular datasets such as MNIST, CIFAR-10, and UTKFace. This framework aims to provide a standardized approach for researchers to assess FL algorithms against consistent baselines.

The authors conclude by emphasizing the practical usability of PROFED and its potential for facilitating realistic evaluations of FL algorithms. Future research directions include investigating the impact of region-level partitioning on learning dynamics compared to conventional non-IID strategies, as well as exploring convergence behavior and robustness to data heterogeneity. Additionally, the authors plan to expand PROFED by incorporating more datasets and baseline algorithms, particularly in the realms of clustered and personalized FL, to enhance its versatility. PROFED is designed to be platform-independent, compatible with any operating system that supports Python 3.12 and above.

Introduction

The introduction of the paper discusses the growing interest in Federated Learning (FL) as a solution to privacy concerns associated with centralized data collection. FL enables the training of a global model while keeping user data decentralized. However, existing research indicates that FL performs well with homogeneously distributed data but struggles with non-independently and identically distributed (non-IID) data. This is particularly evident in urban traffic prediction, where data from geographically proximate areas exhibit strong spatial correlations, leading to better predictive accuracy within the same district compared to different ones.

To address the challenges posed by data heterogeneity, various algorithms such as Scaffold and FedProx have been proposed, typically assuming IID data within client clusters. However, real-world scenarios often reveal that data from nearby devices share more similarities than those from distant ones. While some studies have attempted to cluster clients based on similarity metrics, the lack of standardized benchmarks for evaluating these approaches remains a significant limitation. To fill this gap, the authors introduce PROFED, a novel benchmark tailored for proximity-based non-IID FL. PROFED utilizes established computer vision datasets and data partitioning methods, allowing researchers to manipulate data skewness for more nuanced experimentation and analysis. Its effectiveness is underscored by its application in various scientific contributions.

Discussion

The discussion section of the paper highlights the evolution of benchmarks in federated learning (FL), emphasizing the shift from homogeneous data splits to more complex scenarios that account for data shifts and spatial distributions. Existing frameworks like FedScale and LEAF have made strides in evaluating various aspects of FL, but they often overlook the geographical distribution of devices and the correlation of data with location. To address this gap, the authors introduce PROFED, a benchmark that simulates data splits with varying degrees of skewness across different regions, thereby enabling a more realistic evaluation of FL algorithms.

PROFED is implemented using PyTorch and TorchVision, allowing for the creation of region-aware data partitions that reflect the non-IID nature of real-world data. The framework supports various partitioning methods, including Dirichlet and hard partitioning, and provides fine-grained control over data distribution among devices. The authors conducted experiments using state-of-the-art algorithms (FedAvg, FedProx, and Scaffold) to assess PROFED’s effectiveness, revealing that performance significantly degrades under non-IID conditions, particularly with extreme data skewness. The findings underscore the limitations of current FL approaches in handling heterogeneous data distributions and suggest that further research is needed to develop more robust solutions.

شارك: