تخصيص الرتبة التكيفية للتعديل الدقيق الفعال للمعلمات للنماذج اللغوية
Adaptive Rank Allocation for Federated Parameter-Efficient Fine-Tuning of Language Models

شارك:
المجلة: IEEE Transactions on Computers، المجلد: 75، العدد: 4
DOI: https://doi.org/10.1109/tc.2026.3655161
تاريخ النشر: 2026-01-27
المؤلف: Zhenyun Du وآخرون
الموضوع الرئيسي: تقنيات معالجة اللغة الطبيعية

نظرة عامة

تقدم البحث FedARA، وهو إطار مبتكر مصمم لتعزيز الضبط الدقيق الفيدرالي الفعال من حيث المعلمات (FedPEFT) لنماذج اللغة المدربة مسبقًا (PLMs) في سياق بيئات البيانات غير المتجانسة. تحدد الدراسة قيودين رئيسيين في أساليب FedPEFT الحالية: تدهور الأداء الناتج عن عدم تجانس البيانات عبر الأجهزة وعدم كفاءة الاتصال بسبب تكوينات المعلمات الثابتة. لمعالجة هذه القضايا، يستخدم FedARA تحليل القيم الفردية المقطوعة (SVD) لتحسين تمثيل الميزات بين العملاء، وتخصيص الرتبة الديناميكي لتحسين كفاءة الاتصال، وتقليص الوحدات المعتمدة على الرتبة لتقليل تكاليف الحساب المحلي واستخدام الذاكرة خلال جولات التعلم الفيدرالي.

تظهر التجارب الواسعة أن FedARA يتفوق على الأساليب الأساسية بمتوسط يتراوح بين 6.95% إلى 8.49% عبر مجموعات بيانات ونماذج متنوعة، بينما يحقق أيضًا تحسينًا بمعدل 2.40× في كفاءة الاتصال. بالإضافة إلى ذلك، يقلل الإطار بشكل كبير من إجمالي وقت التدريب واستهلاك الطاقة على الأجهزة الطرفية، مع تخفيضات تصل إلى 48.90% و46.95%، على التوالي. تشير النتائج إلى أن FedARA لا يعزز الأداء تحت ظروف البيانات غير المستقلة وغير المتجانسة فحسب، بل يقلل أيضًا من استخدام الموارد، مما يمهد الطريق للبحوث المستقبلية لتوسيع تطبيقه على نماذج أكبر ودمج تقنيات التكميم لتحقيق مزيد من مكاسب الكفاءة.

مقدمة

تناقش مقدمة الورقة التقدمات الكبيرة التي حققتها نماذج اللغة المدربة مسبقًا (PLMs) في معالجة اللغة الطبيعية (NLP)، مع التأكيد على أهمية كل من مراحل التدريب المسبق والتعديل الدقيق. يتضمن التدريب المسبق التعلم الذاتي المراقب على مجموعات بيانات واسعة لتطوير تمثيلات لغوية عامة، بينما يقوم التعديل الدقيق بتكييف هذه النماذج مع مهام محددة من خلال التعلم المراقب. على الرغم من فعاليتها، تتطلب PLMs تعديلًا دقيقًا باستخدام بيانات محلية، مما يثير مخاوف تتعلق بالخصوصية. يُقترح التعلم الفيدرالي (FL) كحل للحفاظ على خصوصية البيانات من خلال الاحتفاظ بالبيانات الخام على الأجهزة المحلية، ولكن الحجم الكبير لـ PLMs يقدم تحديات في تكاليف الحساب والاتصال.

لمعالجة هذه التحديات، تقدم الورقة الضبط الدقيق الفيدرالي الفعال من حيث المعلمات (FedPEFT)، الذي يجمع بين FL وطرق الضبط الدقيق الفعالة من حيث المعلمات (PEFT) مثل التكيف منخفض الرتبة (LoRA). ومع ذلك، تواجه FedPEFT مشكلات تتعلق بالبيانات غير المستقلة وغير المتجانسة (non-IID)، والتي يمكن أن تؤدي إلى تدهور أداء النموذج، وعدم كفاءة تكوينات الرتبة الثابتة التي لا تتكيف مع حساسيات البيانات المتغيرة. يقترح المؤلفون نهجًا جديدًا، وهو تخصيص الرتبة التكيفية للـ PEFT الفيدرالي (FedARA)، الذي يتضمن ثلاث طرق مبتكرة: تكيف SVD المقطوع، تخصيص الرتبة الديناميكي، وتقليص الوحدات المعتمدة على الرتبة. تهدف هذه الطرق إلى تعزيز الأداء وكفاءة الاتصال مع تقليل تكاليف الموارد، مما يظهر تحسينات كبيرة في الدقة والكفاءة على الأجهزة ذات الموارد المحدودة.

طرق

في هذا القسم، يصف المؤلفون الطرق المستخدمة في بحثهم، مع التركيز على إطار FedARA، الذي يوسع نهج AdaLoRA للتعلم الفيدرالي (FL). تم بناء FedARA على مكتبة Hugging Face Transformers، مما يسهل تكوين المعلمات الديناميكي في سياق فيدرالي ويدعم نماذج معالجة اللغة الطبيعية (NLP) المدربة مسبقًا. يستخدم الإطار آلية اختيار عميل عشوائية، يمكن تعزيزها باستراتيجيات أكثر تعقيدًا، وينسق الاتصال بين العملاء والخادم من خلال المتوسطات الموزونة. بالإضافة إلى ذلك، ينفذ FedARA تقنيات تكميم النموذج، مثل QLoRA وQA-LoRA، لتقليل المتطلبات الحسابية وذاكرة الوصول العشوائي، مما قد يقلل من تكاليف الاتصال.

تشمل مقاييس التقييم للدراسة الدقة النهائية لمهام التصنيف، ودرجات ROUGE لمهام التلخيص، وإجمالي تكاليف الاتصال بعد عدد محدد مسبقًا من جولات FL، مما يضمن مقارنة عادلة للأداء. تشمل المقاييس الثانوية تكاليف الاتصال لكل جولة، تخصيص الرتبة التكيفية، إجمالي وقت التدريب، استخدام الذاكرة، واستهلاك الطاقة على الأجهزة الطرفية، مما يوفر تقييمًا شاملاً لكفاءة النظام وفعاليته. تُجرى التجارب عبر المحاكاة على جهاز كمبيوتر محمول مزود ببطاقة RTX 4070، مع قياس أوقات التدريب المحلية على ثلاث أجهزة طرفية تمثيلية: AGX Orin، Orin Nano، وRaspberry Pi 5. يفترض المؤلفون عرض نطاق ترددي للاتصال يبلغ 1 ميغابايت/ثانية بين الخادم والعملاء، مما يعكس ظروفًا واقعية نموذجية لشبكات إنترنت الأشياء، الواي فاي المنزلي، والشبكات الخلوية.

نقاش

في قسم النقاش من ورقة البحث، يتم تسليط الضوء على عدة ملاحظات رئيسية تتعلق بإطار الضبط الدقيق الفيدرالي الفعال من حيث المعلمات (FedPEFT). أولاً، يتفوق FedPEFT على الضبط الدقيق الكامل الفيدرالي (FedFFT) على الأجهزة الطرفية المتقدمة، محققًا دقة قابلة للمقارنة مع تقليل كبير في وقت التدريب المحلي (1.6×)، وتكاليف الاتصال (200×)، واستخدام الذاكرة (1.5×). ومع ذلك، لا يزال وقت التدريب الإجمالي على الأجهزة المتقدمة غير مستكشف بشكل كافٍ. من الجدير بالذكر أن أداء FedPEFT يتناقص تحت ظروف البيانات غير المستقلة وغير المتجانسة، مع انخفاضات متوسطة في الدقة تصل إلى 14.62% و12.81% على مجموعات بيانات 20News وNews Category، على التوالي. وهذا يشير إلى حاجة ملحة لاستراتيجيات لمعالجة الآثار السلبية للبيانات غير المستقلة وغير المتجانسة.

علاوة على ذلك، يؤثر موضع وحدات الضبط الدقيق الفعالة من حيث المعلمات (PEFT) داخل إطار FedPEFT على الأداء، حيث تؤدي تكوينات المعلمات الثابتة إلى عدم كفاءة الاتصال من خلال نقل المعلمات الأقل تأثيرًا بشكل متكرر. تكشف الدراسة أن الاتصال يصبح عنق الزجاجة الرئيسي على الأجهزة عالية الأداء، بينما يكون الحساب هو العامل المحدد على الأجهزة المنخفضة الأداء مثل Raspberry Pi. تشير هذه النتائج إلى أن الأبحاث المستقبلية يجب أن تركز على تحسين FedPEFT لمختلف الأجهزة الطرفية ومعالجة التحديات الفريدة التي تطرحها البيانات غير المستقلة وغير المتجانسة. يهدف إدخال FedARA، وهو إطار تخصيص الرتبة التكيفية، إلى تعزيز كفاءة الاتصال والتخفيف من قيود الأساليب الحالية من خلال دمج هيكل SVD المقطوع وتخصيص الرتبة الديناميكي.

Journal: IEEE Transactions on Computers, Volume: 75, Issue: 4
DOI: https://doi.org/10.1109/tc.2026.3655161
Publication Date: 2026-01-27
Author(s): Zhenyun Du et al.
Primary Topic: Natural Language Processing Techniques

Overview

The research presents FedARA, an innovative framework designed to enhance Federated Parameter-Efficient Fine-Tuning (FedPEFT) for Pre-trained Language Models (PLMs) in the context of heterogeneous data environments. The study identifies two primary limitations of existing FedPEFT approaches: the performance degradation caused by data heterogeneity across devices and communication inefficiencies due to fixed parameter configurations. To address these issues, FedARA employs truncated Singular Value Decomposition (SVD) for improved feature representation among clients, dynamic rank allocation to optimize communication efficiency, and rank-based module pruning to reduce local computational costs and memory usage during federated learning rounds.

Extensive experiments demonstrate that FedARA outperforms baseline methods by an average of 6.95% to 8.49% across various datasets and models, while also achieving a 2.40× improvement in communication efficiency. Additionally, the framework significantly reduces total training time and energy consumption on edge devices, with reductions of up to 48.90% and 46.95%, respectively. The findings suggest that FedARA not only enhances performance under non-IID data conditions but also minimizes resource usage, paving the way for future research to expand its application to larger models and integrate quantization techniques for further efficiency gains.

Introduction

The introduction of the paper discusses the significant advancements made by Pre-trained Language Models (PLMs) in Natural Language Processing (NLP), emphasizing the importance of both pre-training and fine-tuning stages. Pre-training involves self-supervised learning on extensive datasets to develop general language representations, while fine-tuning adapts these models to specific tasks through supervised learning. Despite their effectiveness, PLMs require fine-tuning with local data, which raises privacy concerns. Federated Learning (FL) is proposed as a solution to maintain data privacy by keeping raw data on local devices, but the large scale of PLMs presents challenges in computation and communication costs.

To address these challenges, the paper introduces Federated Parameter-Efficient Fine-Tuning (FedPEFT), which combines FL with Parameter-Efficient Fine-Tuning (PEFT) methods like Low-Rank Adaptation (LoRA). However, FedPEFT faces issues related to non-independent and identically distributed (non-IID) data, which can degrade model performance, and the inefficiency of fixed-rank configurations that do not adapt to varying data sensitivities. The authors propose a novel approach, Adaptive Rank Allocation for Federated PEFT (FedARA), which incorporates three innovative methods: Truncated SVD adaptation, dynamic rank allocation, and rank-based module pruning. These methods aim to enhance performance and communication efficiency while minimizing resource costs, demonstrating significant improvements in accuracy and efficiency on resource-constrained devices.

Methods

In this section, the authors describe the methods employed in their research, focusing on the FedARA framework, which extends the AdaLoRA approach for federated learning (FL). FedARA, built on the Hugging Face Transformers library, facilitates dynamic parameter configuration in a federated context and supports various pre-trained natural language processing (NLP) models. The framework utilizes a random client selection mechanism, which can be enhanced with more sophisticated strategies, and synchronizes communication between clients and the server through weighted averaging. Additionally, FedARA implements model quantization techniques, such as QLoRA and QA-LoRA, to minimize computational and memory demands, thereby potentially reducing communication costs.

The evaluation metrics for the study include final accuracy for classification tasks, ROUGE scores for summarization tasks, and total communication overhead after a predetermined number of FL rounds, ensuring a fair comparison of performance. Secondary metrics encompass per-round communication costs, adaptive rank allocation, total training time, memory usage, and energy consumption on edge devices, providing a thorough assessment of system efficiency and effectiveness. The experiments are conducted via emulation on a laptop with an RTX 4070 GPU, with local training times measured on three representative edge devices: AGX Orin, Orin Nano, and Raspberry Pi 5. The authors assume a communication bandwidth of 1MB/s between the server and clients, reflecting realistic conditions typical of IoT, home WiFi, and cellular networks.

Discussion

In the discussion section of the research paper, several key observations regarding the Federated Parameter Efficient Fine-Tuning (FedPEFT) framework are highlighted. Firstly, FedPEFT outperforms Federated Full Fine-Tuning (FedFFT) on advanced edge devices, achieving comparable accuracy while significantly reducing local training time (1.6×), communication costs (200×), and memory usage (1.5×). However, the overall training time on advanced devices remains inadequately explored. Notably, the performance of FedPEFT diminishes under non-IID data conditions, with average accuracy reductions of 14.62% and 12.81% on the 20News and News Category datasets, respectively. This indicates a critical need for strategies to address the adverse effects of non-IID data.

Furthermore, the positioning of Parameter Efficient Fine-Tuning (PEFT) modules within the FedPEFT framework affects performance, as fixed parameter configurations lead to inefficient communication by repeatedly transmitting less impactful parameters. The study reveals that communication becomes the primary bottleneck on high-end devices, while computation is the limiting factor on lower-end devices like Raspberry Pi. These findings suggest that future research should focus on optimizing FedPEFT for various edge devices and addressing the unique challenges posed by non-IID data. The introduction of FedARA, an adaptive rank allocation framework, aims to enhance communication efficiency and mitigate the limitations of existing methods by incorporating a truncated SVD structure and dynamic rank allocation.

شارك: