هل نحن ننسى حقًا؟ إعادة فحص نقدية لبروتوكولات تقييم نسيان الآلة
Are we truly forgetting? A critical re-examination of machine unlearning evaluation protocols

شارك:
المجلة: Engineering Applications of Artificial Intelligence، المجلد: 167
DOI: https://doi.org/10.1016/j.engappai.2026.113785
تاريخ النشر: 2026-01-14
المؤلف: YongWoo Kim وآخرون
الموضوع الرئيسي: الصلابة ضد الهجمات في تعلم الآلة

نظرة عامة

تناقش هذه الفقرة مفهوم “إلغاء التعلم الآلي”، الذي يهدف إلى إزالة نقاط بيانات معينة من نموذج مدرب مع الحفاظ على أدائه على البيانات المحتفظ بها، مما يعالج مخاوف الخصوصية والقانون. ينتقد المؤلفون التقييمات الحالية لطرق الإلغاء، التي تعتمد في الغالب على مقاييس قائمة على اللوجيت في سياقات صغيرة النطاق، مما يشير إلى أن هذه التقييمات قد توفر شعورًا مضللًا بالأمان بشأن فعالية الإلغاء في التطبيقات الواقعية.

لمعالجة هذه الفجوة، يقدم البحث إطار تقييم شامل يستخدم مقاييس قائمة على التمثيل لتقييم أداء النموذج الذي تم إلغاء تعلمه في سيناريوهات واسعة النطاق. تشير النتائج إلى أن طرق الإلغاء الحديثة غالبًا ما تضر بالجودة التمثيلية للنموذج أو تغير فقط المصنف، مما يؤدي إلى تحسين الأداء القائم على اللوجيت دون تغييرات كبيرة في التمثيل الأساسي للنموذج. بالإضافة إلى ذلك، يقترح المؤلفون سيناريو تقييم جديد حيث تكون الفئات التي يجب نسيانها مشابهة دلاليًا لفئات المهام اللاحقة، مما يتطلب تباينًا أكبر في تمثيلات الميزات عن النموذج الأصلي. يهدف هذا المعيار إلى إنشاء بروتوكول موحد لتقييم خوارزميات الإلغاء في ظل ظروف أكثر واقعية.

مقدمة

تسلط مقدمة البحث الضوء على التأثير التحويلي للشبكات العصبية عبر مجالات مختلفة، لا سيما في تصنيف الصور، مدفوعة بالتقدم في قدرات النماذج من خلال تدريب مجموعات بيانات واسعة. ومع ذلك، يثير هذا التقدم مخاوف كبيرة بشأن الخصوصية، خاصة فيما يتعلق باستخدام البيانات الحساسة، مما يدفع إلى ظهور “إلغاء التعلم الآلي” كمنطقة بحث حاسمة. يهدف إلغاء التعلم الآلي إلى إزالة تأثير بيانات معينة من النماذج المدربة مسبقًا بشكل فعال مع الاحتفاظ بالمعرفة من بيانات أخرى. الطريقة التقليدية للإلغاء الدقيق، التي تتضمن إعادة تدريب النماذج من الصفر، مكلفة حسابيًا، مما أدى إلى تطوير خوارزميات إلغاء تقريبية تسعى لتحقيق الكفاءة مع الحفاظ على أداء النموذج.

على الرغم من التقدم الأخير، كانت تقييمات طرق الإلغاء محدودة بالمهام الصغيرة النطاق والمقاييس القائمة على اللوجيت، مما يثير القلق بشأن قابليتها للتوسع إلى مجموعات بيانات أكبر وتطبيقات العالم الحقيقي. يقترح البحث إطار تقييم شامل يعالج هذه القيود من خلال دمج تقييمات النماذج التي تم إلغاء تعلمها بناءً على تمثيلات الميزات في إعدادات واسعة النطاق، باستخدام مجموعة بيانات ImageNet-1K بشكل خاص. لا يحدد هذا الإطار بروتوكول تقييم صارم لتصنيف الصور فحسب، بل يقترح أيضًا امتدادات محتملة لمجالات أخرى، مثل الشبكات العصبية الرسومية (GNNs)، مما يبرز الحاجة إلى فهم شامل للمقايضات بين فعالية الإلغاء والضمانات الرسمية للخصوصية مثل الخصوصية التفاضلية (DP). يقدم المؤلفون معيارًا موحدًا يكشف عن أوجه القصور في الطرق الحالية ويضع الأساس لبحث أكثر موثوقية وقابلية للتوسع في إلغاء التعلم الآلي.

طرق

في قسم الطرق، يقدم البحث تقييمًا شاملاً لخوارزميات الإلغاء على نطاق واسع، مع تسليط الضوء على أداء طريقة “تسمية زائفة” (PL). تشير النتائج، الملخصة في الجدول 2، إلى أن PL، التي لا تستخدم المورد الإضافي $D_r$، تحقق أداءً يعادل أو يتفوق على الخوارزميات الحديثة مثل DUCK و CU و SCAR و SCRUB، التي تعتمد على $D_r$. تتناقض هذه النتيجة مع الدراسات السابقة للإلغاء على نطاق صغير، مما يشير إلى أن العديد من الطرق المعتمدة قد تتدهور فعاليتها عند تطبيقها على مجموعات بيانات أكبر. علاوة على ذلك، يوضح الجدول 3 أن PL تتفوق باستمرار على الخوارزميات الأخرى في التقييمات القائمة على التمثيل، محققة أعلى درجات H-LR عبر جميع المعايير الأساسية.

تدعم موثوقية النتائج المزيد من التجارب باستخدام هياكل عظمية مختلفة، بما في ذلك ConvNeXt-T و Swin-T، كما هو موضح في الجدول B.7. تظهر النتائج أن اتجاهات الأداء الملاحظة مع ResNet-50 تُحافظ عبر هياكل مختلفة، مما يعزز موثوقية إطار التقييم المقترح. من الجدير بالذكر أن الدراسة تحدد مشكلات كبيرة في قابلية التوسع مع بعض الطرق، مثل SCRUB، التي تعاني من انخفاض حاد في الأداء مع زيادة عدد الفئات المنسية. وهذا يبرز قيود استراتيجيات التحسين المعقدة في مهام الإلغاء على نطاق واسع ويقترح أن معالجة هذه التحديات قد تتطلب تغييرات أساسية في أهداف التحسين أو هياكل النماذج.

مناقشة

في هذا القسم، يقوم المؤلفون بتقييم نقدي لطرق إلغاء التعلم الآلي الحالية، مع التركيز بشكل خاص على فعاليتها في محو المعلومات المتعلمة على مستوى التمثيل. يبرزون قيودًا كبيرة في أطر التقييم الحالية، التي تعتمد في الغالب على مقاييس قائمة على اللوجيت، مثل دقة النسيان ودقة الاحتفاظ، التي قد لا تعكس التغييرات الداخلية في تمثيلات النماذج بشكل كافٍ. من خلال التجارب التي تتضمن إلغاء تعلم 100 فئة عشوائية في ImageNet-1K باستخدام نماذج مثل ResNet-50، يظهرون أنه بينما تشير التقييمات القائمة على اللوجيت إلى إلغاء تعلم ناجح، تكشف التقييمات القائمة على التمثيل باستخدام محاذاة النواة المركزية (CKA) أن النماذج التي تم إلغاء تعلمها تحتفظ بمزيد من التشابه مع النموذج الأصلي مقارنةً بنظرائها المعاد تدريبها، مما يشير إلى عدم كفاية محو المعلومات المتعلمة.

لمعالجة هذه القصور، يقترح المؤلفون نموذج تقييم جديد يسمى “نسيان الفئات العليا”، الذي يركز على التشابه الدلالي بين الفئات المستهدفة للإلغاء وتلك الموجودة في المهام اللاحقة. يهدف هذا النهج إلى تقييم فعالية خوارزميات الإلغاء بشكل صارم من خلال الحاجة إلى تباين كبير في تمثيلات الميزات. يدعو المؤلفون إلى معيار موحد يدمج كل من المقاييس القائمة على التمثيل وإعدادات التقييم على نطاق واسع، مما يوفر إطارًا أكثر شمولاً لتقييم طرق الإلغاء. تؤكد نتائجهم على ضرورة تحسين تقنيات التقييم التي يمكن أن تعكس بدقة تعقيدات الإلغاء في التطبيقات الواقعية.

القيود

تنبع قيود الإطار المقترح بشكل أساسي من تركيزه على مهام التصنيف، مما يقيد قابليته للتطبيق على النماذج التوليدية أو نماذج الصندوق الأسود. بينما توفر مقاييس مثل محاذاة النواة المركزية (CKA) و k-أقرب الجيران (k-NN) رؤى حول التحولات الدلالية، قد لا تغطي تمامًا التعقيدات المعنية. يجب أن تهدف الأبحاث المستقبلية إلى توسيع الإطار لاستيعاب الإلغاء الديناميكي ومجموعة متنوعة أوسع من أنواع المهام.

بالإضافة إلى ذلك، تم تصميم الإطار للإلغاء على مستوى الفئات، مما يقدم تحديات للإلغاء على مستوى العينات. من غير المحتمل أن يؤدي إزالة نقاط بيانات فردية إلى إحداث تغييرات كبيرة في تمثيل النموذج، مما يجعل مقاييس المستوى التوزيعي مثل CKA غير حساسة بشكل كافٍ. تكافح التقييمات القائمة على اللوجيت أيضًا مع هذه القيود، حيث إن دقة المستوى المهام لا تقيم بشكل فعال تأثير إزالة عينة واحدة. وبالتالي، يبقى تطوير نهج قوي للإلغاء على مستوى العينات مجالًا مهمًا للاستكشاف المستقبلي. علاوة على ذلك، فإن الاعتماد على النموذج المعاد تدريبه كمعيار ذهبي يطرح تحديات حسابية، مما يشير إلى أن تشابه CKA بين النماذج التي تم إلغاء تعلمها والنماذج الأصلية يمكن أن يكون بمثابة وكيل مفيد لتقدير فعالية الإلغاء. ومع ذلك، فإن إنشاء إطار تقييم موثوق لا يعتمد على نموذج أوراكل هو تحدٍ مستمر حاسم.

Journal: Engineering Applications of Artificial Intelligence, Volume: 167
DOI: https://doi.org/10.1016/j.engappai.2026.113785
Publication Date: 2026-01-14
Author(s): YongWoo Kim et al.
Primary Topic: Adversarial Robustness in Machine Learning

Overview

The section discusses the concept of machine unlearning, which aims to remove specific data points from a trained model while preserving its performance on retained data, thereby addressing privacy and legal concerns. The authors critique existing evaluations of unlearning methods, which predominantly rely on logit-based metrics in small-scale contexts, suggesting that such evaluations may provide a misleading sense of security regarding the effectiveness of unlearning in real-world applications.

To address this gap, the paper presents a comprehensive evaluation framework that utilizes representation-based metrics to assess the unlearned model’s performance in large-scale scenarios. The findings indicate that current state-of-the-art unlearning methods often compromise the representational quality of the model or only alter the classifier, resulting in improved logit-based performance without significant changes to the model’s underlying representation. Additionally, the authors propose a novel evaluation scenario where the classes to be forgotten are semantically similar to the downstream task classes, requiring a more substantial divergence in feature representations from the original model. This benchmark aims to establish a standardized protocol for evaluating unlearning algorithms under more realistic conditions.

Introduction

The introduction of the paper highlights the transformative impact of neural networks across various domains, particularly in image classification, driven by advancements in model capabilities through extensive dataset training. However, this progress raises significant privacy concerns, especially regarding the use of sensitive data, prompting the emergence of machine unlearning as a critical area of research. Machine unlearning aims to effectively remove the influence of specific data from pre-trained models while retaining knowledge from other data. The traditional method of exact unlearning, which involves retraining models from scratch, is computationally expensive, leading to the development of approximate unlearning algorithms that strive for efficiency while maintaining model performance.

Despite recent advancements, the evaluation of unlearning methods has been limited to small-scale tasks and logit-based metrics, raising concerns about their scalability to larger datasets and real-world applications. The paper proposes a holistic evaluation framework that addresses these limitations by incorporating assessments of unlearned models based on feature representations in large-scale settings, specifically using the ImageNet-1K dataset. This framework not only establishes a rigorous evaluation protocol for image classification but also suggests potential extensions to other domains, such as Graph Neural Networks (GNNs), emphasizing the need for a comprehensive understanding of the trade-offs between unlearning effectiveness and formal privacy guarantees like Differential Privacy (DP). The authors present a unified benchmark that reveals the shortcomings of existing methods and lays the groundwork for more reliable and scalable research in machine unlearning.

Methods

In the section on methods, the research presents a comprehensive evaluation of large-scale unlearning algorithms, highlighting the performance of the Pseudo Labeling (PL) method. The results, summarized in Table 2, indicate that PL, which does not utilize the additional resource $D_r$, achieves performance that is comparable to or superior to state-of-the-art algorithms such as DUCK, CU, SCAR, and SCRUB, which do rely on $D_r$. This finding contrasts with previous small-scale unlearning studies, suggesting that many established methods may degrade in effectiveness when applied to larger datasets. Furthermore, Table 3 illustrates that PL consistently outperforms other algorithms in representation-based evaluations, achieving the highest H-LR scores across all baselines.

The robustness of the findings is further supported by additional experiments using various backbone architectures, including ConvNeXt-T and Swin-T, as detailed in Table B.7. The results demonstrate that the performance trends observed with ResNet-50 are maintained across different architectures, reinforcing the reliability of the proposed evaluation framework. Notably, the study identifies significant scalability issues with certain methods, such as SCRUB, which experiences a drastic drop in performance as the number of forgetting classes increases. This underscores the limitations of complex optimization strategies in large-scale unlearning tasks and suggests that addressing these challenges may require fundamental changes to the optimization objectives or model architectures.

Discussion

In this section, the authors critically evaluate existing machine unlearning methods, particularly focusing on their effectiveness in erasing learned information at the representation level. They highlight significant limitations in current evaluation frameworks, which predominantly rely on logit-based metrics, such as Forget Accuracy and Retain Accuracy, that may not adequately capture the internal changes in model representations. Through experiments involving the unlearning of 100 random classes in ImageNet-1K using models like ResNet-50, they demonstrate that while logit-based evaluations indicate successful unlearning, representation-based assessments using Centered Kernel Alignment (CKA) reveal that unlearned models retain more similarity to the original model than to retrained counterparts, suggesting insufficient erasure of learned information.

To address these shortcomings, the authors propose a novel evaluation paradigm termed “Top Class-wise Forgetting,” which emphasizes semantic similarity between classes targeted for unlearning and those in downstream tasks. This approach aims to rigorously assess the effectiveness of unlearning algorithms by requiring a substantial divergence in feature representations. The authors advocate for a unified benchmark that incorporates both representation-based metrics and large-scale evaluation settings, thereby providing a more comprehensive framework for assessing unlearning methods. Their findings underscore the necessity for improved evaluation techniques that can accurately reflect the complexities of unlearning in real-world applications.

Limitations

The limitations of the proposed framework primarily stem from its focus on classification tasks, which restricts its applicability to generative or black-box models. While metrics such as Centered Kernel Alignment (CKA) and k-Nearest Neighbors (k-NN) provide insights into semantic shifts, they may not fully encapsulate the complexities involved. Future research should aim to expand the framework to accommodate dynamic unlearning and a wider variety of task types.

Additionally, the framework is tailored for class-wise unlearning, presenting challenges for sample-wise unlearning. The removal of individual data points is unlikely to induce significant changes in the model’s representation, rendering distribution-level metrics like CKA insufficiently sensitive. Logit-based evaluations similarly struggle with this limitation, as task-level accuracy does not effectively assess the impact of single-sample removal. Consequently, developing a robust approach for sample-wise unlearning remains an important area for future exploration. Furthermore, the reliance on the retrained model as a gold standard poses computational challenges, suggesting that CKA similarity between the unlearned and original models could serve as a useful proxy for estimating unlearning effectiveness. Nonetheless, creating a reliable evaluation framework that does not depend on an oracle model is a critical ongoing challenge.

شارك: