XMutant: الفحص القائم على XAI لأنظمة التعلم العميق
XMutant: XAI-based fuzzing for deep learning systems

شارك:
المجلة: Empirical Software Engineering، المجلد: 31، العدد: 4
DOI: https://doi.org/10.1007/s10664-025-10792-1
تاريخ النشر: 2026-03-27
المؤلف: Xingcheng Chen وآخرون
الموضوع الرئيسي: تطبيقات الشبكات العصبية المتقدمة

نظرة عامة

يقدم هذا القسم نظرة عامة على XMutant، وهي تقنية جديدة لتوليد مدخلات تسبب الفشل لأنظمة التعلم العميق (DL)، مما يعزز كفاءة مولدات الاختبار المعتمدة على الدلالات. تعتمد الطرق التقليدية على الاضطرابات العشوائية لمفاهيم المدخلات الدلالية، مما يؤدي غالبًا إلى عدم الكفاءة في تحديد الفشل. على النقيض من ذلك، يستخدم XMutant تقنيات الذكاء الاصطناعي القابل للتفسير (XAI) للاستفادة من الشروحات المحلية للمدخلات، مما يوجه عملية اختبار الفوضى بشكل أكثر فعالية نحو الفشل المحتمل.

أظهر تقييم XMutant عبر تكوينات مختلفة أداؤه المتفوق في تحفيز الفشل في كل من اختبار مستوى النموذج (مثل تحليل المشاعر، التعرف على الأرقام) واختبار مستوى النظام (مثل المساعدة المتقدمة في القيادة). تشير النتائج إلى أن XMutant يمكنه توليد ما يصل إلى 125% من المدخلات المسببة للفشل مقارنة بالطرق الحالية، محققًا ذلك بسرعة تصل إلى سبع مرات. بالإضافة إلى ذلك، تحافظ التقنية على معدل تحقق يتجاوز 89% للمدخلات المولدة، كما أكدت ذلك كل من المراجعين الآليين والبشريين، مما يبرز موثوقيتها وفعاليتها في اختبار أنظمة DL.

مقدمة

تؤكد مقدمة هذه الورقة البحثية على أهمية أنظمة التعلم العميق (DL) في هندسة البرمجيات، لا سيما في المهام المعقدة، مع تسليط الضوء على التحديات الفريدة التي تقدمها للاختبار. على عكس البرمجيات التقليدية، فإن سلوك أنظمة DL أقل توقعًا، مما يتطلب توليد اختبارات فعالة لضمان الموثوقية والصحة. يتضمن ذلك إنشاء بيانات مدخلات متنوعة لتحديد نقاط الضعف المحتملة، خاصة في التطبيقات الحرجة مثل القيادة الذاتية. تناقش الورقة الأساليب الحالية لتوليد حالات الاختبار، مميزة بين تقنيات معالجة المدخلات الخام، التي تركز على التعديلات الطفيفة على المدخلات الحالية، وتوليد الاختبارات الوظيفية، التي تهدف إلى إنشاء مدخلات جديدة تتحدى قدرات تعميم النموذج.

يقترح المؤلفون تقنية اختبار جديدة تُسمى XMutant، والتي تستخدم الذكاء الاصطناعي القابل للتفسير (XAI) لتعزيز توليد الاختبارات المعتمدة على الدلالات. من خلال الاستفادة من الشروحات المحلية لتوقعات نموذج DL، يحدد XMutant المجالات الرئيسية للمدخلات التي تؤثر بشكل كبير على اتخاذ القرار. تتيح هذه الطريقة المستهدفة مزيدًا من الفعالية في تغيير المدخلات، مما يؤدي إلى عدد أكبر من حالات الفشل المسببة وزيادة الكفاءة مقارنة بالطرق الحالية مثل DeepJanus. يُظهر تقييم XMutant عبر أنظمة DL المختلفة قدرته على إنتاج مدخلات صالحة، ضمن التوزيع، تحافظ على معدلات تحقق عالية وحفظ التسمية، مما يعزز مجال اختبار DL من خلال دمج تقنيات XAI.

طرق

في هذا القسم، يناقش المؤلفون مجموعة من طرق الذكاء الاصطناعي القابل للتفسير (XAI) المستخدمة لتعزيز الشفافية وقابلية التفسير لأنظمة التعلم العميق (DL) المعقدة، وهو أمر حاسم للتحقق والثقة والامتثال التنظيمي. يركزون على تطبيقات محددة مثل تحليل المشاعر، التعرف على الأرقام، وأنظمة المساعدة المتقدمة في القيادة (ADAS)، مشيرين إلى الأعمال المعتمدة في اختبار DL. بالنسبة لتحليل المشاعر، يقيم المؤلفون طرق الشرح المحلية بما في ذلك LIME وSmoothGrad وIntegrated Gradients. يُلاحظ أن LIME، على الرغم من فعاليته، يتطلب تكلفة حسابية عالية بسبب الحاجة إلى العديد من الاضطرابات والاستنتاجات. لمعالجة ذلك، يدمج المؤلفون SmoothGrad وIntegrated Gradients، اللذان يخففان من المشكلات المتعلقة بانقطاع التدرجات والتشبع.

بالنسبة للمهام التصويرية مثل التعرف على الأرقام المكتوبة بخط اليد وADAS، يتجه المؤلفون إلى طرق التأثير ونسب البكسل، التي توفر شروحات محلية من خلال تسليط الضوء على المناطق المؤثرة في الصور المدخلة. يستخدمون SmoothGrad وIntegrated Gradients جنبًا إلى جنب مع تقنيات تعتمد على خريطة تنشيط الفئة (CAM)، وخاصة Grad-CAM++، التي تولد خرائط حرارية من خلال دمج المشتقات الجزئية الإيجابية من آخر طبقة تلافيفية من الشبكات العصبية التلافيفية. تهدف هذه الطريقة الشاملة إلى ضمان كفاءة وفعالية توليد الاختبارات في سياق XAI.

النتائج

يقدم قسم “النتائج” في الورقة البحثية النتائج المستمدة من التجارب أو التحليلات التي تم إجراؤها. يوضح النتائج التي توصلت إليها الدراسة، مع تسليط الضوء على نقاط البيانات والاتجاهات المهمة التي لوحظت خلال التحقيق. عادةً ما تكون النتائج مصحوبة بتحليلات إحصائية ذات صلة، والتي قد تشمل قيم p، فترات الثقة، أو مقاييس أخرى تدعم صحة النتائج.

غالبًا ما يتم توضيح النتائج الرئيسية من خلال الجداول أو الرسوم البيانية أو الأشكال، مما يوفر تمثيلًا بصريًا للبيانات. تساعد هذه الوسائل البصرية في توضيح العلاقات أو الأنماط المعقدة التي ظهرت من التحليل. قد يناقش القسم أيضًا تداعيات النتائج فيما يتعلق بالفرضيات الأصلية أو أسئلة البحث المطروحة في بداية الدراسة، مع التأكيد على كيفية مساهمة النتائج في المعرفة الحالية في هذا المجال.

المناقشة

في هذا القسم، تناقش البحث تنفيذ وتقييم XMutant، وهي أداة جديدة للاختبار المركّز لأنظمة التعلم العميق (DL) التي تستخدم تقنيات الذكاء الاصطناعي القابل للتفسير (XAI) لتوليد الاختبارات المعتمدة على الدلالات. يُعتبر XMutant الأول من نوعه الذي يستخدم XAI لتوليد الاختبارات على كل من مستوى النموذج ومستوى النظام، مما يظهر فعالية وكفاءة وصحة ومعدلات حفظ التسمية تفوق الطرق الحالية الرائدة. تعمل الأداة من خلال الاستفادة من التمثيلات الدلالية للمدخلات، التي تجرد البيانات إلى مفاهيم ذات معنى، مما يقلل من مساحة البحث لتوليد الاختبارات ويضمن بقاء المدخلات المولدة ضمن النطاق التشغيلي الصالح لأنظمة DL.

تتوسع المناقشة في منهجية XMutant، التي تتضمن تقنية اختبار مركّز تستخدم الشروحات المحلية لتوقعات النموذج لتوجيه تغييرات المدخلات. تختار الخوارزمية مفاهيم دلالية للتغيير بناءً على المناطق ذات الانتباه العالي التي تشير إليها هذه الشروحات، مما يسمح باضطرابات مستهدفة ذات معنى دلالي. تسلط الورقة أيضًا الضوء على الاختلافات في سيناريوهات الاختبار عبر تطبيقات DL المختلفة، مثل تحليل المشاعر، التعرف على الأرقام، وأنظمة المساعدة المتقدمة في القيادة (ADAS)، كل منها يتطلب تمثيلات دلالية ومعايير فشل مميزة. تهدف الدراسة التجريبية الموضحة في هذا القسم إلى الإجابة على عدة أسئلة بحثية تتعلق بفعالية وكفاءة وصحة XMutant، بالإضافة إلى أدائه المقارن ضد تقنيات معالجة المدخلات التقليدية المعتمدة على التدرجات.

Journal: Empirical Software Engineering, Volume: 31, Issue: 4
DOI: https://doi.org/10.1007/s10664-025-10792-1
Publication Date: 2026-03-27
Author(s): Xingcheng Chen et al.
Primary Topic: Advanced Neural Network Applications

Overview

The section presents an overview of XMutant, a novel technique for generating failure-inducing inputs for Deep Learning (DL) systems, which enhances the efficiency of semantic-based test generators. Traditional methods rely on random perturbations of input semantic concepts, often leading to inefficiencies in identifying failures. In contrast, XMutant utilizes explainable artificial intelligence (XAI) techniques to leverage local explanations of inputs, guiding the fuzz testing process more effectively toward potential failures.

The evaluation of XMutant across various configurations demonstrated its superior performance in triggering failures in both model-level (e.g., sentiment analysis, digit recognition) and system-level testing (e.g., advanced driving assistance). The results indicate that XMutant can generate up to 125% more failure-inducing inputs compared to existing methods, achieving this up to seven times faster. Additionally, the technique maintains a validation rate exceeding 89% for the generated inputs, as confirmed by both automated and human validators, underscoring its reliability and effectiveness in testing DL systems.

Introduction

The introduction of this research paper emphasizes the significance of Deep Learning (DL) systems in software engineering, particularly in complex tasks, while highlighting the unique challenges they present for testing. Unlike traditional software, the behavior of DL systems is less predictable, necessitating effective test generation to ensure reliability and correctness. This involves creating diverse input data to identify potential weaknesses, especially in critical applications like autonomous driving. The paper discusses existing approaches for test case generation, distinguishing between raw input manipulation techniques, which focus on minor alterations to existing inputs, and functional test generation, which aims to create novel inputs that challenge the model’s generalization capabilities.

The authors propose a novel testing technique called XMutant, which utilizes eXplainable Artificial Intelligence (XAI) to enhance semantic-based test generation. By leveraging local explanations of DL model predictions, XMutant identifies key areas of input that significantly influence decision-making. This targeted approach allows for more effective mutation of inputs, resulting in a higher number of failure-inducing cases and improved efficiency compared to existing methods like DeepJanus. The evaluation of XMutant across various DL systems demonstrates its ability to produce valid, in-distribution inputs that maintain high validity and label preservation rates, thereby advancing the field of DL testing through the integration of XAI techniques.

Methods

In this section, the authors discuss various Explainable Artificial Intelligence (XAI) methods employed to enhance the transparency and interpretability of complex deep learning (DL) systems, which is crucial for validation, trust, and regulatory compliance. They focus on specific applications such as sentiment analysis, digit recognition, and advanced driving assistance systems (ADAS), referencing established work in DL testing. For sentiment analysis, the authors evaluate local explanation methods including LIME, SmoothGrad, and Integrated Gradients. LIME, while effective, is noted for its computational expense due to the need for numerous perturbations and inferences. To address this, the authors incorporate SmoothGrad and Integrated Gradients, which mitigate issues related to gradient discontinuity and saturation.

For imagery tasks like handwritten digit recognition and ADAS, the authors turn to saliency and pixel attribution methods, which provide local explanations by highlighting influential regions in input images. They employ SmoothGrad and Integrated Gradients alongside Class Activation Map (CAM)-based techniques, particularly Grad-CAM++, which generates heatmaps by combining positive partial derivatives from the last convolutional layer of convolutional neural networks. This comprehensive approach aims to ensure both the efficiency and effectiveness of test generation in the context of XAI.

Results

The “Results” section of the research paper presents the findings derived from the conducted experiments or analyses. It details the outcomes of the study, highlighting significant data points and trends observed during the investigation. The results are typically accompanied by relevant statistical analyses, which may include p-values, confidence intervals, or other metrics that support the validity of the findings.

Key findings are often illustrated through tables, graphs, or figures, providing a visual representation of the data. These visual aids help to clarify complex relationships or patterns that emerged from the analysis. The section may also discuss the implications of the results in relation to the original hypotheses or research questions posed at the outset of the study, emphasizing how the findings contribute to the existing body of knowledge in the field.

Discussion

In this section, the research discusses the implementation and evaluation of XMutant, a novel tool for focused fuzzing of deep learning (DL) systems that employs explainable artificial intelligence (XAI) techniques for semantic-based test generation. XMutant is the first of its kind to utilize XAI for generating tests at both the model and system levels, demonstrating superior effectiveness, efficiency, validity, and label-preservation rates compared to existing state-of-the-art approaches. The tool operates by leveraging semantic representations of inputs, which abstract the data into meaningful concepts, thereby reducing the search space for test generation and ensuring that generated inputs remain within the valid operational domain of the DL systems.

The discussion further elaborates on the methodology of XMutant, which includes a focused fuzzing technique that utilizes local explanations of model predictions to guide input mutations. The algorithm selects semantic concepts for mutation based on areas of high attention indicated by these explanations, allowing for targeted perturbations that are semantically meaningful. The paper also highlights the differences in testing scenarios across various DL applications, such as sentiment analysis, digit recognition, and advanced driving assistance systems (ADAS), each requiring distinct semantic representations and failure criteria. The empirical study outlined in this section aims to answer several research questions regarding the effectiveness, efficiency, and validity of XMutant, as well as its comparative performance against traditional gradient-guided input manipulation techniques.

شارك: