التخفيف من التحيز في الذكاء الاصطناعي: توليد بيانات عادلة عبر نماذج سببية لصنع قرارات شفافة وقابلة للتفسير
Mitigating bias in artificial intelligence: Fair data generation via causal models for transparent and explainable decision-making

المجلة: Future Generation Computer Systems، المجلد: 155
DOI: https://doi.org/10.1016/j.future.2024.02.023
تاريخ النشر: 2024-02-24
المؤلف: Rubén González-Sendino وآخرون
الموضوع الرئيسي: الذكاء الاصطناعي القابل للتفسير (XAI)

نظرة عامة

تتناول هذه الورقة البحثية القضايا الحرجة المتعلقة بالتحيز وقابلية التفسير في نماذج الذكاء الاصطناعي (AI). تصنف استراتيجيات التخفيف من التحيز الموجودة إلى نهج ما قبل التدريب، التدريب، وما بعد التدريب، وتقدم تقنية جديدة لإنشاء مجموعة بيانات مخففة من التحيز من خلال نموذج سببي مخفف ضمن شبكة بايزي. تشمل المساهمات الرئيسية خوارزمية تدريب جديدة للتخفيف من التحيز للنماذج السببية، ومنهجية ما قبل التدريب لتوليد مجموعات بيانات عادلة، والحفاظ على الميزات الحساسة أثناء التحليل. تؤكد الدراسة على أهمية قابلية التفسير والشفافية في الذكاء الاصطناعي، مقدمة عرضًا تفاعليًا للنتائج وتسهيل النسخ.

تشير النتائج إلى أن استخدام النماذج السببية المخففة يمكن أن يؤدي إلى خوارزميات غير تمييزية تعالج المجموعات المميزة وغير المميزة بشكل مشابه، مع التركيز على الميزات ذات الصلة بدلاً من سمات المجموعة. تسمح خوارزمية التخفيف المقترحة بتخفيف التحيز بشكل متزامن عبر ميزات حساسة متعددة، بينما يمكن استخدام البيانات العادلة الناتجة لتدريب خوارزميات خالية من التمييز. تسلط الورقة الضوء على الحاجة إلى تحسين قابلية التفسير في خوارزميات التخفيف وتقترح العمل المستقبلي في توسيع هذه الطرق لتشمل البيانات غير المهيكلة واستكشاف خرائط أكثر تعقيدًا للميزات الحساسة. تم الإشارة إلى قيد وهو تحدي الكشف التلقائي عن الميزات الحساسة، والذي لا يزال مجالًا مهمًا للبحث المستقبلي.

مقدمة

تسلط مقدمة الورقة البحثية الضوء على الإمكانية المزدوجة للذكاء الاصطناعي (AI) لتحويل المجتمع بشكل إيجابي أو perpetuate existing inequities if not designed with fairness in mind. تؤكد على دور التحيز المعرفي في توليد البيانات ونشر النماذج، مما يمكن أن يؤدي إلى التمييز ضد المجموعات المحرومة. يجادل المؤلفون بأن اختيار خصائص البيانات بناءً على الارتباط، بدلاً من السببية، يزيد من تفاقم هذه القضايا. يقترحون أن تحقيق العدالة في الذكاء الاصطناعي يتطلب خوارزميات تلبي معايير العدالة والأداء، مع استراتيجيات تخفيف التحيز القابلة للتطبيق في مراحل مختلفة من تطوير النموذج.

تهدف الدراسة إلى تطوير نموذج خالٍ من التمييز يعمل بشكل مستقل عن الخصائص الشخصية الحساسة، داعيةً إلى استخدام مجموعات بيانات عادلة وغير متحيزة في التدريب. يحدد المؤلفون قيودين رئيسيين في الأدبيات الحالية المتعلقة بتخفيف التحيز: تحدي معالجة متغيرات متحيزة متعددة في وقت واحد والحاجة إلى فهم شامل للميزات الحساسة طوال عملية النمذجة. يحافظ نهجهم المقترح على الميزات الحساسة، مما يعزز قدرات التدقيق ويسهل إنشاء مجموعات بيانات عادلة. توضح الورقة هيكلها، مشيرةً إلى أن الأقسام التالية ستتناول الأعمال ذات الصلة، والمفاهيم الأساسية، والخوارزمية المقترحة لتخفيف التحيز، والنتائج التجريبية، والاتجاهات المستقبلية. بالإضافة إلى ذلك، تتوفر موارد لإعادة إنتاج الدراسة من خلال مكتبة التعلم العادل للذكاء الاصطناعي (FLAI).

النتائج

تحدد قسم النتائج في الورقة البحثية تقييم نماذج مختلفة فيما يتعلق بمعايير الدقة والعدالة، مع التركيز على ثلاث تحليلات رئيسية. التحليل الأول يقيم أداء نموذج سببي (CM) بدون تخفيف ونموذج سببي مخفف (CMM) بعد تطبيق تقنيات التخفيف المحددة. التحليل الثاني يستخدم قيم SHAP (SHapley Additive exPlanations) لتقييم أهمية الميزات الحساسة. التحليل الثالث يقارن نتائج النماذج مع مقاييس ما قبل التدريب الأخرى لتقييم عدالة البيانات الناتجة.

تشمل النماذج التي تم تقييمها XGBoost (XGB)، المدربة على مجموعة البيانات الأصلية، وXGBoost المخفف (XGBM)، المدرب على مجموعة بيانات عادلة. تشمل مقاييس التقييم للعدالة فرق الفرص المتساوية (EOD)، التأثير المتباين (DI)، فرق التوازن الإحصائي (SPD)، وفارق الاحتمالات (OD)، بينما تشمل مقاييس الأداء الدقة، معدل الإيجابيات الحقيقية (TPR)، معدل الإيجابيات الزائفة (FPR)، معدل السلبيات الزائفة (FNR)، والتنبؤ كإيجابي (PPP). تُلاحظ مجموعات البيانات المستخدمة – Compas وGerman وAdult – لوجود تحيزات متأصلة، مع تحديد الميزات الحساسة على أنها الجنس والعرق. تصنف مجموعة بيانات Adult الأفراد بناءً على عتبة راتب قدرها 50 ألف دولار سنويًا. يتم تقديم نتائج إضافية لمجموعات بيانات أخرى في الملحق ب.

المناقشة

تحدد قسم المناقشة في الورقة البحثية منهجيات مختلفة تهدف إلى تخفيف التحيز في نماذج التعلم الآلي، مصنفة إلى مراحل ما قبل التدريب، التدريب، وما بعد التدريب. في مرحلة ما قبل التدريب، يتم استخدام تقنيات مثل إعادة الوزن، وإعادة أخذ العينات، والمعالجة المسبقة المحسّنة لتعزيز مجموعات البيانات وتقليل التحيزات المتأصلة. تعدل إعادة الوزن أوزان الحالات بناءً على امتياز المجموعة، بينما تعدل إعادة أخذ العينات أحجام مجموعات البيانات لتحقيق توازن في التوزيعات. يتم تسليط الضوء أيضًا على التعلم العادل للتمثيل (LFR) والمعالجة المسبقة المحسّنة (OPP) لفعاليتها في الحفاظ على العدالة أثناء إعداد البيانات.

خلال مرحلة التدريب، يتم استخدام طرق مثل التنظيم والتدريب المعاكس لتقليل التمييز الذي قد تتعلمه النماذج من البيانات المتحيزة. تساعد تقنيات التنظيم في معالجة الإفراط في التكيف وتقليل الفجوات في التنبؤ بين المجموعات، بينما يتضمن إزالة التحيز المعاكس تدريب شبكتين مزدوجتين لتحديد وإزالة التحيزات. تشمل مرحلة ما بعد التدريب خوارزميات مثل فرص متساوية وتصنيف خيار الرفض، والتي تعدل مخرجات النموذج لضمان العدالة عبر مجموعات ديموغرافية مختلفة. يشير المؤلفون إلى أن التقنيات الحالية تركز بشكل أساسي على متغيرات حساسة واحدة، مما يحد من فعاليتها في مجموعات البيانات التي تحتوي على سمات حساسة متعددة. بشكل عام، تؤكد القسم على أهمية تحقيق توازن بين العدالة، وتقليل التمييز، ودقة النموذج لتطوير أنظمة تعلم آلي عادلة.

Journal: Future Generation Computer Systems, Volume: 155
DOI: https://doi.org/10.1016/j.future.2024.02.023
Publication Date: 2024-02-24
Author(s): Rubén González-Sendino et al.
Primary Topic: Explainable Artificial Intelligence (XAI)

Overview

This research paper addresses the critical issues of bias and explainability in Artificial Intelligence (AI) models. It categorizes existing bias mitigation strategies into pre-training, training, and post-training approaches, and introduces a novel technique for creating a mitigated bias dataset through a mitigated causal model within a Bayesian network. Key contributions include a new mitigation training algorithm for causal models, a pretraining methodology for generating fair datasets, and the maintenance of sensitive features during analysis. The study emphasizes the importance of explainability and transparency in AI, providing an interactive demonstration of the findings and facilitating replication.

The findings suggest that employing mitigated causal models can lead to nondiscriminatory algorithms that treat both privileged and underprivileged groups similarly, focusing on relevant features rather than group attributes. The proposed mitigation algorithm allows for simultaneous bias mitigation across multiple sensitive features, while the fair data generated can be used to train discrimination-free algorithms. The paper highlights the need for improved explainability in mitigation algorithms and suggests future work in extending these methods to unstructured data and exploring more complex mappings of sensitive features. A limitation noted is the challenge of automatically detecting sensitive features, which remains a significant area for future research.

Introduction

The introduction of the research paper highlights the dual potential of artificial intelligence (AI) to either transform society positively or perpetuate existing inequities if not designed with fairness in mind. It emphasizes the role of cognitive bias in data generation and model deployment, which can lead to discrimination against underprivileged groups. The authors argue that the selection of data characteristics based on correlation, rather than causation, exacerbates these issues. They propose that achieving fairness in AI requires algorithms that meet both fairness and performance metrics, with bias mitigation strategies applicable at various stages of model development.

The study aims to develop a discrimination-free model that operates independently of sensitive personal characteristics, advocating for the use of fair and unbiased datasets in training. The authors identify two main limitations in existing literature regarding bias mitigation: the challenge of addressing multiple biased variables simultaneously and the need for a comprehensive understanding of sensitive features throughout the modeling process. Their proposed approach preserves sensitive features, enhancing auditing capabilities and facilitating the creation of fair datasets. The paper outlines its structure, indicating subsequent sections will cover related work, background concepts, the proposed algorithm for bias mitigation, experimental results, and future directions. Additionally, resources for reproducing the study are made available through the Fairness Learning Artificial Intelligence (FLAI) library.

Results

The Results section of the research paper outlines the evaluation of various models concerning accuracy and fairness metrics, focusing on three primary analyses. The first analysis assesses the performance of a Causal Model (CM) without mitigation and a Causal Model Mitigated (CMM) after applying specified mitigation techniques. The second analysis employs SHAP (SHapley Additive exPlanations) values to evaluate the importance of sensitive features. The third analysis compares the results of the models with other pre-training metrics to assess the fairness of the generated data.

The models evaluated include XGBoost (XGB), trained on the original dataset, and XGBoost Mitigated (XGBM), trained on a fair dataset. The evaluation metrics for fairness include Equal Opportunity Difference (EOD), Disparate Impact (DI), Statistical Parity Difference (SPD), and Odds Difference (OD), while performance metrics encompass accuracy, True Positive Rate (TPR), False Positive Rate (FPR), False Negative Rate (FNR), and Predicted as Positive (PPP). The datasets utilized—Compas, German, and Adult—are noted for their inherent biases, with sensitive features identified as sex and race. The Adult dataset specifically categorizes individuals based on a salary threshold of $50K/year. Additional results for other datasets are provided in Appendix B.

Discussion

The discussion section of the research paper outlines various methodologies aimed at mitigating bias in machine learning models, categorized into pre-training, training, and post-training phases. In the pre-training phase, techniques such as re-weighting, resampling, and optimized pre-processing are employed to enhance datasets and reduce inherent biases. Re-weighting adjusts the weights of instances based on group privilege, while resampling modifies dataset sizes to balance distributions. Learning Fair Representation (LFR) and Optimized Pre-Processing (OPP) are also highlighted for their effectiveness in maintaining fairness during data preparation.

During the training phase, methods like regularization and adversarial training are utilized to minimize discrimination that models may learn from biased data. Regularization techniques help address overfitting and reduce prediction disparities among groups, while adversarial de-biasing involves training dual networks to identify and eliminate biases. The post-training phase includes algorithms such as Equalized Odds and Reject Option Classification, which adjust model outputs to ensure fairness across different demographic groups. The authors note that existing techniques primarily focus on single sensitive variables, which limits their effectiveness in datasets with multiple sensitive attributes. Overall, the section emphasizes the importance of balancing fairness, discrimination reduction, and model accuracy to develop equitable machine learning systems.