دراسة تجريبية حول التعلم الآلي المدرك للعدالة لمشاكل تقييم الائتمان
An experimental study on fairness-aware machine learning for credit scoring problems

شارك:
المجلة: Digital Finance، المجلد: 8، العدد: 3
DOI: https://doi.org/10.1007/s42521-026-00202-6
تاريخ النشر: 2026-06-16
المؤلف: Huyen Giang Thi Thu وآخرون
الموضوع الرئيسي: التنبؤ بالضغوط المالية والإفلاس

نظرة عامة

إن رقمنة تقييم الائتمان أصبحت ضرورية بشكل متزايد للمؤسسات المالية، خاصة مع استخدام تقنيات التعلم الآلي لتقييم جدارة العملاء الائتمانية. ومع ذلك، غالبًا ما تظهر هذه النماذج تحيزات تتعلق بالسمات المحمية مثل العرق والجنس. تقدم هذه الورقة دراسة تجريبية شاملة تقيم نماذج التعلم الآلي التي تأخذ العدالة في الاعتبار في سياق تقييم الائتمان، مع فحص مجموعات البيانات المالية، والنماذج التنبؤية، ومقاييس العدالة. تشير النتائج إلى أن النماذج التي تأخذ العدالة في الاعتبار تتفوق على نماذج التصنيف التقليدية من خلال تحقيق توازن أكثر ملاءمة بين الدقة التنبؤية والعدالة.

في الاستنتاجات، يبرز المؤلفون التحيزات الكامنة الموجودة في مجموعات بيانات تقييم الائتمان المستخدمة بشكل شائع، كما تم الكشف عنها من خلال تحليل الشبكات البايزية. يقارنون بين المصنفين التقليديين مع مجموعة متنوعة من النماذج التي تأخذ العدالة في الاعتبار عبر أساليب المعالجة المسبقة، والمعالجة أثناء العملية، والمعالجة اللاحقة. تظهر النتائج أن الأساليب التي تأخذ العدالة في الاعتبار تعزز كل من العدالة والدقة مقارنة بالنماذج التقليدية. تشمل اتجاهات البحث المستقبلية توسيع تقييم العدالة عبر سمات محمية متعددة، وتطوير نماذج توليد بيانات اصطناعية عادلة، وإنشاء خوارزميات تصنيف عادلة وقابلة للتفسير لفهم أفضل لمصادر التحيز في كل من الخوارزميات ومجموعات البيانات.

مقدمة

تسلط مقدمة الورقة الضوء على الحاجة الملحة للتحول الرقمي في قطاع البنوك، مدفوعة بعدم كفاءة الأنظمة المصرفية التقليدية وظهور تقنيات متقدمة مثل أتمتة العمليات الروبوتية (RPA)، والبيانات الضخمة، والحوسبة السحابية، وبلوك تشين. يتم التركيز بشكل كبير على أتمتة تقييم الائتمان، وهو عنصر حاسم في إدارة المخاطر للمؤسسات المالية، حيث تم تطبيق طرق التعلم الآلي (ML) بشكل متزايد. ومع ذلك، يشير المؤلفون إلى إمكانية التمييز في اتخاذ القرارات المعتمدة على التعلم الآلي، خاصة فيما يتعلق بالسمات المحمية مثل العرق والجنس، مما يبرز أهمية العدالة في هذه النماذج.

تكشف الأدبيات عن فجوة في تطبيق التعلم الآلي الذي يأخذ العدالة في الاعتبار بشكل خاص داخل قطاع البنوك، لا سيما في تقييم الائتمان. بينما استكشفت عدة دراسات العدالة في التعلم الآلي عبر مجالات متنوعة، تناولت القليل فقط هذه القضية في البنوك، وغالبًا ما كانت الأبحاث الحالية محدودة بمجموعات بيانات معينة أو مقاييس عدالة. يجادل المؤلفون بضرورة مراجعة شاملة لنماذج التعلم الآلي التي تأخذ العدالة في الاعتبار ومقاييس العدالة القابلة للتطبيق على تقييم الائتمان، حيث يمكن أن يؤثر اختيار معيار العدالة بشكل كبير على قرارات الإقراض. تهدف الورقة إلى تقديم نظرة عامة على التعلم الآلي الذي يأخذ العدالة في الاعتبار، وتحليل مجموعات بيانات تقييم الائتمان الشائعة، وتقييم كل من نماذج التصنيف التقليدية وتلك التي تأخذ العدالة في الاعتبار، مما يساهم في النقاش حول ممارسات الإقراض العادلة.

طرق

ت outlines قسم الطرق الإعدادات التجريبية والنتائج التي تم الحصول عليها خلال الدراسة. تم استخدام تكوينات متنوعة لاختبار الفرضيات، بما في ذلك البيئات المسيطر عليها التي سمحت بالتلاعب الدقيق بالمتغيرات. تم توثيق قياسات وظروف محددة لضمان إمكانية إعادة إنتاج النتائج وموثوقيتها.

أظهرت النتائج التجريبية اتجاهات كبيرة تدعم الفرضيات الأولية. كشفت تحليل البيانات عن علاقات بين المتغيرات التي تم التلاعب بها والنتائج الملاحظة، مع دلالة إحصائية تشير إليها قيم p أقل من 0.05. تساهم هذه النتائج في فهم أعمق للآليات الأساسية المعنية وتوفر أساسًا للبحوث المستقبلية في هذا المجال.

نقاش

في هذا القسم، يناقش المؤلفون نماذج التعلم الآلي التي تأخذ العدالة في الاعتبار والمخصصة لتقييم الائتمان، مصنفين إياها إلى ثلاثة نهج: المعالجة المسبقة، والمعالجة أثناء العملية، والمعالجة اللاحقة. يتم تأطير مشكلة تقييم الائتمان كمسألة تصنيف ثنائي، حيث يتم تصنيف المتقدمين إلى فئات مخاطر ‘جيدة’ أو ‘سيئة’ بناءً على سلوكهم المالي. الهدف من النماذج التي تأخذ العدالة في الاعتبار هو تطوير دالة تحويل \( f: D \to Y \) تقلل من الخسارة بينما تقلل من النتائج التمييزية. يبرز المؤلفون تقنيات متنوعة ضمن كل نهج، مثل تمثيلات التعلم العادلة (LFR) ومزيل التأثير المتباين (DIR) للمعالجة المسبقة، وطريقة أغاروال وAdaFair للمعالجة أثناء العملية، وعمليات المعالجة اللاحقة المتساوية (EOP) وعمليات المعالجة اللاحقة المتساوية المعايرة (CEP) للمعالجة اللاحقة.

يتم تقييم هذه النماذج باستخدام مجموعات بيانات متعددة، مما يكشف أن النماذج التي تأخذ العدالة في الاعتبار، وخاصة AdaFair وLFR-MLP، غالبًا ما تتفوق على نماذج التصنيف التقليدية في تحقيق توازن بين الدقة والعدالة. ومع ذلك، يشير المؤلفون إلى القيود في دراستهم، مثل التركيز على السمات المحمية الفردية دون النظر في تقاطع السمات المتعددة، والاعتماد على مجموعات بيانات مستخدمة بشكل شائع قد لا تعكس بالكامل التحيزات المعقدة في التطبيقات المالية. تشمل اتجاهات البحث المستقبلية تقييمًا أوسع للعدالة عبر سمات محمية متعددة وتطوير نماذج أكثر قابلية للتفسير لكشف التحيزات الأساسية في كل من الخوارزميات ومجموعات البيانات.

Journal: Digital Finance, Volume: 8, Issue: 3
DOI: https://doi.org/10.1007/s42521-026-00202-6
Publication Date: 2026-06-16
Author(s): Huyen Giang Thi Thu et al.
Primary Topic: Financial Distress and Bankruptcy Prediction

Overview

The digitalization of credit scoring is increasingly vital for financial institutions, particularly as machine learning techniques are employed to assess customer creditworthiness. However, these models often exhibit biases related to protected attributes such as race and gender. This paper presents a comprehensive experimental study that evaluates fairness-aware machine learning models in the context of credit scoring, examining financial datasets, predictive models, and fairness measures. The findings indicate that fairness-aware models outperform traditional classification models by achieving a more favorable balance between predictive accuracy and fairness.

In the conclusions, the authors highlight the inherent biases present in commonly used credit scoring datasets, as revealed through Bayesian network analysis. They compare traditional classifiers with various fairness-aware models across pre-processing, in-processing, and post-processing approaches. The results demonstrate that fairness-aware methods enhance both fairness and accuracy compared to conventional models. Future research directions include expanding the evaluation of fairness across multiple protected attributes, developing fair synthetic data generation models, and creating fair and explainable classification algorithms to better understand the sources of bias in both algorithms and datasets.

Introduction

The introduction of the paper highlights the pressing need for digital transformation in the banking sector, driven by the inefficiencies of traditional banking systems and the advent of advanced technologies such as Robotic Process Automation (RPA), Big Data, Cloud Computing, and Blockchain. A significant focus is placed on the automation of credit scoring, a critical component of risk management for financial institutions, where machine learning (ML) methods have been increasingly applied. However, the authors note the potential for discrimination in ML-based decision-making, particularly concerning protected attributes like race and gender, underscoring the importance of fairness in these models.

The literature reveals a gap in the application of fairness-aware ML specifically within the banking sector, particularly in credit scoring. While several studies have explored fairness in ML across various domains, only a few have addressed this issue in banking, with existing research often limited to specific datasets or fairness measures. The authors argue for the necessity of a comprehensive review of fairness-aware ML models and fairness metrics applicable to credit scoring, as the choice of fairness criterion can significantly impact lending decisions. The paper aims to provide an overview of fairness-aware ML, analyze popular credit scoring datasets, and evaluate both traditional and fairness-aware classification models, ultimately contributing to the discourse on equitable lending practices.

Methods

The section on Methods outlines the experimental setups and results obtained during the study. Various configurations were employed to test the hypotheses, including controlled environments that allowed for precise manipulation of variables. Specific measurements and conditions were documented to ensure reproducibility and reliability of the findings.

The experimental results demonstrated significant trends that support the initial hypotheses. Data analysis revealed correlations between the manipulated variables and the observed outcomes, with statistical significance indicated by p-values less than 0.05. These findings contribute to a deeper understanding of the underlying mechanisms at play and provide a foundation for future research in the field.

Discussion

In this section, the authors discuss fairness-aware machine learning (ML) models tailored for credit scoring, categorizing them into three approaches: pre-processing, in-processing, and post-processing. The credit scoring problem is framed as a binary classification task, where applicants are classified into ‘good’ or ‘bad’ risk classes based on their financial behavior. The goal of fairness-aware models is to develop a mapping function \( f: D \to Y \) that minimizes loss while reducing discriminatory outcomes. The authors highlight various techniques within each approach, such as Learning Fair Representations (LFR) and Disparate Impact Remover (DIR) for pre-processing, Agarwal’s method and AdaFair for in-processing, and Equalized Odds Post-Processing (EOP) and Calibrated Equalized Odds Post-Processing (CEP) for post-processing.

The evaluation of these models is conducted using multiple datasets, revealing that fairness-aware models, particularly AdaFair and LFR-MLP, often outperform traditional classification models in balancing accuracy and fairness. However, the authors note limitations in their study, such as the focus on individual protected attributes without considering the intersectionality of multiple attributes, and the reliance on commonly used datasets that may not fully encapsulate complex biases in financial applications. Future research directions include a broader evaluation of fairness across multiple protected attributes and the development of more explainable models to uncover underlying biases in both algorithms and datasets.

شارك: