DOI: https://doi.org/10.1007/s13347-026-01043-2
تاريخ النشر: 2026-02-06
المؤلف: Joshua Thong
الموضوع الرئيسي: الأخلاقيات والآثار الاجتماعية للذكاء الاصطناعي
نظرة عامة
في مجال العدالة الخوارزمية، تم مناقشة معيارين بارزين، الاحتمالات المتساوية والمعايرة، على نطاق واسع. ومع ذلك، أظهر كلاينبرغ وآخرون (2017) أن هذه المعايير غير متسقة في السيناريوهات غير التافهة، مما يستدعي رفض واحد على الأقل. استجابةً لهذا التحدي، قدمت إيفا (2022) نسخة معدلة من المعايرة تُعرف بتتبع المعدل الأساسي. على الرغم من نيتها معالجة العدالة، وجد ستيوارت وآخرون (2024) أن تتبع المعدل الأساسي يفشل أيضًا في الحفاظ على الاتساق مع الاحتمالات المتساوية في الحالات غير التافهة.
لتعزيز النقاش حول العدالة الخوارزمية، يقترح المؤلف معيارًا جديدًا يُسمى تتبع المعدل الأساسي المقارن، والذي تم تصميمه ليكون متسقًا مع الاحتمالات المتساوية حتى في الحالات المعقدة. يحتفظ هذا الإطار الجديد، على الرغم من كونه تخفيفًا كبيرًا للمعايير السابقة، بالقدرة على تحديد العيوب التي لا تستطيع مقاييس العدالة الحديثة الأخرى، مثل التمدد بواسطة نيلسن وستيوارت (2024أ) والتباعد بواسطة إيفا (2024)، اكتشافها. تشير النتائج إلى أن تتبع المعدل الأساسي المقارن يمكن أن يكون أداة قيمة في السعي المستمر نحو اتخاذ قرارات خوارزمية عادلة.
مقدمة
تسلط مقدمة هذه الورقة البحثية الضوء على الأهمية المتزايدة للعدالة الخوارزمية في الذكاء الاصطناعي (AI)، خاصة مع تأثير أنظمة الذكاء الاصطناعي بشكل متزايد على عمليات اتخاذ القرار الحاسمة عبر مجالات مختلفة، بما في ذلك العدالة الجنائية والرعاية الصحية والتوظيف. يؤكد المؤلف على ضرورة ضمان أن تكون القرارات المدفوعة بالذكاء الاصطناعي عادلة وغير تمييزية، مما يؤدي إلى اقتراح عدة معايير للعدالة في الأدبيات. تركز الورقة بشكل خاص على ثلاثة معايير: الاحتمالات المتساوية، المعايرة، وتتبع المعدل الأساسي. ومن الجدير بالذكر أن الأعمال السابقة لكل من كلاينبرغ وآخرون (2017) وستيوارت وآخرون (2024) تشير إلى عدم الاتساق بين هذه المعايير في الحالات غير التافهة، مما يستدعي رفض معيار واحد على الأقل.
يقدم المؤلف مفهومًا جديدًا يُسمى تتبع المعدل الأساسي المقارن، وهو نسخة مخففة من تتبع المعدل الأساسي التي تحافظ على الاتساق مع الاحتمالات المتساوية، حتى في السيناريوهات المعقدة. يُعتبر هذا المعيار إضافة قيمة للمعايير الحالية للعدالة، حيث يمكنه تحديد العيوب التي قد تتجاهلها معايير أخرى، مثل التمدد والتباعد. ومع ذلك، يعترف المؤلف أيضًا بأن تتبع المعدل الأساسي المقارن له حدوده، حيث توجد حالات لا يمكنه تشخيصها يمكن أن تعالجها التمدد أو التباعد. تدعو الورقة إلى اعتماد تتبع المعدل الأساسي المقارن جنبًا إلى جنب مع الاحتمالات المتساوية، التمدد، والتباعد لإنشاء إطار شامل للعدالة الخوارزمية. يتم توضيح هيكل الورقة، مع تخصيص الأقسام اللاحقة لصياغة معايير العدالة وتقديم تتبع المعدل الأساسي المقارن، جنبًا إلى جنب مع مناقشة استقلاله المنطقي عن المعايير الأخرى وقدراته التشخيصية.
مناقشة
في هذا القسم، تناقش الورقة الإطار النظري لتقييم العدالة الخوارزمية، مع التركيز بشكل خاص على شروط الاحتمالات المتساوية، المعايرة، وتتبع المعدل الأساسي. يتم تعريف سيناريو التقييم كأربعة عناصر ⟨Ω، Y، π، h⟩، حيث تمثل Ω مجموعة من الأفراد، Y هو متغير عشوائي يشير إلى ما إذا كان الفرد يمتلك خاصية معينة، π هو تقسيم لـ Ω إلى مجموعات، و h هو مقيم يقوم بتعيين درجات المخاطر للأفراد. تبرز الورقة أنه لكي يُعتبر المقيم عادلاً، يجب أن يستوفي شروطًا معينة: تتطلب المعايرة أن تتطابق نسبة الأفراد في مجموعة مع درجة خطر معينة مع المعدل الأساسي الفعلي لتلك المجموعة، بينما يتطلب تتبع المعدل الأساسي أن تعكس متوسط درجات المخاطر للمجموعات المختلفة معدلاتها الأساسية.
يقترح المؤلف تخفيفًا لتتبع المعدل الأساسي، يُسمى تتبع المعدل الأساسي المقارن، والذي يؤكد أنه إذا كانت مجموعة واحدة لديها معدل أساسي أقل من مجموعة أخرى، فيجب أن تتلقى أيضًا درجة خطر متوسطة أقل. يُعتبر هذا الشرط معيارًا ضروريًا، وإن لم يكن كافيًا، للعدالة. تستكشف الورقة أيضًا العلاقات المنطقية بين تتبع المعدل الأساسي المقارن، التمدد، والتباعد، والتي تُعرض جميعها كبدائل أضعف للمعايرة. يؤكد المؤلف أنه على الرغم من أن هذه المعايير متميزة، إلا أنها يمكن أن تخدم معًا كشرط ضروري للعدالة في التقييمات الخوارزمية، مما يسمح بفهم أكثر دقة للعدالة يتناسب مع تعقيدات السيناريوهات الواقعية. تختتم القسم بالتأكيد على أهمية اعتماد معايير متعددة لمعالجة العيوب الكامنة في كل شرط فردي.
DOI: https://doi.org/10.1007/s13347-026-01043-2
Publication Date: 2026-02-06
Author(s): Joshua Thong
Primary Topic: Ethics and Social Impacts of AI
Overview
In the realm of algorithmic fairness, two prominent criteria, Equalized Odds and Calibration, have been widely discussed. However, Kleinberg et al. (2017) demonstrated that these criteria are inconsistent in non-trivial scenarios, necessitating the rejection of at least one. In response to this challenge, Eva (2022) introduced a modified version of Calibration known as Base Rate Tracking. Despite its intent to address fairness, Stewart et al. (2024) found that Base Rate Tracking also fails to maintain consistency with Equalized Odds in non-trivial cases.
To advance the discourse on algorithmic fairness, the author proposes a new criterion termed Comparative Base Rate Tracking, which is designed to be consistent with Equalized Odds even in complex situations. This new framework, while a significant relaxation of previous criteria, retains the ability to identify deficiencies that other recent fairness measures, such as Spanning by Nielsen and Stewart (2024a) and Spacing by Eva (2024), cannot detect. The findings suggest that Comparative Base Rate Tracking could serve as a valuable tool in the ongoing pursuit of equitable algorithmic decision-making.
Introduction
The introduction of this research paper highlights the growing importance of algorithmic fairness in artificial intelligence (AI), particularly as AI systems increasingly influence critical decision-making processes across various domains, including criminal justice, healthcare, and employment. The author emphasizes the necessity of ensuring that AI-driven decisions are fair and non-discriminatory, leading to the proposal of several fairness criteria in the literature. Specifically, the paper focuses on three criteria: Equalized Odds, Calibration, and Base Rate Tracking. Notably, previous work by Kleinberg et al. (2017) and Stewart et al. (2024) indicates inconsistencies among these criteria in non-trivial cases, necessitating the rejection of at least one criterion.
The author introduces a new concept termed Comparative Base Rate Tracking, which is a weakened version of Base Rate Tracking that maintains consistency with Equalized Odds, even in complex scenarios. This criterion is posited as a valuable addition to the existing fairness criteria, as it can identify deficiencies that other criteria, such as Spanning and Spacing, may overlook. However, the author also acknowledges that Comparative Base Rate Tracking has its limitations, as there are cases it cannot diagnose that Spanning or Spacing can address. The paper advocates for the cumulative adoption of Comparative Base Rate Tracking alongside Equalized Odds, Spanning, and Spacing to create a comprehensive framework for algorithmic fairness. The structure of the paper is outlined, with subsequent sections dedicated to the formulation of the fairness criteria and the introduction of Comparative Base Rate Tracking, along with a discussion of its logical independence from the other criteria and its diagnostic capabilities.
Discussion
In this section, the paper discusses the theoretical framework for assessing algorithmic fairness, particularly focusing on the conditions of Equalized Odds, Calibration, and Base Rate Tracking. The assessment scenario is defined as a quadruple ⟨Ω, Y, π, h⟩, where Ω represents a collection of individuals, Y is a random variable indicating whether an individual possesses a certain property, π is a partition of Ω into groups, and h is an assessor assigning risk scores to individuals. The paper highlights that for an assessor to be considered fair, it must satisfy specific conditions: Calibration requires that the proportion of individuals in a group with a given risk score matches the actual base rate of that group, while Base Rate Tracking demands that the average risk scores of different groups reflect their respective base rates.
The author proposes a weakening of Base Rate Tracking, termed Comparative Base Rate Tracking, which maintains that if one group has a lower base rate than another, it should also receive a lower average risk score. This condition is argued to be a necessary, though not sufficient, criterion for fairness. The paper further explores the logical relationships between Comparative Base Rate Tracking, Spanning, and Spacing, all of which are presented as weaker alternatives to Calibration. The author asserts that while these criteria are distinct, they can collectively serve as necessary conditions for fairness in algorithmic assessments, allowing for a more nuanced understanding of fairness that accommodates the complexities of real-world scenarios. The section concludes by emphasizing the importance of adopting multiple criteria to address the deficiencies inherent in each individual condition.
