نموذج قابل للتفسير قائم على المفهوم والحجة للبيانات الجدولية
An interpretable model based on concept and argumentation for tabular data

شارك:
المجلة: Scientific Reports، المجلد: 16، العدد: 1
DOI: https://doi.org/10.1038/s41598-025-30540-1
PMID: https://pubmed.ncbi.nlm.nih.gov/41507242
تاريخ النشر: 2026-01-08
المؤلف: Haixiao Chi وآخرون
الموضوع الرئيسي: الذكاء الاصطناعي القابل للتفسير (XAI)

نظرة عامة

تقدم ورقة البحث إطارًا جديدًا يسمى نموذج المفهوم والحجة (CAM) مصممًا لتعزيز القابلية للتفسير في تطبيقات الذكاء الاصطناعي، لا سيما في المجالات عالية المخاطر مثل الرعاية الصحية، والبنوك، والأمن. غالبًا ما تتجاهل الأساليب التقليدية في تعلم الآلة دمج المعرفة القابلة للفهم من أوصاف البيانات، حيث تركز فقط على البيانات العددية والفئوية. يعالج CAM هذه الفجوة من خلال استخدام طبقة حجة كمية تستخرج وتستخدم المفاهيم من كل من البيانات وأوصافها. يمثل هذا الإطار المعرفة الميدانية بشكل صريح من خلال قواعد حجة قابلة للفهم، مما يسهل اتخاذ قرارات قابلة للتفسير ومتوافقة مع عمليات التفكير البشري.

تشير النتائج إلى أن CAM لا يوفر فقط مسارات تفكير شفافة ومتدرجة ولكن أيضًا يحقق مستوى عالٍ من الفهم وقبول المستخدم في التقييمات البشرية. تتيح التفسيرات الحوارية المستمدة من هياكل التفكير في CAM للمستخدمين فهم المنطق وراء القرارات بشكل فعال. بالإضافة إلى ذلك، تُظهر التجارب التي أُجريت على كل من المعايير مفتوحة المصدر ومجموعات البيانات الواقعية أن CAM يحافظ على أداء تنبؤي تنافسي مع ضمان دقة عالية وموثوقية في مخرجاته. تكمن المساهمة الرئيسية لـ CAM في قدرته على دمج الدلالات النصية مع التفكير الكمي، مما يوفر إطارًا موحدًا وقابلًا للتفسير يدعم اتخاذ قرارات موثوقة في التطبيقات عالية المخاطر دون التضحية بالدقة.

مقدمة

في هذا القسم، يقدم المؤلفون مجموعات البيانات المستخدمة لتقييم أداء CAM (آلية الانتباه السياقي). تشمل الدراسة أربع مجموعات بيانات: معيارين عامين ومجموعتين بيانات صناعية ملكية. يركز المعيار العام الأول، مجموعة بيانات Fico، على تقييم مخاطر الائتمان المستمدة من بيانات مكاتب الائتمان الخاصة بالمستهلكين ضمن قطاع البنوك. يهدف المعيار الثاني، مجموعة بيانات Mimic3 السريرية، إلى توقع الوفيات داخل المستشفى بناءً على البيانات الفسيولوجية المجمعة خلال الـ 24 ساعة الأولى من إقامة المريض في وحدة العناية المركزة (ICU).

بالإضافة إلى هذه المجموعات العامة، يدمج المؤلفون مجموعتين من بيانات مكافحة الاحتيال الملكية، المشار إليهما باسم data1 و data2، والتي تم الحصول عليها من تطبيقات التجارة الإلكترونية الخاصة بـ Alibaba. تُصنف جميع مجموعات البيانات على أنها متوسطة الحجم، تحتوي على ما بين 10,000 و 100,000 عينة. يتم تقديم ملخص لخصائص مجموعات البيانات في الجدول 4، مع توفر تفاصيل إضافية في الملحق.

مناقشة

يقدم قسم المناقشة في الورقة إطار الحجة الكمية (QAF) كنموذج قوي للتفكير تحت الصراع وعدم اليقين، لا سيما في سياق أنظمة الذكاء الاصطناعي القابلة للتفسير. يوسع QAF إطار الحجة ثنائية القطب (BAF) من خلال دمج المعرفة العددية لتحديد قوة العلاقات بين الحجج، مما يمكّن من تمثيل أكثر دقة لعمليات التفكير. يتم صياغة الإطار كرباعية \( \langle A, E, \beta, \omega \rangle \)، حيث تمثل \( A \) مجموعة من الحجج، ويمثل \( E \) حواف موجهة تشير إلى العلاقات بين الحجج، و\( \beta \) تعين درجات أساسية للحجج، و\( \omega \) تعين أوزانًا للحواف تعكس تأثيرها. تسلط الورقة الضوء على تطبيق طرق التفكير المختلفة، بما في ذلك الشبكات العصبية متعددة الطبقات (MLPs)، لتقييم دلالات QAF، مما يسمح بالاستخراج التلقائي للمعرفة الكمية من البيانات الجدولية.

علاوة على ذلك، يحدد القسم عملية تمثيل واكتساب المعرفة ضمن QAF، مع التأكيد على التمييز بين المعرفة على مستوى الإنسان والمعرفة الضمنية المستمدة من البيانات. يقترح المؤلفون نهجًا هرميًا لبناء نموذج تجريد المفهوم (CAM) الذي ينقي ويختار المفاهيم ذات المعنى من خلال سلسلة من طبقات الحجة الكمية (QALs). تتوج هذه العملية بإطار QAF شامل يدعم التفسيرات الحوارية، مما يمكّن المستخدمين من التفاعل مع النموذج وفهم التفكير وراء القرارات. توضح الورقة هذا الإطار من خلال دراسة حالة حول تقييم المخاطر، مما يظهر كيف يمكن لـ QAF تبسيط التفكير المعقد إلى حوارات بديهية، مما يعزز القابلية للتفسير ومشاركة المستخدم في عمليات اتخاذ القرار في الذكاء الاصطناعي.

Journal: Scientific Reports, Volume: 16, Issue: 1
DOI: https://doi.org/10.1038/s41598-025-30540-1
PMID: https://pubmed.ncbi.nlm.nih.gov/41507242
Publication Date: 2026-01-08
Author(s): Haixiao Chi et al.
Primary Topic: Explainable Artificial Intelligence (XAI)

Overview

The research paper introduces a novel framework called the Concept and Argumentation Model (CAM) designed to enhance interpretability in artificial intelligence applications, particularly in high-risk domains such as healthcare, banking, and security. Traditional machine learning approaches often overlook the integration of human-understandable knowledge from data descriptions, focusing solely on numerical and categorical data. CAM addresses this gap by employing a quantitative argumentation layer that extracts and utilizes concepts from both the data and its descriptions. This framework explicitly represents domain knowledge through understandable argumentation rules, facilitating decisions that are interpretable and aligned with human reasoning processes.

The findings indicate that CAM not only provides transparent and stepwise reasoning paths but also achieves a high level of comprehensibility and user acceptance in human evaluations. The dialogical explanations derived from CAM’s reasoning structures allow users to grasp the rationale behind decisions effectively. Additionally, experiments conducted on both open-source benchmarks and real-world datasets demonstrate that CAM maintains competitive predictive performance while ensuring high fidelity and faithfulness in its outputs. The primary contribution of CAM lies in its ability to integrate textual semantics with quantitative reasoning, thereby delivering a unified, interpretable modeling framework that supports trustworthy decision-making in high-stakes applications without sacrificing accuracy.

Introduction

In this section, the authors introduce the datasets utilized to evaluate the performance of the CAM (Contextual Attention Mechanism). The study encompasses four datasets: two public benchmarks and two proprietary industrial datasets. The first public benchmark, the Fico dataset, focuses on credit risk assessment derived from consumer credit bureau data within the banking sector. The second benchmark, the Mimic3 clinical dataset, aims to predict in-hospital mortality based on physiological data collected during the initial 24 hours of a patient’s stay in the intensive care unit (ICU).

In addition to these public datasets, the authors incorporate two proprietary anti-fraud datasets, referred to as data1 and data2, which were sourced from Alibaba’s e-commerce applications. All datasets are characterized as medium-sized, containing between 10,000 and 100,000 samples. A summary of the datasets’ characteristics is presented in Table 4, with additional details available in the Appendix.

Discussion

The discussion section of the paper introduces the quantitative argumentation framework (QAF) as a robust model for reasoning under conflict and uncertainty, particularly in the context of explainable AI systems. QAF extends the bipolar argumentation framework (BAF) by incorporating numerical knowledge to quantify the strength and relationships of arguments, thereby enabling a more nuanced representation of reasoning processes. The framework is formalized as a quadruple \( \langle A, E, \beta, \omega \rangle \), where \( A \) represents a set of arguments, \( E \) denotes directed edges indicating relationships between arguments, \( \beta \) assigns base scores to arguments, and \( \omega \) assigns weights to edges reflecting their influence. The paper highlights the application of various reasoning methods, including multilayer perceptrons (MLPs), to evaluate the semantics of QAF, allowing for the automatic extraction of quantitative knowledge from tabular data.

Furthermore, the section outlines the process of knowledge representation and acquisition within QAF, emphasizing the distinction between human-level knowledge and implicit knowledge derived from data. The authors propose a hierarchical approach to construct a Concept Abstraction Model (CAM) that iteratively refines and selects meaningful concepts through a series of Quantitative Argumentation Layers (QALs). This process culminates in a comprehensive QAF that supports dialogical explanations, enabling users to interact with the model and understand the reasoning behind decisions. The paper illustrates this framework through a case study on risk assessment, demonstrating how the QAF can simplify complex reasoning into intuitive dialogues, thereby enhancing interpretability and user engagement in AI decision-making processes.

شارك: