الذكاء الاصطناعي القابل للتفسير في كشف احتيال بطاقات الائتمان: نماذج قابلة للتفسير وقرارات شفافة لتعزيز الثقة والامتثال في الولايات المتحدة
Explainable AI in Credit Card Fraud Detection: Interpretable Models and Transparent Decision-making for Enhanced Trust and Compliance in the USA

المجلة: Journal of Computer Science and Technology Studies، المجلد: 6، العدد: 2
DOI: https://doi.org/10.32996/jcsts.2024.6.2.1
تاريخ النشر: 2024-04-06
المؤلف: MD Rokibul Hasan وآخرون
الموضوع الرئيسي: الذكاء الاصطناعي في القانون

نظرة عامة

تبحث الدراسة في فعالية تقنيات التعلم الآلي المختلفة في اكتشاف احتيال بطاقات الائتمان، وهي قضية هامة عبر عدة قطاعات مثل الرعاية الصحية والمالية والتجارة الإلكترونية. تم تحليل أربع تقنيات رئيسية: آلة الدعم الناقل (SVM)، والانحدار اللوجستي، وغابة عشوائية، والشبكة العصبية الاصطناعية (ANN). تم تقييم أداء هذه النماذج باستخدام مقاييس تشمل الدقة والاسترجاع والدقة العامة ومقياس F. على الرغم من تحقيق جميع النماذج لمعدلات دقة عالية تصل إلى 99%، إلا أن ذلك كان يُعزى بشكل أساسي إلى خصائص مجموعة البيانات، التي تحتوي على 284,807 سمة مع 492 معاملة احتيالية فقط.

من بين النماذج، حقق SVM أعلى استرجاع عند 89.5، مما يدل على قدرته القوية في تحديد الحالات الإيجابية الحقيقية، بينما سجل نموذج ANN أعلى دقة عند 79.4، مما يُظهر فعاليته في إجراء توقعات دقيقة. تسلط هذه النتائج الضوء على أنه بينما تعتبر الدقة العالية ملحوظة، قد لا تكون المقياس الأكثر إفادة للمقارنة، مما يبرز أهمية النظر في الاسترجاع والدقة عند تقييم أداء النموذج في اكتشاف الاحتيال.

مقدمة

تسلط مقدمة الورقة الضوء على الدور الحاسم لاكتشاف الاحتيال عبر مختلف القطاعات، بما في ذلك التجارة الإلكترونية والمالية والرعاية الصحية، مما يبرز النتائج الواعدة التي تم تحقيقها من خلال تطبيق أنظمة الذكاء الاصطناعي (AI) (بورا، 2022). على الرغم من هذه التقدمات، فإن اعتماد الذكاء الاصطناعي في اكتشاف الاحتيال يعوقه تحديات تتعلق بالامتثال والثقة، ويرجع ذلك أساسًا إلى طبيعة العديد من نماذج الذكاء الاصطناعي التي تشبه الصندوق الأسود. تهدف هذه الورقة إلى تقديم مراجعة شاملة للأدبيات حول فوائد الذكاء الاصطناعي والتعلم الآلي في اكتشاف الاحتيال، مع تقديم توصيات للبحث والتنفيذ في المستقبل.

تركز الدراسة بشكل رئيسي على أهمية الذكاء الاصطناعي القابل للتفسير (XAI) في اكتشاف احتيال بطاقات الائتمان، مما يبرز الحاجة إلى نماذج قابلة للتفسير وعمليات اتخاذ قرارات شفافة لتعزيز الامتثال والثقة. تستكشف الورقة أهمية القابلية للتفسير في أنظمة الذكاء الاصطناعي وتناقش منهجيات مختلفة لتحقيق القابلية للتفسير في اكتشاف الاحتيال. مع استمرار المحتالين في تعديل تكتيكاتهم، تصبح الحاجة إلى تقنيات متقدمة لتحديد ومكافحة الأنشطة الاحتيالية بشكل فعال أكثر وضوحًا. يُعتبر الذكاء الاصطناعي، وخاصة من خلال خوارزميات التعلم الآلي، أداة حيوية لتحليل مجموعات البيانات الكبيرة واكتشاف الأنماط الشاذة التي تشير إلى الاحتيال، على الرغم من المخاوف المتعلقة بشفافية وتعقيد هذه النماذج.

الطرق

توضح قسم المنهجية مجموعة البيانات المستخدمة في البحث، والتي تم الحصول عليها من كاجل وتتكون من معاملات بطاقات الائتمان من حاملي البطاقات الأمريكيين المسجلة على مدى يومين في يناير 2024. تتضمن مجموعة البيانات ما مجموعه 284,807 معاملة، مع تحديد 491 فقط على أنها احتيالية، مما يؤدي إلى عدم توازن كبير في الفئات حيث تمثل المعاملات الاحتيالية حوالي 0.172% من الإجمالي. لمعالجة هذا التوازن والحفاظ على سرية العملاء، تم تحويل بعض الميزات باستخدام تحليل المكونات الرئيسية (PCA)، مما أسفر عن متغيرات مُعلمة من V1 إلى V21، بينما تم الاحتفاظ بسمات أخرى مثل الفئة والوقت والمبلغ في شكلها الأصلي.

لتحليل التجارب، تم استخدام بايثون كلغة برمجة رئيسية لكل من مهام النمذجة والمعالجة المسبقة، مستفيدًا من مكتباتها الواسعة المخصصة لتطبيقات علم البيانات. تشمل المكتبات الرئيسية المستخدمة pandas لمعالجة البيانات وتحليلها، وmatplotlib لتصور البيانات، وNumPy للحسابات العلمية، وscikit-learn لتطوير النماذج. يبرز هذا الإطار المنهجي التزام البحث بممارسات معالجة وتحليل البيانات القوية في سياق اكتشاف الاحتيال.

المناقشة

تسلط قسم المناقشة في ورقة البحث الضوء على القضية الحاسمة للاحتيال في بطاقات الائتمان، والتي تتضمن الاستخدام غير المصرح به لمعلومات البطاقة لإجراء معاملات دون موافقة حامل البطاقة. يتم تحديد أشكال مختلفة من الاحتيال، بما في ذلك سحب النقود والمشتريات التي تتم باستخدام تفاصيل بطاقات مسروقة، مما يبرز الحاجة إلى آليات فعالة للكشف والوقاية. تشير الأدبيات إلى أن اكتشاف احتيال بطاقات الائتمان يستخدم تقنيات وخوارزميات متنوعة لتحديد الأنشطة المشبوهة، وبالتالي حماية المؤسسات المالية والمستهلكين من خسائر كبيرة.

تُبرز أهمية القابلية للتفسير في نماذج الذكاء الاصطناعي (AI) لاكتشاف الاحتيال، حيث تعزز ثقة أصحاب المصلحة والامتثال للمعايير التنظيمية. تُلاحظ نماذج مثل أشجار القرار والأطر التراكمية العامة لشفافيتها وقدرتها على تقديم رؤى حول عمليات اتخاذ القرار. تناقش الورقة أيضًا تطبيق تقنيات التعلم الآلي، وبشكل خاص آلات الدعم الناقل (SVM) والشبكات العصبية الاصطناعية (ANN)، التي تفوقت على الأنظمة التقليدية القائمة على القواعد في اكتشاف الاحتيال. تكشف النتائج أنه بينما حققت جميع النماذج معدلات دقة عالية، تفوق SVM في الاسترجاع، مما يحدد المعاملات الاحتيالية بفعالية، بينما أظهرت ANN دقة متفوقة في التعريفات الإيجابية. تدعو هذه الدراسة إلى دمج هذه النماذج المتقدمة في المؤسسات المالية لتعزيز اكتشاف الاحتيال في الوقت الفعلي وكفاءة العمليات، مما يساهم في الاستقرار الاقتصادي في النهاية.

Journal: Journal of Computer Science and Technology Studies, Volume: 6, Issue: 2
DOI: https://doi.org/10.32996/jcsts.2024.6.2.1
Publication Date: 2024-04-06
Author(s): MD Rokibul Hasan et al.
Primary Topic: Artificial Intelligence in Law

Overview

The research investigates the effectiveness of various Machine Learning techniques in detecting credit card fraud, a significant issue across multiple sectors such as healthcare, finance, and ecommerce. Four primary techniques were analyzed: Support Vector Machine (SVM), Logistic Regression, Random Forest, and Artificial Neural Network (ANN). The performance of these models was assessed using metrics including Precision, Recall, Accuracy, and F-measure. Despite all models achieving high accuracy rates of 99%, this was primarily attributed to the dataset’s characteristics, which contained 284,807 attributes with only 492 fraudulent transactions.

Among the models, SVM achieved the highest recall at 89.5, indicating its strong capability in identifying true positive instances, while the ANN model recorded the highest precision at 79.4, showcasing its effectiveness in making accurate predictions. These findings highlight that while high accuracy is notable, it may not be the most informative metric for comparison, emphasizing the importance of considering recall and precision in evaluating model performance in fraud detection.

Introduction

The introduction of the paper highlights the critical role of fraud detection across various sectors, including e-commerce, finance, and healthcare, emphasizing the promising outcomes achieved through the application of artificial intelligence (AI) systems (Bora, 2022). Despite these advancements, the adoption of AI in fraud detection is hindered by challenges related to compliance and trust, primarily due to the black-box nature of many AI models. This paper aims to provide a comprehensive review of the literature on AI and machine learning’s benefits in fraud detection, while also offering recommendations for future research and implementation.

A key focus of the study is the importance of explainable AI (XAI) in credit fraud detection, which emphasizes the need for interpretable models and transparent decision-making processes to enhance compliance and trust. The paper explores the significance of interpretability in AI systems and discusses various methodologies for achieving explainability in fraud detection. As fraudsters continuously adapt their tactics, the necessity for advanced technologies to effectively identify and combat fraudulent activities becomes increasingly evident. AI, particularly through machine learning algorithms, is positioned as a vital tool for analyzing large datasets and detecting anomalous patterns indicative of fraud, despite concerns regarding the transparency and complexity of these models.

Methods

The methodology section outlines the dataset utilized for the research, which was sourced from Kaggle and consists of credit card transactions from American cardholders recorded over two days in January 2024. The dataset includes a total of 284,807 transactions, with only 491 identified as fraudulent, resulting in a significant class imbalance where fraudulent transactions represent approximately 0.172% of the total. To address this imbalance and maintain client confidentiality, certain features were transformed using Principal Component Analysis (PCA), resulting in variables labeled V1 through V21, while other attributes such as class, time, and amount were retained in their original form.

For the experimental analysis, Python was employed as the primary programming language for both modeling and preprocessing tasks, leveraging its extensive libraries tailored for data science applications. Key libraries utilized include pandas for data manipulation and analysis, matplotlib for data visualization, NumPy for scientific computations, and scikit-learn for model development. This methodological framework underscores the research’s commitment to robust data handling and analysis practices in the context of fraud detection.

Discussion

The discussion section of the research paper highlights the critical issue of credit card fraud, which involves unauthorized use of card information for transactions without the cardholder’s consent. Various forms of fraud, including cash withdrawals and purchases made with stolen card details, are identified, emphasizing the need for effective detection and prevention mechanisms. The literature indicates that credit fraud detection employs diverse techniques and algorithms to identify suspicious activities, thereby protecting financial institutions and consumers from significant losses.

The importance of interpretability in artificial intelligence (AI) models for fraud detection is underscored, as it fosters stakeholder trust and compliance with regulatory standards. Models such as decision trees and generalized additive frameworks are noted for their transparency and ability to provide insights into decision-making processes. The paper also discusses the application of machine learning techniques, specifically Support Vector Machines (SVM) and Artificial Neural Networks (ANN), which outperformed traditional rule-based systems in detecting fraud. The findings reveal that while all models achieved high accuracy rates, SVM excelled in recall, effectively identifying fraudulent transactions, whereas ANN demonstrated superior precision in positive identifications. This research advocates for the integration of these advanced models in financial organizations to enhance real-time fraud detection and operational efficiency, ultimately contributing to economic stability.