نموذج التعلم الفيدرالي لكشف احتيال بطاقات الائتمان مع تقنيات توازن البيانات
Federated learning model for credit card fraud detection with data balancing techniques

المجلة: Neural Computing and Applications، المجلد: 36، العدد: 11
DOI: https://doi.org/10.1007/s00521-023-09410-2
تاريخ النشر: 2024-01-20
المؤلف: Mustafa Abdul Salam وآخرون
الموضوع الرئيسي: تقنيات تصنيف البيانات غير المتوازنة

نظرة عامة

تتناول ورقة البحث القضية الملحة للاحتيال في معاملات بطاقات الائتمان، والتي أدت إلى خسائر مالية كبيرة لكل من المستهلكين والبنوك. وتؤكد على ضرورة وجود أنظمة قوية للكشف عن الاحتيال، خاصة من خلال عدسة الكشف عن احتيال بطاقات الائتمان (CCFD). تحدد الدراسة التحديات الرئيسية في تطوير أنظمة فعالة للكشف عن الاحتيال، ولا سيما القيود المفروضة على تبادل البيانات بين المؤسسات المالية بسبب مخاوف الأمان والخصوصية، مما يعيق طرق الكشف التقليدية عن الاحتيال. للتغلب على هذه التحديات، تقترح الورقة استخدام أطر التعلم الفيدرالي، وخاصة TensorFlow Federated وPyTorch، للكشف عن احتيال بطاقات الائتمان.

بالإضافة إلى ذلك، تسلط الورقة الضوء على المشكلة الحرجة لعدم توازن الفئات في معاملات بطاقات الائتمان، حيث تشكل المعاملات الاحتيالية جزءًا صغيرًا من الإجمالي. لمواجهة ذلك، تجري الدراسة تحليلًا مقارنًا لمجموعة من تقنيات إعادة أخذ العينات الفردية والهجينة التي تهدف إلى تحقيق توازن في مجموعة البيانات. تشير النتائج التجريبية إلى أن طرق إعادة أخذ العينات الهجينة تحقق أداءً متفوقًا لنماذج التعلم الآلي، مثل الغابة العشوائية (RF)، والانحدار اللوجستي، وأقرب الجيران (KNN)، وشجرة القرار (DT)، وGaussian Naive Bayes (NB)، محققة دقة تصل إلى 99.99%، 94.61%، 99.96%، 99.98%، و91.47%، على التوالي. ومن الجدير بالذكر أن مصنف RF يظهر أعلى أداء عبر مقاييس متعددة، بما في ذلك الدقة، والاسترجاع، والدقة، وF-score. علاوة على ذلك، بينما يوفر إطار PyTorch دقة تنبؤ أعلى لنموذج التعلم الفيدرالي مقارنة بـ TensorFlow، فإنه يتطلب موارد حسابية أكبر.

مقدمة

تسلط المقدمة الضوء على الزيادة السريعة في معاملات بطاقات الائتمان المدفوعة بالتقدم في الخدمات الإلكترونية، مثل التجارة الإلكترونية والمدفوعات عبر الهاتف المحمول. أدت هذه الزيادة إلى خسائر مالية كبيرة بسبب احتيال بطاقات الائتمان، والتي بلغت حوالي 18.39 مليار دولار على مستوى العالم في عام 2018، ارتفاعًا من 14.99 مليار دولار في عام 2017. تشير التوقعات إلى أن إجمالي حجم بطاقات الدفع سيصل إلى 57.08 تريليون دولار بحلول عام 2023، مع توقع ارتفاع احتيال البطاقات إلى 35.67 مليار دولار، وقد يصل إلى 40 مليار دولار بحلول عام 2027.

تؤكد الورقة على الطبيعة المزدوجة للأنشطة الاحتيالية، التي يمكن أن تنشأ من بطاقات مسروقة أو معلومات زائفة. استكشف باحثون مختلفون الكشف عن احتيال بطاقات الائتمان، مستخدمين بشكل أساسي نماذج التعلم الآلي الخاضعة للإشراف للتفريق بين المعاملات الاحتيالية والشرعية. على الرغم من هذه التقدمات، تواجه تطوير نظام فعال للكشف عن الاحتيال تحديات كبيرة، بما في ذلك مجموعات البيانات غير الكافية وتوزيع البيانات غير المتوازن.

طرق

في هذا القسم، يستكشف المؤلفون آثار استراتيجيات إعادة أخذ العينات الفردية والهجينة المختلفة على مجموعة بيانات مستخدمة للكشف عن احتيال بطاقات الائتمان. يستخدمون خوارزميات تصنيف متعددة، بما في ذلك أشجار القرار (DT)، وGaussian Naive Bayes (GaussianNB)، والغابة العشوائية (RF)، وأقرب الجيران (KNN)، والانحدار اللوجستي، لتقييم نتائج الأداء. يتم تنفيذ نموذج التعلم الفيدرالي عبر أطر مختلفة لمعالجة مخاوف أمان البيانات والخصوصية.

تم إجراء التجارب باستخدام Python (الإصدار 3) واستفادت من عدة مكتبات مفتوحة المصدر، بما في ذلك Scikit-learn (1.1.3)، وpandas (1.4.4)، وNumPy (1.22.3)، وmatplotlib (3.5.3)، وTensorFlow Federated (0.17.0)، وPyTorch (1.2.0)، وImblearn (0.9.1). كانت البيئة الحسابية تتكون من جهاز كمبيوتر مكتبي مزود بمعالج Intel Core i7 بسرعة 1.80 جيجاهرتز، وذاكرة وصول عشوائي سعتها 16 جيجابايت، ونظام تشغيل Windows 10 64 بت.

نتائج

في قسم النتائج، تقيم الدراسة أداء مصنفات مختلفة، بما في ذلك تقنيات التعلم الآلي التقليدية وشبكة عصبية تلافيفية (CNN)، باستخدام مقاييس مثل الدقة، والاسترجاع، والدقة، والخسارة، ومقياس F1، وإجمالي الوقت الحسابي. يتم تقييم فعالية هذه المصنفات بالتزامن مع تقنيات إعادة أخذ العينات لتعزيز أدائها.

استخدم التحليل المقارن نسبة تدريب-اختبار 80:20، مما يكشف عن رؤى مهمة حول دقة المصنفات. تؤكد النتائج على فعالية المصنفات في سياق طرق إعادة أخذ العينات المطبقة، على الرغم من عدم تقديم نتائج عددية محددة ومقارنات مفصلة في هذا الملخص.

مناقشة

يتناول قسم المناقشة في ورقة البحث التحديات الحرجة في أنظمة الكشف عن احتيال بطاقات الائتمان (CCFDS)، مع التركيز بشكل خاص على نقص البيانات وعدم توازن الفئات. يبرز المؤلفون أن نقص مجموعات البيانات المتاحة للجمهور، بشكل أساسي بسبب مخاوف الخصوصية المتعلقة بالبيانات بين البنوك، يتطلب استخدام نهج التعلم الفيدرالي. يسمح هذه الطريقة للبنوك ببناء نموذج قوي للكشف عن الاحتيال بشكل تعاوني من خلال تجميع التحديثات المحسوبة محليًا دون مشاركة بيانات العملاء الحساسة، مما يحافظ على الخصوصية. بالإضافة إلى ذلك، تؤكد الورقة على مشكلة توزيع الفئات غير المتوازن، حيث تشكل المعاملات الاحتيالية أقل من 2% من إجمالي المعاملات، مما يعقد تحديد أنماط الاحتيال ويؤثر سلبًا على أداء المصنفات.

لتخفيف هذه التحديات، يقترح المؤلفون تقنيات إعادة أخذ العينات المختلفة، بما في ذلك طرق زيادة العينات مثل تقنية زيادة العينات الأقلية الاصطناعية (SMOTE) والعينة الاصطناعية التكيفية (AdaSyn)، بالإضافة إلى طرق تقليل العينات مثل تقليل العينات العشوائية (RUS). تقيم الدراسة فعالية هذه الاستراتيجيات في إعادة أخذ العينات بالتزامن مع عدة خوارزميات للتعلم الآلي والتعلم العميق، بما في ذلك الغابة العشوائية (RF)، وأشجار القرار (DT)، والشبكات العصبية التلافيفية (CNN). تشير النتائج إلى أن تقنيات إعادة أخذ العينات الهجينة يمكن أن تعزز بشكل كبير دقة التصنيف. علاوة على ذلك، توضح الورقة تطوير نموذج تعلم فيدرالي يدمج هذه الطرق في إعادة أخذ العينات، مما يمكّن المؤسسات المالية من الاستفادة من نموذج عالمي مشترك يحسن قدرات الكشف عن الاحتيال مع الحفاظ على أمان البيانات. تؤكد النتائج على أهمية معالجة عدم توازن البيانات ومخاوف الخصوصية في تقدم أنظمة الكشف عن احتيال بطاقات الائتمان.

Journal: Neural Computing and Applications, Volume: 36, Issue: 11
DOI: https://doi.org/10.1007/s00521-023-09410-2
Publication Date: 2024-01-20
Author(s): Mustafa Abdul Salam et al.
Primary Topic: Imbalanced Data Classification Techniques

Overview

The research paper addresses the pressing issue of credit card transaction fraud, which has led to significant financial losses for both consumers and banks. It emphasizes the necessity for robust fraud detection systems, particularly through the lens of credit card fraud detection (CCFD). The study identifies key challenges in developing effective fraud detection systems, notably the restrictions on data sharing among financial institutions due to security and privacy concerns, which hinder traditional fraud detection methods. To overcome these challenges, the paper proposes the use of federated learning frameworks, specifically TensorFlow Federated and PyTorch, for CCFD.

Additionally, the paper highlights the critical problem of class imbalance in credit card transactions, where fraudulent transactions constitute a small fraction of the total. To tackle this, the study conducts a comparative analysis of various individual and hybrid resampling techniques aimed at balancing the dataset. Experimental results indicate that hybrid resampling methods yield superior performance for machine learning classifiers, such as Random Forest (RF), Logistic Regression, K-Nearest Neighbors (KNN), Decision Tree (DT), and Gaussian Naive Bayes (NB), achieving accuracies of 99.99%, 94.61%, 99.96%, 99.98%, and 91.47%, respectively. Notably, the RF classifier demonstrates the highest performance across multiple metrics, including accuracy, recall, precision, and F-score. Furthermore, while the PyTorch framework offers higher prediction accuracy for the federated learning model compared to TensorFlow, it requires more computational resources.

Introduction

The introduction highlights the rapid increase in credit card transactions driven by advancements in electronic services, such as e-commerce and mobile payments. This surge has led to significant financial losses due to credit card fraud, which amounted to approximately $18.39 billion globally in 2018, up from $14.99 billion in 2017. Projections indicate that total payment card volume will reach $57.08 trillion by 2023, with gross card fraud expected to escalate to $35.67 billion, potentially reaching $40 billion by 2027.

The paper emphasizes the dual nature of fraudulent activities, which can stem from stolen cards or false information. Various researchers have explored credit card fraud detection, primarily employing supervised machine learning models to differentiate between fraudulent and legitimate transactions. Despite these advancements, the development of an effective fraud detection system faces significant challenges, including insufficient datasets and skewed distribution of transaction data.

Methods

In this section, the authors investigate the effects of various individual and hybrid resampling strategies on a dataset used for credit card fraud detection. They employ multiple classification algorithms, including Decision Trees (DT), Gaussian Naive Bayes (GaussianNB), Random Forest (RF), K-Nearest Neighbors (KNN), and Logistic Regression, to assess performance outcomes. A federated learning model is implemented across different frameworks to address data security and privacy concerns.

The experiments were conducted using Python (version 3) and leveraged several open-source libraries, including Scikit-learn (1.1.3), pandas (1.4.4), NumPy (1.22.3), matplotlib (3.5.3), TensorFlow Federated (0.17.0), PyTorch (1.2.0), and Imblearn (0.9.1). The computational environment consisted of a desktop computer equipped with an Intel Core i7 1.80 GHz CPU, 16GB of RAM, and a Windows 10 64-bit operating system.

Results

In the Results section, the study evaluates the performance of various classifiers, including traditional machine learning techniques and a Convolutional Neural Network (CNN), utilizing metrics such as precision, recall, accuracy, loss, F1-measure, and total computational time. The effectiveness of these classifiers is assessed in conjunction with resampling techniques to enhance their performance.

The comparative analysis employed an 80:20 training-testing ratio, revealing significant insights into the classifiers’ accuracy. The findings underscore the effectiveness of the classifiers in the context of the applied resampling methods, although specific numerical results and detailed comparisons are not provided in this summary.

Discussion

The discussion section of the research paper addresses critical challenges in credit card fraud detection systems (CCFDS), particularly focusing on dataset insufficiency and class imbalance. The authors highlight that the lack of publicly available datasets, primarily due to data privacy concerns among banks, necessitates the use of a federated learning approach. This method allows banks to collaboratively build a robust fraud detection model by aggregating locally computed updates without sharing sensitive client data, thus preserving privacy. Additionally, the paper emphasizes the issue of skewed class distribution, where fraudulent transactions constitute less than 2% of total transactions, complicating the identification of fraud patterns and adversely affecting classifier performance.

To mitigate these challenges, the authors propose various resampling techniques, including oversampling methods like the Synthetic Minority Over-sampling Technique (SMOTE) and Adaptive Synthetic Sampling (AdaSyn), as well as undersampling methods such as Random Undersampling (RUS). The study evaluates the effectiveness of these resampling strategies in conjunction with multiple machine learning and deep learning algorithms, including Random Forest (RF), Decision Trees (DT), and Convolutional Neural Networks (CNN). The results indicate that hybrid resampling techniques can significantly enhance classification accuracy. Furthermore, the paper outlines the development of a federated learning model that integrates these resampling methods, enabling financial institutions to benefit from a shared global model that improves fraud detection capabilities while maintaining data security. The findings underscore the importance of addressing data imbalance and privacy concerns in the advancement of CCFDS.