كشف الاحتيال في الإيثيريوم عبر نموذج لغة المعاملات المشتركة وتعلم تمثيل الرسوم البيانية
Ethereum fraud detection via joint transaction language model and graph representation learning

شارك:
المجلة: Information Fusion، المجلد: 120
DOI: https://doi.org/10.1016/j.inffus.2025.103074
تاريخ النشر: 2025-03-13
المؤلف: Jianguo Sun وآخرون
الموضوع الرئيسي: تقنيات تصنيف البيانات غير المتوازنة

نظرة عامة

تقدم ورقة البحث TLMG4Eth، وهو إطار مبتكر مصمم لتعزيز اكتشاف الاحتيال في Ethereum من خلال دمج نماذج لغة المعاملات مع الأساليب القائمة على الرسوم البيانية. تقنيات الكشف الحالية، بما في ذلك الشبكات العصبية الرسومية ونماذج التسلسل، لا تعالج بشكل كافٍ الفروق الدلالية وأنماط التشابه الكامنة في بيانات المعاملات. يتغلب TLMG4Eth على هذه القيود من خلال تحويل بيانات المعاملات الرقمية إلى جمل ذات معنى من خلال نموذج لغة المعاملات، مما يسهل تعلم الدلالات الصريحة للمعاملات. كما يقدم رسمًا بيانيًا لتشابه سمات المعاملات لالتقاط تشابهات المعاملات ورسمًا بيانيًا لتفاعل الحسابات لتمثيل الديناميات الهيكلية لشبكة المعاملات.

تستخدم المنهجية المقترحة شبكة انتباه متعددة الرؤوس عميقة لدمج التضمينات الدلالية والتشابه، تليها نهج تدريب مشترك يدمج بشكل متآزر شبكة الانتباه مع رسم تفاعل الحسابات. تشير النتائج إلى أن TLMG4Eth يتفوق بشكل كبير على الأساليب الحالية الرائدة، محققًا تحسينات في الأداء تتراوح بين 10% إلى 20% عبر ثلاثة مجموعات بيانات. لا يقتصر هذا العمل على ريادة تطبيق نماذج اللغة لتوضيح دلالات المعاملات فحسب، بل يؤسس أيضًا إطارًا جديدًا لنمذجة تشابه المعاملات، مما يبرز إمكانيات دمج التحليلات اللغوية والدلالية والهيكلية في تحليلات blockchain واكتشاف الاحتيال.

مقدمة

تسلط مقدمة ورقة البحث هذه الضوء على التأثير التحويلي لتكنولوجيا blockchain، وخاصة Ethereum، التي تسهل التطبيقات اللامركزية والعقود الذكية. على الرغم من مزاياها، فإن صعود Ethereum قد جذب أيضًا أنشطة خبيثة، مع زيادة ملحوظة في الاحتيال المرتبط بالعملات المشفرة، كما يتضح من تقرير Chainalysis 2023 حول جرائم العملات المشفرة، الذي أفاد بأن العناوين غير المشروعة تلقت 39.6 مليار دولار من الأصول المشفرة. تستخدم طرق اكتشاف الاحتيال الحالية بشكل أساسي بيانات المعاملات التاريخية من خلال الشبكات العصبية الرسومية (GNNs) ونماذج التسلسل مثل Transformers. ومع ذلك، فإن هذه الأساليب لها قيود كبيرة، بما في ذلك نقص الدلالات المتعلقة بالمعاملات، ونمذجة غير كافية للتشابهات المعاملات، وتحسين غير فعال بين GNNs ونماذج التسلسل.

لمعالجة هذه التحديات، يقترح المؤلفون TLMG4Eth، وهو إطار جديد يدمج نموذج لغة المعاملات (TLM) مع رسومات المعاملات لتعزيز اكتشاف الاحتيال في Ethereum. تتضمن هذه الطريقة تدريب TLM لتحويل بيانات المعاملات الرقمية إلى تمثيلات دلالية، مما يوضح نوايا المعاملات. بالإضافة إلى ذلك، يستخدم الإطار رسمًا بيانيًا لتشابه سمات المعاملات لالتقاط التشابهات الدلالية ورسمًا بيانيًا لتفاعل الحسابات لنمذجة سلوكيات الحسابات. من خلال استخدام شبكة انتباه متعددة الرؤوس لدمج المعلومات الدلالية والتشابه، وتدريبها بشكل مشترك مع رسم تفاعل الحسابات، يظهر TLMG4Eth تحسينات كبيرة في أداء اكتشاف الاحتيال، محققًا درجات F1 تتجاوز الأساليب الرائدة الحالية بنسبة 10%-20% عبر ثلاثة مجموعات بيانات. كما يقدم المؤلفون مجموعة بيانات جديدة لدعم نتائجهم.

طرق

تركز الطرق المناقشة في هذا القسم على أساليب مختلفة لاكتشاف الأنشطة الاحتيالية في شبكات المعاملات، بشكل أساسي من خلال الأساليب القائمة على الرسوم البيانية، والأساليب القائمة على التسلسل، والأساليب الهجينة. تستخدم الأساليب القائمة على الرسوم البيانية شبكات المعاملات وتستخدم خوارزميات مثل Node2Vec وDeep-Walk لاستخراج الميزات والتصنيف. تشمل المساهمات البارزة استخدام التعلم التبايني تحت الرسم البياني لتحديد عناوين التصيد (TGC) وشبكة الالتفاف الرسومية Chebyshev من النهاية إلى النهاية (PDGNN) التي تتضمن أخذ عينات ديناميكية من الرسوم البيانية وتجميع العقد.

بالمقابل، تعالج الأساليب القائمة على التسلسل قيود الأساليب الرسومية في التعامل مع المعاملات عالية التردد من خلال اعتبارها تدفقات أحداث مرتبة زمنياً. يمثل نموذج BERT4ETH هذه الاستراتيجية، مستفيدًا من بنية Transformer لمعالجة أحداث المعاملات الزمنية ويستخدم نموذج لغة مقنع للتدريب المسبق. تجمع الأساليب الهجينة، مثل TSGN وTTAGN، بين ميزات من كل من نماذج الرسوم البيانية ونماذج التسلسل، مدمجة تقنيات مصنوعة يدويًا وتجميع زمني لتعزيز دقة التصنيف. تعكس هذه المنهجيات المتنوعة التطور المستمر في اكتشاف المعاملات الاحتيالية ضمن الشبكات المعقدة.

مناقشة

تقدم قسم المناقشة في الورقة TLMG4Eth، وهو إطار جديد يدمج نماذج لغة المعاملات مع الأساليب القائمة على الرسوم البيانية لتعزيز فهم بيانات معاملات Ethereum. تتكون البنية من نموذج لغة المعاملات (TLM) الذي ينشئ تمثيلات لغوية لبيانات المعاملات الرقمية، مما يسمح باستخراج التضمينات الدلالية. يتم تحقيق ذلك من خلال تحويل سمات المعاملات إلى رموز وصفية والتقاط الفترات الزمنية بين المعاملات لتحسين قابلية التفسير. تستخدم النموذج BERT لاستخراج التضمينات الدلالية وتبني رسمًا بيانيًا لتشابه سمات المعاملات (TASG) لتحديد الارتباطات العالمية للمعاملات، وهو أمر حاسم لاكتشاف الشذوذ مثل أنشطة التصيد.

تشير النتائج إلى أن TLMG4Eth يتفوق بشكل كبير على النماذج الأساسية عبر ثلاثة مجموعات بيانات، محققًا تحسينات في درجات F1 والدقة المتوازنة بنسبة تقارب 10% إلى 20%. تسلط الدراسة الضوء على أهمية دمج رسوم بيانية لتشابه سمات المعاملات، خاصة باستخدام TF-IDF، مما يعزز أداء النموذج من خلال التقاط المعلومات الرئيسية المتعلقة بالحسابات المحتالة بشكل فعال. بالإضافة إلى ذلك، يظهر التدريب المشترك لشبكة الانتباه متعددة الرؤوس ورسم تفاعل الحسابات أنه أكثر فعالية من استخدام أي من النموذجين بشكل مستقل، مع تحقيق الأداء الأمثل عند إعدادات معينة لمعاملات التبادل. بشكل عام، تؤكد النتائج على إمكانيات دمج التحليلات اللغوية والدلالية والهيكلية في تحليلات blockchain واكتشاف الاحتيال.

Journal: Information Fusion, Volume: 120
DOI: https://doi.org/10.1016/j.inffus.2025.103074
Publication Date: 2025-03-13
Author(s): Jianguo Sun et al.
Primary Topic: Imbalanced Data Classification Techniques

Overview

The research paper presents TLMG4Eth, an innovative framework designed to enhance fraud detection in Ethereum by integrating transaction language models with graph-based methods. Current detection techniques, including graph neural networks and sequence models, inadequately address the semantic nuances and similarity patterns inherent in transaction data. TLMG4Eth overcomes these limitations by first converting numerical transaction data into meaningful sentences through a transaction language model, thereby facilitating the learning of explicit transaction semantics. It also introduces a transaction attribute similarity graph to capture transaction similarities and an account interaction graph to represent the structural dynamics of the transaction network.

The proposed methodology employs a deep multi-head attention network to merge semantic and similarity embeddings, followed by a joint training approach that synergistically integrates the attention network with the account interaction graph. The results indicate that TLMG4Eth significantly outperforms existing state-of-the-art methods, achieving performance improvements of 10% to 20% across three datasets. This work not only pioneers the application of language models for clarifying transaction semantics but also establishes a novel framework for modeling transaction similarity, underscoring the potential of combining linguistic, semantic, and structural analyses in blockchain analytics and fraud detection.

Introduction

The introduction of this research paper highlights the transformative impact of blockchain technology, particularly Ethereum, which facilitates decentralized applications and smart contracts. Despite its advantages, the rise of Ethereum has also attracted malicious activities, with a notable increase in crypto-related fraud, as evidenced by the Chainalysis 2023 Crypto Crime Report, which reported illicit addresses receiving $39.6 billion in cryptoassets. Current fraud detection methods primarily utilize historical transaction data through Graph Neural Networks (GNNs) and sequence models like Transformers. However, these approaches have significant limitations, including a lack of transaction semantics, insufficient modeling of transactional similarities, and ineffective synergistic optimization between GNNs and sequence models.

To address these challenges, the authors propose TLMG4Eth, a novel framework that integrates a transaction language model (TLM) with two transaction graphs to enhance Ethereum fraud detection. This approach involves training a TLM to convert numerical transaction data into semantic representations, thereby clarifying transaction intentions. Additionally, the framework employs a transaction attribute similarity graph to capture semantic similarities and an account interaction graph to model account behaviors. By utilizing a multi-head attention network for the fusion of semantic and similarity information, and jointly training it with the account interaction graph, TLMG4Eth demonstrates significant improvements in fraud detection performance, achieving F1-Scores that surpass existing state-of-the-art methods by 10%-20% across three datasets. The authors also introduce a new dataset to support their findings.

Methods

The methods discussed in the section focus on various approaches to detect fraudulent activities in transaction networks, primarily through graph-based, sequence-based, and hybrid methodologies. Graph-based methods utilize transaction networks and employ algorithms such as Node2Vec and Deep-Walk for feature extraction and classification. Notable contributions include the use of subgraph contrastive learning for phishing address identification (TGC) and an end-to-end Chebyshev Graph Convolutional Network (PDGNN) that incorporates dynamic graph sampling and node aggregation.

In contrast, sequence-based methods address the limitations of graph approaches in handling high-frequency transactions by treating them as time-ordered event streams. The BERT4ETH model exemplifies this strategy, leveraging a Transformer architecture to process chronological transaction events and employing a masked language model for pre-training. Hybrid methods, such as TSGN and TTAGN, combine features from both graph and sequence models, integrating handcrafted techniques and temporal aggregation to enhance classification accuracy. These diverse methodologies reflect the ongoing evolution in the detection of fraudulent transactions within complex networks.

Discussion

The discussion section of the paper presents TLMG4Eth, a novel framework that integrates transaction language models with graph-based methods to enhance the understanding of Ethereum transaction data. The architecture consists of a transaction language model (TLM) that creates linguistic representations of numerical transaction data, allowing for the extraction of semantic embeddings. This is achieved by transforming transaction attributes into descriptive tokens and capturing time intervals between transactions to improve interpretability. The model employs BERT for semantic embedding extraction and constructs a Transaction Attribute Similarity Graph (TASG) to identify global transaction correlations, which is crucial for detecting anomalies such as phishing activities.

The results indicate that TLMG4Eth significantly outperforms baseline models across three datasets, achieving improvements in F1-Scores and Balanced Accuracy by approximately 10% to 20%. The study highlights the importance of incorporating transaction attribute similarity graphs, particularly using TF-IDF, which enhances model performance by effectively capturing key information related to phishing accounts. Additionally, the joint training of the multi-head attention network and account interaction graph is shown to be more effective than using either model independently, with optimal performance achieved at specific trade-off parameter settings. Overall, the findings underscore the potential of combining linguistic, semantic, and structural analyses in blockchain analytics and fraud detection.

شارك: