سمارت كودر-R1: نحو توليد عقود ذكية آمنة وقابلة للتفسير مع تحسين السياسات النسبية المراعية للأمان
SmartCoder-R1: Towards Secure and Explainable Smart Contract Generation with Security-Aware Group Relative Policy Optimization

شارك:
المجلة: Proceedings of the ACM on software engineering.، المجلد: 3
DOI: https://doi.org/10.1145/3808104
تاريخ النشر: 2026-06-30
المؤلف: Lei Yu وآخرون
الموضوع الرئيسي: التكنولوجيا المالية، التمويل الجماعي، المالية الرقمية

نظرة عامة

تقدم ورقة البحث SmartCoder-R1، وهو إطار عمل جديد يهدف إلى تعزيز توليد العقود الذكية الآمنة والقابلة للتفسير باستخدام نماذج اللغة الكبيرة (LLMs). يتناول الإطار التحديات الحرجة المرتبطة بتوليد العقود الذكية بشكل آلي، لا سيما نقص الشفافية وانتشار الثغرات الأمنية في الشيفرة التي تنتجها النماذج الحالية. يستخدم SmartCoder-R1 خط أنابيب من ثلاث مراحل: التدريب المستمر (CPT) لتكييف النموذج مع تعقيدات العقود الذكية، والتعديل الدقيق المراقب لسلسلة الأفكار الطويلة (L-CoT SFT) لغرس التفكير الأمني المنظم، وتحسين سياسة المجموعة الواعية بالأمان (S-GRPO) لتحسين سياسة التوليد بناءً على الامتثال الأمني ومعايير أخرى.

تظهر التقييمات الواسعة ضد 17 نموذجًا متقدمًا أن SmartCoder-R1 يحقق أداءً متفوقًا على معيار يتكون من 756 وظيفة من العالم الحقيقي من 289 عقدًا تم نشرها. يسجل نسبة ComPass تبلغ 87.70%، ونسبة VulRate تبلغ 8.60%، ونسبة SafeAval تبلغ 80.16%، ونسبة FuncRate تبلغ 53.84%، ونسبة FullRate تبلغ 50.53%، مما يمثل تحسينًا بنسبة 45.79% مقارنة بأفضل نموذج أداء، DeepSeek-R1. بالإضافة إلى ذلك، تشير التقييمات البشرية إلى مخرجات تفكير عالية الجودة، مع تقييمات تبلغ 82.7% للوظائف، و85.3% للأمان، و90.7% للوضوح، مما يبرز فعالية الإطار في إنتاج عقود ذكية ليست فقط وظيفية ولكن أيضًا آمنة وقابلة للتفسير.

مقدمة

تسلط مقدمة ورقة البحث هذه الضوء على التبني السريع لتكنولوجيا البلوكشين، مع التأكيد على هيكلها اللامركزي ودور العقود الذكية في إدارة الأصول الرقمية. على الرغم من المزايا، فإن تعقيد العقود الذكية وعدم قابليتها للتغيير يمثلان تحديات أمنية كبيرة، كما يتضح من حادثة DAO، التي أدت إلى خسارة مالية كبيرة بسبب الثغرات في شيفرة العقود الذكية. يحدد المؤلفون فجوة حرجة في نماذج اللغة الكبيرة الحالية (LLMs)، التي غالبًا ما تفتقر إلى فهم عميق لمبادئ الأمان، مما يؤدي إلى توليد شيفرة معيبة يمكن أن تُدخل ثغرات حرجة.

لمعالجة هذه القضايا، يقترح المؤلفون SmartCoder-R1، وهو إطار عمل جديد مصمم لتوليد عقود ذكية آمنة وقابلة للتفسير. يدمج هذا الإطار التدريب المستمر (CPT)، والتعديل الدقيق المراقب لسلسلة الأفكار الطويلة (SFT)، وتحسين سياسة المجموعة الواعية بالأمان (S-GRPO) لتعزيز قدرات التفكير في النموذج وتقليل الثغرات. تظهر النتائج من التجارب الواسعة أن SmartCoder-R1 يحقق نسبة FullRate تبلغ 50.53%، متفوقًا بشكل كبير على النماذج الحالية ومؤسسًا معيارًا جديدًا في توليد العقود الذكية الآمنة. تسهم الورقة بمجموعات بيانات عالية الجودة وتؤكد على أهمية دمج عمليات التفكير في توليد الشيفرة لتحسين الأمان وقابلية التفسير في العقود الذكية.

طرق

في قسم نتائج التجارب، يتم تقييم أداء SmartCoder-R1 مقابل نماذج أساسية مختلفة لمعالجة أسئلة البحث. بالنسبة لسؤال البحث 1 (RQ1)، يظهر SmartCoder-R1 أداءً متفوقًا عبر المقاييس الرئيسية، محققًا درجة قابلية الترجمة (ComPass) تبلغ 87.70%، والتي تتجاوز أفضل نموذج عام (GPT-4.1) بنسبة 4.24% وأفضل نموذج متخصص بنسبة 6.08%. بالإضافة إلى ذلك، يسجل معدل ثغرات منخفض (VulRate) يبلغ 8.60%، وهو أقل بكثير من المنافسين مثل CodeBC وGPT-4.1. يتفوق النموذج أيضًا في مقياس SafeAval (80.16%) وصحة الوظائف (FuncRate تبلغ 53.84%)، مما يؤدي إلى نسبة FullRate تبلغ 50.53%، وهو تحسين بنسبة 45.79% مقارنة بأفضل نتيجة تالية.

بالنسبة لسؤال البحث 2 (RQ2)، تكشف دراسة الإزالة عن مساهمات المكونات الأساسية: التدريب المستمر (CPT)، والتعديل الدقيق المراقب لسلسلة الأفكار الطويلة (L-COT SFT)، وتحسين سياسة المجموعة الواعية بالأمان (S-GRPO). تشير التحليلات إلى أنه بينما ينتج L-COT SFT بمفرده نسبة ComPass تبلغ 80.82%، إلا أنه يفشل في ضمان الأمان، كما يتضح من معدل VulRate المرتفع. يحسن إدخال CPT المقاييس ولكنه لا يعالج مخاوف الأمان بشكل كافٍ. تأتي أكبر تحسينات الأمان من S-GRPO، الذي يقلل من VulRate إلى 8.60% ويزيد من SafeAval وFullRate بشكل كبير. تسلط الدراسة أيضًا الضوء على أهمية L-COT SFT في توفير إطار تفكير منظم يساعد في التنقل عبر قيود الأمان، مما يعزز فعالية S-GRPO. توضح تحليل حساسية المعلمات الحاجة إلى نهج متوازن لتحسين التوازن بين الترجمة، والأمان، وصحة الوظائف، مع تحقيق التكوين المختار لأعلى نسبة FullRate.

نقاش

يستعرض قسم النقاش في ورقة البحث التحديات الحرجة في توليد شيفرة العقود الذكية، مؤكدًا على الحاجة إلى صحة وظيفية، وقابلية الترجمة، والأمان. يجادل المؤلفون بأن نماذج توليد الشيفرة الحالية، مثل Qwen2.5-Coder، غالبًا ما تتجاهل ضرورة التفكير المنظم في الأمان، مما يؤدي إلى ثغرات يمكن أن تؤدي إلى خسائر مالية كبيرة. يبرزون دافعين رئيسيين لعملهم: نقص التفكير الأمني الصريح في النماذج الحالية، مما يترك المطورين مع شيفرة غير قابلة للتحقق، وميول هذه النماذج لتوليد ثغرات حرجة بسبب فهم سطحي لمبادئ العقود الذكية.

لمعالجة هذه القضايا، يقترح المؤلفون نهجًا معززًا بالتفكير لا يولد فقط وظائف العقود الذكية ولكن أيضًا يوفر سلسلة تفكير شفافة تبرر خيارات التصميم. يهدف هذا النهج إلى ضمان أن الشيفرة المولدة تلتزم بأفضل ممارسات الأمان، مثل حماية إعادة الدخول والتحكم المناسب في الوصول. يوضح المؤلفون قيود النماذج الحالية من خلال مقارنتها بـ SmartCoder-R1 المقترح، الذي يظهر أداءً متفوقًا في توليد شيفرة آمنة ووظيفية، كما يتضح من البيانات التجريبية التي تظهر فجوة أداء كبيرة في مقاييس الأمان والوظائف. يبرز هذا الحاجة إلى نماذج قادرة على التفكير الأمني الصريح في مجال تطوير العقود الذكية الذي يتسم بالمخاطر العالية.

Journal: Proceedings of the ACM on software engineering., Volume: 3
DOI: https://doi.org/10.1145/3808104
Publication Date: 2026-06-30
Author(s): Lei Yu et al.
Primary Topic: FinTech, Crowdfunding, Digital Finance

Overview

The research paper presents SmartCoder-R1, a novel framework aimed at enhancing the generation of secure and explainable smart contracts using Large Language Models (LLMs). The framework addresses critical challenges associated with automated smart contract generation, particularly the lack of transparency and the prevalence of security vulnerabilities in code produced by existing models. SmartCoder-R1 employs a three-stage pipeline: Continual Pre-Training (CPT) to tailor the model to smart contract intricacies, Long Chain-of-Thought Supervised Fine-Tuning (L-CoT SFT) for instilling structured security reasoning, and Security-Aware Group Relative Policy Optimization (S-GRPO) to refine the generation policy based on security compliance and other metrics.

Extensive evaluations against 17 state-of-the-art baselines demonstrate that SmartCoder-R1 achieves superior performance on a benchmark comprising 756 real-world functions from 289 deployed contracts. It records a ComPass of 87.70%, a VulRate of 8.60%, a SafeAval of 80.16%, a FuncRate of 53.84%, and a FullRate of 50.53%, marking a 45.79% improvement over the best-performing baseline, DeepSeek-R1. Additionally, human evaluations indicate high-quality reasoning outputs, with ratings of 82.7% for Functionality, 85.3% for Security, and 90.7% for Clarity, underscoring the framework’s effectiveness in producing not only functional but also secure and interpretable smart contracts.

Introduction

The introduction of this research paper highlights the rapid adoption of blockchain technology, emphasizing its decentralized architecture and the role of smart contracts in managing digital assets. Despite the advantages, the complexity and immutability of smart contracts present significant security challenges, as illustrated by the DAO incident, which resulted in a substantial financial loss due to vulnerabilities in smart contract code. The authors identify a critical gap in existing Code Large Language Models (LLMs), which often lack a deep understanding of security principles, leading to flawed code generation that can introduce critical vulnerabilities.

To address these issues, the authors propose SmartCoder-R1, a novel framework designed to generate secure and explainable smart contracts. This framework integrates Continual Pre-Training (CPT), Long Chain-of-Thought Supervised Fine-Tuning (SFT), and Security-Aware Group Relative Policy Optimization (S-GRPO) to enhance the model’s reasoning capabilities and minimize vulnerabilities. The results from extensive experiments demonstrate that SmartCoder-R1 achieves a FullRate of 50.53%, significantly outperforming existing models and establishing a new state of the art in secure smart contract generation. The paper contributes high-quality datasets and emphasizes the importance of integrating reasoning processes into code generation to improve security and explainability in smart contracts.

Methods

In the experimental results section, the performance of SmartCoder-R1 is evaluated against various baselines to address the research questions. For Research Question 1 (RQ1), SmartCoder-R1 demonstrates superior performance across key metrics, achieving a compilability score (ComPass) of 87.70%, which surpasses the best general-purpose model (GPT-4.1) by 4.24% and the best domain-specific model by 6.08%. Additionally, it records a low vulnerability rate (VulRate) of 8.60%, significantly lower than competitors such as CodeBC and GPT-4.1. The model also excels in the SafeAval metric (80.16%) and functional correctness (FuncRate of 53.84%), culminating in a FullRate of 50.53%, which is a 45.79% improvement over the next best result.

For Research Question 2 (RQ2), an ablation study reveals the contributions of core components: Continual Pre-training (CPT), Long Chain-of-Thought Supervised Fine-Tuning (L-COT SFT), and Security-Aware Group Relative Policy Optimization (S-GRPO). The analysis indicates that while L-COT SFT alone yields a ComPass of 80.82%, it fails to ensure security, as evidenced by a high VulRate. The introduction of CPT improves metrics but does not adequately address security concerns. The most significant enhancement in security arises from S-GRPO, which reduces the VulRate to 8.60% and increases SafeAval and FullRate substantially. The study also highlights the importance of L-COT SFT in providing a structured reasoning framework that aids in navigating security constraints, thereby enhancing the effectiveness of S-GRPO. A parameter sensitivity analysis further illustrates the need for a balanced approach to optimize the trade-off between compilation, security, and functional correctness, with the chosen configuration achieving the highest FullRate.

Discussion

The discussion section of the research paper outlines the critical challenges in smart contract code generation, emphasizing the need for functional correctness, compilability, and security. The authors argue that existing code generation models, such as Qwen2.5-Coder, often overlook the necessity for structured reasoning in security, leading to vulnerabilities that can result in significant financial losses. They highlight two primary motivations for their work: the lack of explicit security reasoning in current models, which leaves developers with unverifiable code, and the propensity for these models to generate critical vulnerabilities due to a shallow understanding of smart contract principles.

To address these issues, the authors propose a reasoning-enhanced approach that not only generates smart contract functions but also provides a transparent reasoning chain justifying design choices. This approach aims to ensure that generated code adheres to security best practices, such as reentrancy protection and proper access control. The authors illustrate the limitations of existing models by comparing them to their proposed SmartCoder-R1, which demonstrates superior performance in generating secure and functional code, as evidenced by empirical data showing a significant performance gap in security and functionality metrics. This highlights the necessity for models capable of explicit security reasoning in the high-stakes domain of smart contract development.

شارك: