DOI: https://doi.org/10.1109/lra.2026.3662620
تاريخ النشر: 2026-02-09
المؤلف: Jiyue Tao وآخرون
الموضوع الرئيسي: الصلابة ضد الهجمات في تعلم الآلة
نظرة عامة
تتعلق مشكلة الدفاع عن الهدف (TDP) للمركبات السطحية غير المأهولة (USVs) باعتراض مركبة سطحية غير مأهولة معادية قبل دخولها منطقة الهدف المحددة، خاصة عندما يمتلك المهاجم قدرة مناورة متفوقة. لمعالجة هذا التحدي، يقدم المؤلفون ARBoids، وهو إطار تعلم تعزيز متبقي تكيفي يجمع بين التعلم العميق المعزز (DRL) ونموذج Boids المستوحى من البيولوجيا. في هذا الإطار، يعمل نموذج Boids كخط أساس فعال حسابيًا لتنسيق الوكلاء المتعددين، بينما يتم استخدام DRL لتعلم سياسة متبقية تعمل على تحسين إجراءات المدافعين. تظهر التحقق في محاكاة Gazebo عالية الدقة أن ARBoids يتفوق على استراتيجيات الاعتراض التقليدية، بما في ذلك الطرق المعتمدة على القوة البحتة وسياسات DRL القياسية، ويظهر تكيفًا قويًا مع ملفات تعريف مناورة المهاجمين المختلفة.
في الختام، تقدم الدراسة ARBoids كاستراتيجية تحكم مبتكرة للدفاع ضد المركبات السطحية غير المأهولة المهاجمة ذات القدرة العالية على المناورة، حيث تدمج خوارزمية Soft Actor-Critic (SAC) مع نموذج Boids. لا تعزز هذه الطريقة التنسيق من خلال سياسة خط الأساس فحسب، بل تعمل أيضًا على تحسين إجراءات المدافعين من خلال سياسة متبقية تم تعلمها، مدعومة بدالة مكافأة مصممة خصيصًا وتعلم منهجي لتحسين استقرار التدريب والتعميم. تؤكد التجارب الواسعة على الأداء المتفوق لـ ARBoids مقارنة بالطرق المرجعية. ستركز الأبحاث المستقبلية على تعزيز تعميم ARBoids ضد استراتيجيات الهجوم المتنوعة، وتطوير تقنيات متقدمة لتجنب الاصطدام، والتحقق من النهج في سيناريوهات العالم الحقيقي.
مقدمة
تناقش مقدمة هذه الورقة البحثية الاهتمام المتزايد بالمركبات السطحية غير المأهولة (USVs) داخل مجتمع الروبوتات البحرية، مع التركيز بشكل خاص على مشكلة الدفاع عن الهدف (TDP). تتضمن TDP نشر مركبات سطحية غير مأهولة للدفاع لاعتراض مركبة سطحية غير مأهولة مهاجمة قبل أن تخترق منطقة محددة، وهو أمر حاسم لتطبيقات مثل المراقبة البحرية وأمن الحدود. تنشأ تعقيدات TDP من طبيعتها الديناميكية والمعادية، مما يتطلب اتخاذ قرارات في الوقت الفعلي من المدافعين الذين يجب عليهم توقع مناورات المهاجم المراوغة. بينما حققت الدراسات السابقة تقدمًا، فإنها غالبًا ما تفترض أن المدافعين لديهم قدرة مناورة متساوية أو متفوقة، مما يحد من قابليتها للتطبيق ضد المهاجمين ذوي القدرة العالية على المناورة.
لمعالجة هذه التحديات، تقدم الورقة ARBoids، وهو إطار تعلم تعزيز متبقي (RL) تكيفي يجمع بين التعلم العميق المعزز (DRL) ونموذج Boids المستوحى من البيولوجيا. يستخدم هذا النهج الهجين نموذج Boids كخط أساس فعال حسابيًا للتنسيق بينما يسمح لـ DRL بتنقيح إجراءات المدافعين استجابةً للسيناريوهات المعقدة. يتم التحقق من الإطار المقترح من خلال تجارب واسعة في بيئة محاكاة بحرية تعتمد على Gazebo عالية الدقة، مما يظهر تحسينات كبيرة في معدلات نجاح الاعتراض والتكيف مقارنة بالاستراتيجيات المرجعية. تشمل المساهمات الرئيسية تطوير حل جديد لـ TDP، ووحدة محول لضبط الوزن الديناميكي بين مكونات DRL وBoids، ودليل على الأداء المتفوق وقابلية التعميم عبر سيناريوهات متنوعة.
طرق
في هذا القسم، يوضح المؤلفون إعداد التجارب والنتائج لتقييم إطار ARBoids، الذي يدمج التعلم المعزز والاستراتيجيات التعاونية للدفاع ضد مركبة سطحية غير مأهولة مهاجمة (USV). تشمل معلمات المحاكاة نصف قطر الالتقاط 5 م، ونطاقات الهدف والاستشعار المحددة عند 15 م و60 م، على التوالي، ووقت محاكاة إجمالي قدره 60 ثانية. تستخدم المركبات السطحية غير المأهولة المدافعة خوارزمية Soft Actor-Critic (SAC) مع معلمات تدريب محددة، بينما يستخدم المهاجم استراتيجية مجال محتمل اصطناعي (APF) محسنة للتنقل نحو هدف مع تجنب المدافعين. تعتمد بيئة التدريب على محاكي عددي مخصص يبسط ديناميات السفن إلى ثلاثة درجات من الحرية، تم تنفيذها بلغة بايثون مع PyTorch.
تظهر النتائج التجريبية فعالية إطار ARBoids في سيناريو معادي محاكى، حيث تحقق معدل نجاح (SR) قدره 85% في التقاط المهاجم عبر 100 تجربة. يتكيف المدافعون بنجاح مع استراتيجياتهم في الوقت الفعلي، مع الحفاظ على التشكيل وضبط المواقع لحجب المهاجم. ومع ذلك، تم نسب بعض الإخفاقات إلى تصادمات بين المدافعين واختراق المهاجم للدفاعات. تكشف تحليل مقارن مع الطرق المرجعية، بما في ذلك نموذج Boids القياسي وسياسة SAC العادية، أن ARBoids يتفوق بشكل كبير على هذه الأساليب، حيث يحقق معدل نجاح 90% خلال 0.35 مليون خطوة تدريب، مقارنة بـ 55% للسياسة المتبقية وركود بالقرب من 20% للـ SAC العادي. تؤكد النتائج على مزايا ARBoids من حيث كفاءة العينة، والاستقرار، والأداء العام، مما يبرز أهمية تصميمه الهجين التكيفي.
مناقشة
في هذا البحث، يتناول المؤلفون مشكلة الدفاع عن الهدف المتعدد للمركبات السطحية غير المأهولة، حيث يجب على فريق من المركبات السطحية غير المأهولة المتجانسة (USVs) حماية منطقة محددة من مركبة سطحية غير مأهولة مهاجمة أكثر مرونة. يتم نمذجة ديناميات حركة المركبات السطحية غير المأهولة المدافعة باستخدام منصة محاكاة عالية الدقة، Virtual RobotX (VRX)، ويتم صياغة استراتيجية التحكم كمشكلة تحسين تهدف إلى زيادة احتمال الدفاع الناجح مع الالتزام بقيود الدفع وتدابير تجنب الاصطدام. يجمع إطار ARBoids المقترح بين خوارزمية Soft Actor-Critic (SAC) ونموذج Boids المستوحى من البيولوجيا، مما يسمح بتعديلات سياسة تكيفية بناءً على حالة البيئة. يتيح هذا الدمج للمدافعين تنسيق إجراءاتهم بفعالية أثناء استجابتهم لمناورات المهاجم.
تظهر النتائج التجريبية أن ARBoids يتفوق بشكل كبير على الطرق التقليدية، حيث يحافظ على معدلات نجاح عالية حتى مع زيادة مرونة المهاجم. يظهر الإطار قوة ضد أحجام الفرق المتنوعة والمهاجمين المعتمدين على التعلم، مما يبرز قابليته للتكيف وفعاليته في السيناريوهات الديناميكية. يبرز المؤلفون أهمية التعلم المنهجي في تعزيز استقرار التدريب والتعميم، ويقترحون عملًا مستقبليًا لتحسين قابلية توسيع النموذج وقدرات تجنب الاصطدام. بشكل عام، تقدم الدراسة نهجًا واعدًا لاستراتيجيات الدفاع الذاتية في البيئات البحرية المعقدة، مع آثار على التطبيقات في العالم الحقيقي.
DOI: https://doi.org/10.1109/lra.2026.3662620
Publication Date: 2026-02-09
Author(s): Jiyue Tao et al.
Primary Topic: Adversarial Robustness in Machine Learning
Overview
The target defense problem (TDP) for unmanned surface vehicles (USVs) involves intercepting an adversarial USV before it enters a designated target area, particularly when the attacker possesses superior maneuverability. To address this challenge, the authors introduce ARBoids, an adaptive residual reinforcement learning framework that combines deep reinforcement learning (DRL) with the biologically inspired Boids model. In this framework, the Boids model acts as a computationally efficient baseline for multi-agent coordination, while DRL is employed to learn a residual policy that optimizes the defenders’ actions. Validation in a high-fidelity Gazebo simulation demonstrates that ARBoids outperforms traditional interception strategies, including pure force-based methods and standard DRL policies, and shows strong adaptability to various attacker maneuverability profiles.
In conclusion, the study presents ARBoids as an innovative control strategy for defending against agile attacking USVs, integrating the Soft Actor-Critic (SAC) algorithm with the Boids model. This approach not only enhances coordination through a baseline policy but also optimizes defender actions via a learned residual policy, supported by a tailored reward function and curriculum learning for improved training stability and generalization. Extensive experiments confirm ARBoids’ superior performance compared to benchmark methods. Future research will focus on enhancing the generalization of ARBoids against diverse attack strategies, developing advanced collision avoidance techniques, and validating the approach in real-world scenarios.
Introduction
The introduction of this research paper discusses the growing interest in unmanned surface vehicles (USVs) within the marine robotics community, particularly focusing on the target defense problem (TDP). The TDP involves deploying defending USVs to intercept an attacking USV before it breaches a designated area, which is crucial for applications like maritime surveillance and border security. The complexity of the TDP arises from its dynamic and adversarial nature, necessitating real-time decision-making by defenders who must anticipate the attacker’s evasive maneuvers. While previous studies have made progress, they often assume that defenders have equal or superior agility, which limits their applicability against highly maneuverable attackers.
To address these challenges, the paper introduces ARBoids, an adaptive residual reinforcement learning (RL) framework that combines deep reinforcement learning (DRL) with the biologically inspired Boids model. This hybrid approach utilizes the Boids model as a computationally efficient baseline for coordination while allowing DRL to refine the defenders’ actions in response to complex scenarios. The proposed framework is validated through extensive experiments in a high-fidelity Gazebo-based marine simulation environment, demonstrating significant improvements in interception success rates and adaptability compared to benchmark strategies. Key contributions include the development of a novel TDP solution, an adapter module for dynamic weight adjustment between the DRL and Boids components, and evidence of superior performance and generalizability across varying scenarios.
Methods
In this section, the authors detail the experimental setup and results for evaluating the ARBoids framework, which integrates reinforcement learning and cooperative strategies for defending against an attacking unmanned surface vehicle (USV). The simulation parameters include a capture radius of 5 m, target and sensing ranges set at 15 m and 60 m, respectively, and a total simulation time of 60 seconds. The defending USVs utilize the Soft Actor-Critic (SAC) algorithm with specific training parameters, while the attacker employs an improved artificial potential field (APF) strategy to navigate towards a target while avoiding defenders. The training environment is based on a custom numerical simulator that simplifies vessel dynamics to three degrees of freedom, implemented in Python with PyTorch.
The experimental results demonstrate the effectiveness of the ARBoids framework in a simulated adversarial scenario, achieving a success rate (SR) of 85% in capturing the attacker across 100 trials. The defenders successfully adapt their strategies in real-time, maintaining formation and adjusting positions to block the attacker. However, some failures were attributed to inter-defender collisions and the attacker breaching defenses. A comparative analysis with benchmark methods, including the standard Boids model and a vanilla SAC policy, reveals that ARBoids significantly outperforms these approaches, achieving a 90% SR within 0.35 million training steps, compared to 55% for the residual policy and stagnation near 20% for the vanilla SAC. The results underscore the advantages of ARBoids in terms of sample efficiency, stability, and overall performance, highlighting the importance of its adaptive hybrid design.
Discussion
In this research, the authors address the multi-USV target defense problem, where a team of homogeneous unmanned surface vehicles (USVs) must protect a designated area from a more agile attacking USV. The motion dynamics of the defending USVs are modeled using a high-fidelity simulation platform, Virtual RobotX (VRX), and the control strategy is formulated as an optimization problem aimed at maximizing the probability of successful defense while adhering to thrust constraints and collision avoidance measures. The proposed ARBoids framework combines a soft actor-critic (SAC) algorithm with a biologically inspired Boids model, allowing for adaptive policy adjustments based on the state of the environment. This integration enables the defenders to effectively coordinate their actions while responding to the attacker’s maneuvers.
The experimental results demonstrate that ARBoids significantly outperforms traditional methods, maintaining high success rates even as the attacker’s agility increases. The framework exhibits robustness against varying team sizes and learning-based attackers, showcasing its adaptability and effectiveness in dynamic scenarios. The authors highlight the importance of curriculum learning in enhancing training stability and generalization, and they propose future work to further improve the model’s scalability and collision avoidance capabilities. Overall, the study presents a promising approach to autonomous defense strategies in complex marine environments, with implications for real-world applications.
