DOI: https://doi.org/10.3389/frobt.2025.1697155
PMID: https://pubmed.ncbi.nlm.nih.gov/41694101
تاريخ النشر: 2026-01-30
المؤلف: Sirui Song وآخرون
الموضوع الرئيسي: تعلم التعزيز في الروبوتات
نظرة عامة
تقدم هذه البحث إطار عمل جديد مستوحى من الفيزياء للتعلم العميق المعزز (DRL) يهدف إلى تعزيز تنقل الروبوتات المتنقلة في البيئات المعقدة. يقوم الإطار بنمذجة الروبوت، وجهته، والعقبات كأحمال كهربائية، مستفيدًا من قوى كولومب لإبلاغ دالة المكافأة التي تقدم إشارات جذابة وطاردة. لا يوجه هذا النهج المبتكر سلوك الروبوت بفعالية فحسب، بل يتضمن أيضًا مكافآت استباقية مستمدة من تقسيم LiDAR للمساعدة في تجنب الاصطدامات بشكل استباقي. تم تدريب النموذج في بيئات محاكاة Gazebo وتم نشره بنجاح على TurtleBot v3، مما أظهر تحسينات كبيرة في تقليل الاصطدامات، والحفاظ على مسافة آمنة من العقبات، وسلاسة المسار في كل من السيناريوهات المحاكية والواقعية.
تشير النتائج إلى أن هذه الآلية القائمة على كولومب للمكافآت هي الأولى من نوعها المطبقة على تخطيط المسار والتعلم المعزز، مما يوفر إطار عمل واضح وقابل للتفسير لتخطيط حركة الروبوت. يسمح دمج حدود العقبات بتجنب الاصطدامات المتقدمة، بينما يعزز تصميم المكونات غير المتغيرة للبيئة قدرات تعميم النظام. بشكل عام، يضع الإطار المقترح أساسًا واعدًا للتطبيقات المستقبلية في أنظمة الروبوتات المتعددة، والبيئات المعقدة، والتخطيط التكيفي في الوقت الحقيقي، مما يظهر كل من الفعالية والقدرة القوية على التفسير في تشكيل سلوك الروبوت.
مقدمة
تناقش مقدمة هذه الورقة البحثية التقدم في تنقل الروبوتات المتنقلة باستخدام تقنية LiDAR، مع تسليط الضوء على مزاياها مقارنة بالأنظمة التقليدية المعتمدة على الخرائط العالمية. يمكّن LiDAR من إنشاء خرائط ثلاثية الأبعاد مفصلة في الوقت الحقيقي، والتي تعتبر ضرورية لاتخاذ قرارات فعالة في البيئات الديناميكية. تستعرض الورقة تقنيات تخطيط الحركة المختلفة، بما في ذلك الطرق الهندسية، المعتمدة على العينات، وطرق التحسين، بالإضافة إلى الاتجاه المتزايد لأساليب التعلم الآلي (ML). بشكل خاص، تقارن بين التعلم المراقب، الذي يتنبأ بسياسات التحكم مباشرة من بيانات المستشعر، والتعلم المعزز (RL)، الذي يسمح للروبوتات بتعلم استراتيجيات التنقل المثلى من خلال التفاعل مع بيئتها.
يحدد المؤلفون القيود في طرق التعلم العميق المعزز (DRL) الحالية، لا سيما فيما يتعلق بهياكل المكافآت التي غالبًا ما تفتقر إلى الأساس الفيزيائي ودمج تقنيات تخطيط المسار الكلاسيكية. لمعالجة هذه التحديات، تقترح الورقة خوارزمية جديدة لتخطيط الحركة المستندة إلى DRL مستوحاة من الفيزياء تستخدم قوى كولومب لنمذجة التفاعلات بين الروبوت، وجهته، والعقبات. يهدف هذا النهج إلى تحسين تجنب الاصطدامات وقدرات الوصول إلى الأهداف مع تعزيز التعميم عبر بيئات متنوعة. يتم التحقق من الطريقة المقترحة من خلال المحاكاة والتجارب الواقعية، مما يوضح فعاليتها وقابلية تفسير هياكل المكافآت الخاصة بها. تشمل المساهمات الرئيسية الاستخدام المبتكر لقوى كولومب في تخطيط المسار، ودمج تقسيم العقبات المعتمد على LiDAR، وتصميم مكونات غير متغيرة للبيئة لتعزيز التعميم.
الطرق
تحدد قسم الورقة البحثية الأساليب المستخدمة في تخطيط حركة الروبوت، مميزًا بين الطرق الكلاسيكية والطرق المعتمدة على التعلم المعزز (RL). تستخدم الطرق الكلاسيكية خطوط أنابيب منظمة للتخطيط العالمي، وتجنب العقبات المحلية، وتوليد المسارات، معتمدة على خوارزميات مثل A*، RRT*، وDWA. بينما تكون فعالة في البيئات الثابتة، تواجه هذه الطرق تحديات في الإعدادات الديناميكية بسبب اعتمادها على قواعد صارمة وبناء خرائط واسعة. لقد حسنت التطورات الأخيرة من تجنب الاصطدامات من خلال خوارزميات مصقولة، ومع ذلك لا تزال الطرق الكلاسيكية تعاني من قابلية التطبيق في العالم الحقيقي عند التعامل مع العقبات الديناميكية.
بالمقابل، تشمل طرق تخطيط الحركة المعتمدة على RL استراتيجيات هجينة، من النهاية إلى النهاية، ومتعددة الروبوتات. تجمع الطرق الهجينة بين RL والتقنيات الكلاسيكية لتعزيز التكيف، بينما تتعلم الطرق من النهاية إلى النهاية مباشرة من بيانات المستشعر، مما يظهر أداءً متفوقًا في البيئات غير المرئية. ومع ذلك، تعاني طرق RL من متطلبات تدريب عالية وقدرات تعميم محدودة. تستفيد طرق الروبوتات المتعددة من التدريب المركزي مع التنفيذ اللامركزي، مما يظهر وعدًا في التنسيق ولكنه يواجه تحديات تتعلق بتعقيد التدريب والتفاعلات الديناميكية.
تركز الدراسة بشكل خاص على روبوت متنقل يتنقل من نقطة بداية محددة إلى وجهة مستهدفة مع تجنب العقبات الثابتة والديناميكية. باستخدام إطار عمل DRL مستوحى من الفيزياء، يتم تدريب الروبوت في بيئة محاكاة Gazebo قبل نشره على TurtleBot3 (TB3) المزود بمستشعرات مثل LiDAR بزاوية 360° ووحدة قياس القصور الذاتي (IMU). يبرز إعداد التجربة قدرة الروبوت على تعديل مساره في الوقت الحقيقي لتجنب العقبات، مع التحكم في الحركة من خلال مكونات السرعة الخطية والزوايا.
النتائج
في بيئات المحاكاة، تم تدريب خمسة نماذج على مدى 7,000 دورة، مع رسم مقاييس الأداء في الشكل 10. تشير النتائج إلى أن النموذج W obst أظهر أبطأ تقارب، حيث استقر فقط بعد حوالي 4,800 دورة، وأظهر أداءً غير متسق بعد ذلك. بالمقابل، تقاربت النماذج W C وW C+obst بشكل أسرع، محققة أداءً مستقرًا حول 800 دورة، بينما أظهرت W C+v وW C+v+obst تقاربًا أسرع حتى حوالي 400 دورة. تشير هذه النتائج إلى أن تضمين مصطلح R كولومب يعزز بشكل كبير من كفاءة التدريب، مع تسريع مصطلح R الرؤية من التقارب.
لتقييم قوة وعمومية إطار التعلم العميق المعزز (DRL) المقترح، تم اختبار بيئتين مزدحمتين إضافيتين. تكشف النتائج الملخصة في الجدول 3 عن انخفاض ملحوظ في الأداء عبر جميع النماذج مع زيادة تعقيد البيئة، ويرجع ذلك أساسًا إلى تقليل المساحة القابلة للتنقل واحتمالية احتجاز الوكلاء في الحد الأدنى المحلي. ومع ذلك، حقق النموذج W C+v الموجه بواسطة كولومب والرؤية باستمرار أعلى معدلات نجاح، وأدنى نسب اصطدام، وأقصر متوسط مسافات الأهداف في كلا الإعدادين المزدحمين. حتى في وجود عقبات ديناميكية، حافظ W C+v على أداء متفوق، مما يؤكد أن المكافأة القائمة على كولومب تُحاكي بشكل فعال قوى العقبات والأهداف، بينما تعزز مدخلات الرؤية الوعي المكاني والقدرة على التكيف. تؤكد هذه الاتجاهات قدرة النموذج على التخطيط لمسارات فعالة في البيئات المعقدة، مما يتناقض بشكل حاد مع الأداء المنخفض لنموذج W obst.
المناقشة
في قسم المناقشة من الورقة البحثية، يستكشف المؤلفون الفروق الأساسية والتشابهات بين القوى الجاذبية وقوى كولومب، مع التأكيد على تعبيراتها الرياضية وآثارها على الروبوتات. تُوصف القوة الجاذبية بموجب قانون نيوتن كـ \( F_{\text{gravity}} = \frac{G m_1 m_2}{r^2} \)، وهي دائمًا جذابة وتتناقص مع المسافة. بالمقابل، يُعبر عن قانون كولومب كـ \( F_{12} = k_e \frac{|q_1 q_2|}{r^2} \)، الذي يحكم التفاعلات الكهروستاتيكية بين الجسيمات المشحونة، والتي يمكن أن تكون جذابة أو طاردة اعتمادًا على إشارات الشحنة. يبرز المؤلفون أنه بينما تتبع كلتا القوتين قانون المربع العكسي، فإن القوة الكهروستاتيكية أقوى بكثير على المقاييس الذرية، بينما تهيمن الجاذبية على المسافات الفلكية.
تقدم الورقة خوارزمية الحقل المحتمل الاصطناعي (APF) كنهج كلاسيكي لتخطيط المسار في الوقت الحقيقي، مع الإشارة إلى قيودها، مثل الحساسية لتعديل المعلمات والضعف أمام الحد الأدنى المحلي. لمعالجة هذه القضايا، يقترح المؤلفون إطار عمل جديد للتعلم العميق المعزز (DRL) يستخدم قوى كولومب لإنشاء حقل توجيه حركة سلس ومتسق عالميًا. يتضمن هذا الإطار مكافآت تعتمد على التفاعلات الكهروستاتيكية، مما يعزز قدرات تنقل الروبوت مع تجنب العقبات. يؤكد المؤلفون على أهمية تعميم النموذج في DRL، لا سيما لتوجيه الروبوتات في بيئات متنوعة، ويناقشون تصميم المكافآت المبتكر الخاص بهم، الذي يجمع بين القوى الجذابة والطاردة لتحسين استقرار التعلم وسرعة التقارب. بشكل عام، تقدم البحث تقدمًا كبيرًا في تنقل الروبوتات من خلال دمج المبادئ المستندة إلى الفيزياء في إطار التعلم، مما يظهر كفاءة محسنة وموثوقية في تخطيط المسار وتجنب العقبات.
DOI: https://doi.org/10.3389/frobt.2025.1697155
PMID: https://pubmed.ncbi.nlm.nih.gov/41694101
Publication Date: 2026-01-30
Author(s): Sirui Song et al.
Primary Topic: Reinforcement Learning in Robotics
Overview
This research presents a novel physics-inspired deep reinforcement learning (DRL) framework aimed at enhancing mobile robot navigation in complex environments. The framework models the robot, its destination, and obstacles as electrical charges, utilizing Coulomb forces to inform a reward function that delivers both attractive and repulsive signals. This innovative approach not only guides the robot’s behavior effectively but also incorporates anticipatory rewards derived from LiDAR segmentation to help avoid collisions proactively. The model was trained in Gazebo simulation environments and successfully deployed on a TurtleBot v3, demonstrating significant improvements in collision reduction, safe distance maintenance from obstacles, and trajectory smoothness in both simulated and real-world scenarios.
The findings indicate that this Coulomb-based reward mechanism is the first of its kind applied to path planning and reinforcement learning, providing a clear and interpretable framework for robot motion planning. The integration of obstacle boundaries allows for advanced collision avoidance, while the design of environment-invariant components enhances the system’s generalization capabilities. Overall, the proposed framework lays a promising groundwork for future applications in multirobot systems, complex environments, and real-time adaptive planning, showcasing both effectiveness and strong explainability in shaping robot behavior.
Introduction
The introduction of this research paper discusses the advancements in mobile robot navigation using LiDAR technology, highlighting its advantages over traditional global-map-based systems. LiDAR enables the creation of real-time, detailed 3D maps, which are essential for effective decision-making in dynamic environments. The paper reviews various motion planning techniques, including geometric, sampling-based, and optimization methods, as well as the growing trend of machine learning (ML) approaches. Specifically, it contrasts supervised learning, which predicts control policies directly from sensor data, with reinforcement learning (RL), which allows robots to learn optimal navigation strategies through interaction with their environment.
The authors identify limitations in existing deep reinforcement learning (DRL) methods, particularly regarding reward structures that often lack physical grounding and the integration of classical path-planning techniques. To address these challenges, the paper proposes a novel physics-inspired DRL-based motion planning algorithm that utilizes Coulomb forces to model interactions between the robot, its destination, and obstacles. This approach aims to improve collision avoidance and goal-reaching capabilities while enhancing generalization across various environments. The proposed method is validated through simulations and real-world experiments, demonstrating its effectiveness and the interpretability of its reward structures. Key contributions include the innovative use of Coulomb forces in path planning, the incorporation of LiDAR-based obstacle segmentation, and the design of environment-invariant components to bolster generalization.
Methods
The research paper section outlines the methodologies employed in robot motion planning, distinguishing between classical and reinforcement learning (RL)-based approaches. Classical methods utilize structured pipelines for global planning, local obstacle avoidance, and trajectory generation, relying on algorithms like A*, RRT*, and DWA. While effective in static environments, these methods face challenges in dynamic settings due to their reliance on rigid rules and extensive map construction. Recent advancements have improved collision avoidance through refined algorithms, yet classical methods still struggle with real-world applicability when dealing with dynamic obstacles.
In contrast, RL-based motion planning encompasses hybrid, end-to-end, and multi-robot strategies. Hybrid methods combine RL with classical techniques to enhance adaptability, while end-to-end approaches learn directly from sensor data, demonstrating superior performance in unseen environments. However, RL methods are hindered by high training demands and limited generalization capabilities. Multi-robot approaches leverage centralized training with decentralized execution, showing promise in coordination but facing challenges related to training complexity and dynamic interactions.
The study specifically focuses on a mobile robot navigating from a designated start to a target destination while avoiding static and dynamic obstacles. Utilizing a physics-inspired DRL framework, the robot is trained in a Gazebo simulation environment before deployment on a TurtleBot3 (TB3) equipped with sensors such as a 360° LiDAR and an Inertial Measurement Unit (IMU). The experimental setup emphasizes the robot’s ability to adjust its trajectory in real-time to avoid obstacles, with motion controlled by linear and angular velocity components.
Results
In the simulation environments, five models were trained over 7,000 epochs, with performance metrics plotted in Figure 10. The results indicate that the model W obst exhibited the slowest convergence, stabilizing only after approximately 4,800 epochs, and demonstrated inconsistent performance thereafter. In contrast, models W C and W C+obst converged more rapidly, achieving stable performance around 800 epochs, while W C+v and W C+v+obst showed even faster convergence at approximately 400 epochs. These findings suggest that the inclusion of the R Coulomb term significantly enhances training efficiency, with the R vision term further accelerating convergence.
To assess the robustness and generalization of the proposed deep reinforcement learning (DRL) framework, two additional cluttered environments were tested. The results summarized in Table 3 reveal a notable decline in performance across all models as environmental complexity increased, primarily due to reduced navigable space and the likelihood of agents becoming trapped in local minima. However, the Coulomb-and vision-guided model W C+v consistently achieved the highest success rates, lowest collision ratios, and shortest average goal distances in both cluttered settings. Even in the presence of dynamic obstacles, W C+v maintained superior performance, confirming that the Coulomb-based reward effectively models obstacle and goal forces, while the vision input enhances spatial awareness and adaptability. This trend underscores the model’s capability for efficient path planning in complex environments, contrasting sharply with the diminished performance of the W obst model.
Discussion
In the discussion section of the research paper, the authors explore the fundamental differences and similarities between gravitational and Coulomb forces, emphasizing their mathematical expressions and implications for robotics. Gravitational force, described by Newton’s law as \( F_{\text{gravity}} = \frac{G m_1 m_2}{r^2} \), is always attractive and diminishes with distance. In contrast, Coulomb’s law, expressed as \( F_{12} = k_e \frac{|q_1 q_2|}{r^2} \), governs the electrostatic interactions between charged particles, which can be either attractive or repulsive depending on the charge signs. The authors highlight that while both forces follow an inverse-square law, the electrostatic force is significantly stronger at atomic scales, whereas gravity dominates at astronomical distances.
The paper introduces the Artificial Potential Field (APF) algorithm as a classical approach for real-time path planning, noting its limitations, such as sensitivity to parameter tuning and susceptibility to local minima. To address these issues, the authors propose a novel Deep Reinforcement Learning (DRL) framework that utilizes Coulomb forces to create a smooth, globally consistent motion-guidance field. This framework incorporates rewards based on Coulomb interactions, enhancing the robot’s navigation capabilities while avoiding obstacles. The authors emphasize the importance of model generalization in DRL, particularly for robotic navigation in diverse environments, and discuss their innovative reward design, which combines attractive and repulsive forces to improve learning stability and convergence speed. Overall, the research presents a significant advancement in robot navigation by integrating physics-based principles into a learning framework, demonstrating improved efficiency and robustness in path planning and obstacle avoidance.
