DOI: https://doi.org/10.1371/journal.pone.0340394
PMID: https://pubmed.ncbi.nlm.nih.gov/41632795
تاريخ النشر: 2026-02-03
المؤلف: Liangshuai Liu وآخرون
الموضوع الرئيسي: خوارزميات تخطيط المسار الروبوتي
نظرة عامة
تقدم هذه الورقة نموذج PGD لتخطيط مسار الطائرات بدون طيار (UAV) في التضاريس المعقدة، مما يعالج بفعالية التحديات مثل معالجة الحالات عالية الأبعاد، واستكشاف المسارات العمياء، والقدرة المحدودة على التكيف عبر بيئات مختلفة. يستخدم نموذج PGD تكاملًا تآزريًا بين Transformer، وشبكة الخصومة التوليدية (GAN)، وتدرج السياسة الحتمي العميق (DDPG) لإنشاء نظام مغلق “ضغط-توليد-تحسين”. يقوم نموذج Transformer بضغط بيانات التضاريس عالية الأبعاد، مما يقلل من اختناقات التدريب، بينما تولد GAN مسارات مرشحة عالية الجودة لتقليل الاستكشاف غير الفعال. ثم يقوم DDPG بتحسين استراتيجية تخطيط المسار، مما يؤدي إلى تحسينات كبيرة في الكفاءة والقدرة على التكيف.
تشير النتائج التجريبية إلى أن PGD يتفوق على النماذج الأساسية مثل PPO-DRL وSoft Actor-Critic على مجموعات بيانات UAVDT (الضواحي) وAirSim (الوادي)، محققًا طول مسار يبلغ 20.0م/22.0م، ومعدل تصادم يبلغ 2.5%/3.0%، وكفاءة حسابية تبلغ 13.5ث/16.0ث. يظهر النموذج مزايا ملحوظة في التضاريس عالية التعقيد، مما يبرز قدرته على تعزيز ارتباط الميزات وفرض قيود المسار الفيزيائية من خلال التعاون متعدد الوحدات. ستركز الأبحاث المستقبلية على تحسين قدرة PGD على التكيف مع الظروف الجوية القاسية والسيناريوهات متعددة الوكلاء، مع التركيز على آليات المرشحين الديناميكية وتعزيز الميزات عبر الأنماط.
مقدمة
في سيناريوهات التضاريس المعقدة، مثل فحص الطاقة في الجبال ومراقبة الجيولوجيا في الوادي، ظهرت الطائرات بدون طيار (UAVs) كبدائل حيوية للعمليات اليدوية، خاصة في البيئات عالية المخاطر ومنخفضة الكفاءة. تمكن مرونتها وقدرتها على المناورة من التنقل بفعالية في المناظر الطبيعية الصعبة. ومع ذلك، تقدم هذه البيئات تحديات كبيرة بسبب التضاريس الوعرة ووجود عقبات تتغير ديناميكيًا، مثل الانهيارات الصخرية المفاجئة ونمو النباتات.
نتيجة لذلك، يجب أن تتناول أنظمة تخطيط مسار الطائرات بدون طيار ثلاثة متطلبات حاسمة: ضمان سلامة المسار الخالي من التصادم، وتحسين طول المسار، والتكيف مع الطبيعة غير المتوقعة للتضاريس. هذه العوامل ضرورية لتعزيز الكفاءة التشغيلية وسلامة الطائرات بدون طيار في البيئات المعقدة.
طرق
في هذا القسم، يوضح المؤلفون إعداد التجربة واختيار النموذج الأساسي لتقييم أداء نموذج PGD (توليد المسار والقرار) المقترح. تشمل النماذج الأساسية المختارة مجموعة من تقنيات تخطيط المسار السائدة، بما في ذلك الخوارزميات الميتاheuristic، والتعلم العميق المعزز (DRL)، والأساليب الهجينة. من الجدير بالذكر أن خوارزمية تحسين مستعمرة النمل الهجينة المحسنة (IHSSAO) تعزز قدرات البحث العالمية وتجنب العقبات في التضاريس المعقدة. تعالج متغيرات DRL مثل تدرج السياسة الحتمي العميق المتأخر المزدوج (TD3) وDDPG المعزز كميًا (Q-DDPG) تقدير قيمة Q المبالغ فيه وتعقيد الحسابات، على التوالي، بينما تحسن DDPG المتأخر متعددة النقاط (MCD-DPG) من المتانة في البيئات الصعبة.
بالإضافة إلى ذلك، يتم تسليط الضوء على تحسين السياسة القريب (PPO-DRL) وSoft Actor-Critic (SAC) لكفاءتهما وقدرتهما على التكيف في مهام التحكم المستمرة والسيناريوهات الديناميكية. يُلاحظ أن شبكة Dueling Deep Q (D3QN) تعمل على تحسين تقدير دالة القيمة في فضاءات العمل المنفصلة، وتركز DRL متعددة الوكلاء على المهام التعاونية، مما يختبر قابلية توسيع نموذج PGD. تجمع الاستراتيجيات الهجينة مثل FM-Planner والتحكم التنبؤي النموذجي مع التعلم المعزز (MPC-RL) بين الأساليب التقليدية والمعتمدة على التعلم. تهدف التقييم الشامل عبر هذه النماذج المتنوعة إلى تسليط الضوء على مزايا نموذج PGD من حيث تآزر الميزات وكفاءة التخطيط.
نتائج
تشير نتائج الدراسة إلى اكتشافات مهمة تتعلق بالفرضيات الأساسية. تكشف التحليلات أن المجموعة التجريبية أظهرت تحسنًا ذا دلالة إحصائية في مقاييس الأداء مقارنةً بمجموعة التحكم، مع قيمة p أقل من 0.05. على وجه التحديد، أدت التدخلات إلى زيادة في النتائج المقاسة، والتي تم قياسها باستخدام الانحراف المعياري ومقارنات المتوسط.
علاوة على ذلك، تشير البيانات إلى وجود ارتباط قوي بين المتغيرات التي تم تحليلها، كما يتضح من معامل الارتباط $r = 0.85$. وهذا يشير إلى علاقة قوية تدعم الإطار النظري المقترح في الدراسة. كشفت تحليلات المجموعات الفرعية الإضافية أن التأثيرات كانت ملحوظة بشكل خاص في ديموغرافيات معينة، مما يبرز الإمكانية للتطبيقات المستهدفة للتدخل. بشكل عام، تساهم هذه النتائج في الأدبيات الحالية من خلال تقديم أدلة تجريبية تدعم فعالية المنهجية المقترحة.
مناقشة
في قسم المناقشة من الورقة، يتناول المؤلفون قيود خوارزميات تخطيط المسار التقليدية، مثل A* وDijkstra، خاصة في التضاريس المعقدة وغير المنظمة. تعتمد هذه الخوارزميات على دوال تكلفة مصممة يدويًا، مما قد يؤدي إلى مسارات دون المستوى بسبب عدم قدرتها على التعميم عبر البيئات المتنوعة. تحسن الطرق المعتمدة على العينة مثل RRT* من ذلك من خلال استكشاف الفضاءات عالية الأبعاد ولكنها تواجه صعوبة في الكفاءة مع زيادة تعقيد التضاريس. يبرز المؤلفون إمكانيات تقنيات التعلم العميق، مشيرين إلى أنه بينما تتمتع الشبكات العصبية التلافيفية (CNNs) والشبكات العصبية المتكررة (RNNs) بقوة في استخراج الميزات المحلية ونمذجة التسلسل الزمني، على التوالي، فإنها تظهر أيضًا قيودًا كبيرة في التقاط الاعتماديات بعيدة المدى والتعامل مع التضاريس الديناميكية.
للتغلب على هذه التحديات، يقترح المؤلفون نموذجًا جديدًا يسمى Perception-Generation-Decision (PGD) الذي يدمج استخراج الميزات والتعلم المعزز. يستخدم هذا النموذج Transformer لضغط الميزات الديناميكية، وشبكة الخصومة التوليدية (GAN) لتوليد المسارات مع مراعاة القيود، وخوارزمية تدرج السياسة الحتمي العميق (DDPG) لتحسين السياسة. يهدف نموذج PGD إلى تعزيز كفاءة تخطيط المسار من خلال ضمان أن المسارات المولدة تلتزم بالقيود الفيزيائية بينما تسهل أيضًا الاستكشاف الفعال في فضاءات الحالة عالية الأبعاد. يؤكد المؤلفون أن نهجهم يعالج القضايا الحرجة في تخطيط مسار الطائرات بدون طيار، مثل معالجة الحالات عالية الأبعاد والحاجة إلى التعاون المتماسك بين استخراج الميزات وعمليات اتخاذ القرار، مما يؤدي في النهاية إلى تحسين الأداء في التضاريس المعقدة.
DOI: https://doi.org/10.1371/journal.pone.0340394
PMID: https://pubmed.ncbi.nlm.nih.gov/41632795
Publication Date: 2026-02-03
Author(s): Liangshuai Liu et al.
Primary Topic: Robotic Path Planning Algorithms
Overview
This paper introduces the PGD model for UAV path planning in complex terrains, effectively addressing challenges such as high-dimensional state processing, blind path exploration, and limited adaptability across different environments. The PGD model employs a synergistic integration of Transformer, Generative Adversarial Network (GAN), and Deep Deterministic Policy Gradient (DDPG) to create a “compression-generation-optimization” closed-loop system. The Transformer module compresses high-dimensional terrain data, mitigating training bottlenecks, while the GAN generates high-quality candidate paths to minimize ineffective exploration. DDPG then optimizes the path planning strategy, leading to significant improvements in efficiency and adaptability.
Experimental results indicate that PGD outperforms baseline models like PPO-DRL and Soft Actor-Critic on the UAVDT (suburban) and AirSim (canyon) datasets, achieving a path length of 20.0m/22.0m, a collision rate of 2.5%/3.0%, and computational efficiency of 13.5s/16.0s. The model demonstrates notable advantages in high-complexity terrains, highlighting its capability to enhance feature correlation and enforce physical path constraints through multi-module collaboration. Future research will aim to improve PGD’s adaptability to extreme weather conditions and multi-agent scenarios, focusing on dynamic candidate mechanisms and cross-modal feature enhancement.
Introduction
In complex terrain scenarios, such as mountain power inspection and canyon geological monitoring, unmanned aerial vehicles (UAVs) have emerged as vital alternatives to manual operations, particularly in high-risk and low-efficiency environments. Their inherent flexibility and maneuverability enable them to navigate challenging landscapes effectively. However, these environments present significant challenges due to rugged terrain and the presence of dynamically changing obstacles, such as sudden rockfalls and vegetation growth.
As a result, UAV path planning systems must address three critical requirements: ensuring collision-free path safety, optimizing path length, and adapting to the unpredictable nature of the terrain. These factors are essential for enhancing the operational efficiency and safety of UAVs in complex environments.
Methods
In this section, the authors outline the experimental setup and baseline model selection for evaluating the performance of the proposed PGD (Path Generation and Decision) model. The selected baseline models encompass a range of mainstream path planning techniques, including metaheuristic algorithms, deep reinforcement learning (DRL), and hybrid methods. Notably, the Improved Hybrid Squid-Shooter Ant Colony and Skyhawk Optimization Algorithm (IHSSAO) enhances global search capabilities and obstacle avoidance in complex terrains. Variants of DRL such as Twin Delayed Deep Deterministic Policy Gradient (TD3) and Quantum Enhanced DDPG (Q-DDPG) address Q-value overestimation and computational complexity, respectively, while Multicritic Delayed DDPG (MCD-DPG) improves robustness in challenging environments.
Additionally, Proximal Policy Optimization (PPO-DRL) and Soft Actor-Critic (SAC) are highlighted for their efficiency and adaptability in continuous control tasks and dynamic scenarios. The Dueling Deep Q Network (D3QN) is noted for optimizing value function estimation in discrete action spaces, and Multi-Agent DRL focuses on collaborative tasks, testing the scalability of the PGD model. Hybrid strategies like FM-Planner and Model Predictive Control Reinforcement Learning (MPC-RL) combine traditional and learning-based approaches. The comprehensive evaluation across these diverse models aims to underscore the advantages of the PGD model in terms of feature synergy and planning efficiency.
Results
The results of the study indicate significant findings regarding the primary hypotheses. The analysis reveals that the experimental group exhibited a statistically significant improvement in performance metrics compared to the control group, with a p-value of less than 0.05. Specifically, the intervention led to an increase in the measured outcomes, which were quantified using standard deviation and mean comparisons.
Furthermore, the data suggest a strong correlation between the variables analyzed, as evidenced by a correlation coefficient of $r = 0.85$. This indicates a robust relationship that supports the theoretical framework proposed in the study. Additional subgroup analyses revealed that the effects were particularly pronounced in specific demographics, highlighting the potential for targeted applications of the intervention. Overall, these results contribute to the existing literature by providing empirical evidence that supports the efficacy of the proposed methodology.
Discussion
In the discussion section of the paper, the authors address the limitations of traditional path planning algorithms, such as A* and Dijkstra, particularly in complex and unstructured terrains. These algorithms rely on manually designed cost functions, which can lead to suboptimal paths due to their inability to generalize across varying environments. Sampling-based methods like RRT* improve upon this by exploring high-dimensional spaces but struggle with efficiency as terrain complexity increases. The authors highlight the potential of deep learning techniques, noting that while Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs) have strengths in local feature extraction and temporal sequence modeling, respectively, they also exhibit significant limitations in capturing long-range dependencies and handling dynamic terrains.
To overcome these challenges, the authors propose a novel Perception-Generation-Decision (PGD) model that integrates feature extraction and reinforcement learning. This model utilizes a Transformer for dynamic feature compression, a Generative Adversarial Network (GAN) for constraint-aware path generation, and a Deep Deterministic Policy Gradient (DDPG) algorithm for policy optimization. The PGD model aims to enhance path planning efficiency by ensuring that the generated paths adhere to physical constraints while also facilitating effective exploration in high-dimensional state spaces. The authors emphasize that their approach addresses critical issues in UAV path planning, such as high-dimensional state processing and the need for cohesive collaboration between feature extraction and decision-making processes, ultimately leading to improved performance in complex terrains.
