DOI: https://doi.org/10.3389/fnbot.2025.1582995
PMID: https://pubmed.ncbi.nlm.nih.gov/40313416
تاريخ النشر: 2025-04-17
المؤلف: Mohammed Alshehri وآخرون
الموضوع الرئيسي: التعرف على وضع الجسم والحركة
نظرة عامة
تقدم ورقة البحث إطار عمل للتعلم العميق مصمم للتعرف على أفعال متعددة الأشخاص في مقاطع الفيديو الملتقطة بواسطة الطائرات بدون طيار (UAVs). تتناول الدراسة تحديات كبيرة مثل تشويش الحركة، والخلفيات الديناميكية، وظروف الإضاءة المتغيرة، التي تعقد مهام التعرف على الأفعال. من خلال دمج تقنيات استخراج الميزات المتقدمة ونماذج الشبكات العصبية، يحقق النظام المقترح دقة تعرف تبلغ 91.50% على مجموعة بيانات MOD20 و89.71% على مجموعة بيانات Okutama-Action، مما يظهر فعاليته وقوته في البيئات المعقدة.
تشير النتائج إلى أن مصنف الشبكة العصبية التلافيفية (CNN) يتفوق على النماذج الأخرى، مثل الشبكات العميقة القائمة على الاعتقاد (DBN) والشبكات العصبية المتكررة (RNN)، خاصة في استخراج الميزات المكانية من صور UAV مع التكيف مع التغيرات في المنظور والمقياس. تسلط الدراسة الضوء على إمكانيات الأساليب المعتمدة على الشبكات العصبية في التغلب على قيود تطبيقات UAV في التعرف على أفعال البشر. ستركز الأبحاث المستقبلية على معالجة تحدي الأفعال المحجوبة، التي تكون شائعة في عمليات UAV في العالم الحقيقي، من خلال دمج رؤى زمنية ومعلومات سياقية لتعزيز قوة النظام في السيناريوهات الديناميكية.
مقدمة
تناقش مقدمة ورقة البحث الأهمية المتزايدة للتعرف على أفعال البشر (HAR) ضمن مجال رؤية الكمبيوتر، مع تسليط الضوء على تطبيقاته في مجالات مثل تحليلات الرياضة، والرعاية الصحية، والمراقبة. يتم التأكيد على دمج HAR مع الطائرات بدون طيار (UAVs)، أو الطائرات المسيرة، كتنمية تحويلية، مما يمكّن من جمع البيانات المحسّنة من خلال لقطات جوية عالية الجودة. ومع ذلك، فإن هذا الدمج يقدم تحديات، بما في ذلك زوايا الرؤية المتغيرة وظروف الإضاءة التي تعقد اكتشاف الأفعال، خاصة عندما تظهر الموضوعات أصغر وأقل تمييزًا.
لمعالجة هذه التحديات، يقترح المؤلفون إطار عمل قوي للتعرف على أفعال متعددة الأشخاص في مقاطع الفيديو الملتقطة بواسطة UAV. يتضمن هذا الإطار خط أنابيب معالجة مسبقة لتحسين جودة الفيديو ويستخدم استراتيجية استخراج ميزات متعددة المستويات تجمع بين ميزات الجسم الكامل (مثل أوصاف فورييه وأوصاف AKAZE) مع ميزات قائمة على النقاط الرئيسية مستمدة من التمثيلات الهيكلية للموضوعات البشرية. يستخدم النظام مصنفات تعلم عميقة متقدمة – الشبكات العميقة القائمة على الاعتقاد (DBN)، الشبكات العصبية التلافيفية (CNN)، والشبكات العصبية المتكررة (RNN) – تم تحسينها من خلال الانحدار التدرجي لضمان تصنيف دقيق للأفعال. تظهر المنهجية المقترحة فعاليتها، حيث تحقق دقة تبلغ 91.50% و89.71% على مجموعات البيانات المرجعية MOD20 وOkutama-Action، على التوالي. توضح الورقة هيكلها، مشيرة إلى أن الأقسام التالية ستتناول الأعمال ذات الصلة، والمنهجية، وإعداد التجارب، والنتائج، واتجاهات البحث المستقبلية.
الطرق
تستخدم المنهجية المقترحة للتعرف على أفعال متعددة الأشخاص في مقاطع الفيديو الملتقطة بواسطة UAV نهجًا منظمًا مصممًا لتحديات الصور الجوية. تبدأ العملية بخطوات معالجة مسبقة، بما في ذلك تشويش غاوسي لتقليل الضوضاء، وتحويل إلى تدرج الرمادي، وإزالة الخلفية. يتم استخراج ظلال البشر باستخدام نموذج خليط غاوسي (GMM)، تليها بناء نموذج هيكلي لتمثيل النقاط الرئيسية. يتم إجراء استخراج الميزات من خلال استراتيجيتين: ميزات الجسم الكامل التي تلتقط أنماط الحركة العامة وميزات قائمة على النقاط الرئيسية تركز على ديناميات الحركة. يستخدم النظام محسن انحدار تدرجي لتحقيق تحسين فعال ويصنف الأفعال باستخدام ثلاثة مصنفات تعلم عميقة: الشبكات العميقة القائمة على الاعتقاد (DBN)، الشبكات العصبية التلافيفية (CNN)، والشبكات العصبية المتكررة (RNN). يمكّن هذا الخط الأنبوبي الشامل من التعرف الدقيق على أفعال متعددة الأشخاص عبر سيناريوهات UAV المختلفة.
يتضمن إعداد التجارب وصفًا تفصيليًا لمجموعات البيانات، وتكوين النظام، ومقاييس التقييم. تم استخدام جهاز كمبيوتر يعمل بنظام Windows 10 مزود بمعالج Intel Core i7، وGPU Nvidia Tesla K80، و16 جيجابايت من ذاكرة الوصول العشوائي لتدريب النموذج وتقييمه باستخدام Python 3.6 وواجهة برمجة التطبيقات Keras. تم تقسيم مجموعة البيانات إلى 80% للتدريب و20% للاختبار، مما يضمن تقييمًا قويًا لأداء النموذج. تظهر التحليلات المقارنة ضد التقنيات الحديثة أن النظام المقترح يتفوق على الأساليب الحالية، كما يتضح من معايير الدقة المقدمة في الدراسة. من الجدير بالذكر أن تقنيات استخراج الميزات المتقدمة، وتحسين الانحدار التدرجي، وتصنيف CNN تساهم في الدقة الكبيرة المحققة، كما هو موضح في النتائج من مجموعة بيانات MOD20.
النتائج
في قسم النتائج، أجرى المؤلفون سلسلة من التجارب لتقييم دقة تصنيف النموذج المقترح باستخدام مجموعات البيانات المرجعية. كان الهدف الرئيسي هو تقييم فعالية النموذج من خلال مقارنة أدائه ضد طرق التصنيف الحديثة المختلفة. كان التحليل يهدف إلى تقديم فهم شامل لقدرات النموذج ومزاياه المحتملة في التطبيقات العملية.
المناقشة
تسلط قسم المناقشة في الورقة الضوء على التقدم في التعرف على أفعال البشر من صور UAV، مع التأكيد على دمج أساليب التعلم الآلي والتعلم العميق. تستخدم أساليب التعلم الآلي، مثل الطريقة التي طورها عباس وجلال (2024)، استخراج الميزات ونماذج التصنيف، بما في ذلك YOLOv5 لاكتشاف البشر وتقدير الوضع، تليها تقليل الأبعاد من خلال تحليل التمييز الخطي (LDA) والتصنيف عبر آلات الدعم المتجهة (SVM). أظهرت هذه الطريقة فعالية في التعرف على الأفعال على مجموعة بيانات Drone-Action، مما يبرز إمكانيات أطر التعلم الآلي التقليدية في تطبيقات UAV.
من ناحية أخرى، عززت تقنيات التعلم العميق بشكل كبير قدرة UAVs على تحليل بيانات الفيديو مباشرة، متجاوزة الحاجة إلى استخراج الميزات يدويًا. تعالج أطر بارزة مثل المحول الهجين الانتباه (HAT) وتوافق الميزات الزمنية المستندة إلى المعلومات المتبادلة (MITFAS) التحديات مثل الحجب وزوايا الرؤية المتغيرة، مما يحسن دقة التعرف على الأفعال. تناقش الورقة أيضًا هياكل التعلم العميق المختلفة، بما في ذلك CNNs والمحولات، التي تتفوق في اكتشاف الأجسام المتحركة الصغيرة في البيئات المعقدة. تشير النتائج إلى أنه بينما يقدم التعلم العميق أداءً متفوقًا في التعرف على أفعال متعددة الأشخاص، لا تزال هناك حاجة إلى تحسينات إضافية في الكفاءة الحسابية والقدرة على التكيف مع البيئات المتنوعة لتطبيقات عملية.
DOI: https://doi.org/10.3389/fnbot.2025.1582995
PMID: https://pubmed.ncbi.nlm.nih.gov/40313416
Publication Date: 2025-04-17
Author(s): Mohammed Alshehri et al.
Primary Topic: Human Pose and Action Recognition
Overview
The research paper presents a deep learning framework designed for recognizing multi-person actions in videos captured by Unmanned Aerial Vehicles (UAVs). The study addresses significant challenges such as motion blur, dynamic backgrounds, and varying lighting conditions, which complicate action recognition tasks. By integrating advanced feature extraction techniques and neural network models, the proposed system achieves a recognition accuracy of 91.50% on the MOD20 dataset and 89.71% on the Okutama-Action dataset, demonstrating its effectiveness and robustness in complex environments.
The findings indicate that the convolutional neural network (CNN) classifier outperforms other models, such as Deep Belief Networks (DBN) and Recurrent Neural Networks (RNN), particularly in extracting spatial features from UAV imagery while adapting to changes in perspective and scale. The study highlights the potential of neural network-based methods in overcoming the limitations of UAV applications for human action recognition. Future research will focus on addressing the challenge of occluded actions, which are common in real-world UAV operations, by incorporating time-related insights and contextual information to enhance the system’s robustness in dynamic scenarios.
Introduction
The introduction of the research paper discusses the growing significance of Human Action Recognition (HAR) within the field of computer vision, highlighting its applications in areas such as sports analytics, healthcare, and surveillance. The integration of HAR with Unmanned Aerial Vehicles (UAVs), or drones, is emphasized as a transformative development, enabling enhanced data collection through high-quality aerial footage. However, this integration presents challenges, including variable viewing angles and lighting conditions that complicate action detection, particularly when subjects appear smaller and less distinguishable.
To address these challenges, the authors propose a robust framework for multi-person action recognition in UAV-captured videos. This framework incorporates a preprocessing pipeline to enhance video quality and employs a multi-level feature extraction strategy that combines full-body features (such as Fourier Descriptors and AKAZE descriptors) with keypoint-based features derived from skeletal representations of human subjects. The system utilizes advanced deep learning classifiers—Deep Belief Networks (DBN), Convolutional Neural Networks (CNN), and Recurrent Neural Networks (RNN)—optimized through gradient descent to ensure accurate classification of actions. The proposed methodology demonstrates effectiveness, achieving accuracies of 91.50% and 89.71% on benchmark datasets MOD20 and Okutama-Action, respectively. The paper outlines its structure, indicating subsequent sections will cover related work, methodology, experimental setup, results, and future research directions.
Methods
The proposed methodology for multi-person action recognition in UAV-captured videos employs a structured approach tailored to the challenges of aerial imagery. The process begins with preprocessing steps, including Gaussian blur for noise reduction, grayscale conversion, and background removal. Human silhouettes are extracted using a Gaussian Mixture Model (GMM), followed by the construction of a skeletal model to represent keypoints. Feature extraction is conducted through two strategies: full-body features that capture overall movement patterns and keypoint-based features focusing on motion dynamics. The system utilizes a gradient descent optimizer for efficient optimization and classifies actions using three deep learning classifiers: Deep Belief Networks (DBN), Convolutional Neural Networks (CNN), and Recurrent Neural Networks (RNN). This comprehensive pipeline enables precise multi-person action recognition across various UAV scenarios.
The experimental setup includes a detailed description of the datasets, system configuration, and evaluation metrics. A Windows 10 PC equipped with an Intel Core i7 processor, Nvidia Tesla K80 GPU, and 16 GB of RAM was employed for model training and evaluation using Python 3.6 and the Keras API. The dataset was partitioned into 80% for training and 20% for testing, ensuring a robust assessment of model performance. Comparative analysis against state-of-the-art techniques demonstrates that the proposed system outperforms existing methods, as evidenced by accuracy benchmarks presented in the study. Notably, the advanced feature extraction techniques, gradient descent optimization, and CNN-based classification contribute to the significant accuracy achieved, as highlighted in the results from the MOD20 dataset.
Results
In the Results section, the authors conducted a series of experiments to evaluate the classification accuracy of their proposed model using benchmark datasets. The primary objective was to assess the model’s effectiveness by comparing its performance against various state-of-the-art classification methods. The analysis aimed to provide a comprehensive understanding of the model’s capabilities and its potential advantages in practical applications.
Discussion
The discussion section of the paper highlights advancements in human action recognition from UAV imagery, emphasizing the integration of machine learning and deep learning methodologies. Machine learning approaches, such as the method developed by Abbas and Jalal (2024), utilize feature extraction and classification models, including YOLOv5 for human detection and pose estimation, followed by dimensionality reduction through Linear Discriminant Analysis (LDA) and classification via Support Vector Machines (SVM). This method demonstrated effective action recognition on the Drone-Action dataset, showcasing the potential of conventional machine learning frameworks in UAV applications.
Conversely, deep learning techniques have significantly enhanced the capability of UAVs to analyze video data directly, bypassing the need for manual feature extraction. Notable frameworks like the Hybrid Attention Transformer (HAT) and the Mutual Information-Based Temporal Feature Alignment and Sampling (MITFAS) address challenges such as occlusion and varying viewpoints, improving action recognition accuracy. The paper also discusses various deep learning architectures, including CNNs and transformers, which excel in detecting small moving objects in complex environments. The findings indicate that while deep learning offers superior performance in multi-person action recognition, further improvements in computational efficiency and adaptability to diverse environments remain essential for practical applications.
