DOI: https://doi.org/10.32604/cmc.2025.059224
تاريخ النشر: 2025-01-01
المؤلف: Yawar Abbas وآخرون
الموضوع الرئيسي: الاستشعار عن بعد وتطبيقات LiDAR
نظرة عامة
تقدم ورقة البحث نظامًا جديدًا للتعرف على نشاط الإنسان (HAR) مصممًا لمقاطع الفيديو الملتقطة بواسطة الطائرات بدون طيار، حيث يتناول تحديات كبيرة مثل اختلافات حركة الإنسان، والخلفيات المعقدة، وضباب الحركة، والاعتراضات، وزوايا الكاميرا المحدودة. يقوم النظام المقترح بمعالجة مقاطع الفيديو RGB من خلال تقسيمها أولاً إلى إطارات، وتطبيق تصفية ثنائية لتحسين مقدمة الأجسام مع تقليل ضوضاء الخلفية، وتحويل الصور إلى تدرجات الرمادي. يتم استخدام خوارزمية YOLO (You Only Look Once) لاكتشاف واستخراج أشكال الإنسان، والتي يتم من خلالها اشتقاق ميزات متنوعة – مثل زوايا المفاصل، والإزاحة، والسرعة، وتوزيع التدرجات الموجهة (HOG)، والنقاط ثلاثية الأبعاد، والمسافة الجيوديسية. يتم تحسين هذه الميزات باستخدام تحليل التمييز التربيعي (QDA) وتصنيفها من خلال مصنف عصبي ضبابي (NFC).
تم التحقق من فعالية النظام من خلال تقييمات في العالم الحقيقي على مجموعات بيانات متعددة، حيث حقق معدلات تعرف بلغت 93% لأفعال الطائرات بدون طيار، و97% لإيماءات الطائرات بدون طيار، و81% لأفعال أوكوتاما، مما يظهر دقته الفائقة مقارنة بالطرق الحالية. يقلل دمج QDA بشكل كبير من التصنيف الخاطئ بين فئات الأفعال المتشابهة، بينما تعزز خطوات المعالجة المسبقة جودة الصورة وتخفف من تداخل الخلفية. تهدف الأعمال المستقبلية إلى توسيع قدرات النظام من خلال دمج وظائف إضافية والتحقق منه مع مجموعة أوسع من مجموعات البيانات المرجعية، مما يعزز في النهاية قابليته للتكيف وأدائه في التطبيقات الواقعية.
مقدمة
تتناول مقدمة ورقة البحث هذه تعقيدات التعرف على نشاط الإنسان من مقاطع الفيديو الملتقطة بواسطة الطائرات الجوية غير المأهولة (UAVs). تعتبر هذه القدرة حاسمة عبر مجالات متعددة، بما في ذلك المراقبة بالفيديو، وإعادة التأهيل، وأداء الرياضة، والتفاعل بين الإنسان والروبوت. يبرز المؤلفون الفوائد المحتملة لهذه التكنولوجيا، لا سيما في تحسين نتائج إعادة التأهيل للمرضى وزيادة السلامة في سياقات المراقبة. ومع ذلك، يشيرون أيضًا إلى تحديات كبيرة، مثل اختلافات أوضاع الإنسان، وديناميات الحركة، وندرة بيانات التدريب المعلّمة، مما يعقد تطوير خوارزميات التعرف الفعالة.
للتغلب على هذه التحديات، يقترح المؤلفون طريقة جديدة تستخدم مصنفًا عصبيًا ضبابيًا (NFC) للتعرف على نشاط الإنسان بناءً فقط على فيديو RGB، مما يلغي الحاجة إلى معلومات العمق. تعزز هذه الطريقة الكفاءة في البيئات ذات الموارد المنخفضة مع الحفاظ على معدلات اكتشاف عالية. تشمل المكونات الرئيسية للنظام المقترح اكتشاف الإنسان عبر خوارزمية YOLO، واستخراج الميزات من خلال زوايا المفاصل وتوزيع التدرجات الموجهة (HOG)، والتحسين باستخدام تحليل التمييز التربيعي (QDA). تشير النتائج إلى أن هذه الطريقة تحسن دقة التعرف وتقلل من التصنيف الخاطئ بين الأفعال المتشابهة، مما يعالج تعقيدات حركة الإنسان في البيئات الديناميكية. بشكل عام، يظهر النظام المقترح تقدمًا كبيرًا في التعرف على نشاط الإنسان من مقاطع الفيديو الجوية.
طرق
تتضمن المنهجية المقترحة لتحليل مقاطع الفيديو RGB الملتقطة بواسطة الطائرات بدون طيار نهجًا منهجيًا يبدأ بتحويل الفيديو إلى إطارات فردية. تخضع كل إطار لخطوات معالجة مسبقة تهدف إلى تقليل التعقيد الحسابي، وتعزيز وضوح الأجسام الأمامية، وتقليل ضوضاء الخلفية. يتم إجراء اكتشاف الإنسان باستخدام خوارزمية YOLO (You Only Look Once)، التي تسهل استخراج الهياكل العظمية البشرية وتحديد المعالم الرئيسية المرتبطة بأجزاء الجسم المهمة، مثل الرأس، والرقبة، والكتفين، والمرفقين، والمعصمين، والوركين، والركبتين، والقدمين، وسرة البطن. تعتبر هذه المعالم حاسمة لحساب ميزات متنوعة، بما في ذلك المواقع العادية، وزوايا المفاصل، والإزاحة، والسرعة، وتوزيع التدرجات الموجهة (HOG)، والإحداثيات ثلاثية الأبعاد، والمسافات الجيوديسية. ثم يتم تحسين الميزات المستخرجة باستخدام تحليل التمييز التربيعي (QDA) وتصنيفها من خلال مصنف عصبي ضبابي (NFC) للتعرف على النشاط.
في الإعداد التجريبي، تم استخدام جهاز كمبيوتر محمول مزود بمعالج Intel Core i5 وذاكرة وصول عشوائي سعتها 8 جيجابايت، يعمل بنظام تشغيل Windows 10 بواجهة 64 بت مع PyCharm كبيئة تطوير. تم تقييم نظام التعرف على نشاط الإنسان (HAR) المقترح باستخدام ثلاث مجموعات بيانات: UAV-Gesture، Drone-action، وOkutama-Action. أظهرت النتائج دقة متوسطة بلغت 97% على مجموعة بيانات UAV-Gesture، و93% على مجموعة بيانات Drone-Action، و81% على مجموعة بيانات Okutama-Action. تظهر مقاييس الأداء، بما في ذلك مصفوفة الارتباك لمجموعة بيانات UAV-Gesture، فعالية النظام في التعرف بدقة على الأنشطة البشرية عبر سيناريوهات مختلفة.
مناقشة
تسلط قسم المناقشة في ورقة البحث الضوء على التقدم في التعرف على نشاط الإنسان من خلال تحسين خطوات المعالجة المسبقة وطرق استخراج الميزات المطبقة على بيانات الفيديو الملتقطة بواسطة الطائرات. يعزز النظام المقترح جودة الصورة باستخدام فلتر ثنائي، يحافظ على الحواف بينما يقلل الضوضاء، مما يسهل التعرف على الأفعال بشكل أكثر موثوقية مقارنة بالطرق السابقة التي اعتمدت على تصحيح غاما. تضمن هذه الخطوة المسبقة الحفاظ على التفاصيل المكانية الحرجة، مما يضع أساسًا قويًا لعمليات استخراج الميزات والتصنيف اللاحقة.
فيما يتعلق باكتشاف الإنسان، يقدم الدراسة نهجًا جديدًا يستخدم مقاطع الفيديو RGB بدلاً من البيانات الهيكلية، مما يسمح باستخراج معلومات بصرية أغنى. تستخدم الطريقة بنية تعلم عميق تضم 15 معلمًا رئيسيًا للجسم لالتقاط حركة الإنسان بدقة، مما يعزز دقة الاكتشاف من خلال خوارزمية YOLO. بالإضافة إلى ذلك، فإن دمج مصنف عصبي ضبابي وتحليل التمييز التربيعي (QDA) يحسن بشكل كبير أداء التصنيف من خلال إدارة الشكوك بفعالية وتحسين اختيار الميزات. تظهر النتائج تحسنًا ملحوظًا مقارنة بالنماذج الحالية، مما يشير إلى إمكانية النظام للتطبيقات الواقعية في التعرف على نشاط الإنسان. تهدف الأعمال المستقبلية إلى توسيع قدرات النموذج من خلال دمج مجموعات بيانات وسيناريوهات متنوعة لتعزيز مرونته التشغيلية ودقته بشكل أكبر.
DOI: https://doi.org/10.32604/cmc.2025.059224
Publication Date: 2025-01-01
Author(s): Yawar Abbas et al.
Primary Topic: Remote Sensing and LiDAR Applications
Overview
The research paper presents a novel Human Activity Recognition (HAR) system designed for drone-captured videos, addressing significant challenges such as variations in human motion, complex backgrounds, motion blur, occlusions, and limited camera angles. The proposed system processes RGB videos by first partitioning them into frames, applying bilateral filtering to enhance object foregrounds while minimizing background noise, and converting the images to grayscale. The YOLO (You Only Look Once) algorithm is employed to detect and extract human figures, from which various features—such as joint angles, displacement, velocity, histogram of oriented gradients (HOG), 3D points, and geodesic distance—are derived. These features are optimized using Quadratic Discriminant Analysis (QDA) and classified through a Neuro-Fuzzy Classifier (NFC).
The effectiveness of the system is validated through real-world evaluations on multiple datasets, achieving recognition rates of 93% for drone actions, 97% for UAV gestures, and 81% for Okutama actions, thus demonstrating its superior accuracy compared to existing methods. The integration of QDA significantly reduces misclassification between similar action classes, while the preprocessing steps enhance image quality and mitigate background interference. Future work aims to expand the system’s capabilities by incorporating additional functionalities and validating it with a broader range of benchmark datasets, ultimately enhancing its adaptability and performance in real-world applications.
Introduction
The introduction of this research paper addresses the complexities of recognizing human activity from video footage captured by Unmanned Aerial Vehicles (UAVs). This capability is critical across various fields, including video surveillance, rehabilitation, sports performance, and human-robot interaction. The authors highlight the potential benefits of such technology, particularly in enhancing rehabilitation outcomes for patients and improving safety in surveillance contexts. However, they also note significant challenges, such as variations in human poses, movement dynamics, and the scarcity of labeled training data, which complicate the development of effective recognition algorithms.
To overcome these challenges, the authors propose a novel method that utilizes a Neuro-Fuzzy Classifier (NFC) for human activity recognition based solely on RGB video, eliminating the need for depth information. This approach enhances efficiency in low-resource environments while maintaining high detection rates. Key components of the proposed system include human detection via the YOLO algorithm, feature extraction through joint angles and Histogram of Oriented Gradients (HOG), and optimization using Quadratic Discriminant Analysis (QDA). The findings indicate that this method improves recognition accuracy and reduces misclassification among similar actions, thereby addressing the complexities of human motion in dynamic settings. Overall, the proposed system demonstrates significant advancements in human activity recognition from aerial video footage.
Methods
The proposed methodology for analyzing RGB videos captured by drones involves a systematic approach that begins with converting the video into individual frames. Each frame undergoes preprocessing aimed at reducing computational complexity, enhancing the visibility of foreground objects, and minimizing background noise. Human detection is performed using the YOLO (You Only Look Once) algorithm, which facilitates the extraction of human skeletal structures and identification of key landmarks associated with significant body parts, such as the head, neck, shoulders, elbows, wrists, hips, knees, feet, and belly button. These landmarks are crucial for calculating various features, including normalized positions, joint angles, displacement, velocity, Histogram of Oriented Gradients (HOG), 3D coordinates, and geodesic distances. The extracted features are then optimized using Quadratic Discriminant Analysis (QDA) and classified through a Neuro-Fuzzy Classifier (NFC) for activity recognition.
In the experimental setup, a laptop equipped with an Intel Core i5 processor and 8 GB of RAM was utilized, running a 64-bit Windows 10 operating system with PyCharm as the development environment. The proposed Human Activity Recognition (HAR) system was evaluated using three datasets: UAV-Gesture, Drone-action, and Okutama-Action. The results indicated a mean accuracy of 97% on the UAV-Gesture dataset, 93% on the Drone-Action dataset, and 81% on the Okutama-Action dataset. The performance metrics, including a confusion matrix for the UAV-Gesture dataset, demonstrate the system’s effectiveness in accurately identifying human activities across different scenarios.
Discussion
The discussion section of the research paper highlights advancements in human activity recognition through improved preprocessing and feature extraction methods applied to drone-acquired video data. The proposed system enhances image quality using a bilateral filter, which preserves edges while reducing noise, thereby facilitating more reliable action recognition compared to previous methods that relied on gamma correction. This preprocessing step ensures that critical spatial details are maintained, setting a robust foundation for subsequent feature extraction and classification processes.
In terms of human detection, the study introduces a novel approach that utilizes RGB videos instead of skeletal data, allowing for richer visual information extraction. The method employs a deep learning architecture that incorporates 15 key body landmarks to accurately capture human motion, enhancing detection precision through the YOLO algorithm. Additionally, the integration of a Neuro-Fuzzy Classifier and Quadratic Discriminant Analysis (QDA) significantly improves classification performance by effectively managing uncertainties and optimizing feature selection. The results demonstrate a marked improvement over existing models, indicating the system’s potential for real-world applications in human activity recognition. Future work aims to expand the model’s capabilities by incorporating diverse datasets and scenarios to further enhance its operational flexibility and accuracy.
