DOI: https://doi.org/10.1038/s41598-024-58146-z
PMID: https://pubmed.ncbi.nlm.nih.gov/38580704
تاريخ النشر: 2024-04-05
المؤلف: Chengang Dong وآخرون
الموضوع الرئيسي: التعرف على وضع الجسم والحركة
نظرة عامة
تقدم ورقة البحث نموذجًا جديدًا لتقدير وضعية الإنسان في الوقت الحقيقي (HPE) يسمى CCAM-Person، مبنيًا على إطار عمل YOLOv8. الهدف الرئيسي من هذا النموذج هو تعزيز دقة تقدير وضعية الإنسان في ظل التحديات مثل الحجب الجزئي وحقول الاستقبال المحدودة، والتي يمكن أن تعيق الأداء. لمعالجة هذه القضايا، قام المؤلفون بتحسين مكونات العمود الفقري والعنق لنموذج YOLOv8x-pose وقدموا وحدة انتباه إحداثيات السياق (CCAM). تعزز هذه الوحدة قدرة النموذج على التركيز على الميزات البارزة مع تقليل الضوضاء الخلفية والتخفيف من فشل انحدار النقاط الرئيسية بسبب حجب الأطراف. يظهر النموذج المقترح تحسينات كبيرة، حيث حقق زيادة في الدقة المتوسطة بنسبة 2.8% و3.5% على مجموعتي بيانات MS COCO 2017 وCrowdPose، على التوالي، مقارنة بنموذج YOLOv8x-pose الأساسي.
في الختام، يظهر إطار عمل CCAM-Person تقدمًا واعدًا في اكتشاف الأشخاص وتقدير الوضعية، محققًا نتائج تنافسية في دقة الانحدار وسرعة الاستدلال مقارنة بأساليب HPE الأخرى في الوقت الحقيقي. أدت التحسينات التي تم إجراؤها على بنية نموذج YOLOv8x-pose وقدرات معالجة المعلومات إلى تحقيق اختراقات ملحوظة في الدقة. ستركز الأعمال المستقبلية على تحسين هيكل النموذج وسرعة الاستدلال، مع إمكانية دمج تقنيات حديثة مثل ReXNet وPAGCP pruning. بالإضافة إلى ذلك، يخطط المؤلفون لاستكشاف تطبيق آلية انتباه CCAM في مجالات أخرى، بما في ذلك تقسيم الصور وتتبع الأجسام.
طرق
يقدم البحث نموذجًا جديدًا لتقدير وضعية الإنسان في الوقت الحقيقي (HPE) يسمى CCAM-Person، مصممًا للتغلب على القيود في الأساليب الحالية، مثل عدم دقة تحديد النقاط الرئيسية وفشل تقدير الوضعية بسبب الحجب. يبني النموذج على بنية YOLOv8، مدمجًا الانحدار في الوقت الحقيقي لنقاط رئيسية بشرية جنبًا إلى جنب مع اكتشاف المناطق. يستخدم CCAM-Person نهجًا يشبه التصنيف الثنائي لتحديد جميع النقاط الرئيسية البشرية المحتملة في الصورة، يتبعه خطوة معالجة لاحقة لمطابقة هذه النقاط مع الأجسام البشرية الحقيقية. ينتج النموذج متجهًا يحتوي على معلومات موضعية وثقة لكل من المنطقة البشرية المكتشفة والنقاط الرئيسية المرتبطة، مما يعزز الفعالية العامة لتقدير الوضعية.
تظهر النتائج التجريبية أن CCAM-Person يحقق دقة تقدير تبلغ 74.9% على مجموعة بيانات MS COCO 2017، متفوقًا على العديد من أساليب HPE الحديثة في الوقت الحقيقي بنسبة 2.8% في الدقة و4.2% في معدل الاسترجاع. كما يظهر النموذج أداءً تنافسيًا ضد الأساليب المعتمدة على Vision Transformer من حيث الدقة وسرعة المعالجة، مع الحفاظ على الكفاءة حتى مع أحجام الإدخال الأكبر. بالإضافة إلى ذلك، تم إجراء دراسات إلغاء لتقييم مساهمات وحدات التحسين المختلفة، مما يؤكد فعالية التحسينات المقترحة في تحقيق تقدير وضعية متفوق، خاصة في المشاهد ذات الكثافة السكانية العالية.
مناقشة
تستعرض قسم المناقشة في الورقة مجموعة متنوعة من الأساليب لتقدير وضعية الإنسان (HPE)، مع تسليط الضوء على كل من المنهجيات ذات المرحلتين والمرحلة الواحدة. تركز الأساليب ذات المرحلتين، مثل Alphapose وHRNet، على تحسين تحديد موقع الصندوق المحيط ودقة الوضعية من خلال هياكل معقدة ودمج ميزات عالية الدقة. ومع ذلك، غالبًا ما تتطلب هذه الأساليب موارد حسابية كبيرة وتواجه تحديات تتعلق بضبط المعلمات الفائقة. في المقابل، تهدف الأساليب ذات المرحلة الواحدة مثل OpenPose وHigherHRNet إلى تحقيق الكفاءة من خلال الانحدار المباشر للنقاط الرئيسية من الصور، لكنها تواجه مشكلات مثل الحساسية لجودة الصورة والحاجة إلى مجموعات بيانات تدريب واسعة.
تقدم الورقة أيضًا أطرًا جديدة مثل وحدة انتباه إحداثيات السياق (CCAM) وشبكة هرم الميزات متعددة المسارات (MFPN)، التي تعزز استخراج الميزات والتفاعل عبر مقاييس مختلفة. تحسن CCAM دقة انحدار النقاط الرئيسية من خلال دمج آليات الانتباه القنوي والمكاني، بينما تعمل MFPN على تحسين دمج الميزات للاحتفاظ بالمعلومات الدلالية والتفصيلية. يتم التحقق من هذه التقدمات من خلال التجارب على مجموعات بيانات MS COCO وCrowdPose، مما يظهر تحسينات في مقاييس الأداء مثل الدقة المتوسطة (AP) ومعدل الاسترجاع المتوسط (AR). بشكل عام، تؤدي التحسينات المقترحة إلى زيادة كبيرة في دقة تقدير الوضعية، مما يعالج التحديات التي تطرحها مقاييس الأجسام المتغيرة والحجب في المشاهد المعقدة.
DOI: https://doi.org/10.1038/s41598-024-58146-z
PMID: https://pubmed.ncbi.nlm.nih.gov/38580704
Publication Date: 2024-04-05
Author(s): Chengang Dong et al.
Primary Topic: Human Pose and Action Recognition
Overview
The research paper presents a novel real-time human pose estimation (HPE) model named CCAM-Person, built upon the YOLOv8 framework. The primary aim of this model is to enhance the accuracy of HPE in the presence of challenges such as partial occlusion and limited receptive fields, which can hinder performance. To address these issues, the authors have refined the backbone and neck components of the YOLOv8x-pose model and introduced a context coordinate attention module (CCAM). This module enhances the model’s ability to focus on salient features while minimizing background noise and mitigating key point regression failures due to limb occlusion. The proposed model demonstrates significant improvements, achieving an average precision increase of 2.8% and 3.5% on the MS COCO 2017 and CrowdPose datasets, respectively, compared to the baseline YOLOv8x-pose model.
In conclusion, the CCAM-Person framework shows promising advancements in person detection and pose estimation, achieving competitive results in regression accuracy and inference speed against other real-time HPE methods. The enhancements made to the YOLOv8x-pose model’s architecture and information processing capabilities have led to notable breakthroughs in accuracy. Future work will focus on further optimizing the model structure and inference speed, potentially integrating recent technologies such as ReXNet and PAGCP pruning. Additionally, the authors plan to explore the application of the CCAM attention mechanism in other domains, including image segmentation and object tracking.
Methods
The research presents a novel real-time human pose estimation (HPE) model named CCAM-Person, designed to overcome limitations in existing methods, such as inaccurate keypoint localization and pose estimation failures due to occlusions. The model builds upon the YOLOv8 architecture, integrating real-time regression of human keypoints alongside region detection. CCAM-Person employs a binary classification-like approach to identify all potential human keypoints in an image, followed by a post-processing step to match these keypoints with ground-truth human bodies. The model outputs a vector containing positional and confidence information for both the detected human region and associated keypoints, enhancing the overall effectiveness of pose estimation.
Experimental results demonstrate that CCAM-Person achieves an estimation accuracy of 74.9% on the MS COCO 2017 dataset, outperforming several contemporary real-time HPE methods by 2.8% in accuracy and 4.2% in recall rate. The model also exhibits competitive performance against Vision Transformer-based methods in terms of accuracy and processing speed, maintaining efficiency even with larger input sizes. Additionally, ablation studies are conducted to evaluate the contributions of various enhancement modules, confirming the effectiveness of the proposed improvements in achieving superior pose estimation, particularly in densely populated scenes.
Discussion
The discussion section of the paper reviews various approaches to human pose estimation (HPE), highlighting both two-stage and single-stage methodologies. The two-stage approaches, such as Alphapose and HRNet, focus on improving bounding box localization and pose accuracy through complex architectures and high-resolution feature integration. However, these methods often require substantial computational resources and face challenges related to hyperparameter tuning. In contrast, single-stage methods like OpenPose and HigherHRNet aim for efficiency by regressing keypoints directly from images, yet they struggle with issues like sensitivity to image quality and the need for extensive training datasets.
The paper also introduces novel frameworks such as the Context Coordinate Attention Module (CCAM) and the Multi-Path Feature Pyramid Network (MFPN), which enhance feature extraction and interaction across different scales. The CCAM improves keypoint regression accuracy by integrating channel and spatial attention mechanisms, while the MFPN optimizes feature fusion to retain semantic and detailed information. These advancements are validated through experiments on the MS COCO and CrowdPose datasets, demonstrating improved performance metrics such as Average Precision (AP) and Average Recall (AR). Overall, the proposed enhancements lead to a significant increase in pose estimation accuracy, addressing the challenges posed by varying object scales and occlusions in complex scenes.
