دمج التمثيلات المعتمدة على الهيكل العظمي لتصنيف أوضاع اليوغا بشكل موثوق باستخدام نماذج التعلم العميق
Integrating skeleton based representations for robust yoga pose classification using deep learning models

شارك:
المجلة: Scientific Reports، المجلد: 16، العدد: 1
DOI: https://doi.org/10.1038/s41598-025-23726-0
PMID: https://pubmed.ncbi.nlm.nih.gov/42449118
تاريخ النشر: 2026-07-14
المؤلف: Mohammed Mohi-Ud-Din وآخرون
الموضوع الرئيسي: التعرف على وضع الجسم والحركة

نظرة عامة

تقدم ورقة البحث دراسة شاملة حول تصنيف أوضاع اليوغا الآلي، مع تسليط الضوء على أهمية التعرف الدقيق على الأوضاع لمنع الإصابات المرتبطة بالوضعيات غير الصحيحة. يقدم المؤلفون مجموعة بيانات جديدة، “Yoga-16″، التي تعالج قيود مجموعات البيانات الحالية وتقييم منهجي لثلاثة هياكل تعلم عميق—VGG16 و ResNet50 و Xception—باستخدام ثلاثة أوضاع إدخال مختلفة: الصور الخام، وصور هيكل MediaPipe Pose، وصور هيكل YOLOv8 Pose. تكشف النتائج أن التمثيلات المعتمدة على الهيكل تتفوق بشكل كبير على الصور الخام، حيث حقق VGG16 أعلى دقة تصنيف بنسبة 96.09% عند استخدام مدخلات هيكل MediaPipe Pose. بالإضافة إلى ذلك، تستخدم الدراسة Grad-CAM لتحليل القابلية للتفسير، مما يوفر رؤى حول عمليات اتخاذ القرار للنماذج.

في الختام، تؤسس هذه الأبحاث قاعدة قوية لتصنيف أوضاع اليوغا الآلي، مما يبرز فعالية التمثيلات المعتمدة على الهيكل وهياكل الشبكات العصبية المتقدمة. بينما تظهر مجموعة بيانات Yoga-16 نتائج واعدة، فإن حجمها المحدود يطرح تحديات للتعميم، والتي يتناولها المؤلفون من خلال التحقق المتقاطع. يُوصى بالعمل المستقبلي للتركيز على توسيع مجموعة البيانات، وتعزيز حساسية النموذج للاختلافات الدقيقة في الأوضاع، واستكشاف التعرف على الأوضاع في الوقت الحقيقي والمدخلات متعددة الوسائط. يُقترح دمج تقنيات الواقع المعزز والافتراضي لتعزيز تفاعل المستخدمين وسهولة الوصول، بهدف تطوير أنظمة آلية قوية لتعليم اليوغا الشخصية التي تعزز الرفاهية عبر مجموعات سكانية متنوعة.

مقدمة

تناقش مقدمة ورقة البحث التقدم في تصنيف أوضاع اليوغا، مع تسليط الضوء على الدور الكبير لتقنيات التعلم بالنقل والشبكات العصبية التلافيفية (CNNs) في تقدير أوضاع الإنسان. يصنف المؤلفون الدراسات الحالية إلى ثلاثة نهج رئيسية: (أ) إدخال الصورة المباشر أو عدم الهيكلة، (ب) الهيكلة، و (ج) التعلم بالنقل.

في نهج عدم الهيكلة، يتم استخدام هياكل CNN فقط، بينما يتضمن نهج الهيكلة نماذج CNN مختلفة، مثل ResNet50 و ResNet101 و Xception، بالتزامن مع أطر كشف الكائنات مثل YOLO و MediaPipe و OpenPose. كما قامت بعض الدراسات بدمج النماذج لكل من الهيكلة والتصنيف. يتميز نهج التعلم بالنقل باستخدام نماذج التعلم بالنقل لمهام التصنيف دون الحاجة لاستخراج النقاط الرئيسية، مما يتماشى مع منهجية عدم الهيكلة. يبرز هذا الاستعراض المنظم المنهجيات المتطورة في تقدير أوضاع اليوغا ويضع الأساس لمزيد من الاستكشاف في هذا المجال.

الطرق

توضح قسم المنهجية النهج المستخدم في تصنيف أوضاع اليوغا، مع تفاصيل حول جمع البيانات، واستخراج النقاط الرئيسية لأوضاع الإنسان، وعمليات التدريب لهياكل التعلم العميق. تقيم الدراسة تسع تقنيات متميزة لتصنيف أوضاع اليوغا، كما هو ملخص في الجدول 10، وتستخدم منصة Kaggle السحابية لضمان ظروف تجريبية متسقة، كما هو موضح في الجدول 11.

تم إجراء التجارب عبر ثلاثة أوضاع إدخال: (ط) إدخال الصورة المباشر، (2) إدخال صورة هيكل Mediapipe Pose، و (3) إدخال صورة هيكل YOLOv8 Pose. يسمح هذا الإعداد بإجراء تحليل مقارن للأداء بين سيناريوهات الإدخال المعتمدة على الهيكل وغير المعتمدة على الهيكل. يتم مناقشة أداء نماذج التعلم العميق المختلفة، وبشكل خاص VGG16 و ResNet50 و Xception، فيما يتعلق بهذه الأوضاع الإدخالية، مع تسليط الضوء على فعالية النهج المختلفة في تصنيف أوضاع اليوغا.

المناقشة

في قسم المناقشة، يتم تلخيص دراسات مختلفة حول تصنيف أوضاع اليوغا باستخدام تقنيات التعلم العميق، مع تسليط الضوء على التقدم في إنشاء مجموعات البيانات، وهياكل النماذج، ومقاييس الأداء. قدم Verma وآخرون مجموعة بيانات تضم 28,400 صورة عبر 82 وضعية يوغا، محققين دقة تصنيف بنسبة 79.35% مع متغيرات DenseNet الهرمية. عدل Imran وآخرون النماذج المدربة مسبقًا، ووجدوا أن نموذج Xception حقق أعلى دقة بنسبة 82.52% لـ 82 فئة و 92.42% لـ 6 فئات عُليا. طور Byeon وآخرون نموذجًا جماعيًا، EIR2TNet، محققين دقة 95.34% على مجموعة بيانات تضم 51,000 صورة، بينما جمع Liaqat وآخرون بين هياكل CNN و LSTM، محققين دقة تزيد عن 98% مع نموذج هجين.

علاوة على ذلك، تناقش القسم تأثير الهيكلة على تصنيف الأوضاع، حيث أظهر Garg وآخرون دقة محسنة باستخدام الصور المهيكلة، خاصة مع نموذج YogaConvo2d الذي حقق 99.62%. قارن Wadhwa وآخرون طرق كشف الأوضاع، ووجدوا أن MediaPipe تتفوق على OpenPose. استخدم Agrawal وآخرون نهج الهيكلة مع Random Forest، محققين دقة 99.04%. تختتم المناقشة برؤى حول التعلم بالنقل، حيث أفاد Long وآخرون بدقة قصوى بلغت 98.43% باستخدام نموذج TL-MobileNet-DA، مما يبرز فعالية تعزيز البيانات وضبط النموذج في تحسين أداء التصنيف عبر أوضاع اليوغا المختلفة.

Journal: Scientific Reports, Volume: 16, Issue: 1
DOI: https://doi.org/10.1038/s41598-025-23726-0
PMID: https://pubmed.ncbi.nlm.nih.gov/42449118
Publication Date: 2026-07-14
Author(s): Mohammed Mohi-Ud-Din et al.
Primary Topic: Human Pose and Action Recognition

Overview

The research paper presents a comprehensive study on automated yoga pose classification, highlighting the importance of accurate pose recognition to prevent injuries associated with incorrect postures. The authors introduce a new dataset, “Yoga-16,” which addresses the limitations of existing datasets and systematically evaluate three deep learning architectures—VGG16, ResNet50, and Xception—using three different input modalities: raw images, MediaPipe Pose skeleton images, and YOLOv8 Pose skeleton images. The findings reveal that skeleton-based representations significantly outperform raw images, with VGG16 achieving the highest classification accuracy of 96.09% when using MediaPipe Pose skeleton inputs. Additionally, the study employs Grad-CAM for interpretability analysis, providing insights into the decision-making processes of the models.

In conclusion, this research lays a strong foundation for automated yoga pose classification, underscoring the effectiveness of skeleton-based representations and advanced neural network architectures. While the Yoga-16 dataset demonstrates promising results, its limited size poses challenges for generalization, which the authors address through cross-validation. Future work is recommended to focus on expanding the dataset, enhancing model sensitivity to subtle pose distinctions, and exploring real-time pose recognition and multimodal inputs. The integration of augmented and virtual reality technologies is suggested to enhance user engagement and accessibility, ultimately aiming to develop robust automated systems for personalized yoga instruction that promote wellness across diverse populations.

Introduction

The introduction of the research paper discusses the advancements in yoga pose classification, highlighting the significant role of transfer learning techniques and Convolutional Neural Networks (CNNs) in human pose estimation. The authors categorize existing studies into three main approaches: (a) direct image input or non-skeletonization, (b) skeletonization, and (c) transfer learning.

In the non-skeletonization approach, only CNN architectures are utilized, while the skeletonization approach incorporates various CNN models, such as ResNet50, ResNet101, and Xception, in conjunction with object detection frameworks like YOLO, MediaPipe, and OpenPose. Some studies have also combined models for both skeletonization and classification. The transfer learning approach is characterized by the use of transfer learning models for classification tasks without the need for keypoint extraction, aligning with the non-skeletonization methodology. This structured review underscores the evolving methodologies in yoga pose estimation and sets the stage for further exploration in the field.

Methods

The methodology section outlines the approach for classifying yoga poses, detailing the data collection, human pose keypoint extraction, and training processes for the deep learning architecture. The study evaluates nine distinct yoga pose classification techniques, as summarized in Table 10, and utilizes Kaggle’s cloud-based platform for consistent experimental conditions, as indicated in Table 11.

The experiments were conducted across three input modalities: (i) Direct Image Input, (ii) Mediapipe Pose Skeleton Image Input, and (iii) YOLOv8 Pose Skeleton Image Input. This setup allows for a comparative analysis of performance between skeleton-based and non-skeleton-based input scenarios. The performance of various deep learning models, specifically VGG16, ResNet50, and Xception, is discussed in relation to these input modalities, highlighting the effectiveness of different approaches in yoga pose classification.

Discussion

In the discussion section, various studies on yoga pose classification using deep learning techniques are summarized, highlighting advancements in dataset creation, model architecture, and performance metrics. Verma et al. introduced a dataset of 28,400 images across 82 yoga poses, achieving a classification accuracy of 79.35% with hierarchical DenseNet variants. Imran et al. modified pre-trained models, finding that the Xception model yielded the highest accuracy of 82.52% for 82 classes and 92.42% for 6 super classes. Byeon et al. developed an ensemble model, EIR2TNet, achieving 95.34% accuracy on a dataset of 51,000 images, while Liaqat et al. combined CNN and LSTM architectures, achieving over 98% accuracy with a hybrid model.

Further, the section discusses the impact of skeletonization on pose classification, with Garg et al. demonstrating improved accuracy using skeletonized images, particularly with the YogaConvo2d model achieving 99.62%. Wadhwa et al. compared pose detection methods, finding MediaPipe superior to OpenPose. Agrawal et al. utilized a skeletonization approach with Random Forest, achieving 99.04% accuracy. The discussion concludes with insights into transfer learning, where Long et al. reported a top accuracy of 98.43% using the TL-MobileNet-DA model, emphasizing the effectiveness of data augmentation and model fine-tuning in enhancing classification performance across various yoga poses.

شارك: