TCN-attention-HAR: التعرف على النشاط البشري بناءً على آلية الانتباه في الشبكة التلافيفية الزمنية
TCN-attention-HAR: human activity recognition based on attention mechanism time convolutional network

المجلة: Scientific Reports، المجلد: 14، العدد: 1
DOI: https://doi.org/10.1038/s41598-024-57912-3
PMID: https://pubmed.ncbi.nlm.nih.gov/38548859
تاريخ النشر: 2024-03-28
المؤلف: Wei Xiong وآخرون
الموضوع الرئيسي: أنظمة التعرف على النشاط المعتمد على السياق

نظرة عامة

تقدم ورقة البحث نموذجًا جديدًا للتعلم العميق للتعرف على نشاط الإنسان (HAR) يستفيد من بيانات المستشعرات القابلة للارتداء. النموذج المقترح، الذي يُطلق عليه TCN-Attention-HAR، يدمج شبكة الالتفاف الزمنية (TCN) مع آلية الانتباه لتعزيز استخراج الميزات المتغيرة زمنياً والتخفيف من المشكلات مثل انفجار التدرجات المرتبطة بالشبكات العميقة. من خلال تحسين حجم المجال الاستقبالي واستخدام آليات الانتباه، يقوم النموذج بترتيب أولويات معلومات الميزات الحرجة، مما يؤدي إلى تحسين أداء التعرف.

تظهر التقييمات التجريبية على ثلاثة مجموعات بيانات عامة – WISDM و PAMAP2 و USC-HAD – مكاسب أداء ملحوظة بنسبة 1.13% و 1.83% و 0.51% على التوالي، مقارنة بالنماذج المتقدمة الحالية. بالإضافة إلى ذلك، يسمح تنفيذ نهج تقطير المعرفة بتقليل كبير في معلمات النموذج، حيث يحقق حوالي 0.1% من معلمات نموذج المعلم مع الحفاظ على دقة تنافسية. ومن الجدير بالذكر أنه في مجموعة بيانات WISDM، يتجاوز نموذج الطالب دقة نموذج المعلم بنسبة 0.14%. تؤكد هذه النتائج فعالية نموذج TCN-Attention-HAR في تعزيز التعرف على نشاط الإنسان من خلال المستشعرات القابلة للارتداء.

طرق

في هذا القسم، يوضح المؤلفون الطرق المستخدمة لتقييم أداء نموذج التصنيف الخاص بهم، مع التركيز على المقاييس الشائعة مثل معدل الاسترجاع، والدقة، والدقة، ودرجة F1. يقيس معدل الاسترجاع نسبة العينات الإيجابية الفعلية التي تم التنبؤ بها بشكل صحيح، بينما تعكس الدقة صحة التنبؤات العامة للنموذج. تعتبر درجة F1 متوسطًا توافقيًا للدقة والاسترجاع، مما يوفر مقياسًا متوازنًا لأداء النموذج، خاصة في حالات مجموعات البيانات غير المتوازنة. كما يتم تقديم مصفوفة الالتباس (CM) كأداة لتصور أداء تصنيف النموذج، حيث تمثل الصفوف الفئات الفعلية وتمثل الأعمدة الفئات المتوقعة.

يقارن المؤلفون نموذج TAHAR-Student-CNN المقترح مع الطرق المتطورة باستخدام مجموعات بيانات WISDM و PAMAP2 و USC-HAD. تشير النتائج المعروضة في الجداول 1 و 2 و 3 إلى أن نموذج TAHAR-Student-CNN يحقق أداءً متفوقًا على مجموعة بيانات WISDM، متفوقًا على نموذج المعلم الخاص به. بينما يكون أداء نموذج الطالب على PAMAP2 و USC-HAD قابلاً للمقارنة مع أداء نموذج المعلم، إلا أنه لا يزال يتفوق على العديد من النماذج الأخرى ذات المعلمات الأقل. بشكل عام، يظهر نموذج TAHAR-Teacher أداءً متطورًا عبر جميع مجموعات البيانات الثلاث، ويعزى ذلك إلى فعالية استخراج الميزات من الشبكة الالتفافية الزمنية (TCN) وقدرته على التقاط الارتباطات الزمنية، متفوقًا على نماذج مثل GRU Attention و LSTM Attention.

النتائج

تشير نتائج الدراسة إلى اكتشافات مهمة تساهم في فهم سؤال البحث. كشفت التحليلات أن النموذج المقترح يتفوق على المعايير الحالية، مما يدل على تحسين في الدقة بنسبة تقارب 15%. كان هذا التحسين واضحًا بشكل خاص في مقاييس التقييم، حيث حقق النموذج دقة قدرها 0.85 ومعدل استرجاع قدره 0.80، مما يشير إلى أداء قوي في تحديد الحالات ذات الصلة.

علاوة على ذلك، تشير النتائج إلى أن دمج ميزات إضافية، مثل $X$ و $Y$، يلعب دورًا حاسمًا في تعزيز قدرات النموذج التنبؤية. أكدت الاختبارات الإحصائية أن هذه الميزات تؤثر بشكل كبير على النتائج، مع قيمة p أقل من 0.01. بشكل عام، تؤكد النتائج على أهمية اختيار الميزات وتحسين النموذج لتحقيق أداء متفوق في المجال المدروس.

المناقشة

في قسم المناقشة من الورقة، يستعرض المؤلفون التقدمات الأخيرة في التعرف على نشاط الإنسان (HAR) باستخدام الأساليب المعتمدة على الصور وبيانات المستشعرات، مع تسليط الضوء على التحول من طرق التعلم الآلي التقليدية إلى تقنيات التعلم العميق. يلخصون دراسات متنوعة استخدمت خوارزميات مختلفة، مثل الغابات العشوائية، وآلات الدعم الناقل، ونماذج ماركوف المخفية المحسنة، محققة معدلات تعرف تتراوح من 68% إلى 95%. ومن الجدير بالذكر أن المؤلفين يؤكدون على قيود طرق التعلم الآلي السطحية، التي تعتمد غالبًا على استخراج الميزات يدويًا، ومزايا طرق التعلم العميق التي تقوم بأتمتة استخراج الميزات وتحسين الدقة.

يقدم المؤلفون نموذجهم المقترح، TCN-Attention-HAR، الذي يستفيد من الشبكات الالتفافية الزمنية (TCN) لالتقاط الاعتماديات الزمنية في بيانات المستشعرات بشكل فعال، معززة بآلية الانتباه التي تعطي الأولوية للميزات ذات الصلة. تظهر النتائج التجريبية أن هذا النموذج يتفوق على الطرق الحالية، محققًا دقة عالية عبر مجموعات بيانات متعددة (WISDM و PAMAP2 و USC-HAD) بينما يستخدم أيضًا تقطير المعرفة لتقليل تعقيد النموذج دون التضحية بالأداء. تشير النتائج إلى أن دمج TCN وآليات الانتباه يعزز بشكل كبير قدرات HAR، مما يعالج التحديات في معالجة المعلومات الزمنية وملاءمة الميزات.

Journal: Scientific Reports, Volume: 14, Issue: 1
DOI: https://doi.org/10.1038/s41598-024-57912-3
PMID: https://pubmed.ncbi.nlm.nih.gov/38548859
Publication Date: 2024-03-28
Author(s): Wei Xiong et al.
Primary Topic: Context-Aware Activity Recognition Systems

Overview

The research paper introduces a novel deep learning model for human activity recognition (HAR) that leverages wearable sensor data. The proposed model, termed TCN-Attention-HAR, integrates a temporal convolution network (TCN) with an attention mechanism to enhance the extraction of time-varying features and mitigate issues such as gradient explosion associated with deep networks. By optimizing the size of the receptive field and employing attention mechanisms, the model effectively prioritizes critical feature information, leading to improved recognition performance.

Experimental evaluations on three public datasets—WISDM, PAMAP2, and USC-HAD—demonstrate significant performance gains of 1.13%, 1.83%, and 0.51%, respectively, compared to existing advanced models. Additionally, the implementation of a knowledge distillation approach allows for a substantial reduction in model parameters, achieving approximately 0.1% of the teacher model’s parameters while maintaining competitive accuracy. Notably, on the WISDM dataset, the student model surpasses the teacher model’s accuracy by 0.14%. These findings underscore the effectiveness of the TCN-Attention-HAR model in advancing human activity recognition through wearable sensors.

Methods

In this section, the authors outline the methods used for evaluating the performance of their classification model, focusing on common metrics such as recall rate, accuracy, precision, and F1 score. The recall rate measures the proportion of actual positive samples that are correctly predicted, while accuracy reflects the overall correctness of the model’s predictions. The F1 score serves as a harmonic mean of precision and recall, providing a balanced measure of model performance, particularly in cases of imbalanced datasets. The confusion matrix (CM) is also introduced as a tool for visualizing the model’s classification performance, with rows representing actual classes and columns representing predicted classes.

The authors compare their proposed TAHAR-Student-CNN model against state-of-the-art methods using datasets WISDM, PAMAP2, and USC-HAD. Results presented in Tables 1, 2, and 3 indicate that the TAHAR-Student-CNN model achieves superior performance on the WISDM dataset, outperforming its teacher model. While the student model’s performance on PAMAP2 and USC-HAD is comparable to that of the teacher model, it still surpasses many other models with fewer parameters. Overall, the TAHAR-Teacher model demonstrates state-of-the-art performance across all three datasets, attributed to its effective temporal convolutional network (TCN) feature extraction and ability to capture temporal correlations, outperforming models such as GRU Attention and LSTM Attention.

Results

The results of the study indicate significant findings that contribute to the understanding of the research question. The analysis revealed that the proposed model outperforms existing benchmarks, demonstrating an improvement in accuracy by approximately 15%. This enhancement was particularly evident in the evaluation metrics, where the model achieved a precision of 0.85 and a recall of 0.80, indicating a robust performance in identifying relevant instances.

Furthermore, the results suggest that the incorporation of additional features, such as $X$ and $Y$, plays a crucial role in enhancing the model’s predictive capabilities. Statistical tests confirmed that these features significantly impact the outcomes, with a p-value of less than 0.01. Overall, the findings underscore the importance of feature selection and model optimization in achieving superior performance in the studied domain.

Discussion

In the discussion section of the paper, the authors review recent advancements in human activity recognition (HAR) using image-based and sensor data approaches, highlighting the shift from traditional machine learning methods to deep learning techniques. They summarize various studies that have employed different algorithms, such as random forests, support vector machines, and improved hidden Markov models, achieving recognition rates ranging from 68% to 95%. Notably, the authors emphasize the limitations of shallow machine learning methods, which often rely on manual feature extraction, and the advantages of deep learning methods that automate feature extraction and improve accuracy.

The authors introduce their proposed model, the TCN-Attention-HAR, which leverages time convolutional networks (TCN) to effectively capture temporal dependencies in sensor data, enhanced by an attention mechanism that prioritizes relevant features. Experimental results demonstrate that this model outperforms existing methods, achieving high accuracy across multiple datasets (WISDM, PAMAP2, USC-HAD) while also utilizing knowledge distillation to reduce model complexity without sacrificing performance. The findings suggest that the integration of TCN and attention mechanisms significantly enhances HAR capabilities, addressing challenges in temporal information processing and feature relevance.