شبكة CNN ثلاثية الأبعاد خفيفة الوزن للتعرف على أفعال الإنسان المعتمدة على الأحداث مع إمكانية الحفاظ على الخصوصية
A Lightweight 3D-CNN for Event-Based Human Action Recognition With Privacy-Preserving Potential

شارك:
المجلة: IEEE Access، المجلد: 14
DOI: https://doi.org/10.1109/access.2026.3660117
تاريخ النشر: 2026-01-01
المؤلف: Mehdi Sefidgar Dilmaghani وآخرون
الموضوع الرئيسي: التعرف على وضع الجسم والحركة

نظرة عامة

تقدم هذه الورقة البحثية شبكة عصبية تلافيفية ثلاثية الأبعاد خفيفة الوزن (3D-CNN) مصممة للتعرف على نشاط الإنسان (HAR) باستخدام بيانات الرؤية المعتمدة على الأحداث، والتي توفر بديلاً يحافظ على الخصوصية مقارنة بالكاميرات التقليدية المعتمدة على الإطارات. النموذج المقترح يلتقط الديناميات المكانية والزمانية بفعالية مع الحفاظ على هيكل مدمج مناسب للنشر على الأجهزة الطرفية. للتخفيف من عدم توازن الفئات وتحسين التعميم، يقوم المؤلفون بتنفيذ خسارة مركزة مع إعادة وزن الفئات واستراتيجيات تعزيز البيانات المستهدفة. تم تدريب النموذج على مجموعة بيانات مركبة من مجموعة بيانات منزل تويوتا الذكي وETRI، محققًا درجة F1 تبلغ 0.9415 ودقة إجمالية تبلغ 94.17%، متجاوزًا الهياكل المعروفة لشبكات 3D-CNN مثل C3D وResNet3D وMC3_18 بنسبة تصل إلى 3%.

في الختام، تُظهر بنية 3D-CNN الخفيفة الوزن مزايا كبيرة مقارنة بأساليب HAR التقليدية، خاصة في معالجة مخاوف الخصوصية المرتبطة بالأنظمة المعتمدة على الإطارات. التصميم المدمج للنموذج، جنبًا إلى جنب مع التعامل الفعال مع عدم توازن الفئات وكفاءة التدريب العالية، يجعله مناسبًا للنشر على الأجهزة الطرفية. تشير النتائج التجريبية إلى قدرات قوية على التعميم، خاصة في تمييز الأنشطة المتشابهة، ووقت التدريب القصير للنموذج يعزز من عمليته. ستستكشف الأبحاث المستقبلية آليات انتباه متقدمة وطرق الشبكات العصبية المتفجرة لتحسين دقة التعرف وكفاءة الحساب. بشكل عام، تقدم هذه العمل حلاً قويًا يحافظ على الخصوصية للتعرف على النشاط في الوقت الحقيقي، مع تطبيقات محتملة في الرعاية الصحية والمراقبة والبيئات الذكية.

مقدمة

تسلط مقدمة هذه الورقة البحثية الضوء على أهمية التعرف على نشاط الإنسان (HAR) في تطبيقات متنوعة، خاصة في تعزيز رعاية الفئات الضعيفة مع معالجة مخاوف الخصوصية المرتبطة بأساليب المراقبة التقليدية. يعتمد HAR على فهم الديناميات المكانية والزمانية للأفعال، مما يقدم تحديات في تكنولوجيا المستشعرات وكفاءة الخوارزميات. تستخدم أنظمة HAR الحالية بشكل أساسي الكاميرات المعتمدة على الإطارات، مما يؤدي إلى متطلبات حسابية كبيرة ومشاكل في الخصوصية بسبب تسجيل المعلومات الشخصية في الإعدادات الخاصة.

للتغلب على هذه القيود، تقترح الورقة استخدام كاميرات الأحداث، التي تلتقط فقط التغيرات في كثافة البكسل، مما يحافظ على الخصوصية بينما يمكّن من التعرف الفعال على الأفعال. يشير المؤلفون إلى أن أساليب HAR الحالية تركز بشكل أساسي على الشبكات العصبية التلافيفية (CNNs) المصممة للبيانات المعتمدة على الإطارات، والتي تكافح مع المعلومات الزمنية وغالبًا ما تتطلب هياكل معقدة. استجابةً لذلك، تقدم الورقة شبكة عصبية تلافيفية ثلاثية الأبعاد (3D-CNN) مصممة لبيانات كاميرات الأحداث، قادرة على التقاط كل من الميزات المكانية والزمانية بكفاءة. يهدف النموذج المقترح إلى تقديم حلول HAR دقيقة ومدمجة وتحافظ على الخصوصية، مع معالجة الحاجة المتزايدة لخوارزميات خفيفة الوزن مناسبة للتطبيقات في الوقت الحقيقي.

طرق

في قسم المنهجية، توضح الورقة خوارزمية جديدة للتعرف على نشاط الإنسان (HAR) باستخدام شبكة عصبية تلافيفية ثلاثية الأبعاد (3D-CNN). يتميز هذا النهج عن الأساليب التقليدية، التي اعتمدت على ميزات مكانية وزمانية مصممة يدويًا وكانت حساسة لمجموعة متنوعة من العوامل البيئية. تم تصميم 3D-CNN خصيصًا لتحليل كل من المعلومات المكانية والزمانية في وقت واحد، مما يجعلها فعالة بشكل خاص في معالجة تدفقات الأحداث الخام الناتجة عن المستشعرات العصبية. تسهل البنية الخفيفة للشبكة تدريبًا أسرع ونشرًا أسهل على الأجهزة الطرفية، بينما تقدم أيضًا حماية محسّنة للخصوصية مقارنة بالكاميرات التقليدية المعتمدة على الإطارات.

تشمل المنهجية المقترحة خط أنابيب منظم يتضمن معالجة البيانات المدخلة إلى مصفوفات ثنائية الأبعاد، والتي تشبه الإطارات، قبل إدخالها في نموذج 3D-CNN. لتحسين تمثيل الميزات، يتم دمج آلية الانتباه الذاتي جنبًا إلى جنب مع الطبقات التلافيفية. بالإضافة إلى ذلك، تستخدم الخوارزمية خسارة مركزة واستراتيجيات إعادة وزن الفئات لمعالجة عدم توازن الفئات في بيانات التدريب. بشكل عام، تهدف المنهجية إلى الاستفادة من نقاط القوة في تقنيات التعلم العميق الحديثة مع معالجة قيود أساليب HAR السابقة.

نتائج

في قسم النتائج، أظهر النموذج المقترح مقاييس أداء مثيرة للإعجاب، محققًا دقة اختبار تبلغ 94.17% ودرجة F1 للاختبار تبلغ 0.9415 على بيانات غير مرئية، مما يشير إلى قدرات قوية على التعميم. سجل النموذج الأمثل، المحفوظ عند العصر 648، درجة F1 للتحقق تبلغ 0.9409، مما يشير إلى القوة دون الإفراط في التكيف. استغرق عملية التدريب 322 دقيقة، وهو أكثر كفاءة من اثنين من ثلاثة هياكل 3D-CNN المقارنة، مما يجعله مناسبًا للتطبيقات في الوقت الحقيقي. تشير رسوم بيانية لخسارة التدريب والتحقق إلى اتجاه ثابت دون الإفراط في التكيف بشكل كبير، مما يؤكده أيضًا مصفوفة الالتباس، التي تظهر معدلات إيجابية حقيقية عالية عبر جميع فئات النشاط، بما في ذلك الطهي، والشرب، والأكل، والقيام، والجلوس، وغسل الأطباق.

تكشف النتائج النوعية عن حالات من التصنيف الخاطئ، كما هو موضح في الشكل 4، حيث يكافح حتى المراقبون البشر لتحديد الأنشطة الصحيحة من الإطارات الفيديو المقدمة. على الرغم من هذه التصنيفات الخاطئة، فإن الأداء العام للنموذج وقدرته على التقاط الميزات المكانية والزمانية بدقة عبر فئات متشابهة بصريًا أو زمانيًا جديرة بالملاحظة، مما يبرز فعاليتها في التعامل مع مهام التعرف على الأنشطة المعقدة.

مناقشة

تسلط قسم المناقشة في الورقة البحثية الضوء على تطور والتحديات الحالية في التعرف على نشاط الإنسان (HAR) ضمن مجال رؤية الكمبيوتر. يؤكد على الانتقال من الميزات المصممة تقليديًا إلى أساليب التعلم العميق، خاصة استخدام هياكل متنوعة مثل CNNs ذات التدفقين، وشبكات القطاعات الزمنية (TSN)، وConvNets ثلاثية الأبعاد المنفوخة (I3D). بينما حققت هذه النماذج دقة مثيرة للإعجاب على المعايير القياسية – حيث أبلغت عن دقة تصل إلى 94.9% على UCF-101 – إلا أنها غالبًا ما تتطلب موارد حسابية كبيرة وتثير مخاوف كبيرة بشأن الخصوصية بسبب اعتمادها على البيانات المرئية التي تحتوي على معلومات شخصية. وهذا يستدعي استكشاف طرق بديلة، مثل الأنظمة المعتمدة على RF والرؤية المعتمدة على الأحداث، التي تقدم دقة تنافسية دون المساس بخصوصية المستخدم.

تحدد الورقة الفجوات الحرجة في الأدبيات، خاصة فيما يتعلق بالتوازن بين الخصوصية والدقة، وكفاءة النمذجة المكانية والزمانية، والحاجة إلى مجموعات بيانات أكبر وأكثر واقعية للتعرف على نشاط الإنسان المعتمد على الأحداث. على الرغم من أن طرق RF والرادار تظهر وعدًا في الحفاظ على الخصوصية مع تحقيق دقة معقولة، إلا أنها لا تزال تكافح مع التعرف التفصيلي على الأفعال عبر مشاهد متنوعة. علاوة على ذلك، بينما اكتسبت الأساليب المعتمدة على الأحداث زخمًا، غالبًا ما تتضمن نماذج معقدة تتطلب حسابات كثيفة، مما يحد من نشرها على الأجهزة الطرفية. يدعو المؤلفون إلى مزيد من البحث في الهياكل الخفيفة الوزن ودمج عدة أنماط استشعار لتعزيز كل من الأداء والخصوصية في تطبيقات HAR.

القيود

تحدد الأبحاث عدة قيود قد تؤثر على أداء النموذج المقترح. أولاً، تتعرض فعالية النموذج للخطر في السيناريوهات التي تتميز بعدم توازن الفئات الشديد، مما قد يؤدي إلى نتائج تصنيف منحازة. بالإضافة إلى ذلك، فإن تصنيف أنماط الحركة الدقيقة أو المتداخلة يمثل تحديًا كبيرًا، مما يشير إلى أن النموذج قد يكافح للتمييز بين الفئات ذات الصلة الوثيقة بشكل فعال.

علاوة على ذلك، بينما يوفر دمج آليات الانتباه الذاتي بعض التحسينات في الأداء، فإن هذه المكاسب هامشية. كما أن تنفيذ نماذج المحولات الكاملة مقيد أكثر بسبب قيود الموارد الحسابية، مما يشير إلى أن قابلية التوسع والكفاءة تظل اعتبارات حاسمة للتطورات المستقبلية في هذا المجال.

Journal: IEEE Access, Volume: 14
DOI: https://doi.org/10.1109/access.2026.3660117
Publication Date: 2026-01-01
Author(s): Mehdi Sefidgar Dilmaghani et al.
Primary Topic: Human Pose and Action Recognition

Overview

This research paper introduces a lightweight three-dimensional convolutional neural network (3D-CNN) designed for human activity recognition (HAR) using event-based vision data, which offers a privacy-preserving alternative to traditional frame-based cameras. The proposed model effectively captures spatial and temporal dynamics while maintaining a compact architecture suitable for edge deployment. To mitigate class imbalance and improve generalization, the authors implement focal loss with class reweighting and targeted data augmentation strategies. The model is trained on a composite dataset from the Toyota Smart Home and ETRI datasets, achieving an F1-score of 0.9415 and an overall accuracy of 94.17%, surpassing established 3D-CNN architectures such as C3D, ResNet3D, and MC3_18 by up to 3%.

In conclusion, the lightweight 3D-CNN architecture demonstrates significant advantages over traditional HAR methods, particularly in addressing privacy concerns associated with frame-based systems. The model’s compact design, combined with effective handling of class imbalance and high training efficiency, makes it suitable for deployment on edge devices. Experimental results indicate strong generalization capabilities, especially in distinguishing similar activities, and the model’s short training time further enhances its practicality. Future research will explore advanced attention mechanisms and spiking neural network approaches to improve recognition accuracy and computational efficiency. Overall, this work presents a robust, privacy-preserving solution for real-time HAR, with potential applications in healthcare, surveillance, and smart environments.

Introduction

The introduction of this research paper highlights the significance of Human Action Recognition (HAR) in various applications, particularly in enhancing the care of vulnerable populations while addressing privacy concerns associated with traditional monitoring methods. HAR relies on understanding both spatial and temporal dynamics of actions, which presents challenges in sensor technology and algorithm efficiency. Current HAR systems predominantly utilize frame-based cameras, leading to substantial computational demands and privacy issues due to the recording of personal information in private settings.

To overcome these limitations, the paper proposes the use of event cameras, which capture only changes in pixel intensity, thus preserving privacy while enabling effective action recognition. The authors note that existing HAR methods primarily focus on convolutional neural networks (CNNs) designed for frame-based data, which struggle with temporal information and often require complex architectures. In response, the paper introduces a three-dimensional convolutional neural network (3D-CNN) tailored for event camera data, capable of capturing both spatial and temporal features efficiently. The proposed model aims to deliver accurate, compact, and privacy-preserving HAR solutions, addressing the growing need for lightweight algorithms suitable for real-time applications.

Methods

In the section on methodology, the paper outlines a novel algorithm for human activity recognition (HAR) utilizing a 3D Convolutional Neural Network (3D-CNN). This approach is distinguished from traditional methods, which relied on hand-crafted spatio-temporal features and were sensitive to various environmental factors. The 3D-CNN is specifically designed to analyze both spatial and temporal information simultaneously, making it particularly effective for processing raw event streams generated by neuromorphic sensors. The lightweight architecture of the network facilitates faster training and easier deployment on edge devices, while also offering enhanced privacy protection compared to conventional frame-based cameras.

The proposed methodology involves a structured pipeline that includes preprocessing the input data into 2D matrices, which resemble frames, before feeding them into the 3D-CNN model. To improve feature representation, a self-attention mechanism is integrated alongside the convolutional layers. Additionally, the algorithm employs focal loss and class reweighting strategies to address class imbalance in the training data. Overall, the methodology aims to leverage the strengths of modern deep learning techniques while addressing the limitations of earlier HAR approaches.

Results

In the results section, the proposed model demonstrated impressive performance metrics, achieving a test accuracy of 94.17% and a test F1-score of 0.9415 on unseen data, indicating strong generalization capabilities. The optimal model, saved at epoch 648, recorded a validation F1-score of 0.9409, suggesting robustness without overfitting. The training process took 322 minutes, which is more efficient than two of the three benchmarked 3D-CNN architectures, making it suitable for real-time applications. The training and validation loss graphs indicate a consistent trend without significant overfitting, further corroborated by the confusion matrix, which shows high true positive rates across all activity classes, including Cooking, Drinking, Eating, Getting up, Sitting down, and Washing up the dishes.

Qualitative results reveal instances of misclassification, as illustrated in Figure 4, where even human observers struggle to identify the correct activities from the presented video frames. Despite these misclassifications, the model’s overall performance and ability to accurately capture spatiotemporal features across visually or temporally similar classes are noteworthy, highlighting its effectiveness in handling complex activity recognition tasks.

Discussion

The discussion section of the research paper highlights the evolution and current challenges in Human Action Recognition (HAR) within the field of computer vision. It emphasizes the transition from traditional engineered features to deep learning approaches, particularly the use of various architectures such as Two-Stream CNNs, Temporal Segment Networks (TSN), and Inflated 3D ConvNets (I3D). While these models have achieved impressive accuracy on standard benchmarks—reporting up to 94.9% on UCF-101—they often require substantial computational resources and raise significant privacy concerns due to their reliance on visual data containing personal information. This necessitates the exploration of alternative methods, such as RF-based systems and event-based vision, which offer competitive accuracy without compromising user privacy.

The paper identifies critical gaps in the literature, particularly regarding the trade-offs between privacy and accuracy, the efficiency of spatio-temporal modeling, and the need for larger, more realistic datasets for event-based HAR. Although RF and radar methods show promise in maintaining privacy while achieving reasonable accuracy, they still struggle with detailed action recognition across diverse scenes. Furthermore, while event-based approaches have gained traction, they often involve complex models that are computationally demanding, thus limiting their deployment on edge devices. The authors call for further research into lightweight architectures and the integration of multiple sensing modalities to enhance both performance and privacy in HAR applications.

Limitations

The research identifies several limitations that may affect the performance of the proposed model. Firstly, the model’s efficacy is compromised in scenarios characterized by extreme class imbalance, which can lead to skewed classification results. Additionally, the classification of subtle or overlapping motion patterns presents a significant challenge, indicating that the model may struggle to differentiate between closely related classes effectively.

Moreover, while the incorporation of self-attention mechanisms provides some improvements in performance, these gains are marginal. The implementation of full transformer models is further constrained by computational resource limitations, suggesting that scalability and efficiency remain critical considerations for future developments in this area.

شارك: