DOI: https://doi.org/10.3389/fncom.2025.1543643
PMID: https://pubmed.ncbi.nlm.nih.gov/40386804
تاريخ النشر: 2025-05-02
المؤلف: Zhizhong Xing وآخرون
الموضوع الرئيسي: إعادة تأهيل السكتة الدماغية والتعافي
نظرة عامة
تتناول هذه الدراسة القضية الحرجة لإعادة تأهيل وظيفة اليد لكبار السن، لا سيما في سياق تدابير الوقاية من الأوبئة المستمرة وشيخوخة السكان. تقدم طريقة جديدة تستخدم نموذج تعلم عميق ثلاثي الأبعاد لتحليل بيانات سحابة النقاط من أجهزة الاستشعار بالليزر لتحليل ميزات سطح الإيماءات بدون تلامس. تدمج الدراسة تقنيات أساسية، بما في ذلك جمع سحابة نقاط سطح اليد واستخراج الميزات المحلية، مما أدى إلى نظام تحليل الإيماءات الذي حقق دقة تعرف متوسطة مثيرة للإعجاب بلغت 88.72%. تؤكد هذه النتائج على إمكانيات تقنيات إعادة التأهيل الذكية غير التلامسية لتعزيز سلامة وراحة التفاعل للمستخدمين من كبار السن ومرضى إعادة التأهيل.
تتفصل الدراسة أيضًا في بناء وأداء النموذج المقترح، المسمى GSFAN، الذي يلتقط بفعالية الهياكل المكانية والتفاصيل المحلية للإيماءات من خلال الالتفاف الحدي. لوحظ الأداء الأمثل مع ثلاث طبقات من الالتفاف الحدي، وكان البعد 128 لسحابة نقاط سطح اليد قد حسّن بشكل كبير من دقة التعرف. أظهرت التجارب المقارنة تفوق GSFAN على نماذج معالجة سحابة النقاط الحالية، محققة دقة أعلى ومعدلات ارتباك أقل في تصنيف 32 سحابة نقاط سطح اليد. تسلط النتائج الضوء على القيمة العملية لـ GSFAN في إعادة التأهيل الذكي، مع خطط للبحث المستقبلي لتقييم ومقارنة النماذج المتقدمة في معالجة سحابة النقاط لتعزيز قوة النتائج.
مقدمة
تتناول مقدمة هذه الورقة البحثية التحدي المزدوج الحرج المتمثل في تعزيز الصحة العامة ورفاهية كبار السن ومرضى إعادة التأهيل في ظل تطبيع تدابير الوقاية من الأوبئة العالمية وشيخوخة السكان. تمثل طرق إعادة التأهيل التقليدية، التي تعتمد غالبًا على العمليات القائمة على التلامس، مخاطر العدوى وغير كافية لتلبية الطلب المتزايد على إعادة تأهيل وظيفة اليد. يقترح المؤلفون نظام تعرف على الإيماءات غير التلامسي الذي يستفيد من بيانات سحابة النقاط من أجهزة الاستشعار بالليزر ونماذج التعلم العميق ثلاثية الأبعاد لتسهيل تكنولوجيا إعادة التأهيل الذكية. يهدف هذا النظام إلى تحسين كفاءة علاجات إعادة التأهيل مع تقليل الاعتماد على المعالجين المحترفين.
تستخدم طريقة التعرف على الإيماءات المقترحة تقنيات متقدمة مثل جهاز LEAP وMediaPipe، التي تتفوق في التقاط وتحليل حركات اليد. من خلال استخدام أجهزة الاستشعار بالليزر للقياس غير التلامسي عالي الدقة، تسعى الدراسة إلى التغلب على القيود في طرق تقييم إعادة التأهيل الحالية، التي تعتمد غالبًا على التقييمات الذاتية من قبل المعالجين. يسمح دمج تقنيات معالجة بيانات سحابة النقاط، بما في ذلك PointNet وPointNet++، باستخراج فعال لميزات سطح الإيماءات، مما يعالج تعقيدات تحليل الإيماءات. في النهاية، تهدف هذه الدراسة إلى توفير أداة قوية لمؤسسات إعادة التأهيل، مما يعزز دقة وكفاءة العلاج مع تقليل مخاطر العدوى، وفتح التطبيقات المحتملة للواقع المعزز وواجهات التحكم بالإيماءات.
طرق
في هذا القسم، يوضح المؤلفون الطرق المستخدمة في التعرف على الإيماءات في أنظمة إعادة التأهيل الافتراضية لوظيفة اليد، مع التأكيد على التحديات التي تطرحها درجة الحرية العالية في مفاصل اليد البشرية. يقدمون شبكة تحليل ميزات سطح الإيماءات (GSFAN) المستندة إلى DGCNN، التي تستخدم بيانات سحابة النقاط – تمثيلات ثلاثية الأبعاد للإيماءات اليدوية التي تم الحصول عليها عبر أجهزة قياس غير تلامسية مثل أجهزة الاستشعار بالليزر. تم تقسيم مجموعة بيانات سحابة نقاط سطح اليد إلى مجموعات تدريب واختبار، مع ضمان العشوائية والتوازن بين فئات الإيماءات لتقليل التحيز في GSFAN.
كانت تقييم أداء النموذج يعتمد بشكل أساسي على الدقة، وهو أمر ذو صلة خاصة لتقييم موثوقية التعرف على الإيماءات في سياقات إعادة التأهيل لكبار السن. استخدم المؤلفون أيضًا مصفوفات الارتباك لتحليل مفصل لنتائج التصنيف. لتعزيز تعميم النموذج، اختبروا GSFAN مع سحابات نقاط سطح اليد المتنوعة التي تم جمعها من مختلف المجربين. تم إجراء التجارب على إعداد حوسبة عالي الأداء، بما في ذلك معالج Intel Core i9 من الجيل الثالث عشر وبطاقة الرسوميات NVIDIA GeForce RTX 4080، باستخدام Python وTensorFlow لتنفيذات التعلم العميق. تم اختيار دالة تفعيل Leaky ReLU لمعالجة القضايا المتعلقة بالتدرجات المتلاشية، مما يضمن بقاء الخلايا العصبية نشطة أثناء التدريب.
نتائج
تسلط نتائج الدراسة الضوء على التأثير الكبير لعدد طبقات الالتفاف الحدي وأبعاد بيانات سحابة نقاط سطح اليد على أداء التعرف على الإيماءات. من خلال سلسلة من التجارب المضبوطة، وُجد أن زيادة عدد طبقات الالتفاف الحدي من واحدة إلى ثلاث يعزز قدرة النموذج على استخراج الميزات، مما يحسن دقة التعرف. ومع ذلك، يمكن أن يؤدي عدد مفرط من الطبقات إلى تدهور الأداء بسبب الإفراط في تنعيم الميزات، مما يقلل من قوة التمييز للنموذج. لوحظ الأداء الأمثل مع ثلاث طبقات، حيث لم تؤدِ الزيادات الإضافية إلى فوائد كبيرة وأضفت تعقيدًا حسابيًا.
بالإضافة إلى ذلك، فحصت الدراسة آثار تعزيز التجريد البعدي على دقة التعرف. أشارت النتائج إلى أنه مع زيادة أبعاد بيانات سحابة النقاط من 32 إلى 128، تحسن أداء النموذج بسبب تعزيز قدرات التمثيل. ومع ذلك، بعد حد معين، أدت الزيادات الإضافية في الأبعاد إلى عوائد متناقصة وإمكانية الإفراط في التكيف، مما أدى إلى انخفاض الدقة. تم تحديد الأبعاد المثلى للتعرف الفعال على الإيماءات حول 128، مما يشير إلى أن نهجًا متوازنًا لكل من طبقات الالتفاف الحدي والأبعاد أمر حاسم لتعظيم أداء نماذج التعرف على الإيماءات في سياقات التفاعل بين الإنسان والآلة.
مناقشة
في هذا البحث، يتم التأكيد على جمع ومعالجة سحابات نقاط سطح اليد باستخدام تكنولوجيا الكشف بالليزر وتحديد المدى كخطوة أساسية لتحليل ميزات اليد ذات الصلة بإعادة التأهيل الوظيفي. شملت الدراسة التقاط سحابات النقاط من أيدي أربعة مشاركين تحت ظروف إضاءة مضبوطة لضمان جودة البيانات. تم تبرير استخدام رادار الليزر Direct Time of Flight (DToF) بسبب طبيعته غير التلامسية، ودقته العالية، ومتانته ضد التداخل البيئي. أدت عملية جمع البيانات إلى مجموعة بيانات متنوعة من 32 إيماءة متميزة، تم تحقيقها من خلال تغيير أوضاع وأحوال الأصابع. بعد الجمع، خضعت بيانات سحابة النقاط الخام لمعالجة دقيقة باستخدام CloudCompare لإزالة الضوضاء والمعلومات غير ذات الصلة، مما يضمن إدخال بيانات عالية الجودة للتحليلات اللاحقة.
تقدم الدراسة نموذجًا جديدًا، GSFAN، الذي يستخدم الالتفاف الحدي لالتقاط الهياكل المحلية والميزات الهندسية من سحابات نقاط سطح اليد بفعالية. من خلال تمثيل سحابة النقاط كرسوم بيانية غير موجهة، يقوم GSFAN بتجميع الميزات من العقد المجاورة لتعزيز تمثيل كل نقطة. تشمل بنية النموذج شبكة عصبية متعددة الطبقات تقوم بمزيد من تجريد أبعاد البيانات، مما يحسن دقة التعرف على الإيماءات. أظهرت النتائج التجريبية أن GSFAN تفوق على النماذج الحالية مثل PointNet وPointNet++ من حيث الدقة، محققًا معدل تعرف متوسط يبلغ 88.72%. تشير النتائج إلى أن قدرة GSFAN على التقاط التفاصيل المحلية وتعزيز المعلومات البعدية تسهم بشكل كبير في أدائه المتفوق في التعرف على إيماءات اليد، مما يوفر تداعيات واعدة للتفاعل غير التلامسي بين الإنسان والآلة في سياقات إعادة التأهيل. ستركز الأعمال المستقبلية على معالجة القيود المتعلقة بحجم العينة والأداء في الوقت الحقيقي، بالإضافة إلى استكشاف آليات التدريب الشخصية لتكييف النموذج مع المستخدمين الفرديين.
DOI: https://doi.org/10.3389/fncom.2025.1543643
PMID: https://pubmed.ncbi.nlm.nih.gov/40386804
Publication Date: 2025-05-02
Author(s): Zhizhong Xing et al.
Primary Topic: Stroke Rehabilitation and Recovery
Overview
This research addresses the critical issue of hand function rehabilitation for the elderly, particularly in the context of ongoing epidemic prevention measures and an aging population. It introduces a novel method utilizing a 3D deep learning model to analyze laser sensor point cloud data for non-contact gesture surface feature analysis. The study integrates essential technologies, including hand surface point cloud collection and local feature extraction, resulting in a gesture analysis system that achieved an impressive average recognition accuracy of 88.72%. These findings underscore the potential of non-contact intelligent rehabilitation technologies to enhance interaction safety and comfort for elderly users and rehabilitation patients.
The research further details the construction and performance of the proposed model, termed GSFAN, which effectively captures spatial structures and local details of gestures through edge convolution. Optimal performance was observed with three edge convolution layers, and a dimensionality of 128 for the hand surface point cloud significantly improved recognition accuracy. Comparative experiments demonstrated GSFAN’s superiority over existing point cloud processing models, achieving higher accuracy and lower confusion rates in classifying 32 hand surface point clouds. The results highlight GSFAN’s practical value in intelligent rehabilitation, with plans for future research to systematically evaluate and compare advanced models in point cloud processing to enhance the robustness of findings.
Introduction
The introduction of this research paper addresses the critical dual challenge of enhancing public health and the well-being of elderly and rehabilitation patients amidst the normalization of global epidemic prevention measures and an aging population. Traditional rehabilitation methods, which often rely on contact-based operations, pose risks of infection and are insufficient to meet the growing demand for hand function rehabilitation. The authors propose a non-contact gesture recognition system that leverages laser sensor point cloud data and 3D deep learning models to facilitate intelligent rehabilitation technology. This system aims to improve the efficiency of rehabilitation treatments while reducing reliance on professional therapists.
The proposed gesture recognition method utilizes advanced technologies such as the LEAP device and MediaPipe, which excel in capturing and analyzing hand movements. By employing laser sensors for high-precision, non-contact measurement, the research seeks to overcome limitations in existing rehabilitation assessment methods, which often depend on subjective evaluations by therapists. The integration of point cloud data processing techniques, including PointNet and PointNet++, allows for effective extraction of gesture surface features, addressing the complexities of gesture analysis. Ultimately, this research aims to provide a robust tool for rehabilitation institutions, enhancing treatment accuracy and efficiency while minimizing infection risks, and potentially extending applications to augmented reality and gesture control interfaces.
Methods
In this section, the authors outline the methods employed for gesture recognition in hand function virtual rehabilitation systems, emphasizing the challenges posed by the high degree of freedom in human hand joints. They introduce a gesture surface feature analysis network (GSFAN) based on DGCNN, which utilizes point cloud data—3D representations of hand gestures obtained via non-contact measurement devices like laser sensors. The dataset of hand surface point clouds was divided into training and testing sets, ensuring randomness and balance among gesture categories to mitigate bias in GSFAN.
The evaluation of the model’s performance was primarily based on accuracy, which is particularly relevant for assessing the reliability of gesture recognition in rehabilitation contexts for the elderly. The authors also employed confusion matrices for a detailed analysis of classification results. To enhance the model’s generalization, they tested GSFAN with diverse hand surface point clouds collected from different experimenters. The experiments were conducted on a high-performance computing setup, including a 13th Gen Intel Core i9 processor and an NVIDIA GeForce RTX 4080 graphics card, utilizing Python and TensorFlow for deep learning implementations. The Leaky ReLU activation function was chosen to address issues related to vanishing gradients, ensuring that neurons remain active during training.
Results
The results of the study highlight the significant impact of the number of edge convolution layers and the dimensionality of hand surface point cloud data on gesture recognition performance. Through a series of controlled experiments, it was found that increasing the number of edge convolution layers from one to three enhances the model’s ability to extract features, thereby improving recognition accuracy. However, an excessive number of layers can lead to performance degradation due to over-smoothing of features, which diminishes the model’s discriminative power. The optimal performance was observed with three layers, beyond which further increases did not yield significant benefits and added computational complexity.
Additionally, the study examined the effects of dimensional abstraction enhancement on recognition accuracy. Results indicated that as the dimensionality of the point cloud data increased from 32 to 128, the model’s performance improved due to enhanced representation capabilities. However, beyond a certain threshold, further increases in dimensionality led to diminishing returns and potential overfitting, resulting in decreased accuracy. The optimal dimensionality for effective gesture recognition was identified around 128, suggesting that a balanced approach to both edge convolution layers and dimensionality is crucial for maximizing the performance of gesture recognition models in human-machine interaction contexts.
Discussion
In this research, the collection and processing of hand surface point clouds using laser detection and ranging technology is emphasized as a foundational step for analyzing hand features relevant to functional rehabilitation. The study involved capturing point clouds from the hands of four participants under controlled lighting conditions to ensure data quality. The use of Direct Time of Flight (DToF) laser radar was justified due to its non-contact nature, high precision, and robustness against environmental interference. The data collection process resulted in a diverse dataset of 32 distinct gestures, achieved by varying finger postures and states. Post-collection, the raw point cloud data underwent rigorous processing using CloudCompare to eliminate noise and irrelevant information, ensuring high-quality input for subsequent analyses.
The research introduces a novel model, GSFAN, which employs edge convolution to effectively capture local structures and geometric features from the hand surface point clouds. By representing the point cloud as an undirected graph, GSFAN aggregates features from neighboring nodes to enhance the representation of each point. The model’s architecture includes a multilayer perceptron that further abstracts the dimensionality of the data, improving gesture recognition accuracy. Experimental results demonstrated that GSFAN outperformed existing models like PointNet and PointNet++ in accuracy, achieving an average recognition rate of 88.72%. The findings suggest that GSFAN’s ability to capture local details and enhance dimensional information significantly contributes to its superior performance in hand gesture recognition, offering promising implications for non-contact human-machine interaction in rehabilitation contexts. Future work will focus on addressing limitations related to sample size and real-time performance, as well as exploring personalized training mechanisms to adapt the model to individual users.
