تقييم تجمد المشي باستخدام وحدات القياس القابلة للحركة والتعلم العميق: تأثير المهام، وحالات الأدوية، والتوقفات
Freezing of gait assessment with inertial measurement units and deep learning: effect of tasks, medication states, and stops

المجلة: Journal of NeuroEngineering and Rehabilitation، المجلد: 21، العدد: 1
DOI: https://doi.org/10.1186/s12984-024-01320-1
PMID: https://pubmed.ncbi.nlm.nih.gov/38350964
تاريخ النشر: 2024-02-13
المؤلف: Po-Kai Yang وآخرون
الموضوع الرئيسي: التوازن، والمشي، والوقاية من السقوط

نظرة عامة

تبحث ورقة البحث في نموذج تقييم آلي لتجميد المشي (FOG) في مرض باركنسون (PD) باستخدام التعلم العميق (DL) ووحدات القياس الحركية (IMUs). تعتمد طرق تقييم FOG التقليدية على الفحص البصري، مما يستغرق وقتًا طويلاً ويكون ذاتيًا. تتناول الدراسة التحديات المتعلقة بتنوع المشي بسبب تأثيرات الأدوية والحاجة إلى التمييز بين FOG والحركات النموذجية. تم استخدام شبكة التلافيف الزمنية (TCN) لتحليل بيانات من اثني عشر مريضًا مصابًا بمرض باركنسون خلال مهام مختلفة تحفز FOG تحت حالات دوائية مختلفة، مع وبدون توقف إرادي.

تشير النتائج إلى أن نموذج TCN أظهر توافقًا قويًا مع التعليقات من الخبراء، محققًا معامل ارتباط داخلي (ICC) قدره 0.92 بالنسبة لنسبة الوقت المجمد (%TF) و0.95 لعدد حلقات التجميد (#FOG) في التجارب بدون توقف. بينما واجهت النماذج المدربة على مهام محددة تحفز FOG صعوبة في التعميم على مهام جديدة، أظهرت النماذج المدربة على حالات الأدوية قدرة أفضل على التكيف. في التجارب التي شملت التوقف، حافظ النموذج على توافق معتدل (ICC %TF = 0.95؛ ICC #FOG = 0.79) فقط عندما كانت بيانات التوقف جزءًا من مجموعة التدريب. تشير النتائج إلى أن نموذج DL واحد يمكنه تقييم FOG بفعالية عبر مهام وحالات دوائية مختلفة، مما يمهد الطريق لأبحاث مستقبلية حول الكشف الآلي عن FOG في السيناريوهات الحياتية الواقعية.

مقدمة

يؤثر مرض باركنسون (PD) بشكل كبير على المشي، حيث يؤثر تجميد المشي (FOG) على حوالي 70% من المرضى. يتميز FOG بتوقف مؤقت أو تقليل في الحركة، يظهر بأشكال مختلفة مثل الارتعاش، والانزلاق، والجمود الكامل. يشكل هذا العرض مخاطر السقوط ويؤثر سلبًا على جودة حياة المرضى. غالبًا ما تفشل طرق التقييم الحالية، بما في ذلك الاستبيانات الذاتية والمهام القياسية مثل اختبار “القيام والذهاب” (TUG) والدوران بزاوية 360 درجة، في تقديم مقاييس موضوعية لشدة FOG بسبب تحيز الذاكرة وتنوع نتائج الاختبارات. وبالتالي، هناك اهتمام متزايد في الأساليب الآلية، وخاصة تقنيات التعلم العميق (DL)، لتعزيز الكشف عن FOG.

تتناول هذه الدراسة عدة فجوات في الأدبيات الحالية بشأن الكشف عن FOG باستخدام نماذج DL. تقدم نموذجًا جديدًا لتقييم FOG تلقائيًا استنادًا إلى وحدات القياس الحركية (IMUs) خلال مهام TUG و360Turn. تقيم الدراسة ما إذا كان التدريب المحدد للمهام أو الأدوية يحسن أداء الكشف وتقييم قدرة النموذج على التعميم على مهام أو حالات دوائية غير مرئية. بالإضافة إلى ذلك، تحقق من تأثير تضمين فترات التوقف الطوعية في بيانات التدريب على دقة الكشف عن FOG. تشير النتائج إلى أنه بينما يمكن للنموذج التعميم بفعالية عبر المهام وحالات الأدوية، فإنه يواجه صعوبة مع الظروف غير المرئية، مما يبرز ضرورة وجود بيانات تدريب متنوعة. تساهم الدراسة في المجال من خلال إثبات جدوى تقييم FOG التلقائي خلال مهمة 360Turn وإقامة توافق قوي مع التقييمات من الخبراء عبر ظروف مختلفة.

طرق البحث

في هذه الدراسة، هدفنا إلى تعزيز تقييم تجميد المشي (FOG) في البيئات السريرية حيث لا يتم عادةً تضمين التوقف. أولاً، قمنا بتقييم الأداء العام والنسبى لنموذج عام تم تدريبه على مهام تحفز FOG القياسية، مع استبعاد سيناريوهات التوقف. بعد ذلك، بحثنا في قدرة النموذج على التعميم عبر مهام مختلفة تحفز FOG وحالات دوائية مختلفة. شمل ذلك تحليل تأثير تضمين أو استبعاد بيانات التدريب من مهام أو حالات دوائية محددة على فعالية النموذج في الكشف عن FOG.

النتائج

يقدم قسم النتائج تحليلًا شاملاً لخصائص مجموعة البيانات وأداء النماذج المختلفة في الكشف عن تجميد المشي (FOG) تحت ظروف مختلفة. يتم فحص مستويين رئيسيين من تقييم FOG: القياسات السريرية والسيناريوهات الحياتية. تشير النتائج إلى أن نموذج_سريري، المدرب على مهام محددة تحفز FOG وحالات دوائية، حقق درجات متغيرة من Segment-F1 وSample-F1 عبر ظروف مختلفة. على سبيل المثال، حقق نموذج_TUG درجة Segment-F1 قدرها 0.70، بينما سجل نموذج_360Turn 0.53، مع ملاحظة دلالة إحصائية لنموذج TUG (p < 0.005). كشف التحليل الإضافي لتعميم النموذج أن النماذج المحددة للمهام أدت بشكل مختلف عند تطبيقها على مهام غير مرئية. برز نموذج TCN كالأكثر فعالية، محققًا أعلى درجة F1 مقارنة بنماذج التعلم الآلي المعتمدة على الميزات، على الرغم من عدم العثور على فرق كبير بين نماذج TCN وLSTM. تؤكد هذه النتائج على أهمية ظروف تدريب النموذج وتبرز أداء TCN المتفوق في الكشف عن FOG، كما هو موضح في إعدادات المعلمات الفائقة المقدمة في الجدول 11.

المناقشة

في هذه الدراسة، تم تجنيد 12 مريضًا مصابًا بمرض باركنسون (PD) للتحقيق في الكشف عن حلقات تجميد المشي (FOG) من خلال تقييمات مختلفة، بما في ذلك اختبار “القيام والذهاب” (TUG) واختبار الدوران بزاوية 360 درجة، وكلاهما تم أداؤه تحت ظروف المهام المزدوجة. تم تطوير نموذج الكشف عن FOG باستخدام بنية مكونة من عنصرين: كتلة توقع أولية تستخدم شبكة تلافيف زمنية (TCN) لتسجيلات FOG من إشارات وحدات القياس الحركية (IMU)، وكتلة تحسين التوقع لتقليل أخطاء التقسيم الزائد. تم تقييم النموذج باستخدام نهج التحقق المتقاطع “ترك موضوع واحد خارجًا”، مما يضمن تقييم أداء النموذج على مواضيع غير مرئية، وهو أمر حاسم للتطبيق السريري.

أشارت النتائج إلى أن النموذج أظهر توافقًا قويًا مع التعليقات من الخبراء من حيث نسبة الوقت المجمد (%TF) وعدد حلقات FOG المكتشفة (#FOG)، مع معاملات ارتباط داخلي (ICCs) قدرها 0.92 و0.95، على التوالي، للنموذج السريري. ومع ذلك، كان النموذج يميل إلى المبالغة في تقدير حدوث FOG، خاصة في التمييز بين FOG واضطرابات المشي ذات الصلة مثل الانزلاق. عند مقارنة أداء النماذج المدربة تحت ظروف مختلفة (مثل مع وبدون فترات توقف)، أظهر النموذج المدرب على التجارب التي تشمل التوقف (نموذج_توقف) توافقًا قويًا مع التعليقات من الخبراء بالنسبة لـ %TF (ICC = 0.95) ولكن توافقًا معتدلًا بالنسبة لـ #FOG (ICC = 0.79). بشكل عام، تؤكد النتائج على إمكانية النموذج في الكشف عن FOG في كل من الإعدادات السريرية والحياتية اليومية، بينما تبرز أيضًا التحديات التي تطرحها تنوع المهام وفترات التوقف في تحديد حلقات FOG بدقة.

Journal: Journal of NeuroEngineering and Rehabilitation, Volume: 21, Issue: 1
DOI: https://doi.org/10.1186/s12984-024-01320-1
PMID: https://pubmed.ncbi.nlm.nih.gov/38350964
Publication Date: 2024-02-13
Author(s): Po-Kai Yang et al.
Primary Topic: Balance, Gait, and Falls Prevention

Overview

The research paper investigates an automated assessment model for freezing of gait (FOG) in Parkinson’s Disease (PD) using deep learning (DL) and inertial measurement units (IMUs). Traditional FOG evaluation methods rely on visual inspection, which is time-consuming and subjective. The study addresses the challenges of gait variability due to medication effects and the need to differentiate FOG from typical movements. A temporal convolutional network (TCN) was employed to analyze data from twelve PD patients during various FOG-provoking tasks under different medication states, with and without volitional stopping.

The results indicate that the TCN model demonstrated strong agreement with expert annotations, achieving an intra-class correlation coefficient (ICC) of 0.92 for the percentage of time frozen (%TF) and 0.95 for the number of freezing episodes (#FOG) in trials without stopping. While models trained on specific FOG-provoking tasks struggled to generalize to new tasks, those trained on medication states showed better adaptability. In trials that included stopping, the model maintained moderate agreement (ICC %TF = 0.95; ICC #FOG = 0.79) only when stopping data was part of the training set. The findings suggest that a single DL model can effectively assess FOG across various tasks and medication states, paving the way for future research on automated FOG detection in real-life scenarios.

Introduction

Parkinson’s disease (PD) significantly impacts gait, with freezing of gait (FOG) affecting approximately 70% of patients. FOG is characterized by a temporary cessation or reduction in movement, manifesting in various forms such as trembling, shuffling, and complete akinesia. This symptom poses risks of falls and negatively affects patients’ quality of life. Current assessment methods, including subjective questionnaires and standardized tasks like the timed-up-and-go (TUG) and 360-degree turning, often fail to provide objective measures of FOG severity due to recall bias and variability in test results. Consequently, there is a growing interest in automated approaches, particularly deep learning (DL) techniques, to enhance FOG detection.

This research addresses several gaps in the existing literature regarding FOG detection using DL models. It introduces a new model for automatic FOG assessment based on inertial measurement units (IMUs) during TUG and 360Turn tasks. The study evaluates whether task-specific or medication-specific training improves detection performance and assesses the model’s generalizability to unseen tasks or medication states. Additionally, it investigates the impact of including voluntary stopping periods in the training data on FOG detection accuracy. The findings indicate that while the model can effectively generalize across tasks and medication states, it struggles with unseen conditions, underscoring the necessity for diverse training data. The study contributes to the field by demonstrating the feasibility of automatic FOG assessment during the 360Turn task and establishing strong agreement with expert evaluations across various conditions.

Methods

In this study, we aimed to enhance the assessment of freezing of gait (FOG) in clinical environments where stopping is not typically incorporated. We first evaluated the overall and relative performance of a generic model that was trained on standardized FOG-provoking tasks, specifically excluding stopping scenarios. Subsequently, we investigated the model’s ability to generalize across various FOG-provoking tasks and different medication states. This involved analyzing the impact of including or excluding training data from specific tasks or medication states on the model’s effectiveness in detecting FOG.

Results

The results section presents a comprehensive analysis of the dataset characteristics and the performance of various models in detecting Freezing of Gait (FOG) under different conditions. Two primary levels of FOG assessment are examined: clinical measurements and real-life scenarios. The findings indicate that the Model_Clinical, trained on specific FOG-provoking tasks and medication states, yielded varying Segment-F1 and Sample-F1 scores across different conditions. For instance, the Model_TUG achieved a Segment-F1 score of 0.70, while the Model_360Turn scored 0.53, with statistical significance noted for the TUG model (p < 0.005). Further analysis of model generalization revealed that task-specific models performed differently when applied to unseen tasks. The TCN model emerged as the most effective, achieving the highest F1 score compared to feature-based machine learning models, although no significant difference was found between the TCN and LSTM models. These results underscore the importance of model training conditions and highlight the TCN's superior performance in FOG detection, as detailed in the hyperparameter settings provided in Table 11.

Discussion

In this study, 12 patients with Parkinson’s Disease (PD) were recruited to investigate the detection of freezing of gait (FOG) episodes through various assessments, including the Timed Up and Go (TUG) test and a 360-degree turn test, both performed under dual-task conditions. The FOG detection model was developed using a two-component architecture: an initial prediction block utilizing a Temporal Convolutional Network (TCN) for FOG annotations from inertial measurement unit (IMU) signals, and a prediction refinement block to reduce over-segmentation errors. The model was evaluated using a leave-one-subject-out cross-validation approach, ensuring that the model’s performance was assessed on unseen subjects, which is crucial for clinical applicability.

The results indicated that the model demonstrated strong agreement with expert annotations in terms of percentage time-frozen (%TF) and the number of detected FOG episodes (#FOG), with intra-class correlation coefficients (ICCs) of 0.92 and 0.95, respectively, for the clinical model. However, the model tended to overestimate FOG occurrences, particularly in distinguishing between FOG and related gait disturbances like shuffling. When comparing the performance of models trained under different conditions (e.g., with and without stopping periods), the model trained on trials including stopping (Model_Stop) showed a strong agreement with expert annotations for %TF (ICC = 0.95) but a moderately strong agreement for #FOG (ICC = 0.79). Overall, the findings underscore the model’s potential for FOG detection in both clinical and daily life settings, while also highlighting the challenges posed by task variability and stopping periods in accurately identifying FOG episodes.