نظام مدمج يعتمد على الانتباه المكون من خطوتين للكشف عن الخرف القائم على الكلام
A Two-Step Attention-Based Feature Combination Cross-Attention System for Speech-Based Dementia Detection

شارك:
المجلة: IEEE Transactions on Audio Speech and Language Processing، المجلد: 33
DOI: https://doi.org/10.1109/taslpro.2025.3533363
تاريخ النشر: 2025-01-01
المؤلف: Yilin Pan وآخرون
الموضوع الرئيسي: التعرف على العواطف والمزاج

نظرة عامة

تبحث الورقة البحثية في استخدام نماذج التعلم الذاتي (SSL)، وتحديداً BERT و wav2vec 2.0 (w2v)، للكشف عن الخرف من خلال تحليل الكلام. تؤكد على أهمية الكشف المبكر وإمكانات اللغة المنطوقة كأداة تشخيصية. يقدم المؤلفون نظام دمج الميزات المعتمد على الانتباه ثنائي الخطوة (TSAC-ATT)، الذي يعزز استخراج ودمج الميزات الصوتية واللغوية. تشير نتائجهم إلى أن الميزات الصوتية فقط المستمدة من نموذج w2v يمكن أن تحقق أداءً عالياً عبر مجموعات بيانات متعددة، مما يتحدى الفكرة السائدة بأن الأنظمة متعددة الوسائط أفضل.

في الخاتمة، يبرز المؤلفون أن أداء أنظمة التعرف التلقائي على الكلام (ASR) لا يتوافق دائماً مع فعالية التصنيف اللاحق. ومن الجدير بالذكر أن نظام w2v، على الرغم من معدل خطأ الكلمات (WER) الأعلى، تفوق على نظام ASR الكلاسيكي في الكشف عن الخرف. وهذا يشير إلى أن WER وحده ليس مقياساً كافياً لتقييم أنظمة ASR في هذا السياق. تدعو الورقة إلى مزيد من الاستكشاف للأسباب وراء نجاح نظام w2v وتقترح تطوير أنظمة SSL أكثر انسيابية لاستخراج المعلومات الصوتية واللغوية بشكل مباشر. ستأخذ الأبحاث المستقبلية أيضًا في الاعتبار التطبيقات عبر مجموعات البيانات وعبر اللغات لتعزيز قوة وملاءمة أنظمة الكشف عن الخرف.

الطرق

في قسم الطرق، يصف المؤلفون الإعداد التجريبي لتقييم نظام TSAC-ATT المقترح من خلال منهجية التحقق المتقاطع (CV). تعتبر هذه الطريقة حاسمة لتقييم أداء النظام وقوته. يوضح القسم أيضًا أنظمة التعرف التلقائي على الكلام (ASR) المستخدمة في نسخ الصوت، والتي تعد مكونًا أساسيًا في عملية التقييم.

بالإضافة إلى ذلك، يحدد المؤلفون المعلمات المستخدمة في أنظمة استخراج الميزات المختلفة، بما في ذلك الميزات اللغوية والصوتية، بالإضافة إلى أنظمة الدمج التي تدمج هذه الميزات. توفر هذه الوصف الشامل للمنهجيات والأنظمة المستخدمة إطارًا واضحًا لفهم التصميم التجريبي والتحليل اللاحق لفعالية نظام TSAC-ATT.

النتائج

في قسم النتائج، يتم تقييم أداء الأنظمة المقترحة من خلال مقارنتين رئيسيتين: النموذج اللغوي فقط مقابل مصادر النسخ المختلفة وكفاءة إطار TSAC ونظام TSAC-ATT. في القسم VI-A، يتم استخدام نموذج BERT الأساسي المعدل لتقييم دقة النسخ الناتجة عن أنظمة ASR الكلاسيكية و w2v مقابل النسخ التي تم إنشاؤها يدويًا. تشير النتائج إلى أنه بينما يظهر نظام ASR w2v معدل خطأ كلمات (WER) أعلى بنسبة 48.43% مقارنةً بـ 33.19% لنظام ASR الكلاسيكي، فإنه يحقق درجة F أعلى بنسبة 80.06% على مجموعة بيانات DB، مما يشير إلى أن نسخ w2v تحتوي على معلومات قيمة لتصنيف AD. ومن الجدير بالذكر أن نسخ w2v تتفوق على النسخ اليدوية في مجموعة بيانات IVA، حيث تحقق درجة F تبلغ 86.01% مقابل 83.65%.

في القسم VI-B، يتم تحليل كفاءة إطار TSAC، مما يظهر أنه يتفوق باستمرار أو يتطابق مع أفضل النتائج من طرق دمج الميزات المختلفة عبر جميع مجموعات البيانات. يعزز نظام TSAC-ATT الأداء من خلال آلية الانتباه المتقاطع، مما يؤدي إلى تحسين درجة F من 79.13% إلى 81.24% على مجموعة بيانات ADReSS. بشكل عام، تشير النتائج إلى أن إطار TSAC هو طريقة قوية لمعالجة الميزات الصوتية، وأن نظام TSAC-ATT يتفوق بشكل كبير على كل من النموذج اللغوي فقط والطرق الصوتية السابقة، محققًا درجة F تبلغ 82.54% على مجموعة بيانات DB، والتي تتجاوز معايير الأبحاث السابقة.

المناقشة

تتناول قسم المناقشة في الورقة البحثية المشهد الحالي لأساليب الكشف عن الخرف المعتمدة على اللغة المنطوقة، مصنفة إياها إلى أساليب قائمة على الصوت، وأخرى قائمة على اللغة، وأخرى مدمجة صوتية-لغوية. تبرز تطور الميزات اللغوية التي تلتقط الفروق الدقيقة في اللغة المتأثرة بالخرف، مثل تصنيف أجزاء الكلام، ونسبة النوع إلى الرمز، وتعقيد التركيب. تشير الورقة إلى التقدم الكبير الذي تم إحرازه من خلال تطبيق نماذج التعلم العميق، وخاصة BERT، الذي أظهر تفوقه على الطرق السابقة في استخراج المعلومات اللغوية لمهام الكشف عن الخرف. يشير المؤلفون إلى أعمالهم السابقة، التي حققت نتائج متقدمة باستخدام نظام يعتمد على الانتباه الهرمي، ويؤكدون على الاستمرار في أهمية BERT في أبحاثهم الحالية.

على الجانب الصوتي، تناقش الورقة كيف يمكن أن تشير الخصائص الصوتية إلى الخرف قبل فترة طويلة من التشخيص الرسمي، مع ميزات مثل معاملات تردد ميل (MFCC) ومعدل الكلام التي يتم استخدامها بشكل شائع. يذكر المؤلفون دمج تقنيات التعلم الذاتي (SSL)، مثل نموذج w2v، الذي أظهر وعدًا في استخراج الميزات الصوتية. يتناول القسم أيضًا الجانب الحاسم لدمج الميزات، حيث يمكن أن يعزز دمج البيانات الصوتية واللغوية دقة الكشف. يقترح المؤلفون آلية انتباه متقاطع لتحسين عملية الدمج، مشيرين إلى أنه بينما تدعو بعض الدراسات إلى النمذجة المشتركة لهذه الوسائط، تشير دراسات أخرى إلى أن التحليلات المنفصلة قد تكون كافية. تهدف الورقة إلى استكشاف ضرورة ومنهجية دمج هذه المجموعات من الميزات في نظام TSAC-ATT المقترح، الذي يجمع بين المتجهات الصوتية من w2v مع الميزات اللغوية من BERT من خلال إطار عمل يعتمد على الانتباه ثنائي الخطوة.

Journal: IEEE Transactions on Audio Speech and Language Processing, Volume: 33
DOI: https://doi.org/10.1109/taslpro.2025.3533363
Publication Date: 2025-01-01
Author(s): Yilin Pan et al.
Primary Topic: Emotion and Mood Recognition

Overview

The research paper investigates the use of self-supervised learning (SSL) models, specifically BERT and wav2vec 2.0 (w2v), for detecting dementia through speech analysis. It emphasizes the importance of early detection and the potential of spoken language as a diagnostic tool. The authors introduce the Two-Step Attention-based Feature Combination Cross-attention system (TSAC-ATT), which enhances the extraction and fusion of acoustic and linguistic features. Their findings indicate that acoustic-only features derived from the w2v model can achieve high performance across multiple datasets, challenging the prevailing notion that multimodal systems are superior.

In the conclusion, the authors highlight that the performance of automatic speech recognition (ASR) systems does not always correlate with downstream classification effectiveness. Notably, the w2v system, despite a higher word error rate (WER), outperformed a classic ASR system in dementia detection. This suggests that WER alone is not a sufficient metric for evaluating ASR systems in this context. The paper calls for further exploration into the reasons behind the w2v system’s success and proposes the development of more streamlined SSL systems for direct acoustic and linguistic information extraction. Future research will also consider cross-corpora and cross-language applications to enhance the robustness and applicability of dementia detection systems.

Methods

In the Methods section, the authors describe the experimental setup for evaluating the proposed TSAC-ATT system through a cross-validation (CV) methodology. This approach is critical for assessing the system’s performance and robustness. The section further details the automatic speech recognition (ASR) systems employed for audio transcription, which serve as a foundational component of the evaluation process.

Additionally, the authors outline the parameters used for various feature extraction systems, including linguistic and acoustic features, as well as the fusion systems that integrate these features. This comprehensive description of the methodologies and systems used provides a clear framework for understanding the experimental design and the subsequent analysis of the TSAC-ATT system’s effectiveness.

Results

In the Results section, the performance of the proposed systems is evaluated through two main comparisons: the linguistic-only model against various transcript sources and the efficiency of the TSAC framework and TSAC-ATT system. In Section VI-A, the fine-tuned BERT base model is employed to assess the accuracy of transcripts generated by classic and w2v ASR systems against manually created transcripts. The findings indicate that while the w2v ASR system exhibits a higher word error rate (WER) of 48.43% compared to the classic ASR’s 33.19%, it achieves a superior F-score of 80.06% on the DB dataset, suggesting that the w2v transcripts contain valuable information for AD classification. Notably, the w2v transcripts outperform manual transcripts on the IVA dataset, achieving an F-score of 86.01% versus 83.65%.

In Section VI-B, the TSAC framework’s efficiency is analyzed, demonstrating that it consistently outperforms or matches the best results from various feature combination methods across all datasets. The TSAC-ATT system further enhances performance through a cross-attention mechanism, leading to an F-score improvement from 79.13% to 81.24% on the ADReSS dataset. Overall, the results indicate that the TSAC framework is a robust acoustic feature processing method, and the TSAC-ATT system significantly outperforms both the linguistic-only model and previous acoustic-only methods, achieving an F-score of 82.54% on the DB dataset, which surpasses prior research benchmarks.

Discussion

The discussion section of the research paper outlines the current landscape of dementia detection methodologies based on spoken language, categorizing them into acoustic-based, linguistic-based, and acoustic-linguistic fused approaches. It highlights the evolution of linguistic features that capture the nuances of language affected by dementia, such as Part-of-Speech tagging, Type-Token Ratio, and syntactic complexity. The paper notes the significant advancements made through the application of deep learning models, particularly BERT, which has been shown to outperform previous methods in extracting linguistic information for dementia detection tasks. The authors reference their own prior work, which achieved state-of-the-art results using a hierarchical attention-based system, and emphasize the continued relevance of BERT in their current research.

On the acoustic side, the paper discusses how vocal characteristics can indicate dementia long before formal diagnosis, with features like Mel Frequency Cepstral Coefficients (MFCC) and speech rate being commonly utilized. The authors mention the integration of self-supervised learning (SSL) techniques, such as the w2v model, which has shown promise in extracting acoustic features. The section also addresses the critical aspect of feature fusion, where the integration of acoustic and linguistic data can enhance detection accuracy. The authors propose a cross-attention mechanism to improve the fusion process, suggesting that while some studies advocate for joint modeling of these modalities, others indicate that separate analyses may suffice. The paper aims to explore the necessity and methodology of integrating these feature sets in their proposed TSAC-ATT system, which combines acoustic vectors from w2v with linguistic features from BERT through a two-step attention-based framework.

شارك: