DOI: https://doi.org/10.3389/fcomp.2025.1629725
تاريخ النشر: 2025-08-22
المؤلف: Yupei Li وآخرون
الموضوع الرئيسي: الصحة النفسية من خلال الكتابة
نظرة عامة
تتناول هذه البحث قيود طرق الكشف عن الاكتئاب الحالية، وخاصة عدم كفاية الشبكات العصبية العميقة (DNNs) ونماذج اللغة الكبيرة (LLMs) عند الاعتماد فقط على البيانات النصية. مع الاعتراف بأن الاكتئاب غالبًا ما يُعبر عنه من خلال نغمة الصوت والسلوك، يقترح المؤلفون نهجًا متعدد الوسائط باستخدام مجموعة بيانات DAIC-WOZ. يقومون باستخراج ميزات الصوت باستخدام نموذج Wav2Vec المدرب مسبقًا ودمجها مع نماذج اللغة الكبيرة المعتمدة على النص. بالإضافة إلى ذلك، يقدمون استراتيجية جديدة لإدماج الخبرة النفسية في نماذج اللغة الكبيرة من خلال إطار عمل الأسئلة والأجوبة، مما يعزز بشكل كبير دقة التشخيص.
تشير النتائج التجريبية إلى أن هذا النهج المبتكر يؤدي إلى تحسينات كبيرة في مقاييس الأداء، حيث تحقق متوسط الخطأ المطلق (MAE) حوالي 0.35 وجذر متوسط مربع الخطأ (RMSE) حوالي 1.36، متفوقًا على المعايير المعمول بها. تسلط النتائج الضوء على إمكانيات نماذج اللغة الكبيرة كأدوات فعالة لتقييم الصحة النفسية، مما يقترح سبلًا للبحث المستقبلي قد تشمل تدريب نماذج اللغة الكبيرة على نطاق كامل ودمج أوصاف صوتية أغنى لتعزيز قدرات النموذج في تفسير أعراض الاكتئاب.
مقدمة
تسلط مقدمة هذه الورقة البحثية الضوء على الأهمية المتزايدة لاستخدام الذكاء الاصطناعي (AI) في تشخيص الاضطرابات العاطفية، وخاصة الاكتئاب. اعتمدت طرق التشخيص التقليدية بشكل كبير على التقييمات الذاتية للأطباء، مما أدى إلى تباين في التشخيصات. لتعزيز الاتساق والدقة، تدعو الورقة إلى دمج التقييمات الآلية المدعومة بالبيانات الكبيرة وتقنيات الذكاء الاصطناعي. أظهرت الدراسات السابقة أن الجمع بين الذكاء الاصطناعي والخبرة السريرية يمكن أن يحسن من التعرف على الاكتئاب، مما يزيد من ثقة الأطباء ورضا المرضى. ومع ذلك، لا تزال هناك تحديات، بما في ذلك الشكوك من المرضى بشأن موثوقية الذكاء الاصطناعي وقيود الأنظمة الحالية، التي غالبًا ما تفتقر إلى التكامل مع المعرفة النفسية.
تستعرض الورقة التقدمات المهمة في الذكاء الاصطناعي للكشف عن الاكتئاب، خاصة من خلال الشبكات العصبية العميقة (DNNs) والنهج متعدد الوسائط. على الرغم من النتائج الواعدة من نماذج مختلفة، مثل الشبكات الذاكرة الطويلة القصيرة (LSTM) والشبكات ذات الانتباه المزدوج، لا تزال قيود الأداء قائمة بسبب القدرات التعليمية المحدودة لهذه الهياكل. يؤكد المؤلفون على الحاجة إلى دمج متعدد الوسائط لتحقيق فهم شامل لأعراض الاكتئاب، حيث غالبًا ما تفشل النهج الأحادية في التقاط تعقيد التقييمات العاطفية. يقترحون خط أنابيب جديد يستفيد من نماذج اللغة الكبيرة (LLMs) للغة المنطوقة في التعرف على الاكتئاب، مع دمج المعرفة النفسية المهنية لتعزيز الموثوقية والصلاحية السريرية. تشير النتائج الأولية إلى أن هذا النهج يتفوق بشكل كبير على النماذج الأساسية الحالية على مجموعة بيانات DAIC-WOZ، مما يعالج الفجوات الحرجة في مشهد البحث الحالي.
طرق البحث
في السنوات الأخيرة، استخدمت معالجة اللغة الطبيعية (NLP) بشكل متزايد نماذج اللغة العميقة للكشف عن الاكتئاب القائم على النص، متماشية مع الممارسات السريرية التي تستخدم الإشارات اللغوية لتشخيص الاكتئاب. من الجدير بالذكر أن فالستار وآخرين (2013) كانوا روادًا في دمج أنماط النص في الكشف عن الاكتئاب متعدد الوسائط من خلال دمج نصوص التعرف على الكلام التلقائي (ASR) مع بيانات الصوت والفيديو. أظهرت الدراسات اللاحقة، مثل تلك التي أجراها أوجونلي وآخرون (2024)، فعالية النماذج الهجينة، حيث حققت درجات F1 بنسبة 69% و76% على مجموعات بيانات وسائل التواصل الاجتماعي من خلال دمج Sentence-BERT وطرق التجميع جنبًا إلى جنب مع مؤشرات المشاعر المعتمدة على المعجم. بالمثل، استخدم سيفامانيكاندان وآخرون (2022) نماذج التحويل مثل RoBERTa على مجموعة بيانات من مهمة تكنولوجيا اللغة للمساواة والتنوع والشمول (LT-EDI) 2022، محققين درجة F1 قدرها 0.457 في مهمة تصنيف ثلاثية الفئات، مما يبرز فعالية هياكل التحويل في هذا المجال.
تتجه الاتجاهات في الكشف عن الاكتئاب القائم على النص بشكل متزايد نحو تحسين نماذج اللغة المدربة مسبقًا باستخدام بيانات موسومة، كما يتضح من الدراسات التي تجمع بين تقنيات التعلم العميق مع العلامات اللغوية أو المعاجم النفسية اللغوية (مثل، ليو وآخرون، 2023). قدم كاثان وآخرون (2022b) ميزات نصية بديلة مثل مقياس التنشيط السلوكي للاكتئاب – الشكل القصير (BADSSF) ومقياس دراسات الاكتئاب لمركز الأوبئة (CESD). بشكل عام، تشير الأبحاث إلى تحول كبير نحو نماذج قائمة على التحويل تركز على مؤشرات لغوية محددة وأنماط جمل مرتبطة بالحالات العاطفية، مما يعزز أداء أنظمة الكشف عن الاكتئاب.
النتائج
في هذا القسم، يقدم المؤلفون نتائج تجاربهم، والتي تم تلخيصها في الجدول 2. استخدموا متوسط الخطأ المطلق (MAE) وجذر متوسط مربع الخطأ (RMSE) كمقاييس رئيسية لتقييم أداء النموذج. تشير النتائج إلى أن طريقتهم المعتمدة على الصوت تتفوق بشكل كبير على المعايير التي تم وضعها في تحدي AVEC 2016 (فالستار وآخرون، 2016)، مما يوضح فعالية ميزات الصوت في تعزيز الدقة التنبؤية. على وجه التحديد، تحسن MAE للصوت من 5.373 إلى 5.356، بينما انخفض للنص من 6.342 إلى 5.354. كما أظهرت RMSE تحسنًا، حيث انخفض الصوت من 6.733 إلى 6.713 والنص من 8.891 إلى 6.429، مما يشير إلى أن إدخال المعرفة له تأثير أكثر وضوحًا على التنبؤات المعتمدة على النص.
علاوة على ذلك، تشير النتائج إلى أنه بينما يؤدي دمج الصوت والنص مع إدخال المعرفة إلى أداء تنافسي، إلا أنه لا يؤدي دائمًا إلى مكاسب كبيرة. قد يُعزى ذلك إلى التحيزات الكامنة في القدرات التعليمية لنماذج اللغة الكبيرة (LLMs) والقيود التي تفرضها عدد المعلمات القابلة للتدريب في إطار LoRA. استكشف المؤلفون أيضًا فعالية إدخال المعرفة من خلال توليد خمسين سؤالًا تتعلق بتقييم الاكتئاب، والتي تم تقييمها بواسطة نموذجين – أحدهما تم تدريبه مع أزواج الأسئلة والأجوبة والآخر بدون. تم تقييم استجابات PsycholexLLAMA بشكل إضافي لتقييم تأثير إدخال المعرفة على فهم النموذج للكشف عن الاكتئاب، مما يشير إلى إمكانية تحسين الأداء من خلال الدمج الفعال للمعلومات متعددة الوسائط.
المناقشة
تسلط قسم المناقشة في الورقة البحثية الضوء على التقدمات في الكشف عن الاكتئاب من خلال الأساليب المعتمدة على الصوت، ومتعددة الوسائط، ونماذج اللغة الكبيرة (LLM). يعتمد الكشف المعتمد على الصوت على ميزات صوتية مثل تباين النغمة ومعدل الكلام، مع نماذج مثل wav2vec 2.0 التي تحقق دقة عالية (96.5%) في تصنيف الاكتئاب. تظل الطرق التقليدية، مثل مصنفات SVM، فعالة أيضًا، مما يشير إلى نهج مزدوج في هذا المجال. أظهرت النماذج متعددة الوسائط، التي تدمج البيانات النصية والصوتية والمرئية، أداءً محسنًا مقارنة بالطرق الأحادية، كما يتضح من الدراسات التي تستخدم مجموعات بيانات مثل AVEC 2017 ومدونات الفيديو على يوتيوب.
لقد حول ظهور نماذج اللغة الكبيرة (LLMs) التعرف على الاكتئاب، حيث أظهرت نماذج مثل GPT-3.5 وEmoLLMs تحسينات كبيرة في الدقة (حتى 90.2%) في تقييم مخاطر الاكتئاب. تشير الأبحاث إلى أن دمج المعرفة النفسية في نماذج اللغة الكبيرة يعزز أدائها، كما يتضح من نماذج مثل PsychoLexLLaMA وDORIS، التي تدمج الرؤى السريرية لتحسين التقييم العاطفي. تؤكد الورقة على إمكانيات نماذج اللغة الكبيرة عند تحسينها باستخدام المعرفة الخاصة بالمجال، مما يقترح أن تركز الأبحاث المستقبلية على تطوير مجموعات بيانات موثقة سريريًا واستكشاف تدريب نماذج اللغة الكبيرة على نطاق كامل لسد الفجوة بين المعرفة النظرية والتطبيق العملي في تقييمات الصحة النفسية.
DOI: https://doi.org/10.3389/fcomp.2025.1629725
Publication Date: 2025-08-22
Author(s): Yupei Li et al.
Primary Topic: Mental Health via Writing
Overview
The research addresses the limitations of current depression detection methods, particularly the inadequacy of deep neural networks (DNNs) and large language models (LLMs) when relying solely on textual data. Recognizing that depression is often conveyed through vocal tone and behavior, the authors propose a multimodal approach utilizing the DAIC-WOZ dataset. They extract audio features using the pre-trained Wav2Vec model and integrate these with text-based LLMs. Additionally, they introduce a novel strategy to incorporate psychological expertise into LLMs through a question-and-answer framework, significantly enhancing diagnostic accuracy.
Empirical results indicate that this innovative approach leads to substantial improvements in performance metrics, achieving a Mean Absolute Error (MAE) of approximately 0.35 and a Root Mean Square Error (RMSE) of around 1.36, outperforming established baselines. The findings highlight the potential of LLMs as effective tools for mental health assessment, suggesting avenues for future research that may include full-scale LLM training and the integration of richer acoustic descriptions to further enhance the model’s capabilities in interpreting depressive symptoms.
Introduction
The introduction of this research paper highlights the growing importance of utilizing artificial intelligence (AI) in the diagnosis of emotional disorders, particularly depression. Traditional diagnostic methods have relied heavily on the subjective assessments of clinicians, leading to variability in diagnoses. To enhance consistency and accuracy, the paper advocates for the integration of automated assessments powered by large-scale data and AI techniques. Previous studies have shown that combining AI with clinical expertise can improve depression recognition, thereby increasing clinician confidence and patient satisfaction. However, challenges remain, including skepticism from patients regarding AI’s reliability and the limitations of current AI systems, which often lack integration with psychological knowledge.
The paper reviews significant advancements in AI for depression detection, particularly through deep neural networks (DNNs) and multimodal approaches. Despite promising results from various models, such as Long Short-Term Memory (LSTM) networks and Dual Attention Networks, performance limitations persist due to the constrained learning capacities of these architectures. The authors emphasize the need for multimodal fusion to achieve a comprehensive understanding of depressive symptoms, as unimodal approaches often fail to capture the complexity of emotional assessments. They propose a novel pipeline that leverages large language models (LLMs) for spoken language in depression recognition, incorporating professional psychological knowledge to enhance reliability and clinical validity. Preliminary results indicate that this approach significantly outperforms existing baseline models on the DAIC-WOZ dataset, addressing critical gaps in the current research landscape.
Methods
In recent years, Natural Language Processing (NLP) has increasingly leveraged deep language models for text-based depression detection, aligning with clinical practices that utilize linguistic cues to diagnose depression. Notably, Valstar et al. (2013) pioneered the integration of text modalities in multimodal depression detection by combining automatic speech recognition (ASR) transcripts with audio and video data. Subsequent studies, such as those by Ogunleye et al. (2024), demonstrated the efficacy of hybrid models, achieving F1 scores of 69% and 76% on social media datasets by incorporating Sentence-BERT and ensemble methods alongside lexicon-based sentiment indicators. Similarly, Sivamanikandan et al. (2022) utilized transformer models like RoBERTa on a dataset from the Language Technology for Equality, Diversity, and Inclusion (LT-EDI) 2022 task, achieving an F1 score of 0.457 in a three-class classification task, underscoring the effectiveness of transformer architectures in this domain.
The trend in text-based depression detection is increasingly oriented towards fine-tuning pretrained language models with labeled data, as evidenced by studies that combine deep learning techniques with linguistic markers or psycholinguistic lexicons (e.g., Lyu et al., 2023). Kathan et al. (2022b) introduced alternative text-based features such as the Behavioral Activation for Depression Scale-Short Form (BADSSF) and the Center for Epidemiologic Studies Depression Scale (CESD). Overall, the research indicates a significant shift towards transformer-based models that emphasize specific lexical indicators and sentence patterns associated with emotional states, enhancing the performance of depression detection systems.
Results
In this section, the authors present the results of their experiments, which are summarized in Table 2. They utilized Mean Absolute Error (MAE) and Root Mean Square Error (RMSE) as primary metrics to evaluate model performance. The findings indicate that their audio-based method significantly outperforms the baseline established in the AVEC 2016 challenge (Valstar et al., 2016), demonstrating the efficacy of audio features in enhancing predictive accuracy. Specifically, the MAE for audio improved from 5.373 to 5.356, while for text, it decreased from 6.342 to 5.354. The RMSE also showed improvements, with audio dropping from 6.733 to 6.713 and text from 8.891 to 6.429, suggesting that knowledge injection has a more pronounced effect on text-based predictions.
Furthermore, the results imply that while combining audio and text with knowledge injection yields competitive performance, it does not always lead to substantial gains. This may be attributed to the inherent biases in the learning capabilities of large language models (LLMs) and the limitations posed by the number of trainable parameters in the LoRA framework. The authors also explored the effectiveness of knowledge injection by generating fifty questions related to depression assessment, which were evaluated by two models—one trained with Question & Answer pairs and one without. The scoring of responses by PsycholexLLAMA further assessed the impact of knowledge injection on the model’s understanding of depression detection, indicating the potential for improved performance through effective integration of multimodal information.
Discussion
The discussion section of the research paper highlights advancements in depression detection through audio-based, multimodal, and large language model (LLM) approaches. Audio-based detection leverages acoustic features such as pitch variability and speech rate, with models like wav2vec 2.0 achieving high accuracy (96.5%) in classifying depression. Traditional methods, such as SVM classifiers, also remain effective, indicating a dual approach in the field. Multimodal models, which integrate textual, audio, and visual data, have shown improved performance over unimodal methods, as evidenced by studies utilizing datasets like AVEC 2017 and YouTube vlogs.
The emergence of LLMs has further transformed depression recognition, with models like GPT-3.5 and EmoLLMs demonstrating significant accuracy improvements (up to 90.2%) in assessing depression risk. Research indicates that incorporating psychological knowledge into LLMs enhances their performance, as seen with models like PsychoLexLLaMA and DORIS, which integrate clinical insights for better emotional assessment. The paper emphasizes the potential of LLMs when fine-tuned with domain-specific knowledge, suggesting that future research should focus on developing clinically annotated datasets and exploring full-scale LLM training to bridge the gap between theoretical knowledge and practical application in mental health assessments.
