نمذجة أعراض الاكتئاب من نصوص وسائل التواصل الاجتماعي: نهج تعلم شبه مشرف مدفوع بنموذج اللغة الكبير
Depression symptoms modelling from social media text: an LLM driven semi-supervised learning approach

المجلة: Language Resources and Evaluation، المجلد: 58، العدد: 3
DOI: https://doi.org/10.1007/s10579-024-09720-4
تاريخ النشر: 2024-04-04
المؤلف: Nawshad Farruque وآخرون
الموضوع الرئيسي: الصحة النفسية من خلال الكتابة

نظرة عامة

في هذا القسم، يتناول المؤلفون تحدي اكتشاف أعراض الاكتئاب (DSD) باستخدام لغة وسائل التواصل الاجتماعي التي ينشئها المستخدمون، مشيرين إلى غياب مجموعة بيانات مناسبة تلتقط كل من الرؤى السريرية وتوزيع الأعراض من الأفراد الذين أعلنوا عن اكتئابهم. يقترحون إطار عمل للتعلم شبه المراقب (SSL) يدمج نموذج تعلم مراقب، يستخدم نموذج لغة متقدم تم تدريبه مسبقًا وتم ضبطه على مجموعة بيانات DSD تم وضع علامات عليها من قبل الأطباء، مع نموذج تعلم بدون عينة. تم تصميم هذا الإطار لاستخراج عينات ذات صلة من مستودع كبير من التغريدات الاكتئابية (DTR)، والذي تم إنشاؤه من خطوط زمنية على تويتر لمستخدمين أعلنوا عن اكتئابهم، بما في ذلك بيانات من مجموعة بيانات مرجعية كبيرة لاكتشاف الاكتئاب على مستوى المستخدم.

يؤكد المؤلفون أن مجموعة البيانات التي تم وضع علامات عليها من قبل الأطباء هي الأكبر المتاحة، وأن DTR تحافظ بشكل فعال على توزيع أعراض الاكتئاب الموجودة في التغريدات المعلنة. يوضحون عملية إعادة التدريب التكرارية لنموذج DSD الأولي باستخدام البيانات التي تم جمعها، ويناقشون معايير التوقف والقيود على نهج SSL. تشير النتائج إلى أن مجموعة البيانات النهائية المنتجة هي الأكبر من نوعها، وأن كلا من نماذج DSD واكتشاف منشورات الاكتئاب المدربة على هذه المجموعة تظهر دقة محسنة بشكل كبير مقارنةً بإصداراتها الأولية.

مقدمة

تسلط المقدمة الضوء على انتشار الاكتئاب بشكل كبير بين البالغين الشباب، مشيرة إلى أن حوالي 75% من حالات القبول النفسي في البلدان المتقدمة تشمل هذه الفئة، مع كون الانتحار هو السبب الرابع الرئيسي للوفاة المرتبط بالاكتئاب غير المعالج. غالبًا ما تكون طرق الاستطلاع التقليدية لفحص الاكتئاب غير فعالة بسبب التحيزات المعرفية، مما يبرز الحاجة إلى تدخلات مبتكرة وفي الوقت الحقيقي. إن ارتفاع وسائل التواصل الاجتماعي كمنصة للبالغين الشباب لمشاركة التجارب الشخصية يمثل فرصة للتعرف المبكر على أعراض الاكتئاب من خلال تحليل منشورات المستخدمين.

تشير الأبحاث إلى أن الأفراد الذين يعانون من الاكتئاب، والذين قد يظهرون عادة انسحابًا اجتماعيًا، غالبًا ما يتفاعلون أكثر مع وسائل التواصل الاجتماعي للتواصل مع الآخرين وطلب الدعم. لقد نجحت الدراسات السابقة في تحديد العلامات اللغوية للاكتئاب في لغة وسائل التواصل الاجتماعي، باستخدام ميزات مثل n-grams وتحليل المشاعر. ومع ذلك، ركزت معظم الأبحاث الحالية على مجموعات بيانات عامة من منصات مثل تويتر ورديت، مما يفتقر إلى نهج شامل للنمذجة السريرية. تهدف هذه الدراسة إلى معالجة هذه الفجوة من خلال إنشاء مجموعة بيانات تم وضع علامات عليها من قبل الأطباء من تغريدات المستخدمين الذين أعلنوا عن اكتئابهم، مستفيدة من نماذج اللغة المتقدمة لتعزيز اكتشاف أعراض الاكتئاب.

الطرق

في هذا القسم، يوضح المؤلفون منهجيتهم لنمذجة أعراض الاكتئاب باستخدام نهج من جزئين. الجزء الأول يتضمن تجميع مجموعة بيانات تم وضع علامات عليها من قبل الأطباء مستمدة من تغريدات المستخدمين الذين أعلنوا عن اكتئابهم. تم وضع علامات على هذه المجموعة بمساعدة طبيب لضمان الدقة والملاءمة. الجزء الثاني يستخدم التعلم شبه المراقب (SSL) لتطوير وتحسين نماذج اكتشاف أعراض الاكتئاب (DSD) وتوقع الاكتئاب (DPD) من خلال جمع البيانات بشكل تكراري وإعادة التدريب.

تشمل الإعدادات التجريبية جمع البيانات المستمر وإعادة تدريب كل من نماذج DSD وDPD، مع تقييم مقاييس الأداء بعد كل تكرار. بالنسبة لمهمة DSD، التي تتناول مشكلة متعددة الفئات ومتعددة التسميات، يبلغ المؤلفون عن دقة ماكرو ومتوسط دقة مرجحة، والاسترجاع، ودرجات F1، بينما لمهمة التصنيف الثنائي DPD، يركزون على دقة ماكرو ومتوسط الاسترجاع ودرجات F1. تم تخصيص مجموعة فرعية من مجموعة البيانات التي تم وضع علامات عليها من قبل الأطباء كمجموعة اختبار لمهمة DSD، وتم تخصيص 10% من مجموعة تدريب DPD-Human لاختبار مهمة DPD. يتم توثيق نتائج عملية SSL في الأقسام اللاحقة، مع تسليط الضوء على تحسينات أداء النماذج عبر التكرارات (المشار إليها في الجداول 3 و4).

النتائج

في هذا القسم، يقيم المؤلفون فعالية أطر التعلم الذاتي (SSL) الخاصة بهم عبر ثلاثة أبعاد متميزة. يهدف التحليل إلى تقديم فهم شامل لأداء الأطر، مع تسليط الضوء على المقاييس والنتائج الرئيسية التي تظهر قدراتها.

تشير النتائج إلى أن أطر SSL المقترحة تظهر تحسينات كبيرة في مهام مختلفة مقارنة بالنماذج الأساسية. يتم الإبلاغ عن مقاييس أداء محددة، مثل الدقة ومعدلات التقارب، مما يظهر مزايا نهج SSL في تعزيز قوة النموذج والتعميم. بشكل عام، تؤكد النتائج على إمكانيات هذه الأطر في تقدم مجال التعلم الآلي.

المناقشة

في هذا القسم، يناقش المؤلفون تطوير نموذج قوي لاكتشاف أعراض الاكتئاب من نصوص وسائل التواصل الاجتماعي، تحديدًا من خلال إنشاء مجموعة بيانات تم وضع علامات عليها من قبل الأطباء. على عكس الدراسات السابقة التي اعتمدت على تغريدات تم جمعها بناءً على الكلمات الرئيسية، يتبنى هذا البحث نهج “في الموقع”، مستفيدًا من منشورات المستخدمين الذين يعلنون عن حالة اكتئابهم. تم تجميع مجموعة البيانات، التي تُسمى تغريدات مرشحة للاكتئاب، من مصادر موثوقة وتم تصفيتها باستخدام نموذج اكتشاف منشورات الاكتئاب (DPD)، مما أسفر عن أكبر مجموعة بيانات من نوعها. يؤكد المؤلفون على عملية وضع العلامات الدقيقة، التي تشمل نظام تصويت للأغلبية يدمج تفضيلات الأطباء لتعزيز دقة العلامات.

تشمل مهمة وضع العلامات تصنيف التغريدات إلى عشرة أعراض اكتئاب متميزة، بالإضافة إلى تسميات لـ “لا دليل على الاكتئاب”، و”دليل على الاكتئاب”، و”هراء”. يبلغ المؤلفون عن درجة موثوقية بين المراجعين عالية تبلغ 83%، مما يشير إلى اتساق في عملية وضع العلامات. كما يوضحون إطار التعلم شبه المراقب (SSL) الخاص بهم، الذي يقوم بتحسين النموذج بشكل تكراري من خلال جمع بيانات إضافية وإعادة تدريبه، مما يؤدي في النهاية إلى زيادة كبيرة في حجم مجموعة البيانات وتحسينات في دقة النموذج. يحقق النموذج النهائي درجة ماكرو-F1 تبلغ 45%، مما يظهر فعالية نهجهم في نمذجة أعراض الاكتئاب من نصوص وسائل التواصل الاجتماعي.

القيود

تسلط قيود الدراسة الضوء على قلقين رئيسيين بشأن مجموعة البيانات وعملية وضع العلامات. أولاً، مجموعة البيانات الإجمالية صغيرة نسبيًا، مما يؤدي إلى عدم كفاية البيانات لتدريب واختبار بعض التسميات. قد تؤثر هذه القيود على قوة وعمومية نتائج النموذج. ثانيًا، تفتقر عملية الجمع التكراري إلى وضع علامات مستمر من قبل البشر أو استراتيجية “إنسان في الحلقة”، وذلك أساسًا بسبب التكاليف العالية المرتبطة بإشراك الخبراء في دورات متعددة من وضع العلامات.

تكشف مقاييس الأداء المقدمة في الدراسة عن تباين كبير عبر تسميات أعراض الاكتئاب المختلفة، كما هو موضح في الجدول 20. على سبيل المثال، حقق النموذج دقة قدرها 0.00 للتسميات “فقدان المتعة” و”عدم القدرة على اتخاذ القرار”، مما يشير إلى فشل كامل في تحديد هذه الأعراض بدقة. بالمقابل، أظهرت التسمية “التعب” دقة مثالية قدرها 1.00 ولكن استرجاعًا منخفضًا قدره 0.17، مما يشير إلى أنه بينما يمكن للنموذج تحديد بعض الحالات بشكل صحيح، فإنه يفوت عددًا كبيرًا من الحالات الحقيقية. تشير درجة F1 العامة المتوسطة البالغة 0.45 والمتوسط المرجح البالغ 0.56 إلى أنه على الرغم من تحقيق بعض التحسينات، لا يزال أداء النموذج غير متسق عبر أعراض مختلفة.

Journal: Language Resources and Evaluation, Volume: 58, Issue: 3
DOI: https://doi.org/10.1007/s10579-024-09720-4
Publication Date: 2024-04-04
Author(s): Nawshad Farruque et al.
Primary Topic: Mental Health via Writing

Overview

In this section, the authors address the challenge of detecting depression symptoms (DSD) using user-generated social media language, highlighting the absence of a suitable dataset that captures both clinical insights and the symptom distribution from self-disclosed depressed individuals. They propose a semi-supervised learning (SSL) framework that integrates a supervised learning model, which utilizes a state-of-the-art pre-trained language model fine-tuned on a clinician-annotated DSD dataset, with a Zero-Shot learning model. This framework is designed to extract relevant samples from a large repository of depressive tweets (DTR), which is constructed from Twitter timelines of self-disclosed depressed users, including data from a significant benchmark dataset for user-level depression detection.

The authors emphasize that their clinician-annotated dataset is the largest available, and the DTR effectively maintains the distribution of depression symptoms found in self-disclosed tweets. They detail the iterative retraining process of the initial DSD model with the harvested data, discussing the stopping criteria and limitations of the SSL approach. The results indicate that the final dataset produced is the largest of its kind, and both the DSD and Depression Post Detection models trained on this dataset demonstrate significantly improved accuracy compared to their initial versions.

Introduction

The introduction highlights the significant prevalence of depression among young adults, noting that approximately 75% of psychiatric admissions in developed countries involve this demographic, with suicide being the fourth leading cause of death linked to untreated depression. Traditional survey methods for depression screening are often ineffective due to cognitive biases, underscoring the need for innovative, real-time interventions. The rise of social media as a platform for young adults to share personal experiences presents an opportunity for early identification of depression symptoms through analysis of users’ posts.

Research indicates that individuals with depression, who may typically exhibit social withdrawal, often engage more with social media to connect with others and seek support. Previous studies have successfully identified linguistic markers of depression in social media language, utilizing features such as n-grams and sentiment analysis. However, most existing research has focused on public datasets from platforms like Twitter and Reddit, lacking a comprehensive clinical modeling approach. This study aims to address this gap by creating a clinician-annotated dataset from tweets of self-disclosed depressed users, leveraging advanced language models to enhance the detection of depression symptoms.

Methods

In this section, the authors outline their methodology for modeling depression symptoms using a two-part approach. The first part involves the curation of a clinician-annotated dataset derived from tweets by self-disclosed depressed users. This dataset is annotated with the assistance of a clinician to ensure accuracy and relevance. The second part employs Semi-supervised Learning (SSL) to develop and refine Depression Symptom Detection (DSD) and Depression Prediction (DPD) models through iterative data harvesting and retraining.

The experimental setup includes continuous data harvesting and retraining of both DSD and DPD models, with performance metrics evaluated after each iteration. For the DSD task, which addresses a multi-class multi-label problem, the authors report Macro and Weighted-Averaged Precision, Recall, and F1 scores, while for the binary classification DPD task, they focus on Macro-Averaged Precision, Recall, and F1 scores. A subset of the clinician-annotated dataset is designated as a test set for the DSD task, and 10% of the DPD-Human training set is reserved for testing the DPD task. The results of the SSL process are documented in subsequent sections, highlighting the models’ performance improvements across iterations (referenced in Tables 3 and 4).

Results

In this section, the authors evaluate the effectiveness of their self-supervised learning (SSL) frameworks across three distinct dimensions. The analysis aims to provide a comprehensive understanding of the frameworks’ performance, highlighting key metrics and outcomes that demonstrate their capabilities.

The results indicate that the proposed SSL frameworks exhibit significant improvements in various tasks compared to baseline models. Specific performance metrics, such as accuracy and convergence rates, are reported, showcasing the advantages of the SSL approaches in enhancing model robustness and generalization. Overall, the findings underscore the potential of these frameworks in advancing the field of machine learning.

Discussion

In this section, the authors discuss the development of a robust model for detecting depression symptoms from social media text, specifically through the creation of a clinician-annotated dataset. Unlike previous studies that relied on keyword-based crawled tweets, this research adopts an “in-situ” approach, leveraging posts from users who self-disclose their depression status. The dataset, termed the Depression-Candidate-Tweets, is curated from verified sources and filtered using a Depression Post Detection (DPD) model, resulting in the largest dataset of its kind. The authors emphasize the rigorous annotation process, which includes a majority voting scheme that incorporates clinician preferences to enhance label accuracy.

The annotation task involves categorizing tweets into ten distinct depression symptoms, along with labels for “No Evidence of Depression,” “Evidence of Depression,” and “Gibberish.” The authors report a high inter-annotator reliability score of 83%, indicating consistency in the annotation process. They further detail their semi-supervised learning (SSL) framework, which iteratively refines the model by harvesting additional data and retraining it, ultimately leading to a significant increase in dataset size and improvements in model accuracy. The final model achieves a Macro-F1 score of 45%, demonstrating the effectiveness of their approach in modeling depression symptoms from social media text.

Limitations

The limitations of the study highlight two primary concerns regarding the dataset and the annotation process. Firstly, the overall dataset is relatively small, which results in insufficient data for training and testing certain labels. This limitation may affect the robustness and generalizability of the model’s findings. Secondly, the iterative harvesting process lacks continuous human annotation or a human-in-the-loop strategy, primarily due to the high costs associated with involving experts in multiple cycles of annotation.

The performance metrics presented in the study reveal significant variability across different depression symptom labels, as shown in Table 20. For instance, the model achieved a precision of 0.00 for the labels “Anhedonia” and “Indecisiveness,” indicating a complete failure to identify these symptoms accurately. In contrast, the label “Fatigue” demonstrated a perfect precision of 1.00 but a low recall of 0.17, suggesting that while the model can identify some instances correctly, it misses a substantial number of true cases. The overall macro-average F1-score of 0.45 and weighted average of 0.56 indicate that while some improvements were made, the model’s performance remains inconsistent across different symptoms.