الكشف المبكر عن قضايا الصحة النفسية باستخدام منشورات وسائل التواصل الاجتماعي بناءً على نموذج الانتباه العميق متعدد الأنماط-BiLSTM
Multi-modal deep-attention-BiLSTM based early detection of mental health issues using social media posts

شارك:
المجلة: Scientific Reports، المجلد: 15، العدد: 1
DOI: https://doi.org/10.1038/s41598-025-19141-0
PMID: https://pubmed.ncbi.nlm.nih.gov/41062710
تاريخ النشر: 2025-10-08
المؤلف: Qasim Bin Saeed وآخرون
الموضوع الرئيسي: الصحة النفسية من خلال الكتابة

نظرة عامة

في هذه الورقة، يقدم المؤلفون إطار عمل جديد للتعلم العميق متعدد الأنماط مصمم للكشف المبكر وتصنيف مشاكل الصحة النفسية باستخدام منشورات وسائل التواصل الاجتماعي. يدمج الإطار بيانات نصية من منصات مثل Reddit وTwitter مع معلومات زمنية من خلال شبكة الذاكرة طويلة وقصيرة المدى ثنائية الاتجاه (BiLSTM)، ووحدات الذاكرة طويلة وقصيرة المدى (LSTM)، وآلية انتباه عبر الأنماط. تتناول هذه الطريقة دمج أنماط البيانات المختلفة، والتي لم يتم استكشافها بشكل كافٍ في أبحاث الصحة النفسية، وتظهر أداءً متفوقًا مقارنةً بالنماذج التقليدية التي تعتمد فقط على النصوص.

تظهر تقييمات النموذج المقترح على مجموعة بيانات كبيرة من العالم الحقيقي عدة نتائج رئيسية: (أ) إن تضمين الميزات الزمنية يعزز بشكل كبير من أداء التصنيف، (ب) يتفوق النموذج على الهياكل الأساسية التي تعتمد فقط على البيانات النصية، مثل BiLSTM وLSTM والشبكات العصبية التلافيفية (CNN)، من حيث الدقة والاسترجاع ودرجة F1، (ج) تعمل آلية الانتباه عبر الأنماط على إعطاء الأولوية للميزات ذات الصلة، مما يعزز دقة التنبؤ، و(د) تعتبر المتغيرات الزمنية، بما في ذلك أوقات وأنماط النشر، ضرورية لفهم سلوك المستخدم فيما يتعلق بالصحة النفسية. ومع ذلك، تعترف الدراسة بالقيود مثل عدم توازن الفئات الذي يؤثر على التعميم، والمخاوف الأخلاقية المتعلقة بالخصوصية وشفافية النموذج، والاحتمال المحتمل لاستبعاد السرديات الحيوية من خلال تصفية المحتوى غير المناسب. بشكل عام، تؤكد هذه الأبحاث على وعد الأساليب متعددة الأنماط في تعزيز تصنيف الصحة النفسية وتضع الأساس للتطورات المستقبلية في هذا المجال.

النتائج

تظهر نتائج الدراسة فعالية النموذج المقترح في تصنيف اضطرابات الصحة النفسية باستخدام مدخلات متعددة الأنماط، كما يتضح من مقاييس الأداء المختلفة بما في ذلك خسارة الاختبار، الدقة، الاسترجاع، ودرجة F1. حقق النموذج أدنى خسارة اختبار تبلغ 0.7130 ودقة اختبار تبلغ 75.96%، مما يشير إلى قدرات تعميم قوية. تكشف مقاييس الأداء التفصيلية عن دقة تبلغ 0.7407، واسترجاع يبلغ 0.7455، ودرجة F1 تبلغ 0.7376، مما يعكس أداءً متوازنًا عبر معايير التقييم مع تقليل فعال للنتائج الإيجابية والسلبية الكاذبة.

توضح مصفوفة الالتباس أداء النموذج عبر فئات الصحة النفسية المختلفة، مشيرة إلى أن القلق واضطراب الشخصية الحدية (BPD) أظهرا أعلى دقة تصنيف، بينما أظهرت الفصام والمرض العقلي العام أداءً أضعف بسبب تداخل الميزات والتمثيل المحتمل غير الكافي في مجموعة البيانات. من الجدير بالذكر أن BPD حقق دقة تبلغ 0.8943 ودرجة F1 تبلغ 0.8300، مما يشير إلى تحديد دقيق مع الحد الأدنى من النتائج الإيجابية الكاذبة. في المقابل، سجل الفصام أدنى المقاييس (الدقة 0.2747، الاسترجاع 0.3777، F1 0.3181)، مما يشير إلى تحديات في تمييز هذه الفئة. بشكل عام، يظهر النموذج قوة في مهام التصنيف متعددة الفئات لتنبؤ الصحة النفسية، مستفيدًا بشكل كبير من دمج الميزات النصية والزمنية، بما في ذلك آلية الانتباه التي تعزز دقة التصنيف.

المناقشة

يجمع نموذج التعلم العميق المقترح، DABLNet، بين الميزات النصية والزمنية لتعزيز الكشف المبكر عن اضطرابات الصحة النفسية من خلال تحليل وسائل التواصل الاجتماعي. يتكون الهيكل من أربعة مكونات رئيسية: وحدة التحليل النصي (TAM-I)، وحدة التحليل الزمني (TAM-II)، وحدة الانتباه عبر الأنماط (CAM)، وطبقة التصنيف. تستخدم TAM-I شبكة الذاكرة طويلة وقصيرة المدى ثنائية الاتجاه (BiLSTM) لالتقاط المعاني السياقية العميقة من النصوص التي ينشئها المستخدم، مع معالجة التسلسلات في كلا الاتجاهين الأمامي والخلفي. يتم تلخيص المخرجات باستخدام التجميع الأقصى لإنشاء تمثيل نصي قوي. في هذه الأثناء، تستخدم TAM-II LSTM لتحليل الميزات الزمنية، مثل أوقات النشر وتكراراتها، والتي تعتبر حاسمة لفهم أنماط السلوك المرتبطة بالصحة النفسية.

تدمج آلية الانتباه عبر الأنماط التمثيلات النصية والزمنية، مما يسمح للنموذج بالتركيز ديناميكيًا على الميزات الأكثر معلوماتية من كل نمط. يعد هذا الدمج حاسمًا للتصنيف متعدد الأنماط القوي، حيث يعزز قدرة النموذج على تمييز التفاعلات المعقدة بين المؤشرات اللغوية والسلوكية للصحة النفسية. يتم تحقيق التصنيف النهائي من خلال طبقة كثيفة تقوم بربط التمثيل المدمج بحالات الصحة النفسية المحددة، باستخدام دالة تنشيط softmax لإنتاج احتمالات الفئات. يتم تدريب النموذج على مجموعة بيانات مأخوذة من Reddit، والتي توفر سياقًا غنيًا لتحليل المناقشات المتعلقة بالصحة النفسية، ويستخدم مقاييس التقييم مثل الدقة، الاسترجاع، درجة F1، والدقة لتقييم الأداء بشكل شامل.

Journal: Scientific Reports, Volume: 15, Issue: 1
DOI: https://doi.org/10.1038/s41598-025-19141-0
PMID: https://pubmed.ncbi.nlm.nih.gov/41062710
Publication Date: 2025-10-08
Author(s): Qasim Bin Saeed et al.
Primary Topic: Mental Health via Writing

Overview

In this paper, the authors present a novel multi-modal deep learning framework designed for the early detection and categorization of mental health issues using social media posts. The framework integrates textual data from platforms like Reddit and Twitter with temporal information through a Bidirectional Long Short-Term Memory (BiLSTM) network, Long Short-Term Memory (LSTM) units, and a cross-modal attention mechanism. This approach addresses the integration of different data modalities, which has been underexplored in mental health research, and demonstrates superior performance compared to traditional text-only models.

The evaluation of the proposed model on a large, real-world dataset reveals several key findings: (a) the inclusion of temporal features significantly enhances classification performance, (b) the model surpasses baseline architectures that rely solely on textual data, such as BiLSTM, LSTM, and Convolutional Neural Networks (CNN), in terms of precision, recall, and F1-Score, (c) the cross-modal attention mechanism effectively prioritizes relevant features, thereby boosting predictive accuracy, and (d) temporal variables, including posting times and patterns, are essential for understanding user behavior in relation to mental health. However, the study acknowledges limitations such as class imbalance affecting generalization, ethical concerns regarding privacy and model transparency, and the potential exclusion of critical narratives by filtering out NSFW content. Overall, this research underscores the promise of multi-modal approaches in enhancing mental health categorization and lays the groundwork for future developments in the field.

Results

The results of the study demonstrate the effectiveness of the proposed model in categorizing mental health disorders using multimodal inputs, as evidenced by various performance metrics including test loss, accuracy, precision, recall, and F1-Score. The model achieved a minimum test loss of 0.7130 and a test accuracy of 75.96%, indicating strong generalization capabilities. Detailed performance metrics reveal a precision of 0.7407, recall of 0.7455, and an F1-score of 0.7376, reflecting a balanced performance across evaluation criteria while effectively minimizing false positives and negatives.

The confusion matrix further elucidates the model’s performance across different mental health classes, highlighting that anxiety and borderline personality disorder (BPD) exhibited the highest classification accuracy, while schizophrenia and general mental illness showed weaker performance due to overlapping features and potential underrepresentation in the dataset. Notably, BPD achieved an accuracy of 0.8943 and an F1-score of 0.8300, indicating precise identification with minimal false positives. In contrast, schizophrenia recorded the lowest metrics (precision 0.2747, recall 0.3777, F1 0.3181), suggesting challenges in distinguishing this class. Overall, the model demonstrates robustness in multi-class classification tasks for mental health prediction, significantly benefiting from the integration of textual and temporal features, including an attention mechanism that enhances classification accuracy.

Discussion

The proposed deep learning model, DABLNet, integrates both textual and temporal features to enhance early detection of mental health disorders through social media analysis. The architecture consists of four main components: the Textual Analysis Module (TAM-I), Temporal Analysis Module (TAM-II), Cross-Modal Attention Module (CAM), and a Classification layer. TAM-I employs a Bidirectional Long Short-Term Memory (BiLSTM) network to capture deep contextual meanings from user-generated text, processing sequences in both forward and backward directions. The output is summarized using max-pooling to create a robust textual embedding. Meanwhile, TAM-II utilizes an LSTM to analyze temporal features, such as posting times and frequencies, which are critical for understanding behavioral patterns linked to mental health.

The Cross-Modal Attention mechanism fuses the textual and temporal embeddings, allowing the model to dynamically focus on the most informative features from each modality. This integration is crucial for robust multi-modal classification, as it enhances the model’s ability to discern complex interactions between linguistic and behavioral indicators of mental health. The final classification is achieved through a dense layer that maps the fused representation to specific mental health conditions, utilizing a softmax activation function to produce class probabilities. The model is trained on a dataset sourced from Reddit, which provides a rich context for analyzing mental health discussions, and employs evaluation metrics such as Precision, Recall, F1-Score, and Accuracy to assess performance comprehensively.

شارك: