DOI: https://doi.org/10.2196/58129
PMID: https://pubmed.ncbi.nlm.nih.gov/38876484
تاريخ النشر: 2024-06-14
المؤلف: Christine Lee وآخرون
الموضوع الرئيسي: الصحة النفسية من خلال الكتابة
نظرة عامة
تستكشف هذه الدراسة فعالية GPT-4 من OpenAI في التنبؤ بأزمات الصحة النفسية، وبشكل خاص الأفكار الانتحارية مع خطة، باستخدام بيانات من منصة الصحة النفسية عن بُعد الوطنية. شملت الأبحاث تحليل بيانات الاستقبال غير المعروفة الهوية من 260 مريضًا أبلغوا عن أفكار انتحارية ومجموعة تحكم من 200 مريض بدون مثل هذه الأفكار. تم تكليف الأطباء (ثلاثة علماء نفس وثلاثة أطباء نفسيين) وGPT-4 بالتنبؤ بالأفكار الانتحارية بناءً على الشكاوى الرئيسية المبلغ عنها ذاتيًا وتاريخ محاولات الانتحار للمرضى. أشارت النتائج إلى أنه بينما حقق GPT-4 دقة بنسبة 61.5%، إلا أنه كان أقل في الخصوصية مقارنة بالأطباء، مع متوسطات بلغت 70.5% و50.1% للنموذج، على التوالي.
تشير النتائج إلى أن أداء GPT-4 قابل للمقارنة مع أداء الأطباء المدربين في بعض المقاييس، خاصة عندما يتم توفير سياق إضافي حول محاولات الانتحار السابقة. ومع ذلك، يبرز انخفاض خصوصية النموذج الحاجة إلى الحذر قبل تنفيذ مثل هذه الأدوات الذكية في البيئات السريرية. تؤكد الدراسة على إمكانية نماذج اللغة الكبيرة في تعزيز تحديد المرضى ذوي المخاطر العالية وتحسين الرعاية في الوقت المناسب، بينما تدعو أيضًا إلى مزيد من البحث لمعالجة السلامة والتحيز والاعتبارات الأخلاقية في تطبيق الذكاء الاصطناعي في رعاية الصحة النفسية. تعتبر هذه الأبحاث خطوة أساسية نحو وضع معايير للاستخدام المسؤول لتقنيات الذكاء الاصطناعي في الكشف عن الحالات النفسية والتنبؤ بها.
مقدمة
تسلط المقدمة الضوء على أزمة الصحة العامة المتزايدة المتعلقة بالانتحار، خاصة في الولايات المتحدة، حيث أصبحت السبب الثاني الرئيسي للوفاة بين البالغين الذين تتراوح أعمارهم بين 18-45 عامًا. في عام 2021، توفي ما يقرب من 50,000 فرد بسبب الانتحار، مما يمثل أعلى معدل في عقود. يتزامن هذا الارتفاع في معدلات الانتحار مع توسع غير كافٍ في قوة العمل في الرعاية الصحية السلوكية، مما يعيق الوصول الفوري إلى الرعاية الأساسية لاكتشاف مخاطر الانتحار والوقاية منها. تعقد عوامل الخطر المختلفة، بما في ذلك محاولات الانتحار السابقة، والتشخيصات النفسية، والعزلة الاجتماعية، من التنبؤ بمخاطر الانتحار. تؤكد الأبحاث الحديثة على مسارات الأفكار إلى الأفعال، مشيرة إلى أن الأفراد الذين لديهم أفكار انتحارية (SI) – التي تُعرف بأنها أفكار الانتحار – هم أكثر عرضة بشكل كبير لمحاولة أو إكمال الانتحار.
تعتبر انتشار أفكار الانتحار مقلقة، حيث أفاد حوالي 12.3 مليون بالغ أمريكي عن مثل هذه الأفكار في عام 2021. تشير الدراسات إلى أن أفكار الانتحار هي مؤشر أكثر حساسية لمخاطر الانتحار مدى الحياة مقارنة بالمخاطر الوشيكة، مع احتمال 29% لمحاولات الانتحار بين أولئك الذين يعانون من أفكار الانتحار. علاوة على ذلك، فإن الأفراد الذين لديهم أفكار انتحارية متكررة هم أكثر عرضة بشكل كبير لمحاولة أو الموت بسبب الانتحار في فترة زمنية قصيرة. تناقش المقدمة أيضًا تطبيق الذكاء الاصطناعي (AI) في تقييم الصحة النفسية، خاصة من خلال تقنيات معالجة اللغة الطبيعية (NLP) وتعلم الآلة (ML)، التي كانت فعالة في تحديد السلوك الانتحاري من السجلات الطبية الإلكترونية وتواصل المرضى. أظهرت التطورات الأخيرة في الذكاء الاصطناعي، بما في ذلك نماذج اللغة الكبيرة (LLMs) مثل GPT-4 وMed-PaLM 2، كفاءة طبية وفعالية في تقييم الوظائف النفسية من النصوص التي ينتجها المرضى، مما يشير إلى آفاق واعدة لتعزيز تقييم مخاطر الانتحار واستراتيجيات التدخل.
الطرق
اتبعت الدراسة المعايير الأخلاقية الموضحة في إعلان هلسنكي وحصلت على موافقة من مجلس المراجعة المؤسسية لـ WCG بموجب البروتوكول رقم #20240207. يضمن ذلك أن البحث تم إجراؤه مع الاعتبارات الأخلاقية اللازمة والإشراف، بما يتماشى مع الإرشادات المعمول بها لأبحاث البشر. تشير الموافقة إلى الالتزام بالحفاظ على سلامة المشاركين ونزاهتهم طوال فترة الدراسة.
النتائج
في هذه الدراسة، تم تحديد 260 مريضًا كمن يستوفون معايير الإدراج للأفكار الانتحارية (SI) مع خطة، حيث كان 140 إيجابيًا عند الاستقبال و120 بعد الاستقبال. تم أيضًا تحليل مجموعة عشوائية من 200 مريض سلبية بالنسبة للأفكار الانتحارية. تشير النتائج إلى أن GPT-4 يمكنه التنبؤ بفعالية بمخاطر الأفكار الانتحارية مع خطة باستخدام نصوص حرة تم إنشاؤها بواسطة المرضى، محققًا مستويات أداء قابلة للمقارنة مع الأطباء على الرغم من عدم تدريبه بشكل صريح لهذه المهمة. بينما أظهر GPT-4 حساسية عالية، كانت خصوصيته ودقته أقل من تلك الخاصة بالخبراء السريريين، مما يعكس التحديات الكامنة في تحديد الأفكار الانتحارية من النصوص فقط.
تسلط الدراسة الضوء على قدرة GPT-4 على استخراج مؤشرات سريرية ذات مغزى، مثل اليأس، مما قد يعزز تصنيف الأزمات. ومن الجدير بالذكر أن GPT-4 أكمل تقييمات لـ 460 عينة نصية في أقل من 10 دقائق، أسرع بكثير من تقييم متوسط الأطباء الذي يستغرق 3 ساعات. هذه الكفاءة ذات صلة خاصة بالنظر إلى نقص قوة العمل في الرعاية الصحية السلوكية الحالية وارتفاع معدلات الانتحار، مما يشير إلى أن طرق الفحص المدعومة بالتكنولوجيا مثل GPT-4 يمكن أن تسهل التدخلات الأسرع. إن الدمج المسؤول لمثل هذه النماذج في الممارسة السريرية أمر ضروري، مما يضمن أن التنبؤات تتبعها مراجعات من الأطباء لتوفير السياق والبيانات السريرية الإضافية. بشكل عام، يمثل GPT-4 أداة واعدة لتحسين توقيت الرعاية في اكتشاف مخاطر الانتحار.
المناقشة
في هذه الدراسة، كان هدف المؤلفين هو تقييم فعالية نموذج اللغة الكبير (LLM)، وبشكل خاص GPT-4 من OpenAI، في التنبؤ بالأفكار الانتحارية مع خطة بين المرضى باستخدام منصة الصحة النفسية عن بُعد الوطنية. تم مقارنة تنبؤات النموذج بتلك التي قدمها أطباء الصحة النفسية ذوو الخبرة بناءً على النصوص التي أنشأها المرضى من استمارات الاستقبال. شملت التحليل مجموعة من المرضى الذين أيدوا الأفكار الانتحارية عند الاستقبال أو أثناء العلاج، بالإضافة إلى مجموعة تحكم لم تؤيد مثل هذه الأفكار. أشارت النتائج الرئيسية إلى أنه بينما أظهر GPT-4 دقة قابلة للمقارنة مع الأطباء في التنبؤ بالأفكار الانتحارية عند الاستقبال (دقة 0.671 مقابل 0.702 للأطباء)، إلا أنه أظهر خصوصية ودقة أقل. ومن الجدير بالذكر أن أداء النموذج تراجع عند التنبؤ بالأفكار الانتحارية المستقبلية بعد الاستقبال، مما يبرز التحديات في التنبؤ بمثل هذه النتائج.
كما قامت الدراسة بتقييم تأثير تاريخ المرضى الإضافي، مثل محاولات الانتحار السابقة، على أداء الأطباء والنموذج. أدى دمج هذه المعلومات إلى تحسين الدقة التنبؤية العامة للأطباء ولكن كان له تأثيرات مختلطة على أداء GPT-4. كانت درجة الاتفاق بين GPT-4 والأطباء متوسطة، حيث تشير قيم كابا لكوهين إلى مستويات متفاوتة من الاتفاق عبر سيناريوهات مختلفة. علاوة على ذلك، كشفت تحليل مؤشرات المخاطر التي حددها GPT-4 أن بعض العوامل، مثل الشعور باليأس، كانت أكثر انتشارًا في حالات الأفكار الانتحارية. أظهر الأطباء حساسية أقل في الحالات التي لم يتم فيها تحديد مؤشرات خطر، مما يبرز أهمية هذه المؤشرات في التقييمات السريرية. بشكل عام، تشير النتائج إلى أنه بينما يمكن أن تساعد نماذج اللغة الكبيرة مثل GPT-4 في تحديد المرضى المعرضين للخطر، ينبغي أن تكمل هذه النماذج، بدلاً من استبدال، الحكم السريري.
القيود
تعترف الدراسة بعدة قيود يجب معالجتها قبل أن يمكن تعميم النتائج لتنبؤ وتحديد مخاطر الانتحار. بينما تُظهر الأبحاث كيف يمكن تخصيص نماذج اللغة الكبيرة (LLMs) لتقييمات نفسية محددة، فإنها تؤكد على أنه لا ينبغي اعتبار هذه النماذج حلولًا نهائية. تشمل القيود الرئيسية الحاجة إلى بيانات من مجموعات سكانية أكبر وأكثر تنوعًا، واستكشاف نماذج LLM بديلة، خاصة تلك المصممة للتطبيقات الطبية، والتحقيق بشكل أعمق في تأثير تصميم المحفزات على الأداء.
تسلط الدراسة الضوء على التحديات الكامنة في التنبؤ بالانتحار، مشيرة إلى أن عدم اليقين في الدقة غالبًا ما يؤدي إلى معدل مرتفع من الإيجابيات الكاذبة، خاصة مع نموذج GPT-4، الذي أظهر معدلًا أعلى من الإيجابيات الكاذبة مقارنة بالأطباء عند استخدام نص الشكوى الرئيسية فقط. ومع ذلك، أظهر GPT-4 أيضًا معدلًا أقل من السلبيات الكاذبة، مما يشير إلى إمكانيته في تعزيز الوعي بالمخاطر بين الأطباء ذوي الخبرة الأقل. علاوة على ذلك، تؤكد الأبحاث على أهمية معالجة التحيزات الموجودة في بيانات تدريب نماذج LLM، حيث قد تؤثر هذه التحيزات على جودة التنبؤ عبر مجموعات فرعية مختلفة من المرضى. يُشجع على إجراء أبحاث مستقبلية لتحسين هذه النماذج وزيادة توافقها مع التقييمات السريرية الخبيرة.
DOI: https://doi.org/10.2196/58129
PMID: https://pubmed.ncbi.nlm.nih.gov/38876484
Publication Date: 2024-06-14
Author(s): Christine Lee et al.
Primary Topic: Mental Health via Writing
Overview
This study investigates the efficacy of OpenAI’s GPT-4 in predicting mental health crises, specifically suicidal ideation with a plan, using data from a national telemental health platform. The research involved analyzing de-identified intake data from 260 patients who reported suicidal ideation and a control group of 200 patients without such ideation. Clinicians (three psychologists and three psychiatrists) and GPT-4 were tasked with predicting suicidal ideation based on patients’ self-reported chief complaints and suicide attempt histories. The results indicated that while GPT-4 achieved an accuracy of 61.5%, it fell short in specificity compared to clinicians, with averages of 70.5% and 50.1% for the model, respectively.
The findings suggest that GPT-4’s performance is comparable to that of trained clinicians in some metrics, particularly when additional context about previous suicide attempts is provided. However, the model’s lower specificity highlights the need for caution before implementing such AI tools in clinical settings. The study emphasizes the potential of large language models to enhance the identification of high-risk patients and improve timely care, while also calling for further research to address safety, bias, and ethical considerations in the application of AI in mental health care. This research serves as a foundational step toward establishing standards for the responsible use of AI technologies in detecting and predicting mental health conditions.
Introduction
The introduction highlights the escalating public health crisis of suicide, particularly in the United States, where it has become the second leading cause of death among adults aged 18-45. In 2021, nearly 50,000 individuals died by suicide, marking the highest rate in decades. This rise in suicide rates coincides with an insufficient expansion of the behavioral healthcare workforce, which hampers timely access to essential care for suicide risk detection and prevention. Various risk factors, including previous suicide attempts, psychiatric diagnoses, and social isolation, complicate the prediction of suicide risk. Recent research emphasizes the ideation-to-action pathways, indicating that individuals with suicidal ideation (SI)—defined as thoughts of suicide—are significantly more likely to attempt or complete suicide.
The prevalence of SI is concerning, with approximately 12.3 million American adults reporting such thoughts in 2021. Studies suggest that SI is a more sensitive predictor of lifetime suicide risk compared to imminent risk, with a 29% probability of suicide attempts among those experiencing SI. Furthermore, individuals with frequent SI are substantially more likely to attempt or die by suicide within a short timeframe. The introduction also discusses the application of Artificial Intelligence (AI) in mental health assessment, particularly through natural language processing (NLP) and machine learning (ML) techniques, which have been effective in identifying suicidal behavior from electronic medical records and patient communications. Recent advancements in AI, including large language models (LLMs) like GPT-4 and Med-PaLM 2, have demonstrated medical competency and effectiveness in evaluating psychiatric functioning from patient-generated text, suggesting promising avenues for enhancing suicide risk assessment and intervention strategies.
Methods
The study adhered to the ethical standards outlined in the Declaration of Helsinki and received approval from the Institutional Review Board of WCG under Protocol #20240207. This ensures that the research was conducted with the necessary ethical considerations and oversight, aligning with established guidelines for human subjects research. The approval indicates a commitment to maintaining participant safety and integrity throughout the study.
Results
In this study, 260 patients were identified as meeting the inclusion criteria for suicidal ideation (SI) with a plan, with 140 positive at intake and 120 post-intake. A random subset of 200 patients negative for SI was also analyzed. The findings indicate that GPT-4 can effectively predict the risk of SI with a plan using patient-generated free text, achieving performance levels comparable to clinicians despite not being explicitly trained for this task. While GPT-4 demonstrated high sensitivity, its specificity and precision were lower than those of clinical experts, reflecting the inherent challenges in identifying SI from text alone.
The study highlights GPT-4’s capability to extract meaningful clinical indicators, such as hopelessness, which could enhance crisis triaging. Notably, GPT-4 completed evaluations of 460 text samples in under 10 minutes, significantly faster than the average clinician’s 3-hour assessment. This efficiency is particularly relevant given the current behavioral healthcare workforce shortage and rising suicide rates, suggesting that technology-enabled screening methods like GPT-4 could facilitate quicker interventions. The responsible integration of such models into clinical practice is essential, ensuring that predictions are followed by clinician reviews to provide context and additional clinical data. Overall, GPT-4 presents a promising tool for improving the timeliness of care in suicide risk detection.
Discussion
In this study, the authors aimed to evaluate the efficacy of a Large Language Model (LLM), specifically OpenAI’s GPT-4, in predicting suicidal ideation with a plan among patients using a national telemental health platform. The model’s predictions were compared against those made by experienced mental health clinicians based on patient-generated text from intake forms. The analysis included a cohort of patients who endorsed suicidal ideation at intake or during treatment, as well as a control group that did not endorse such ideation. Key findings indicated that while GPT-4 demonstrated comparable accuracy to clinicians in predicting suicidal ideation at intake (accuracy of 0.671 vs. 0.702 for clinicians), it exhibited lower specificity and precision. Notably, the model’s performance declined when predicting future suicidal ideation post-intake, highlighting the challenges in forecasting such outcomes.
The study also assessed the impact of additional patient history, such as prior suicide attempts, on both clinician and model performance. Incorporating this information improved overall predictive accuracy for clinicians but had mixed effects on GPT-4’s performance. The inter-rater agreement between GPT-4 and clinicians was moderate, with Cohen’s kappa values indicating varying levels of agreement across different scenarios. Furthermore, the analysis of risk indicators identified by GPT-4 revealed that certain factors, such as a sense of hopelessness, were more prevalent in cases of suicidal ideation. Clinicians showed lower sensitivity in cases where no risk indicators were identified, underscoring the importance of such indicators in clinical assessments. Overall, the findings suggest that while LLMs like GPT-4 can assist in identifying at-risk patients, they should complement, rather than replace, clinical judgment.
Limitations
The study acknowledges several limitations that must be addressed before the findings can be generalized for predicting and identifying suicidal risk. While the research demonstrates how large language models (LLMs) can be customized for specific psychiatric assessments, it emphasizes that these models should not be viewed as definitive solutions. Key limitations include the need for data from larger and more diverse populations, the exploration of alternative LLMs, particularly those designed for medical applications, and a deeper investigation into the impact of prompt design on performance.
The study highlights the inherent challenges in predicting suicidality, noting that precision uncertainties often lead to a high rate of false positives, particularly with the GPT-4 model, which exhibited a higher false positive rate compared to clinicians when only chief complaint text was used. However, GPT-4 also demonstrated a lower rate of false negatives, suggesting its potential utility in enhancing awareness of risk among clinicians with less expertise. Furthermore, the research underscores the importance of addressing biases present in the training data of LLMs, as these biases may affect prediction quality across different patient subpopulations. Future research is encouraged to refine these models and improve their alignment with expert clinical assessments.
