تطوير والتحقق من صحة مقاييس تقييم نماذج اللغة الكبيرة لجلسات العلاج النفسي المنقولة تلقائيًا
Development and validation of large language model rating scales for automatically transcribed psychological therapy sessions

شارك:
المجلة: Scientific Reports، المجلد: 15، العدد: 1
DOI: https://doi.org/10.1038/s41598-025-14923-y
PMID: https://pubmed.ncbi.nlm.nih.gov/40796797
تاريخ النشر: 2025-08-12
المؤلف: Steffen T. Eberhardt وآخرون
الموضوع الرئيسي: الصحة النفسية من خلال الكتابة

نظرة عامة

تقدم البحث منهجية جديدة لتطوير مقاييس التقييم باستخدام نماذج اللغة الكبيرة (LLMs) لتقييم البنى الكامنة في النصوص، مع التركيز بشكل خاص على مشاركة المرضى في جلسات العلاج. استخدمت الدراسة مقاطع فيديو تم نسخها تلقائيًا من 1,131 جلسة علاجية تضم 155 مريضًا، وتم تحليلها من خلال إطار عمل DISCOVER. قامت Llama 3.1 8B LLM بتقييم 120 عنصرًا من عناصر المشاركة، مما أدى إلى الحصول على درجة إجمالية مستمدة من متوسط أعلى ثماني تقييمات. أظهر التقييم النفسي قياس موثوقية قوية (ω = 0.953) ومؤشرات ملاءمة مقبولة (CFI = 0.968، SRMR = 0.022)، على الرغم من أن RMSEA كان مرتفعًا قليلاً عند 0.108. تم دعم الصلاحية من خلال ارتباطات كبيرة مع مختلف محددات المشاركة والعمليات والنتائج، مما يدل على قوة مقياس تقييم LLM عبر إعادة أخذ العينات باستخدام bootstrap والتحقق المتقاطع.

تؤسس الدراسة خط أنابيب شامل للنسخ شبه الآلي، وتوليد العناصر، والاختيار، مما يعزز قابلية التوسع وكفاءة التقييمات النفسية. أنتج مقياس تقييم LLM، الذي تمثل في تطوير مقياس LLEAP لمشاركة المرضى، نتائج نفسية موثوقة مع ارتباطات ذات مغزى مع الاستبيانات المعتمدة. لا تقدم هذه الطريقة أداة قابلة للتوسع وآلية للتقييم النفسي فحسب، بل تضمن أيضًا حماية البيانات السرية، متكيفة مع التقدم في تكنولوجيا LLM لتحسين الفعالية بشكل أكبر. بشكل عام، تسهم هذه الأبحاث بشكل كبير في المجال من خلال تقديم إطار عمل جديد لقياس البنى النفسية ضمن بيانات النصوص.

النتائج

تشير النتائج إلى أن تطبيق نموذج اللغة الكبير (LLM) لتحليل 120 عنصرًا من عناصر المشاركة جنبًا إلى جنب مع 4,574 مقطعًا من 1,131 نصًا تطلب حوالي 60 ساعة من وقت المعالجة. من الجدير بالذكر أن العناصر الثمانية المختارة لمقياس LLEAP تم تحديدها باستمرار كأكثر العناصر التي تم اختيارها بشكل متكرر عبر 1,000 عينة bootstrap و20 تكرارًا من التحقق المتقاطع ثلاثي الأضعاف، مما يدل على قوة وموثوقية عملية الاختيار. تؤكد هذه النتائج فعالية LLM في تحديد العناصر الرئيسية للمشاركة ضمن مجموعة البيانات.

المناقشة

تحدد قسم المناقشة المنهجية الشاملة المستخدمة في تطوير وتقييم مقياس التقييم القائم على LLM لقياس مشاركة المرضى في العلاج النفسي. بدأت العملية بنسخ وتقسيم جلسات العلاج، تلتها إنشاء مجموعة أولية من العناصر تضم 120 عنصرًا مستمدة من الأطر النظرية. تم استخدام خط أنابيب اختيار نفسي لضمان موثوقية وصلاحية العناصر، مع دمج استراتيجيات مثل التحقق المتقاطع ثلاثي الأضعاف المتكرر وتصحيح التفاؤل bootstrap للتخفيف من الإفراط في التكيف وتعزيز القابلية للتعميم. تألفت العينة من 1,131 نصًا من 155 مريضًا، مع مجموعة متنوعة من التشخيصات، واتبعت جميع الإجراءات المعايير الأخلاقية، مما يضمن سرية المرضى طوال الدراسة.

شمل تقييم المقياس القائم على LLM تقييمات نفسية صارمة، بما في ذلك توزيع المقياس، وتحليل الموثوقية، وتحليل الصلاحية. تم تأكيد موثوقية المقياس من خلال تحليل العوامل التأكيدية متعددة المستويات (MCFA)، مما أسفر عن مؤشرات ملاءمة مرضية ومقاييس اتساق داخلي. تم تقييم الصلاحية من خلال تحليلات الارتباط والنمذجة متعددة المستويات، مما يظهر علاقات كبيرة بين درجات LLEAP والمقاييس المعتمدة للتحقق. كما أن استخدام تصحيح التفاؤل bootstrap والتحقق المتقاطع المتكرر k-fold عزز من خصائص المقياس النفسية، مما يضمن نتائج قوية وقابلة للتعميم. بشكل عام، تؤكد دقة منهجية الدراسة والالتزام بالمعايير الأخلاقية على إمكانية مقياس التقييم القائم على LLM في تعزيز فهم مشاركة المرضى في السياقات العلاجية.

Journal: Scientific Reports, Volume: 15, Issue: 1
DOI: https://doi.org/10.1038/s41598-025-14923-y
PMID: https://pubmed.ncbi.nlm.nih.gov/40796797
Publication Date: 2025-08-12
Author(s): Steffen T. Eberhardt et al.
Primary Topic: Mental Health via Writing

Overview

The research presents a novel methodology for developing rating scales using Large Language Models (LLMs) to assess latent constructs in text, specifically focusing on patient engagement in therapy sessions. The study utilized automatically transcribed videos from 1,131 therapy sessions involving 155 patients, analyzed through the DISCOVER framework. The Llama 3.1 8B LLM rated 120 engagement items, culminating in a total score derived from the average of the top eight ratings. Psychometric evaluation revealed strong reliability (ω = 0.953) and acceptable fit indices (CFI = 0.968, SRMR = 0.022), although the RMSEA was slightly elevated at 0.108. Validity was supported by significant correlations with various engagement determinants, processes, and outcomes, demonstrating the robustness of the LLM rating scale across bootstrap resampling and cross-validation.

The study establishes a comprehensive pipeline for semi-automated transcription, item generation, and selection, enhancing the scalability and efficiency of psychological assessments. The LLM rating scale, exemplified by the development of the LLEAP measure for patient engagement, produced psychometrically sound results with meaningful correlations to established questionnaires. This approach not only offers a scalable and automated tool for psychological assessment but also ensures the protection of confidential data, adapting to advancements in LLM technology to further improve efficacy. Overall, this research contributes significantly to the field by providing a new framework for measuring psychological constructs within text data.

Results

The results indicate that the application of the large language model (LLM) to analyze 120 engagement items alongside 4,574 segments from 1,131 transcripts required approximately 60 hours of processing time. Notably, the eight items chosen for the LLEAP were consistently identified as the most frequently selected across 1,000 bootstrap samples and 20 iterations of 3-fold cross-validation, demonstrating the robustness and reliability of the selection process. These findings underscore the effectiveness of the LLM in identifying key engagement items within the dataset.

Discussion

The discussion section outlines the comprehensive methodology employed in developing and evaluating the LLM-based rating scale for measuring patient engagement in psychological therapy. The process began with the transcription and segmentation of therapy sessions, followed by the creation of an initial item pool of 120 items derived from theoretical frameworks. A psychometric selection pipeline was utilized to ensure the reliability and validity of the items, incorporating strategies such as repeated 3-fold cross-validation and bootstrap optimism correction to mitigate overfitting and enhance generalizability. The sample consisted of 1,131 transcripts from 155 patients, with a diverse range of diagnoses, and all procedures adhered to ethical standards, ensuring patient confidentiality throughout the study.

The evaluation of the LLM-based scale involved rigorous psychometric assessments, including scale distribution, reliability, and validity analyses. The scale’s reliability was confirmed through Multilevel Confirmatory Factor Analysis (MCFA), yielding satisfactory fit indices and internal consistency metrics. Validity was assessed through correlation analyses and multilevel modeling, demonstrating significant relationships between the LLEAP scores and established validation scales. The use of bootstrap optimism correction and repeated k-fold cross-validation further validated the scale’s psychometric properties, ensuring robust and generalizable findings. Overall, the study’s methodological rigor and adherence to ethical standards underscore the potential of the LLM-based rating scale in enhancing the understanding of patient engagement in therapeutic contexts.

شارك: