DOI: https://doi.org/10.1016/j.invent.2025.100841
PMID: https://pubmed.ncbi.nlm.nih.gov/40525210
تاريخ النشر: 2025-06-03
المؤلف: Ludwig Franke Föyen وآخرون
الموضوع الرئيسي: التدخلات الرقمية في الصحة النفسية
نظرة عامة
تشير نتائج البحث إلى أن النصائح النفسية التي تم إنشاؤها بواسطة الذكاء الاصطناعي تظهر جودة علمية وتعاطفًا معرفيًا يمكن مقارنته بتلك التي يقدمها خبراء في النصوص المكتوبة غير المتزامنة، بينما تظهر تعاطفًا عاطفيًا وتحفيزيًا متفوقًا. على الرغم من هذه القوة، يكشف الدراسة عن تحيز كبير في التفضيلات بناءً على التأليف المدرك؛ حيث تم تقييم النصائح المنسوبة إلى خبراء بشريين بشكل أكثر إيجابية باستمرار، بغض النظر عن مصدرها الفعلي. وهذا يشير إلى أن التحيزات المتعلقة بالتأليف المدرك قد تعيق قبول الذكاء الاصطناعي في بيئات الرعاية الصحية النفسية.
يدعو المؤلفون إلى مزيد من التحقيق في هذه التحيزات عبر سياقات مختلفة وتطوير استراتيجيات للتخفيف من تأثيرها. بالإضافة إلى ذلك، يؤكدون على الحاجة إلى أبحاث مستقبلية لاستكشاف كيف يمكن أن تعزز أدوات الذكاء الاصطناعي اتخاذ القرارات السريرية وتحسين عمليات العلاج ضمن مجال الرعاية الصحية النفسية.
مقدمة
تقدم دمج الذكاء الاصطناعي (AI) في تقديم النصائح النفسية من أنظمة مبكرة مثل ELIZA، التي كانت تحاكي التفاعلات العلاجية، إلى نماذج اللغة الكبيرة المعاصرة (LLMs) وعوامل معالجة اللغة الطبيعية (NLP). على الرغم من إمكانية هذه الأدوات الحديثة في معالجة الفجوة في الرعاية التي يواجهها ما يصل إلى 70% من الأفراد الذين يعانون من مرض نفسي، لا تزال الأدلة الأولية بشأن فعاليتها مختلطة. أظهرت تحليل شامل أن الدعم للنصائح المعتمدة على الذكاء الاصطناعي في تخفيف أعراض الاكتئاب والقلق كان ضعيفًا، مما يبرز التحيزات ونقص التجارب عالية الجودة. ومع ذلك، أظهرت تدخلات محددة مثل “تيس”، “ووبوت”، و”ويسا” وعدًا في تقليل الأعراض، خاصة بين المستخدمين المتفاعلين.
بينما يمكن لوكلاء المحادثة جذب المستخدمين بشكل فعال، لا تزال الأسئلة حول ملاءمتها كبدائل للمعالجين البشريين قائمة. تم ملاحظة قيود في دقة وعمق الاستجابات التي تم إنشاؤها بواسطة الذكاء الاصطناعي، على الرغم من أن بعض الدراسات تشير إلى أن المستخدمين قد يرون هذه الاستجابات أكثر تعاطفًا من تلك التي يقدمها الخبراء. قد تعزز القدرات المتطورة للنماذج الأحدث قدرتها على نقل التعاطف والفروق اللغوية. ومع ذلك، تظل المخاوف الأخلاقية بشأن نشر الذكاء الاصطناعي في سياقات الرعاية الصحية الحساسة حرجة، مما يتطلب مزيدًا من البحث لتقييم جودة النصائح التي تم إنشاؤها بواسطة الذكاء الاصطناعي مقارنةً بمساهمات الخبراء، خاصة من حيث الصرامة العلمية والانخراط التعاطفي.
الطرق
يستعرض قسم “الطرق” المواد المستخدمة في الدراسة، موضحًا المكونات والأدوات المحددة المستخدمة لجمع البيانات وتحليلها. يتضمن ذلك وصفًا للإعداد التجريبي، وأنواع العينات أو الموضوعات المعنية، وأي معدات أو برامج ذات صلة تم استخدامها لتسهيل عملية البحث. كما يتم مناقشة معايير اختيار المواد، لضمان توافقها مع أهداف الدراسة وفرضياتها.
بالإضافة إلى ذلك، قد يتوسع القسم في البروتوكولات المتبعة أثناء التجربة، بما في ذلك أي ضوابط تم تنفيذها لضمان موثوقية وصلاحية النتائج. من خلال تقديم نظرة شاملة على المواد والأساليب، يؤسس هذا القسم أساسًا لإمكانية إعادة إنتاج الدراسة وتفسير نتائجها.
النتائج
في هذه الدراسة، قام ثلاثة وأربعون من الأطباء النفسيين المرخصين بتقييم 208 استجابات، تتكون من 104 أزواج من النصائح التي تم إنشاؤها بواسطة الذكاء الاصطناعي وتلك التي كتبها خبراء. قام كل طبيب نفسي بتقييم متوسط 2.4 زوج، مع الإشارة إلى أن النصائح التي تم إنشاؤها بواسطة الذكاء الاصطناعي حصلت على تقييمات متساوية أو أكثر إيجابية في مقاييس مختلفة. بينما لم تكن الفروقات في الجودة العلمية (p = .10) والتعاطف المعرفي (p = .08) ذات دلالة إحصائية، تم تقييم استجابات الذكاء الاصطناعي بشكل أعلى بشكل ملحوظ من حيث التعاطف العاطفي (β = 0.59، p = .02) والتعاطف التحفيزي (β = 0.61، p = .02). أظهر تحليل لاحق أن الدراسة يمكن أن تكشف عن أحجام تأثير معتدلة ولكنها تفتقر إلى الحساسية تجاه التأثيرات الأصغر، مما قد يفسر النتائج غير الدالة.
واجه الأطباء النفسيون صعوبة في التمييز بين استجابات الذكاء الاصطناعي واستجابات الخبراء، حيث حققوا دقة تبلغ 45% فقط (اختبار χ²، p = .27). ومع ذلك، تم تقييم استجابات الخبراء بشكل أكثر إيجابية عبر جميع المقاييس، بما في ذلك الجودة العلمية (β = -1.89، p < .001) وأشكال مختلفة من التعاطف (عاطفي: β = -3.62، p < .001؛ معرفي: β = -3.02، p < .001؛ تحفيزي: β = -2.74، p < .001). ومن الجدير بالذكر أن المشاركين أظهروا تفضيلًا قويًا لاستجابات الذكاء الاصطناعي بشكل عام، متأثرين بالتأليف المدرك، حيث فضل 93.55% النصائح المنسوبة إلى الخبراء بغض النظر عن التأليف الفعلي. كانت التفاعلات بين التأليف المدرك والفعلية ذات دلالة (β = -12.29، p = .001)، مما يبرز تأثير الخبرة المدركة على تفضيلات الأطباء النفسيين.
المناقشة
تستكشف الدراسة جودة وتعاطف النصائح النفسية التي تم إنشاؤها بواسطة نموذج الذكاء الاصطناعي (GPT-4) مقارنة بتلك التي يقدمها خبراء الصحة النفسية المرخصون. تهدف إلى تقييم ما إذا كان بإمكان الذكاء الاصطناعي إنتاج نصائح بجودة علمية وتعاطف مماثل، وقدرة الأطباء النفسيين على تحديد مصدر النصيحة، وكيف يؤثر التأليف المدرك على تقييماتهم. شملت الأبحاث 43 طبيبًا نفسيًا مرخصًا قاموا بتقييم النصائح من كل من مصادر الذكاء الاصطناعي والخبراء، مما يكشف أنه بينما تم تقييم النصائح التي تم إنشاؤها بواسطة الذكاء الاصطناعي بشكل مماثل أو حتى متفوق من حيث التعاطف والجودة العلمية، أظهر الأطباء النفسيون تحيزًا لصالح المحتوى الذي كتبه الخبراء عندما تم إخفاء التأليف.
تشير النتائج إلى أن الذكاء الاصطناعي يمكنه تقليد التواصل التعاطفي بشكل فعال، مما يتحدى الفكرة القائلة بأن التعاطف هو سمة بشرية بحتة. على الرغم من قدرات الذكاء الاصطناعي، كانت تفضيلات الأطباء النفسيين متأثرة بشكل كبير بتصوراتهم حول التأليف، حيث حصلت النصائح من الخبراء على تقييمات أعلى. وهذا يشير إلى أن التحيزات المتعلقة بالخبرة المدركة قد تعيق قبول أدوات الذكاء الاصطناعي في سياقات الصحة النفسية، على الرغم من إمكاناتها المثبتة. تبرز الدراسة الحاجة إلى مزيد من الاستكشاف لدور الذكاء الاصطناعي في الرعاية الصحية النفسية، خاصة كأداة داعمة للمعالجين، وتؤكد على أهمية معالجة التحيزات في تقييم المحتوى الذي تم إنشاؤه بواسطة الذكاء الاصطناعي لتعزيز دمجه في الممارسة السريرية.
القيود
تقدم الدراسة عدة نقاط قوة، بما في ذلك رؤى تجريبية حول الفروق في التعاطف والجودة العلمية والتفضيلات بين النصائح التي تم إنشاؤها بواسطة الذكاء الاصطناعي وتلك التي يقدمها الخبراء، مما يساهم في فهم دقيق لقدرة الذكاء الاصطناعي على محاكاة الذكاء العاطفي للخبراء. لقد ساهم إخفاء هوية المشاركين عن التأليف في تقليل التحيز بشكل فعال، مما سمح بتقييم أوضح لتأثيرات التأليف المدرك مقابل الفعلي. بالإضافة إلى ذلك، فإن مشاركة الأطباء النفسيين كمقيمين تعزز تقييم أداء الذكاء الاصطناعي مقارنةً بالتقييمات المستندة إلى المستهلك. تلتزم الأبحاث بالمعايير الأخلاقية ومبادئ العلوم المفتوحة، مع توفر المواد على OSF.io، مما يعزز الشفافية وإمكانية التكرار.
ومع ذلك، تحتوي الدراسة أيضًا على قيود ملحوظة قد تؤثر على صلاحيتها وقابليتها للتعميم. يثير الاعتماد على نصوص كتبها بشكل أساسي خبير واحد (67%) مخاوف بشأن تمثيل تقييمات النصائح من الخبراء فيما يتعلق بالتعاطف والجودة العلمية، مما يبرز ضرورة تنوع التأليف في التحقيقات المستقبلية. علاوة على ذلك، قد يؤدي الإبلاغ الذاتي عن المؤهلات المهنية للمقيمين، التي لم يتم التحقق منها، إلى إدخال تحيز من خلال السماح للأفراد غير المؤهلين بالتأثير على النتائج. أخيرًا، يحد استخدام عينة ملائمة ونقص المعلومات التفصيلية حول خصائص العينة من قابلية تعميم النتائج، حيث قد لا تعكس هذه الطريقة بدقة السكان المستهدفين الأوسع.
DOI: https://doi.org/10.1016/j.invent.2025.100841
PMID: https://pubmed.ncbi.nlm.nih.gov/40525210
Publication Date: 2025-06-03
Author(s): Ludwig Franke Föyen et al.
Primary Topic: Digital Mental Health Interventions
Overview
The research findings indicate that AI-generated psychological advice exhibits a scientific quality and cognitive empathy comparable to that of expert asynchronous written advice, while demonstrating superior emotional and motivational empathy. Despite these strengths, the study reveals a significant bias in preferences based on perceived authorship; advice attributed to human experts was consistently rated more favorably, regardless of its actual source. This suggests that biases related to perceived authorship could hinder the acceptance of AI in mental healthcare settings.
The authors call for further investigation into these biases across various contexts and the development of strategies to mitigate their impact. Additionally, they emphasize the need for future research to explore how AI tools can enhance clinical decision-making and improve treatment processes within the mental healthcare domain.
Introduction
The integration of artificial intelligence (AI) in providing psychological advice has progressed from early systems like ELIZA, which simulated therapeutic interactions, to contemporary large language models (LLMs) and natural language processing (NLP) agents. Despite the potential of these modern tools to address the care gap faced by up to 70% of individuals with mental illness, preliminary evidence regarding their effectiveness remains mixed. A meta-analysis indicated weak support for AI-based treatments in alleviating symptoms of depression and anxiety, highlighting biases and a lack of high-quality trials. However, specific interventions such as “Tess,” “Woebot,” and “Wysa” have shown promise in reducing symptoms, particularly among engaged users.
While conversational agents can engage users effectively, questions about their suitability as substitutes for human therapists persist. Limitations in the accuracy and depth of AI-generated responses have been noted, although some studies suggest that users may perceive these responses as more empathetic than those from experts. The evolving capabilities of newer models may enhance their ability to convey empathy and linguistic nuance. Nonetheless, ethical concerns regarding the deployment of AI in sensitive healthcare contexts remain critical, necessitating further research to evaluate the quality of AI-generated advice compared to expert input, particularly in terms of scientific rigor and empathetic engagement.
Methods
The “Methods” section outlines the materials utilized in the study, detailing the specific components and tools employed for data collection and analysis. This includes a description of the experimental setup, the types of samples or subjects involved, and any relevant equipment or software used to facilitate the research process. The selection criteria for materials are also discussed, ensuring that they align with the study’s objectives and hypotheses.
Additionally, the section may elaborate on the protocols followed during experimentation, including any controls implemented to ensure the reliability and validity of the results. By providing a comprehensive overview of the materials and methodologies, this section establishes a foundation for the reproducibility of the study and the interpretation of its findings.
Results
In this study, forty-three licensed mental health clinicians evaluated 208 responses, comprising 104 pairs of AI-generated and expert-authored advice. Each clinician rated an average of 2.4 pairs, with findings indicating that AI-generated advice received equal or more favorable ratings in various measures. While differences in scientific quality (p = .10) and cognitive empathy (p = .08) were not statistically significant, AI responses were rated significantly higher for emotional empathy (β = 0.59, p = .02) and motivational empathy (β = 0.61, p = .02). A post-hoc analysis revealed that the study could detect moderate effect sizes but lacked sensitivity to smaller effects, which may account for the non-significant results.
Clinicians struggled to distinguish between AI and expert responses, achieving only 45% accuracy (χ² test, p = .27). However, expert responses were rated significantly more favorably across all measures, including scientific quality (β = -1.89, p < .001) and various forms of empathy (emotional: β = -3.62, p < .001; cognitive: β = -3.02, p < .001; motivational: β = -2.74, p < .001). Notably, participants showed a strong preference for AI-authored responses overall, influenced by perceived authorship, with 93.55% favoring perceived expert advice regardless of actual authorship. The interaction between perceived and actual authorship was significant (β = -12.29, p = .001), highlighting the impact of perceived expertise on clinician preferences.
Discussion
The study investigates the quality and empathy of psychological advice generated by an AI model (GPT-4) compared to that provided by licensed mental health experts. It aims to assess whether AI can produce advice of similar scientific quality and empathy, the ability of clinicians to identify the source of the advice, and how perceived authorship influences their evaluations. The research involved 43 licensed clinicians who rated advice from both AI and expert sources, revealing that while AI-generated advice was rated comparably or even superior in terms of empathy and scientific quality, clinicians exhibited a bias favoring expert-authored content when authorship was concealed.
Findings indicate that AI can effectively mimic empathetic communication, challenging the notion that empathy is solely a human trait. Despite the AI’s capabilities, clinicians’ preferences were significantly influenced by their perceptions of authorship, with expert advice receiving higher ratings. This suggests that biases related to perceived expertise may hinder the acceptance of AI tools in mental health contexts, despite their demonstrated potential. The study highlights the need for further exploration of AI’s role in mental healthcare, particularly as a supportive tool for therapists, and emphasizes the importance of addressing biases in the evaluation of AI-generated content to enhance its integration into clinical practice.
Limitations
The study presents several strengths, including empirical insights into the differences in empathy, scientific quality, and preferences between AI-generated and expert advice, which contribute to a nuanced understanding of AI’s ability to emulate expert emotional intelligence. The blinding of participants to authorship effectively mitigated bias, allowing for a clearer assessment of perceived versus actual authorship effects. Additionally, the involvement of clinicians as raters enhances the evaluation of AI’s performance compared to consumer-based assessments. The research adheres to ethical standards and open science principles, with materials accessible on OSF.io, promoting transparency and replicability.
However, the study also has notable limitations that may impact its validity and generalizability. The reliance on texts authored predominantly by a single expert (67%) raises concerns about the representativeness of expert advice evaluations regarding empathy and scientific quality, highlighting the necessity for diverse authorship in future investigations. Furthermore, the self-reported professional credentials of raters, which were not verified, could introduce bias by allowing unqualified individuals to influence the findings. Lastly, the use of convenience sampling and the lack of detailed information about sample characteristics restrict the generalizability of the results, as this method may not accurately reflect the broader target population.
