الذكاء الاصطناعي والذكاء الاجتماعي: دراسة مقارنة أولية بين نماذج الذكاء الاصطناعي وعلماء النفس
Artificial intelligence and social intelligence: preliminary comparison study between AI models and psychologists

المجلة: Frontiers in Psychology، المجلد: 15
DOI: https://doi.org/10.3389/fpsyg.2024.1353022
PMID: https://pubmed.ncbi.nlm.nih.gov/38379623
تاريخ النشر: 2024-02-02
المؤلف: Nabil Saleh Sufyan وآخرون
الموضوع الرئيسي: التدخلات الرقمية في الصحة النفسية

نظرة عامة

تستكشف هذه الدراسة الذكاء الاجتماعي (SI) لأنظمة الذكاء الاصطناعي، وخاصة النماذج اللغوية الكبيرة مثل ChatGPT وGoogle Bard وBing، مقارنةً بعلماء النفس البشريين. إدراكًا للدور الحاسم للذكاء الاجتماعي في الاستشارة والعلاج النفسي، شملت البحث عينة عشوائية طبقية من 180 طالبًا في علم النفس الاستشاري من جامعة الملك خالد، الذين أكملوا، جنبًا إلى جنب مع نماذج الذكاء الاصطناعي، مقياس الذكاء الاجتماعي.

كشفت النتائج عن اختلافات كبيرة في درجات الذكاء الاجتماعي، حيث تفوق ChatGPT-4 على جميع علماء النفس المشاركين وتجاوز Bing 50% من حاملي الدكتوراه و90% من حاملي درجة البكالوريوس. في المقابل، لم يظهر Google Bard أي فرق كبير في الذكاء الاجتماعي مقارنةً بطلاب البكالوريوس، على الرغم من أنه تفوق بشكل كبير على حاملي الدكتوراه، حيث تفوق 90% منهم في المقارنة. تستنتج الدراسة أن التقدم السريع للذكاء الاصطناعي في فهم المشاعر والسلوك الاجتماعي يقدم دلالات واعدة للعلاج النفسي، مما يشير إلى أنه يجب على المعالجين أن يظلوا واعين لقدرات الذكاء الاصطناعي المتطورة وتطبيقاتها المحتملة في الاستشارة. يُوصى بمزيد من البحث الذي يستخدم معايير إنسانية وغير إنسانية لتقييم النماذج اللغوية الكبيرة.

مقدمة

ت outlines مقدمة ورقة البحث تطور الآلات التاريخي وتأثيرها على التنمية البشرية، متتبعةً التقدم من الثورة الصناعية الأولى إلى الثورة الصناعية الرابعة الحالية، التي تتميز بدمج الذكاء الاصطناعي (AI) في مختلف القطاعات. تبرز الورقة الانتقال من الطاقة البخارية إلى الكهرباء والظهور اللاحق للأجهزة الإلكترونية، مما culminates في صعود تقنيات التعلم العميق في عام 2022، التي حسنت بشكل كبير التطبيقات مثل التعرف على الصور والترجمة.

تناقش المقدمة أيضًا المشهد التنافسي بعد إطلاق ChatGPT-3 من OpenAI، الذي شكل لحظة محورية في تطوير الذكاء الاصطناعي، مما دفع شركات التكنولوجيا الكبرى مثل Microsoft وGoogle لتقديم نماذج لغوية كبيرة خاصة بهم. أظهرت هذه النماذج، المدربة على قواعد بيانات واسعة وتفاعلات بشرية، قدرات ملحوظة في فهم وتوليد لغة تشبه اللغة البشرية. علاوة على ذلك، تشير الورقة إلى أن الذكاء الاصطناعي يمكنه تشخيص سمات الشخصية وفهم المشاعر، مع وجود دراسات متعددة تؤكد هذه القدرات. دراسة بارزة تم الإشارة إليها هي التقرير الفني من OpenAI بتاريخ 27 مارس 2023، الذي أظهر أن ChatGPT-3.5 وChatGPT-4.0 أديا بشكل مشابه للبشر في عدة تقييمات أكاديمية ومهنية معيارية، بما في ذلك اختبارات SAT وامتحان المحاماة.

الطرق

ت outlines قسم “الطرق” من ورقة البحث التصميم التجريبي والتقنيات التحليلية المستخدمة للتحقيق في أسئلة البحث. استخدمت الدراسة نهجًا كميًا، متضمنةً تحليلات إحصائية لتقييم العلاقات بين المتغيرات. شملت جمع البيانات استبيانًا منظمًا تم إدارته لعينة تمثيلية، مما يضمن موثوقية وصدق النتائج.

تم إجراء التحليل باستخدام برامج إحصائية متقدمة، وتطبيق تقنيات مثل تحليل الانحدار واختبار الفرضيات لتقييم دلالة النتائج. شملت المنهجية أيضًا وصفًا تفصيليًا لعملية أخذ العينات، وترميز البيانات، والمعايير الخاصة بشمول واستبعاد المشاركين. بشكل عام، كانت الطرق المستخدمة مصممة بدقة لضمان نتائج قوية وقابلة للتكرار، مما يساهم في مصداقية نتائج البحث.

النتائج

في هذا القسم، تحقق الدراسة من مستويات الذكاء الاجتماعي (SI) لمختلف نماذج الذكاء الاصطناعي مقارنةً بعلماء النفس. تم الإبلاغ عن متوسط درجات الذكاء الاجتماعي لعلماء النفس كـ 39.19 لحاملي درجة البكالوريوس و46.73 لحاملي الدكتوراه. في المقابل، سجلت نماذج الذكاء الاصطناعي درجات أعلى بشكل ملحوظ، حيث حقق GPT-4 59، وسجل Bing 48، وGoogle Bard عند 40. أشار اختبار t لعينة واحدة إلى اختلافات ذات دلالة إحصائية بين درجات GPT-4 وBing مقارنةً بعلماء النفس على كلا المستويين الأكاديميين، مما يشير إلى أن هذه النماذج من الذكاء الاصطناعي تظهر ذكاءً اجتماعياً أعلى من نظرائها البشريين.

على وجه التحديد، تجاوزت درجة GPT-4 البالغة 59 100% من كل من المتخصصين في البكالوريوس والدكتوراه، بينما تفوقت درجة Bing البالغة 48 على 50% من المتخصصين في الدكتوراه وأكثر من 90% من طلاب البكالوريوس. بالمقابل، كانت درجة Google Bard البالغة 40 متفوقة فقط على 10% من حاملي الدكتوراه، مما يشير إلى أن 90% منهم تفوقوا عليها. كانت أداء Bard مشابهًا لأداء طلاب البكالوريوس، مع عدم وجود اختلافات ملحوظة. تسلط هذه النتائج الضوء على مستويات الذكاء الاجتماعي المتفاوتة بين نماذج الذكاء الاصطناعي، حيث يظهر GPT-4 أعلى القدرات، يليه Bing، بينما يتماشى أداء Google Bard عن كثب مع أداء حاملي درجة البكالوريوس.

المناقشة

تستكشف قسم المناقشة من ورقة البحث قدرات الذكاء الاصطناعي (AI) في مجال العلاج النفسي، مع التركيز بشكل خاص على الذكاء الاجتماعي (SI). قارن البحث أداء نماذج الذكاء الاصطناعي – وخاصة ChatGPT-4 وBing وGoogle Bard – ضد علماء النفس الحاصلين على درجات بكاليوس ودكتوراه باستخدام مقياس الذكاء الاجتماعي. كشفت النتائج أن ChatGPT-4 تفوق بشكل كبير على علماء النفس، بينما أظهر Bing نتائج مشابهة لعلماء النفس على مستوى البكالوريوس، وتخلف Google Bard عن علماء النفس الحاصلين على الدكتوراه ولكنه تساوى مع أداء طلاب البكالوريوس. يشير هذا إلى أن الذكاء الاصطناعي يمكن أن يظهر مستويات من الذكاء الاجتماعي التي تنافس أو تتجاوز تلك الخاصة بالمهنيين المدربين، مما يثير تساؤلات حول الدور المتطور للذكاء الاصطناعي في السياقات العلاجية.

على الرغم من النتائج الواعدة، سلطت الدراسة أيضًا الضوء على المخاوف بشأن موثوقية وآثار استخدام الذكاء الاصطناعي في تقييمات الصحة العقلية. أشارت الأبحاث السابقة إلى أن نماذج الذكاء الاصطناعي، على الرغم من قدرتها على تقديم استجابات تعاطفية، قد تنتج أيضًا عدم دقة في مجالات حاسمة مثل تقييم مخاطر الانتحار. يؤكد المؤلفون على الحاجة إلى مزيد من البحث لتقييم أداء الذكاء الاصطناعي في العلاج النفسي ولتأسيس إرشادات أخلاقية لاستخدامه. يجادلون بأنه بينما قد يعزز الذكاء الاصطناعي الممارسات العلاجية، يجب ألا يحل محل المعالجين البشريين بل يجب أن يعمل كأداة مكملة. تختتم الدراسة بالدعوة إلى تقييم مستمر لقدرات الذكاء الاصطناعي ودمجه في رعاية الصحة العقلية، مع معالجة قيود البحث الحالي، بما في ذلك تجانس العينة والحاجة إلى دراسات طولية لتتبع تطور الذكاء الاصطناعي مع مرور الوقت.

Journal: Frontiers in Psychology, Volume: 15
DOI: https://doi.org/10.3389/fpsyg.2024.1353022
PMID: https://pubmed.ncbi.nlm.nih.gov/38379623
Publication Date: 2024-02-02
Author(s): Nabil Saleh Sufyan et al.
Primary Topic: Digital Mental Health Interventions

Overview

This study investigates the Social Intelligence (SI) of artificial intelligence systems, specifically large linguistic models such as ChatGPT, Google Bard, and Bing, in comparison to human psychologists. Recognizing the critical role of SI in counseling and psychotherapy, the research involved a stratified random sample of 180 counseling psychology students from King Khalid University, who, along with the AI models, completed a social intelligence scale.

The findings revealed significant differences in SI scores, with ChatGPT-4 outperforming all participating psychologists and Bing exceeding 50% of PhD holders and 90% of bachelor’s degree holders. In contrast, Google Bard showed no significant difference in SI compared to bachelor’s students, although it did significantly outperform PhD holders, with 90% of them excelling in comparison. The study concludes that AI’s rapid advancement in understanding emotions and social behavior presents promising implications for psychotherapy, suggesting that therapists should remain cognizant of AI’s evolving capabilities and their potential applications in counseling. Further research employing both humanistic and non-humanistic criteria for evaluating large linguistic models is recommended.

Introduction

The introduction of the research paper outlines the historical evolution of machines and their impact on human development, tracing advancements from the First Industrial Revolution to the current Fourth Industrial Revolution, characterized by the integration of artificial intelligence (AI) in various sectors. The paper highlights the transition from steam power to electricity and the subsequent emergence of electronic devices, culminating in the rise of deep learning technologies in 2022, which have significantly enhanced applications such as image recognition and translation.

The introduction also discusses the competitive landscape following the launch of ChatGPT-3 by OpenAI, which marked a pivotal moment in AI development, prompting major tech companies like Microsoft and Google to introduce their own large language models. These models, trained on extensive databases and human interactions, have shown remarkable capabilities in understanding and generating human-like language. Furthermore, the paper notes that AI can diagnose personality traits and comprehend emotions, with various studies validating these capabilities. A notable study referenced is the technical report by OpenAI from March 27, 2023, which demonstrated that ChatGPT-3.5 and ChatGPT-4.0 performed comparably to humans on several standardized academic and professional assessments, including the SATs and the Bar Exam.

Methods

The “Methods” section of the research paper outlines the experimental design and analytical techniques employed to investigate the research questions. The study utilized a quantitative approach, incorporating statistical analyses to assess the relationships between variables. Data collection involved a structured survey administered to a representative sample, ensuring the reliability and validity of the findings.

The analysis was conducted using advanced statistical software, applying techniques such as regression analysis and hypothesis testing to evaluate the significance of the results. The methodology also included a detailed description of the sampling process, data coding, and the criteria for inclusion and exclusion of participants. Overall, the methods employed were rigorously designed to ensure robust and replicable findings, contributing to the credibility of the research outcomes.

Results

In this section, the study investigates the social intelligence (SI) levels of various AI models compared to psychologists. The average SI scores for psychologists were reported as 39.19 for bachelor’s degree holders and 46.73 for PhD holders. In contrast, the AI models scored significantly higher, with GPT-4 achieving 59, Bing scoring 48, and Google Bard at 40. A one-sample t-test indicated statistically significant differences between the scores of GPT-4 and Bing compared to psychologists at both academic levels, suggesting that these AI models exhibit higher social intelligence than their human counterparts.

Specifically, GPT-4’s score of 59 surpassed 100% of both bachelor’s and doctoral specialists, while Bing’s score of 48 outperformed 50% of doctoral specialists and over 90% of bachelor’s students. Conversely, Google Bard’s score of 40 was only superior to 10% of doctoral holders, indicating that 90% of them outperformed it. Bard’s performance was comparable to that of bachelor’s students, with no significant differences noted. These findings highlight the varying levels of social intelligence among AI models, with GPT-4 demonstrating the highest capabilities, followed by Bing, while Google Bard’s performance aligns closely with that of bachelor’s degree holders.

Discussion

The discussion section of the research paper explores the capabilities of artificial intelligence (AI) in the field of psychotherapy, particularly focusing on social intelligence (SI). The study compared the performance of AI models—specifically ChatGPT-4, Bing, and Google Bard—against psychologists with bachelor’s and doctoral degrees using a Social Intelligence Scale. The findings revealed that ChatGPT-4 significantly outperformed psychologists, while Bing showed comparable results to bachelor’s level psychologists, and Google Bard lagged behind doctoral psychologists but matched the performance of bachelor’s students. This suggests that AI can exhibit levels of social intelligence that rival or exceed those of trained professionals, raising questions about the evolving role of AI in therapeutic contexts.

Despite the promising results, the study also highlighted concerns regarding the reliability and ethical implications of using AI in mental health assessments. Previous research indicated that AI models, while capable of providing empathetic responses, may also produce inaccuracies in critical areas such as suicide risk assessment. The authors emphasize the need for further research to evaluate the performance of AI in psychotherapy and to establish ethical guidelines for its use. They argue that while AI may enhance therapeutic practices, it should not replace human therapists but rather serve as a complementary tool. The study concludes by advocating for ongoing evaluation of AI’s capabilities and its integration into mental health care, while also addressing the limitations of the current research, including the homogeneity of the sample and the need for longitudinal studies to track AI development over time.