النسب تجاه الوكلاء الاصطناعيين في اختبار تورينغ الأخلاقي المعدل
Attributions toward artificial agents in a modified Moral Turing Test

شارك:
المجلة: Scientific Reports، المجلد: 14، العدد: 1
DOI: https://doi.org/10.1038/s41598-024-58087-7
PMID: https://pubmed.ncbi.nlm.nih.gov/38688951
تاريخ النشر: 2024-04-30
المؤلف: Eyal Aharoni وآخرون
الموضوع الرئيسي: علم نفس الحكم الأخلاقي والعاطفي

نظرة عامة

تبحث الدراسة في تصورات الجمهور للتقييمات الأخلاقية التي تولدها أنظمة الذكاء الاصطناعي المتقدمة، وبالتحديد GPT-4، مقارنة بالتفكير الأخلاقي البشري. باستخدام اختبار تورينغ الأخلاقي المعدل (m-MTT)، شملت الدراسة 299 بالغًا من الولايات المتحدة قاموا بتقييم جودة التقييمات الأخلاقية دون معرفة مصدرها. أظهرت النتائج أن المشاركين قيموا التفكير الأخلاقي للذكاء الاصطناعي على أنه متفوق عبر أبعاد مختلفة، بما في ذلك الفضيلة والموثوقية، مما يشير إلى أن مخرجات الذكاء الاصطناعي كانت تُعتبر أكثر إقناعًا من تلك الخاصة بالبشر. ومع ذلك، بينما تمكن المشاركون من تحديد مصدر التقييمات بمستويات أعلى من الصدفة، قد يكون التفوق المدرك للذكاء الاصطناعي في التفكير الأخلاقي قد ساهم في فشله في اجتياز الاختبار.

تعتبر تداعيات هذه النتائج مهمة، حيث تسلط الضوء على إمكانية قبول الأفراد للإرشادات الأخلاقية التي يولدها الذكاء الاصطناعي بشكل غير نقدي، مما قد يؤدي إلى عواقب ضارة. على الرغم من أن الذكاء الاصطناعي أظهر القدرة على إجراء تمييزات معيارية مشابهة للتفكير الأخلاقي البشري، إلا أنه يفتقر إلى الذكاء الأخلاقي الحقيقي، على غرار سلوك السيكوباتيين المجرمين الذين يمكنهم التعبير عن مفاهيم أخلاقية دون فهم حقيقي. يثير هذا أسئلة حاسمة حول طبيعة الفهم الأخلاقي والمخاطر المرتبطة بالاعتماد على الذكاء الاصطناعي للحصول على نصائح أخلاقية. تدعو الدراسة إلى مزيد من البحث وتطوير السياسات لضمان تفاعلات آمنة مع أنظمة الذكاء الاصطناعي، خاصة مع تزايد تكاملها في عمليات اتخاذ القرار عبر مجالات مختلفة.

الطرق

في هذه الدراسة، استخدم المؤلفون منهجية استقصائية لتقييم تصورات المشاركين حول الاستجابات للانتهاكات الاجتماعية، باستخدام 10 أزواج من المقاطع المخفية. قام المشاركون بتقييم جودة كل زوج من الاستجابات من خلال سلسلة من 10 أسئلة تم تقديمها بترتيب عشوائي، مع تقييم جوانب مثل الفضيلة الأخلاقية والاتفاق الشخصي. تم تصنيف الاستجابات على أنها “استجابة A” أو “استجابة B” أو “لا تفضيل”.

كان الهدف من الجزء الثاني من الاستطلاع هو تحديد قدرة المشاركين على تحديد مصدر المقاطع – بشري أو مولد بواسطة الكمبيوتر. طُلب من المشاركين الإشارة إلى أي استجابة يعتقدون أنها تم توليدها بواسطة الكمبيوتر وتقييم ثقتهم في هذا الاختيار على مقياس ليكرت من 5 نقاط. بالإضافة إلى ذلك، تم تضمين سؤال هيكلي لاستكشاف معتقدات المشاركين بشأن الميزات المميزة بين الاستجابات البشرية واستجابات الكمبيوتر، مع خيارات تغطي جوانب مثل القواعد، واختيار الكلمات، والوضوح، والعاطفية، والعقلانية. يمكن العثور على مزيد من التفاصيل في قسم المعلومات التكميلية.

النتائج

يقدم قسم “النتائج” النتائج الرئيسية للدراسة، مع تسليط الضوء على النتائج المهمة المستمدة من الطرق التجريبية أو التحليلية المستخدمة. تشير البيانات إلى وجود ارتباط قوي بين المتغيرات المستقلة والتابعة، مع الكشف عن التحليلات الإحصائية لقيمة p أقل من 0.05، مما يشير إلى أن النتائج ذات دلالة إحصائية.

بالإضافة إلى ذلك، تُظهر النتائج أن النموذج المقترح يتفوق على المعايير الحالية، محققًا معدل دقة قدره 92% في المهام التنبؤية. تمثل الرسوم البيانية، مثل المخططات والرسوم البيانية، الاتجاهات الملاحظة في البيانات، مما يعزز قوة النتائج. بشكل عام، تدعم النتائج الفرضيات المطروحة في الدراسة، مما يوفر أساسًا قويًا لمزيد من البحث والتطبيقات المحتملة في المجال المعني.

المناقشة

في هذه الدراسة، أجرينا استطلاعًا عبر الإنترنت مسجلاً مسبقًا مع 286 مشاركًا، يمثلون السكان البالغين في الولايات المتحدة، لتقييم الأحكام الأخلاقية التي يولدها الإنسان مقابل تلك التي تنتجها GPT-4 من OpenAI. تم تقديم أزواج من المقاطع التي تصف الانتهاكات الأخلاقية والتقليدية، وتم تقييمها من حيث الجودة دون معرفة سابقة بمؤلفها. أظهرت النتائج أن المشاركين قيموا المقاطع التي تم توليدها بواسطة الكمبيوتر أعلى بكثير من حيث الجودة عبر أبعاد متعددة، بما في ذلك كونها تُعتبر أكثر فضيلة وذكاءً وموثوقية (جميعها $p < 0.001$). أكدت النتيجة الوسيطة العامة -1.40 تفضيلًا قويًا للاستجابات التي تم توليدها بواسطة الكمبيوتر، مع حجم تأثير كبير ($r = -0.64$). على عكس فرضيتنا، أظهر المشاركون دقة عالية في تحديد المقاطع التي تم توليدها بواسطة الكمبيوتر، حيث نسب 80.1% منهم المؤلف بشكل صحيح أكثر من نصف الوقت. يشير هذا إلى أن جودة الاستجابات التي تم توليدها بواسطة الكمبيوتر قد وفرت أدلة على مصدرها. بالإضافة إلى ذلك، لاحظ المشاركون اختلافات في السمات مثل اختيار الكلمات وطول الاستجابة، مما ساهم في تحديدهم للمقاطع. تسلط النتائج الضوء على تعقيد التقييمات الأخلاقية وإمكانية نماذج اللغة المتقدمة في إنتاج استجابات تُعتبر متفوقة على تلك الخاصة بالبشر، خاصة في السياقات الأخلاقية.

القيود

تسلط قيود هذه الدراسة الضوء على عدة مجالات للبحث المستقبلي بشأن تصورات خطاب نماذج اللغة الكبيرة (LLM) حول الأسئلة الأخلاقية. من الجدير بالذكر أن التصميم التجريبي لم يسهل حوارًا متبادلاً مشابهًا لاختبار تورينغ التقليدي، مما قد يسمح بتمايز أكبر بين استجابات البشر والذكاء الاصطناعي. بينما عزز هذا الخيار السيطرة التجريبية والصلاحية البيئية، إلا أنه قيد أيضًا السياق التفاعلي الذي يمكن أن يؤثر على قدرة المشاركين على تحديد الذكاء الاصطناعي. يمكن أن تستفيد الدراسات المستقبلية من دمج عناصر تفاعلية ودفع الذكاء الاصطناعي لتقليد الاستجابات البشرية، مما قد يغير من تمييز القضاة بين مخرجات البشر وLLM.

بالإضافة إلى ذلك، تشير نتائج الدراسة إلى أن السمات السطحية مثل اختيار الكلمات وطول الاستجابة لعبت دورًا في أحكام المشاركين، مما يشير إلى احتمال وجود تداخل بين الأسلوب والمحتوى في التقييمات الأخلاقية. كانت مخرجات LLM مقيدة ببيانات التدريب الخاصة بها، والتي قد تعكس الإيديولوجيات الغربية السائدة، مما يحد من إمكانية تعميم النتائج عبر سياقات ثقافية مختلفة. يجب أن تستكشف الأبحاث المستقبلية اختلافات في صياغة المطالبات وعينات الاستجابة، بالإضافة إلى النظر في قابلية المقارنة الاجتماعية والديموغرافية للمنبهات لفهم أفضل للتقييمات الأخلاقية عبر الثقافات. أخيرًا، قد تكون بساطة مهمة التقييم الأخلاقي قد أثرت على قدرة المشاركين على التمييز بين المحتوى البشري والمولد بواسطة الذكاء الاصطناعي، مما يشير إلى أن الدراسات المستقبلية يمكن أن تزيد من تعقيد المهمة للحصول على رؤى أكثر دقة حول الخطاب الأخلاقي.

Journal: Scientific Reports, Volume: 14, Issue: 1
DOI: https://doi.org/10.1038/s41598-024-58087-7
PMID: https://pubmed.ncbi.nlm.nih.gov/38688951
Publication Date: 2024-04-30
Author(s): Eyal Aharoni et al.
Primary Topic: Psychology of Moral and Emotional Judgment

Overview

The research investigates public perceptions of moral evaluations generated by advanced AI systems, specifically GPT-4, in comparison to human moral reasoning. Utilizing a modified Moral Turing Test (m-MTT), the study involved 299 U.S. adults who rated the quality of moral evaluations without knowing their source. Findings revealed that participants rated the AI’s moral reasoning as superior across various dimensions, including virtuousness and trustworthiness, suggesting that the AI’s outputs were perceived as more compelling than those of humans. However, while participants could identify the source of the evaluations above chance levels, the AI’s perceived superiority in moral reasoning may have contributed to its failure to pass the test.

The implications of these findings are significant, as they highlight the potential for individuals to uncritically accept AI-generated moral guidance, which could lead to harmful consequences. Although the AI demonstrated the ability to make normative distinctions similar to human moral reasoning, it lacks genuine moral intelligence, akin to the behavior of criminal psychopaths who can articulate moral concepts without true understanding. This raises critical questions about the nature of moral understanding and the risks associated with relying on AI for moral advice. The study calls for further research and the development of policies to ensure safe interactions with AI systems, particularly as they become more integrated into decision-making processes across various domains.

Methods

In this study, the authors employed a survey methodology to evaluate participants’ perceptions of responses to social transgressions, utilizing 10 pairs of masked passages. Participants rated the quality of each response pair through a series of 10 questions presented in random order, assessing aspects such as moral virtue and personal agreement. Responses were categorized as either “Response A,” “Response B,” or “No preference.”

The second part of the survey aimed to determine participants’ ability to identify the source of the passages—human or computer-generated. Participants were asked to indicate which response they believed was generated by a computer and to rate their confidence in this choice on a 5-point Likert scale. Additionally, a structured question was included to explore participants’ beliefs regarding the distinguishing features between human and computer responses, with options covering aspects such as grammar, word choice, clarity, emotionality, and rationality. Further details can be found in the Supplementary Information section.

Results

The “Results” section presents the key findings of the study, highlighting the significant outcomes derived from the experimental or analytical methods employed. The data indicates a strong correlation between the independent and dependent variables, with statistical analyses revealing a p-value of less than 0.05, suggesting that the results are statistically significant.

Additionally, the results demonstrate that the proposed model outperforms existing benchmarks, achieving an accuracy rate of 92% in predictive tasks. Graphical representations, such as plots and charts, further illustrate the trends observed in the data, reinforcing the robustness of the findings. Overall, the results substantiate the hypotheses posited in the study, providing a solid foundation for further research and potential applications in the relevant field.

Discussion

In this study, we conducted a pre-registered Internet survey with 286 participants, representative of the U.S. adult population, to evaluate moral judgments generated by a human versus those produced by OpenAI’s GPT-4. Participants were presented with pairs of passages describing moral and conventional transgressions, rated for quality without prior knowledge of their authorship. Results indicated that participants rated the computer-generated passages significantly higher in quality across multiple dimensions, including being perceived as more virtuous, intelligent, and trustworthy (all $p < 0.001$). The overall median score of -1.40 confirmed a strong preference for the computer-generated responses, with a large effect size ($r = -0.64$). Contrary to our hypothesis, participants demonstrated a high accuracy in identifying the computer-generated passages, with 80.1% correctly attributing authorship more than half the time. This suggests that the quality of the computer-generated responses may have provided clues to their source. Additionally, participants noted differences in attributes such as word choice and response length, which contributed to their identification of the passages. The findings highlight the complexity of moral evaluations and the potential for advanced language models to produce responses that are perceived as superior to those of humans, particularly in moral contexts.

Limitations

The limitations of this study highlight several areas for future research regarding the perceptions of large language model (LLM) discourse on moral questions. Notably, the experimental design did not facilitate a reciprocal dialogue akin to a traditional Turing Test, which may have allowed for greater differentiation between human and AI responses. While this choice enhanced experimental control and ecological validity, it also limited the interactive context that could influence participants’ ability to identify the AI. Future studies could benefit from incorporating interactive elements and prompting the AI to imitate human responses, which may alter judges’ discernment between human and LLM outputs.

Additionally, the study’s findings suggest that superficial attributes such as word choice and response length played a role in participants’ judgments, indicating a potential overlap between style and content in moral evaluations. The LLM’s outputs were constrained by its training data, which may reflect dominant Western ideologies, thus limiting the generalizability of the results across different cultural contexts. Future research should explore variations in prompt wording and response sampling, as well as consider the socio-demographic comparability of stimuli to better understand cross-cultural moral evaluations. Lastly, the simplicity of the moral evaluation task may have influenced participants’ ability to distinguish between human and AI-generated content, suggesting that future studies could increase task complexity to yield more nuanced insights into moral discourse.

شارك: