DOI: https://doi.org/10.1177/17456916231214460
PMID: https://pubmed.ncbi.nlm.nih.gov/38165766
تاريخ النشر: 2024-01-02
المؤلف: Max Pellert وآخرون
الموضوع الرئيسي: نمذجة الموضوعات
نظرة عامة
يتناول هذا القسم الدور التحويلي لنماذج اللغة الكبيرة (LLMs) في معالجة البيانات التي ينتجها البشر وتطبيقاتها عبر مجالات متنوعة. تُعرف هذه النماذج باسم “نماذج الأساس”، وتعمل كأساس عام للأنظمة الحاسوبية المستقبلية التي يمكن ضبطها لمهام محددة. يبرز المؤلفون إمكانية استخدام الجرد النفسي، الذي تم تصميمه تقليديًا لتقييم الصفات غير المعرفية لدى البشر، كأدوات تشخيصية لتقييم خصائص مشابهة في LLMs. من خلال استخدام التصنيف بدون أمثلة، يمكن لـ LLMs الاستجابة لعناصر الاستبيان، مما يسمح للباحثين بتجميع الاستجابات في درجات مقياس تعكس صفات مثل القابلية للتوافق.
يؤكد المؤلفون على تعقيد وغموض كيفية اكتساب LLMs للصفات من مجموعات النصوص التي تم تدريبها عليها، والتي تتضمن معتقدات وتحاملات مؤلفين بشريين متنوعين. بينما يرسمون أوجه التشابه بين اكتساب الصفات لدى البشر وLLMs، يحذرون من إضفاء الطابع الإنساني على نماذج الذكاء الاصطناعي، مشيرين إلى أن LLMs تعمل كآلات توقع مع نطاق أضيق من الصفات مقارنةً بالمناظر النفسية الغنية للبشر. يختتم القسم بالتأكيد على الحاجة إلى مزيد من البحث في “علم النفس الاصطناعي”، وهو مجال متعدد التخصصات يدمج علم النفس واللغويات وعلوم الحاسوب، مع تسليط الضوء على الآثار الاجتماعية والاقتصادية الكبيرة للذكاء الاصطناعي مع استمرار تطوره.
طرق
في هذا القسم، يحدد المؤلفون طرقًا متنوعة لتقييم نماذج اللغة الكبيرة (LLMs) من خلال الجرد النفسي. النهج الرئيسي الذي تم مناقشته هو نمذجة اللغة المقنعة (MLM)، حيث يتم تدريب النماذج مسبقًا على مجموعات نصية واسعة من خلال توقع الرموز المقنعة داخل التسلسلات. يسمح هذا التدريب غير المراقب بتقييم النماذج باستخدام تسلسلات مثل “أنا [MASK]” لاشتقاق احتمالات لصفات مختلفة مرتبطة بالسمات النفسية. طريقة أخرى تم تسليط الضوء عليها هي توقع الكلمة التالية، المستخدمة عادة في النماذج التوليدية مثل GPT، والتي تتوقع الكلمات التالية بناءً على المحفزات المعطاة. هذه الطريقة، على الرغم من مرونتها، تقدم تحديات مثل الحاجة إلى هندسة محفزات دقيقة والطبيعة العشوائية لمخرجات النموذج.
بالإضافة إلى ذلك، يصف المؤلفون التصنيف بدون أمثلة، الذي يستفيد من النماذج التي تم ضبطها على مهام الاستدلال اللغوي (NLI) لتصنيف الاستجابات بدون أمثلة سابقة. هذه الطريقة مفيدة بشكل خاص للتقييمات النفسية، حيث تبسط التقييم وتتجنب المشكلات المتعلقة بالنماذج التوليدية. يوضح المؤلفون منهجيتهم باستخدام جرد الخمسة الكبار (BFI) وأدوات نفسية أخرى لاشتقاق ملفات نفسية لـ LLMs، مما يمكّن من المقارنات مع المعايير البشرية وبين نماذج مختلفة. يؤكدون أنه بينما لا تهدف هذه التقييمات إلى قياس أداء النموذج، فإنها تقدم رؤى قيمة حول الصفات النفسية التي تظهرها LLMs، مما قد يكشف عن تحيزات أو خصائص مثيرة للجدل متأصلة في هذه النماذج.
نقاش
يتتبع قسم النقاش في الورقة التقاطع التاريخي بين علم النفس الاصطناعي والذكاء الاصطناعي (AI)، مسلطًا الضوء على المحاولات المبكرة لتقييم قدرات الذكاء الاصطناعي من خلال اختبارات معرفية، مثل مقياس ويشسلر لذكاء البالغين واختبار ستانفورد-بينه. ومن الجدير بالذكر أن شخصيات أساسية مثل توماس جي. إيفانز وألين نيويل أكدت على أهمية دمج مبادئ علم النفس في تطوير الذكاء الاصطناعي. ومع ذلك، تعرضت نماذج الذكاء الاصطناعي المبكرة للانتقاد بسبب تركيزها على “الإدراك البارد”، حيث كانت تفتقر إلى التعقيد العاطفي والدافع. في المقابل، أظهرت التطورات الأخيرة في نماذج اللغة الكبيرة (LLMs) تحسينات كبيرة في معالجة اللغة الطبيعية، مما يمكّن هذه النماذج من إظهار خصائص مرتبطة بـ “الإدراك الحار”، مثل الاستجابات العاطفية والسلوكيات الشبيهة بالبشر.
تدعو الورقة إلى نهج شامل لعلم النفس الاصطناعي يشمل كل من الصفات المعرفية وغير المعرفية، بما في ذلك الشخصية والقيم والمعتقدات الأخلاقية. من خلال الاستفادة من الجرد النفسي المعتمد، يقترح المؤلفون طرقًا لتقييم الملفات النفسية لـ LLMs، باستخدام تقنيات مثل التصنيف بدون أمثلة. تشير النتائج الأولية إلى أن LLMs تظهر عمومًا صفات شخصية متوازنة، حيث تسجل درجات منخفضة في الخصائص غير المرغوب فيها المرتبطة بالتتاراد المظلم، وتقترح تحيزات محتملة في المعايير الأخلاقية ومعتقدات الجنس المشفرة في هذه النماذج. يدعو المؤلفون إلى مزيد من الدراسات التجريبية لاستكشاف آثار هذه الصفات على سلوك LLMs واتخاذ القرارات، مؤكدين على الحاجة إلى أنظمة ذكاء اصطناعي مسؤولة تأخذ في الاعتبار التعقيدات النفسية لهذه النماذج.
DOI: https://doi.org/10.1177/17456916231214460
PMID: https://pubmed.ncbi.nlm.nih.gov/38165766
Publication Date: 2024-01-02
Author(s): Max Pellert et al.
Primary Topic: Topic Modeling
Overview
The section discusses the transformative role of large language models (LLMs) in processing human-generated data and their applications across various domains. These models, referred to as “foundation models,” serve as a general-purpose basis for future computational systems that can be fine-tuned for specific tasks. The authors highlight the potential of using psychometric inventories, traditionally designed to assess noncognitive traits in humans, as diagnostic tools for evaluating similar characteristics in LLMs. By employing zero-shot classification, LLMs can respond to questionnaire items, allowing researchers to aggregate responses into scale scores that reflect traits such as agreeableness.
The authors emphasize the complexity and opacity of how LLMs acquire traits from the text corpora they are trained on, which embed the beliefs and biases of diverse human authors. While drawing parallels between human and LLM trait acquisition, they caution against anthropomorphizing AI models, noting that LLMs operate as prediction machines with a narrower range of traits compared to the rich psychological landscape of humans. The section concludes by underscoring the need for further research in “AI psychometrics,” an interdisciplinary field that merges psychology, linguistics, and computer science, highlighting the significant social and economic implications of AI as it continues to evolve.
Methods
In this section, the authors outline various methods for assessing large language models (LLMs) through psychometric inventories. The primary approach discussed is masked language modeling (MLM), where models are pretrained on extensive text corpora by predicting masked tokens within sequences. This unsupervised training allows for the evaluation of models using sequences like “I am [MASK]” to derive probabilities for various adjectives linked to psychological traits. Another method highlighted is next-word prediction, commonly used in generative models like GPT, which predicts subsequent words based on given prompts. This method, while flexible, introduces challenges such as the need for careful prompt engineering and the stochastic nature of model outputs.
Additionally, the authors describe zero-shot classification, which leverages models fine-tuned on natural language inference (NLI) tasks to classify responses without prior examples. This approach is particularly useful for psychometric assessments, as it simplifies evaluation and avoids issues related to generative models. The authors illustrate their methodology using the Big Five Inventory (BFI) and other psychometric tools to derive psychological profiles of LLMs, enabling comparisons with human norms and among different models. They emphasize that while these assessments do not aim to benchmark model performance, they provide valuable insights into the psychological traits exhibited by LLMs, potentially revealing biases or controversial characteristics inherent in these models.
Discussion
The discussion section of the paper traces the historical intersection of psychometrics and artificial intelligence (AI), highlighting early attempts to assess AI capabilities through cognitive tests, such as the Wechsler Adult Intelligence Scale and the Stanford-Binet test. Notably, foundational figures like Thomas G. Evans and Allen Newell emphasized the importance of integrating psychometric principles into AI development. However, early AI models were criticized for their focus on “cold cognition,” lacking emotional and motivational complexity. In contrast, recent advancements in large language models (LLMs) have demonstrated significant improvements in natural language processing, enabling these models to exhibit characteristics associated with “hot cognition,” such as emotional responses and human-like behaviors.
The paper advocates for a comprehensive approach to AI psychometrics that encompasses both cognitive and noncognitive traits, including personality, values, and moral beliefs. By leveraging established psychometric inventories, the authors propose methods for assessing LLMs’ psychological profiles, using techniques like zero-shot classification. Initial findings indicate that LLMs generally display balanced personality traits, scoring low on undesirable characteristics associated with the dark tetrad, and suggest potential biases in moral norms and gender beliefs encoded in these models. The authors call for further empirical studies to explore the implications of these traits on LLM behavior and decision-making, emphasizing the need for responsible AI systems that account for the psychological complexities of these models.
