DOI: https://doi.org/10.3389/fpsyg.2026.1782184
PMID: https://pubmed.ncbi.nlm.nih.gov/41853835
تاريخ النشر: 2026-03-03
المؤلف: Pegah Imannezhad وآخرون
الموضوع الرئيسي: تطور التعلم لدى الأطفال والحيوانات
نظرة عامة
تستكشف هذه الدراسة أحكام الاحتمالات لنموذج اللغة الكبير (LLM) GPT-5، لا سيما فيما يتعلق بقواعد الاحتمالات الكلاسيكية، بما في ذلك أخطاء الاقتران، وأخطاء الانفصال، وانتهاكات التكامل الثنائي. من خلال استخدام مجموعة بيانات كبيرة من أحكام الاحتمالات البشرية، حيث أظهر المشاركون أخطاء متنوعة، قام الباحثون بتقييم أداء GPT-5 في مهام مماثلة مع ملفات تعريف المشاركين المقارنة.
كشفت النتائج أن GPT-5 ارتكب في الغالب أخطاء اقتران أو انفصال فردية فقط وأظهر التزامًا شبه كامل بقيود التكامل الثنائي. ومن الجدير بالذكر أن أنماط استجابته كانت أكثر اتساقًا مع النماذج الكمومية الاحتمالية المبكرة مقارنة بالنماذج الأحدث التي تأخذ في الاعتبار الضوضاء. تشير هذه النتائج إلى أن GPT-5 يظهر نهجًا أكثر اتساقًا وتماسكًا داخليًا في التفكير الاحتمالي مقارنة بالمشاركين البشريين الساذجين، مما يبرز إمكانيته في المجالات التي تتطلب تقييمات احتمالية.
مقدمة
تناقش مقدمة الورقة التأثير التحويلي لنماذج اللغة الكبيرة (LLMs) على معالجة اللغة الطبيعية، مع تسليط الضوء على تطبيقاتها في توليد النصوص، والترجمة، وأنظمة الذكاء الاصطناعي التفاعلية. مع تكامل نماذج اللغة الكبيرة في المهام الإدراكية واتخاذ القرار، هناك حاجة متزايدة لتقييم نقاط قوتها وقيودها، لا سيما في الاستدلال الاحتمالي. تم تصميم التطورات الأخيرة، مثل GPT-5 من OpenAI، للتفكير المعقد ودعم القرار، مما أدى إلى زيادة الاعتماد على هذه النماذج للإرشاد الأكاديمي والمهني. ومع ذلك، لا تزال المخاوف بشأن خصوصية البيانات وجودة النصائح المقدمة من LLMs غير مفحوصة إلى حد كبير.
تعتمد فعالية LLMs في تقديم نصائح موثوقة على بعدين رئيسيين: دقة معرفتها الخاصة بالمجال والتزامها بالمعايير النمطية لصنع القرار العقلاني. البعد الأول يعتمد على السياق، بينما يتعامل الثاني مع مناقشات أوسع في علم الإدراك حول عقلانية الاستدلالات الاحتمالية. يعتمد الاستدلال الاحتمالي النمطي على مبادئ بايزي، التي تقضي بأن أحكام الاحتمالات يجب أن تلبي معايير مثل الإضافة والتماسك. على الرغم من أن التفكير البشري غالبًا ما يتماشى مع هذه المبادئ، إلا أنه يظهر أيضًا انحيازات منهجية، كما يتضح من الأخطاء المعرفية الموثقة جيدًا. تؤكد هذه الثنائية على ضرورة التقييم النقدي لقدرات LLMs في التفكير الاحتمالي.
الطرق
في دراستهم، استخدم هوانغ وآخرون (2025) نموذجًا تجريبيًا منظمًا للتحقيق في التفكير الاحتمالي المتعلق بالانتخابات الرئاسية الأمريكية لعام 2020. تم تكليف المشاركين بتقديم 78 حكمًا احتماليًا بشأن النتائج المحتملة لفوز المرشحين في ولايات معينة، تم تنظيمها في مجموعتين: أوهايو-ميزوري-ميشيغان (المجموعة 1) وجورجيا-مونتانا-نيفادا (المجموعة 2). تضمنت كل مجموعة تركيبات متنوعة من الأحداث—الأحداث الهامشية، الاقترانات، الانفصالات، والشرطيات—بمجموع 78 عنصرًا، مع توازن إضافي للتخفيف من تأثيرات الترتيب. استخدم المشاركون شريط تمرير مستمر للتعبير عن أحكامهم، وطبقت الدراسة تصميمًا بين الموضوعات 2 × 2 للتلاعب بمجموعة الولايات وتعقيدها، مع التركيز على ظروف منخفضة التعقيد لتفسير أوضح.
تضمن التلاعب بالتعقيد هيكلة كتل الأحداث المركبة إما حسب أزواج الولايات (تعقيد منخفض) أو عن طريق خلط الأزواج (تعقيد مرتفع)، وكان الهدف من ذلك التأثير على تماسك المشاركين الاحتمالي. لم يجد هوانغ وآخرون أي تأثيرات منهجية تتعلق بالتعقيد، مما أدى إلى القرار بالتركيز على ظروف التعقيد المنخفض (T1LC و T2LC) للتحليلات اللاحقة. أسفر هذا النهج عن عينة من 553 شخصية ديموغرافية متميزة من 284 مشاركًا في T1LC و269 في T2LC، مما يوفر مجموعة بيانات قوية للتحقيق الحالي باستخدام GPT-5.
النتائج
تشير نتائج الدراسة إلى وجود فرق كبير في أداء التفكير الاحتمالي بين المشاركين البشريين وشخصيات GPT-5 المتطابقة، كما تم تقييمه من خلال سلسلة من اختبارات t لعينة مزدوجة. أظهر GPT-5 التزامًا قويًا بالمبادئ الاحتمالية، مع معدلات أقل بشكل ملحوظ من أخطاء التفكير الكلاسيكية مقارنة بالبشر. على وجه التحديد، من حيث أداء التكامل، كانت انحرافات GPT-5 ضئيلة (المتوسط $M = 0.00004$، الانحراف المعياري $SD = 0.00045$)، مما يشير إلى سلوك قريب من النمط، بينما أظهر المشاركون البشر انحرافات أكبر وأكثر تنوعًا (المتوسط $M = 0.248$، الانحراف المعياري $SD = 0.238$).
كشفت التحليلات الإضافية للتكامل المشترك عن انحراف أكثر وضوحًا، حيث ظلت انحرافات GPT-5 ضئيلة (المتوسط $M = 0.00108$، الانحراف المعياري $SD = 0.00820$). في المقابل، انحرفت أحكام البشر بشكل كبير عن القيمة النظرية، بمتوسط يزيد عن نقطة احتمال كاملة (المتوسط $M = 1.250$، الانحراف المعياري $SD = 0.656$). تؤكد هذه النتائج على قدرات التفكير الاحتمالي المتفوقة لـ GPT-5 مقارنة بالمشاركين البشر.
المناقشة
تسلط قسم المناقشة في ورقة البحث الضوء على نتائج مهمة تتعلق بأخطاء التفكير الاحتمالي الكلاسيكية، لا سيما أخطاء الاقتران والانفصال، بالإضافة إلى انتهاكات التكامل الثنائي. هذه الأخطاء موثقة جيدًا في الحكم البشري، مما يظهر انحرافات منهجية عن نظرية الاحتمالات النمطية. خطأ الاقتران، الذي يوضحه عمل تفرسكي وكاهنمان (1983) حول احتمال أن يتجاوز الاقتران احتمال مكوناته، هو قوي بشكل خاص عبر سياقات متنوعة. وبالمثل، فإن خطأ الانفصال، حيث يُحكم على احتمال الانفصال بأنه أقل من أحد الأحداث المكونة له، يتعارض أيضًا مع المسلمات الأساسية للاحتمالات. تكشف انتهاكات التكامل الثنائي، على الرغم من أنها أقل مناقشة، عن قيود حرجة في التفكير الاحتمالي البشري، خاصة في السياقات الغامضة.
تستكشف الورقة أيضًا أداء GPT-5، وهو نموذج لغوي متقدم، فيما يتعلق بهذه الانحيازات الاحتمالية. تشير النتائج الأولية إلى أن GPT-5 يظهر التزامًا أعلى بمبدأ التكامل الثنائي مقارنة بالمشاركين البشر، مما يدل على تماسك أكثر منهجية في أحكامه الاحتمالية. ومع ذلك، لا يزال أداء النموذج في أخطاء الاقتران والانفصال أقل وضوحًا، حيث أظهرت النماذج السابقة ميولًا لارتكاب هذه الأخطاء. تهدف الدراسة إلى تقييم ما إذا كان تفكير GPT-5 يتماشى مع أنماط الحكم الاحتمالي البشري أو ينحرف بطرق فريدة، وبالتالي تقييم موثوقيته كنظام دعم للقرار. هذه الاستفسارات في الوقت المناسب، نظرًا للاعتماد المتزايد على نماذج اللغة للحصول على النصائح في مجالات متنوعة، مما يبرز ضرورة وجود أطر تقييم مبدئية في تطبيقاتها.
DOI: https://doi.org/10.3389/fpsyg.2026.1782184
PMID: https://pubmed.ncbi.nlm.nih.gov/41853835
Publication Date: 2026-03-03
Author(s): Pegah Imannezhad et al.
Primary Topic: Child and Animal Learning Development
Overview
This study investigates the probability judgments of the Large Language Model (LLM) GPT-5, particularly in relation to classical probability rules, including conjunction fallacies, disjunction fallacies, and binary complementarity violations. By utilizing a substantial dataset of human probabilistic judgments, where participants exhibited various fallacies, the researchers assessed GPT-5’s performance on analogous tasks with comparable participant profiles.
The results revealed that GPT-5 predominantly committed only single conjunction or disjunction fallacies and demonstrated near-perfect adherence to binary complementarity constraints. Notably, its response patterns were more consistent with early quantum-probabilistic models than with newer models that account for noise. These findings indicate that GPT-5 exhibits a more coherent and internally consistent approach to probabilistic reasoning compared to naïve human participants, highlighting its potential utility in domains requiring probabilistic assessments.
Introduction
The introduction of the paper discusses the transformative impact of Large Language Models (LLMs) on natural language processing, highlighting their applications in text generation, translation, and interactive AI systems. As LLMs become integral to cognitive tasks and decision-making, there is a growing need to evaluate their strengths and limitations, particularly in probabilistic inference. Recent advancements, such as OpenAI’s GPT-5, are designed for complex reasoning and decision support, leading to increased reliance on these models for academic and professional guidance. However, concerns about data privacy and the quality of advice provided by LLMs remain largely unexamined.
The effectiveness of LLMs in offering reliable advice hinges on two key dimensions: the accuracy of their domain-specific knowledge and their adherence to normative standards of rational decision-making. The first dimension is context-dependent, while the second engages with broader cognitive science debates regarding the rationality of probabilistic inferences. Normative probabilistic inference is based on Bayesian principles, which dictate that probabilistic judgments should meet criteria such as additivity and coherence. Although human reasoning often aligns with these principles, it also exhibits systematic biases, as evidenced by well-documented cognitive fallacies. This duality underscores the necessity for a critical assessment of LLMs’ capabilities in probabilistic reasoning.
Methods
In their study, Huang et al. (2025) employed a structured experimental paradigm to investigate probabilistic reasoning related to the 2020 U.S. presidential election. Participants were tasked with providing 78 probability judgments regarding the potential outcomes of candidates winning specific states, organized into two triplets: Ohio-Missouri-Michigan (Triplet 1) and Georgia-Montana-Nevada (Triplet 2). Each triplet included various combinations of events—marginal events, conjunctions, disjunctions, and conditionals—totaling 78 items, with additional counterbalancing to mitigate order effects. Participants utilized a continuous slider to express their judgments, and the study implemented a 2 × 2 between-subjects design to manipulate State Triplet and Complexity, with a focus on low-complexity conditions for clearer interpretation.
The complexity manipulation involved structuring composite-event blocks either by state pair (low complexity) or by intermixing pairs (high complexity), which was intended to influence the participants’ probabilistic coherence. Huang et al. found no systematic effects related to complexity, leading to the decision to focus on the low-complexity conditions (T1LC and T2LC) for subsequent analyses. This approach resulted in a sample of 553 distinct demographic personas from 284 participants in T1LC and 269 in T2LC, providing a robust dataset for the current investigation utilizing GPT-5.
Results
The results of the study indicate a significant difference in probabilistic reasoning performance between human participants and matched GPT-5 personas, as assessed through a series of paired-samples t-tests. GPT-5 exhibited a robust adherence to probabilistic principles, with notably lower rates of classical reasoning fallacies compared to humans. Specifically, in terms of complementarity performance, GPT-5’s deviations were minimal (mean $M = 0.00004$, standard deviation $SD = 0.00045$), suggesting near-normative behavior, while human participants displayed larger and more variable deviations (mean $M = 0.248$, standard deviation $SD = 0.238$).
Further analysis of joint complementarity revealed an even more pronounced divergence, with GPT-5’s deviations remaining negligible (mean $M = 0.00108$, standard deviation $SD = 0.00820$). In contrast, human judgments deviated significantly from the theoretical value, averaging more than a full probability point (mean $M = 1.250$, standard deviation $SD = 0.656$). These findings underscore the superior probabilistic reasoning capabilities of GPT-5 relative to human participants.
Discussion
The discussion section of the research paper highlights significant findings regarding classical probabilistic reasoning fallacies, specifically the conjunction and disjunction fallacies, as well as violations of binary complementarity. These fallacies are well-documented in human judgment, demonstrating systematic deviations from normative probability theory. The conjunction fallacy, illustrated by Tversky and Kahneman’s (1983) work on the probability of a conjunction exceeding that of its components, is particularly robust across various contexts. Similarly, the disjunction fallacy, where the probability of a disjunction is judged to be lower than that of one of its constituent events, also contradicts fundamental axioms of probability. Violations of binary complementarity, though less frequently discussed, reveal critical constraints in human probabilistic reasoning, especially in ambiguous contexts.
The paper further explores the performance of GPT-5, a state-of-the-art language model, in relation to these probabilistic biases. Initial findings suggest that GPT-5 exhibits a higher adherence to the principle of binary complementarity compared to human participants, indicating a more systematic coherence in its probability judgments. However, the model’s performance on conjunction and disjunction fallacies remains less clear, with earlier models showing tendencies to commit these errors. The study aims to evaluate whether GPT-5’s reasoning aligns with human patterns of probabilistic judgment or diverges in unique ways, thereby assessing its reliability as a decision-support system. This inquiry is timely, given the increasing reliance on language models for advice in various domains, underscoring the necessity for principled evaluation frameworks in their application.
