DOI: https://doi.org/10.1038/s41598-026-62286-9
PMID: https://pubmed.ncbi.nlm.nih.gov/42469328
تاريخ النشر: 2026-07-17
المؤلف: Azza Bouleimen وآخرون
الموضوع الرئيسي: الصحة النفسية من خلال الكتابة
نظرة عامة
تبحث الدراسة في قدرة نماذج اللغة الكبيرة (LLMs) على محاكاة المحادثات البشرية على منصات التواصل الاجتماعي، من خلال تقييم المحادثات التي تم إنشاؤها بواسطة Llama 3 70B و GPT-4o. من خلال مقارنة هذه الحوارات التي تم إنشاؤها بواسطة LLM مع المحادثات الحقيقية المستمدة من Reddit، وجدت الدراسة أن المشاركين أخطأوا في تحديد المحتوى الذي تم إنشاؤه بواسطة LLM على أنه من صنع الإنسان بنسبة 39% من الوقت. ومن الجدير بالذكر أن المحادثات التي أنتجها Llama 3 تم التعرف عليها على أنها تم إنشاؤها بواسطة الذكاء الاصطناعي فقط بنسبة 56% من الوقت، مما يشير إلى أداء أفضل قليلاً من التخمين العشوائي. وهذا يشير إلى أن LLMs يمكن أن تخلق تفاعلات واقعية بما فيه الكفاية على وسائل التواصل الاجتماعي، مما يثير القلق بشأن إمكانية إساءة استخدامها في إنشاء محتوى غير أصيل.
تسلط الاستنتاجات المستخلصة من الدراسة الضوء على ثلاث نقاط رئيسية: أولاً، تظهر LLMs قدرة كبيرة على محاكاة الخطاب البشري، متجاوزة عتبة اختبار تورينغ؛ ثانياً، اختيار النموذج أمر حاسم، حيث يظهر Llama 3 70B أداءً متفوقًا في إنشاء محادثات شبيهة بالبشر مقارنةً بـ GPT-4o؛ وثالثًا، بينما يمكن لـ LLMs تقليد الحوار العام بشكل مقنع، فإنها تواجه تحديات في تكرار المناقشات المشحونة عاطفياً والاستقطاب، وهي جوانب أساسية من تفاعلات وسائل التواصل الاجتماعي. تؤكد هذه النتائج على التطبيقات الواعدة لـ LLMs في نمذجة الوكلاء والآثار الأخلاقية لاستخدامها في المحاكاة الاجتماعية.
مقدمة
تناقش مقدمة ورقة البحث تطبيق نمذجة الوكلاء (ABM) في فهم السلوكيات الجماعية مثل الاستقطاب السياسي، والفصل، والديناميات الثقافية. بينما يمكن أن توضح ABM الشروط الدنيا اللازمة لظهور الظواهر الاجتماعية، فإنها غالبًا ما تبسط السلوك البشري، مما يحد من فائدتها في صنع السياسات. لتعزيز التطبيق العملي لـ ABM، من الضروري أن تستند سلوكيات الأفراد وتفاعلات الوكلاء إلى بيانات تجريبية. ومع ذلك، فإن التحديات مثل ندرة البيانات وتغير السلوك البشري تعيق معايرة النموذج والقدرات التنبؤية.
تقدم التطورات الأخيرة في الذكاء الاصطناعي التوليدي، وخاصة نماذج اللغة الكبيرة (LLMs)، حلولاً واعدة لهذه التحديات من خلال تمكين محاكاة أكثر دقة للسلوك الجماعي. يمكن أن تكمل LLMs البيانات التجريبية وتسهّل تفاعلات الوكلاء الواقعية، مما يجعلها أدوات قيمة لدراسة الديناميات الاجتماعية، خاصة على منصات التواصل الاجتماعي. نظرًا للمخاوف الأخلاقية المحيطة بالتجارب واسعة النطاق على هذه المنصات، توفر نمذجة الوكلاء التوليدية بيئة خاضعة للتحكم لاختبار استراتيجيات التدخل. تسلط الورقة الضوء على الأهمية المتزايدة لنمذجة الوكلاء التوليدية في محاكاة ديناميات وسائل التواصل الاجتماعي، مشيرةً إلى الدراسات المبكرة التي تظهر قدرة LLMs على تكرار الشبكات الاجتماعية عبر الإنترنت الشبيهة بالبشر وتقييم تأثير الخوارزميات المختلفة على تفاعلات المستخدمين. على الرغم من التقدم، يعترف المؤلفون بأن العديد من الأسئلة لا تزال قائمة بشأن الاستخدام الفعال لنمذجة الوكلاء التوليدية في هذا السياق.
الطرق
في هذه الدراسة، صمم الباحثون تجربة لمقارنة المحادثات البشرية وتلك التي تم إنشاؤها بواسطة الذكاء الاصطناعي باستخدام بيانات حقيقية مستمدة من Reddit. تضمنت إعداد التجربة إنشاء واجهة قائمة على الويب تحاكي تخطيط Reddit، مما يسمح للمشاركين بالتفاعل مع كلا النوعين من المحادثات. تم تسهيل مهمة التوضيح بواسطة مكتبة Potato 50 Python، التي تم تخصيصها لتلبية المتطلبات المحددة للدراسة.
بدأ المشاركون التجربة بمراجعة صفحة معلومات توضح المهمة، التي تضمنت التمييز بين المحادثات التي تم إنشاؤها بواسطة الذكاء الاصطناعي وتلك البشرية. بعد ذلك، أكملوا استبيان ما قبل الفحص لتقديم الموافقة المستنيرة والإجابة على أسئلة فحص الانتباه. فقط أولئك الذين اجتازوا هذه الفحوصات انتقلوا إلى مرحلة التوضيح، حيث قاموا بأداء ستة توضيحات واختبار انتباه إضافي. لضمان تجربة متنوعة، قام الباحثون بتحسين مكتبة Potato بخوارزمية أخذ عينات مخصصة، مما يمنع المشاركين من مواجهة نفس منشور Reddit عدة مرات، بغض النظر عن اختلافات طول المحادثة، أو نموذج اللغة (LLM)، أو إعدادات درجة الحرارة.
النتائج
في هذه الدراسة، ساهم 251 مشاركًا في تحليل 1741 توضيحًا يقارن بين المناقشات الحقيقية والمحاكاة. بعد استبعاد بيانات 46 مشاركًا لم يجتازوا فحوصات الانتباه، تضمنت مجموعة البيانات النهائية 1205 توضيحات من 203 مشاركين، مما يلبي حجم العينة المطلوب لتحقيق قوة إحصائية بنسبة 80%. أظهرت النتائج معدل نجاح قدره 0.61، مما يعني أن المشاركين حددوا المحادثات البشرية بشكل صحيح في 61% من الحالات، بينما تم إدراك 39% من المناقشات التي تم إنشاؤها بواسطة نماذج اللغة الكبيرة (LLMs) بشكل خاطئ على أنها من صنع الإنسان.
المناقشة
في هذه الدراسة، استكشفنا قدرة نماذج اللغة الكبيرة (LLMs) على تكرار المحادثات الشبيهة بالبشر على وسائل التواصل الاجتماعي، باستخدام بيانات حقيقية من Reddit. تشير نتائجنا إلى أن المشاركين لم يتمكنوا من تحديد المحادثات التي تم إنشاؤها بواسطة الذكاء الاصطناعي بنسبة 39% من الوقت، متجاوزين العتبة البالغة 30% المرتبطة باختبار تورينغ. ومن الجدير بالذكر أن المحادثات التي تم إنشاؤها بواسطة Llama 3 70B كانت أكثر تحديًا بشكل ملحوظ للمشاركين لتمييزها عن المحادثات البشرية مقارنةً بتلك التي أنتجها GPT-4o، مما يشير إلى أن Llama 3 قد تحاكي بشكل أفضل النغمة غير الرسمية والبنية النموذجية للخطاب على Reddit. ومع ذلك، كشفت الدراسة أيضًا أن طول المحادثة لم يؤثر بشكل كبير على معدلات الكشف، على الرغم من أنه تم ملاحظة علاقة عكسية-على شكل حرف U، حيث بلغ النجاح ذروته عند طول ثمانية تعليقات قبل أن ينخفض.
سلط التحليل النوعي لتعليقات المشاركين الضوء على الاختلافات الرئيسية بين المحادثات التي تم إنشاؤها بواسطة الذكاء الاصطناعي وتلك البشرية، حيث أشار العديد إلى أن ردود الذكاء الاصطناعي كانت تميل إلى أن تكون رسمية للغاية، تفتقر إلى التعقيد العاطفي، وتفتقر إلى الأصالة. غالبًا ما أشار المشاركون إلى ميزات أسلوب النص، مثل الأدب والإطناب، كدلائل على المحتوى الذي تم إنشاؤه بواسطة الذكاء الاصطناعي. تؤكد هذه النتائج على قيود نماذج اللغة الكبيرة الحالية في التقاط تعقيدات التفاعل البشري، خاصة في السياقات التي تنطوي على شدة عاطفية أو صراع. للمضي قدمًا، تؤكد الدراسة على الحاجة إلى تحسين تصميم المطالبات واختيار النموذج لتعزيز واقعية المحادثات التي تم إنشاؤها بواسطة الذكاء الاصطناعي، خاصة للتطبيقات في نمذجة الوكلاء (ABM) التي تهدف إلى محاكاة الديناميات الاجتماعية مثل الاستقطاب والمعلومات المضللة.
DOI: https://doi.org/10.1038/s41598-026-62286-9
PMID: https://pubmed.ncbi.nlm.nih.gov/42469328
Publication Date: 2026-07-17
Author(s): Azza Bouleimen et al.
Primary Topic: Mental Health via Writing
Overview
The research investigates the capability of Large Language Models (LLMs) to simulate human conversations on social media platforms, specifically through an evaluation of conversations generated by Llama 3 70B and GPT-4o. By comparing these LLM-generated dialogues with authentic conversations sourced from Reddit, the study found that participants misidentified LLM-generated content as human-created 39% of the time. Notably, conversations produced by Llama 3 were recognized as AI-generated only 56% of the time, indicating a performance that is only marginally better than random guessing. This suggests that LLMs can create sufficiently realistic social media interactions, raising concerns about their potential misuse in generating inauthentic content.
The conclusions drawn from the study highlight three main points: first, LLMs demonstrate a significant ability to simulate human discourse, surpassing the Turing Test threshold; second, the choice of model is crucial, with Llama 3 70B showing superior performance in generating human-like conversations compared to GPT-4o; and third, while LLMs can convincingly mimic general dialogue, they face challenges in replicating emotionally charged discussions and polarization, which are essential aspects of social media interactions. These findings underscore both the promising applications of LLMs in agent-based modeling and the ethical implications of their use in social simulations.
Introduction
The introduction of the research paper discusses the application of Agent-Based Modeling (ABM) in understanding collective behaviors such as political polarization, segregation, and cultural dynamics. While ABM can illustrate the minimal conditions necessary for social phenomena to emerge, it often oversimplifies human behavior, which limits its utility in policy-making. To enhance the practical applicability of ABM, it is crucial to ground individual behavior and agent interactions in empirical data. However, challenges such as data scarcity and the variability of human behavior hinder model calibration and predictive capabilities.
Recent advancements in generative AI, particularly Large Language Models (LLMs), offer promising solutions to these challenges by enabling more nuanced simulations of collective behavior. LLMs can complement empirical data and facilitate realistic agent interactions, making them valuable tools for studying social dynamics, especially on social media platforms. Given the ethical concerns surrounding large-scale experimentation on these platforms, generative ABM provides a controlled environment for testing intervention strategies. The paper highlights the increasing relevance of generative ABM in simulating social media dynamics, noting early studies that demonstrate LLMs’ ability to replicate human-like online social networks and assess the impact of different algorithms on user interactions. Despite the progress, the authors acknowledge that many questions remain regarding the effective use of generative ABM in this context.
Methods
In this study, the researchers designed an experiment to compare human and AI-generated conversations using real data sourced from Reddit. The experimental setup involved creating a web-based interface that mimicked Reddit’s layout, allowing participants to engage with both types of conversations. The annotation task was facilitated by the Potato 50 Python library, which was customized to meet the specific requirements of the study.
Participants began the experiment by reviewing an information page that outlined the task, which involved distinguishing between AI-generated and human conversations. Following this, they completed a pre-screening survey to provide informed consent and answer attention-check questions. Only those who passed these checks proceeded to the annotation phase, where they performed six annotations and an additional attention test. To ensure a diverse experience, the researchers enhanced the Potato library with a custom sampling algorithm, preventing participants from encountering the same Reddit post multiple times, regardless of variations in conversation length, language model (LLM), or temperature settings.
Results
In this study, 251 participants contributed to the analysis of 1741 annotations comparing authentic and simulated discussions. After excluding data from 46 participants who did not pass attention checks, the final dataset comprised 1205 annotations from 203 participants, meeting the required sample size for achieving 80% statistical power. The results indicated a success rate of 0.61, meaning that participants correctly identified human conversations in 61% of cases, while 39% of the discussions generated by large language models (LLMs) were incorrectly perceived as human-made.
Discussion
In this study, we explored the ability of large language models (LLMs) to replicate human-like conversations on social media, specifically using real Reddit data. Our findings indicate that participants were unable to identify AI-generated conversations 39% of the time, surpassing the 30% threshold associated with the Turing Test. Notably, conversations generated by Llama 3 70B were significantly more challenging for participants to distinguish from human conversations compared to those produced by GPT-4o, suggesting that Llama 3 may better emulate the informal tone and structure typical of Reddit discourse. However, the study also revealed that conversation length did not significantly influence detection rates, although an inverse-U relationship was observed, where success peaked at a length of eight comments before declining.
Qualitative analysis of participant feedback highlighted key differences between AI and human conversations, with many noting that AI responses tended to be overly formal, lacking in emotional nuance, and devoid of authenticity. Participants frequently cited textual style features, such as politeness and verbosity, as indicators of AI-generated content. These findings underscore the limitations of current LLMs in capturing the complexities of human interaction, particularly in contexts involving emotional intensity or conflict. Moving forward, the study emphasizes the need for improved prompt design and model selection to enhance the realism of AI-generated conversations, especially for applications in agent-based modeling (ABM) that aim to simulate social dynamics like polarization and misinformation.
