DOI: https://doi.org/10.1002/advs.202524142
PMID: https://pubmed.ncbi.nlm.nih.gov/41873462
تاريخ النشر: 2026-03-24
المؤلف: Silvia Rondini وآخرون
الموضوع الرئيسي: الإبداع في التعليم وعلوم الأعصاب
نظرة عامة
تستكشف هذه الدراسة الأداء الإبداعي المقارن في توليد الصور بين المشاركين البشريين (الفنانين البصريين وغير الفنانين) ونموذج الذكاء الاصطناعي التوليدي (GenAI) تحت ظروف مختلفة من المدخلات البشرية. تكشف النتائج عن تسلسل هرمي مميز للإبداع: يتفوق الفنانون البصريون على غير الفنانين، الذين بدورهم أكثر إبداعًا من GenAI المستوحى من البشر، مع عرض GenAI الموجه ذاتيًا أقل مستوى من الإبداع. ومن الجدير بالذكر أن الدراسة تسلط الضوء على أن زيادة الإرشاد البشري تعزز بشكل كبير الناتج الإبداعي لـ GenAI، مما يجعله أقرب إلى ذلك الخاص بغير الفنانين. علاوة على ذلك، أظهرت تقييمات الإبداع من قبل المقيمين البشريين ونموذج الذكاء الاصطناعي (GPT-4o) أنماط حكم متباينة، مما يشير إلى أن GenAI يواجه تحديات فريدة في الإبداع البصري تختلف عن تلك الموجودة في المهام اللغوية.
تؤكد الاستنتاجات على أن قيود GenAI في الإبداع البصري تنبع من اختلافات أساسية في الإدراك والوعي مقارنة بالبشر، لا سيما فيما يتعلق بالحساسية لتعقيد العالم الحقيقي. تؤكد الدراسة على ضرورة وجود مقاييس تقييم أوسع لالتقاط تعقيدات الإبداع في الوكلاء الحاسوبية. بينما يظهر GenAI إمكانات للتعبير الإبداعي عند توجيهه، إلا أنه لا يزال لا يحقق أداءً مستقلاً يمكن مقارنته بالبشر، مما يشير إلى أن الإبداع البشري متجذر بعمق في العمليات التطورية التي تتفاعل مع البيئات المعقدة.
مقدمة
تناقش مقدمة هذه الورقة البحثية الاهتمام المتزايد في نماذج الذكاء الاصطناعي التوليدي (GenAI)، وخاصة نماذج اللغة الكبيرة (LLMs)، التي أظهرت قدرات كبيرة في تقليد الإبداع البشري عبر المجالات اللفظية والبصرية. تُعرف هذه النماذج غالبًا باسم “الموسى الاصطناعية” أو “محركات الدهشة”، وقد تم تقييمها بشكل أساسي من خلال تقييمات التفكير المتباين (DT)، التي تقيس الإبداع بناءً على أبعاد مثل المرونة، والطلاقة، والأصالة، والتفصيل. بينما أظهرت LLMs أنها تتفوق على البشر في العديد من مهام DT، توجد استثناءات حيث لا يزال الأفراد ذوو الإبداع العالي يتفوقون على هذه النماذج. علاوة على ذلك، كانت LLMs فعالة في تقييم الإبداع في كل من مخرجات البشر وGenAI، حيث تتوافق بشكل جيد مع أحكام البشر.
على الرغم من التقدم في أبحاث الإبداع القائم على اللغة، لا يزال الإبداع البصري غير مستكشف بشكل كافٍ، لا سيما في سياق توليد الصور. تقدم الورقة مفهوم الصور الذهنية الإبداعية (CMI)، الذي يشير إلى التوليد الداخلي للصور الأصلية، مما يبرز أهميته في فهم الإبداع بما يتجاوز المخرجات اللفظية. تهدف الدراسة إلى التحقيق في قدرات الخيال البصري لنموذج الانسياب المستقر مقارنة بالمشاركين البشريين، مع التركيز على كل من إنتاج وتقييم الصور الذهنية الإبداعية. من خلال استخدام درجات متفاوتة من الإرشاد البشري في التحفيز، تسعى البحث إلى توضيح الفروق في العمليات الإبداعية بين البشر وGenAI، مع معالجة قضية النية في إبداع GenAI. تتكون الدراسة من أربع مراحل، بما في ذلك توليد الرسومات الإبداعية وتقييم هذه الصور من قبل كل من المقيمين البشريين وGenAI عبر أبعاد إبداعية متعددة.
الطرق
في هذه الدراسة، تم استخدام تحليل العوامل (FA) لتقييم درجات الإبداع للصور التي تم إنشاؤها بواسطة البشر مقابل تلك التي أنتجها الذكاء الاصطناعي التوليدي (GenAI). كشفت التحليلات أن الصور التي أنشأها البشر حصلت على درجات أعلى بكثير من صور GenAI (χ²(1) = 229.71، p < .001). أظهرت التحقيقات الإضافية باستخدام نماذج التأثيرات المختلطة الخطية أن فئة الصورة أثرت بشكل كبير على درجات الإبداع العامة (χ²(3) = 457.03، p < .001). من الجدير بالذكر أن الصور من مجموعة الفنانين البصريين حققت أعلى درجات الإبداع، تليها غير الفنانين، وGenAI المستوحى من البشر، وصور GenAI الموجه ذاتيًا. أكدت المقارنات الزوجية وجود اختلافات كبيرة بين الفئات، لا سيما بين الفنانين البصريين وغير الفنانين (التقدير = 0.42، p < .001)، والفنانين البصريين وHI-GenAI (التقدير = 0.51، p < .001)، والفنانين البصريين وSG-GenAI (التقدير = 0.97، p < .001). كانت الاختلافات الوحيدة غير المهمة بين غير الفنانين وHI-GenAI (التقدير = 0.09، p = .198)، مما يشير إلى أن تحفيز GenAI بأفكار بشرية يمكن أن يعزز إبداعه إلى مستويات قابلة للمقارنة مع البشر غير الخبراء. استخدمت الدراسة نسخة معدلة من TCIA، حيث تم تقديم 12 محفزًا تجريديًا للمشاركين، نصفها كان أصليًا والنصف الآخر تم إنشاؤه حديثًا. بالإضافة إلى معايير TCIA الأصلية لـ "الحيوية" و"الأصالة"، دمجت البحث مقاييس لـ "الإعجاب" و"الجماليات" و"الفضول" لتقييم الإبداع (مفصل في القسم S1). تم الحصول على الموافقة الأخلاقية من لجنة أخلاقيات البحث السريري بمستشفى بيلفيتج الجامعي، وقدم المشاركون موافقة مستنيرة وتلقوا تعويضًا ماليًا عن مشاركتهم.
النتائج
في قسم النتائج، يقدم المؤلفون أمثلة على مخرجات الرسم، موضحة في الشكل 2. تم تقييم تقييمات الإبداع لمجموعة الصور، التي تم إنشاؤها في المرحلتين I وII، من قبل كل من المقيمين البشريين ومقيمين GenAI (GPT-4o) في المراحل اللاحقة III وIV. تم إجراء تحليل هذه التقييمات على مرحلتين: أولاً، تم تحليل التقييمات من المقيمين البشريين بشكل فردي، تلاها تحليل مقارن مع التقييمات المقدمة من نظام GenAI. يسمح هذا النهج المزدوج بفهم شامل للإبداع المدرك للصور المولدة عبر وجهات نظر تقييمية مختلفة.
المناقشة
تستكشف الدراسة الإبداع البصري المقارن للفنانين البشريين وغير الفنانين مقابل نماذج الذكاء الاصطناعي التوليدي (GenAI)، مع التركيز بشكل خاص على اختبار قدرات الصور الإبداعية (TCIA). تشير النتائج إلى أن الصور التي أنشأها البشر، وخاصة من الفنانين البصريين، تم تقييمها على أنها أكثر إبداعًا بشكل كبير من تلك التي أنتجها GenAI، مع ملاحظة تسلسل هرمي واضح: الفنانون البصريون > غير الفنانين > GenAI المستوحى من البشر > GenAI الموجه ذاتيًا. من الجدير بالذكر أن دمج الأفكار البشرية في التحفيزات لـ GenAI عزز بشكل كبير الإبداع المدرك، مما يشير إلى أن الإرشاد البشري أمر حاسم لتحسين مخرجات GenAI.
في المقابل، انحرفت تقييمات نموذج GPT-4o عن التقييمات البشرية، حيث أظهرت تمييزًا أقل عبر فئات الصور ودرجات أعلى بشكل عام لصور GenAI. يسلط هذا التباين الضوء على قيود GenAI في الإبداع البصري، حيث يكافح لتكرار الطبيعة الدقيقة والحساسة للسياق للإبداع البشري. تشير النتائج إلى أنه بينما يمكن لـ GenAI إنتاج مخرجات إبداعية تحت ظروف معينة، فإنه يفتقر إلى القدرات المستقلة والتكيفية المتأصلة في الإدراك البشري، الذي يتشكل من خلال العمليات التطورية والتفاعلات مع العالم الحقيقي. بشكل عام، تؤكد الدراسة على ضرورة وجود مقاييس تقييم أوسع لتقييم الإبداع بدقة في الوكلاء الحاسوبية، كاشفة عن اختلافات أساسية بين العمليات الإبداعية البشرية وGenAI.
DOI: https://doi.org/10.1002/advs.202524142
PMID: https://pubmed.ncbi.nlm.nih.gov/41873462
Publication Date: 2026-03-24
Author(s): Silvia Rondini et al.
Primary Topic: Creativity in Education and Neuroscience
Overview
This study investigates the comparative creative performance in image generation between human participants (Visual Artists and Non-Artists) and a Generative AI (GenAI) model under varying conditions of human input. The findings reveal a distinct creativity hierarchy: Visual Artists outperform Non-Artists, who in turn are more creative than the Human-Inspired GenAI, with the Self-Guided GenAI exhibiting the least creativity. Notably, the study highlights that increased human guidance significantly enhances the creative output of GenAI, bringing it closer to that of Non-Artists. Furthermore, the evaluation of creativity by human raters and the AI model (GPT-4o) demonstrated divergent judgment patterns, indicating that GenAI encounters unique challenges in visual creativity that differ from those in language tasks.
The conclusions emphasize that GenAI’s limitations in visual creativity stem from fundamental differences in cognition and perception compared to humans, particularly regarding sensitivity to real-world complexity. The study underscores the necessity for broader evaluative measures to capture the intricacies of creativity in computational agents. While GenAI shows potential for creative expression when guided, it does not yet achieve autonomous performance comparable to humans, suggesting that human creativity is deeply rooted in evolutionary processes that interact with complex environments.
Introduction
The introduction of this research paper discusses the burgeoning interest in Generative AI (GenAI) models, particularly Large Language Models (LLMs), which have demonstrated significant capabilities in mimicking human creativity across verbal and visual domains. These models, often referred to as “Artificial Muses” or “Engines of Wow,” have been evaluated primarily through divergent thinking (DT) assessments, which measure creativity based on dimensions such as Flexibility, Fluency, Originality, and Elaboration. While LLMs have shown to outperform humans in many DT tasks, exceptions exist where highly creative individuals still surpass these models. Moreover, LLMs have been effective in assessing creativity in both human and GenAI outputs, correlating well with human judgments.
Despite the advancements in language-based creativity research, visual creativity remains underexplored, particularly in the context of image generation. The paper introduces the concept of Creative Mental Imagery (CMI), which refers to the internal generation of original images, highlighting its significance in understanding creativity beyond verbal outputs. The study aims to investigate the visual imagination abilities of a Stable Diffusion model compared to human participants, focusing on both the production and evaluation of creative mental imagery. By employing varying degrees of human guidance in prompting, the research seeks to elucidate the differences in creative processes between humans and GenAI, while also addressing the issue of intentionality in GenAI creativity. The study comprises four phases, including the generation of creative drawings and the assessment of these images by both human and GenAI raters across multiple creativity dimensions.
Methods
In this study, Factor Analysis (FA) was employed to assess the Creativity scores of images generated by humans versus those produced by Generative AI (GenAI). The analysis revealed that human-generated images received significantly higher scores than GenAI images (χ²(1) = 229.71, p < .001). Further investigation using linear mixed-effects models indicated that the image category significantly influenced overall Creativity scores (χ²(3) = 457.03, p < .001). Notably, images from the Visual Artists group achieved the highest Creativity scores, followed by Non-Artists, Human-Inspired GenAI, and Self-Guided GenAI images. Pairwise comparisons confirmed significant differences between categories, particularly between Visual Artists and Non-Artists (estimate = 0.42, p < .001), Visual Artists and HI-GenAI (estimate = 0.51, p < .001), and Visual Artists and SG-GenAI (estimate = 0.97, p < .001). The only non-significant difference was between Non-Artists and HI-GenAI (estimate = 0.09, p = .198), suggesting that prompting GenAI with human ideas can enhance its creativity to levels comparable to non-expert humans. The study utilized an adapted version of the TCIA, presenting participants with 12 abstract stimuli, half of which were original and half newly created. In addition to the original TCIA criteria of "Vividness" and "Originality," the research incorporated measures of "Liking," "Aesthetics," and "Curiosity" to evaluate creativity (detailed in Section S1). Ethical approval was obtained from the Clinical Research Ethics Committee of the Bellvitge University Hospital, and participants provided informed consent and received monetary compensation for their involvement.
Results
In the Results section, the authors present examples of drawing outputs, illustrated in Figure 2. The creativity ratings for the image dataset, generated in phases I and II, were assessed by both human raters and GenAI (GPT-4o) raters in subsequent phases III and IV. The analysis of these ratings was conducted in two stages: first, the ratings from human evaluators were analyzed individually, followed by a comparative analysis with the ratings provided by the GenAI system. This dual approach allows for a comprehensive understanding of the perceived creativity of the generated images across different evaluative perspectives.
Discussion
The study investigates the comparative visual creativity of human artists and non-artists against generative AI (GenAI) models, specifically focusing on the Test of Creative Imagery Abilities (TCIA). Results indicate that human-generated images, particularly from Visual Artists, were rated significantly more creative than those produced by GenAI, with a clear hierarchy observed: Visual Artists > Non-Artists > Human-Inspired GenAI > Self-Guided GenAI. Notably, the incorporation of human ideas into prompts for GenAI significantly enhanced perceived creativity, suggesting that human guidance is crucial for improving GenAI outputs.
In contrast, ratings from the GPT-4o model diverged from human assessments, exhibiting less discrimination across image categories and generally higher scores for GenAI images. This discrepancy highlights the limitations of GenAI in visual creativity, as it struggles to replicate the nuanced, context-sensitive nature of human creativity. The findings suggest that while GenAI can produce creative outputs under specific conditions, it lacks the autonomous, adaptive capabilities inherent to human cognition, which is shaped by evolutionary processes and real-world interactions. Overall, the study underscores the necessity for broader evaluative measures to accurately assess creativity in computational agents, revealing fundamental differences between human and GenAI creative processes.
