الرؤية المنتشرة
Diffused Seeing

شارك:
المجلة: Media theory.، المجلد: 8، العدد: 1
DOI: https://doi.org/10.70064/mt.v8i1.1075
تاريخ النشر: 2024-06-11
المؤلف: Joanna Zylinska
الموضوع الرئيسي: السيبرنتيك والتكنولوجيا في المجتمع

نظرة عامة

تستكشف هذه المقالة العلاقة المتطورة بين الإدراك والفهم لدى البشر والآلات، من خلال عدسة تقنيات الذكاء الاصطناعي التوليدي، مع التركيز على Stable Diffusion ومنافسيها، DALL•E 2 وMidjourney. تبدأ بتحليل البنية الفوتوغرافية التي تدعم هذه النماذج التوليدية وتغوص في مفهوم “الانتشار” كآلية أساسية في عملية توليد النص إلى صورة.

تثير المناقشة أسئلة حاسمة بشأن عدم الاستقرار والذوبان المحتمل للأطر المعرفية والسياسية الحالية في ضوء التقدم في الذكاء الاصطناعي التوليدي. ومعالجة المخاوف من النقاد الذين يشبهون هذه التطورات بالعدمية أو الفاشية، تستكشف المقالة إمكانية العثور على فرص تحررية ضمن المشهد الاجتماعي والتقني الحالي. ومن الجدير بالذكر أن الصور تُستخدم ليس فقط لأغراض توضيحية ولكن أيضًا لدعم حجج المقالة.

مناقشة

تستكشف قسم المناقشة في الورقة العلاقة المتطورة بين الإدراك والفهم في سياق الذكاء الاصطناعي التوليدي، مع التركيز بشكل خاص على دور الصور الفوتوغرافية في نماذج التعلم الآلي. يفترض المؤلف أن الإدراك الحديث، المتأثر بالتطور التاريخي للتصوير الفوتوغرافي، قد توسع من خلال مختلف التقدم التكنولوجي، مما أدى إلى استخدام الذكاء الاصطناعي التوليدي لإنشاء صور فوتوغرافية واقعية. توضح دراسة حالة Stable Diffusion من Stability AI كيف تعتمد هذه النماذج بشكل كبير على البيانات الفوتوغرافية، مما يثير أسئلة حاسمة حول التمثيل والآثار المعرفية لمثل هذه التقنيات. إن دمج عمليات الانتشار في هذه النماذج يعمل كآلية تقنية وإطار مجازي، مما يشير إلى تحول نحو فهم أكثر عدم استقرارًا و permeable للصور ومعانيها.

علاوة على ذلك، تتناول الورقة النقاشات الفلسفية المحيطة بمفهوم الفهم في الذكاء الاصطناعي، مع تسليط الضوء على الانقسام بين العلماء بشأن ما إذا كانت النماذج التوليدية يمكن أن “تفهم” اللغة أو الصور حقًا. بينما يتبنى بعض الباحثين وجهة نظر وظيفية، arguing that behavior resembling understanding qualifies as such, يعتقد آخرون أن هذه النماذج تحاكي الفهم فقط دون إدراك حقيقي. ينتقد المؤلف الافتراض السائد بأن الذكاء الاصطناعي التوليدي يجب أن ينتج واقعية، مقترحًا بدلاً من ذلك أن المخرجات “غير الصحيحة” أو الخيالية لهذه النماذج يمكن أن تثير حوارات جديدة حول طبيعة التمثيل والعلاقة بين الكلمات والصور. تدعو هذه النظرة إلى إعادة تقييم كيفية تفاعلنا مع التقنيات التوليدية، مع التأكيد على إمكاناتها كأدوات للاستكشاف بدلاً من كونها مجرد مكررات للمعايير البصرية الحالية.

Journal: Media theory., Volume: 8, Issue: 1
DOI: https://doi.org/10.70064/mt.v8i1.1075
Publication Date: 2024-06-11
Author(s): Joanna Zylinska
Primary Topic: Cybernetics and Technology in Society

Overview

This article investigates the evolving relationship between perception and comprehension in both humans and machines, specifically through the lens of generative AI technologies, with a focus on Stable Diffusion and its competitors, DALL•E 2 and Midjourney. It begins by analyzing the photographic infrastructure that supports these generative models and delves into the concept of ‘diffusion’ as a fundamental mechanism in the text-to-image generation process.

The discussion raises critical questions regarding the instability and potential dissolution of existing epistemological and political frameworks in light of advancements in generative AI. Addressing concerns from critics who liken these developments to nihilism or fascism, the article explores the possibility of finding emancipatory opportunities within the current socio-technical landscape. Notably, images are employed not merely for illustrative purposes but also to substantiate the article’s arguments.

Discussion

The discussion section of the paper explores the evolving relationship between perception and understanding in the context of generative AI, particularly focusing on the role of photographic imagery in machine learning models. The author posits that modern perception, influenced by the historical development of photography, has expanded through various technological advancements, culminating in the use of generative AI to create photorealistic images. The case study of Stability AI’s Stable Diffusion illustrates how these models rely heavily on photographic data, raising critical questions about representation and the epistemological implications of such technologies. The incorporation of diffusion processes in these models serves as both a technical mechanism and a metaphorical framework, suggesting a shift towards a more unstable and permeable understanding of images and their meanings.

Furthermore, the paper delves into the philosophical debates surrounding the notion of understanding in AI, highlighting a divide among scholars regarding whether generative models can truly “understand” language or images. While some researchers adopt a functionalist perspective, arguing that behavior resembling understanding qualifies as such, others contend that these models merely simulate understanding without genuine comprehension. The author critiques the prevailing assumption that generative AI should produce verisimilitude, suggesting instead that the “incorrect” or fantastical outputs of these models can provoke new dialogues about the nature of representation and the relationship between words and images. This perspective invites a reevaluation of how we engage with generative technologies, emphasizing their potential as tools for exploration rather than mere replicators of existing visual norms.

شارك: