وضع الأسس لاستخدام صور الذكاء الاصطناعي التوليدي في البحث المعماري: هل تنقل الصور المساحات والأجواء المقصودة؟
Laying the foundations for using generative AI images in architectural research: do images convey the intended spaces and ambiances?

شارك:
المجلة: Architectural Intelligence، المجلد: 3، العدد: 1
DOI: https://doi.org/10.1007/s44223-024-00076-x
تاريخ النشر: 2024-10-28
المؤلف: Azuka Odiah وآخرون
الموضوع الرئيسي: دراسات العمارة وتاريخ الفن

نظرة عامة

تبحث ورقة البحث في قدرات نماذج الذكاء الاصطناعي التوليدي (GenAI)، وبشكل خاص DALL-E2، في التصميم المعماري، مع التركيز على كفاءتها في توليد الصور التي تمثل بدقة المساحات والأجواء المقصودة. من خلال مقارنة الصور المولدة لثمانية عشر بيئة منزلية متميزة، تم تصميم كل منها لاستحضار ثلاث أجواء مختلفة، مع الصور الفوتوغرافية الحقيقية المستمدة من صور جوجل، تقيم الدراسة الواقعية المدركة ودقة هذه الصور المولدة بواسطة الذكاء الاصطناعي. تشير النتائج إلى أنه بينما يتفوق DALL-E2 في تصوير الدقة المكانية، فإن نقل الأجواء المحددة يمثل تحديًا أكبر.

في الختام، تكشف الدراسة أن فعالية الصورة في تمثيل مساحة أو جو معين تعتمد إلى حد كبير على واقعيتها. تسلط هذه الرؤية الضوء على الحاجة إلى تقييم مستمر لأدوات الذكاء الاصطناعي في العمارة، مما يبرز نقاط قوتها والمجالات التي تتطلب مزيدًا من التطوير. تسهم هذه الأبحاث في فهم دور الذكاء الاصطناعي في الممارسة المعمارية، داعية إلى تحقيق توازن بين التقدم التكنولوجي والتصميم الموجه نحو الإنسان. كما تؤكد على أهمية التقييمات الدقيقة للتحقق من فعالية نماذج GenAI، مما يضمن دمجها في العمليات الإبداعية التي تعطي الأولوية لكل من الوظائف والارتباط العاطفي.

مقدمة

تسلط المقدمة الضوء على التأثير التحويلي لنماذج الذكاء الاصطناعي التوليدي (GenAI) على مختلف الصناعات، وخاصة العمارة. تركز على تقييم قدرات نموذج DALL-E2 في نقل المساحات المعمارية والأجواء، مقارنة أدائه بنموذج بديل. يتم تعريف GenAI بقدرته على توليد محتوى يحاكي بشكل وثيق الإبداع البشري عبر تنسيقات متعددة، بما في ذلك النصوص والصور، من خلال تقنيات التعلم الآلي المتقدمة، وخاصة التعلم العميق. تستخدم هذه النماذج، مثل المحولات المدربة مسبقًا التوليدية (GPT)، مجموعات بيانات واسعة لإنتاج مخرجات تشبه البشر، مما يحدث ثورة في عملية التصميم المعماري من الفكرة إلى التصور والتسليم.

الطرق

استخدمت الدراسة منهجية من مرحلتين للتحقيق في الأجواء المرغوبة في المساحات السكنية. في المرحلة الأولى، بعنوان “مصادر الصور”، تم الحصول على الصور من خلال الذكاء الاصطناعي التوليدي (GenAI) وعمليات بحث صور جوجل (GI). المرحلة الثانية، “تقييم الصور”، تضمنت مراقبين بشريين يقيمون هذه الصور. تم إبلاغ اختيار المساحات والأجواء من خلال أبحاث سابقة طورت جرد الأجواء المرغوبة في المنزل المثالي (IDAIH)، الذي حدد الأجواء الأكثر طلبًا عبر 18 مساحة سكنية (Graham et al., 2015). يتم تقديم ملخص لهذه المساحات، مع الأجواء الثلاثة الأكثر طلبًا لكل منها، في الجدول 1، الذي يوضح أيضًا عدد الصور المختارة لكل مجموعة من المساحة-الجو من كل من GenAI وGI، مع كون GI بمثابة معيار للمقارنة.

شارك المشاركون في الدراسة من خلال استبيان عبر الإنترنت مستضاف على Qualtrics. بعد الحصول على موافقة مستنيرة، أكملوا تقييمًا موجزًا لأبعاد الشخصية الخمسة الكبرى (جرد الشخصية المكون من عشرة عناصر [TIPI]، Gosling et al., 2003)، على الرغم من أن هذه البيانات لم يتم تحليلها في الدراسة الحالية. بعد ذلك، قام المشاركون بتقييم الصور، وتقديم معلومات ديموغرافية، وكان لديهم خيار تقديم ملاحظات حول مرحلة تقييم الصور. حصلت الدراسة على موافقة من مجلس مراجعة المؤسسات بجامعة تكساس في أوستن واتبعت جميع الإرشادات الأخلاقية ذات الصلة. هذه الأبحاث هي جزء من بروتوكول تجريبي أوسع يهدف إلى تحليل الروابط المنهجية بين العناصر الفيزيائية للمساحة وأجوائها المقابلة.

النتائج

في قسم النتائج، تؤسس الدراسة معيارًا من خلال استخدام الصور المستمدة من GI لتقييم أداء الصور المولدة بواسطة DALL-E2. تهدف هذه التحليل المقارن إلى تقييم جودة وفعالية الصور المولدة بواسطة DALL-E2 مقابل المعيار المحدد الذي توفره صور GI. تعتبر النتائج من هذا التقييم حاسمة لفهم قدرات وقيود DALL-E2 في مهام توليد الصور.

المناقشة

تسلط قسم المناقشة في ورقة البحث الضوء على الإمكانيات التحويلية للذكاء الاصطناعي التوليدي (GenAI) في التصميم المعماري والبحث، مع التأكيد على قدرته على تعزيز عمليات التصور والتصور. تتيح أدوات GenAI، مثل DALL-E2 وMidJourney، للمعماريين استكشاف تنوعات التصميم بسرعة، مما يسهل عملية إبداعية أكثر ديناميكية وتوسعًا. على الرغم من هذه التقدمات، تكشف الدراسة أنه بينما يمكن لـ GenAI توليد صور تصور المساحات المقصودة بدقة مكانية تبلغ 76%، فإن فعاليتها في نقل الأجواء المحددة محدودة، حيث تحقق فقط 25% دقة. يشير هذا إلى وجود فجوة كبيرة في قدرة GenAI على فهم وتمثيل الجوانب العاطفية والوظيفية الدقيقة للتصميم المعماري بشكل كامل.

تؤكد النتائج أيضًا على التباين في الدقة عبر أنواع مختلفة من المساحات، حيث يتم التعرف على المناطق المشتركة مثل المطابخ وغرف الطعام بدقة أكبر مقارنة بالمساحات الأقل تكرارًا مثل غرف الوسائط وغرف الجلوس. قد تنشأ هذه الفجوة من الميزات المتداخلة لهذه المساحات، مما يعقد تمييزها. تدعو الدراسة إلى مزيد من الأبحاث لتأسيس أطر تقييم قوية لمخرجات GenAI، بهدف سد الفجوة بين القدرات التقنية للذكاء الاصطناعي والمتطلبات العملية للتصميم المعماري الموجه نحو الإنسان. بشكل عام، بينما يحمل GenAI وعدًا بالابتكار في العمارة، فإن قيوده الحالية تتطلب تحسينًا مستمرًا لتعزيز قابليته في نقل الأبعاد المكانية والعاطفية بشكل فعال.

القيود

تقدم الدراسة عدة قيود تؤثر على قابلية تعميم وموثوقية نتائجها. تعتبر إحدى القضايا الرئيسية هي التطور السريع لتقنيات الذكاء الاصطناعي، التي يمكن أن تتجاوز دورة البحث، مما يجعل النتائج قد تكون غير صالحة عند نشرها. على سبيل المثال، بينما بدأت الأبحاث باستخدام DALL-E2، فإن إصدار DALL-E3 خلال مرحلة تقديم الدراسة يبرز هذه التحديات. يدعو المؤلفون إلى إنشاء بنى تحتية بحثية يمكن أن تتكيف مع هذه التقدمات السريعة، وقد جعلوا موادهم ورمزهم متاحة للجمهور لتسهيل المقارنات المستقبلية في الواقعية والدقة.

بالإضافة إلى ذلك، واجهت الدراسة صعوبات في الحصول على صور جوجل (GI) التي تتوافق مع معايير الأجواء المحددة، مما أدى إلى تفاوت في عدد الصور بين الذكاء الاصطناعي التوليدي (GenAI) وGI (540 مقابل 242). على الرغم من أن تحليلًا لاحقًا أشار إلى أن الفروق في الدقة لم تكن ناتجة فقط عن هذه الفجوة، إلا أن المؤلفين يوصون بتحسين تقنيات استرجاع الصور وتطوير قواعد بيانات مخصصة لتعزيز التمثيل. علاوة على ذلك، فإن اعتماد الدراسة على الحكم البشري لاختيار الصور يقدم تباينًا، حاول المؤلفون التخفيف منه من خلال قياسات اتفاقية القضاة. يجب أن تستكشف الأبحاث المستقبلية تصاميم بديلة، مثل الأساليب داخل الموضوع، للتحكم بشكل أفضل في الفروق الفردية بين القضاة. أخيرًا، قد يحد تركيز الدراسة على السياقات المعمارية الموجهة نحو الولايات المتحدة من قابليتها للتطبيق عبر تفسيرات ثقافية متنوعة للمساحة والأجواء، مما يشير إلى الحاجة إلى اعتبارات ثقافية أوسع في التحقيقات المستقبلية.

Journal: Architectural Intelligence, Volume: 3, Issue: 1
DOI: https://doi.org/10.1007/s44223-024-00076-x
Publication Date: 2024-10-28
Author(s): Azuka Odiah et al.
Primary Topic: Architecture and Art History Studies

Overview

The research paper investigates the capabilities of Generative Artificial Intelligence (GenAI) models, specifically DALL-E2, in architectural design, focusing on their proficiency in generating images that accurately represent intended spaces and ambiances. By comparing generated images of eighteen distinct home environments, each designed to evoke three different ambiances, against real photographs sourced from Google Images, the study evaluates the perceived realism and accuracy of these AI-generated images. The findings indicate that while DALL-E2 excels in depicting spatial accuracy, conveying specific ambiances poses a greater challenge.

In conclusion, the study reveals that the effectiveness of an image in representing a particular space or ambiance is largely independent of its realism. This insight underscores the need for ongoing evaluation of AI tools in architecture, highlighting both their strengths and areas requiring further development. The research contributes to the understanding of AI’s role in architectural practice, advocating for a balance between technological advancements and human-centric design. It emphasizes the importance of rigorous assessments to validate the effectiveness of GenAI models, ensuring their integration into creative processes that prioritize both functionality and emotional resonance.

Introduction

The introduction highlights the transformative impact of Generative Artificial Intelligence (GenAI) models on various industries, particularly architecture. It focuses on evaluating the capabilities of the DALL-E2 model in conveying architectural spaces and ambiances, comparing its performance to an alternative model. GenAI is defined by its ability to generate content that closely mimics human creation across multiple formats, including text and images, through advanced machine learning techniques, especially deep learning. These models, such as Generative Pre-trained Transformers (GPT), utilize extensive datasets to produce human-like outputs, thereby revolutionizing the architectural design process from conception to visualization and delivery.

Methods

The study employed a two-phase methodology to investigate the desired ambiances in residential spaces. In Phase 1, titled “Image Sourcing,” images were obtained through Generative AI (GenAI) and Google Image (GI) searches. Phase 2, “Judging Images,” involved human observers evaluating these images. The selection of spaces and ambiances was informed by prior research that developed the Inventory of Desired Ambiances in the Ideal Home (IDAIH), which identified the most sought-after ambiances across 18 residential spaces (Graham et al., 2015). A summary of these spaces, along with the three most desired ambiances for each, is presented in Table 1, which also details the number of images selected for each space-ambiance combination from both GenAI and GI, with GI serving as a benchmark for comparison.

Participants engaged in the study via an online survey hosted on Qualtrics. Following informed consent, they completed a brief assessment of the Big Five personality dimensions (Ten Item Personality Inventory [TIPI], Gosling et al., 2003), although this data is not analyzed in the current study. Subsequently, participants judged the images, provided demographic information, and had the option to give feedback on the image judgment phase. The study received approval from the University of Texas at Austin Institutional Review Board and adhered to all relevant ethical guidelines. This research is part of a broader experimental protocol aimed at analyzing systematic connections between physical elements of space and their corresponding ambiances.

Results

In the Results section, the study establishes a benchmark by utilizing images sourced from GI to evaluate the performance of images generated by DALL-E2. This comparative analysis aims to assess the quality and effectiveness of the DALL-E2 generated images against the established standard provided by GI images. The findings from this evaluation are critical for understanding the capabilities and limitations of DALL-E2 in image generation tasks.

Discussion

The discussion section of the research paper highlights the transformative potential of generative AI (GenAI) in architectural design and research, emphasizing its ability to enhance visualization and conceptualization processes. GenAI tools, such as DALL-E2 and MidJourney, allow architects to rapidly explore diverse design variations, facilitating a more dynamic and expansive creative process. Despite these advancements, the study reveals that while GenAI can generate images that depict intended spaces with a space accuracy of 76%, its effectiveness in conveying specific ambiances is limited, achieving only 25% accuracy. This indicates a significant gap in the ability of GenAI to fully understand and represent the nuanced emotional and functional aspects of architectural design.

The findings also underscore the variability in accuracy across different types of spaces, with common areas like kitchens and dining rooms being more accurately identified compared to less frequently encountered spaces such as media rooms and sitting rooms. This discrepancy may arise from the overlapping features of these spaces, which complicate their differentiation. The study calls for further research to establish robust evaluation frameworks for GenAI outputs, aiming to bridge the gap between AI’s technical capabilities and the practical requirements of human-centric architectural design. Overall, while GenAI holds promise for innovation in architecture, its current limitations necessitate ongoing refinement to enhance its applicability in conveying both spatial and emotional dimensions effectively.

Limitations

The study presents several limitations that impact the generalizability and robustness of its findings. A primary concern is the rapid evolution of AI technologies, which can outpace the research cycle, rendering findings potentially obsolete by the time of publication. For instance, while the research began with DALL-E2, the release of DALL-E3 during the study’s submission phase highlights this challenge. The authors advocate for the establishment of research infrastructures that can adapt to such rapid advancements, and they have made their materials and code publicly available to facilitate future comparisons in realism and accuracy.

Additionally, the study faced difficulties in sourcing Google Images (GI) that matched specific ambiance criteria, leading to a disparity in the number of images between Generative AI (GenAI) and GI (540 vs. 242). Although a post-hoc analysis indicated that the accuracy differences were not solely due to this discrepancy, the authors recommend improving image retrieval techniques and developing dedicated databases to enhance representativeness. Furthermore, the study’s reliance on human judgment for image selection introduces variability, which the authors attempted to mitigate through interjudge agreement measures. Future research should explore alternative designs, such as within-subject approaches, to better control for individual differences among judges. Lastly, the study’s focus on US-centric architectural contexts may limit its applicability across diverse cultural interpretations of space and ambiance, suggesting a need for broader cultural considerations in future investigations.

شارك: