ما وراء النص إلى نص: نظرة عامة على الذكاء الاصطناعي متعدد الوسائط والتوليدي للتعليم باستخدام نمذجة الموضوعات
Beyond Text-to-Text: An Overview of Multimodal and Generative Artificial Intelligence for Education Using Topic Modeling

شارك:
المجلة: Proceedings of the 40th ACM/SIGAPP Symposium on Applied Computing
DOI: https://doi.org/10.1145/3672608.3707764
تاريخ النشر: 2025-03-31
المؤلف: Ville Heilala وآخرون
الموضوع الرئيسي: طرق التحليل الحاسوبي والنصي

نظرة عامة

يقدم قسم ورقة البحث لمحة عامة عن دور الذكاء الاصطناعي التوليدي (GenAI) في التعليم، مع تسليط الضوء على هيمنة نماذج اللغة الكبيرة (LLMs) مثل ChatGPT في الدراسات الحالية، مع الإشارة إلى الإمكانيات غير المستكشفة للقدرات متعددة الوسائط، مثل تقنيات تحويل النص إلى كلام وتحويل النص إلى صورة. من خلال بحث أدبي شامل حدد 4,175 مقالة، استخدمت الدراسة نمذجة الموضوعات لاستخراج 38 موضوعًا قابلًا للتفسير، والتي تم تنظيمها في 14 مجالًا موضوعيًا.

تكشف النتائج عن تركيز كبير على نماذج النص إلى نص ضمن السياقات التعليمية، مما يشير إلى فجوة بحثية بشأن الوسائط الأخرى. تشير هذه الفجوة إلى أن الإمكانات الأوسع للنهج متعددة الوسائط في التعليم لا يتم تحقيقها بالكامل. تؤكد الدراسة على ضرورة إجراء المزيد من الأبحاث الشاملة التي تشمل مختلف وسائط الذكاء الاصطناعي والمستويات التعليمية، مما يبرز الفرص لاستكشاف تقنيات متعددة الوسائط لتعزيز التأثير التحويلي للذكاء الاصطناعي في البيئات التعليمية.

مقدمة

تناقش مقدمة هذه الورقة البحثية الإمكانات التحويلية للذكاء الاصطناعي (AI) في التعليم، وخاصة من خلال تقنيات الذكاء الاصطناعي التوليدي. تسلط الضوء على التطور المستمر للنماذج التعليمية، من التعليم 2.0 إلى التعليم 6.0، وتؤكد على الحاجة إلى دمج مسؤول للذكاء الاصطناعي لإعادة تشكيل الممارسات التعليمية الأساسية. تؤكد إرشادات اليونسكو بشأن الذكاء الاصطناعي التوليدي على التغييرات الكبيرة التي يمكن أن تجلبها هذه التقنيات إلى الأنظمة التعليمية الراسخة، مما يقدم فرصًا وتحديات تتطلب أطرًا تنظيمية، مثل قانون الذكاء الاصطناعي في الاتحاد الأوروبي.

تحدد الورقة فجوة في مشهد البحث بشأن النهج متعددة الوسائط والذكاء الاصطناعي التوليدي في التعليم، مشيرة إلى أنه بينما تظهر حلول بيداغوجية مدفوعة بالذكاء الاصطناعي، فإن التوزيع غير المتساوي للبحث وتأثير قادة السوق يعقدان الوضع. يتم الاعتراف بالروابط التاريخية بين الذكاء الاصطناعي والتعليم، جنبًا إلى جنب مع التقدم في القدرات متعددة الوسائط التي تمكنت من خلال التعلم العميق. يدعو المؤلفون إلى إجراء تحقيق شامل في تطبيقات الذكاء الاصطناعي التوليدي في التعليم، مع معالجة التحديات مثل القضايا الأخلاقية والحاجة إلى أدوات ذكاء اصطناعي أكثر تعقيدًا وتعدد التخصصات. تهدف الدراسة إلى دمج الأدبيات الموجودة باستخدام نمذجة الموضوعات لاستكشاف مشهد البحث على مستوى عالٍ لهذه التقنيات في السياقات التعليمية.

الطرق

في هذا القسم، يوضح المؤلفون منهجيتهم لاختيار الأدبيات ذات الصلة حول الذكاء الاصطناعي التوليدي في التعليم. استخدموا استراتيجية بحث شاملة باستخدام سلسلة البحث التالية: (التعليم أو الطالب أو المدرسة) AND (‘الذكاء الاصطناعي التوليدي’ أو ‘الذكاء الاصطناعي التوليدي’ أو genai أو LLM أو ‘نموذج اللغة الكبير’ أو ‘تحويل النص إلى نص’ أو ‘تحويل النص إلى فيديو’ أو ‘تحويل النص إلى صوت’) مع دمج وسائط مختلفة مثل النص والفيديو والصوت والمزيد. من المهم أن البحث استبعد أسماء خدمات محددة لتقليل التحيز. أسفر البحث الأولي عن 4,475 نتيجة، والتي تم تصفيتها لاحقًا إلى 4,175 مقالة بعد استبعاد تلك التي تفتقر إلى معرف كائن الوثيقة (DOI) أو الملخص.

يتكون الجسم الناتج بشكل أساسي من منشورات حديثة، حيث تم نشر 95% من المقالات بعد عام 2014. يعود أقدم مقال ذي صلة إلى عام 1983، مما يبرز نهجًا حسابيًا لدعم تحويل النص إلى كلام يستهدف الطلاب ذوي الإعاقة البصرية. تؤكد هذه الأعمال الأساسية على الاهتمام المستمر في الاستفادة من التكنولوجيا لتحسين النتائج التعليمية. استخدم المؤلفون مجموعة البيانات المنسقة هذه لتحليلات نمذجة الموضوعات اللاحقة، كما هو موضح في الشكل 1.

النتائج

يقدم قسم “النتائج” من ورقة البحث النتائج الرئيسية المستمدة من التجارب أو التحليلات التي تم إجراؤها. يوضح النتائج التي توصلت إليها الدراسة، مع تسليط الضوء على نقاط البيانات والاتجاهات المهمة التي لوحظت. غالبًا ما تكون النتائج مصحوبة بتحليلات إحصائية ذات صلة، والتي قد تشمل قيم p، وفترات الثقة، أو معاملات الارتباط، لدعم النتائج.

بالإضافة إلى ذلك، قد يتضمن القسم تمثيلات بصرية مثل الرسوم البيانية أو الجداول التي توضح البيانات بوضوح، مما يسمح بتفسير أسهل للنتائج. عادةً ما يتم مناقشة تداعيات هذه النتائج فيما يتعلق بأسئلة البحث المطروحة في بداية الدراسة، مما يوفر أساسًا لمزيد من المناقشة في الأقسام اللاحقة. بشكل عام، يخدم هذا القسم لنقل الأدلة التجريبية التي تدعم فرضيات أو أهداف الدراسة.

المناقشة

في هذه الدراسة، استخدم المؤلفون نهج BERTopic لإجراء نمذجة الموضوعات على مجموعة من الملخصات المتعلقة بالذكاء الاصطناعي التوليدي في التعليم، مما أدى في النهاية إلى تحديد 38 موضوع بحثي متميز. تضمنت المنهجية إنشاء تضمينات قائمة على المحولات، وتقليل الأبعاد باستخدام UMAP، والتجميع باستخدام HDBSCAN، تليها بحث عن المعلمات لتحسين المعلمات الفائقة بناءً على مؤشر التحقق من التجميع القائم على الكثافة (DBCV). أسفر النموذج النهائي عن 54 موضوعًا، حيث تم تخصيص 85% من الملخصات لهذه الموضوعات، بينما تم ترك 15% غير مخصصة لتجنب إدخال أخطاء. تم دمج الموضوعات في 14 منطقة تركيز موضوعية، مما يكشف أن OpenAI ChatGPT هو الموضوع الأكثر انتشارًا في البحث التعليمي، مما يبرز اهتمامًا كبيرًا بنماذج اللغة الكبيرة (LLMs).

تشير النتائج إلى الحاجة إلى استكشاف أوسع لتقنيات الذكاء الاصطناعي متعددة الوسائط بخلاف LLMs، مثل تحويل النص إلى كلام وتحويل النص إلى صورة، والتي أظهرت إمكانات في تعزيز التعلم الشخصي والإبداع. تؤكد الدراسة على أهمية معالجة القضايا المتعلقة بنزاهة الأكاديمية المرتبطة بالذكاء الاصطناعي التوليدي، فضلاً عن ضرورة انخراط المعلمين مع مجموعة متنوعة من أدوات الذكاء الاصطناعي لإثراء الممارسات البيداغوجية. تشمل قيود البحث نظرة عامة عامة على الجسم، والفجوات المحتملة في التقاط جميع التقنيات ذات الصلة، ونقص التركيز على وجهات نظر المعلمين. يجب أن تهدف الأبحاث المستقبلية إلى فهم أكثر دقة لدور الذكاء الاصطناعي التوليدي في التعليم، خاصة عبر مستويات تعليمية ومواضيع مختلفة.

Journal: Proceedings of the 40th ACM/SIGAPP Symposium on Applied Computing
DOI: https://doi.org/10.1145/3672608.3707764
Publication Date: 2025-03-31
Author(s): Ville Heilala et al.
Primary Topic: Computational and Text Analysis Methods

Overview

The research paper section provides an overview of the role of generative artificial intelligence (GenAI) in education, highlighting the dominance of large language models (LLMs) like ChatGPT in current studies while noting the underexplored potential of multimodal capabilities, such as text-to-speech and text-to-image technologies. Through an extensive literature search that identified 4,175 articles, the study employed topic modeling to extract 38 interpretable topics, which were organized into 14 thematic areas.

The findings reveal a significant emphasis on text-to-text models within educational contexts, indicating a research gap regarding other modalities. This imbalance suggests that the broader potential of multimodal approaches in education is not being fully realized. The study underscores the necessity for more comprehensive research that encompasses various AI modalities and educational levels, thereby highlighting opportunities for future exploration of multimodal technologies to enhance the transformative impact of artificial intelligence in educational settings.

Introduction

The introduction of this research paper discusses the transformative potential of artificial intelligence (AI) in education, particularly through generative AI technologies. It highlights the ongoing evolution of educational paradigms, from Education 2.0 to Education 6.0, and emphasizes the need for responsible integration of AI to reshape foundational educational practices. The UNESCO Guidance on generative AI underscores the significant changes these technologies could bring to established educational systems, presenting both opportunities and challenges that necessitate regulatory frameworks, such as the EU’s Artificial Intelligence Act.

The paper identifies a gap in the research landscape regarding multimodal approaches and generative AI in education, noting that while various AI-driven pedagogical solutions are emerging, the uneven distribution of research and the influence of market leaders complicate the situation. Historical ties between AI and education are acknowledged, alongside the advancements in multimodal capabilities enabled by deep learning. The authors call for a comprehensive investigation into the applications of generative AI in education, addressing challenges like ethical concerns and the need for more sophisticated, interdisciplinary AI tools. The study aims to synthesize existing literature using topic modeling to explore the high-level research landscape of these technologies in educational contexts.

Methods

In this section, the authors detail their methodology for selecting relevant literature on generative artificial intelligence in education. They employed a comprehensive search strategy using the following search string: (education OR student OR school) AND (‘generative ai’ OR ‘generative artificial intelligence’ OR genai OR LLM OR ‘large language model’ OR ‘text-to-text’ OR ‘text-to-video’ OR ‘text-to-audio’) combined with various modalities such as text, video, audio, and more. Importantly, the search excluded specific service names to minimize bias. The initial search yielded 4,475 results, which were subsequently refined to 4,175 articles after filtering out those lacking a Document Object Identifier (DOI) or abstract.

The resulting corpus primarily consists of recent publications, with 95% of the articles published after 2014. The earliest relevant article dates back to 1983, highlighting a computational approach for text-to-speech support aimed at visually impaired students. This foundational work underscores the long-standing interest in leveraging technology to enhance educational outcomes. The authors utilized this curated dataset for subsequent topic modeling analyses, as illustrated in Figure 1.

Results

The “Results” section of the research paper presents the key findings derived from the conducted experiments or analyses. It details the outcomes of the study, highlighting significant data points and trends observed. The results are often accompanied by relevant statistical analyses, which may include p-values, confidence intervals, or correlation coefficients, to substantiate the findings.

Additionally, the section may include visual representations such as graphs or tables that illustrate the data clearly, allowing for easier interpretation of the results. The implications of these findings are typically discussed in relation to the research questions posed at the outset of the study, providing a foundation for further discussion in subsequent sections. Overall, this section serves to convey the empirical evidence that supports the study’s hypotheses or objectives.

Discussion

In this study, the authors employed the BERTopic approach to conduct topic modeling on a corpus of abstracts related to generative AI in education, ultimately identifying 38 distinct research topics. The methodology involved creating transformer-based embeddings, dimensionality reduction using UMAP, and clustering with HDBSCAN, followed by a parameter search to optimize hyperparameters based on the Density-based Clustering Validation (DBCV) index. The final model yielded 54 topics, with 85% of the abstracts assigned to these topics, while 15% were left unassigned to avoid introducing errors. The topics were synthesized into 14 thematic focus areas, revealing that OpenAI ChatGPT is the most prevalent subject in educational research, highlighting a significant interest in large language models (LLMs).

The findings indicate a need for broader exploration of multimodal AI technologies beyond LLMs, such as text-to-speech and text-to-image, which have shown potential in enhancing personalized learning and creativity. The study emphasizes the importance of addressing academic integrity concerns associated with generative AI, as well as the necessity for educators to engage with a diverse range of AI tools to enrich pedagogical practices. Limitations of the research include a general overview of the corpus, potential gaps in capturing all relevant technologies, and a lack of focus on educators’ perspectives. Future research should aim for a more nuanced understanding of generative AI’s role in education, particularly across different educational levels and themes.

شارك: