DOI: https://doi.org/10.1344/der.2024.45.124-140
تاريخ النشر: 2024-07-01
المؤلف: Tuğra Karademir Coşkun وآخرون
الموضوع الرئيسي: التقييم وتقييم الأداء
نظرة عامة
تستكشف هذه الدراسة الفروقات بين تقييمات المعلمين وأنظمة تقييم الذكاء الاصطناعي (AI) في امتحانات الجامعات، والتي تشمل مجموعة متنوعة من أنواع الامتحانات، بما في ذلك التقييمات العددية، اللفظية، القائمة على المشاريع، التقليدية، النظرية، والعملية. باستخدام عينة معيارية، تستخدم البحث تحليل بلاند-ألتمن وتحليل معامل الارتباط داخل الفئة (ICC) لتقييم أداء تقييمات الذكاء الاصطناعي والمعلمين عبر هذه التنسيقات المتنوعة.
تكشف النتائج عن مستوى عالٍ من الاتفاق بشكل عام بين درجات الامتحانات الكلية من تقييمات الذكاء الاصطناعي وتقييمات المعلمين. ومع ذلك، تختلف الاتساق بشكل كبير عبر أنواع الامتحانات: لوحظ اتساق متوسط في الامتحانات المعتمدة على الرؤية، واتساق منخفض في امتحانات الفيديو، واتساق مرتفع في امتحانات الاختبار، واتساق منخفض في الامتحانات التقليدية. تعتبر هذه النتائج حاسمة لتحديد المجالات التي يمكن أن يعزز فيها الذكاء الاصطناعي التقييم التعليمي أو يحتاج إلى مزيد من التحسين، مما يساهم في تطوير أدوات تقييم أكثر دقة وإنصافًا.
مقدمة
تسلط مقدمة هذه الورقة البحثية الضوء على الدور التحويلي للذكاء الاصطناعي (AI) في التعليم، مع التأكيد على إمكانيته في تعزيز الكفاءة والإنتاجية وتجارب التعلم الشخصية. تمكن تقنيات الذكاء الاصطناعي المعلمين من تخصيص التعليم وفقًا لاحتياجات الطلاب الفردية، وأتمتة المهام، وتحسين استخدام البيانات، مما يحسن من اكتساب المعرفة والنتائج التعليمية. يُلاحظ أن دمج الذكاء الاصطناعي في منهجيات التدريس وعمليات التقييم يعزز من التنمية المعرفية والإبداع، ويزيد من استخدام الموارد، مما يؤدي في النهاية إلى بيئات تعلم أكثر جذبًا.
علاوة على ذلك، تناقش الورقة الدور الحاسم للذكاء الاصطناعي في عمليات التقييم، حيث يمكن أن يعزز دقة التقييم ويوفر تغذية راجعة في الوقت المناسب، مما يحسن من نتائج تعلم الطلاب. تُعترف أدوات الذكاء الاصطناعي بقدرتها على تحليل مجموعات بيانات كبيرة، وتحديد احتياجات الطلاب، ودعم التعليم الشخصي. ومع ذلك، تتناول المقدمة أيضًا التحديات مثل التحيزات الخوارزمية والحاجة إلى الشفافية في اتخاذ قرارات الذكاء الاصطناعي. تؤكد الورقة على أهمية دمج الخبرة البشرية مع قدرات الذكاء الاصطناعي لضمان تقييمات عادلة وموضوعية، خاصة في التقييمات ذات المخاطر العالية، مع مراعاة الآثار الأخلاقية وموثوقية النتائج التي ينتجها الذكاء الاصطناعي.
الطرق
تستخدم الدراسة مجموعة متنوعة من تنسيقات الامتحانات المجمعة من أقسام أكاديمية مختلفة لتحليل عمليات تقييم المعلمين وأدوات الذكاء الاصطناعي (AI) في تقييم الأعمال الطلابية متعددة التخصصات. تعتبر هذه المقاربة المنهجية ضرورية لفهم كيفية دمج الذكاء الاصطناعي ومعالجة المعلومات عبر تخصصات مختلفة، مما يقيم فعاليته في دعم وتعزيز التعليم متعدد التخصصات. تصنف الدراسة الامتحانات إلى عدة أنواع: امتحانات دورات عددية ولفظية، أسئلة مفتوحة، اختبارات متعددة الخيارات، تقييمات قائمة على المشاريع، امتحانات بتنسيق بصري، تقييمات قائمة على الفيديو، وامتحانات عملية مقابل نظرية. كل فئة تخدم غرضًا مميزًا، مما يسمح بمقارنة شاملة للمهارات المعرفية، والإبداع، واستراتيجيات حل المشكلات، وتقييم المعرفة متعددة التخصصات.
تقدم الجدول 1 توزيع هذه التنسيقات الامتحانية، كاشفًا أن الامتحانات التقليدية هي الأكثر انتشارًا، حيث تمثل 36.4% من البيانات التي تم تحليلها، تليها اختبارات متعددة الخيارات بنسبة 16.1%. تشمل امتحانات المشاريع بتنسيق الفيديو 18 طالبًا (15.3%)، بينما تشارك مشاريع بتنسيق الملصقات 38 طالبًا (32.2%). يبرز هذا التوزيع الطرق المتنوعة للتقييم المستخدمة ويبرز أهمية فحص كيفية تقييم كل من المعلمين والذكاء الاصطناعي لأداء الطلاب عبر سياقات وتخصصات مختلفة.
النتائج
في هذا القسم، تُعرض نتائج تقييم إجابات الطلاب وامتحانات المشاريع، مع تسليط الضوء على استخدام أدوات الذكاء الاصطناعي للتقييم. تم مقارنة استجابات الطلاب مع الإجابات الصحيحة المحدثة، مما يكشف عن الحاجة لإعادة حساب الدرجات بناءً على البيانات الأخيرة. استخدم تقييم امتحانات المشاريع، وخاصة الملصقات، الإضافة “بحث الويب” من GPT-4، التي تدمج معالجة اللغة الطبيعية مع عمليات البحث عبر الإنترنت في الوقت الحقيقي للتحقق من دقة المعلومات المقدمة من قبل الطلاب. تضمن هذه الطريقة أن الادعاءات المقدمة في الملصقات مدعومة بمصادر موثوقة، مما يعزز مصداقية التقييمات.
تركز النتائج على الاتساق بين الدرجات المخصصة من قبل أدوات الذكاء الاصطناعي وتلك التي منحها المعلمون عبر تنسيقات امتحانية متنوعة، بما في ذلك الاختبارات التقليدية والتقييمات القائمة على المشاريع. تم إجراء تحليل بلاند-ألتمن لتقييم الفروقات بين هذه طرق التقييم، مع تفاصيل النتائج في الجدول 3 ومبين في الشكل 1. يهدف هذا التحليل إلى تقديم رؤى حول موثوقية تقييمات الذكاء الاصطناعي مقارنة بممارسات التقييم التقليدية في البيئات الأكاديمية.
المناقشة
ت outlines قسم المناقشة في الورقة البحثية هدف الدراسة لتقييم الاتساق بين الدرجات المخصصة من قبل المعلمين وتلك التي تولدها أدوات الذكاء الاصطناعي (AI) عبر تنسيقات امتحانية متنوعة، بما في ذلك الاختبارات التقليدية، المشاريع، الفيديوهات، والملصقات. تهدف الدراسة إلى ربط طرق التقييم التعليمية التقليدية مع تقنيات الذكاء الاصطناعي المبتكرة، وتقييم موثوقية ودقة الذكاء الاصطناعي في البيئات التعليمية. من خلال التحقيق في الفروقات والاتساقات في التقييمات، تسعى الدراسة إلى تقديم رؤى حول كيفية تعزيز الذكاء الاصطناعي لعمليات التقييم التقليدية، مما قد يؤدي إلى تقييمات أكثر موضوعية وإنصافًا تحسن من نتائج التعلم.
تستخدم الدراسة تصميمًا مقطعيًا لتحليل البيانات المجمعة من مجموعات عينة مختلفة ضمن إطار زمني محدد، مما يسمح بفحص متزامن لعدة متغيرات، مثل درجات التقييم من كل من المعلمين وأدوات الذكاء الاصطناعي. تستخدم الدراسة أدوات ذكاء اصطناعي متقدمة، تحديدًا ChatGPT-4 من OpenAI، لتقييم أنواع امتحانات مختلفة، مستفيدة من قدراتها في معالجة اللغة الطبيعية لتقييم استجابات الطلاب وفقًا لمعايير محددة. تُستخدم طرق إحصائية، بما في ذلك تحليل بلاند-ألتمن ومعامل الارتباط داخل الفئة (ICC)، لتقييم اتساق الدرجات بين المقيمين البشريين والذكاء الاصطناعي، مما يوفر تقييمًا شاملاً لفعالية الذكاء الاصطناعي في التقييمات التعليمية. تعتبر هذه الدراسة مهمة في تقدم الممارسات التعليمية من خلال استكشاف إمكانيات الذكاء الاصطناعي لتكملة طرق التقييم التقليدية، مما يعزز تجارب التعلم الشخصية والفعالة لطلاب الجامعات.
القيود
تستكشف الدراسة التفاعل بين الذكاء الاصطناعي (AI) وتقييمات المعلمين في تقييمات أوراق الامتحانات، كاشفة عن كل من الفوائد المحتملة والقيود الجوهرية لدمج الذكاء الاصطناعي في السياقات التعليمية. بينما تؤكد النتائج على فعالية أنظمة الذكاء الاصطناعي عند استخدامها بالتزامن مع المعلمين البشريين، فإنها تسلط الضوء أيضًا على ضرورة المزيد من الاستكشاف لتعزيز قدرات هذه الأنظمة.
تعتبر إحدى القيود المهمة الملحوظة هي الاعتماد على برنامج ذكاء اصطناعي واحد، مما قد يقيد تنوع نتائج التقييم وصلاحية النتائج بشكل عام. لمعالجة ذلك، يجب أن تتضمن الأبحاث المستقبلية أدوات ذكاء اصطناعي متعددة لتوفير فهم أوسع وأكثر دقة لدور الذكاء الاصطناعي في التقييمات التعليمية. بالإضافة إلى ذلك، قد يؤدي اتساق مطالبات التقييم إلى تحريف أداء الذكاء الاصطناعي، مما يشير إلى أن الدراسات المستقبلية يجب أن تختلف في المطالبات لتقييم كفاءة تقييم الذكاء الاصطناعي عبر أنواع الأسئلة المختلفة.
DOI: https://doi.org/10.1344/der.2024.45.124-140
Publication Date: 2024-07-01
Author(s): Tuğra Karademir Coşkun et al.
Primary Topic: Evaluation and Performance Assessment
Overview
This study investigates the differences between teacher evaluations and artificial intelligence (AI) assessment systems in university examinations, encompassing a variety of exam types, including numerical, verbal, project-based, traditional, theoretical, and practical assessments. Utilizing criterion sampling, the research employs Bland-Altman Analysis and Intraclass Correlation Coefficient (ICC) analyses to evaluate the performance of AI and teacher assessments across these diverse formats.
The findings reveal a generally high level of agreement between total exam scores from AI and teacher evaluations. However, the consistency varies significantly across exam types: medium consistency is noted in visually based exams, low consistency in video exams, high consistency in test exams, and low consistency in traditional exams. These results are pivotal for identifying areas where AI can enhance educational assessment or requires further refinement, ultimately contributing to the development of more accurate and equitable evaluation tools.
Introduction
The introduction of this research paper highlights the transformative role of Artificial Intelligence (AI) in education, emphasizing its potential to enhance efficiency, productivity, and personalized learning experiences. AI technologies enable educators to tailor instruction to individual student needs, automate tasks, and optimize data usage, thereby improving knowledge acquisition and educational outcomes. The integration of AI into teaching methodologies and assessment processes is noted to foster cognitive development, creativity, and resource maximization, ultimately leading to more engaging learning environments.
Furthermore, the paper discusses the critical role of AI in evaluation processes, where it can enhance assessment accuracy and provide timely feedback, thereby improving student learning outcomes. AI tools are recognized for their ability to analyze large datasets, identify student needs, and support personalized instruction. However, the introduction also addresses challenges such as algorithmic biases and the need for transparency in AI decision-making. The paper underscores the importance of combining human expertise with AI capabilities to ensure fair and objective evaluations, particularly in high-stakes assessments, while also considering ethical implications and the reliability of AI-generated results.
Methods
The research employs a diverse array of exam formats collected from various academic departments to analyze the evaluation processes of instructors and artificial intelligence (AI) tools in assessing multidisciplinary student work. This methodological approach is essential for understanding how AI integrates and processes information across different disciplines, thereby evaluating its effectiveness in supporting and promoting multidisciplinary education. The study categorizes exams into several types: numerical and verbal course exams, open-ended questions, multiple-choice tests, project-based assessments, visual format exams, video-based evaluations, and practical versus theoretical exams. Each category serves a distinct purpose, allowing for a comprehensive comparison of cognitive skills, creativity, problem-solving strategies, and interdisciplinary knowledge assessment.
Table 1 presents the distribution of these exam formats, revealing that classical exams are the most prevalent, accounting for 36.4% of the analyzed data, followed by multiple-choice tests at 16.1%. Project-based exams in video format involve 18 students (15.3%), while poster format projects engage 38 students (32.2%). This distribution underscores the varied assessment methods employed and highlights the significance of examining how both instructors and AI evaluate student performance across different contexts and disciplines.
Results
In this section, the evaluation results of student answers and project exams are presented, highlighting the use of AI tools for assessment. The student responses were compared against updated correct answers, revealing the need for recalculation of scores based on the latest data. The evaluation of project exams, particularly posters, utilized the “Web Search” plugin of GPT-4, which integrates natural language processing with real-time internet searches to verify the accuracy of information presented by students. This method ensures that the claims made in the posters are corroborated by reliable sources, enhancing the credibility of the evaluations.
The findings focus on the consistency between scores assigned by AI tools and those given by human instructors across various exam formats, including traditional tests and project-based assessments. A Bland-Altman analysis was conducted to assess the discrepancies between these scoring methods, with results detailed in Table 3 and illustrated in Figure 1. This analysis aims to provide insights into the reliability of AI evaluations compared to traditional grading practices in academic settings.
Discussion
The discussion section of the research paper outlines the study’s objective to evaluate the consistency between scores assigned by instructors and those generated by artificial intelligence (AI) tools across various exam formats, including classical tests, projects, videos, and posters. The research aims to bridge traditional educational assessment methods with innovative AI technologies, assessing the reliability and accuracy of AI in educational settings. By investigating discrepancies and consistencies in evaluations, the study seeks to provide insights into how AI can enhance traditional assessment processes, potentially leading to more objective and equitable evaluations that improve learning outcomes.
The research employs a cross-sectional design to analyze data collected from different sample groups within a specific timeframe, allowing for the simultaneous examination of multiple variables, such as assessment scores from both instructors and AI tools. The study utilizes advanced AI tools, specifically OpenAI’s ChatGPT-4, for grading various exam types, leveraging its natural language processing capabilities to evaluate student responses against established criteria. Statistical methods, including Bland-Altman analysis and Interclass Correlation Coefficient (ICC), are employed to assess the consistency of scores between human evaluators and AI, providing a comprehensive evaluation of the effectiveness of AI in educational assessments. This research is significant in advancing educational practices by exploring the potential of AI to complement traditional grading methods, thereby fostering personalized and effective learning experiences for university students.
Limitations
The study investigates the interplay between artificial intelligence (AI) and instructor evaluations in exam paper assessments, revealing both the potential benefits and inherent limitations of AI integration in educational contexts. While the findings underscore the effectiveness of AI systems when used in conjunction with human instructors, they also highlight the necessity for further exploration to enhance these systems’ capabilities.
One significant limitation noted is the reliance on a single AI software, which may restrict the diversity of assessment outcomes and the overall validity of the results. To address this, future research should incorporate multiple AI tools to provide a broader and more nuanced understanding of AI’s role in educational assessments. Additionally, the consistency of evaluation prompts may skew AI performance, suggesting that future studies should vary prompts to better assess AI’s evaluation proficiency across different question types.
