مقارنة منهجية لنماذج اللغة الكبيرة لتقييم الواجبات الآلي في تعليم البرمجة: استكشاف أهمية الهيكل والبائع
A systematic comparison of Large Language Models for automated assignment assessment in programming education: Exploring the importance of architecture and vendor

شارك:
المجلة: Computers and Education Open، المجلد: 10
DOI: https://doi.org/10.1016/j.caeo.2026.100364
تاريخ النشر: 2026-04-20
المؤلف: Marcin Jukiewicz
الموضوع الرئيسي: تعليم وتعلم البرمجة

نظرة عامة

تقدم هذه الورقة البحثية تقييمًا منهجيًا لـ 18 نموذجًا متقدمًا من نماذج اللغة الكبيرة (LLMs) لتقييم البرمجة بشكل آلي، كاشفة عن اختلافات كبيرة في فلسفات التقييم بين النماذج. تصنف الدراسة النماذج إلى ثلاثة أنماط متميزة: ليبرالية (مثل، GPT-4o، Claude-haiku-3.5)، والتي تميل إلى منح درجات أعلى؛ مقيدة (مثل، DeepSeek-Reasoner، GPT-4.1-nano)، التي تفضل الدرجات المنخفضة؛ ونماذج متوازنة تستخدم تقييمات متوسطة. كما تحدد التحليل أيضًا تجمعات محددة حسب البائع، مما يشير إلى أن النماذج من نفس المزود (OpenAI، Google، Anthropic، DeepSeek) تظهر أنماط تقييم مشابهة، مما يقترح وجود أساليب خوارزمية مشتركة.

تسلط النتائج الضوء على اتفاق محدود بين الدرجات التي تولدها النماذج وتقييمات المعلمين البشريين، حيث حقق النموذج الأكثر أداءً، Claude-haiku-3.5، موثوقية معتدلة فقط. ومن الجدير بالذكر أن المعلمين منحوا باستمرار درجات أعلى من جميع نماذج الذكاء الاصطناعي، مما يشير إلى اختلافات أساسية في أساليب التقييم. تؤكد الدراسة أيضًا على تأثير حجم النموذج على جودة التقييم، حيث تتفوق النسخ الكاملة على نظيراتها “الصغيرة” و”النانوية”. تؤكد هذه النتائج على أهمية توافق اختيار النموذج مع الأهداف التربوية والحفاظ على إشراف بشري في التقييمات الآلية لضمان الدقة والعدالة في السياقات التعليمية.

مقدمة

تناقش مقدمة هذه الورقة البحثية الإمكانات التحويلية للذكاء الاصطناعي التوليدي (GenAI)، وخاصة نماذج اللغة الكبيرة (LLMs)، في مجال التعليم. منذ إطلاق ChatGPT في أواخر عام 2022، كان هناك زيادة في الأدبيات التي تسلط الضوء على فوائد نماذج اللغة الكبيرة، مثل تجارب التعلم الشخصية وأتمتة المهام الإدارية للمعلمين. ومن الجدير بالذكر أن أتمتة تقييم الطلاب تم تحديدها كتطبيق واعد، حيث تشير الدراسات إلى أن نماذج اللغة الكبيرة يمكن أن تحقق دقة تقييم مقارنة بالمقيمين البشريين تحت ظروف معينة، وخاصة للمهام المنظمة. ومع ذلك، تؤكد الورقة أيضًا على التحديات الكبيرة، بما في ذلك المخاوف بشأن النزاهة الأكاديمية، وتطوير التفكير النقدي، وتباين أداء التقييم عبر نماذج ومهام مختلفة.

يحدد المؤلفون فرضيات رئيسية بشأن سلوكيات التقييم لمختلف نماذج اللغة الكبيرة، مقترحين أن هناك اختلافات في توزيع الدرجات، ومتوسط الدرجات، وأنماط التقييم بين النماذج. يقترحون أن النماذج ذات الهياكل المماثلة قد تظهر ميول تقييم مرتبطة وأن الاتفاق بين النماذج في المهام المحددة يجب أن يتجاوز مستويات الصدفة. تهدف الدراسة إلى تحليل منهجي لهذه السلوكيات التقييمية، مقدمة رؤى تجريبية حول فعالية وموثوقية نماذج اللغة الكبيرة في التقييم التعليمي. هذه الأبحاث ذات صلة خاصة حيث تتنقل المؤسسات التعليمية في دمج تقنيات الذكاء الاصطناعي بينما تسعى للحفاظ على الجودة الأكاديمية والعدالة، مما يبرز الحاجة إلى تقييم مستمر للنماذج الناشئة في ضوء التقدم السريع في هذا المجال.

الطرق

توضح قسم الطرق في الدراسة إطار جمع البيانات والتحليل المستخدم على مدى أربع سنوات في دورة برمجة تمهيدية لطلاب العلوم المعرفية. تم تحليل ما مجموعه حوالي 6,081 زوجًا كاملًا من الأسئلة والأجوبة، مستمدة من حوالي 6,500 سجل مهمة تم إنشاؤها من تسعة اختبارات برمجة تُجرى سنويًا. تم توضيح هيكل الدورة والمواضيع المغطاة في منشور سابق (Jukiewicz، 2024)، مع تقديم ملخص في الجدول 1.

تقوم الدراسة بمقارنة أداء 18 نموذجًا من نماذج اللغة الكبيرة (LLMs) من أربعة بائعين رائدين، مع التركيز على أحدث إصداراتها المتاحة للجمهور اعتبارًا من أغسطس 2025. تشمل هذه المقارنة كل من النسخ الخفيفة (مثل، الصغيرة، النانوية، الخفيفة) والنماذج الرائدة الكاملة (مثل، opus، pro، gpt-5)، مما يضمن تمثيلًا شاملاً لهياكل نماذج اللغة الكبيرة المعاصرة. تشمل النماذج المحددة التي تم تقييمها إصدارات مختلفة من Anthropic وDeepseek وGoogle، كما هو مدرج في النص.

النتائج

في هذا القسم، يقدم المؤلفون النتائج من تحقيقهم في سلوكيات تقييم نماذج اللغة الكبيرة (LLMs). يوضحون الأساليب التحليلية المستخدمة لتقييم توزيع الدرجات الممنوحة من قبل النماذج، مع تسليط الضوء على قياس ميول التقييم والاتساق. علاوة على ذلك، تحدد الدراسة أنماط تقييم متميزة وتجمعات من النماذج التي تظهر أنماط تقييم مشابهة، مما يوفر نظرة شاملة على أداء النماذج.

تمهد النتائج الطريق لمقارنة دقيقة بين التقييمات المعتمدة على النماذج والتقييمات البشرية، مما يشير إلى أن فهم هذه السلوكيات التقييمية أمر حاسم لتقييم فعالية وموثوقية نماذج اللغة الكبيرة في السياقات التعليمية. يعمل الفصل كأساس لاستكشاف المزيد من الآثار المترتبة على هذه النتائج بشأن أنظمة التقييم الآلي.

المناقشة

تسلط قسم المناقشة في الورقة البحثية الضوء على اختلافات كبيرة في سلوكيات التقييم بين نماذج اللغة المختلفة، مؤكدة الفرضيات المتعلقة بتباين توزيع الدرجات ومتوسط الدرجات. ومن الجدير بالذكر أن نماذج مثل GPT-4o وClaude-Haiku-3.5 أظهرت ميلاً لمنح درجات أعلى، بينما كانت نماذج أخرى مثل DeepSeek-Reasoner وGPT-4.1-Nano أكثر تقييدًا، حيث تصدر درجات أقل بشكل متكرر. تشير هذه الاستقطابات إلى اختلافات جوهرية في عتبات النماذج للاعتراف بصحة الحلول. بالإضافة إلى ذلك، أظهرت نماذج مثل Claude-Sonnet-4 وGemini-2.0-Flash-Lite نهجًا متوازنًا، حيث تمنح غالبًا درجات متوسطة، مما قد يكون مفيدًا في البيئات التعليمية التي تقدر الصحة الجزئية.

كشفت تحليلات الارتباط والتجمع أن النماذج من نفس البائع تميل إلى مشاركة فلسفات تقييم مشابهة، مع أعلى الارتباطات الملحوظة ضمن عائلات GPT-5 وGemini. حددت الدراسة ستة تجمعات متميزة بناءً على أنماط التقييم، تتراوح من تلك التي تفضل الدرجات القصوى إلى تلك التي تتبنى استراتيجية تقييم أكثر حذرًا. تؤكد هذه النتائج على أهمية فهم ميول التقييم للنماذج المختلفة، حيث يمكن أن تؤثر بشكل كبير على نتائج الطلاب في أنظمة التقييم الآلي. تتماشى النتائج مع الأدبيات الحالية، التي تشير أيضًا إلى اختلافات منهجية في ميول التقييم بين النماذج المختلفة، مما يعزز فكرة أن اختيار النموذج يجب أن يأخذ في الاعتبار هذه الأنماط السلوكية لتحسين التقييمات التعليمية.

Journal: Computers and Education Open, Volume: 10
DOI: https://doi.org/10.1016/j.caeo.2026.100364
Publication Date: 2026-04-20
Author(s): Marcin Jukiewicz
Primary Topic: Teaching and Learning Programming

Overview

This research paper presents a systematic evaluation of 18 state-of-the-art Large Language Models (LLMs) for automated grading in programming education, revealing significant differences in grading philosophies among the models. The study categorizes the models into three distinct styles: liberal (e.g., GPT-4o, Claude-haiku-3.5), which tend to assign higher scores; restrictive (e.g., DeepSeek-Reasoner, GPT-4.1-nano), which favor lower grades; and balanced models that utilize intermediate scoring. The analysis also identifies vendor-specific clustering, indicating that models from the same provider (OpenAI, Google, Anthropic, DeepSeek) exhibit similar grading patterns, suggesting shared algorithmic approaches.

The findings highlight a limited agreement between model-generated grades and human teacher assessments, with even the highest-performing model, Claude-haiku-3.5, achieving only moderate reliability. Notably, teachers consistently assigned higher grades than all AI models, pointing to fundamental differences in evaluation approaches. The study further emphasizes the impact of model size on grading quality, with full-scale versions outperforming their “mini” and “nano” counterparts. These results underscore the importance of aligning model selection with pedagogical goals and maintaining human oversight in automated assessments to ensure accuracy and fairness in educational contexts.

Introduction

The introduction of this research paper discusses the transformative potential of generative artificial intelligence (GenAI), particularly large language models (LLMs), in the field of education. Since the launch of ChatGPT in late 2022, there has been a surge in literature highlighting the benefits of LLMs, such as personalized learning experiences and the automation of administrative tasks for educators. Notably, the automation of student assessment is identified as a promising application, with studies indicating that LLMs can achieve grading accuracy comparable to human evaluators under specific conditions, particularly for structured tasks. However, the paper also emphasizes significant challenges, including concerns about academic integrity, critical thinking development, and the variability in grading performance across different models and tasks.

The authors outline key hypotheses regarding the grading behaviors of various LLMs, suggesting that differences in grade distribution, mean scores, and grading patterns exist among models. They propose that models with similar architectures may exhibit correlated grading tendencies and that agreement between models on well-defined tasks should exceed chance levels. The study aims to systematically analyze these grading behaviors, providing empirical insights into the effectiveness and reliability of LLMs in educational assessment. This research is particularly relevant as educational institutions navigate the integration of AI technologies while striving to maintain academic quality and equity, underscoring the need for ongoing evaluation of emerging models in light of rapid advancements in the field.

Methods

The Methods section of the study details the data collection and analysis framework employed over four years in an introductory programming course for cognitive science students. A total of approximately 6,081 complete question-answer pairs were analyzed, derived from around 6,500 task records generated from nine programming tests administered annually. The course structure and topics covered are elaborated in a prior publication (Jukiewicz, 2024), with a summary provided in Table 1.

The study further compares the performance of 18 Large Language Models (LLMs) from four leading vendors, focusing on their most recent publicly available versions as of August 2025. This comparison includes both lightweight variants (e.g., mini, nano, lite) and full-scale flagship models (e.g., opus, pro, gpt-5), ensuring a comprehensive representation of contemporary LLM architectures. The specific models evaluated include various releases from Anthropic, Deepseek, and Google, as listed in the text.

Results

In this section, the authors present the findings from their investigation into the grading behavior of Large Language Models (LLMs). They detail the analytical methods employed to assess the distribution of scores assigned by the models, highlighting the quantification of grading tendencies and consistency. Furthermore, the study identifies distinct evaluation styles and clusters of models that exhibit similar grading patterns, providing a comprehensive overview of the models’ performance.

The results pave the way for a nuanced comparison between model-based assessments and human evaluations, suggesting that understanding these grading behaviors is crucial for evaluating the efficacy and reliability of LLMs in educational contexts. The chapter serves as a foundation for further exploration of the implications of these findings on automated grading systems.

Discussion

The discussion section of the research paper highlights significant differences in grading behaviors among various language models, confirming hypotheses regarding the variability in grade distributions and mean scores. Notably, models such as GPT-4o and Claude-Haiku-3.5 exhibited a tendency to assign higher grades, while others like DeepSeek-Reasoner and GPT-4.1-Nano were more restrictive, frequently issuing lower scores. This polarization indicates inherent differences in the models’ thresholds for recognizing solution correctness. Additionally, models like Claude-Sonnet-4 and Gemini-2.0-Flash-Lite demonstrated a balanced approach, often awarding intermediate scores, which may be beneficial in educational settings that value partial correctness.

Correlation and clustering analyses revealed that models from the same vendor tend to share similar grading philosophies, with the highest correlations observed within the GPT-5 and Gemini families. The study identified six distinct clusters based on grading styles, ranging from those favoring maximum scores to those employing a more cautious evaluation strategy. These findings underscore the importance of understanding the grading tendencies of different models, as they can significantly impact student outcomes in automated grading systems. The results align with existing literature, which also notes systematic differences in grading tendencies among various models, reinforcing the notion that model selection should consider these behavioral patterns to enhance educational assessments.

شارك: