التقييم التوأمي: نهج لتعزيز صلاحية التقييم في عصر الذكاء الاصطناعي التوليدي
Assessment twins: An approach for strengthening assessment validity in the age of generative AI

شارك:
المجلة: Journal of Applied Learning & Teaching، المجلد: 9، العدد: 2
DOI: https://doi.org/10.37074/jalt.2026.9.2.3
تاريخ النشر: 2026-06-03
المؤلف: Jasper Roe وآخرون
الموضوع الرئيسي: تقييم الطلاب والتغذية الراجعة

نظرة عامة

تناقش الورقة التأثير التحويلي للذكاء الاصطناعي التوليدي (GenAI) على تقييمات التعليم العالي، مع تسليط الضوء على الحاجة الملحة لإعادة التصميم للحفاظ على نزاهة وموثوقية التقييم. تقدم مفهوم “توأم التقييم”، الذي يتكون من مكونين تقييميين مترابطين يستهدفان نفس نتائج التعلم من خلال أوضاع مختلفة من الأدلة. يتم تأطير هذا النهج ضمن إطار موحد للموثوقية لميسيك، مع معالجة التهديدات المختلفة للموثوقية التي تطرحها GenAI، بما في ذلك المحتوى، الهيكل، العواقب، القابلية للتعميم، والموثوقية الخارجية. من خلال استخدام تنسيقات تقييم مكملة—مثل دمج المقالات مع الدفاعات الشفوية أو العروض العملية—يهدف استراتيجية التوأم إلى تعزيز الموثوقية من خلال مثلث الأدلة.

يقترح المؤلفون عملية تصميم من ثلاث خطوات لتنفيذ توائم التقييم: تحديد نقاط الضعف، مواءمة النتائج، واختيار المهام المكملة مع مخططات تقييم مترابطة. بينما يعترفون بالتحديات مثل متطلبات الموارد ومخاوف العدالة، تجادل الورقة بأن توائم التقييم يمكن أن تدعم بفعالية نتائج التعلم المعنوية وتتكيف مع أحجام المجموعات المتنوعة. يدعو المؤلفون إلى مزيد من البحث والتجريب لتحسين هذا النموذج، موضحين أنه بديل تقدمي لممارسات التقييم المدفوعة بالمراقبة. في النهاية، يتماشى مفهوم توائم التقييم مع الدعوات الأوسع لإعادة تصور مشهد التقييم الذي يعطي الأولوية للقيمة التربوية ويعالج عدم المساواة النظامية في سياق تقدمات GenAI.

مقدمة

تتناول مقدمة الورقة التأثير الكبير للذكاء الاصطناعي التوليدي (GenAI) على ممارسات التقييم في التعليم العالي، خاصة في ضوء المخاوف المتعلقة بالنزاهة الأكاديمية. لقد أثار ظهور نماذج GenAI المتقدمة قضايا حول موثوقية تنسيقات التقييم التقليدية، مثل المقالات المنزلية والعروض التقديمية، التي يمكن الآن إكمالها بسهولة بواسطة هذه التقنيات. لقد دفع هذا التحول إلى إعادة تقييم هدف التقييم ودور التعليم العالي، مع استمرار النقاشات حول ما إذا كان GenAI سيعزز أو يقوض الممارسات التعليمية في النهاية. يعبر النقاد عن مخاوف بشأن الانخفاض المحتمل في قدرة المتعلمين على اتخاذ القرار ومهارات التفكير النقدي بسبب الاعتماد على أدوات GenAI.

استجابةً لهذه التحديات، يقترح المؤلفون نهجًا جديدًا يسمى “توأم التقييم”، والذي يتضمن دمج المهام المعرضة لـ GenAI مع مهام مكملة، أقل تعرضًا، تقيم نفس نتائج التعلم. يهدف هذا الاستراتيجية إلى تعزيز موثوقية التقييم من خلال توفير بيانات تأكيدية والحفاظ على القيمة التربوية للتنسيقات التقليدية. علاوة على ذلك، يتماشى مع الدعوات إلى المزيد من طرق التقييم التعاونية ومتعددة الأشكال في عصر GenAI. توضح الورقة هيكل المناقشة، بدءًا من مراجعة الأدبيات حول GenAI وموثوقية التقييم، تليها استكشاف مفصل لمفهوم توأم التقييم، وإرشادات التنفيذ العملي، ونظر في قيود الإطار.

مناقشة

تتناول قسم المناقشة في الورقة التحديات التي تطرحها GenAI في التقييمات الأكاديمية، خاصة فيما يتعلق بموثوقية عمل الطلاب. لقد جعل ظهور GenAI من الصعب اكتشاف الانتحال، مما أدى إلى مخاوف بشأن ما إذا كانت التقييمات تعكس بدقة تعلم الطلاب. لقد أثبتت طرق الكشف التقليدية عن الذكاء الاصطناعي عدم كفاءتها، مما أدى إلى دعوة لاستراتيجيات بديلة تعزز الثقة بين الطلاب والمؤسسات. تشمل الحلول المقترحة تضمين محو الأمية في الذكاء الاصطناعي في المناهج الدراسية، ودمج مهام التأمل الذاتي، وتطوير أطر مثل مقياس قابلية التقييم-GenAI PANDORA.

يقدم المؤلفون مفهوم “توأم التقييم”، الذي يتضمن دمج تقييمات تستهدف نفس نتائج التعلم ولكن تستخدم أوضاع مختلفة من الأدلة لتعزيز الموثوقية. يستند هذا النهج إلى نموذج موحد لميسيك للموثوقية، الذي يبرز أهمية موثوقية البناء. من خلال تنفيذ توائم التقييم، يمكن للمعلمين تقليل المخاطر المرتبطة بـ GenAI، مما يضمن أن تظل التقييمات ذات مغزى وموثوقة. يوضح المؤلفون عملية عملية من ثلاث خطوات لتطوير توائم التقييم، مع التأكيد على الحاجة إلى اعتبار دقيق لنتائج التعلم والترابط بين المهام المزدوجة. يعترفون بالطبيعة المستهلكة للموارد لهذا النهج، خاصة في المجموعات الكبيرة، ويبرزون أهمية التصميم العادل لتجنب تفاقم عدم المساواة التعليمية القائمة. يتم تشجيع البحث المستقبلي للتحقق تجريبيًا من فعالية توائم التقييم عبر سياقات تعليمية متنوعة.

القيود

تسلط قسم القيود الضوء على غياب البيانات التجريبية التي تدعم فعالية الإطار المقترح الذي يتضمن توائم التقييم. بينما يفترض الإطار أن الفجوات بين مكونات التوأم يمكن أن تشير إلى التعلم، قد تنشأ هذه التناقضات أيضًا من عوامل مشروعة مثل القلق، وتطور المهارات غير المتساوي، أو مستويات الراحة المتفاوتة مع تنسيقات التقييم المختلفة، كما أشار سترويفن وآخرون (2005). علاوة على ذلك، يؤكد شو وكريسب (2012) أن تنسيقات التقييم المختلفة قد تؤدي إلى نتائج تعلم متميزة بدلاً من متكافئة.

يعترف المؤلفون بالدعوة التي قدمها إيلشال وبادير (2025) لنهج هجينة تدمج طرق التقييم التقليدية مع المشاريع المدعومة بالذكاء الاصطناعي، ومع ذلك، يلاحظون أن مثل هذه النماذج لا تزال غير مفحوصة إلى حد كبير. على الرغم من هذه القيود، يؤكد المؤلفون أن توائم التقييم تحمل قيمة محتملة، مما يحفز النقاشات النقدية حول الموثوقية، والعدالة، والثقة في سياق تقدم أدوات الذكاء الاصطناعي التوليدي. يدعون إلى رؤية توائم التقييم ليس كحل نهائي ولكن كنهج تجريبي مهم في التقييم التعليمي.

Journal: Journal of Applied Learning & Teaching, Volume: 9, Issue: 2
DOI: https://doi.org/10.37074/jalt.2026.9.2.3
Publication Date: 2026-06-03
Author(s): Jasper Roe et al.
Primary Topic: Student Assessment and Feedback

Overview

The paper discusses the transformative impact of Generative Artificial Intelligence (GenAI) on higher education assessments, highlighting the urgent need for redesign to maintain assessment integrity and validity. It introduces the concept of “assessment twins,” which consists of two interconnected assessment components targeting the same learning outcomes through distinct modes of evidence. This approach is framed within Messick’s unified validity framework, addressing various validity threats posed by GenAI, including content, structural, consequential, generalizability, and external validity. By employing complementary assessment formats—such as pairing essays with oral defenses or practical demonstrations—the twin strategy aims to enhance validity through triangulation of evidence.

The authors propose a three-step design process for implementing assessment twins: identifying vulnerabilities, aligning outcomes, and selecting complementary tasks with interdependent marking schemes. While acknowledging challenges such as resource demands and equity concerns, the paper argues that assessment twins can effectively support meaningful learning outcomes and adapt to diverse cohort sizes. The authors advocate for further research and experimentation to refine this model, positioning it as a progressive alternative to surveillance-driven assessment practices. Ultimately, the assessment twins concept aligns with broader calls for a reimagined assessment landscape that prioritizes pedagogical value and addresses systemic inequalities in the context of GenAI advancements.

Introduction

The introduction of the paper addresses the significant impact of Generative Artificial Intelligence (GenAI) on assessment practices in higher education, particularly in light of concerns regarding academic integrity. The emergence of advanced GenAI models has raised issues about the reliability of traditional assessment formats, such as take-home essays and presentations, which can now be easily completed by these technologies. This shift has prompted a reevaluation of the purpose of assessment and the role of higher education, with ongoing debates about whether GenAI will ultimately enhance or undermine educational practices. Critics express concerns about potential reductions in learner agency and critical thinking skills due to reliance on GenAI tools.

In response to these challenges, the authors propose a novel approach termed “assessment twins,” which involves pairing GenAI-vulnerable tasks with complementary, less vulnerable tasks that assess the same learning outcomes. This strategy aims to enhance assessment validity by providing confirmatory data and preserving the pedagogical value of traditional formats. Furthermore, it aligns with calls for more collaborative and multimodal assessment methods in the GenAI era. The paper outlines the structure of the discussion, beginning with a literature review on GenAI and assessment validity, followed by a detailed exploration of the assessment twin concept, practical implementation guidelines, and a consideration of the framework’s limitations.

Discussion

The discussion section of the paper addresses the challenges posed by Generative AI (GenAI) in academic assessments, particularly concerning the validity of student work. The emergence of GenAI has complicated the detection of plagiarism, leading to concerns about whether assessments accurately reflect student learning. Traditional AI detection methods have proven inadequate, prompting a call for alternative strategies that foster trust between students and institutions. Proposed solutions include embedding AI literacy in curricula, incorporating self-reflection tasks, and developing frameworks like the Assessment-GenAI Susceptibility Rubric PANDORA.

The authors introduce the concept of “assessment twins,” which involves pairing assessments that target the same learning outcomes but utilize different modes of evidence to enhance validity. This approach is grounded in Messick’s unified model of validity, which emphasizes the importance of construct validity. By implementing assessment twins, educators can mitigate the risks associated with GenAI, ensuring that assessments remain meaningful and reliable. The authors outline a practical three-step process for developing assessment twins, emphasizing the need for careful consideration of learning outcomes and the interdependence of paired tasks. They acknowledge the resource-intensive nature of this approach, particularly in larger cohorts, and highlight the importance of equitable design to avoid exacerbating existing educational inequalities. Future research is encouraged to empirically validate the effectiveness of assessment twins across diverse educational contexts.

Limitations

The section on limitations highlights the absence of empirical data supporting the effectiveness of the proposed framework involving assessment twins. While the framework posits that discrepancies between twin components can indicate learning, these inconsistencies may also stem from legitimate factors such as anxiety, uneven skill development, or varying comfort levels with different assessment formats, as noted by Struyven et al. (2005). Furthermore, Shaw and Crisp (2012) emphasize that different assessment formats may yield distinct learning outcomes rather than equivalent ones.

The authors acknowledge the call by Elshall and Badir (2025) for hybrid approaches that integrate traditional assessment methods with AI-assisted projects, yet they note that such models remain largely unexamined. Despite these limitations, the authors assert that assessment twins hold potential value, prompting critical discussions about validity, fairness, and trust in the context of advancing generative AI tools. They advocate for viewing assessment twins not as a definitive solution but as an important experimental approach in educational assessment.

شارك: