التنبؤ بدرجات المقرر في برنامج التعليم العالي لعلوم الحاسوب من خلال مجموعة من دوال الخسارة في نموذج LSTM
Prediction of Course Grades in Computer Science Higher Education Program via a Combination of Loss Functions in LSTM Model

شارك:
المجلة: IEEE Access، المجلد: 12
DOI: https://doi.org/10.1109/access.2024.3351186
تاريخ النشر: 2024-01-01
المؤلف: Anahita Ghazvini وآخرون
الموضوع الرئيسي: طرق التنبؤ بسوق الأسهم

نظرة عامة

تستكشف هذه الدراسة تطبيق النمذجة التنبؤية في التعليم العالي، مع التركيز بشكل خاص على التنبؤ بالدرجات المبكرة لتحديد التحديات التعليمية وتعزيز التعلم الشخصي في برامج علوم الكمبيوتر. تسلط الأبحاث الضوء على فعالية الشبكات العصبية التكرارية (RNNs) في التقاط العلاقات المعقدة داخل بيانات تعلم الطلاب، مع معالجة مشكلة تلاشي التدرج التي غالبًا ما تعيق أداء RNN. للتخفيف من هذه المشكلة، يقترح المؤلفون دالة خسارة جديدة تُسمى MSECosine، والتي تجمع بين متوسط الخطأ التربيعي (MSE) وLogCosh، مما يؤدي إلى تحسين استقرار النموذج وأدائه. يُظهر النموذج المحسن، الذي يُطلق عليه eLSTM، دقة ملحوظة تبلغ 0.6191% ومعدل خطأ منخفض يبلغ 0.1738، متفوقًا على دالات الخسارة التقليدية وهياكل التعلم العميق الأخرى.

تؤكد النتائج على مزايا نماذج السلاسل الزمنية العميقة مقارنة بالطرق التقليدية في التنبؤ بالنتائج التعليمية. من خلال سلسلة من التجارب التي تستخدم 29 مجموعة ميزات مستمدة من البيانات الديموغرافية، وأنشطة التعلم، وبيانات التقييم، تؤسس الدراسة إطارًا شاملاً لفهم العوامل التي تؤثر على الأداء الأكاديمي. لا تعالج دالة خسارة MSECosine المقترحة مشكلة تلاشي التدرج فحسب، بل تعزز أيضًا الدقة التنبؤية العامة لنموذج eLSTM. تشمل اتجاهات البحث المستقبلية تحسين دالات الخسارة، واستكشاف تقنيات التنظيم، وتقييم قابلية تعميم دالة خسارة MSECosine عبر سياقات تعليمية متنوعة، مما يساهم في تقدم منهجيات التنبؤ بالدرجات المبكرة.

مقدمة

تؤكد مقدمة هذه الورقة البحثية على الحاجة الملحة للكشف المبكر عن الطلاب المعرضين لخطر الفشل الأكاديمي، مع تسليط الضوء على أهمية التنبؤ بالأداء الأكاديمي من خلال البيانات التسلسلية المستمدة من أنشطة الطلاب. تُلاحظ طرق التنبؤ بالسلاسل الزمنية التقليدية، مثل النماذج الذاتية الانحدار وتقنيات التنعيم الأسّي، لقيودها في التقاط الأنماط المعقدة وغير الخطية واعتمادها على مجموعات بيانات كبيرة للدقة. بالمقابل، تُقدم نماذج التعلم العميق، وخاصة الشبكات العصبية طويلة وقصيرة الأمد (LSTM)، كبدائل متفوقة قادرة على معالجة هذه التحديات من خلال نمذجة العلاقات المعقدة داخل البيانات بشكل فعال.

تقترح الدراسة دالة خسارة جديدة تُسمى MSECosine، والتي تجمع بين متوسط الخطأ التربيعي (MSE) وLogCosh للتخفيف من المشكلات المتعلقة بتضخيم الخطأ خلال مرحلة التدريب. تهدف هذه المقاربة المبتكرة إلى تحسين أداء نماذج السلاسل الزمنية العميقة في التنبؤ بدرجات الطلاب. يجادل المؤلفون بأنه بينما ركزت الأدبيات الحالية على العلاقات الزمنية وتقنيات البوابة في نماذج LSTM، لم يتم معالجة المشكلة الكبيرة لتضخيم الخطأ بشكل شامل. من خلال الاستفادة من مزايا دالة خسارة MSECosine المقترحة، تسعى الأبحاث إلى تحسين دقة التنبؤات المبكرة للدرجات، مما يساهم في المجال الأوسع لتحليل البيانات التعليمية.

الطرق

تم هيكلة منهجية هذه الدراسة في خمس مراحل متميزة: 1) جمع البيانات، 2) تنفيذ نماذج اختيار الميزات، 3) نمذجة نماذج السلاسل الزمنية العميقة التنبؤية، 4) مقاييس التقييم لأداء النموذج، و5) تحسين أفضل نموذج باستخدام دالة خسارة MSECosine المقترحة. تضمنت مرحلة جمع البيانات إنشاء مجموعة بيانات تم جمعها ذاتيًا، والتي شملت دمج مجموعات بيانات ‘LMS’ و’التقييم’ وتطبيق تقنيات المعالجة المسبقة.

في المرحلة الثانية، تم تنفيذ 29 نموذجًا لاختيار الميزات لتقييم أهميتها. ركزت المرحلة الثالثة على تطوير إطار عمل يستخدم أربعة نماذج سلاسل زمنية مختلفة: الشبكة العصبية متعددة الطبقات (MLP)، الشبكة العصبية التلافيفية (CNN)، الشبكة العصبية طويلة وقصيرة الأمد (LSTM)، ونموذج هجين CNN-LSTM. أوضحت المرحلة الرابعة مقاييس التقييم المستخدمة لقياس أداء كل نموذج بشكل كمي. أخيرًا، توضح المرحلة الخامسة سير العمل لنموذج LSTM المحسن، الذي يتضمن دالة خسارة MSECosine لتحسين الدقة التنبؤية.

النتائج

في هذا القسم، يقدم المؤلفون نتائج ثلاث تجارب تهدف إلى تقييم أساليب السلاسل الزمنية العميقة. تقيّم التجربة الأولى أداء 29 نموذجًا مصممًا لمجموعات الميزات، مما يوفر رؤى حول فعاليتها في التقاط الأنماط الزمنية. تركز التجربة الثانية على تحليل مقارن لطريقة Long Short-Term Memory المحسنة (eLSTM) مقابل الأساليب التقليدية للسلاسل الزمنية العميقة، مع تسليط الضوء على مزايا eLSTM من حيث الدقة التنبؤية والكفاءة. أخيرًا، يتم تحليل أداء النماذج المختلفة على مجموعات الميزات، مما يقدم نظرة شاملة حول كيفية أداء منهجيات مختلفة تحت الظروف المحددة. تساهم هذه النتائج في فهم فعالية النموذج في التنبؤ بالسلاسل الزمنية العميقة.

المناقشة

في هذا القسم، تناقش الدراسة عمليات جمع البيانات وإعدادها لمجموعتين رئيسيتين من البيانات المستخدمة في الدراسة: مجموعة بيانات ‘التقييم’ ومجموعة بيانات ‘LMS’. تتكون مجموعة بيانات ‘التقييم’ من 4,819 سجلًا من 40 دورة دراسية على مدى ثلاثة فصول دراسية، تحتوي على 24 سمة تتعلق بالبيانات الديموغرافية للطلاب والأداء. تشمل مجموعة بيانات ‘LMS’ 11,895 حالة مع 7 سمات توضح تفاعلات الطلاب مع نظام إدارة التعلم. بعد دمج هذه المجموعات بناءً على أرقام تسجيل الطلاب، تتكون مجموعة البيانات المدمجة من 3,721 حالة و21 سمة ذات صلة، تركز على أداء الطلاب ومشاركتهم.

تضمنت خطوات المعالجة المسبقة استخدام LabelEncoder لتحويل البيانات الفئوية إلى تنسيق عددي وتطبيع MinMaxScaler لضمان أن جميع الميزات على نفس المقياس، وهو أمر حاسم لأداء نماذج التعلم العميق. أنشأت الدراسة ميزات جديدة لتمثيل مشاركة الطلاب بناءً على تكرار وصولهم إلى LMS. بعد ذلك، تم تطوير أربعة نماذج للتعلم العميق (MLP، CNN، LSTM، وCNN-LSTM) للتنبؤ بدرجات الطلاب، مع اتباع نهج منهجي لتدريب النموذج وتقييمه. تم تقييم أداء هذه النماذج باستخدام مقاييس مثل الدقة، ومتوسط الخطأ التربيعي (MSE)، وجذر متوسط الخطأ التربيعي (RMSE)، مما أدى إلى اختيار نموذج LSTM كالأكثر فعالية للتنبؤ بالدرجات المبكرة. تهدف التحسينات المقترحة، بما في ذلك دالة خسارة MSECosine، إلى تحسين دقة التنبؤ من خلال معالجة المشكلات المتعلقة بحساسية الخطأ واستقرار التقارب.

Journal: IEEE Access, Volume: 12
DOI: https://doi.org/10.1109/access.2024.3351186
Publication Date: 2024-01-01
Author(s): Anahita Ghazvini et al.
Primary Topic: Stock Market Forecasting Methods

Overview

This study investigates the application of predictive modeling in higher education, particularly focusing on early grade prediction to identify learning challenges and enhance personalized learning in computer science programs. The research highlights the effectiveness of Recurrent Neural Networks (RNNs) in capturing complex relationships within student learning data, while addressing the vanishing gradient problem that often hampers RNN performance. To mitigate this issue, the authors propose a novel MSECosine loss function, which combines Mean Square Error (MSE) and LogCosh, resulting in improved model stability and performance. The enhanced model, termed eLSTM, demonstrates a significant accuracy of 0.6191% and a reduced error rate of 0.1738, outperforming traditional loss functions and other deep learning architectures.

The findings underscore the advantages of deep time series models over conventional methods for predicting educational outcomes. Through a series of experiments utilizing 29 feature sets derived from demographic, learning activity, and assessment data, the study establishes a comprehensive framework for understanding the factors influencing academic performance. The proposed MSECosine loss function not only addresses the vanishing gradient issue but also enhances the overall predictive accuracy of the eLSTM model. Future research directions include refining loss functions, exploring regularization techniques, and assessing the generalizability of the MSECosine loss function across diverse educational contexts, thereby contributing to the advancement of early grade prediction methodologies.

Introduction

The introduction of this research paper emphasizes the critical need for early detection of students at risk of academic failure, highlighting the importance of predicting academic performance through sequential data derived from students’ activities. Traditional time series forecasting methods, such as autoregressive and exponential smoothing approaches, are noted for their limitations in capturing complex, non-linear patterns and their reliance on large datasets for accuracy. In contrast, deep learning models, particularly Long Short-Term Memory (LSTM) networks, are presented as superior alternatives capable of addressing these challenges by effectively modeling intricate relationships within the data.

The study proposes a novel loss function, termed MSECosine, which combines Mean Square Error (MSE) and LogCosh to mitigate issues related to error magnification during the training phase. This innovative approach aims to enhance the performance of deep time series models in predicting student grades. The authors argue that while existing literature has focused on temporal relationships and gating techniques in LSTM models, the significant problem of error magnification has not been thoroughly addressed. By leveraging the advantages of the proposed MSECosine loss function, the research seeks to improve the accuracy of early grade predictions, thereby contributing to the broader field of educational data analysis.

Methods

The methodology of this research is structured into five distinct phases: 1) Data Collection, 2) Implementation of feature selection models, 3) Modeling of predictive deep time series models, 4) Evaluation metrics for model performance, and 5) Enhancement of the best model using the proposed MSECosine loss function. The data collection phase involved the creation of a self-collected dataset, which included merging ‘LMS’ and ‘Assessment’ datasets and applying preprocessing techniques.

In the second phase, 29 feature selection models were implemented to assess their importance. The third phase focused on developing a framework utilizing four different time series models: Multi-Layer Perceptron (MLP), Convolutional Neural Network (CNN), Long Short-Term Memory (LSTM), and a hybrid CNN-LSTM model. The fourth phase outlined the evaluation metrics employed to quantitatively measure the performance of each model. Finally, the fifth phase illustrated the workflow of the enhanced LSTM model, incorporating the MSECosine loss function to improve predictive accuracy.

Results

In this section, the authors present the results of three experiments aimed at evaluating deep time series approaches. The first experiment assesses the performance of 29 designed feature set models, providing insights into their effectiveness in capturing temporal patterns. The second experiment focuses on a comparative analysis of the enhanced Long Short-Term Memory (eLSTM) method against traditional deep time series approaches, highlighting the advantages of the eLSTM in terms of predictive accuracy and efficiency. Lastly, the performance of various models on the feature sets is analyzed, offering a comprehensive overview of how different methodologies perform under the specified conditions. These findings contribute to the understanding of model efficacy in deep time series forecasting.

Discussion

In this section, the research discusses the data collection and preparation processes for two primary datasets used in the study: the ‘Assessment’ dataset and the ‘LMS’ dataset. The ‘Assessment’ dataset consists of 4,819 records from 40 courses over three semesters, containing 24 attributes related to student demographics and performance. The ‘LMS’ dataset includes 11,895 instances with 7 attributes detailing student interactions with the learning management system. After integrating these datasets based on student matriculation numbers, the combined dataset comprises 3,721 instances and 21 relevant attributes, focusing on student performance and engagement.

The pre-processing steps involved using LabelEncoder to convert categorical data into numerical format and MinMaxScaler normalization to ensure all features are on a similar scale, which is crucial for the performance of deep learning models. The study generated new features to represent student engagement based on their LMS access frequency. Subsequently, four deep learning models (MLP, CNN, LSTM, and CNN-LSTM) were developed for predicting student grades, with a systematic approach to model training and evaluation. The performance of these models was assessed using metrics such as accuracy, mean squared error (MSE), and root mean squared error (RMSE), leading to the selection of the LSTM model as the most effective for early grade prediction. The proposed enhancements, including the MSECosine loss function, aim to improve prediction accuracy by addressing issues related to error sensitivity and convergence stability.

شارك: