DOI: https://doi.org/10.1038/s41598-024-63629-0
PMID: https://pubmed.ncbi.nlm.nih.gov/38839872
تاريخ النشر: 2024-06-05
المؤلف: Maria Psyridou وآخرون
الموضوع الرئيسي: الذكاء الاصطناعي في الرعاية الصحية
نظرة عامة
تسلط هذه القسم الضوء على الدور الحاسم للتعليم في مكافحة الفقر وتعزيز النمو الاقتصادي، بينما تعالج أيضًا التحدي الكبير الذي تطرحه معدلات التسرب من المدرسة. ركزت الأبحاث السابقة في توقع التسرب بشكل أساسي على البيانات قصيرة الأجل، مما حد من فعاليتها. على النقيض من ذلك، تستخدم هذه الدراسة مجموعة بيانات طولية شاملة تمتد على 13 عامًا، تتبع الطلاب من رياض الأطفال حتى الصف التاسع. تتضمن الأبحاث مجموعة واسعة من المعايير، بما في ذلك الأداء الأكاديمي، المهارات المعرفية، الدافع، السلوك، والرفاهية، إلى جانب سجلات التسرب.
تظهر نماذج التعلم الآلي التي تم تطويرها في هذه الدراسة قدرات تصنيف واعدة، حيث حققت متوسط منطقة تحت المنحنى (AUC) قدره 0.61 مع بيانات تصل إلى الصف السادس و AUC محسّن قدره 0.65 مع بيانات تمتد إلى الصف التاسع. تؤكد النتائج على ضرورة جمع البيانات المستمرة وتحليلات مستقلة إضافية لتعزيز دقة النماذج. في النهاية، تقترح الدراسة أن هذه النماذج يمكن أن تكون مفيدة في مساعدة المعلمين على تحديد الطلاب المعرضين للخطر بشكل استباقي، مما يساهم في إبلاغ الاستراتيجيات لتحسين الاحتفاظ بالطلاب والنتائج التعليمية.
طرق
في هذه الدراسة، تم تطوير نماذج التعلم الآلي والتحقق من صحتها باستخدام مجموعة بيانات طولية تمتد على 13 عامًا لتصنيف الطلاب المتسربين من التعليم الثانوي. تم استخدام أربعة خوارزميات تصنيف تحت إشراف: الغابة العشوائية المتوازنة (B-RandomForest)، مجموعة سهلة (Adaboost Ensemble)، RSBoost (Adaboost)، وشجرة القرار باستخدام التجميع. تم تقييم أداء هذه النماذج من خلال التحقق المتقاطع بستة طيات، مما يضمن مقاييس تقييم قوية.
لتحديد فعالية كل مصنف، تم إنشاء مصفوفات الارتباك، مما سمح بتحليل مفصل لدقة التصنيف ومعدلات الخطأ. يتم توضيح سير العمل المنهجي للبحث في الشكل 1، مما يوفر تمثيلًا بصريًا للعمليات المعنية في تدريب النماذج والتحقق من صحتها.
نتائج
في هذه الدراسة، تم تحليل مجموعة بيانات طولية تمتد على 13 عامًا لتصنيف حالات التسرب من التعليم الثانوي وحالات عدم التسرب باستخدام تقنيات التعلم الآلي. شملت مجموعة البيانات مقاييس تعليمية متنوعة، بما في ذلك المهارات الأكاديمية والمعرفية، الدافع، السلوك، والرفاهية، مع التركيز على البيانات حتى الصف التاسع (الأعمار 15-16) والصف السادس (الأعمار 12-13). تم تدريب أربعة مصنفات لمعالجة عدم التوازن في حالة التسرب: الغابة العشوائية المتوازنة (B-RandomForest)، مجموعة سهلة (E-Ensemble)، RSBoost (B-Boosting)، وشجرة القرار باستخدام التجميع (B-Bagging). برز مصنف B-RandomForest كالأكثر فعالية، حيث حقق دقة متوازنة متوسطة قدرها 0.61، ومعدل استرجاع قدره 0.60 للمتسربين، وخصوصية قدرها 0.62 لغير المتسربين.
أشارت تقييمات الأداء، التي أجريت عبر التحقق المتقاطع بستة طيات، إلى أنه بينما تفوق B-Bagging في الخصوصية (0.78)، إلا أنه كان قصيرًا في تصنيف الإيجابيات الحقيقية (0.32). كما أظهر مصنف B-RandomForest متوسط منطقة تحت المنحنى (AUC) قدره 0.65، مما يعكس قدرة تمييز جيدة. كشفت تحليل أهمية الميزات أن أفضل المتنبئين ينتمون بشكل أساسي إلى مجالات المهارات المعرفية والنتائج الأكاديمية، مع كون النتائج الأكاديمية الأكثر أهمية. تضمنت الميزات الرئيسية الطلاقة في القراءة ومهارات الفهم عبر عدة صفوف، والمهارات الحسابية، والتقييمات المعرفية مثل التسمية الآلية السريعة (RAN) والمفردات من رياض الأطفال. تتوفر أوصاف الميزات التفصيلية في المواد التكميلية.
مناقشة
تسلط قسم المناقشة في هذه الدراسة الضوء على التقدم الكبير الذي تم إحرازه في توقع تسرب الطلاب من التعليم الثانوي باستخدام مجموعة بيانات طولية شاملة تمتد على 13 عامًا، والتي تغطي من رياض الأطفال حتى نهاية التعليم الثانوي. استخدمت الأبحاث مصنف الغابة العشوائية المتوازنة (B-RandomForest)، مما يوضح فعاليته في توقع حالات التسرب وعدم التسرب في وقت مبكر يصل إلى الصف السادس، مع منطقة تحت المنحنى (AUC) قدرها 0.61. تعتبر هذه القدرة التنبؤية حاسمة للتعرف المبكر على الطلاب المعرضين للخطر، مما قد يسمح بالتدخلات في الوقت المناسب لتعزيز الاحتفاظ بالطلاب ونجاحهم.
تشدد الدراسة على أهمية استخدام بيانات السجل لتصور التسرب ليس فقط كحدث فردي ولكن كقياس أوسع للتقدم التعليمي، وهو أمر ذو صلة خاصة بالطلاب الذين قد لا يتخرجون في الوقت المحدد. شمل التحليل أكثر من 300 ميزة عبر مجالات متنوعة، مثل الخلفية الأسرية، والعوامل الفردية، والنتائج الأكاديمية، مع معالجة تحديات البيانات عالية الأبعاد من خلال تقنيات التعلم الآلي. على الرغم من القيود المتعلقة بالبيانات المفقودة وعدم توازن الفئات، تمثل النتائج خطوة كبيرة إلى الأمام في البحث التعليمي، مما يوفر رؤى وأدوات قيمة للمعلمين وصانعي السياسات لدعم الطلاب المعرضين للخطر بشكل فعال. يتم تشجيع الأبحاث المستقبلية لاستكشاف استراتيجيات بديلة للتعامل مع البيانات المفقودة وتوسيع مجموعة الميزات لتحسين الدقة التنبؤية.
DOI: https://doi.org/10.1038/s41598-024-63629-0
PMID: https://pubmed.ncbi.nlm.nih.gov/38839872
Publication Date: 2024-06-05
Author(s): Maria Psyridou et al.
Primary Topic: Artificial Intelligence in Healthcare
Overview
This section highlights the critical role of education in combating poverty and fostering economic growth, while also addressing the significant challenge posed by school dropout rates. Previous research in dropout prediction has primarily focused on short-term data, limiting its effectiveness. In contrast, this study utilizes a comprehensive 13-year longitudinal dataset, tracking students from kindergarten through Grade 9. The research incorporates a wide array of parameters, including academic performance, cognitive skills, motivation, behavior, and wellbeing, alongside dropout records.
The machine learning models developed in this study exhibit promising classification capabilities, achieving a mean area under the curve (AUC) of 0.61 with data up to Grade 6 and an improved AUC of 0.65 with data extending to Grade 9. The findings underscore the necessity for ongoing data collection and further independent analyses to enhance model accuracy. Ultimately, the study suggests that these models could be instrumental in aiding educators to identify at-risk students proactively, thereby informing strategies for improving student retention and educational outcomes.
Methods
In this study, machine learning models were developed and validated using a longitudinal dataset spanning 13 years to classify upper secondary school dropouts. Four supervised classification algorithms were employed: Balanced Random Forest (B-RandomForest), Easy Ensemble (Adaboost Ensemble), RSBoost (Adaboost), and Bagging Decision Tree. The performance of these models was assessed through six-fold cross-validation, ensuring robust evaluation metrics.
To quantify the effectiveness of each classifier, confusion matrices were generated, allowing for a detailed analysis of classification accuracy and error rates. The methodological workflow of the research is illustrated in Figure 1, providing a visual representation of the processes involved in model training and validation.
Results
In this study, a 13-year longitudinal dataset was analyzed to classify upper secondary education dropout and non-dropout statuses using machine learning techniques. The dataset encompassed various educational metrics, including academic and cognitive skills, motivation, behavior, and well-being, focusing on data up to Grade 9 (ages 15-16) and Grade 6 (ages 12-13). Four classifiers were trained to address the imbalance in dropout status: Balanced Random Forest (B-RandomForest), Easy Ensemble (E-Ensemble), RSBoost (B-Boosting), and Bagging Decision Tree (B-Bagging). The B-RandomForest classifier emerged as the most effective, achieving a balanced mean accuracy of 0.61, a recall rate of 0.60 for dropouts, and a specificity of 0.62 for non-dropouts.
The performance evaluation, conducted via six-fold cross-validation, indicated that while B-Bagging excelled in specificity (0.78), it fell short in true positive classification (0.32). The B-RandomForest classifier also demonstrated a mean area under the curve (AUC) of 0.65, reflecting good discriminative ability. Feature importance analysis revealed that the top-ranked predictors predominantly belonged to the domains of cognitive skills and academic outcomes, with academic outcomes being the most significant. Key features included reading fluency and comprehension skills across multiple grades, arithmetic skills, and cognitive assessments such as rapid automatized naming (RAN) and vocabulary from kindergarten. Detailed feature descriptions are available in the supplementary materials.
Discussion
The discussion section of this study highlights the significant advancements made in predicting upper secondary school dropout using a comprehensive 13-year longitudinal dataset, which spans from kindergarten to the end of upper secondary education. The research employed the Balanced Random Forest (B-RandomForest) classifier, demonstrating its effectiveness in predicting dropout and non-dropout cases as early as Grade 6, with an area under the curve (AUC) of 0.61. This predictive capability is crucial for early identification of at-risk students, potentially allowing for timely interventions to enhance student retention and success.
The study emphasizes the importance of utilizing register data to conceptualize dropout not merely as a singular event but as a broader measure of educational progression, which is particularly relevant for students who may not graduate on time. The analysis included over 300 features across various domains, such as family background, individual factors, and academic outcomes, addressing the challenges of high-dimensional data through machine learning techniques. Despite limitations related to missing data and class imbalance, the findings represent a significant step forward in educational research, providing valuable insights and tools for educators and policymakers to support at-risk students effectively. Future research is encouraged to explore alternative strategies for handling missing data and to expand the feature set for improved predictive accuracy.
