DOI: https://doi.org/10.1186/s40536-024-00192-0
تاريخ النشر: 2024-03-07
المؤلف: Francis L. Huang
الموضوع الرئيسي: طرق إحصائية واستدلال بايزي
نظرة عامة
تتناول ورقة البحث الاستخدام المتزايد للتقييمات واسعة النطاق (LSAs) في التعليم وتبرز التطبيق المحدود للنماذج متعددة المستويات (MLMs) لتحليل بيانات LSA باستخدام R. يُعزى هذا القيد إلى التعقيدات المرتبطة بدمج القيم المحتملة والتحليلات الموزونة. لمعالجة هذه الفجوة، يقدم المؤلفون وظائف محسّنة ضمن حزمة WeMix (Bailey et al., 2023) التي تسهل التجميع التلقائي للقيم المحتملة، ومقارنات النماذج من خلال اختبارات نسبة الاحتمالية، وتصدير النتائج بتنسيقات مختلفة مثل Word وHTML.
في الختام، توضح المخطوطة القدرات الموسعة لحزمة WeMix، والتي تعتبر أساسية لتحليل بيانات LSA بشكل فعال التي تستخدم القيم المحتملة. لا تدعم الوظائف الجديدة مقارنات النماذج فحسب، بل تسهل أيضًا عملية الإخراج لتسهيل التحرير. تتضمن الورقة نتائج مفصلة من تحليلات MLM ذات المستويين باستخدام كل من SAS وحزمة EdSurvey في R، مما يوفر نتائج إحصائية قوية تؤكد على أهمية هذه التحسينات في أبحاث التقييم التعليمي.
النتائج
في هذا القسم، يناقش المؤلفون تطبيق اختبارات نسبة الاحتمالية (LRT) لتقييم ملاءمة النماذج متعددة المستويات من خلال مقارنة النماذج المتداخلة. يقيم LRT الفرق في إحصائيات الانحراف، المحسوبة كـ Δd = -2LL_{FULL} – (-2LL_{REDUCED})، باستخدام إحصائية χ² مع درجات الحرية (k) التي تتوافق مع الفرق في المعلمات المقدرة. تشير النتيجة المهمة إلى أن النموذج الكامل يوفر ملاءمة أفضل، بينما تشير النتيجة غير المهمة إلى أن نموذجًا أبسط قد يكون كافيًا. يبرز المؤلفون تعقيد أساليب التجميع لاختبارات LRT ويقدمون إحصائية D²، التي قدمها Li et al. (1991) وشرحها Schafer (1997)، والتي تجمع إحصائيات χ² من نماذج متعددة تم تحليلها باستخدام قيم محتملة مختلفة.
يتم حساب إحصائية D² كـ \( D^2 = \frac{d_{k-m+1}}{m-1} r_m \) حيث \( d \) هو متوسط إحصائية Δd، و \( r_m \) يأخذ في الاعتبار الزيادة النسبية في التباين بسبب القيم المفقودة. يستخدم تقييم إحصائية D² توزيع F، مع تعديل لعدد التقديرات. بينما قد تؤدي إحصائية D² إلى معدلات أعلى من خطأ النوع الأول في عينات أصغر، فإن هذا أقل قلقًا في مجموعات البيانات الكبيرة النموذجية للتقييمات واسعة النطاق (LSAs). يحذر المؤلفون أيضًا من الاعتماد فقط على مقاييس معيار المعلومات، مثل معيار معلومات أكايك (AIC)، لاختيار النماذج، مشيرين إلى أن تجميع قيم AIC من مجموعات بيانات متعددة قد لا يؤدي إلى نتائج موثوقة. يشيرون إلى محاكاة تشير إلى أن طرق التجميع المختلفة لـ AIC تؤدي بشكل مشابه ولكنهم يبرزون المخاطر المحتملة لاستخدام هذه المعايير لتقييم ملاءمة النموذج.
المناقشة
يتناول قسم المناقشة في الورقة التعقيدات المرتبطة بتحليل بيانات التقييمات واسعة النطاق (LSAs)، مع التركيز بشكل خاص على آثار أوزان العينة والقيم المحتملة. تستخدم LSAs طرق أخذ عينات متعددة المراحل، مما يستلزم استخدام أوزان العينة لضمان أن التحليلات تعكس بدقة تقديرات السكان بدلاً من خصائص العينة. تبرز الورقة النقاش المستمر حول ضرورة وتأثير هذه الأوزان، مع توصيات لاستخدامها لتقليل التحيز في تقديرات النقاط والأخطاء المعيارية. تؤكد على أن النماذج متعددة المستويات يمكن أن تتضمن أوزانًا على مستويات مختلفة، بما يتماشى مع تصميم العينة، وتقترح طرقًا متنوعة لتطبيقها.
بالإضافة إلى ذلك، يناقش القسم مفهوم القيم المحتملة، التي تُستخدم لمعالجة مشكلة التقييمات غير المكتملة، حيث لا يكمل الطلاب الاختبار بالكامل. تُشتق هذه القيم من توزيع القدرة المقدرة، وتعمل كقياسات مستعارة لإنجاز الطلاب ويجب التعامل معها بشكل مناسب لتجنب التقليل من تقديرات الأخطاء المعيارية. يدعو المؤلفون إلى ملاءمة النماذج عدة مرات باستخدام كل قيمة محتملة وتجميع النتائج وفقًا لقواعد روبن لأخذ التباين في الاعتبار. تختتم الورقة بتقديم وظائف جديدة ضمن حزمة WeMix لـ R، تهدف إلى تبسيط تحليل بيانات LSA باستخدام النماذج متعددة المستويات وتسهيل تجميع النتائج من النماذج التي تستخدم القيم المحتملة.
DOI: https://doi.org/10.1186/s40536-024-00192-0
Publication Date: 2024-03-07
Author(s): Francis L. Huang
Primary Topic: Statistical Methods and Bayesian Inference
Overview
The research paper discusses the increasing utilization of large-scale assessments (LSAs) in education and highlights the limited application of multilevel models (MLMs) for analyzing LSA data using R. This limitation is attributed to the complexities involved in integrating plausible values and weighted analyses. To address this gap, the authors introduce enhanced functions within the WeMix package (Bailey et al., 2023) that facilitate the automatic pooling of plausible values, model comparisons through likelihood ratio tests, and the exportation of results in various formats such as Word and HTML.
In conclusion, the manuscript demonstrates the expanded capabilities of the WeMix package, which are essential for effectively analyzing LSA data that employs plausible values. The newly introduced functions not only support model comparisons but also streamline the output process for easier editing. The paper includes detailed results from two-level MLM analyses using both SAS and the EdSurvey package in R, providing robust statistical findings that underscore the significance of these enhancements in educational assessment research.
Results
In this section, the authors discuss the application of likelihood ratio tests (LRT) for evaluating the fit of multilevel models by comparing nested models. The LRT assesses the difference in deviance statistics, calculated as Δd = -2LL_{FULL} – (-2LL_{REDUCED}), using a χ² statistic with degrees of freedom (k) corresponding to the difference in estimated parameters. A significant result indicates that the full model provides a better fit, while a non-significant result suggests that a simpler model may be adequate. The authors highlight the complexity of pooling approaches for LRTs and present the D² statistic, introduced by Li et al. (1991) and elaborated by Schafer (1997), which pools χ² statistics from multiple models analyzed with different plausible values.
The D² statistic is computed as \( D^2 = \frac{d_{k-m+1}}{m-1} r_m \) where \( d \) is the average Δd statistic, and \( r_m \) accounts for the relative increase in variance due to missing values. The evaluation of the D² statistic employs an F distribution, adjusting for the number of imputations. While the D² statistic may yield higher Type I error rates in smaller samples, this is less of a concern in large datasets typical of large-scale assessments (LSAs). The authors also caution against relying solely on information criterion measures, such as the Akaike information criterion (AIC), for model selection, noting that pooling AIC values from multiple datasets may not yield reliable results. They reference simulations indicating that various pooling methods for AIC perform similarly but emphasize the potential pitfalls of using these criteria for model fit assessment.
Discussion
The discussion section of the paper addresses the complexities involved in analyzing Large-Scale Assessments (LSAs) data, particularly focusing on the implications of sampling weights and plausible values. LSAs utilize multistage sampling methods, which necessitate the use of sampling weights to ensure that analyses accurately reflect population estimates rather than sample characteristics. The paper highlights the ongoing debate regarding the necessity and impact of these weights, with recommendations for their use to mitigate bias in point estimates and standard errors. It emphasizes that multilevel models can incorporate weights at different levels, aligning with the sampling design, and suggests various approaches for their application.
Additionally, the section discusses the concept of plausible values, which are employed to address the issue of incomplete assessments, as students do not complete the entire test. These values, derived from an estimated ability distribution, serve as imputed measures of student achievement and must be handled appropriately to avoid underestimating standard errors. The authors advocate for fitting models multiple times using each plausible value and pooling the results according to Rubin’s rules to account for variability. The paper concludes by introducing new functions within the WeMix package for R, aimed at simplifying the analysis of LSA data using multilevel models and facilitating the pooling of results from models that utilize plausible values.
