متوسط القيم الذاتية المعاقبة: التمديد لمقارنة النماذج المتداخلة وتقييمات مونت كارلو
Penalized eigenvalue block averaging: Extension to nested model comparison and Monte Carlo evaluations

شارك:
المجلة: Behavior Research Methods، المجلد: 58، العدد: 4
DOI: https://doi.org/10.3758/s13428-026-02968-4
PMID: https://pubmed.ncbi.nlm.nih.gov/41952013
تاريخ النشر: 2026-04-08
المؤلف: Njål Foldnes وآخرون
الموضوع الرئيسي: المنهجيات النفسية والاختبار

نظرة عامة

تناقش هذه section تقييم جودة المطابقة والنماذج المتداخلة متعددة المجموعات في تحليل العوامل التأكيدي (CFA) تحت ظروف غير طبيعية، وهو جانب حاسم في علم النفس القياسي. توسع الدراسة إجراء متوسط القيم الذاتية المعاقب (pEBA)، الذي تم إظهار فعاليته في التحكم في معدلات الخطأ من النوع الأول، من خلال اختباره عبر أبعاد أعلى لكل من المتجهات الكامنة والملاحظة. تقيم الأبحاث متغيرات pEBA المختلفة جنبًا إلى جنب مع إحصائيات الاختبار التقليدية، وتحلل أدائها من حيث التحكم في خطأ النوع الأول والقوة الإحصائية. ومن الجدير بالذكر أن أفضل طريقة لأداء اختبار جودة المطابقة تم تحديدها على أنها pEBA مع أربعة كتل باستخدام إحصائية RLS لبراون ومقدّر مصفوفة التباين غير المنحازة. بالنسبة لثبات القياس، حقق pEBA مع كتل فردية وإحصائية الاحتمالية القصوى (ML) نتائج مثالية.

تسلط الخاتمة الضوء على دراسة شاملة لمونت كارلو تشمل 525 حالة لتقييم 38 إجراء اختبار لـ CFA. تشير النتائج إلى أن طرق pEBA، وخاصة تلك التي تستخدم أربعة كتل لاختبار جودة المطابقة وكتل فردية لمقارنات النماذج المتداخلة، تفوقت بشكل كبير على الاختبارات التقليدية مثل اختبارات ساتورا-بينتلر والاختبارات المعدلة والمزاحة من حيث التحكم في خطأ النوع الأول والقوة. تؤكد هذه الأبحاث فعالية نهج pEBA في تعزيز قوة الاختبارات الإحصائية في العلوم الاجتماعية والسلوكية، مع تنفيذ الإجراءات في حزمة R الجديدة المطورة semTests.

النتائج

يوفر قسم النتائج نظرة شاملة على معدلات رفض الاختبار عبر ظروف مختلفة، مع قوائم مفصلة وتصويرات متاحة في المواد التكميلية. يركز المؤلفون على مقاييس الأداء المجمعة لفئات معينة من الاهتمام، مع تسليط الضوء على النتائج الرئيسية التي تتماشى مع البيانات غير المجمعة المقدمة في الجداول التكميلية. بالإضافة إلى ذلك، تم تضمين تمثيل بصري لهذه النتائج غير المجمعة في Study1_Unaggregated_plots المشار إليها، مما يعزز اتساق الأنماط الملاحظة عبر ظروف مختلفة.

المناقشة

في هذا القسم، يناقش المؤلفون اختبار جودة المطابقة في نماذج العوامل التأكيدية، مع التركيز على إطار اختبار التقارب الذي أنشأه ساتورا (1989) وبولين (1989). يوضحون هيكل نموذج العامل التأكيدي، الذي يربط المتغيرات الملاحظة \( \mathbf{x} \) بالعوامل الكامنة \( \mathbf{f} \) من خلال معادلة خطية، ويشرحون طريقة تقدير الاحتمالية القصوى لنظرية الطبيعية (NTML) المستخدمة لاشتقاق مصفوفة التباين المفترضة للنموذج. يؤكد المؤلفون على أهمية اختبار الفرضية الصفرية لنموذج محدد بشكل صحيح مقابل بديل عدم تحديد النموذج، باستخدام إحصائيات الملاءمة مثل \( T_{ML} \) و \( T_{RLS} \). يشيرون إلى أنه تحت الظروف الطبيعية، تتقارب هذه الإحصائيات إلى توزيع كاي-تربيع، بينما تتطلب البيانات غير الطبيعية نهج مجموعات مربعة موزونة، مما يبرز الحاجة إلى تقدير دقيق للأوزان \( \lambda_j \) لتصحيح قيم p.

يقدم المؤلفون أيضًا اختبارات جودة المطابقة المستندة إلى القيم الذاتية المختلفة ويناقشون تصحيح ساتورا-بينتلر (SB) كتقريب مستقر للأوزان. يستكشفون متغيرات جديدة من الإجراءات المعروفة، بما في ذلك متوسط القيم الذاتية المعاقب (pEBA) واختبار المعدل والمزاح (SS)، ويقارنون أدائها من حيث التحكم في خطأ النوع الأول والقوة. يختتم القسم بوصف منهجية الدراسة، التي تشمل ثلاث دراسات لمونت كارلو مصممة لتقييم اختبار جودة المطابقة ومقارنات النماذج المتداخلة تحت ظروف مختلفة، بما في ذلك أحجام العينات والعائلات التوزيعية. تهدف النتائج إلى تقديم رؤى حول قوة وفعالية إجراءات الاختبار المختلفة في التطبيقات العملية، لا سيما في أبحاث العلوم الاجتماعية.

القيود

تسلط قيود هذه الدراسة الضوء على عدة مجالات رئيسية للبحث المستقبلي. تستند النتائج إلى محاكاة تشمل ست فئات توزيع غير طبيعية، والتي قد لا تعمم على توزيعات متعددة المتغيرات الأخرى. حافظت التوزيعات المستخدمة على انحراف وتفرطح متطابقين عبر جميع المتغيرات، وهو شرط من غير المحتمل أن يتحقق في البيانات الواقعية. بالإضافة إلى ذلك، ركزت الدراسة على نموذج خمسة عوامل محدد لتقييم جودة المطابقة ونموذج عامل واحد للمقارنات المتداخلة، مما يشير إلى أن الأعمال المستقبلية يجب أن تستكشف مجموعة أوسع من تكوينات وأنواع النماذج، لا سيما ضمن نمذجة المعادلات الهيكلية.

علاوة على ذلك، لم تتناول الأبحاث البيانات المفقودة أو البيانات الفئوية الترتيبية، على الرغم من أن المؤلفين يشيرون إلى أن إجراءات الاختبار المعاقبة يمكن أن تتكيف مع هذه السيناريوهات. يُوصى أيضًا بالتحقيق في أداء طريقة المربعات الصغرى القوية (RLS) تحت أحجام عينات متغيرة، لا سيما فيما يتعلق بالشرط الطبيعي، حيث وجدت الدراسة الحالية أداءً مرضيًا حتى في ظروف صعبة. أخيرًا، يقترح المؤلفون أن إطار العقوبة قد يعزز جودة الاستدلال عبر فئة أوسع من النماذج حيث تتقارب إحصائيات الاختبار إلى مزيج من متغيرات كاي-تربيع، لا سيما في الإعدادات عالية الأبعاد.

Journal: Behavior Research Methods, Volume: 58, Issue: 4
DOI: https://doi.org/10.3758/s13428-026-02968-4
PMID: https://pubmed.ncbi.nlm.nih.gov/41952013
Publication Date: 2026-04-08
Author(s): Njål Foldnes et al.
Primary Topic: Psychometric Methodologies and Testing

Overview

This section discusses the evaluation of goodness-of-fit and multi-group nested models in confirmatory factor analysis (CFA) under non-normal conditions, a critical aspect in psychometrics. The study extends the penalized eigenvalue block averaging (pEBA) procedure, previously shown to control type I error rates effectively, by testing it across higher dimensions for both latent and observed vectors. The research assesses various pEBA variants alongside traditional test statistics, analyzing their performance in terms of type I error control and statistical power. Notably, the best-performing method for goodness-of-fit testing was identified as pEBA with four blocks using the Browne’s RLS statistic and an unbiased asymptotic covariance matrix estimator. For measurement invariance, pEBA with singleton blocks and the maximum likelihood (ML) statistic yielded optimal results.

The conclusion highlights a comprehensive Monte Carlo study involving 525 conditions to evaluate 38 test procedures for CFA. The findings indicate that the pEBA methods, particularly those utilizing four blocks for goodness-of-fit and singleton blocks for nested model comparisons, significantly outperformed traditional tests like the Satorra-Bentler and scaled-and-shifted tests in both type I error control and power. This research underscores the efficacy of the pEBA approach in enhancing the robustness of statistical testing in social and behavioral sciences, with the procedures implemented in the newly developed R package semTests.

Results

The results section provides a comprehensive overview of test rejection rates across various conditions, with detailed lists and visualizations available in the supplementary material. The authors focus on aggregated performance metrics for specific substrata of interest, highlighting key findings that align with the unaggregated data presented in the supplementary tables. Additionally, a visual representation of these unaggregated results is included in the referenced Study1_Unaggregated_plots, reinforcing the consistency of the observed patterns across different conditions.

Discussion

In this section, the authors discuss the testing of goodness-of-fit in confirmatory factor models, focusing on the asymptotic testing framework established by Satorra (1989) and Bollen (1989). They outline the confirmatory factor model’s structure, which relates observed variables \( \mathbf{x} \) to latent factors \( \mathbf{f} \) through a linear equation, and detail the normal-theory maximum likelihood (NTML) estimation method used to derive the model-implied covariance matrix. The authors emphasize the importance of testing the null hypothesis of a correctly specified model against the alternative of model misspecification, utilizing fit statistics such as \( T_{ML} \) and \( T_{RLS} \). They note that under normality, these statistics converge to a chi-square distribution, while non-normal data necessitates a weighted sum of squares approach, highlighting the need for accurate estimation of weights \( \lambda_j \) to correct p-values.

The authors further introduce various eigenvalue-based goodness-of-fit tests and discuss the Satorra-Bentler (SB) correction as a stabilizing approximation for the weights. They explore new variants of established procedures, including penalized eigenvalue block averaging (pEBA) and the scaled and shifted (SS) test, comparing their performance in terms of type I error control and power. The section concludes with a description of the study’s methodology, which includes three Monte Carlo studies designed to evaluate goodness-of-fit testing and nested model comparisons under various conditions, including sample sizes and distributional families. The findings aim to provide insights into the robustness and efficacy of different testing procedures in practical applications, particularly in social sciences research.

Limitations

The limitations of this study highlight several key areas for future research. The findings are based on simulations that encompass six non-normal distributional classes, which may not generalize to other multivariate distributions. The employed distributions maintained identical skewness and kurtosis across all variables, a condition that is unlikely to be met in real-world data. Additionally, the study focused on a specific five-factor model for goodness-of-fit evaluations and a one-factor model for nested comparisons, suggesting that future work should explore a broader array of model configurations and types, particularly within structural equation modeling.

Moreover, the research did not address missing data or ordinal-categorical data, although the authors note that the penalized test procedures could potentially be adapted to accommodate these scenarios. Investigating the performance of the robust least squares (RLS) method under varying sample sizes, particularly in relation to normality, is also recommended, as the current study found satisfactory performance even in challenging conditions. Lastly, the authors suggest that the penalization framework may enhance inference quality across a wider class of models where test statistics converge to a mixture of chi-square variates, particularly in high-dimensional settings.

شارك: