اختبار ويلش t أكثر حساسية لانتهاكات العالم الحقيقي لافتراضات التوزيع من اختبار t للطالب، لكن الانحدار اللوجستي أكثر قوة من كليهما
Welch’s t test is more sensitive to real world violations of distributional assumptions than student’s t test but logistic regression is more robust than either

المجلة: Statistical Papers، المجلد: 65، العدد: 6
DOI: https://doi.org/10.1007/s00362-024-01531-7
تاريخ النشر: 2024-03-04
المؤلف: David Curtis
الموضوع الرئيسي: طرق إحصائية في التجارب السريرية

نظرة عامة

في هذا القسم، ينتقد المؤلفون أداء اختبار ويلش t في سياق اختبارات الحمل المثقلة الموزونة على مستوى الجين للكشف عن الارتباطات بين المتغيرات النادرة والصدفية. يبرزون أن اختبار ويلش، الذي يُوصى به عادةً لمعالجة فرضية تساوي التباينات، يمكن أن يظهر معدل خطأ من النوع الأول مرتفعًا عندما يتم انتهاك فرضية التوزيع الطبيعي، خاصة في الحالات التي يتبع فيها البيانات توزيع بواسون وأحجام العينات غير متساوية. تشير هذه النتيجة إلى أنه يجب على الباحثين توخي الحذر عند تطبيق اختبار ويلش في هذه الظروف، حيث قد يؤدي ذلك إلى انحياز منهجي وقيم p منخفضة بشكل مفرط.

على النقيض من ذلك، يُبلغ المؤلفون أن اختبار t لستودنت يؤدي بشكل أكثر موثوقية في هذه السيناريوهات. ومع ذلك، يدعون إلى استخدام اختبار نسبة الاحتمالات بناءً على نماذج الانحدار اللوجستي كبديل أفضل لاختبار الفروق في التوزيعات بين عينتين. تؤكد هذه الدراسة، التي أُجريت باستخدام موارد بنك البيانات البريطاني، على أهمية اختيار الأساليب الإحصائية المناسبة في دراسات الارتباط الجيني لتجنب النتائج المضللة.

مقدمة

في المقدمة، يناقش المؤلفون تفضيل اختبار ويلش t على اختبار t لستودنت واختبار U لويلكوك، مشيرين إلى أبحاث راش وآخرين (2011) وديلكر وآخرين (2017، 2022). يبرزون أن اختبار t لستودنت يمكن أن يظهر معدل خطأ من النوع الأول مرتفعًا عندما تكون أحجام العينات غير متساوية وتختلف الانحرافات المعيارية، بينما قد يكون لاختبار U لويلكوك قوة أقل. يجادل ديلكر وآخرون بأن اختبار ويلش t عمومًا أكثر قوة، خاصة في السيناريوهات التي يتم فيها انتهاك فرضية تساوي الانحرافات المعيارية. توضح مثال أن نسبة الانحراف المعياري 2 وأحجام العينات 40 و60، يمكن أن يصل معدل خطأ النوع الأول لاختبار t لستودنت إلى 0.083 عند α = 0.05، بينما يحافظ اختبار ويلش t على المعدل المتوقع 0.05 في نفس الظروف.

يستكشف المؤلفون أيضًا قيود اختبار ويلش t، خاصة في التطبيقات الواقعية، مثل اختبار الحمل المثقل للمتغيرات الجينية النادرة المرتبطة بالصدفية في بنك البيانات البريطاني. يشيرون إلى أن التنفيذ الافتراضي لاختبار ويلش t في R أنتج قيم p منخفضة للغاية لعدة جينات، والتي كانت أكثر توافقًا مع النتائج من اختبار نسبة الاحتمالات عندما تم استخدام اختبار t لستودنت بدلاً من ذلك. أدت هذه الفجوة إلى قيام المؤلفين بإجراء محاكاة تقارن أداء اختبار ويلش t واختبار t لستودنت والانحدار اللوجستي، مع التركيز على السيناريوهات ذات أحجام العينات والتوزيعات المتغيرة. تمهد المقدمة الطريق لتحقيق أعمق في الخصائص الإحصائية لهذه الاختبارات في سياق دراسات الارتباط الجيني.

النتائج

تشير نتائج التحليلات الاستكشافية إلى أن أداء الاختبارات الإحصائية الثلاثة—اختبار ويلش t، اختبار t لستودنت، واختبار نسبة الاحتمالات—يختلف بشكل كبير تحت ظروف مختلفة. عندما تكون أحجام العينات متساوية أو عندما يتم تطبيق التوزيع الطبيعي، تؤدي جميع الاختبارات بشكل كافٍ. ومع ذلك، مع أحجام عينات غير متساوية وتوزيع بواسون، يظهر اختبار ويلش t معدل خطأ من النوع الأول مرتفعًا، خاصة عند $\alpha = 0.05$، بمعدل 0.082 مقارنة بـ 0.046 لاختبار t لستودنت و0.056 لاختبار نسبة الاحتمالات. هذا الاختبار أيضًا متحيز نحو تحديد المتوسطات المنخفضة في العينات الصغيرة كدالة ذات دلالة إحصائية، مما يؤدي إلى قيمة p المعيارية المتوسطة (SLP) تبلغ -0.23، والتي تنحرف عن القيمة المتوقعة 0.

تكشف التحليلات الإضافية أن اختبار ويلش t عرضة لإنتاج قيم p منخفضة للغاية، حيث أسفر 878 من 100,000 محاكاة عن SLP أقل من -4، مما يتجاوز بكثير التكرارات المتوقعة بالصدفة. على النقيض من ذلك، يظهر اختبار t لستودنت واختبار نسبة الاحتمالات أداءً أكثر استقرارًا، مع قيم متطرفة أقل. من الجدير بالذكر أن زيادة حجم العينة تحسن أداء اختبار ويلش t، مما يقلل من معدل خطأ النوع الأول إلى 0.055 ويعدل متوسط SLP ليكون أقرب إلى القيمة المرغوبة. ومع ذلك، لا يزال اختبار t لستودنت يظهر عيوبًا، خاصة في تضخيم SLPs الإيجابية وتخفيض SLPs السلبية، مما يشير إلى أنه قد يكون أقل من اختبار نسبة الاحتمالات في بعض السيناريوهات.

المناقشة

في هذه المناقشة، يتم تقييم أداء اختبار ويلش t واختبار t لستودنت بشكل نقدي استنادًا إلى المحاكاة التي تتضمن أحجام عينات وتوزيعات مختلفة. تشير النتائج إلى أنه بينما تؤدي كلا الاختبارين عمومًا بشكل كافٍ، يظهر اختبار ويلش t معدل خطأ من النوع الأول مثير للقلق يبلغ 0.078 عند تطبيقه على أحجام عينات غير متساوية مع توزيع بواسون، مما يتناقض مع المعدلات المقبولة للاختبارات الأخرى. تكشف الصيغ الأساسية لتقدير الخطأ المعياري للاختلاف في المتوسطات أن اختبار ويلش، الذي يقسم التباينات على أحجام العينات، يمكن أن يؤدي إلى انحياز منهجي عندما يكون لدى العينة الأكبر درجة أعلى، مما يؤدي إلى تضخيم إحصائيات t. على النقيض من ذلك، يظهر اختبار t لستودنت تخفيضًا طفيفًا في احتمالات اللوغاريتم المعياري السلبية (SLPs) وتضخيم SLPs الإيجابية، ولكن دون إنتاج قيم p متطرفة.

تشير النتائج إلى أنه يجب على الباحثين توخي الحذر عند استخدام اختبار ويلش t في هذه الظروف، حيث قد لا يكون مناسبًا لمجموعات البيانات التي تشبه تلك التي تم محاكاتها. نظرًا لأن التنفيذ الافتراضي لاختبارات t في R هو اختبار ويلش، يدعو المؤلفون إلى اختيار أكثر دقة بين الاختبارين، خاصة في السيناريوهات التي تتواجد فيها توزيعات بواسون وأحجام عينات غير متساوية. علاوة على ذلك، يتم تسليط الضوء على اختبار نسبة الاحتمالات، القائم على الانحدار اللوجستي، كبديل محتمل أفضل، حيث يظهر معدلات خطأ من النوع الأول مقبولة وقوة أكبر في ظل ظروف معينة. يوصي المؤلفون بمزيد من التحقيق في مزايا وعيوب هذه الأساليب الإحصائية لإبلاغ أفضل الممارسات في البحث.

Journal: Statistical Papers, Volume: 65, Issue: 6
DOI: https://doi.org/10.1007/s00362-024-01531-7
Publication Date: 2024-03-04
Author(s): David Curtis
Primary Topic: Statistical Methods in Clinical Trials

Overview

In this section, the authors critique the performance of Welch’s t test in the context of gene-wise weighted burden tests for detecting associations between rare variants and psoriasis. They highlight that Welch’s test, which is typically recommended to address the assumption of equal variances, can exhibit an inflated Type I error rate when the assumption of normality is violated, particularly in cases where the data follows a Poisson distribution and sample sizes are unequal. This finding suggests that researchers should be cautious when applying Welch’s test under these conditions, as it may lead to systematic bias and excessively low p-values.

In contrast, the authors report that Student’s t test performs more reliably in these scenarios. However, they advocate for the use of a likelihood ratio test based on logistic regression models as a superior alternative for testing differences in distributions between two samples. This research, conducted using the UK Biobank Resource, underscores the importance of selecting appropriate statistical methods in genetic association studies to avoid misleading results.

Introduction

In the introduction, the authors discuss the preference for Welch’s t test over Student’s t test and Wilcoxon’s U test, citing research by Rasch et al. (2011) and Delacre et al. (2017, 2022). They highlight that Student’s t test can exhibit an inflated Type I error rate when sample sizes are unequal and standard deviations differ, while Wilcoxon’s U test may have lower power. Delacre et al. argue that Welch’s t test is generally more robust, particularly in scenarios where the assumption of equal standard deviations is violated. An example illustrates that with a standard deviation ratio of 2 and sample sizes of 40 and 60, the Type I error rate for Student’s t test can reach 0.083 at α = 0.05, whereas Welch’s t test maintains the expected rate of 0.05 under the same conditions.

The authors further explore the limitations of Welch’s t test, particularly in real-world applications, such as a weighted burden test of rare genetic variants associated with psoriasis in the UK Biobank. They note that the default implementation of Welch’s t test in R produced extremely low p values for several genes, which were more consistent with results from a likelihood ratio test when Student’s t test was employed instead. This discrepancy prompted the authors to conduct simulations comparing the performance of Welch’s t test, Student’s t test, and logistic regression, focusing on scenarios with varying sample sizes and distributions. The introduction sets the stage for a deeper investigation into the statistical properties of these tests in the context of genetic association studies.

Results

The results of the exploratory analyses indicate that the performance of the three statistical tests—Welch’s t test, Student’s t test, and the likelihood ratio test—varies significantly under different conditions. When sample sizes are equal or when the normal distribution is applied, all tests perform adequately. However, with unequal sample sizes and a Poisson distribution, Welch’s t test exhibits an inflated Type I error rate, particularly at $\alpha = 0.05$, with a rate of 0.082 compared to 0.046 for Student’s t test and 0.056 for the likelihood ratio test. This test is also biased towards identifying lower means in smaller samples as statistically significant, resulting in a mean standardized log p-value (SLP) of -0.23, which deviates from the expected value of 0.

Further analysis reveals that Welch’s t test is prone to producing extreme low p-values, with 878 out of 100,000 simulations yielding an SLP less than -4, far exceeding the expected occurrences by chance. In contrast, Student’s t test and the likelihood ratio test show more stable performance, with fewer extreme values. Notably, increasing the sample size improves the performance of Welch’s t test, reducing its Type I error rate to 0.055 and adjusting the mean SLP closer to the desired value. However, the Student’s t test still demonstrates deficiencies, particularly in the inflation of positive SLPs and deflation of negative SLPs, suggesting that it may be inferior to the likelihood ratio test in certain scenarios.

Discussion

In this discussion, the performance of Welch’s t test and Student’s t test is critically evaluated based on simulations involving different sample sizes and distributions. The results indicate that while both tests generally perform adequately, Welch’s t test exhibits a concerning Type I error rate of 0.078 when applied to unequal sample sizes with a Poisson distribution, contrasting with the acceptable rates of the other tests. The underlying formulas for estimating the standard error of the mean difference reveal that Welch’s test, which divides variances by sample sizes, can lead to a systematic bias when the larger sample has a higher mean score, resulting in inflated t statistics. In contrast, Student’s t test shows a slight deflation of negative standardized log probabilities (SLPs) and inflation of positive SLPs, but without generating extreme p-values.

The findings suggest that researchers should be cautious when using Welch’s t test under these conditions, as it may not be suitable for datasets resembling those simulated. Given that the default implementation of t tests in R is Welch’s test, the authors advocate for a more deliberate choice between the two tests, particularly in scenarios where Poisson distributions and unequal sample sizes are present. Furthermore, the likelihood ratio test, based on logistic regression, is highlighted as a potentially superior alternative, demonstrating acceptable Type I error rates and greater power under certain conditions. The authors recommend further investigation into the advantages and disadvantages of these statistical methods to inform best practices in research.