عامل التوحيد الأمثل لتباين التوزيع الطبيعي: منظور تاريخي وتوصيات للممارسة
The Optimal Standardization Factor for the Variance of the Normal Distribution: A Historical Perspective and Recommendations for Practice

شارك:
المجلة: Journal of Statistical Theory and Practice، المجلد: 20، العدد: 2
DOI: https://doi.org/10.1007/s42519-026-00567-2
تاريخ النشر: 2026-03-23
المؤلف: Dominic Edelmann
الموضوع الرئيسي: تقدير التوزيع الإحصائي وتطبيقاته

نظرة عامة

هذا القسم من ورقة البحث يقدم نظرة عامة على مقدرات التباين للتوزيع الطبيعي، مع تسليط الضوء على انتشار المقدّر غير المنحاز، المعروف بتصحيح بيسل. يستعرض المؤلفون عدة مقدرات بديلة تنشأ من استخدام عوامل توحيد مختلفة لمجموع المربعات.

تشير النتائج إلى أنه لا يوجد مبرر قوي لتفضيل مقدر التباين غير المنحاز على هذه البدائل، مما يوحي بأن هناك مقدرات أكثر فعالية. تعتبر تداعيات هذه النتائج ذات صلة خاصة بالتطبيقات التي تتضمن أحجام عينات صغيرة، حيث يمكن أن يؤثر اختيار المقدّر بشكل كبير على دقة تقدير التباين.

مقدمة

تناقش مقدمة هذه الورقة البحثية أهمية التباين والانحراف المعياري كقياسات رئيسية للتشتت في الإحصاءات. تعرف التباين لمتغير عشوائي \( X \) على أنه \( \text{Var}(X) = E[(X – E[X])^2] \) وتلاحظ أن المقدّر الشائع للتباين هو \( \frac{1}{n-1} \sum_{i=1}^{n} (X_i – \mu)^2 \)، حيث \( \mu \) هو متوسط العينة. على الرغم من تطوير عدة مقدرات بديلة على مدار القرن الماضي، إلا أن معظمها تم تجاهله في التطبيقات العملية، مع بعض الاستثناءات المذكورة في الأدبيات الإحصائية.

تهدف الورقة إلى تسليط الضوء على هذه المقدرات البديلة، خاصة في سياق التوزيع الطبيعي، حيث تجادل بأن الاستخدام السائد لمقدر \( \frac{1}{n-1} S \) يعتمد إلى حد كبير على التقليد بدلاً من مبررات إحصائية قوية. تقدم المقدرات من الشكل \( \sigma^2 = \frac{1}{c_n} S \)، حيث \( c_n > 0 \) هو ثابت يعتمد على حجم العينة، وتؤكد أن هذه المقدرات غير حساسة للترجمة ومتوافقة مع المقياس. يؤكد المؤلفون أنه بينما لا يكون المقدّر التقليدي ضارًا بالضرورة، فإن رؤاهم يمكن أن تعزز تقدير المقياس، خاصة لأحجام العينات الصغيرة، وأن بعض السيناريوهات قد تستفيد من استخدام مقدرات بديلة.

مناقشة

في هذه المناقشة، يقدم المؤلفون تحليلًا شاملاً لمختلف المقدرات لتباين متغير عشوائي موزع بشكل طبيعي، مع التركيز على قيود المقدّر غير المنحاز التقليدي $\sigma^2_{UBV} = \frac{1}{n-1} S^2$. يبرزون أنه على الرغم من أن هذا المقدّر مستخدم على نطاق واسع، إلا أنه يميل إلى التقليل من تقدير الانحراف المعياري الحقيقي، خاصة في أحجام العينات الصغيرة. يقترح المؤلفون بدائل، مثل مقدر الاحتمالية القصوى $\sigma^2_{MLE} = \frac{1}{n} S^2$ ومقدر خطأ اللوغاريتم المربع $\sigma^2_{logsq} = \frac{1}{2} \exp(-\psi(\frac{n-1}{2})) S^2$، والتي يمكن أن توفر تقديرات أكثر موثوقية تحت ظروف معينة.

تنتقد الورقة التركيز التقليدي على عدم الانحياز في التقدير، داعية إلى نهج قائم على دالة الخسارة التي تقلل من الخسارة المتوقعة بدلاً من الالتزام الصارم بمعايير عدم الانحياز. يجادل المؤلفون بأن دوال الخسارة يجب أن تظهر تناظرًا في معاقبة التقدير المفرط والتقدير الناقص، وهي خاصية تلبيها دوال خطأ اللوغاريتم المربع وخطأ اللوغاريتم المطلق. ويخلصون إلى أن مقدر خطأ اللوغاريتم المربع هو المفضل بسبب ملاءمته الرياضية وتناسقه عبر قوى مختلفة من الانحراف المعياري، موصين باستخدامه في التطبيقات العملية، خاصة في السيناريوهات ذات العينات الصغيرة حيث قد تفشل المقدرات التقليدية.

Journal: Journal of Statistical Theory and Practice, Volume: 20, Issue: 2
DOI: https://doi.org/10.1007/s42519-026-00567-2
Publication Date: 2026-03-23
Author(s): Dominic Edelmann
Primary Topic: Statistical Distribution Estimation and Applications

Overview

This section of the research paper provides an overview of variance estimators for the normal distribution, highlighting the prevalence of the unbiased estimator, known as Bessel’s correction. The authors review several alternative estimators that arise from employing different standardization factors for the sum of squares.

The findings indicate that there is no strong justification for preferring the unbiased variance estimator over these alternatives, suggesting that more effective estimators exist. The implications of these results are particularly relevant for applications involving small sample sizes, where the choice of estimator can significantly impact the accuracy of variance estimation.

Introduction

The introduction of this research paper discusses the significance of variance and standard deviation as primary measures of dispersion in statistics. It defines the variance of a random variable \( X \) as \( \text{Var}(X) = E[(X – E[X])^2] \) and notes that the common estimator for variance is \( \frac{1}{n-1} \sum_{i=1}^{n} (X_i – \mu)^2 \), where \( \mu \) is the sample mean. Despite the development of various alternative estimators over the past century, most have been overlooked in practical applications, with a few exceptions noted in statistical literature.

The paper aims to highlight these alternative estimators, particularly in the context of the normal distribution, arguing that the prevalent use of the \( \frac{1}{n-1} S \) estimator is largely based on tradition rather than strong statistical justification. It introduces estimators of the form \( \sigma^2 = \frac{1}{c_n} S \), where \( c_n > 0 \) is a constant dependent on sample size, and emphasizes that these estimators are translation-invariant and scale-equivariant. The authors assert that while the conventional estimator is not necessarily harmful, their insights could enhance scale estimation, especially for small sample sizes, and that certain scenarios may benefit from employing alternative estimators.

Discussion

In this discussion, the authors provide a comprehensive analysis of various estimators for the variance of a normally distributed random variable, emphasizing the limitations of the conventional unbiased estimator $\sigma^2_{UBV} = \frac{1}{n-1} S^2$. They highlight that while this estimator is widely used, it tends to underestimate the true standard deviation, especially in small sample sizes. The authors propose alternatives, such as the maximum likelihood estimator $\sigma^2_{MLE} = \frac{1}{n} S^2$ and the squared log-error estimator $\sigma^2_{logsq} = \frac{1}{2} \exp(-\psi(\frac{n-1}{2})) S^2$, which can provide more reliable estimates under certain conditions.

The paper critiques the traditional focus on unbiasedness in estimation, advocating for a loss-function-based approach that minimizes expected loss rather than strictly adhering to unbiased criteria. The authors argue that loss functions should exhibit symmetry in penalizing overestimation and underestimation, a property satisfied by the squared log-error and absolute log-error functions. They conclude that the squared log-error estimator is preferable due to its mathematical convenience and consistency across different powers of the standard deviation, recommending its use in practical applications, particularly in small-sample scenarios where conventional estimators may fail.

شارك: