نهج متعدد الحدود برنشتاين لتقدير دوال التوزيع التراكمي في وجود بيانات مفقودة
A Bernstein polynomial approach for the estimation of cumulative distribution functions in the presence of missing data

شارك:
المجلة: Statistical Papers، المجلد: 67، العدد: 4
DOI: https://doi.org/10.1007/s00362-026-01851-w
تاريخ النشر: 2026-05-30
المؤلف: Zhenyun Du وآخرون
الموضوع الرئيسي: الهيدرولوجيا وتحليل الجفاف

نظرة عامة

تبحث هذه الدراسة في التقدير غير المعلمي لدوال التوزيع التراكمي الأحادية المتغيرة (CDFs) في سياق البيانات المفقودة بشكل عشوائي. يقدم المؤلفون مقدرات تعزز دالة التوزيع التراكمي التجريبية (IPW) باستخدام مشغل برنشتاين، مما يؤدي إلى منحنيات أحادية ومحدودة تتكيف مع دعم البيانات. يتم تحليل نوعين من المقدرات: مقدر زائف يستخدم الميل المعروف ومقدر قابل للتطبيق يقدر الميل بشكل غير معلمي من المتغيرات المساعدة المنفصلة، وهو أكثر قابلية للتطبيق في السيناريوهات الواقعية. تستخلص الدراسة انحياز النقطة وتوسعات التباين لكلا المقدرين، وتحدد الدرجة المثلى متعددة الحدود \( m \) فيما يتعلق بالخطأ التربيعي المتوسط المتكامل، وتظهر التوزيع الطبيعي الأسيمبتي.

اكتشاف مهم هو أن المقدر القابل للتطبيق يظهر تباينًا أقل مقارنة بالمقدر الزائف، مع تعديل بواسطة مصطلح تصحيح غير سالب. يقترح المؤلفون طريقة اختيار الدرجة بناءً على التحقق المتقاطع بأقل المربعات. تشير محاكاة مونت كارلو إلى أن كلا المقدرين الملساء باستخدام برنشتاين يتفوقان على نظرائهما غير الملساء وإصدار الكثافة الأساسية IPW تحت نماذج معينة، خاصة لأحجام العينات الصغيرة إلى المتوسطة. يتم توضيح التطبيق العملي للطريقة من خلال تحليل بيانات سكر البلازما الصائم من مسح NHANES 2017-2018، مع توفر جميع الشيفرات اللازمة على GitHub لإعادة الإنتاج.

مقدمة

تؤكد مقدمة الورقة على أهمية دوال التوزيع التراكمي (CDFs) وعكسها في الاحتمالات والإحصاءات، مع تسليط الضوء على دورها في توفير فهم شامل لتوزيعات المتغيرات العشوائية بخلاف الإحصائيات الملخصة مثل المتوسط والتباين. يعد التقدير الدقيق لدالة التوزيع التراكمي أمرًا ضروريًا عبر مجالات مختلفة، بما في ذلك تحليل البقاء، وهندسة الموثوقية، والاقتصاد، وإدارة المخاطر. يتم الاعتراف بالتحديات التي تطرحها البيانات غير المكتملة في التحليلات الإحصائية، خاصة في الدراسات الكبيرة والتجارب السريرية، حيث يمكن أن تعقد الملاحظات المفقودة تقدير دوال التوزيع التراكمي.

لمعالجة هذه التحديات، تقترح الورقة مقدرًا جديدًا غير معلمي لدالة التوزيع التراكمي يجمع بين وزن الاحتمال العكسي (IPW) مع تنعيم متعدد الحدود باستخدام برنشتاين، مستهدفًا بشكل خاص البيانات المفقودة بشكل عشوائي (MAR). لا يصحح هذا النهج فقط الانحياز المرتبط بفقدان MAR، بل يضمن أيضًا أن التقديرات الناتجة سلسة وتلتزم بقيود الشكل لدوال التوزيع التراكمي. توضح المقدمة هيكل الورقة، موضحة الإطار الإحصائي، وصياغة إعداد MAR، والأقسام اللاحقة التي ستقدم النتائج النظرية، ومحاكاة مونت كارلو، والتطبيقات العملية، مما يساهم في الأدبيات الحالية حول تقدير دوال التوزيع التراكمي في وجود بيانات مفقودة.

النتائج

في هذا القسم، يبحث المؤلفون في الخصائص النظرية للمقدرات المقترحة باستخدام برنشتاين. يستخلصون توسعات أسيمبتي لكل من الانحياز النقاطي والتباين، مما يوفر رؤى حول أداء المقدرات مع زيادة أحجام العينات. علاوة على ذلك، تحدد الدراسة الدرجة المثلى متعددة الحدود \( m \) التي تقلل من الخطأ التربيعي المتوسط (MSE) والخطأ التربيعي المتوسط المتكامل (MISE). كما يثبت المؤلفون التوزيع الطبيعي الأسيمبتي لكل من المقدرين الزائف والقابل للتطبيق، كما هو محدد في المعادلة (2.4)، مما يؤكد موثوقية هذه المقدرات في سياقات العينات الكبيرة.

المناقشة

في هذا القسم، يناقش المؤلفون تطبيق مشغل برنشتاين لتقدير دالة التوزيع التراكمي (CDF) لمتغير استجابة مستمر يتعرض للفقد، باستخدام متغير مساعد منفصل تم ملاحظته بالكامل. يوفر مشغل برنشتاين، المحدد للوظائف المستمرة على الفترة [0، 1]، سلسلة من الحدود المتعددة المحدودة التي تتقارب بشكل موحد إلى الدالة المستهدفة. يقترح المؤلفون مقدرين لدالة التوزيع التراكمي: أحدهما يفترض وجود درجات ميل معروفة والآخر يستخدم التقدير غير المعلمي لهذه الدرجات. يتم بناء المقدر الأول، المسمى \( F_n(y) \)، باستخدام وزن الاحتمال العكسي (IPW) مع الميل المعروف، بينما يعتمد المقدر الثاني، \( F_n(y) \)، على تقديرات الميل المستمدة من البيانات الملاحظة.

تثبت الورقة الخصائص الأسيمبتي لكل من المقدرين، بما في ذلك تحليلات الانحياز والتباين. يظهر المؤلفون أن انحياز المقدر الملساء \( F_{n,m}(y) \) ناتج أساسًا عن عملية التنعيم باستخدام برنشتاين، مع معدل انحياز قياسي من الرتبة \( m^{-1} \). كما يتم ملاحظة تقليل التباين، خاصة بالنسبة للمقدر القابل للتطبيق، الذي يستفيد من التفاعل بين التنعيم وتقدير الميل غير المعلمي. يتم اشتقاق الخطأ التربيعي المتوسط (MSE) والخطأ التربيعي المتوسط المتكامل (MISE)، مما يكشف أن الدرجة المثلى متعددة الحدود \( m \) توازن بين الانحياز والتباين، مما يحقق معدلات تقارب محسنة. يختتم القسم بمناقشة التوزيع الطبيعي الأسيمبتي لكل من المقدرين، وهو أمر حاسم لبناء فترات الثقة، ويبرز قوة مقدر برنشتاين في سيناريوهات المحاكاة المختلفة.

Journal: Statistical Papers, Volume: 67, Issue: 4
DOI: https://doi.org/10.1007/s00362-026-01851-w
Publication Date: 2026-05-30
Author(s): Zhenyun Du et al.
Primary Topic: Hydrology and Drought Analysis

Overview

This research investigates nonparametric estimation of univariate cumulative distribution functions (CDFs) in the context of data missing at random. The authors introduce estimators that enhance the inverse probability weighted (IPW) empirical CDF using the Bernstein operator, resulting in monotonic and bounded curves that adapt to the support of the data. Two types of estimators are analyzed: a pseudo estimator utilizing known propensities and a feasible estimator that estimates propensities nonparametrically from discrete auxiliary variables, which is more applicable in real-world scenarios. The study derives pointwise bias and variance expansions for both estimators, establishes the optimal polynomial degree \( m \) concerning the mean integrated squared error, and demonstrates asymptotic normality.

A significant finding is that the feasible estimator exhibits a lower variance compared to the pseudo estimator, adjusted by a nonnegative correction term. The authors propose a degree selection method based on least-squares cross-validation. Monte Carlo simulations indicate that both Bernstein-smoothed estimators outperform their unsmoothed counterparts and the integrated version of the IPW kernel density estimator under specific models, particularly for small to moderate sample sizes. The practical application of the method is illustrated through an analysis of fasting plasma glucose data from the 2017-2018 NHANES survey, with all necessary code available on GitHub for reproducibility.

Introduction

The introduction of the paper emphasizes the significance of cumulative distribution functions (CDFs) and their inverses in probability and statistics, highlighting their role in providing a comprehensive understanding of random variable distributions beyond summary statistics like mean and variance. Accurate CDF estimation is essential across various fields, including survival analysis, reliability engineering, economics, and risk management. The challenges posed by incomplete data in statistical analyses are acknowledged, particularly in large studies and clinical trials, where missing observations can complicate the estimation of CDFs.

To address these challenges, the paper proposes a novel nonparametric CDF estimator that combines inverse probability weighting (IPW) with Bernstein polynomial smoothing, specifically targeting data that is missing at random (MAR). This approach not only corrects bias associated with MAR missingness but also ensures that the resulting estimates are smooth and adhere to the shape constraints of CDFs. The introduction outlines the structure of the paper, detailing the statistical framework, the formulation of the MAR setting, and the subsequent sections that will present theoretical results, Monte Carlo simulations, and practical applications, thereby contributing to the existing literature on CDF estimation in the presence of missing data.

Results

In this section, the authors investigate the theoretical properties of the proposed Bernstein estimators. They derive asymptotic expansions for both the pointwise bias and variance, providing insights into the estimators’ performance as sample sizes increase. Furthermore, the study identifies the optimal polynomial degree \( m \) that minimizes the mean squared error (MSE) and the mean integrated squared error (MISE). The authors also establish the asymptotic normality of both the pseudo and feasible estimators, as defined in equation (2.4), thereby confirming the reliability of these estimators in large-sample contexts.

Discussion

In this section, the authors discuss the application of the Bernstein operator to estimate the cumulative distribution function (CDF) of a continuous response variable subject to missingness, utilizing a fully observed discrete auxiliary variable. The Bernstein operator, defined for continuous functions on the interval [0, 1], provides a sequence of bounded polynomials that converge uniformly to the target function. The authors propose two estimators for the CDF: one assuming known propensity scores and another using nonparametric estimation of these scores. The first estimator, denoted as \( F_n(y) \), is constructed using inverse probability weighting (IPW) with known propensities, while the second estimator, \( F_n(y) \), relies on estimated propensities derived from the observed data.

The paper establishes the asymptotic properties of both estimators, including bias and variance analyses. The authors demonstrate that the bias of the smoothed estimator \( F_{n,m}(y) \) is primarily due to the Bernstein smoothing process, with a standard bias rate of order \( m^{-1} \). Variance reductions are also noted, particularly for the feasible estimator, which benefits from the interaction between smoothing and nonparametric propensity estimation. The mean squared error (MSE) and mean integrated squared error (MISE) are derived, revealing that the optimal polynomial degree \( m \) balances bias and variance, achieving improved convergence rates. The section concludes with a discussion of the asymptotic normality of both estimators, which is crucial for constructing confidence intervals, and highlights the robustness of the Bernstein estimator in various simulation scenarios.

شارك: