DOI: https://doi.org/10.1038/s41598-024-67189-1
PMID: https://pubmed.ncbi.nlm.nih.gov/39019932
تاريخ النشر: 2024-07-17
المؤلف: Sohaib Ahmad وآخرون
الموضوع الرئيسي: تقنيات أخذ العينات والتقدير
نظرة عامة
تقدم هذه المقالة مقدرًا محسّنًا للوسيط السكاني باستخدام معلومات مساعدة ضمن إطار أخذ العينات العشوائية البسيطة. يستخرج المؤلفون تعبيرات للانحياز ومتوسط مربع الخطأ (MSE) للمقدر المقترح حتى تقريب من الدرجة الأولى ويحددون تقديرات الاحتمالية القصوى (MLE) للمعلمات المثلى. يتم تقييم أداء المقدر الجديد بدقة مقارنة بالمقدرات الموجودة باستخدام MSE كمعيار، مما يظهر تفوقه من حيث انخفاض MSE وزيادة النسبة المئوية للكفاءة النسبية (PRE) من خلال كل من البيانات التجريبية ودراسات المحاكاة.
تستخدم الدراسة ثلاث مجموعات بيانات حقيقية وتقوم بإجراء محاكاة تشمل تجمعات تم إنشاؤها من توزيع طبيعي، مع أحجام عينات تبلغ 50 و150 و250. تشير النتائج إلى أن المقدر المقترح يتفوق باستمرار على الطرق الموجودة، محققًا الحد الأدنى من MSE والحد الأقصى من PRE. تدعم النتائج بأرقام عددية مقدمة في جداول، مما يبرز كفاءة المقدر للتطبيقات العملية في تقدير الوسيط السكاني المحدود. يدعو المؤلفون إلى اعتماد مقدراتهم المقترحة في الاستطلاعات المستقبلية ويقترحون أن المنهجية يمكن توسيعها لتقدير المتوسطات السكانية تحت أطر أخذ العينات الطبقية والعشوائية النظامية.
الطرق
في هذا القسم، يحدد المؤلفون المنهجية المستخدمة في دراستهم، مع التركيز على تجمع سكاني يُشار إليه بـ \( W = (W_1, W_2, \ldots, W_N) \)، يتكون من \( N \) وحدة متميزة. يتم تمثيل المتغير المدروس والمتغيرات المساعدة بـ \( Y_i \) و \( X_i \)، على التوالي، للوحدة \( i \). يتم سحب عينة بحجم \( n \) من التجمع السكاني \( W \)، ويتم الإشارة إلى الوسيطين لكل من المتغير المدروس والمتغيرات المساعدة بـ \( M_y \) و \( M_x \). يتم تمثيل دوال كثافة الاحتمال المرتبطة بهذه الوسيطات بـ \( f_{y}(M_y) \) و \( f_{x}(M_x) \).
يتم تعريف معامل الارتباط بين الوسيطين \( M_y \) و \( M_x \) كـ \( \rho_{yx} \) ويتم التعبير عنه رياضيًا كـ \( \rho(M_y, M_x) = 4 P_{11} – 1 \)، حيث \( P_{11} \) هو الاحتمال المشترك \( P(Y \leq M_y \cap X \leq M_x) \). توفر هذه الصياغة مقياسًا كميًا للعلاقة بين المتغيرات المدروسة والمساعدة، مما يسهل المزيد من التحليل لتداخلها.
المناقشة
في هذا القسم، يناقش المؤلفون مجموعة من المقدرات الموجودة للوسيط السكاني، مقارنةً بانحيازها ومتوسط مربع الخطأ (MSE) مع مقدر مقترح جديد يتضمن معلومات مساعدة. تشمل المقدرات الموجودة مقدر الوسيط غير المنحاز، ومقدر النسبة، ومقدر النوع النسبة الأسية، والعديد من مقدرات الفرق، كل منها مع تعبيراته الخاصة للانحياز وMSE. يبرز المؤلفون القيم المثلى للمعلمات التي تقلل من MSE لهذه المقدرات، مقدّمين صياغات رياضية مفصلة.
المقدر المقترح، المشار إليه بـ \( M^*_y \)، مصمم لتعزيز الكفاءة من خلال استخدام المتغيرات المساعدة في سياق أخذ العينات العشوائية البسيطة. يستخرج المؤلفون الانحياز وMSE لهذا المقدر، مما يظهر تحسين تكيفه وكفاءته مقارنة بالطرق الموجودة. تكشف الدراسات التجريبية باستخدام مجموعات بيانات فعلية ودراسات المحاكاة مع تجمعات تم إنشاؤها أن المقدر المقترح يحقق باستمرار MSE أقل وكفاءة نسبية أعلى. تشير النتائج إلى أن المقدر المقترح هو الأفضل للتطبيقات العملية في تقدير الوسيطات السكانية، ويوصي المؤلفون باستخدامه في الاستطلاعات المستقبلية، مع إمكانية التوسع في تقدير المتوسطات السكانية تحت طرق أخذ عينات مختلفة.
DOI: https://doi.org/10.1038/s41598-024-67189-1
PMID: https://pubmed.ncbi.nlm.nih.gov/39019932
Publication Date: 2024-07-17
Author(s): Sohaib Ahmad et al.
Primary Topic: Survey Sampling and Estimation Techniques
Overview
This article presents an improved estimator for the population median utilizing auxiliary information within the framework of simple random sampling. The authors derive expressions for the bias and mean square error (MSE) of the proposed estimator up to first-order approximation and identify the maximum likelihood estimates (MLE) of the optimal parameters. The performance of the new estimator is rigorously evaluated against existing estimators using MSE as a benchmark, demonstrating its superiority in terms of lower MSE and higher percentage relative efficiency (PRE) through both empirical data and simulation studies.
The study employs three real datasets and conducts a simulation involving populations generated from a normal distribution, with sample sizes of 50, 150, and 250. Results indicate that the suggested estimator consistently outperforms existing methods, achieving the minimum MSE and maximum PRE. The findings are substantiated by numerical results presented in tables, highlighting the estimator’s efficiency for practical applications in estimating the finite population median. The authors advocate for the adoption of their proposed estimators in future surveys and suggest that the methodology can be extended to estimate population means under stratified and systematic random sampling frameworks.
Methods
In this section, the authors outline the methodology employed in their study, focusing on a population denoted as \( W = (W_1, W_2, \ldots, W_N) \), consisting of \( N \) distinct units. The study variable and auxiliary variables are represented by \( Y_i \) and \( X_i \), respectively, for the \( i \)-th unit. A sample of size \( n \) is drawn from the population \( W \), and the medians of both the study and auxiliary variables are denoted as \( M_y \) and \( M_x \). The probability density functions associated with these medians are represented as \( f_{y}(M_y) \) and \( f_{x}(M_x) \).
The correlation coefficient between the medians \( M_y \) and \( M_x \) is defined as \( \rho_{yx} \) and is mathematically expressed as \( \rho(M_y, M_x) = 4 P_{11} – 1 \), where \( P_{11} \) is the joint probability \( P(Y \leq M_y \cap X \leq M_x) \). This formulation provides a quantitative measure of the relationship between the study and auxiliary variables, facilitating further analysis of their interdependence.
Discussion
In this section, the authors discuss various existing estimators for population median, comparing their bias and mean square error (MSE) with a newly proposed estimator that incorporates auxiliary information. The existing estimators include the unbiased median estimator, ratio estimator, exponential ratio-type estimator, and several difference-type estimators, each with their respective expressions for bias and MSE. The authors highlight the optimal values for parameters that minimize the MSE for these estimators, providing detailed mathematical formulations.
The proposed estimator, denoted as \( M^*_y \), is designed to enhance efficiency by utilizing auxiliary variables in a simple random sampling context. The authors derive the bias and MSE for this estimator, demonstrating its improved adaptability and efficiency compared to existing methods. Empirical studies using actual datasets and simulation studies with generated populations reveal that the proposed estimator consistently achieves lower MSE and higher percentage relative efficiency. The findings suggest that the proposed estimator is superior for practical applications in estimating population medians, and the authors recommend its use in future surveys, with potential extensions to estimating population means under different sampling methods.
