إطار تقدير لدمج العينات الاحتمالية وغير الاحتمالية
An Estimation Framework for Combining Probability and Non-probability Samples

شارك:
المجلة: Journal of Data Science
DOI: https://doi.org/10.6339/26-jds1234
تاريخ النشر: 2026-01-01
المؤلف: Mahmoud Elkasabi وآخرون
الموضوع الرئيسي: طرق إحصائية واستدلال بايزي

نظرة عامة

تتناول الورقة البحثية زيادة اعتماد تصاميم العينة الهجينة في أبحاث الاستطلاع، لا سيما لدراسة السكان النادرين أو صعب الوصول إليهم. تواجه عينات الاحتمالية التقليدية تحديات مثل التكاليف العالية والكفاءة المنخفضة، مما يجعل الأساليب الهجينة أكثر عملية. يستخدم المؤلفون بيانات من المسح الوطني للصحة والتغذية (NHANES) لتوضيح إطار عمل لإجراء تقديرات سكانية من خلال دمج البيانات من كل من عينة احتمالية متعددة المراحل وعينة غير احتمالية مأخوذة من نفس وحدات العينة الأولية (PSUs). يقدمون طرق تحليلية للتعامل مع مثل هذه العينات الهجينة، مع التأكيد على أهمية المتغيرات المرافقة ونتائج الاستطلاع الملاحظة في كلا نوعي العينة.

في الخاتمة، يقدم المؤلفون دراسة حالة توضح التطبيق العملي لإطار العمل المقترح. يظهرون أن التوظيف المستهدف غير الاحتمالي يمكن أن يعالج بشكل فعال أوجه القصور في التغطية، مثل التمثيل الناقص للشباب غير المواطنين في العينة الاحتمالية، مع الحفاظ على التوافق مع تصميم الاستطلاع الأصلي. تشير النتائج إلى أن التعديلات المناسبة في الوزن يمكن أن تقلل بشكل كبير من تحيزات الاختيار وتحسن الدقة، كما يتضح من فترات الثقة الضيقة دون إدخال تحيز قابل للاكتشاف. تسلط الدراسة الضوء على الحاجة إلى مزيد من البحث حول تقدير التباين واستراتيجيات التوظيف المثلى للعينات غير الاحتمالية ضمن التصاميم الهجينة، مما يشير إلى أن العمل المستقبلي يمكن أن يعزز الأسس المنهجية الموضوعة في هذه المقالة.

مقدمة

في مقدمة هذه الورقة، يتناول المؤلفون التحديات التي يواجهها باحثو الاستطلاعات عند دراسة السكان النادرين، مشيرين إلى قضايا مثل التكاليف العالية للفحص وانخفاض القوة الإحصائية بسبب أحجام العينات الصغيرة. يقترحون إطار عمل هجيني يدمج البيانات من كل من العينات الاحتمالية وغير الاحتمالية لتعزيز دقة التقديرات لهذه المجموعات الممثلة تمثيلاً ناقصًا. يجمع هذا النهج بين الصرامة الإحصائية لعينات الاحتمالية مع كفاءة الأساليب غير الاحتمالية، مثل العينات المريحة وعينات كرة الثلج.

يتم هيكلة الإطار في مرحلتين رئيسيتين: مرحلة التقدير، التي تركز على اشتقاق التقديرات من العينات غير الاحتمالية ودمجها مع العينات الاحتمالية باستخدام تقنيات إحصائية متنوعة، ومرحلة التقييم، التي تقيم أداء هذه الأساليب. يؤكد المؤلفون على التطبيق العملي لإطارهم من خلال توفير كود R لتوليد أوزان الاستطلاع وتقنيات التقدير الأخرى. تهدف الورقة إلى تقديم سير عمل تشغيلي واضح للممارسين، مع ترك القضايا الأكثر تعقيدًا، مثل ملاءمة النموذج وتقدير التباين، للبحث المستقبلي. ستتناول الأقسام اللاحقة هيكل البيانات، وطرق دمج العينات، ومقاييس تقييم الأداء، ودراسة حالة لتوضيح تطبيق الإطار.

الطرق

في هذا القسم، يحدد المؤلفون طرق دمج البيانات غير الاحتمالية مع البيانات المعتمدة على الاحتمالية، مع التأكيد على أهمية تقييم دقة التقديرات المدمجة. تتضمن عملية التقييم مقارنة هذه التقديرات مع المعايير المرجعية، التي يجب أن تكون مرتبطة ارتباطًا وثيقًا بالمتغيرات الرئيسية ذات الاهتمام. على سبيل المثال، عند تقييم المتغيرات المتعلقة بالصحة، يُوصى بمعايير مثل سلوك التدخين وضغط الدم. يتم اقتراح عدة مقاييس لهذا التقييم، بما في ذلك الفرق بين التقدير المركب ($\hat{Y}_{com}$) والمعيار المرجعي ($Y_R$)، والفرق المطلق، والتحيز النسبي، ومؤشرات تداخل فترات الثقة، مما يسمح بتقييم شامل لدقة وموثوقية التقديرات المدمجة.

كما يصف المؤلفون ثلاث طرق لحساب أوزان الاستطلاع لعينة B، باستخدام مجموعة بيانات NHANES كمرجع. تشمل هذه الطرق وزن الاحتمالية العكسية (IPW) بناءً على ستة متغيرات مساعدة، والمعايرة إلى إجماليات السكان لهذه المتغيرات، وإصدار معاير من IPW. يتم تبرير اختيار NHANES على عينة A من خلال التغطية غير المكتملة للأشخاص غير المواطنين الشباب. بالإضافة إلى ذلك، يدمج المؤلفون مؤشرات الطبقات والعناقيد لأخذ التداخلات التصميمية في الاعتبار، ويلاحظون إمكانية اشتقاق أوزان تحليل غير صفرية للعينة غير الاحتمالية، والتي يمكن تعديلها لإنشاء أوزان زائفة.

النتائج

تشير النتائج إلى وجود اختلافات ديموغرافية كبيرة بين عينتين، A و B، مقارنة ببيانات NHANES. تتكون عينة A، حسب التصميم، من عدد أقل من الأفراد الشباب (تتراوح أعمارهم بين 18-34 عامًا) ونسبة أعلى من البالغين الأكبر سنًا (الأعمار 35-49، 50-64، و65 عامًا فأكثر). تتميز هذه العينة بانخفاض احتمال أن تكون من أصل إسباني أو مكسيكي أمريكي، وأن تكون حاصلة على شهادة الثانوية العامة أو أعلى، وأن يكون لديها أطفال صغار في المنزل. على العكس، تميل عينة B نحو الأفراد الأصغر سنًا وتضم نسبة أعلى من أولئك من خلفيات عرقية متنوعة، والأفراد الذين لديهم أقل من تعليم الثانوية العامة، والأسر التي لديها أطفال صغار. ومن الجدير بالذكر أن عينة B تُظهر أيضًا انتشارًا أقل للذكور، والأفراد المتزوجين، والبيض غير الإسبان.

يكشف التحليل أن تقديرات وزن الاحتمالية العكسية (IPW) من عينة B تتماشى بشكل أقرب مع تقديرات NHANES مقارنة بالتقديرات غير الموزونة. تقديرات IPW المعايرة من عينة B متسقة بشكل خاص مع NHANES، حيث تم تعديلها باستخدام المتغيرات المساعدة ذات الصلة. من المهم، بينما لا تختلف تقديرات النقطة من العينات A و B والعينة المدمجة إحصائيًا عن تقديرات NHANES، تُظهر العينة المدمجة فترات ثقة أضيق، مما يعزز الدقة دون إدخال تحيز. وهذا يشير إلى أن دمج العينات يزيد بشكل فعال من حجم العينة التحليلية، مما يحسن كفاءة التقديرات مع الحفاظ على قابلية المقارنة مع معيار NHANES. تتوفر مقاييس أداء إضافية في المواد التكميلية.

المناقشة

في هذا القسم، يناقش المؤلفون إطار عمل لدمج البيانات يجمع بين عينة احتمالية متعددة المراحل مصنفة (عينة A) مع عينة غير احتمالية (عينة B) لتعزيز تقدير النتائج لسكان نادرين مستهدفين. يستفيد الدمج من المتغيرات المساعدة المشتركة \(X\) والمتغيرات الناتجة \(Y\) الموجودة في كلا العينتين، مما يسمح بإنشاء مجموعة بيانات مدمجة مع وزن تحليل جديد \(w_c\). يؤكد المؤلفون على أهمية تطوير أوزان الاستطلاع لعينة B لتسهيل تقديرات إحصائية دقيقة، مشيرين إلى أن الأدبيات الحالية تركز بشكل أساسي على دمج العينات غير الاحتمالية مع العينات الاحتمالية حيث تكون المتغيرات الناتجة حصرية للأولى.

يصنف المؤلفون طرقًا متنوعة لتقدير أوزان الاستطلاع للعينات غير الاحتمالية، بما في ذلك تعديلات درجة الميل (PSA)، وزن المعايرة، وزن النواة، والتقدير الجماعي. يبرزون أنه بينما تتوافق تقنيات PSA والمعايرة مع هيكل بياناتهم، فإن التقدير الجماعي ومطابقة العينات غير مناسبين بسبب افتراضاتهما وقيودهما المتأصلة. تختتم القسم بمناقشة حول دمج البيانات من كلا العينتين، مقترحين مقدرًا مدمجًا \( \hat{Y}_{com} \) يعتمد على عامل مركب \( \alpha \) لتقليل التباين، مع معالجة التحديات المرتبطة بالتحيزات المحتملة في عينة B. يوضح المؤلفون منهجيتهم من خلال دراسة حالة تحاكي عينة مصنفة تستهدف البالغين غير المواطنين، مما يبرز ضرورة دمج البيانات غير الاحتمالية لتحسين التقديرات للفئات الفرعية الممثلة تمثيلاً ناقصًا.

Journal: Journal of Data Science
DOI: https://doi.org/10.6339/26-jds1234
Publication Date: 2026-01-01
Author(s): Mahmoud Elkasabi et al.
Primary Topic: Statistical Methods and Bayesian Inference

Overview

The research paper discusses the increasing adoption of hybrid sampling designs in survey research, particularly for studying rare or hard-to-reach populations. Traditional probability sampling faces challenges such as high costs and low efficiency, making hybrid approaches more practical. The authors utilize data from the National Health and Nutrition Examination Survey (NHANES) to illustrate a framework for making population estimates by integrating data from both a stratified, multistage probability sample and a non-probability sample drawn from the same primary sampling units (PSUs). They provide analytical methods for handling such hybrid samples, emphasizing the importance of covariates and survey outcomes observed in both sample types.

In the conclusion, the authors present a case study that demonstrates the practical application of their proposed framework. They show that targeted nonprobability recruitment can effectively address coverage deficiencies, such as the under-representation of younger non-citizens in the probability sample, while maintaining alignment with the original survey design. The results indicate that appropriate weighting adjustments can significantly reduce selection biases and improve precision, as evidenced by narrower confidence intervals without introducing detectable bias. The study highlights the need for further research on variance estimation and optimal recruitment strategies for nonprobability samples within hybrid designs, suggesting that future work could enhance the methodological foundations laid out in this article.

Introduction

In the introduction of this paper, the authors address the challenges faced by survey researchers when studying rare populations, highlighting issues such as high screening costs and reduced statistical power due to small sample sizes. They propose a hybrid framework that integrates data from both probability and non-probability samples to enhance the estimation precision for these underrepresented groups. This approach combines the statistical rigor of probability sampling with the efficiency of non-probability methods, such as convenience and snowball sampling.

The framework is structured into two main stages: the estimation stage, which focuses on deriving estimates from non-probability samples and integrating them with probability samples using various statistical techniques, and the evaluation stage, which assesses the performance of these methods. The authors emphasize the practical application of their framework by providing R code for generating survey weights and other estimation techniques. The paper aims to offer a clear operational workflow for practitioners, while leaving more complex issues, such as model fit and variance estimation, for future research. Subsequent sections will elaborate on the data structure, methods for combining samples, performance evaluation measures, and a case study to illustrate the framework’s application.

Methods

In this section, the authors outline methods for integrating non-probability data with probability-based data, emphasizing the importance of evaluating the accuracy of combined estimates. The evaluation process involves comparing these estimates against reference benchmarks, which should be closely correlated with the key variables of interest. For instance, when assessing health-related variables, benchmarks such as smoking behavior and blood pressure are recommended. Several measures are proposed for this evaluation, including the difference between the composite estimate ($\hat{Y}_{com}$) and the benchmark ($Y_R$), absolute difference, relative bias, and confidence interval overlap indicators, allowing for a comprehensive assessment of the accuracy and reliability of the combined estimates.

The authors also describe three methods for calculating survey weights for Sample B, utilizing the NHANES dataset as a reference. These methods include Inverse Probability Weighting (IPW) based on six auxiliary variables, calibration to population totals of these variables, and a calibrated version of IPW. The choice of NHANES over Sample A is justified by the latter’s incomplete coverage of young non-citizens. Additionally, the authors incorporate strata and cluster indicators to account for design overlaps, and they note the potential for deriving non-null analysis weights for the non-probability sample, which can be adjusted to create pseudo-weights.

Results

The results indicate significant demographic differences between two samples, A and B, compared to the NHANES data. Sample A, by design, comprises fewer young individuals (aged 18-34 years) and a higher proportion of older adults (ages 35-49, 50-64, and 65 and older). This sample is characterized by a lower likelihood of being Hispanic or Mexican American, having a high school degree or higher, and having young children in the household. Conversely, Sample B is skewed towards younger individuals and includes a higher proportion of those from diverse racial backgrounds, individuals with less than a high school education, and households with young children. Notably, Sample B also shows a lower prevalence of males, married individuals, and non-Hispanic Whites.

The analysis reveals that the Inverse Probability Weighting (IPW) estimates from Sample B align more closely with NHANES estimates than the unweighted estimates. The calibrated IPW estimates from Sample B are particularly consistent with NHANES, as they were adjusted using relevant auxiliary variables. Importantly, while point estimates from Samples A, B, and the combined sample do not statistically differ from NHANES estimates, the combined sample demonstrates narrower confidence intervals, enhancing precision without introducing bias. This suggests that merging the samples effectively increases the analytic sample size, thereby improving the efficiency of the estimates while maintaining comparability with the NHANES benchmark. Additional performance measures are available in the supplementary materials.

Discussion

In this section, the authors discuss a data integration framework that combines a stratified multistage probability sample (Sample A) with a non-probability sample (Sample B) to enhance the estimation of outcomes for a targeted rare population. The integration leverages shared auxiliary variables \(X\) and outcome variables \(Y\) present in both samples, allowing for the generation of a combined dataset with a new analysis weight \(w_c\). The authors emphasize the importance of developing survey weights for Sample B to facilitate accurate statistical estimates, noting that existing literature primarily focuses on integrating non-probability samples with probability samples where outcome variables are exclusive to the former.

The authors categorize various methods for estimating survey weights for non-probability samples, including propensity score adjustments (PSA), calibration weighting, kernel weighting, and mass imputation. They highlight that while PSA and calibration techniques are compatible with their data structure, mass imputation and sample matching are not suitable due to their inherent assumptions and limitations. The section concludes with a discussion on combining data from both samples, proposing a combined estimator \( \hat{Y}_{com} \) that relies on a composite factor \( \alpha \) to minimize variance, while also addressing the challenges associated with potential biases in Sample B. The authors illustrate their methodology through a case study simulating a stratified sample targeting non-citizen adults, underscoring the necessity of integrating non-probability data to improve estimates for underrepresented subpopulations.

شارك: