الاستدلال بعد الاختيار في تحليل الكون المتعدد (PIMA): إطار استدلالي قائم على اختبار درجة قلب الإشارة
Post-selection Inference in Multiverse Analysis (PIMA): An Inferential Framework Based on the Sign Flipping Score Test

شارك:
المجلة: Psychometrika، المجلد: 89، العدد: 2
DOI: https://doi.org/10.1007/s11336-024-09973-6
PMID: https://pubmed.ncbi.nlm.nih.gov/38664342
تاريخ النشر: 2024-04-25
المؤلف: Paolo Girardi وآخرون
الموضوع الرئيسي: طرق إحصائية في التجارب السريرية

نظرة عامة

في هذه الورقة، يقدم المؤلفون نهج الاستدلال بعد الاختيار لتحليل الكون المتعدد (PIMA)، معالجين القيود المفروضة على الأساليب الحالية في تحليل الكون المتعدد وتحليل منحنى التحديد. بينما قدمت هذه الأساليب السابقة رؤى وصفية بشكل أساسي، يوفر PIMA إطارًا استدلاليًا قويًا يستوعب مجموعة واسعة من مواصفات البيانات والنماذج الخطية العامة. من خلال استخدام إجراء إعادة أخذ عينات مشروط، يمكّن PIMA الباحثين من اختبار الفرضية الصفرية بشأن العلاقة بين المتنبئين والنتائج، مع الحفاظ على تحكم قوي في معدل الخطأ العائلي (FWER). وهذا يسمح بتحديد التأثيرات المهمة عبر نماذج مختلفة ضمن إطار الكون المتعدد.

يظهر المؤلفون تطبيق PIMA من خلال دراسة حالة تحلل التردد الذاتي تجاه لقاح COVID-19 في إيطاليا، مما يبرز فائدته العملية. لا يسهل PIMA فقط التحكم الضعيف والقوي في FWER ولكن أيضًا يقدم حدًا أدنى من الثقة لنسبة الاكتشافات الحقيقية بين النماذج المختبرة. يؤكد الإطار على أهمية تأصيل التحليلات في نموذج نظري محدد جيدًا، خاصة في مجالات مثل علم الأوبئة، حيث تكون مواصفات المتغيرات المربكة حاسمة. بشكل عام، يمثل PIMA تقدمًا كبيرًا في تحليل الكون المتعدد، حيث يوفر للباحثين أداة شاملة لاتخاذ قرارات استدلالية مستنيرة مع الاعتراف بالتعقيدات الكامنة في تحليل البيانات.

مقدمة

تناقش مقدمة هذه الورقة البحثية التعقيدات والتحديات الكامنة في تحليل البيانات في العالم الحقيقي، مسلطة الضوء على تعدد الخيارات القابلة للدفاع التي يواجهها الباحثون في كل خطوة، من تحويل البيانات إلى اختيار النموذج. تؤدي هذه المرونة إلى “كون متعدد” من مجموعات البيانات والنتائج الإحصائية، مما يثير المخاوف بشأن ممارسات مثل التلاعب بالـ p، حيث قد يقوم الباحثون بالإبلاغ الانتقائي عن النتائج المهمة لتعزيز فرص النشر. للتخفيف من هذه القضايا، يتم اقتراح حلين رئيسيين: التسجيل المسبق لخطط التحليل واستخدام تحليل الكون المتعدد، الذي يتضمن استكشاف جميع خيارات معالجة البيانات المعقولة وتحليل كل مجموعة بيانات لتقييم متانة النتائج.

تقدم الورقة نهج الاستدلال بعد الاختيار لتحليل الكون المتعدد (PIMA)، وهي طريقة جديدة مصممة لتوفير استدلالات إحصائية موثوقة ضمن إطار الكون المتعدد. يسمح PIMA للباحثين باختبار الفرضيات المتعلقة بعلاقة المتنبئين بالنتائج مع التحكم في معدلات الخطأ من النوع الأول، حتى بعد اختيار النماذج بناءً على النتائج الملاحظة. تقدم الطريقة قيمة p عالمية عبر جميع النماذج، وقيم p معدلة فردية لنماذج محددة، وحدًا أدنى من الثقة لنسبة الاكتشاف الحقيقي (TDP). هذا النهج المرن قابل للتطبيق في سياقات متنوعة ويهدف إلى تعزيز الشفافية والمصداقية للنتائج العلمية مع معالجة أزمة القابلية للتكرار في علم النفس الكمي. توضح الورقة هيكل الأقسام التالية، التي تفصل الإطار وإجراءات الاختبار وتطبيقات PIMA.

النتائج

في هذا القسم، يقدم المؤلفون نتائج من نموذج ثنائي الحدين بارامتري يستخدم متنبئين خطيين واثنين من متغيرات التباين المركزية الصفرية لتحليل متغير زمني من ثلاثة مستويات. يشمل إطار تحليل الكون المتعدد 243 نموذجًا مشتقة من ثلاثة تحويلات لكل متنبئ مستمر عبر ثلاثة مقارنات: قبل الإغلاق، قبل-بعد، والإغلاق-بعد. تشير النتائج الوصفية إلى أن مجموعات الاختبارات تنتج قيم p أقل من 0.05؛ ومع ذلك، يحذر المؤلفون من الادعاء بالأهمية الإحصائية بسبب احتمال ارتفاع معدلات الإيجابيات الكاذبة.

يكشف التحليل الاستدلالي عن نتائج مهمة لمقارنات ما بعد-قبل وما بعد-الإغلاق عبر 81 نموذجًا، بينما لا تصل مقارنة الإغلاق-قبل إلى الأهمية. يؤكد المؤلفون أن تصحيح ما بعد الاختبار يأخذ في الاعتبار المقارنات الثلاثة بناءً على النماذج المجمعة. لتحسين اختيار النموذج، يقومون بتعديل تصحيح التعددية على مستوى النموذج الفردي، كما هو موضح في الشكل 8، مع تقديم نتائج مفصلة في المواد التكميلية. أخيرًا، تلخص الجدول 5 عدد الاكتشافات الحقيقية ونسبة الاكتشاف الحقيقي (TDP) لكل مقارنة، مشيرة إلى أن جميع النماذج تظهر اختلافات مهمة لمقارنة ما بعد-قبل بعد تطبيق تصحيحات التعددية.

المناقشة

في هذا القسم، يناقش المؤلفون تنفيذ اختبار درجة الانعكاس في سياق تحليل الكون المتعدد، حيث يتم تقييم نماذج متعددة لتقييم تأثير متنبئ على متغير استجابة. يوفر اختبار درجة الانعكاس، كما طوره هيميريك وآخرون (2020) ودي سانتيس وآخرون (2022)، إطارًا قويًا لاختبار الفرضية الصفرية العالمية التي تفيد بأن المتنبئ ليس له تأثير عبر مواصفات نماذج مختلفة. يؤكد المؤلفون على التحديات التي تطرحها الاعتماديات بين إحصائيات الاختبار الأحادي في بيئة متعددة المتغيرات ويقترحون نهجًا قائمًا على إعادة أخذ العينات لاشتقاق إحصائية اختبار دقيقة تقاربيًا للفرضية العالمية.

يحدد المؤلفون هيكل النموذج باستخدام النماذج الخطية العامة (GLMs) وي outline إطار اختبار الفرضيات للنماذج الفردية باستخدام اختبار درجة الانعكاس. يقدمون الصيغ الرياضية اللازمة، بما في ذلك الدرجة الفعالة ومساهماتها، ويحددون الشروط التي بموجبها يحافظ الاختبار على التحكم في معدلات الخطأ من النوع الأول. يختتم القسم بمناقشة كيفية توسيع اختبار درجة الانعكاس ليشمل إطار الكون المتعدد، مما يسمح بجمع النتائج من نماذج متعددة لاشتقاق إحصائية اختبار عالمية. لا يسهل هذا النهج اختبار الفرضيات عبر مواصفات نماذج مختلفة فحسب، بل يمكّن أيضًا من الاستدلال بعد الاختيار، مما يوفر قيم p المعدلة وحدود الثقة لنسبة النماذج التي يكون فيها المتنبئ مهمًا، وبالتالي معالجة القضايا المحتملة للاستدلال الانتقائي في البحث.

Journal: Psychometrika, Volume: 89, Issue: 2
DOI: https://doi.org/10.1007/s11336-024-09973-6
PMID: https://pubmed.ncbi.nlm.nih.gov/38664342
Publication Date: 2024-04-25
Author(s): Paolo Girardi et al.
Primary Topic: Statistical Methods in Clinical Trials

Overview

In this paper, the authors introduce the Post-selection Inference approach to Multiverse Analysis (PIMA), addressing the limitations of existing methods in multiverse analysis and specification curve analysis. While these previous approaches primarily offered descriptive insights, PIMA provides a robust inferential framework that accommodates a wide range of data specifications and generalized linear models. By employing a conditional resampling procedure, PIMA enables researchers to test the null hypothesis regarding the association between predictors and outcomes, while maintaining strong control over the family-wise error rate (FWER). This allows for the identification of significant effects across various models within the multiverse framework.

The authors demonstrate the application of PIMA through a case study analyzing self-reported hesitancy towards the COVID-19 vaccine in Italy, highlighting its practical utility. PIMA not only facilitates weak and strong control of the FWER but also offers a lower confidence bound for the proportion of true discoveries among tested models. The framework emphasizes the importance of grounding analyses in a well-defined theoretical model, particularly in fields like epidemiology, where the specification of confounders is critical. Overall, PIMA represents a significant advancement in multiverse analysis, providing researchers with a comprehensive tool for making informed inferential decisions while acknowledging the complexities inherent in data analysis.

Introduction

The introduction of this research paper discusses the complexities and challenges inherent in real-world data analysis, highlighting the multitude of defensible choices researchers face at each step, from data transformation to model selection. This flexibility leads to a “multiverse” of datasets and statistical results, raising concerns about practices such as p-hacking, where researchers may selectively report significant findings to enhance publication chances. To mitigate these issues, two primary solutions are proposed: preregistration of analysis plans and the use of multiverse analysis, which involves exploring all reasonable data processing options and analyzing each dataset to assess the robustness of results.

The paper introduces the Post-selection Inference approach to Multiverse Analysis (PIMA), a novel method designed to provide reliable statistical inferences within the multiverse framework. PIMA allows researchers to test hypotheses regarding the association of predictors with outcomes while controlling for Type I error rates, even after selecting models based on observed results. The method offers a global p-value across all models, individual adjusted p-values for specific models, and a lower confidence bound for the true discovery proportion (TDP). This flexible approach is applicable to various contexts and aims to enhance the transparency and credibility of scientific findings while addressing the replicability crisis in quantitative psychology. The paper outlines the structure of the subsequent sections, which detail the framework, testing procedures, and applications of PIMA.

Results

In this section, the authors present findings from a parametric binomial model utilizing linear predictors and two zero-centered contrast variables to analyze a three-level Period variable. The multiverse analysis framework encompasses 243 models derived from three transformations of each continuous predictor across three comparisons: Pre-Lockdown, Pre-Post, and Lockdown-Post. Descriptive results indicate that clusters of tests yield p-values below 0.05; however, the authors caution against claiming statistical significance due to potential high false positive rates.

Inferential analysis reveals significant results for the Post-Pre and Post-Lockdown comparisons across the 81 models, while the Lockdown-Pre comparison does not reach significance. The authors emphasize that the post hoc correction accounts for the three comparisons based on the combined models. To refine model selection, they adjust multiplicity correction at the individual model level, as illustrated in Figure 8, with detailed results provided in supplementary materials. Finally, Table 5 summarizes the number of true discoveries and the True Discovery Proportion (TDP) for each comparison, indicating that all models show significant differences for the Post-Pre comparison after applying multiplicity corrections.

Discussion

In this section, the authors discuss the implementation of the sign-flip score test within the context of multiverse analysis, where multiple models are evaluated to assess the effect of a predictor on a response variable. The sign-flip score test, as developed by Hemerik et al. (2020) and De Santis et al. (2022), provides a robust framework for testing the global null hypothesis that a predictor has no effect across various model specifications. The authors emphasize the challenges posed by the dependencies among univariate test statistics in a multivariate setting and propose a resampling-based approach to derive an asymptotically exact test statistic for the global hypothesis.

The authors specify the model structure using Generalized Linear Models (GLMs) and outline the hypothesis testing framework for individual models using the sign-flip score test. They introduce the necessary mathematical formulations, including the effective score and its contributions, and establish the conditions under which the test maintains control over Type I error rates. The section culminates in a discussion of how the sign-flip score test can be extended to a multiverse framework, allowing for the combination of results from multiple models to derive a global test statistic. This approach not only facilitates hypothesis testing across various model specifications but also enables post-selection inference, providing adjusted p-values and confidence bounds for the proportion of models where the predictor is significant, thereby addressing potential issues of selective inference in research.

شارك: