Abstract
Many empirical analyses involve performing multiple statistical tests or analytic strategies and subsequently reporting the most favourable result. Such data-dependent selection produces model selection bias. Here, we provide a causal framework for understanding data-dependent selection procedures as producing bias similar to that which arises from collider stratification. We conceptualize the dataset as a single realization of an N-dimensional random vector, rather than N independent samples. Consequently, estimators, test statistics, P-values, and other derived quantities are random variables generated by the sampling process. Conditioning on a selection criterion based on such quantities is therefore conditioning on a post-data variable, which can induce associations between variables in the analytic sample that were independent in the underlying data-generating process. We illustrate this framework using directed acyclic graphs across several settings, including selective reporting of treatment effects, subgroup analyses, interim stopping decisions, genome-wide association studies, and Mendelian randomization. For example, selecting instrumental variables based on their observed association with an exposure not only exaggerates estimated instrument strength, but can also induce associations between the selected variables and unmeasured confounders, violating exchangeability assumptions. This framework provides a unified interpretation of data-driven analytic practices, including selective reporting, variable selection, and P-hacking.
About the Speaker
Nasir is a Wellcome Trust PhD Fellow at the University of Cambridge, MRC Biostatistics Unit. His work centres around the development and application of methods for making robust causal inferences from large-scale data. He is currently on a grant as a visiting researcher at the HKU Faculty of Dentistry, where he is applying novel causal disparity decomposition methods to oral health datasets.
