Using principal components as covariates in propensity score weighting: a simulation study
摘要
Propensity score weighting (PSW) is widely employed to reduce confounding bias in observational studies. The specification of the propensity score (PS) model—particularly the selection of covariates—significantly affects the accuracy of treatment effect estimates. Although conventional guidance recommends an “all-inclusive” approach that incorporates all available covariates, this strategy can be impractical in high-dimensional datasets and may introduce computational problems such as multicollinearity and complete or quasi-complete separation. This study introduces and evaluates three alternative strategies that leverage principal component analysis (PCA) to mitigate these challenges.
MethodsA Monte Carlo simulation was conducted to compare the traditional all-inclusive approach against three PCA-based covariate selection strategies in terms of model stability and bias reduction.
ResultsWe found that PCA-based strategies consistently produced more stable PS models and more accurate, less-biased treatment effect estimates—particularly when the number of correlated covariates was large. Specifically, incorporating the minimum set of principal components needed to explain approximately 40% of total covariate variance provided the optimal trade-off between dimensionality reduction and confounding adjustment.
ConclusionsThe study findings highlight PCA as a practical and computationally efficient solution for improving PS estimation in complex observational studies. Effective covariate selection for PS models requires identifying variables that are associated with both treatment assignment and the outcome.