Federated propensity score matching: a method for estimating average treatment effect based on data in Islands
摘要
Propensity Score Matching (PSM) is a popular method for estimating Average Treatment Effect (ATE). However, the growing focus on data security has led to the enactment of laws and regulations that hinder data replication, which posed a major obstacle to constructing large-scale datasets to implement matching, resulting in the emergence of “data islands”. Consequently, conducting matching based on data within these islands has become a valuable research topic with significant theoretical and practical importance. In this paper, a Federated Propensity Score Matching (FPSM) method is proposed for estimating ATEs based on data in islands. In the method, a Federated Random Forest (FRF) algorithm is first proposed to estimate the propensity score of samples stored in islands. Then, according to the propensity scores, the nearest-neighbor matching algorithm is extended into Federated Learning (FL) framework to construct the matching sets. Furthermore, the quality of the constructed matching sets is evaluated, and the ATE is estimated under FL framework without exposing private information. Moreover, we extend the proposed FPSM to solve the causal inference problems considering dynamic panel data and propose Federated Propensity Score Matching-Difference in Differences (FPSM-DID). Finally, the experiments are conducted based on three datasets, which illustrate the feasibility and validity of the proposed methods.