Selective guided domain adaptation for improving the proxy means test for poverty targeting
摘要
The Proxy Means Test (PMT) is widely used for poverty targeting, estimating household per capita consumption expenditure based on survey data from a prior year. However, PMT suffers from high estimation errors, as well as substantial inclusion and exclusion errors. We identify three key challenges affecting PMT performance: small dataset sizes leading to “data gaps”, the limited predicted power of collected household variables, and distribution changes. In our previous work, we compared existing statistical methods used in the PMT with a strong baseline machine learning method (XGBoost) and proposed a new domain adaptation method (TrAdaBoost.JQL) for the PMT, an approach that uses data from previous years (sources) to predict the current year’s data (target) using a small subset of households (guide) from the current year. However, TrAdaBoost.JQL works under the simplifying assumption that the true quantiles of guide instances are known. In this paper, we relax this assumption and introduce selective guided domain adaptation (SGDA), a novel method for systematically selecting a small subset of target instances to jointly minimise estimation error and inclusion and exclusion errors, where the true quantiles of the guide instances are not known. We empirically validate SGDA across multiple semi-urban and urban districts in Indonesia, showing that use of the selectively chosen guide improves a range of domain adaptation methods on the PMT, and that selective guided domain adaptation outperforms domain adaptation with randomly chosen guides.