Challenges with Dual Use of Auxiliary Variable: An Extended Regression Estimator for Finite Population Mean
摘要
In survey sampling, regression estimators based on auxiliary information are widely used to improve the efficiency of finite population mean estimation. While recent studies have advocated the dual use of auxiliary variables through rank and empirical distribution function (EDF) transformations, our findings indicate that such approaches do not yield meaningful efficiency gains. In this paper, we propose an extended regression estimator that incorporates X, its rank, and EDF, with several special cases formulated to systematically evaluate the contribution of these transformations. Expressions for the mean squared errors (MSEs) are derived, and extensive simulations are conducted under symmetric (Normal), skewed (Gamma), and heavy-tailed (Log-normal) populations across high, medium, and low correlation regimes. Additional validation is provided using five real population datasets. Both theoretical and empirical results consistently show that estimators using rank or EDF transformations perform nearly identically to the benchmark regression estimator based solely on X, with percent relative efficiencies (PREs) remaining narrowly around 100%. Any modest deviations observed in practice are best regarded as random fluctuations rather than systematic improvements. These findings demonstrate that rank and EDF transformations, being monotonic functions of X, are redundant in efficiency terms and add unnecessary complexity. The broader implication is that parsimonious regression estimators based on the original auxiliary variable remain the most efficient and practical choice in survey sampling, with rank and EDF transformations offering, at best, limited robustness in special cases involving outliers or non-linear associations.