Experimental Design and Hyperoptimization
摘要
The model-building process has thus far been conceptualized as a linear and sequential series of steps. Initially, an ABT was constructed with explanatory and outcome variables. Subsequently, the ABT was partitioned into a training (treatment) and a validation (control) partition. The variable selection process was initiated, beginning with the univariable procedure and subsequently followed by the collinearity reduction method. The original variables were then replaced with their corresponding WoEs, and a multivariable selection procedure was executed, resulting in the initial multivariable model. The objective is to demonstrate that the final multivariable model represents the optimal solution within a feasible region, constrained by a set of predetermined conditions that were defined concurrently with the modeling process. For instance, it is plausible that the subset of variables incorporated into the final model would have been different had the varclus correlation method been selected in place of the ICRM. Furthermore, alternative grouping options in the IG node may result in a different selection of variables (due to the influence of the rejection rate on the stringency of the grouping constraints) and may also increase the discriminatory power of some variables over others, simply by selecting a better cutoff point during the grouping process. Consequently, the feasible region within which the optimal model is contained is defined by the specific set of options, whether methods, parameters, or values, selected during the model-building process. It can thus be concluded that there must be a different optimal solution for each possible feasible region. The process of defining the number of feasible regions to explore is known as experimental design, while the process of defining the parameters that constrain each feasible region is known as hyperoptimization.