Compressive Sensing
摘要
In the high-dimensional setting, we’re essentially looking at the same linear model \(Y = \mathbb {X} \beta + \varepsilon \) with \(\mathbb {X} \in \mathbb {R}^{n \times d}\) . However, we now expect the number of features d is much larger than its sample size n. Under the high-dimensional setting, the ordinary least squares estimator will have troubles. If the features are linearly independent, we have \(\mathrm {rank} (\mathbb {X}) = n\) . Then, \(\mathbb {X} \widehat {\beta }^{\mathrm {LS}} = P_{\mathbb {X}} Y = Y\) , i.e., the ordinary least squares will overfit. Therefore, we need to invoke the parsimonious principle and introduce the following sparse linear model.