This paper considers the optimization problem \(\begin{aligned} \min _{X \in \mathcal {F}_v} f( {X}) + \lambda \Vert X\Vert _1, \end{aligned}\) where f is smooth, \(\mathcal {F}_v = \{X \in \mathbb {R}^{n \times q}: X^T X = I_q, v \in \textrm{span}(X)\}\) , and v is a given positive vector. The clustering models including but not limited to the models used by k-means, community detection, and normalized cut can be reformulated as such optimization problems. It is proven that the domain \(\mathcal {F}_v\) forms a compact embedded submanifold of \(\mathbb {R}^{n \times q}\) and optimization-related tools including a family of computationally efficient retractions and an orthonormal basis of any normal space of \(\mathcal {F}_v\) are derived. A Riemannian proximal gradient method that allows an adaptive step size is proposed. The proposed Riemannian proximal gradient method solves its subproblem inexactly and still guarantees its global convergence. Numerical experiments on community detection in networks and normalized cut for image segmentation are used to demonstrate the performance of the proposed method.