Philippe Schlenker’s recent semantic system for music states formal necessary conditions for a musical snippet \(\mathcal M\) to denote a given situation \(\mathcal S\) . According to these conditions, some features of the music and the scene must evolve in a parallel way through time. In this article, I raise the question of the syntax–semantic interface in music, which has not been investigated in previous works. I argue that Schlenker’s “linear” conditions are not sufficient and that the denotation relation also obeys some structural conditions: both \(\mathcal M\) and \(\mathcal S\) exhibit a tree–structure and these two structures must match in a way or another. After investigating original examples showing that structural conditions are needed, I present an assortment of such conditions in the formalism of rooted trees. Some of these conditions constrain the trees \(\mathcal M\) and \(\mathcal S\) in a symmetric way (meaning that \(\mathcal M\) and \(\mathcal S\) play the same formal role), while others are asymmetric, and both possibilities are investigated. Finally, I examine logical and entailment links between the different conditions stated, leaving future corpus and empirical studies to decide which is the most relevant.