Current terrain segmentation research is limited to fixed modalities and closed label sets. This paper proposes the AMTerrain method, which combines arbitrary modality processing with open vocabulary learning capabilities. We use a Unified Coding structure to encode arbitrary modality data and introduce two groups of learnable tokens, Score Embeds and Modal Embeds, to achieve unified representation of arbitrary modality data. Meanwhile, we use Customized vision-prompts generated by DeepSeek-R1 to enhance the model’s open vocabulary learning ability. The number of parameters of this method is comparable to that of the method with the fewest parameters, and its performance has been improved by 1.02%, 1.86%, and 2.58% respectively under three experimental settings: all-modality, modality-agnostic, and open-vocabulary. Experiments show that our model can fully exploit the value of each modality, has strong adaptability to arbitrary modality inputs, and possesses certain open vocabulary learning capabilities.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

AMTerrain: Research on Arbitrary-Modal Terrain Segmentation Based on Text Guidance

  • Yuqian Wang,
  • Xuefu Xiang,
  • Yongcun Wu,
  • Peng Quan,
  • Yunzhi Luo,
  • Lu Zhang

摘要

Current terrain segmentation research is limited to fixed modalities and closed label sets. This paper proposes the AMTerrain method, which combines arbitrary modality processing with open vocabulary learning capabilities. We use a Unified Coding structure to encode arbitrary modality data and introduce two groups of learnable tokens, Score Embeds and Modal Embeds, to achieve unified representation of arbitrary modality data. Meanwhile, we use Customized vision-prompts generated by DeepSeek-R1 to enhance the model’s open vocabulary learning ability. The number of parameters of this method is comparable to that of the method with the fewest parameters, and its performance has been improved by 1.02%, 1.86%, and 2.58% respectively under three experimental settings: all-modality, modality-agnostic, and open-vocabulary. Experiments show that our model can fully exploit the value of each modality, has strong adaptability to arbitrary modality inputs, and possesses certain open vocabulary learning capabilities.