In the dynamically progressing field of computer vision, the task of Bird’s-Eye-View (BEV), which offers an encompassing top-down perspective by integrating data from various sensors, has become indispensable, particularly for autonomous driving systems. Adapting BEV models to downstream datasets presents a significant challenge due to the scarcity of effective fine-tuning approaches. To surmount these challenges, we propose Few-Shot Visual Prompt (ViPro), a modular Plug-and-Play methodology that facilitates the transfer of models across datasets with minimal data usage, marking a pioneering approach in the BEV sphere. Our methodology employs a two-phase pixel-wise prompting strategy: initially, a subset of source dataset is simpled to train prototype visual prompts that capture the general features of images. Subsequently, the dataset is categorized into clusters to distinguish different scenes, enhancing scene-specific adaptability. This approach is further augmented by token-wise visual prompts within the Cross-View Transformer, bolstering the model’s performance and flexibility while significantly reducing susceptibility to overfitting specific environmental factors.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ViPro-BEV: Few-Shot Visual Prompting for Bird’s-Eye-View Perception

  • Guorong Yuan,
  • Huaibo Huang,
  • Qihang Fan

摘要

In the dynamically progressing field of computer vision, the task of Bird’s-Eye-View (BEV), which offers an encompassing top-down perspective by integrating data from various sensors, has become indispensable, particularly for autonomous driving systems. Adapting BEV models to downstream datasets presents a significant challenge due to the scarcity of effective fine-tuning approaches. To surmount these challenges, we propose Few-Shot Visual Prompt (ViPro), a modular Plug-and-Play methodology that facilitates the transfer of models across datasets with minimal data usage, marking a pioneering approach in the BEV sphere. Our methodology employs a two-phase pixel-wise prompting strategy: initially, a subset of source dataset is simpled to train prototype visual prompts that capture the general features of images. Subsequently, the dataset is categorized into clusters to distinguish different scenes, enhancing scene-specific adaptability. This approach is further augmented by token-wise visual prompts within the Cross-View Transformer, bolstering the model’s performance and flexibility while significantly reducing susceptibility to overfitting specific environmental factors.