<p>Vision-Language Models (VLMs) have emerged as a powerful paradigm for multi-stage robotic manipulation. Although recent frameworks like ReflectVLM have introduced visual reflection via dynamic models, they often suffer from ‘logical stagnation’—repeatedly executing failed actions because they cannot perceive subtle physical blocks. This paper proposes ReflectVLM+, an advanced framework that integrates quantitative stagnation detection and data refinement techniques to overcome these limitations. First, we introduce an embedding-based Visual Stagnation Detection module using CLIP, which enables the model to assess visual progress and trigger proactive replanning numerically. Second, we propose a Trajectory Loop Purging algorithm that identifies and removes inverse action pairs from self-training data, ensuring the model learns only the optimal paths for task completion. Experimental results demonstrate that ReflectVLM+ achieves a 7% improvement in success rate over the baseline, even under hardware-constrained settings with only 400 trajectories. Furthermore, the proposed framework enhances operational efficiency by reducing the average inference time, demonstrating that quantitative reflection can optimize decision-making processes by early termination of redundant planning cycles. Our findings suggest that quantitative reflection and rigorous data pruning are essential for grounding high-level VLM reasoning into precise physical execution.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ReflectVLM+: Enhancing VLM-based Robotic Planning via Quantitative Stagnation Detection and Trajectory Refinement

  • Hyeon-Sung Choi,
  • Jong-Eun Ha

摘要

Vision-Language Models (VLMs) have emerged as a powerful paradigm for multi-stage robotic manipulation. Although recent frameworks like ReflectVLM have introduced visual reflection via dynamic models, they often suffer from ‘logical stagnation’—repeatedly executing failed actions because they cannot perceive subtle physical blocks. This paper proposes ReflectVLM+, an advanced framework that integrates quantitative stagnation detection and data refinement techniques to overcome these limitations. First, we introduce an embedding-based Visual Stagnation Detection module using CLIP, which enables the model to assess visual progress and trigger proactive replanning numerically. Second, we propose a Trajectory Loop Purging algorithm that identifies and removes inverse action pairs from self-training data, ensuring the model learns only the optimal paths for task completion. Experimental results demonstrate that ReflectVLM+ achieves a 7% improvement in success rate over the baseline, even under hardware-constrained settings with only 400 trajectories. Furthermore, the proposed framework enhances operational efficiency by reducing the average inference time, demonstrating that quantitative reflection can optimize decision-making processes by early termination of redundant planning cycles. Our findings suggest that quantitative reflection and rigorous data pruning are essential for grounding high-level VLM reasoning into precise physical execution.