<p>Vision-Language-Action (VLA) models have emerged as the dominant paradigm in embodied artificial intelligence (AI), as it allows agents to visually perceive their environment, linguistically interpret their verbal commands or natural language instructions and perform goal-oriented actions. Unfortunately, there is currently no existing review which provides empirical, quantitative, or metadata-driven evidence to illustrate the structural dynamics of the VLA research ecosystem. This study addresses that gap through a Systematic Quantitative Research Landscape Analysis (SQRLA) of 80 peer-reviewed VLA research articles published over the period 2022 to 2025, employing seven analytical dimensions including: publication growth trajectory modeling, dissemination channel profiling, citation impact analysis with K-Means trajectory clustering, lexical and word collocation analysis via the Log-Likelihood Ratio statistic, organizational contributor profiling, and institutional co-authorship network analysis along with structural hole detection using Burt’s constraint metric. Our key findings indicate an exceptionally rapid expansion of the VLA research field, characterized by a substantial increase in the number of publications between 2022 and 2025; a strong dominance of conference-based dissemination with emerging journal consolidation; a highly collaborative ecosystem driven by strong academic leadership, increasing industry contributions, and extensive cross-institutional partnerships; a median normalized citation rate 2.2 times higher for industry co-authored VLA papers than academic-only papers; evidence of conceptual maturation in 2024–2025, marked by the emergence of a coherent safety-and-deployment cluster and robot-assisted surgery as a statistically salient new application domain; a structural finding that VLA literature is measurably task-oriented over method-oriented; and a structural capability gap whereby European robotics institutions having critical expertise in haptics and compliant control remain isolated from the dominant collaboration network. These findings offer the first objective, quantitative or data-driven characterization of the VLA landscape and provide actionable evidence for researchers, funding bodies, and the broader community on publication strategy, collaboration priorities, and under-invested research areas.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Towards vision-language-action systems: A statistical and objective study of recent advances

  • Mrityunjay Bayan,
  • Rupam Bhattacharyya

摘要

Vision-Language-Action (VLA) models have emerged as the dominant paradigm in embodied artificial intelligence (AI), as it allows agents to visually perceive their environment, linguistically interpret their verbal commands or natural language instructions and perform goal-oriented actions. Unfortunately, there is currently no existing review which provides empirical, quantitative, or metadata-driven evidence to illustrate the structural dynamics of the VLA research ecosystem. This study addresses that gap through a Systematic Quantitative Research Landscape Analysis (SQRLA) of 80 peer-reviewed VLA research articles published over the period 2022 to 2025, employing seven analytical dimensions including: publication growth trajectory modeling, dissemination channel profiling, citation impact analysis with K-Means trajectory clustering, lexical and word collocation analysis via the Log-Likelihood Ratio statistic, organizational contributor profiling, and institutional co-authorship network analysis along with structural hole detection using Burt’s constraint metric. Our key findings indicate an exceptionally rapid expansion of the VLA research field, characterized by a substantial increase in the number of publications between 2022 and 2025; a strong dominance of conference-based dissemination with emerging journal consolidation; a highly collaborative ecosystem driven by strong academic leadership, increasing industry contributions, and extensive cross-institutional partnerships; a median normalized citation rate 2.2 times higher for industry co-authored VLA papers than academic-only papers; evidence of conceptual maturation in 2024–2025, marked by the emergence of a coherent safety-and-deployment cluster and robot-assisted surgery as a statistically salient new application domain; a structural finding that VLA literature is measurably task-oriented over method-oriented; and a structural capability gap whereby European robotics institutions having critical expertise in haptics and compliant control remain isolated from the dominant collaboration network. These findings offer the first objective, quantitative or data-driven characterization of the VLA landscape and provide actionable evidence for researchers, funding bodies, and the broader community on publication strategy, collaboration priorities, and under-invested research areas.