Journals
  Publication Years
  Keywords
Search within results Open Search
Please wait a minute...
For Selected: Toggle Thumbnails
Survey on action representation and generation strategies in vision-language-action models for embodied intelligence
Wentao ZHANG, Aolan SUN, Xiaoyang QU, Xulong ZHANG, Jianzong WANG
Journal of Computer Applications    2026, 46 (8): 2383-2393.   DOI: 10.11772/j.issn.1001-9081.2025070900
Abstract118)   HTML8)    PDF (1186KB)(71)       Save

Vision-Language-Action (VLA) models are critical pathway to realize embodied intelligence, with their core of seamless transformation of multimodal perception and understanding into concrete actions in the physical world. However, action representation and generation strategies of VLA models, serving as the pivotal bridge between “perception” and “execution”, face complex challenges among the high-dimensional continuous spaces, the diversity of action modalities, and the stringent demands of real-time robotic control. Therefore, this survey sorted and summarized a systematic review of the evolution, key methodologies, and future directions of action representation and generation strategies in VLA models. At the same time, we analyzed discrete and continuous action representations in detail, as well as autoregressive, non-autoregressive, and hybrid action generation strategies, highlighting their inherent trade-offs in terms of action precision, generation diversity, and reasoning efficiency. In addition, we summed up emerging high-efficiency strategies for real-time control, such as hybrid generation architecture. Finally, we presented a summary of the current technological landscape and prospected frontier challenges and research opportunities of VLA models integrating with world models and the development of cross-robot generalized representations, aiming to provide a reference for constructing more general and efficient embodied agents.

Table and Figures | Reference | Related Articles | Metrics