《计算机应用》唯一官方网站 ›› 2026, Vol. 46 ›› Issue (8): 2383-2393.DOI: 10.11772/j.issn.1001-9081.2025070900

• 人工智能 •    下一篇

面向具身智能的视觉-语言-动作模型动作表征和生成策略综述

张文涛1,2, 孙奥兰1, 瞿晓阳1, 张旭龙1, 王健宗1()   

  1. 1.平安科技(深圳)有限公司,广东 深圳 518046
    2.清华大学深圳国际研究生院,广东 深圳 518055
  • 收稿日期:2025-08-07 修回日期:2025-09-19 接受日期:2025-09-22 发布日期:2025-11-05 出版日期:2026-08-10
  • 通讯作者: 王健宗
  • 作者简介:张文涛(2001—),男,江西上饶人,硕士研究生,主要研究方向:大模型、具身智能
    孙奥兰(1996—),女,河北唐山人,工程师,硕士,主要研究方向:大模型、信号处理、具身智能
    瞿晓阳(1988—),男,湖北随州人,博士,CCF会员,主要研究方向:大模型、体系结构、具身智能
    张旭龙(1988—),男,河南许昌人,博士,CCF会员,主要研究方向:大模型、具身智能、跨模态智能计算
    第一联系人:王健宗(1983—),男,湖北天门人,正高级工程师,博士,CCF高级会员,主要研究方向:大模型、联邦学习、深度学习。
  • 基金资助:
    深港联合资助项目(A类)(SGDX20240115103359001)

Survey on action representation and generation strategies in vision-language-action models for embodied intelligence

Wentao ZHANG1,2, Aolan SUN1, Xiaoyang QU1, Xulong ZHANG1, Jianzong WANG1()   

  1. 1.Ping An Technology (Shenzhen) Company Limited,Shenzhen Guangdong 518046,China
    2.Tsinghua Shenzhen International Graduate School,Shenzhen Guangdong 518055,China
  • Received:2025-08-07 Revised:2025-09-19 Accepted:2025-09-22 Online:2025-11-05 Published:2026-08-10
  • Contact: Jianzong WANG
  • About author:ZHANG Wentao, born in 2001, M. S. candidate. His research interests include large models, embodied intelligence.
    SUN Aolan, born in 1996, M. S., engineer. Her research interests include large models, signal processing, embodied intelligence.
    QU Xiaoyang, born in 1988, Ph. D. His research interests include large models, architecture, embodied intelligence.
    ZHANG Xulong, born in 1988, Ph. D. His research interests include large models, embodied intelligence, cross-modal intelligent computing.
  • Supported by:
    Shenzhen-Hong Kong Joint Funding Project (Category A)(SGDX20240115103359001)

摘要:

视觉-语言-动作(VLA)模型是实现具身智能的核心路径,它们的核心是将多模态感知理解无缝转化为物理世界的具体行动。然而,VLA模型的动作表征与生成策略作为连接“感知”与“执行”的枢纽环节,面临着高维连续空间、动作多样性与机器人实时控制需求间的复杂挑战。因此,系统性地梳理和总结VLA模型中动作表征和生成策略的演进脉络、核心技术与未来方向;同时,详细剖析离散和连续两种动作表征方式以及自回归、非自回归和混合动作生成策略,并深入探讨它们在动作精度、生成多样性与推理效率之间的内在权衡;此外,总结面向实时控制的新兴高效策略,如混合生成架构。最后,通过比较分析,对现有技术图景进行总结,并展望未来VLA模型在与世界模型结合和跨机器人形态通用表征等方向上的前沿挑战与研究机遇,旨在为构建更通用且更高效的具身智能体提供参考。

关键词: 具身智能, 视觉-语言-动作模型, 动作表征, 生成策略

Abstract:

Vision-Language-Action (VLA) models are critical pathway to realize embodied intelligence, with their core of seamless transformation of multimodal perception and understanding into concrete actions in the physical world. However, action representation and generation strategies of VLA models, serving as the pivotal bridge between “perception” and “execution”, face complex challenges among the high-dimensional continuous spaces, the diversity of action modalities, and the stringent demands of real-time robotic control. Therefore, this survey sorted and summarized a systematic review of the evolution, key methodologies, and future directions of action representation and generation strategies in VLA models. At the same time, we analyzed discrete and continuous action representations in detail, as well as autoregressive, non-autoregressive, and hybrid action generation strategies, highlighting their inherent trade-offs in terms of action precision, generation diversity, and reasoning efficiency. In addition, we summed up emerging high-efficiency strategies for real-time control, such as hybrid generation architecture. Finally, we presented a summary of the current technological landscape and prospected frontier challenges and research opportunities of VLA models integrating with world models and the development of cross-robot generalized representations, aiming to provide a reference for constructing more general and efficient embodied agents.

Key words: embodied intelligence, Visual-Language-Action (VLA) model, action representation, generation strategy

中图分类号: