《计算机应用》唯一官方网站 ›› 2026, Vol. 46 ›› Issue (8): 2383-2393.DOI: 10.11772/j.issn.1001-9081.2025070900
• 人工智能 • 下一篇
张文涛1,2, 孙奥兰1, 瞿晓阳1, 张旭龙1, 王健宗1(
)
收稿日期:2025-08-07
修回日期:2025-09-19
接受日期:2025-09-22
发布日期:2025-11-05
出版日期:2026-08-10
通讯作者:
王健宗
作者简介:张文涛(2001—),男,江西上饶人,硕士研究生,主要研究方向:大模型、具身智能基金资助:
Wentao ZHANG1,2, Aolan SUN1, Xiaoyang QU1, Xulong ZHANG1, Jianzong WANG1(
)
Received:2025-08-07
Revised:2025-09-19
Accepted:2025-09-22
Online:2025-11-05
Published:2026-08-10
Contact:
Jianzong WANG
About author:ZHANG Wentao, born in 2001, M. S. candidate. His research interests include large models, embodied intelligence.Supported by:摘要:
视觉-语言-动作(VLA)模型是实现具身智能的核心路径,它们的核心是将多模态感知理解无缝转化为物理世界的具体行动。然而,VLA模型的动作表征与生成策略作为连接“感知”与“执行”的枢纽环节,面临着高维连续空间、动作多样性与机器人实时控制需求间的复杂挑战。因此,系统性地梳理和总结VLA模型中动作表征和生成策略的演进脉络、核心技术与未来方向;同时,详细剖析离散和连续两种动作表征方式以及自回归、非自回归和混合动作生成策略,并深入探讨它们在动作精度、生成多样性与推理效率之间的内在权衡;此外,总结面向实时控制的新兴高效策略,如混合生成架构。最后,通过比较分析,对现有技术图景进行总结,并展望未来VLA模型在与世界模型结合和跨机器人形态通用表征等方向上的前沿挑战与研究机遇,旨在为构建更通用且更高效的具身智能体提供参考。
中图分类号:
张文涛, 孙奥兰, 瞿晓阳, 张旭龙, 王健宗. 面向具身智能的视觉-语言-动作模型动作表征和生成策略综述[J]. 计算机应用, 2026, 46(8): 2383-2393.
Wentao ZHANG, Aolan SUN, Xiaoyang QU, Xulong ZHANG, Jianzong WANG. Survey on action representation and generation strategies in vision-language-action models for embodied intelligence[J]. Journal of Computer Applications, 2026, 46(8): 2383-2393.
| 年份 | 模型 | 核心范式 | 平台 | 任务领域 | 动作空间类型 | 动作 维度 | 离散 区间 | 存在的问题 |
|---|---|---|---|---|---|---|---|---|
| 2022 | RT-1[ | 模仿学习 | 移动机械臂 | 移动操作 | 末端执行器位置、姿态+移动盘 | 11 | 256 | 模仿学习上限、泛化局限 |
| 2022 | Gato[ | 通用监督学习 | Sawyer机械臂等 | 机器人操作、游戏、对话 | 末端执行器速度控制+夹爪控制 | 5 | 1 024 | 上下文长度限制、推理速度慢 |
| 2023 | RT-2[ | VLM共同微调 | 移动机械臂 | 语义驱动的操作 | 末端执行器位置、姿态+移动盘 | 11 | 256 | 物理技能局限、计算成本高 |
| 2023 | Q-Transformer[ | 离线强化学习 | 移动机械臂 | 多任务操作 | 末端执行器位置、姿态、夹爪 | 8 | 256 | 奖励函数局限、高维动作局限 |
| 2024 | OpenVLA[ | VLM微调 | 多种机械臂 | 跨物理形态的操作 | 末端执行器位置、姿态+夹爪 | 7 | 256 | 仅支持单图像、推理效率低 |
| 2025 | Humanoid-VLA[ | 语言-运动对齐 | 人形机器人 | 移动-操作 | 全身运动姿态 | 24 | 1 024 | 数据质量数量有限、依赖底层RL策略 |
| 2025 | JARVIS-VLA[ | ActVLP | 虚拟智能体 | 游戏操作 | 键盘与鼠标 | — | 51 | 推理慢、与顶尖人类玩家差距大 |
表1 典型模型所用的离散动作表征
Tab. 1 Discretized action representation used in typical models
| 年份 | 模型 | 核心范式 | 平台 | 任务领域 | 动作空间类型 | 动作 维度 | 离散 区间 | 存在的问题 |
|---|---|---|---|---|---|---|---|---|
| 2022 | RT-1[ | 模仿学习 | 移动机械臂 | 移动操作 | 末端执行器位置、姿态+移动盘 | 11 | 256 | 模仿学习上限、泛化局限 |
| 2022 | Gato[ | 通用监督学习 | Sawyer机械臂等 | 机器人操作、游戏、对话 | 末端执行器速度控制+夹爪控制 | 5 | 1 024 | 上下文长度限制、推理速度慢 |
| 2023 | RT-2[ | VLM共同微调 | 移动机械臂 | 语义驱动的操作 | 末端执行器位置、姿态+移动盘 | 11 | 256 | 物理技能局限、计算成本高 |
| 2023 | Q-Transformer[ | 离线强化学习 | 移动机械臂 | 多任务操作 | 末端执行器位置、姿态、夹爪 | 8 | 256 | 奖励函数局限、高维动作局限 |
| 2024 | OpenVLA[ | VLM微调 | 多种机械臂 | 跨物理形态的操作 | 末端执行器位置、姿态+夹爪 | 7 | 256 | 仅支持单图像、推理效率低 |
| 2025 | Humanoid-VLA[ | 语言-运动对齐 | 人形机器人 | 移动-操作 | 全身运动姿态 | 24 | 1 024 | 数据质量数量有限、依赖底层RL策略 |
| 2025 | JARVIS-VLA[ | ActVLP | 虚拟智能体 | 游戏操作 | 键盘与鼠标 | — | 51 | 推理慢、与顶尖人类玩家差距大 |
| 年份 | 代表模型 | 核心范式 | 平台 | 任务领域 | 动作维度 | 表征方法类型 | 存在的问题 |
|---|---|---|---|---|---|---|---|
| 2023 | ACT[ | 模仿学习 | 机械臂 | 精细双臂操作 | 14 | 条件变分 | 硬件局限、感知挑战 |
| 2024 | Octo[ | 模仿学习 | 机械臂 | 跨物理形态的通用操作 | 7/14 | 条件扩散 | 手腕摄像头处理不佳、依赖演示数据 |
| 2024 | VLM微调 | 机械臂、移动机器人 | 高灵巧度、长时程操作 | 18 | 条件流匹配 | 严重依赖大规模、部分不开源的高质量演示数据 | |
| 2025 | HybridVLA[ | 协同训练 | 机械臂 | 通用桌面操作 | 7/14 | 混合生成 | 推理速度受限 |
| 2025 | DexVLA[ | 具身课程学习 | 机械臂、灵巧手机器人 | 跨物理形态的灵巧操作 | — | 多头扩散 | 在涉及频繁接触的复杂场景中适用性有限 |
表2 典型模型所用的连续动作表征
Tab. 2 Continuous action representations for typical models
| 年份 | 代表模型 | 核心范式 | 平台 | 任务领域 | 动作维度 | 表征方法类型 | 存在的问题 |
|---|---|---|---|---|---|---|---|
| 2023 | ACT[ | 模仿学习 | 机械臂 | 精细双臂操作 | 14 | 条件变分 | 硬件局限、感知挑战 |
| 2024 | Octo[ | 模仿学习 | 机械臂 | 跨物理形态的通用操作 | 7/14 | 条件扩散 | 手腕摄像头处理不佳、依赖演示数据 |
| 2024 | VLM微调 | 机械臂、移动机器人 | 高灵巧度、长时程操作 | 18 | 条件流匹配 | 严重依赖大规模、部分不开源的高质量演示数据 | |
| 2025 | HybridVLA[ | 协同训练 | 机械臂 | 通用桌面操作 | 7/14 | 混合生成 | 推理速度受限 |
| 2025 | DexVLA[ | 具身课程学习 | 机械臂、灵巧手机器人 | 跨物理形态的灵巧操作 | — | 多头扩散 | 在涉及频繁接触的复杂场景中适用性有限 |
| 动作类型 | VLA模型 | 平均成功率 |
|---|---|---|
| 连续 | Diffusion Policy[ | 72.4 |
| Octo[ | 75.1 | |
| DiT Policy[ | 82.4 | |
| OpenVLA-OFT[ | 95.4 | |
| 94.2 | ||
| 离散 | OpenVLA[ | 76.5 |
| WorldVLA[ | 79.1 |
表3 LIBERO数据集典型VLA模型评估 (%)
Tab. 3 Evaluation of typical VLA models on LIBERO dataset
| 动作类型 | VLA模型 | 平均成功率 |
|---|---|---|
| 连续 | Diffusion Policy[ | 72.4 |
| Octo[ | 75.1 | |
| DiT Policy[ | 82.4 | |
| OpenVLA-OFT[ | 95.4 | |
| 94.2 | ||
| 离散 | OpenVLA[ | 76.5 |
| WorldVLA[ | 79.1 |
| 动作类型 | VLA模型 | 平均成功率 |
|---|---|---|
| 连续 | Octo-Base [ | 16.8 |
| 70.1 | ||
| 离散 | RT-1[ | 6.8 |
| TraceVLA[ | 42.0 | |
| RT-1-X[ | 53.4 | |
| RT-2-X[ | 60.7 | |
| OpenVLA[ | 27.7 |
表4 Open X-Embodiment数据集典型VLA模型评估 (%)
Tab. 4 Evaluation of typical VLA models on Open X-Embodiment
| 动作类型 | VLA模型 | 平均成功率 |
|---|---|---|
| 连续 | Octo-Base [ | 16.8 |
| 70.1 | ||
| 离散 | RT-1[ | 6.8 |
| TraceVLA[ | 42.0 | |
| RT-1-X[ | 53.4 | |
| RT-2-X[ | 60.7 | |
| OpenVLA[ | 27.7 |
| [1] | Roumeliotis K I, Tselikas N D. ChatGPT and Open-AI models: a preliminary review[J]. Future Internet, 2023, 15(6): No.192. |
| [2] | Stone A, Xiao T, Lu Y, et al. Open-world object manipulation using pre-trained vision-language models[J]. Journal of Machine Learning Research, 2023, 229: 3397-3417. |
| [3] | Brohan A, Brown N, Carbajal J, et al. RT-1: robotics Transformer for real-world control at scale[EB/OL]. (2023-07) [2025-10-30].. |
| [4] | Brohan A, Brown N, Carbajal J, et al. RT-2: vision-language-action models transfer web knowledge to robotic control[J]. Journal of Machine Learning Research, 2023, 229: 2165-2183. |
| [5] | Kim M J, Pertsch K, Karamcheti S, et al. OpenVLA: an open-source vision-language-action model[J]. Journal of Machine Learning Research, 2025, 270: 2679-2713. |
| [6] | Ghosh D, Walke H, Pertsch K, et al. Octo: an open-source generalist robot policy[EB/OL]. (2024-05-20) [2025-10-30].. |
| [7] | Sapkota R, Cao Y, Roumeliotis K I, et al. Vision-language-action models: concepts, progress, applications and challenges[PP/OL]. V1. arXiv (2025-05-07) [2025-05-10].. |
| [8] | 赵博涛,亢祖衡,瞿晓阳,等. 基于多模态大模型的具身智能体研究进展与展望[J]. 大数据, 2025, 11(3): 108-138. |
| Zhao Botao, Kang Zuheng, Qu Xiaoyang, et al. Review and emerging trends of embodied agent based on multimodal large language models[J]. Big Data Research, 2025, 11(3): 108-138. | |
| [9] | Zhong Y, Bai F, Cai S, et al. A survey on vision-language-action models: an action tokenization perspective[PP/OL]. V1. arXiv (2025-07-02) [2025-07-08].. |
| [10] | Liu Y, Chen W, Bai Y, et al. Aligning cyber space with physical world: a comprehensive survey on embodied AI[J]. IEEE/ASME Transactions on Mechatronics, 2025, 30(6): 7253-7274. |
| [11] | Han K, Wang Y, Chen H, et al. A survey on Vision Transformer[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023, 45(1): 87-110. |
| [12] | Liu H, Li C, Wu Q, et al. Visual instruction tuning[C]// NeurIPS 2023. Red Hook: Curran Associates Inc., 2023: 34892-34916. |
| [13] | Kim M J, Finn C, Liang P. Fine-tuning vision-language-action models: optimizing speed and success[EB/OL]. (2025-02-27) [2025-07-02].. |
| [14] | Zhang Z, Zheng K, Chen Z, et al. GRAPE: generalizing robot policy via preference alignment[PP/OL]. V2. arXiv (2025-02-04) [2025-07-02].. |
| [15] | Katiyar N. A model-driven framework for domain-specific adaptation of time series forecasting pipeline[D/OL]. [2025-07-02].. |
| [16] | Shridhar M, Manuelli L, Fox D. CLIPort: what and where pathways for robotic manipulation[J]. Journal of Machine Learning Research, 2022, 164: 894-906. |
| [17] | NVIDIA. GR00T N1: an open foundation model for generalist humanoid robots[EB/OL]. (2025-03-18) [2025-06-08].. |
| [18] | Zhao T Z, Kumar V, Levine S, et al. Learning fine-grained bimanual manipulation with low-cost hardware[EB/OL]. (2023-04-23) [2025-10-30].. |
| [19] | Chi C, Xu Z, Feng S, et al. Diffusion policy: visuomotor policy learning via action diffusion[J]. The International Journal of Robotics Research, 2023, 44(10/11): 1684-1704. |
| [20] | Pertsch K, Stachowicz K, Ichter B, et al. FAST: efficient action tokenization for vision-language-action models[PP/OL]. (2025-02-06) [2025-07-08].. |
| [21] | Gao J, Belkhale S, Dasari S, et al. A taxonomy for evaluating generalist robot manipulation policies[J]. IEEE Robotics and Automation Letters, 2026, 11(3): 3182-3189. |
| [22] | Chen B, Xu Z, Kirmani S, et al. SpatialVLM: endowing vision-language models with spatial reasoning capabilities[C]// CVPR 2024. Piscataway: IEEE, 2024: 14455-14465. |
| [23] | Fan C, Jia X, Sun Y, et al. Interleave-VLA: enhancing robot manipulation with interleaved image-text instructions[PP/OL]. V2. arXiv (2025-08-08) [2025-06-08].. |
| [24] | Cheng H, Xiao E, Yu C, et al. Manipulation facing threats: evaluating physical vulnerabilities in end-to-end vision language action models[PP/OL]. V2. arXiv (2024-11-04) [2025-07-10].. |
| [25] | Firoozi R, Tucker J, Tian S, et al. Foundation models in robotics: applications, challenges, and the future[J]. The International Journal of Robotics Research, 2025, 44(5): 701-739. |
| [26] | Hu Y, Xie Q, Jain V, et al. Toward general-purpose robots via foundation models: a survey and meta-analysis[PP/OL]. V3. arXiv (2025-02-06) [2025-07-08].. |
| [27] | Reed S, Żołna K, Parisotto E, et al. A generalist agent[PP/OL]. V3. arXiv (2022-11-11) [2025-07-08].. |
| [28] | Chebotar Y, Vuong Q, Hausman K, et al. Q-Transformer: scalable offline reinforcement learning via autoregressive Q-functions[J]. Journal of Machine Learning Research, 2023, 229: 3909-3928. |
| [29] | Ding P, Ma J, Tong X, et al. Humanoid-VLA: towards universal humanoid control with visual integration[PP/OL]. V2. arXiv (2025-02-21) [2025-07-08].. |
| [30] | Li M, Wang Z, He K, et al. Jarvis-VLA: post-training large-scale vision language models to play visual games with keyboards and mouse[C]// Findings of the Association for Computational Linguistics: ACL 2025. Stroudsburg: ACL, 2025: 17878-17899. |
| [31] | Gu Z, Li J, Shen W, et al. Humanoid locomotion and manipulation: current progress and challenges in control, planning, and learning[J]. IEEE/ASME Transactions on Mechatronics, 2026, 31(2): 2300-2330. |
| [32] | Nie Y, Li L, Gan Z, et al. MLP architectures for vision-and-language modeling: an empirical study[PP/OL]. V1. arXiv (2021-12-08) [2025-06-05].. |
| [33] | Li X, Wang S, Chen C, et al. RoboFlamingo-Plus: fusion of depth and RGB perception with vision-language models for enhanced robotic manipulation[C]// RCAR 2025. Piscataway: IEEE, 2025: 311-316. |
| [34] | Ren A Z, Lidard J, Ankile L L, et al. Diffusion policy policy optimization[PP/OL]. V3. arXiv (2024-12-09) [2025-07-08].. |
| [35] | Guo Y, Zhang J, Chen X, et al. Improving vision-language-action model with online reinforcement learning[C]// ICRA 2025. Piscataway: IEEE, 2025: 15665-15672. |
| [36] | Chiang H T L, Xu Z, Fu Z, et al. Mobility VLA: multimodal instruction navigation with long-context VLMs and topological graphs[PP/OL]. V2. arXiv (2014-07-12) [2025-07-10].. |
| [37] | Black K, Brown N, Driess D, et al. π 0 : a vision-language-action flow model for general robot control[PP/OL]. V3. arXiv (2024-11-13) [2025-07-08]. . |
| [38] | Liu J, Chen H, An P, et al. HybridVLA: collaborative diffusion and autoregression in a unified vision-language-action model[PP/OL]. V3. arXiv (2025-06-23) [2025-07-08].. |
| [39] | Wen J, Zhu Y, Li J, et al. DexVLA: vision-language model with plug-in diffusion expert for general robot control[PP/OL]. V1. arXiv (2025-02-06) [2025-06-20].. |
| [40] | Zhao Q, Lu Y, Kim M J, et al. Cot-VLA: visual chain-of-thought reasoning for vision-language-action models[C]// CVPR 2025. Piscataway: IEEE, 2025: 1702-1713. |
| [41] | Li Z, Wu X, Du H, et al. Benchmark evaluations, applications, and challenges of large vision language models: a survey[PP/OL]. V4. arXiv (2025-03-13) [2025-06-18].. |
| [42] | Gbagbe K F, Cabrera M A, Alabbas A, et al. Bi-VLA: vision-language-action model-based system for bimanual robotic dexterous manipulations[C]// SMC 2024. Piscataway: IEEE, 2024: 2864-2869. |
| [43] | Xiang T Y, Jin A Q, Zhou X H, et al. VLA model-expert collaboration for bi-directional manipulation learning[PP/OL]. V1. arXiv (2025-03-06) [2025-08-01].. |
| [44] | Driess D, Springenberg J T, Ichter B, et al. Knowledge insulating vision-language-action models: train fast, run fast, generalize better[PP/OL]. V1. arXiv (2025-03-29) [2025-06-08].. |
| [45] | Liu S, Wu L, Li B, et al. RDT-1B: a diffusion foundation model for bimanual manipulation[PP/OL]. V2. arXiv (2025-03-01) [2025-07-08].. |
| [46] | Devlin J, Chang M W, Lee K, et al. BERT: pre-training of deep bidirectional Transformers for language understanding[C]// NAACL-HLT 2019, Volume 1 (Long and Short Papers). Stroudsburg: ACL, 2019: 4171-4186. |
| [47] | Radford A, Wu J, Child R, et al. Language models are unsupervised multitask learners[EB/OL]. [2025-06-08].. |
| [48] | Brown T, Mann B, Ryder N, et al. Language models are few-shot learners[C]// NeurIPS 2020. Red Hook: Curran Associates Inc., 2020: 1877-1901. |
| [49] | Jiang Y, Gupta A, Zhang Z, et al. VIMA: general robot manipulation with multimodal prompts[EB/OL]. (2023-02-02) [2025-07-08].. |
| [50] | Xu Z, Wu K, Wen J, et al. A survey on robotics with foundation models: toward embodied AI[PP/OL]. V1. arXiv (2024-02-04) [2025-05-08].. |
| [51] | Zhou Z, Zhu Y, Zhu M, et al. ChatVLA: unified multimodal understanding and robot control with vision-language-action model[PP/OL]. V2. arXiv (2025-02-21) [2025-06-10].. |
| [52] | Guruprasad P, Sikka H, Song J, et al. Benchmarking vision, language, & action models on robotic learning tasks[PP/OL]. V2. arXiv (2025-12-08) [2025-06-08].. |
| [53] | O’Neill A, Rehman A, Maddukuri A, et al. Open X-Embodiment: robotic learning datasets and RT-X models: Open X-Embodiment Collaboration0 [C]// ICRA 2024. Piscataway: IEEE, 2024: 6892-6903. |
| [54] | Ke T W, Gkanatsios N, Fragkiadaki K. 3D Diffuser Actor: policy diffusion with 3D scene representations[PP/OL]. V3. arXiv (2024-07-25) [2025-06-08].. |
| [55] | Ding P, Zhao H, Zhang W, et al. QUAR-VLA: vision-language-action model for quadruped robots[C]// ECCV 2024, LNCS 15063. Cham: Springer, 2024: 352-367. |
| [56] | Ju C, Wang H, Cheng H, et al. Turbo: informativity-driven acceleration plug-in for vision-language large models[C]// ECCV 2024, LNCS 15104. Cham: Springer, 2024: 436-455. |
| [57] | Zhang T, Hu Y, Cui H, et al. A universal semantic-geometric representation for robotic manipulation[J]. Journal of Machine Learning Research, 2023, 229: 3342-3363. |
| [58] | Xu S, Wang Y, Xia C, et al. VLA-Cache: towards efficient vision-language-action model via adaptive token caching in robotic manipulation[PP/OL]. V1. arXiv (2025-02-04) [2025-07-01].. |
| [59] | Ho J, Jain A, Abbeel P. Denoising diffusion probabilistic models[C]// NeurIPS 2020. Red Hook: Curran Associates Inc., 2020: 6840-6851. |
| [60] | Nichol A, Dhariwal P. Improved denoising diffusion probabilistic models[J]. Journal of Machine Learning Research, 2021, 139: 8162-8171. |
| [61] | Choi J, Kim S, Jeong Y, et al. ILVR: conditioning method for denoising diffusion probabilistic models[C]// ICCV 2021. Piscataway: IEEE, 2021: 14347-14356. |
| [62] | Reuss M, Yağmurlu Ö E, Wenzel F, et al. Multimodal diffusion Transformer: learning versatile behavior from multimodal goals[EB/OL]. (2024-06-08) [2025-06-08].. |
| [63] | Peebles W, Xie S. Scalable diffusion models with Transformers[C]// ICCV 2023. Piscataway: IEEE, 2023: 4172-4182. |
| [64] | Li Q, Liang Y, Wang Z, et al. CogACT: a foundational vision-language-action model for synergizing cognition and action in robotic manipulation[PP/OL]. V1. arXiv (2024-11-29) [2025-06-08].. |
| [65] | Wu J, Zhong M, Xing S, et al. VisionLLM v2: an end-to-end generalist multimodal large language model for hundreds of vision-language tasks[C]// NeurIPS 2024. Red Hook: Curran Associates Inc., 2024: 69925-69975. |
| [66] | Esser P, Rombach R, Ommer B. Taming Transformers for high-resolution image synthesis[C]// CVPR 2021. Piscataway: IEEE, 2021: 12868-12878. |
| [67] | Bolya D, Huang P Y, Sun P, et al. Perception encoder: the best visual embeddings are not at the output of the network[PP/OL]. V2. arXiv (2025-04-28) [2025-06-10].. |
| [68] | Lipman Y, Chen R T Q, Ben-Hamu H, et al. Flow matching for generative modeling[EB/OL]. (2023-02-02) [2025-07-08].. |
| [69] | Gat I, Remez T, Shaul N, et al. Discrete flow matching[C]// NeurIPS 2024. Red Hook: Curran Associates Inc., 2024: 133345-133385. |
| [70] | Deng S, Yan M, Wei S, et al. GraspVLA: a grasping foundation model pre-trained on billion-scale synthetic action data[PP/OL]. V3. arXiv (2025-08-27) [2025-09-08].. |
| [71] | Wen J, Zhu Y, Li J, et al. TinyVLA: towards fast, data-efficient vision-language-action models for robotic manipulation[J]. IEEE Robotics and Automation Letters, 2025, 10(4): 3988-3995. |
| [72] | Zhang B, Zhang Y, Ji J, et al. SafeVLA: towards safety alignment of vision-language-action model via constrained learning[PP/OL]. V2. arXiv (2025-05-31) [2025-06-08].. |
| [73] | Li X, Liu M, Zhang H, et al. Vision-language foundation models as effective robot imitators[PP/OL]. V3. arXiv (2024-02-05) [2025-07-08].. |
| [74] | Chen L, Lu K, Rajeswaran A, et al. Decision Transformer: reinforcement learning via sequence modeling[C]// NeurIPS 2021. Red Hook: Curran Associates Inc., 2021: 15084-15097. |
| [75] | Liu B, Zhu Y, Gao C, et al. LIBERO: benchmarking knowledge transfer for lifelong robot learning[C]// NeurIPS 2023. Red Hook: Curran Associates Inc., 2023: 44776-44791. |
| [76] | Hou Z, Zhang T, Xiong Y, et al. Diffusion Transformer policy[PP/OL]. V6. arXiv (2025-03-23) [2025-07-12].. |
| [77] | Cen J, Yu C, Yuan H, et al. WorldVLA: towards autoregressive action world model[PP/OL]. V1. arXiv (2025-06-26) [2025-09-12].. |
| [78] | Zheng R, Liang Y, Huang S, et al. TraceVLA: visual trace prompting enhances spatial-temporal awareness for generalist robotic policies[PP/OL]. V3. arXiv (2025-06-05) [2025-06-18].. |
| [79] | Li X, Li P, Liu M, et al. Towards generalist robot policies: what matters in building vision-language-action models[PP/OL]. V3. arXiv (2025-12-24) [2025-07-08].. |
| [80] | Intelligence Physical. π 0.5 : a vision-language-action model with open-world generalization[PP/OL]. V1. arXiv (2025-03-22) [2025-06-08].. |
| [81] | Fan L, Chen K, Xu Z, et al. Language reasoning in vision-language-action model for robotic grasping[C]// CAC 2024. Piscataway: IEEE, 2024: 6656-6661. |
| [82] | Ma Y, Song Z, Zhuang Y, et al. A survey on vision-language-action models for embodied AI[J]. IEEE Transactions on Neural Networks and Learning Systems, 2026(Early Access): 1-21. |
| [83] | Zhou X, Han X, Yang F, et al. OpenDriveVLA: towards end-to-end autonomous driving with large vision language action model[PP/OL]. V1. arXiv (2025-03-30) [2025-07-15].. |
| [84] | Zhen H, Qiu X, Chen P, et al. 3D-VLA: a 3D vision-language-action generative world model[J]. Proceedings of Machine Learning Research, 2024, 235: 61229-61245. |
| [85] | Huang J, Yong S, Ma X, et al. An embodied generalist agent in 3D world[J]. Proceedings of Machine Learning Research, 2024, 235: 20413-20451. |
| [86] | Patel D, Eghbalzadeh H, Kamra N, et al. Pretrained language models as visual planners for human assistance[C]// ICCV 2023. Piscataway: IEEE, 2023: 15256-15268. |
| [87] | Zhou Y, Wang S, Dai S, et al. CHOP: mobile operating assistant with constrained high-frequency optimized subtask planning[PP/OL]. V1. arXiv (2025-03-05) [2025-06-20].. |
| [88] | Wang Z, Zhou Z, Song J, et al. Towards testing and evaluating vision-language-action models for robotic manipulation: an empirical study[PP/OL]. V1. arXiv (2024-09-19) [2025-06-18].. |
| [1] | 张仲华, 赵福媛, 郭钧枫, 赵高长. 柯西自适应回溯搜索与最小二乘支持向量机的集成预测模型[J]. 《计算机应用》唯一官方网站, 2022, 42(6): 1829-1836. |
| [2] | 王中玉, 曾国辉, 黄勃, 方志军. 改进A*算法的机器人全局最优路径规划[J]. 计算机应用, 2019, 39(9): 2517-2522. |
| [3] | 霍莉萍,郭成城. Web Server应答负载的生成策略[J]. 计算机应用, 2005, 25(06): 1458-1460. |
| 阅读次数 | ||||||
|
全文 |
|
|||||
|
摘要 |
|
|||||