#WorldVLA
WorldVLA: Towards Autoregressive Action World Model | Discussion
WorldVLA: Towards Autoregressive Action World Model
We present WorldVLA, an autoregressive action world model that unifies action and image understanding and generation. Our WorldVLA intergrates Vision-Language-Action (VLA) model and world model in one single framework. The world model predicts future images by leveraging both action and image understanding, with the purpose of learning the underlying physics of the environment to improve action generation. Meanwhile, the action model generates the subsequent actions based on image observations, aiding in visual understanding and in turn helps visual generation of the world model. We demonstrate that WorldVLA outperforms standalone action and world models, highlighting the mutual enhancement between the world model and the action model. In addition, we find that the performance of the action model deteriorates when generating sequences of actions in an autoregressive manner. This phenomenon can be attributed to the model's limited generalization capability for action prediction, leading to the propagation of errors from earlier actions to subsequent ones. To address this issue, we propose an attention mask strategy that selectively masks prior actions during the generation of the current action, which shows significant performance improvement in the action chunk generation task.
arxiv.org
June 30, 2025 at 1:40 AM
[26/30] 185 Likes, 5 Comments, 3 Posts
2506.21539, cs․RO | cs․AI, 26 Jun 2025

🆕WorldVLA: Towards Autoregressive Action World Model

Jun Cen, Chaohui Yu, Hangjie Yuan, Yuming Jiang, Siteng Huang, Jiayan Guo, Xin Li, Yibing Song, Hao Luo, Fan Wang, Deli Zhao, Hao Chen
July 2, 2025 at 12:06 AM
阿里的达摩院提出的WorldVLA模型超厉害👏首次把世界模型和动作模型融合进全自回归模型,统一了文本、图片和动作的理解与生成

它的双向增强机制超创新,还有动作注意力掩码策略缓解错误传播。在LIBERO测试中,抓取成功率和视频质量大幅提升,为具身智能开辟新路径,感兴趣的朋友可以看一下论文 arxiv.org/pdf/2506.21539
July 7, 2025 at 4:10 AM
2506.21539
アクションとイメージの理解と生成を統合する自己回帰的アクション世界モデル、WorldVLAを紹介する。我々のWorldVLAは、VLA(Vision-Language-Action)モデルとワールドモデルを1つのフレームワークに統合している。ワールドモデルは、アクションとイメージ理解の両方を活用することで、未来のイメージを予測...
July 2, 2025 at 12:07 AM
(3/3) 24 Likes, 2 Comments, 29 Jun 2025, Hacker News
WorldVLA: Towards Autoregressive Action World Model | Hacker News
news.ycombinator.com
July 2, 2025 at 12:07 AM
(2/3) 34 Likes, 3 Comments, 27 Jun 2025, Hugging Face
Paper page - WorldVLA: Towards Autoregressive Action World Model
Join the discussion on this paper page
huggingface.co
July 2, 2025 at 12:07 AM
Jun Cen, Chaohui Yu, Hangjie Yuan, Yuming Jiang, Siteng Huang, Jiayan Guo, Xin Li, Yibing Song, Hao Luo, Fan Wang, Deli Zhao, Hao Chen
WorldVLA: Towards Autoregressive Action World Model
https://arxiv.org/abs/2506.21539
June 27, 2025 at 4:12 AM
WorldVLA:迈向自回归动作世界模型的未来探索

https://qian.cx/posts/69EAB38F-90C8-444D-A354-2FD9986847A4
September 30, 2025 at 2:06 AM
WorldVLA: инновационная модель для предсказания и генерации действий в робототехнике

https://kripta.biz/posts/21C014D2-0732-4134-9BB3-BCC3DDF6D067
September 30, 2025 at 2:06 AM
WorldVLA: Towards Autoregressive Action World Model
https://arxiv.org/abs/2506.21539
[comments] [8 points]
June 30, 2025 at 2:03 AM
Jun Cen, Chaohui Yu, Hangjie Yuan, Yuming Jiang, Siteng Huang, Jiayan Guo, Xin Li, Yibing Song, Hao Luo, Fan Wang, Deli Zhao, Hao Chen: WorldVLA: Towards Autoregressive Action World Model https://arxiv.org/abs/2506.21539 https://arxiv.org/pdf/2506.21539 https://arxiv.org/html/2506.21539
June 27, 2025 at 6:34 AM
WorldVLA: Towards Autoregressive Action World Model
L: https://arxiv.org/abs/2506.21539
C: https://news.ycombinator.com/item?id=44417725
posted on 2025.06.29 at 19:51:57 (c=0, p=3)
June 30, 2025 at 1:25 AM
WorldVLA: Towards Autoregressive Action World Model

https://arxiv.org/abs/2506.21539
June 30, 2025 at 2:00 AM
@rohanpaul_ai https://x.com/rohanpaul_ai/status/1939000339853910203 #x-rohanpaul_ai

Vision-Language-Action (VLA) models lack deep action understanding, and world models cannot generate actions.

Autoregressive action generation suffers from accumulated errors.

WorldVLA unifies actio...
June 28, 2025 at 5:00 PM
WorldVLA: Towards Autoregressive Action World Model
Chaohui Yu, Deli Zhao et al.
Paper
Details
#WorldVLA #AutoregressiveActionModel #MachineLearningResearch
June 30, 2025 at 9:03 AM
알리바바 WorldVLA 대박! 🤖 카메라로 본 것과 음성 명령을 동시에 이해해서 로봇이 스스로 행동을 결정해. 구글, 엔비디아 모델과 경쟁하는 Physical AI인데, 오류 누적 문제도 해결했대. 앞으로 로봇이 더 똑똑해질 듯! 완전 기대되지? 😉
blog.naver.com/jack0604/223...
August 18, 2025 at 10:01 PM