#Vision-Language-Action
Ant Group trained a robot vision-language-action model on 20,000 hours of robot data.
January 29, 2026 at 12:01 AM
Sharpa has a new vision-tactile-language-action model, and they demoed it live at CES. Video is autonomous and some of the most impressive dexterous manipulation I have ever seen.
January 8, 2026 at 2:50 PM
⭐ The first foundational model available on @LeRobotHF ⭐

Pi0 is the most advanced Vision Language Action model. It takes natural language commands as input and directly output autonomous behavior.

It was trained by @physical_int and ported to pytorch by @m_olbap
👇🧵
February 4, 2025 at 5:07 PM
We don't want vision-language-action models to remain bounded by human performance, but training these large models with reinforcement learning is difficult. How can we handle this? Short blog post: open.substack.com/pub/itcanthi...
Paper notes: Improving Vision-Language-Action Model with Online Reinforcement Learning
Many vision-language-action models really don’t work that well.
open.substack.com
February 6, 2025 at 2:39 PM
Vision-Language-Action models are the foundation of a new wave of generalist robots: networks that take in images from robot cameras (vision) and instructions (language) and produce robot trajectories. We are seeing a remarkable convergence in how these work; more: open.substack.com/pub/itcanthi...
Vision-Language-Action Models and the Search for a Generalist Robot Policy
VLAs are general-purpose robotics models. But how are VLAs doing in the real world. and which ones are people using?
open.substack.com
August 27, 2025 at 11:42 PM
📣 Excited to share our #CVPR2025 Spotlight paper and my internship project @wayve: SimLingo.
A Vision-Language-Action (VLA) model that achieves state-of-the-art driving performance with language capabilities.

Code: github.com/RenzKa/simli...
Paper: arxiv.org/abs/2503.09594
May 8, 2025 at 3:25 PM
AdaDE enhances vision-language-action policies by converting dense networks into efficient MoE structures, enabling dynamic expert deactivation that retains 95.7% success in complex robotic tasks while reducing model size by 40%. https://arxiv.org/abs/2609.16503
Dense to MoE Adaptation for Compact Vision Language Action Policies
ArXiv link for Dense to MoE Adaptation for Compact Vision Language Action Policies
arxiv.org
September 22, 2026 at 5:20 PM
AdaDE transforms dense vision-language-action policies into compact mixture-of-experts models, achieving 50% parameter reduction while maintaining performance in robotic tasks. It addresses deployment challenges for resource-constrained robotics. https://arxiv.org/abs/2609.16503
Dense to MoE Adaptation for Compact Vision Language Action Policies
ArXiv link for Dense to MoE Adaptation for Compact Vision Language Action Policies
arxiv.org
September 22, 2026 at 2:10 PM
A bunch of companies are pushing really hard to get on-device inference working for local LLMs, but the thing is, if a cellphone can run a chatbot that same cheap hardware can run vision/language/action models for an automated drone that picks targets based on, oh, idk, regional dialect
My colleague Mr. Sterling reminds me that while I may drop into fond geriatric nostalgia at the drop of a pink plastic carabiner clip, the global street finding its own military uses for user-operated drones has very dire potential.
May 7, 2026 at 6:02 PM
They find that pretrained Vision-Language-Action (VLA) models are surprisingly resistant to forgetting!

Zero forgetting, or even positive backward transfer, is possible with simple experience replay.
March 10, 2026 at 4:35 AM
*That mindless automaton never made a metaphysically-correct moral decision to commit that offense

*To be technically precise, it maneuvered its 6-DoF end-effector in the Vision-Language-Action Model token-space
December 18, 2025 at 7:24 AM
Turning vision language models into vision language action models or leveraging computer vision for robotics
September 16, 2025 at 3:25 PM
This study introduces Direction-Scale Decomposition (DSD) for action representation in vision-language-action models, improving efficiency by separating motion direction from speed. Experiments show DSD boosts robotic manipulation success rates, marking progress. https://arxiv.org/abs/2609.28865
Direction-Scale Decomposition in Action Representation: Rethinking What to Tokenize for Vision-Language-Action Models
ArXiv link for Direction-Scale Decomposition in Action Representation: Rethinking What to Tokenize for Vision-Language-Action Models
arxiv.org
September 25, 2026 at 7:20 PM
Fast robotics needs fast AI inference.

InstinctFlash is an open-source inference runtime for robotics and Vision-Language-Action models.

Interesting direction for physical AI, real-time robotics and edge inference.

Full breakdown → ToolSpire

#RoboticsAI #PhysicalAI #OpenSource
September 27, 2026 at 4:25 PM
TANDEM merges task and motion planning with human teleoperation, minimizing human involvement in robot demos. This strategy boosts data collection, yielding a 2.9x increase in successful demos and improving task success rates for effective robot learning. https://arxiv.org/abs/2609.28314
TANDEM: Task and Motion Planning with As-Needed Demonstrations for Efficient Vision-Language-Action Model Fine-tuning
ArXiv link for TANDEM: Task and Motion Planning with As-Needed Demonstrations for Efficient Vision-Language-Action Model Fine-tuning
arxiv.org
September 25, 2026 at 2:20 AM
*I'm starting to think that "Vision-Language-Action Models really are "robot imiitators," but they might "imitate" robots better than previous robots have ever been "robots"
July 1, 2025 at 6:26 AM
Vision Language Action models are neat. Conceptually they appear to serve as a general solution to what used to be a whole bunch of bespoke robotics problems. I am sure there is much nuance, but seeing the capabilities of coding agents, I'm eager to see embodiment in action. arxiv.org/pdf/2405.14093
August 26, 2025 at 9:29 PM
*I'd be mildly alarmed by Vision-Language-Action AI if it wasn't being built by Google, who will promptly strangle any Google initiative that doesn't buy yachts for its shareholders #KilledByGoogle
June 27, 2025 at 6:18 AM
I have updated my tutorial on making Vision Language Action models. This tutorial starts with a basic Transformer and walks people through the steps to transform it into a full VLA that uses PaliGemma as the pretrained VLM. Links below.
February 9, 2026 at 2:15 PM
Together, we dig into the research behind that idea, from code-as-policies and vision-language-action models to the harnesses and evals needed to make these systems work in the real world.

00:00 — What Are Robot-Use Agents?
02:11 — From Vision-Language-Action Models to General LLMs
September 26, 2026 at 3:22 PM
*I suspect that LLMs will become dead-media in pretty short order, replaced by "Large Multimodal Models" and "Vision-Action-Language Models," and then "we" won't be "perpetuating" them any more than "we" "perpetuate" floppy-disks
i'm simply answering that question: how do these perpetuate? We perpetuate them. It isn't good, but people are doing a lot of the heaving lifting of meaning with "AI" at the moment. Their "agency" still needs to propagate and have meaning to stick around.
*Yeah? How are you gonna "deny their agency" when they start murdering each other over precious voltage-shortages 🤖☠️
December 15, 2025 at 7:44 AM
What can you build with a fully open robotics model in a weekend? 🤖

Robotics engineer @0xbinh.bsky.social used MolmoAct 2, our open vision-language-action model, in the voice-controlled robot that won @southparkcommons.bsky.social's AI hackathon.

Watch our interview with him ↓ 🎥
July 8, 2026 at 6:39 PM
Yeah but language models can't produce AGI

That's why we're adding vision encoders to them and heartlessly throwing them to the RL gulag and maybe soon grafting a language-action decoder to the backbone too
September 15, 2026 at 9:41 AM
Understanding the Impact of Geometric Foundation Models on Vision-Language-Action Models

Yang,Lin,Martin-Martin,Labrie,Gayaka,Kuo,Carlone

Put VGGT into VLA.
Early fusion (before VLM) > late fusion (after VLM)
Impact of 3D (VGGT) highest in single camera setup (!?!)

arxiv.org/abs/2605.24642
May 27, 2026 at 1:46 PM