ATI-VLA: Action-Centric Predictive
Vision–Language–Action Models via
Actionable Alignment Then Adaptive Injection

Yijie Zhu1,2, Rui Shao1,5,*, Jie He1, Wei Li1, Bo Zhao2, Yelin Wang2,
Xiaochen Yuan3, Tao Tan3, Miao Zhang1, Xiaojiang Peng4, Zitong Yu2,6,*
1Harbin Institute of Technology, Shenzhen 2Institute for Artificial Intelligence, Great Bay University 3Macao Polytechnic University
4Shenzhen Technology University 5Shenzhen Loop Area Institute
6Dongguan Key Laboratory for Intelligence and Information Technology

* Corresponding authors: Rui Shao and Zitong Yu

NeurIPS 2026

Abstract

Predictive Vision–Language–Action (VLA) models aim to improve robotic manipulation via future observation or world dynamics forecasting. However, existing approaches often fail to realize this potential and underperform direct action prediction models. We argue that these limitations stem from modality misalignment between observations and actions, together with joint optimization conflicts that drive learning away from an action-centric objective. To this end, we introduce ATI-VLA, an Action-Centric Predictive Vision–Language–Action framework via Actionable Alignment Then Adaptive Injection. Specifically, it follows a two-step design: 1) Actionable Representation Alignment via a Shared Codebook. It aligns predictive observation and action representations by mapping both modalities into a shared discrete latent space via a unified codebook, making predictive observation latents readily usable for action generation and mitigating modality misalignment. 2) Action-Centric Adaptive Injection of Predictive Latents. Building upon this, it then injects predictive observation latents into action decoding as explicit predictive priors via a lightweight adaptive side-path, enabling adaptive predictive guidance under a single action-centric objective. Extensive experiments on both simulation and real-world robotic tasks demonstrate that ATI-VLA achieves state-of-the-art performance with faster convergence.

Comparison of general VLA, predictive VLA, and ATI-VLA, with analysis of observation–action modality gaps and training convergence.

Method

ATI-VLA follows an align-then-inject design: first learn action-grounded predictive representations, then use them to guide action generation under a single action objective.

Stage 1Actionable Representation Alignment

Observation and action encoders use a shared codebook and a unified decoder to predict future observations and actions; the latent visualization shows improved cross-modal alignment.
The current observation and action are encoded and quantized through a shared discrete codebook. A unified decoder jointly predicts future observations and actions, learning action-grounded predictive prototypes through prediction rather than input reconstruction. The visualization on the right illustrates how observation and action latents become aligned during training.

Stage 2Action-Centric Adaptive Injection

A frozen observation encoder and codebook supply predictive latents to a lightweight side-path, which adaptively injects them into selected action-decoder layers.
After alignment, the alignment module is frozen. A lightweight adaptive side-path transforms predictive observation latents and injects them into selected LLM layers to guide action decoding. The action policy and side-path are optimized solely under the action objective.

Experiments

We evaluate ATI-VLA on LIBERO, RoboTwin 2.0, and real-world bimanual manipulation with Galaxea R1 Lite and AgileX Cobot Magic.

LIBERO

Scroll to view the full table, or tap to open the image.

Results on LIBERO. ATI-VLA achieves a 97.9% average success rate across the Spatial, Object, Goal, and Long suites, including 96.2% on LIBERO-Long. The table reproduces the results reported in the paper.

RoboTwin 2.0

Scroll to view the full table, or tap to open the image.

Results on RoboTwin 2.0. ATI-VLA achieves a 72.3% average success rate, improving over our reproduced OpenVLA-OFT baseline by 13.5 percentage points under the same experimental settings.

Real-World Evaluation

Scroll to view the full table, or tap to open the image.

Results across two robot platforms. ATI-VLA achieves average task-completion rates of 70% on Galaxea R1 Lite and 72% on AgileX Cobot Magic across five long-horizon tasks, including OOD T-shirt folding. Rates measure completion of the entire task sequence.

Convergence and Efficiency

Training Convergence

Scroll to view all four plots, or tap to open the image.

Effect of actionable alignment. ATI-VLA converges faster and achieves stronger performance than the variant without alignment across the four LIBERO suites, supporting the role of alignment in making predictive representations useful for action learning.

Performance–Efficiency Trade-off

Scroll to view the full table, or tap to open the image.

Performance–efficiency trade-off. ATI-VLA improves task performance with only 0.7% additional FLOPs over OpenVLA-OFT, while reducing total training time from 30.2 to 16.2 hours (46.4%). Training time includes both alignment and action-policy training.

Real-World Demos

Examples of long-horizon bimanual manipulation with ATI-VLA.

Object Packing

Instruction: Pack the object on the table into the bag.

T-shirt Folding (OOD)

Instruction: Fold the T-shirt.
Generalization to an unseen garment appearance and size.

Conclusion

ATI-VLA addresses the observation–action modality gap and joint optimization conflicts in predictive VLA through an align-then-inject framework. A shared codebook learns action-grounded predictive representations, which are then adaptively injected into action decoding under a single action objective. Experiments across simulation benchmarks and two real-world platforms demonstrate improved task performance and faster convergence.