Robotic manipulation is not simply mapping an observation to an action — it requires inferring the task-relevant transition and realizing that change with an executable action pattern. UniMPA closes this transition realizability gap by unifying future prediction, bidirectional vision–action memory, and action generation on a shared action-grounded transition interface, forming an anticipate—ground—refine process.
Recent advances in Vision-Language-Action (VLA) models have improved robotic manipulation, yet observation-to-action learning remains limited by a fundamental transition realizability gap, manifested in three tightly coupled problems: (i) Transition ambiguity — visually similar observations may correspond to different manipulation phases and imply different subsequent transitions; (ii) Prediction–execution mismatch — a visually plausible predicted future does not necessarily correspond to a physically realizable transition; (iii) Experience–realization mismatch — a historically executable action pattern may not realize the intended transition in the current scene and therefore requires context-aware adaptation.
We propose UniMPA, a Unified Memory-Prediction-Action model that addresses these problems through a shared action-grounded transition interface. Persistent-Selective Future Prediction resolves transition ambiguity by modeling intended future state evolution: a persistent latent stream tracks task-level progress while a transition-critical pixel stream selectively resolves fine-grained interaction changes. A temporal Visual-Action Memory Bank then grounds the predicted transition in historically realized visual–action experience, and an Action-Visual Memory Bank retrieves visually grounded action prototypes, with Prototype-Biased Flow shifting the flow source toward a historically supported action manifold for context-aware refinement.
Extensive experiments show robust and generalizable manipulation with only 25–50% of the training epochs of the $\pi_{0.5}$ baseline, outperforming it by 1.7, 11.7, 18.5, and 12.6 percentage points on LIBERO, LIBERO-Plus, RoboTwin 2.0 Hard, and real-world tasks, respectively.
Visually similar observations may imply different transitions (a); a visually plausible future may not be physically realizable and must be grounded in executable experience (b); and an action prototype must be adapted to the current scene rather than replayed as-is (c). Existing prediction paradigms decouple predicted futures from historical executability (d), while existing memory paradigms rely on observation-centric history without transition–action correspondence (e). UniMPA (f) unifies all three through an action-grounded transition interface.
Three coupled designs, one shared transition interface.
Persistent latent supervision encodes semantic progress and long-term state evolution into pre-head transition tokens, while a Trigger Gate activates pixel decoding only at transition-critical moments — adding interaction-time detail without dense-reconstruction bias.
The future-supervised transition representation queries temporally aligned visual–action trajectories. Retrieval is conditioned on the expected state change rather than frame-level appearance, using historically realized action evolution as executable evidence.
Action history together with its aligned visual evolution retrieves a visually grounded action prototype, preserving the transition semantics needed for phase-consistent experience reuse.
Instead of replaying the retrieved prototype, UniMPA shifts the flow source toward a historically supported action manifold and lets the policy refine it for the current observation, instruction, and robot state.
UniMPA couples future-supervised transition modeling, bidirectional visual–action memory, and action generation through an action-grounded transition.
The bidirectional memory bank is pretrained with future-oriented reconstruction, retrieval-simulation, and cross-modal alignment objectives so that keys index transitions and values preserve temporally evolved vision–action dynamics.
Memory banks are frozen. Persistent latent prediction, trigger-gated pixel prediction, and flow-matching action generation are jointly optimized on top of the $\pi_{0.5}$ backbone (PaliGemma-2B VLM + 311M Gemma action expert + an 18-layer World Expert).
Latent and pixel decoding heads are dropped. Transition tokens and the retrieved action-manifold prior remain, and 10 Euler flow-matching steps produce the action chunk.
Seven task suites evaluated on a 23-DoF mobile bimanual platform, GALAXEA R1 Lite, with 25 independent trials per task.
Four simulation benchmarks and seven real-world task suites.
Success rate over 2,000 rollouts on four suites. UniMPA reaches state of the art using only 25% of the training epochs required by $\pi_0$ / $\pi_{0.5}$.
| Method | Spatial | Object | Goal | Long | Avg. |
|---|---|---|---|---|---|
| OpenVLA [CoRL'24] | 84.7 | 88.4 | 79.2 | 53.7 | 76.5 |
| OpenVLA-OFT [RSS'25] | 97.6 | 98.4 | 97.9 | 94.5 | 97.1 |
| $\pi_0$ [RSS'25] | 96.8 | 98.8 | 95.8 | 85.2 | 94.2 |
| $\pi_{0.5}$ [CoRL'25] | 98.8 | 98.2 | 98.0 | 92.4 | 96.9 |
| DreamVLA [NeurIPS'25] | 97.5 | 94.0 | 89.5 | 89.5 | 92.6 |
| VLA-JEPA [arXiv'26] | 96.2 | 99.6 | 97.2 | 95.8 | 97.2 |
| Fast-WAM [arXiv'26] | 98.2 | 100.0 | 97.0 | 95.2 | 97.6 |
| X-VLA [ICLR'26] | 98.2 | 98.6 | 97.8 | 97.6 | 98.1 |
| UniMPA | 99.6 | 99.8 | 98.8 | 96.0 | 98.6 |
Zero-shot robustness across seven perturbation dimensions (10,030 rollouts). UniMPA ranks first in both average success rate and smallest degradation $\Delta$.
| Method | Origin | Camera | Robot | Lang. | Light | Backgr. | Noise | Layout | Avg. ↑ | $\Delta$ ↓ |
|---|---|---|---|---|---|---|---|---|---|---|
| OpenVLA-OFT | 97.1 | 56.4 | 31.9 | 79.5 | 88.7 | 93.3 | 75.8 | 74.2 | 69.6 | 27.5 |
| $\pi_0$ | 94.2 | 13.8 | 6.0 | 58.8 | 85.0 | 81.4 | 79.0 | 68.9 | 53.6 | 40.6 |
| $\pi_{0.5}$ | 96.9 | 59.7 | 65.5 | 75.3 | 87.0 | 82.4 | 72.1 | 80.3 | 73.6 | 23.3 |
| X-VLA | 98.1 | 23.4 | 89.7 | 75.7 | 88.2 | 96.0 | 62.7 | 71.8 | 71.4 | 26.7 |
| DreamVLA | 92.6 | 65.0 | 40.8 | 63.5 | 85.7 | 82.6 | 84.9 | 74.0 | 69.9 | 22.7 |
| VLA-JEPA | 97.2 | 63.3 | 67.1 | 85.4 | 95.6 | 93.6 | 66.3 | 85.1 | 79.5 | 17.7 |
| FutureVLA | 98.3 | 59.7 | 66.0 | 88.2 | 97.4 | 97.8 | 77.3 | 82.6 | 79.7 | 18.6 |
| Cosmos Policy | 98.5 | 69.6 | 51.0 | 89.6 | 97.7 | 85.7 | 87.3 | 83.7 | 79.7 | 18.8 |
| MemoryVLA++ | 98.4 | 36.4 | 68.9 | 88.7 | 93.8 | 90.6 | 63.5 | 83.8 | 73.1 | 25.3 |
| Libra-VLA | 97.2 | 68.9 | 48.8 | 92.7 | 97.9 | 93.4 | 86.3 | 77.5 | 79.5 | 17.7 |
| UniMPA | 98.6 | 75.7 | 73.0 | 84.6 | 96.8 | 94.5 | 90.9 | 87.3 | 85.3 | 13.3 |
Per-suite averages under perturbation: Spatial 91.0, Object 89.7, Goal 84.3, Long 76.5.
All policies trained on the Clean setting and evaluated zero-shot on the Randomized (Hard) setting, 100 rollouts per task. UniMPA averages 58.2%, ranking first on 8/11 tasks.
| Method | Adjust B | Dump BB | Grab R | Move PA | Open L | Place BF | Place EC | Place OSc | Place PS | Shake BH | Turn S |
|---|---|---|---|---|---|---|---|---|---|---|---|
| DP3 [RSS'24] | 3% | 53% | 2% | 3% | 7% | 18% | 1% | 0% | 2% | 25% | 8% |
| $\pi_0$ [RSS'25] | 56% | 24% | 80% | 22% | 46% | 4% | 11% | 0% | 7% | 51% | 23% |
| $\pi_{0.5}$ [CoRL'25] | 54% | 30% | 63% | 14% | 37% | 45% | 53% | 20% | 7% | 94% | 20% |
| RDT [ICLR'25] | 75% | 32% | 43% | 11% | 32% | 27% | 7% | 0% | 6% | 51% | 15% |
| KAM-WM [arXiv'26] | – | 47% | 80% | 21% | 69% | 40% | 6% | 1% | 15% | 55% | 10% |
| BagelVLA [RSS'26] | 14% | 51% | 41% | 30% | 37% | 11% | 34% | 0% | 2% | 73% | 30% |
| HALO [ICML'26] | 9% | 28% | 57% | 53% | 37% | 37% | 28% | 5% | 10% | 66% | 27% |
| UniMPA | 80% | 62% | 84% | 46% | 67% | 40% | 57% | 38% | 37% | 95% | 34% |
Language-conditioned manipulation with implicit instructions. UniMPA reaches the highest average 44.0%, +4.3 points over $\pi_{0.5}$.
| Method | Painting | Book | Drink | Tube | Condiment | Flower |
|---|---|---|---|---|---|---|
| Octo [RSS'24] | 6.2 | 0.0 | 0.0 | 1.5 | 3.1 | 1.5 |
| OpenVLA [CoRL'24] | 40.2 | 7.7 | 8.5 | 7.7 | 12.4 | 13.9 |
| RDT [ICLR'25] | 35.2 | 3.1 | 7.7 | 12.4 | 21.5 | 21.5 |
| $\pi_{0.5}$ [CoRL'25] | 30.0 | 54.0 | 42.0 | 36.0 | 56.0 | 20.0 |
| VLA-Cache [NeurIPS'25] | 32.0 | 42.9 | 32.0 | 30.0 | 42.0 | 10.0 |
| VLA-IAP [arXiv'26] | 36.0 | 40.4 | 42.0 | 41.7 | 55.0 | 12.0 |
| UniMPA | 38.0 | 59.1 | 51.0 | 41.7 | 52.0 | 22.0 |
GALAXEA R1 Lite — 21 tasks across seven suites: UniMPA 77.7% TSR / 86.3% CSR vs. 65.3% / 75.6% for $\pi_{0.5}$ and 45.7% / 58.1% for OpenVLA-OFT.
AgileX Cobot Magic — seven representative tasks, 25 trials each:
| Method | Fruit Placement | Microwave | Block Stacking | Towel Folding | Bimanual Transfer | Conveyor Interception | Stack Recovery | Average | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| TSR | CSR | TSR | CSR | TSR | CSR | TSR | CSR | TSR | CSR | TSR | CSR | TSR | CSR | TSR | CSR | |
| OpenVLA-OFT† | 32.0 | 50.0 | 48.0 | 62.7 | 40.0 | 59.0 | 44.0 | 58.7 | 36.0 | 57.0 | 28.0 | 44.0 | 32.0 | 55.2 | 37.1 | 55.2 |
| $\pi_{0.5}$† | 52.0 | 67.0 | 72.0 | 84.0 | 68.0 | 78.0 | 72.0 | 80.0 | 72.0 | 81.0 | 40.0 | 60.0 | 60.0 | 74.4 | 62.3 | 74.9 |
| UniMPA | 64.0 | 80.0 | 80.0 | 90.7 | 76.0 | 90.0 | 76.0 | 88.0 | 80.0 | 86.7 | 76.0 | 85.0 | 72.0 | 84.8 | 74.9 | 86.4 |
† denotes our reproduced results. TSR: Task Success Rate. CSR: Cumulative Success Rate.
LIBERO avg. / real-world TSR when using latent-only prediction — persistent and selective streams are complementary.
Removing memory entirely. Future prediction alone does not specify an executable realization.
Removing temporal modeling, which is what keeps retrieval phase-consistent (LIBERO-Long drops 6.0).
Directly copying the retrieved action instead of Prototype-Biased Flow — history cannot be replayed as-is.
Replacing future-oriented memory pretraining with current-state reconstruction.
Random triggering instead of the combined latent + action Trigger Gate.
@article{li2026unimpa,
title = {UniMPA: A Unified Memory-Prediction-Action Model via Action-Grounded Transition Modeling},
author = {Li, Wei and Shao, Rui and He, Jie and Zhang, Lingsen and Liu, Ziwei and Nie, Liqiang},
journal = {arXiv preprint},
year = {2026}
}