Vision-Language-Action · Research Project

World-Calibrated
Proposal-to-Action Flow

ProAct turns recent robot motion into a scene-aware proposal, calibrates it with prospective world reasoning, and refines it into executable actions.

Jie He1,*, Wei Li1,*, Junwen Tong2,*, Rui Shao1,†, Wei-Shi Zheng3, Liqiang Nie1

1Harbin Institute of Technology, Shenzhen · 2ZTE Corporation · 3Sun Yat-sen University
*Equal contribution · †Corresponding author

Start from motion.
Reason about the world.

Flow-based VLA policies typically start every action chunk from an uninformed isotropic Gaussian. ProAct changes the starting point itself: recent executed motion becomes a soft local hypothesis, while a proposal-aware World Expert decides how that hypothesis should be centered, scaled, and shaped for the current scene.

Robot manipulation is locally continuous even when task-level behavior is multimodal. The most recent executed actions therefore contain useful information about the direction and scale of the next motion. However, simply extrapolating this motion is unreliable near contact, reorientation, gripper actuation, and task-phase transitions.

ProAct treats motion continuity as a proposal rather than a commitment. It evaluates the proposal against a task-consistent future, then uses that compatibility to shape the source from which the final action trajectory is generated.

ProAct motivation
Recent motion provides continuity; world reasoning provides correction.

Propose → Calibrate → Refine

The three experts preserve the flexibility of flow matching while making its source distribution predictive and world-aware.

Given the current visual-language context and a history of executed actions, ProAct factorizes action generation into three connected stages. Each stage has a distinct role: estimate a plausible motion, understand whether it is compatible with the scene, and finally produce the executable action chunk.

The calibrated source is proposal-centered and anisotropic. Its refinement extent controls how far the policy can depart from recent motion, while its low-rank geometry allocates correction along coupled translation, rotation, and gripper directions.

ProAct framework
ProAct connects local motion continuity with future-aware visual reasoning before the final flow-matching refinement.
01

Proposal Expert

Maintains an exponentially smoothed motion estimate, extrapolates it over the action horizon, and grounds the proposal in VLM features with a shallow endpoint flow step.

02

World Expert

Uses the proposal as a soft motion prior while predicting the task-consistent future latent with a frozen V-JEPA 2 target and estimating proposal-specific calibration parameters.

03

Action Expert

Samples from the calibrated anisotropic source and performs the final flow-matching transport, retaining the base VLA's ability to correct an imperfect proposal.

More capable.
Less denoising.

ProAct is evaluated across simulation and real-world manipulation tasks. The proposal-centered source improves both task performance and inference efficiency compared with the isotropic baseline.

On simulated benchmarks, ProAct reaches 98.4% average success on LIBERO, 86.8% on LIBERO-Plus, and 60.3% on the randomized hard subset of RoboTwin 2.0. On physical robots, it achieves strong task and cumulative success rates across GALAXEA R1 Lite and AgileX Cobot Magic. The efficiency gains come from starting flow refinement closer to the relevant action manifold, allowing the model to use half as many denoising steps as π0.5.

98.4%LIBERO average
86.8%LIBERO-Plus average
60.3%RoboTwin randomized
50%fewer denoising steps
25.8%lower latency
34.8%higher throughput
GALAXEA R1 Lite78.7% TSR · 86.1% CSR
AgileX Cobot Magic74.7% TSR · 82.3% CSR

Simulation benchmarks

Average performance against representative vision-language-action baselines.

MethodSpatialObjectGoalLongLIBERO Avg. ↑ CameraRobotLanguageLightBackgroundNoiseLayoutLIBERO-Plus ↑
OpenVLA-OFT97.698.497.994.597.156.431.979.588.793.375.874.269.6
π0.598.898.298.092.496.959.765.575.387.082.472.180.373.6
WorldVLA87.696.283.460.081.80.127.941.643.717.110.938.025.0
UniVLA95.498.893.694.095.41.846.269.669.081.021.231.942.9
VLA-Adapter97.899.297.295.097.336.237.974.670.676.158.069.759.1
DreamVLA97.594.089.589.592.665.040.863.585.782.684.974.069.9
VLA-JEPA96.299.697.295.897.263.367.185.495.693.666.385.179.5
Fast-WAM98.2100.098.295.297.634.055.188.990.044.933.373.459.0
HoloBrain—————66.349.065.994.993.373.178.272.6
ProAct99.499.698.496.098.476.180.387.497.895.590.286.286.8

RoboTwin 2.0 · randomized hard

Zero-shot randomized evaluation with 100 rollouts per task.

MethodPut BottlesDual ShoesDiverse BottlesBurger FriesOpen LaptopBlocks RGBDual BottlesDump BinStamp SealBread BasketAvg. ↑
π01306446512244411.8
π0.5973703535642235628.6
Spatial Forcing4123871157477025334.6
X-WAM21726366365540164028.3
X-VLA20123035164514125226.3
Abot-M0261575564231968194734.3
EventVLA3119723686431516.7
Xiaomi Robotics-02191227248284932320.4
FastWAM00000001000.1
GalaxeaVLA391132131214023216.4
AHA-WAM100091127034.2
DP3210118701530110.2
starVLA02151122204056.1
ProAct6939578476466378296260.3

Real-world manipulation

Average Task Success Rate (TSR) and Cumulative Success Rate (CSR) over six tasks.

MethodGALAXEA TSR ↑GALAXEA CSR ↑AgileX TSR ↑AgileX CSR ↑
OpenVLA-OFT42.755.637.155.2
π050.762.250.062.0
π0.564.774.062.374.9
ProAct78.786.174.782.3

What makes
ProAct work?

We examine whether the learned proposals and calibrations are interpretable, and how proposal-centered transport improves action generation across simulated and real-world tasks.

01

Visual attention of the Proposal Expert

Attention maps on LIBERO (top) and RoboTwin 2.0 (bottom). Warmer colors highlight manipulated objects, interaction regions, and receptacles, showing that the Proposal Expert remains focused on task-relevant scene cues.

Visual attention maps of the Proposal Expert on LIBERO and RoboTwin 2.0
02

Calibrated source correlation

Action samples and correlation ellipses visualize the calibrated source geometry. Nearly circular contours indicate weak correlation, while tilted ellipses reveal positive or negative coupling between action dimensions.

Calibrated source correlation visualization
03

World-calibrated proposal refinement

In the left panel, the proposal is refined in 3D translation space. In the right panel, normalized trajectories and calibrated support are shown across seven action dimensions: stable dimensions remain close to the proposal, while incompatible ones bend toward the demonstrated future.

World-calibrated proposal refinement visualization
04

Per-checkpoint real-world ablations

Per-checkpoint success is reported across the six GALAXEA R1 Lite tasks. The three panels ablate motion proposal (left), world calibration (middle), and source calibration (right); the full ProAct model obtains the highest bars, especially at later transition and terminal checkpoints.

Per-checkpoint real-world ablation visualization
Additional ProAct analysis visualizations
Additional examples of proposal refinement across tasks and execution phases.

From proposal
to action.

Representative rollouts show ProAct handling object retrieval, articulated drawers, contact-rich wiping, and repeated color-based placement across different task phases.

These tasks deliberately include situations where local continuation is not enough: the robot must re-ground after opening a drawer, switch from grasping to placement, maintain contact during wiping, or preserve object identity over a sequence of placements. The videos illustrate how proposal, calibration, and refinement work together in closed-loop control.

Drawer Toy Retrieval

Drawer Manipulation

Table Wiping

Color Sorting

Pot Lid Retrieval

Block Sorting