arXiv:2609.23578

AR-WAM: A Visual-Conditioned Agent-Ready World Action Model for Robotic Manipulation

1The Hong Kong University of Science and Technology, 2Shenzhen University of Advanced Technology, 3Sangfor Technologies Inc., 4Mininglamp Technology, 5MMLab, The Chinese University of Hong Kong,
6The University of British Columbia, 7Astribot
*Equal contribution †Corresponding authors
0.5B
parameters, frozen visual encoder,
no language encoder
87.2%
average success on RoboTwin 2.0,
best 84.2% in randomized setting
14.1 ms
inference latency per action chunk,
4.4× faster than π0.5
+36.7%
real-robot long-horizon success
over the strongest baseline

Abstract

As AI agents become increasingly capable, agent-driven robotic control is emerging as a compelling paradigm. However, prevailing vision-language-action (VLA) models and world action models (WAMs) still rely on natural-language instructions to specify manipulation tasks, an ill-suited interface for agent-driven control: referentially ambiguous, spatially imprecise, redundant with the agent’s inherent language understanding, and entangling intent with execution.

We present AR-WAM, a visual-conditioned, agent-ready world action model that replaces language with two complementary conditions: a visual grounding prompt (a bounding box of the target) denoting the interaction object and location, and a learnable operation token dictating the atomic skill to execute. Our compact 0.5B-parameter model, with a frozen pretrained visual encoder and no language encoder, predicts scene evolution within compact latent states while decoding actions, exposing the policy’s intent through explicit, supervisable reasoning signals. A model-agnostic compatibility layer provides three primitives (detect, execute, and query) so that local VLMs or online agent APIs can drive the policy directly, with long-horizon memory and closed-loop error recovery delegated to the agent side.

On RoboTwin 2.0, RMBench, and a real Astribot S1 dual-arm platform, AR-WAM matches the strongest baselines on standard manipulation (87.2% average success) and outperforms them on memory-dependent and real-robot long-horizon tasks, improving success rates by 5.9% and 36.7%, respectively, while maintaining the lowest inference latency (14.1 ms).

Method Overview

AR-WAM method overview

An agent grounds the target and issues a visual prompt g together with an atomic operation s. A compact 0.5B world action model (a 0.2B transformer shared by a reasoning expert and a world action expert, plus a frozen 0.3B vision encoder) predicts an action chunk together with dual-level foresight signals. The reasoning expert runs once per control cycle, and the world action expert then performs multi-step denoising while reusing its KV cache.

Visual Prompt, Not Language

A bounding-box visual grounding prompt specifies where to act, and a learnable operation token specifies which atomic skill to perform (pick, place, click, handover, …). Spatial conditioning replaces text: unambiguous, precise, and free of redundant language understanding inside the policy.

Reasoning Expert + Dual-Level Foresight

A reasoning expert externalizes task intent as explicit, supervisable signals (target mask, end-effector poses, and a task abstract), while visual and semantic foresight predictions keep action generation future-aware. Both experts share one compact transformer.

Agent-Ready Compatibility Layer

Three model-agnostic primitives (detect, execute, and query) let any VLM/LLM agent drive the policy directly. Long-horizon memory and closed-loop failure recovery are delegated to the agent, keeping the policy Markovian and compact.

Dual-Level Foresight

Future-aware action generation in compact latent form

A reasoning expert runs once per control cycle and externalizes task intent as explicit, supervisable signals: a target mask, end-effector poses, and a task abstract. On top of it, dual-level foresight keeps action generation future-aware: an offline foresight-gist extractor compresses the temporal evolution between two endpoint frames into a compact latent gist, with no actions or intermediate frames, and the policy is aligned to predict both the visual foresight of the future scene and the semantic foresight of the target within it. Removing foresight alignment costs 8.6% average success, and replacing the compact gist with raw future patch features drops 10.8%, below even the foresight-free variant. Reasoning attends to the object and the target; foresight attends to where motion will occur.

Foresight prediction visualization

Current RGB, foresight RGB, and reasoning masks for the current and foreseen scene, on RoboTwin (Handover Block) and RMBench (Cover Blocks).

Agent-Policy Closed Loop

Event-triggered agent invocation through three primitives

detect execute query

The agent is invoked only on trigger events: end-effector displacement, gripper open/close, or a timeout. It grounds the target with its own perception (detect), rolls out the policy until the next event (execute), and receives structured feedback for its next reasoning step (query). Commands and feedback accumulate in the agent’s history, so long-horizon memory lives entirely on the agent side. On a grasp failure or a dropped object, the agent re-localizes the target and retries under a fixed budget, and the policy’s decoded intent helps diagnose whether a failure stems from wrong intent or improper execution.

Failure recovery through the agent-policy loop

Failure recovery through the agent-policy loop: the agent detects the failure, re-issues an updated visual prompt, and retries, with no hardwired recovery behaviors.

Recovery rollouts on RoboTwin 2.0 (×5 speed): after a failed grasp, the agent re-grounds the target and issues a fresh visual prompt to retry.

Experiments

RoboTwin 2.0 and RMBench in simulation, plus a real Astribot S1 dual-arm platform

RoboTwin 2.0

Ten dual-arm long-horizon tasks composing rich combinations of atomic skills, with instructions issued by a benchmark expert equipped with an object detector. AR-WAM attains the best average success rate in the randomized setting (84.2%) and stays within 0.6% of the strongest baseline in the clean setting (87.2% vs. 87.8%), with visual prompts instead of language instructions. Replacing the visual prompt with text under identical training drops average success by 13.0%.

Task π0.5 X-VLA Motus Fast-WAM LiLa-WAM AR-WAM (Ours)
CR CR CR CR CR CR
Blocks Ranking Size4926677475639498928810094
Handover Block665773378673998096909892
Handover Mic989700786310010010098100100
Hanging Mug181723273838655656445462
Place Can Basket626249528176726780729082
Place Dual Shoes757579889387888860548678
Put Bottles Dustbin847974778179938292949692
Put Object Cabinet807946488871948292927870
Scan Object726514366766968694907686
Stack Bowls Three777176867987778688829486
Average68.162.850.152.576.670.387.882.585.080.487.284.2

Success rates (%) over 50 trials per task, per setting. C: clean setting; R: randomized setting. Bold: best; underlined: second best.

Rollouts on all ten tasks (clean vs. randomized), with the agent-issued bounding box and the operation token shown alongside the policy’s reasoning readout.

RMBench

Memory-dependent tasks that test the cooperation between the policy and the agent, driven by off-the-shelf agents (a locally deployed Qwen3.8-27B, AR-Q, and the online API Kimi-K3, AR-K) through the model-agnostic interface. On the memory-heavy M(n) tasks, AR-WAM reaches 52.5% / 63.5% average success vs. 38.7% for the strongest memory-augmented baseline, thanks to cleanly delegating memory management to the upstream agent.

Task π0.5 Fast-WAM Goal2Skill Mem-0 Cortex LiLa-WAM AR-K AR-Q
Observe and Pick Up908414104484
Rearrange Blocks1303889100206286
Put Back Block110–90100145888
Swap Blocks240–6799106882
Swap T157–1463181816
M(1) Tasks Average14.41.423.052.875.214.450.071.2
Battery Try1620462837243852
Blocks Ranking Try6266018–301834
Cover Blocks00–68–07278
Press Button001002008290
M(n) Tasks Average5.511.538.728.528.513.552.563.5
Total Average10.45.932.442.061.914.051.167.8

Success rates (%) over 50 trials per task. AR-K / AR-Q: AR-WAM integrated with Kimi-K3 / Qwen3.8-27B as the agent. –: not applicable / not reported. Bold: best; underlined: second best.

The agent-policy closed loop on memory-dependent tasks: the agent reasons over its history and issues a bounding box plus an atomic skill; the policy executes and reports back.

Real-World Evaluation

Three challenging long-horizon tasks on a real Astribot S1 dual-arm platform, all requiring agent-side reasoning and memory: pressing bells a specified number of times, picking up a specified bottle twice, and placing an egg into a specified cell. For each task, the simulation-trained policy is fine-tuned on 100 real-world demonstrations and driven by a locally deployed Qwen3.8-27B. AR-WAM achieves the highest success on all three tasks (71.7% average vs. 35.0% for π0.5).

Real-robot rollouts on Astribot S1
Task π0.5 LiLa-WAM AR-WAM (Ours)
Press bells X times51070
Pick up X bottle twice454065
Place egg into X cell554580
Average35.031.771.7

Success rates (%) over 20 trials per task. Bold: best.

Real-robot rollouts of the three tasks (×5 speed). The red box is the agent-issued visual prompt.

Efficiency

Measured on a single RTX 4090 in bf16 with identical inputs. AR-WAM attains the lowest end-to-end latency (14.09 ms per action chunk), 4.4× faster than π0.5 and 22.7× faster than Fast-WAM, benefiting from the shared reasoning-denoising forward pass and KV-cache reuse.

Method Encoder & Misc. KV Cache ODE (per step) End-to-end
π0.519.7423.362.11 (×9)62.09
Fast-WAM27.5552.5626.58 (×9)319.33
LiLa-WAM4.77–1.02 (×10)14.97
AR-WAM (Ours)4.413.740.66 (×9)14.09

Per-stage inference latency (ms). ODE: latency per denoising step, with the number of steps in parentheses. Bold: fastest; underlined: second fastest.

BibTeX

@article{jiang2026arwam,
  title={AR-WAM: A Visual-Conditioned Agent-Ready World Action Model for Robotic Manipulation},
  author={Jiang, Yicheng and Gan, Zesen and Wang, Xiaobo and He, Tianlun and Zhao, Chenxu and Wu, Minghui and Wang, Xinyue and Wang, Jiaxu and He, Junhao and Wang, Jianan and Shao, Qiming},
  journal={arXiv preprint arXiv:2609.23578},
  year={2026}
}