LaST1.0: A Unified Latent World-Action Model with Control-Centric SpatioTemporal Reasoning

Anonymous Authors

LaST₁.₀ design roadmap with future-latent targets, encoder-free visual input, LaST-Attention, and summary results
Design Roadmap of LaST1.0. The model combines geometry-aware interaction-centric future-latent supervision, vision-encoder-free design, and structured 4D spatiotemporal attention.

Overview video

LaST1.0 overview

A short walkthrough of the control-centric objective, unified architecture, and updated evaluation evidence.

Overview

Control-centric world modeling

A

Interaction-centric targets

Instead of reconstructing an observation-complete video future, LaST1.0 supervises a compact latent representation of the evolving robot-object interaction from future wrist views.

future wrist viewsVGGT-Ω targets
B

Encoder-free visual input

A lightweight patch embedder projects multi-view images directly into the shared backbone, preserving fine-grained visual cues needed for contact-rich manipulation.

image patchesshared context
C

Structured spatiotemporal attention

LaST-Attention applies 4D-RoPE across sequence, height, width, and relative time to enable structured spatiotemporal fusion of multimodal tokens.

4D-RoPELaST-Attention

Method

Method overview

Multi-view image patches, language, and timestep information enter one decoder-only backbone. LaST-Attention organizes the multimodal sequence in four coordinates, and task-routed experts separate context, latent, and action computation.

Full LaST₁.₀ architecture, decoder block, and 4D-RoPE design
Full LaST1.0 architecture from the current manuscript.
01 / Visual input

Direct patch projection

A two-layer convolutional projector preserves local visual evidence without a deep vision encoder.

02 / Future target

Wrist-world latents

Frozen VGGT-Ω produces geometry-aware interaction-centric latent targets from future wrist-view observations.

03 / Attention

LaST-Attention + 4D-RoPE

Sequence, height, width, and relative-time subspaces preserve the structure of multimodal tokens.

04 / Capacity

Task-routed MoE

Context, latent, and action tokens use dedicated experts while action-only deployment stays efficient.

Training

Actions + future latents

Two flow-matching branches shape a shared, physically informed context representation.

Deployment

Actions only

The latent pathway and VGGT-Ω teacher are omitted for efficient closed-loop control.

Simulation results

Simulation benchmarks

RLBench

All methods are trained in the multi-task setting, and we report average success rates (%).

Models Close boxClose laptop lidToilet seat downSweep to dustpan Close fridgePhone on baseTake umbrellaTake frame off hanger Place wine at rackWater plantsAverage
OpenVLA-OFT96.070.080.050.084.038.044.020.038.026.054.6
π0.590.096.094.084.0100.018.028.086.076.038.071.0
HybridVLA92.096.0100.090.096.044.058.068.048.044.073.6
Fast-WAM94.090.096.074.0100.030.066.056.08.014.062.8
Cosmos Policy94.094.0100.060.076.078.080.070.054.058.076.4
LaST096.092.0100.082.086.074.078.066.080.064.081.8
LaST1.0 (Ours)100.0100.0100.096.0100.086.078.066.098.096.092.0

LIBERO and LIBERO-Plus — Comparison with VLA and World-Action Models

Success Rate (%) across four LIBERO suites and seven zero-shot LIBERO-Plus perturbations. Bold marks the best result in each column.

Method LIBERO LIBERO-Plus · Zero-shot
SpatialObjectGoalLongAverage CameraRobotLanguageLightBackgroundNoiseLayoutAverage
OpenVLA-OFT97.698.497.994.597.156.431.979.588.797.375.874.270.0
π0.598.898.298.092.496.978.473.680.896.294.189.084.584.4
StarVLA-α99.099.898.594.197.948.763.486.895.894.675.080.276.1
Fast-WAM98.2100.097.095.297.616.444.568.978.253.737.760.750.0
Cosmos Policy98.1100.098.297.698.575.863.381.796.588.992.782.282.2
LingBot-VA98.599.697.298.598.540.983.086.482.353.164.476.269.5
VLA-JEPA96.299.697.295.897.263.367.185.495.693.666.385.178.0
LaST099.299.698.095.698.181.436.580.286.485.779.868.973.2
LaST1.0 (Ours)98.599.899.198.599.077.968.989.494.793.197.387.386.3

LIBERO results are averaged across three evaluation runs. LIBERO-Plus directly evaluates weights trained on LIBERO without additional fine-tuning.

Higher Overall Accuracy

LaST1.0 reaches a 99.0% LIBERO average and maintains 98.5% on the Long suite, matching the best long-horizon result in the comparison.

Stronger Generalization

Without LIBERO-Plus fine-tuning, LaST1.0 obtains the best overall average of 86.3%, including 97.3% under sensor noise.

Efficient Deployment

The latent branch serves as training-time supervision and is omitted during deployment. The manuscript reports 54.7 ms action inference on an RTX 4090.

Visual representation

Effect of visual representations

On RLBench, LaST1.0's encoder-free visual interface reaches 92.0% average success. A SigLIP2 encoder reaches 84.8%, while a reconstruction-oriented VAE reaches 87.4%.

  • +7.2 pts over the SigLIP2 encoder
  • +4.6 pts over the reconstruction-oriented VAE
Action-to-vision attention maps comparing encoder-free, ViT, and VAE visual interfaces
Action-to-vision attention maps. Unlike ViT and VAE approaches, the encoder-free LaST1.0 interface keeps attention focused on task-critical regions.

Real-world manipulation

Real-world evaluation

69.3%average stage success

57.0%average task success

54.7 msaction inference on RTX 4090

Videos

Five multi-stage tasks

All five demonstrations are shown side by side for direct comparison of the long-horizon tasks.

Robot dataIn-domain
In-domain

Dual-arm Franka · Grippers

Put Object in Bag

In-domain

Dual-arm Franka · Grippers

Put Object in Drawer

In-domain

Dual-arm Tianji · Grippers

Parcel-Box Packing

In-domain

Dual-arm Tianji · Grippers

Assemble Truck Toy

In-domain

Dual-arm Tianji · Sharpa Hands

Stack Cups

Videos

Put Object in Bag

Train with a facial cleanser in the original setup; test without fine-tuning under object, background, and target-position shifts.

Bag generalizationOOD conditions
Evaluation setting

Object: replace the facial cleanser with a banana of different shape and texture. Background: scatter distractors across the workspace. Position: move the target outside the training distribution.

Unseen object

bag_object

Novel object

banana
Unseen background

bag_background

Background distractors

Unseen position

bag_position

Target-position shift

OOD

Abstract

Abstract

World-Action Models couple action learning with future prediction, but reconstruction-oriented video modeling allocates capacity to appearance details incidental to control. LaST1.0 instead predicts compact, geometry-aware future latents from short-horizon future wrist views, concentrating supervision on evolving robot-object interactions.

An encoder-free visual interface preserves patch-level evidence, while LaST-Attention with 4D-RoPE structures multimodal fusion across sequence, height, width, and relative time. A task-routed mixture of experts separates context, latent, and action computation. The latent branch provides training-time regularization and is omitted for efficient action-only deployment.