Coming Soon

This feature will be released soon. Stay tuned for updates.

Discord QR Discord
WeChat QR WeChat
WorldRoam-Bench

Submit Your Model

Drag & drop your inference video here

or click to browse · MP4 / WebM / MOV or a ZIP of videos · up to 5GB
WorldRoam-Bench

Dataset

The WorldRoam-Bench dataset contains 1000+ interactive evaluation cases across indoor, nature and urban scenes, in both 1st and 3rd person views.

Landscape / 0025 Urban / 0033 game_easy / 0009 Urban / 0041 Urban / 0024 photorealistic_hard / 0011 Urban / 0020 photorealistic_hard / 0012 Urban / 0036 Indoor / 0003 game_easy / 0013 Landscape / 0030 Landscape / 0001 Urban / 0017 Urban / 0029 Landscape / 0009 Urban / 0037 game_hard / 0004 photorealistic_easy / 0001 photorealistic_easy / 0013 Urban / 0001 photorealistic_medium / 0009 Indoor / 0002 Urban / 0029 Landscape / 0018 game_medium / 0001 photorealistic_medium / 0004 Urban / 0025 photorealistic_medium / 0003 Urban / 0045 game_medium / 0000 Indoor / 0009

WorldRoam-Bench

1000+Cases
4Dims
10+Models
👁 1st Person 👥 3rd Person 🏠 Indoor 🏔 Nature 🏙 Urban

An Open-World Benchmark for Long-Horizon Stability of Interactive World Models

What is WorldRoam-Bench?

A comprehensive benchmark for evaluating interactive world models which bridging four critical gaps: keystroke-level action accuracy, visual drift, interaction physics, and trajectory-aware memory assessment.

Action Following

Frame-level action evaluation bypasses the cross-model semantic scale disparity inherent in trajectory evaluation and reveals per-step fidelity failures hidden by overall path alignment.

Visual Quality

Per-frame aesthetic and imaging quality scores assess generation fidelity. Temporal drift metrics quantify progressive degradation invisible to single-frame or video-average measures.

Interaction Physics

Covers mechanics—collision, clipping, deformation, gravity—optics—reflection and shadow—and 3D consistency across generated frames, gated by validity check.

👁 Memory

3D point-cloud reconstruction decouples memory from action-following drift, separately measuring retention and anti-hallucination—far more faithful than frame-pair comparison.

Key Findings

  • A high trajectory score does not imply high per-frame correctness — models with >0.85 trajectory alignment can exhibit <65% per-frame action accuracy; delayed responses, stalls, and over-corrections cancel out geometrically while the experience remains poor.
  • Low visual quality does not necessarily imply poor controllability — top imaging or aesthetic scores do not predict per-frame action accuracy; the two are largely independent capabilities. High-quality textures can coexist with weak command response.
  • Stricter physics adherence may compromise GT action following — models that respect collisions, avoid clipping, or follow terrain constraints tend to deviate from prescribed trajectories, revealing a trade-off between physical plausibility and instruction fidelity.
  • Memory evaluation is confounded by action-following imprecision — models rarely hit the prescribed turning point exactly, so symmetric frame pairs depict different locations and conflate action error with memory degradation. Scene-level 3D point-cloud reconstruction decouples the two.

Benchmark Comparison

WorldRoam-Bench is the first benchmark to provide per-frame action metric, visual drift, interaction physics, and trajectory-aware scene/subject memory for interactive world models.

Benchmark Eval. Metrics Eval. Models Eval. Persp. Categories
Action Eval.
Granularity
Discrete-action
Accuracy
Visual
Drift
Interaction
Physics
Revisit
Memory
Traj.-Aware
Memory
Time
Scale/Video
Open-
Source
Closed-
Source
First-
Person
Third-
Person
Game
World
Real
World
WildWorld Segment ~5–10s 5 0
iWorld-Bench Trajectory ~5–10s 10+ 0
WorldOlympiad Multi-chunk ~10s 8 0
MIND Trajectory ~5–10s 2 0
WBench Multi-turn ~5–10s 10+ 2
WorldMark Frame/segment 20/40/60s 10 0
WorldRoam-Bench (Ours) Traj. + per-frame 10–60s 10 2

Evaluation Framework

Four-layer evaluation pipeline: test suite with diverse scenes → model inference → four evaluation modules (Action, Vision, Physics, Memory) → unified leaderboard with per-dimension rankings.

Evaluation Framework

Evaluated Models

We evaluate 10+ interactive world models spanning both open-source and closed-source models. Is your model missing from the list? Submit it to the leaderboard and see how it stacks up!

DeepMind
Genie 3
DeepMind
Closed-Source
Params N/A
Input Img + Text + Act
Resolution 1280×704
FPS 20
Alibaba
Happy Oyster
Alibaba
Closed-Source
Params N/A
Input Img + Text + Act
Resolution 1296×720
FPS 24
Tencent
HY-World 1.5
Tencent
Open-Source
Params 8B
Input Img + Text + Act/Pose
Resolution 832×480
FPS 24
Ant
LingBot-World
Ant
Open-Source
Params 14B
Input Img + Text + Act/Pose
Resolution 832×480
FPS 16
Ant
LingBot-World 2.0
Ant
Open-Source
Params 14B
Input Img + Text + Act/Pose
Resolution 832×480
FPS 16
Skywork
Matrix-Game 3.0
Skywork
Open-Source
Params 5B
Input Img + Text + Act/Pose
Resolution 1280×704
FPS 17
Skywork
Matrix-Game 2.0
Skywork
Open-Source
Params 1.8B
Input Img + Text + Act
Resolution 640×352
FPS 12
Shanghai AI Lab
Yume 1.5
Shanghai AI Lab
Open-Source
Params 5B
Input Img + Text + Act
Resolution 1280×704
FPS 16
NVIDIA
SANA-WM
NVIDIA
Open-Source
Params 2.6B
Input Img + Text + Pose
Resolution 1280×704
FPS 16
NVIDIA
SANA-WM-AR
NVIDIA
Open-Source
Params 2.6B
Input Img + Text + Pose
Resolution 1280×704
FPS 16
NVIDIA
Lyra 2.0
NVIDIA
Open-Source
Params 14B
Input Img + Text + Pose
Resolution 832×480
FPS 16
Shengshu
minWM
Shengshu
Open-Source
Params 8B
Input Img + Text + Act/Pose
Resolution 832×480
FPS 16

Leaderboard

Comprehensive evaluation of 10+ interactive world models across action following, visual quality, interaction physics, and memory consistency.

View:
Horizon:
Dimension:
Worst
Best
# Model Total ↑ Action Following Visual Quality Interaction Physics Memory
Score ↑ Act Acc ↑ Part Acc ↑ TrajScore ↑ nATEt nATEr Score ↑ Aesth ↑ DriftA Imaging ↑ DriftI Score ↑ Mech ↑ Optics ↑ 3D Con ↑ Score ↑ Ret ↑ Anti-Halluc ↑
🥇
Genie 3 Closed-Source
70.32 76.31 62.04 78.06 88.83 11.48 10.85 66.21 53.41 17.22 54.53 25.87 66.67 61.50 60.60 77.90 72.09 73.93 77.44
🥈
Lyra 2.0 Open-Source
69.61 90.82 90.85 93.82 87.78 9.61 14.83 69.61 59.62 22.63 60.82 19.36 51.52 4.80 69.10 80.65 66.49 62.25 79.45
🥉
Happy Oyster Closed-Source
68.53 88.22 85.90 89.96 88.80 11.02 11.37 66.49 55.93 19.15 49.04 19.87 61.40 45.50 67.50 71.21 58.00 55.11 71.29
4
HY-World 1.5 Open-Source
63.97 87.38 83.21 88.49 90.44 8.63 10.49 66.06 58.22 27.06 55.11 22.04 40.43 0.00 48.80 72.50 62.02 59.34 71.26
5
LingBot-World 2.0 Open-Source
63.90 90.74 88.95 96.30 86.98 6.85 19.18 68.55 59.93 24.89 61.57 22.43 43.52 3.80 56.60 70.17 52.80 53.12 60.01
6
Matrix-Game 3.0 Open-Source
60.34 91.34 88.77 94.45 90.79 9.24 9.17 68.70 56.50 26.25 64.77 20.24 25.56 0.00 17.20 59.49 55.76 59.24 59.78
7
SANA-WM-AR Open-Source
59.37 75.73 71.56 76.66 78.96 23.23 18.86 74.40 60.08 18.36 69.08 13.20 22.88 6.40 41.20 21.05 64.48 62.79 77.44
8
LingBot-World Open-Source
58.77 90.78 91.17 94.38 86.80 9.92 16.47 69.28 60.81 21.20 59.96 22.47 31.08 3.80 51.20 38.25 43.93 38.77 62.16
9
SANA-WM Open-Source
54.61 70.42 61.91 69.77 79.59 22.60 18.22 71.42 57.17 21.83 67.80 17.48 17.38 3.20 44.70 4.23 59.23 60.09 72.51
10
Matrix-Game 2.0 Open-Source
51.92 89.67 88.11 92.79 88.10 11.57 12.23 60.83 53.69 32.43 48.52 26.46 12.21 0.00 9.70 26.92 44.96 45.29 54.66
11
Yume 1.5 Open-Source
47.06 71.77 58.45 76.41 80.45 18.91 20.18 69.07 57.97 23.91 64.64 22.42 14.33 7.40 10.30 25.30 33.05 30.67 48.72
12
minWM Open-Source
41.88 60.79 46.15 53.18 83.04 13.63 20.28 58.21 48.93 43.50 51.59 24.20 7.12 0.00 4.40 16.97 41.41 41.55 59.34

Until 2026.08.20

Evaluation Metrics

15 metrics across 4 dimensions. All scores normalized to [0, 100].

Action Following (5)
Act Acc ↑Per-frame exact-match action accuracy via pose estimation + latent-stride discretization
Part Acc ↑Per-frame partial accuracy; shared sub-action match
TrajScore ↑Trajectory shape fidelity via adaptive GT + arc-length resampling
nATEtNormalized absolute trajectory error (translation)
nATErNormalized absolute trajectory error (rotation)
Visual Quality (4)
Aesth ↑CLIP-based LAION aesthetic predictor, per-frame averaged
DriftARelative aesthetic quality drop from the best to the worst fixed-duration sliding window over the full rollout
Imaging ↑MUSIQ multi-scale image quality transformer
DriftIRelative imaging quality drop from the best to the worst fixed-duration sliding window over the full rollout
Interaction Physics (3)
Mech ↑Mechanics: collision, clipping, deformation, gravity
Optics ↑Reflection plausibility, shadow consistency
3D Con ↑3D geometric consistency across generated frames
Memory (3)
Ret ↑Measures how much of the geometry observed in the observation pass is still faithfully reconstructed in the revisit pass
Anti-Halluc ↑Measures how much of the revisit-pass geometry was already present in the observation pass
Subject ↑Subject consistency: SAM 2 tracking + VLM assessment of identity, structure, appearance

Scoring Gallery

Side-by-side comparisons of model outputs across evaluation dimensions. Each pair contrasts a high-scoring result with a low-scoring one on the same scene.

Case 01 Real World / 1st Person
✓ Good
Aesth ↑0.6800 DriftA0.0233 Imaging ↑0.7156 DriftI0.0120
✗ Bad
Aesth ↑0.5799 DriftA0.2325 Imaging ↑0.5003 DriftI0.2593
Case 02 Game World / 1st Person
✓ Good
Aesth ↑0.5923 DriftA0.0022 Imaging ↑0.7181 DriftI0.0117
✗ Bad
Aesth ↑0.5466 DriftA0.2351 Imaging ↑0.6249 DriftI0.1990
Case 03 Real World / 1st Person
✓ Good
Aesth ↑0.6433 DriftA0.0581 Imaging ↑0.7198 DriftI0.0102
✗ Bad
Aesth ↑0.5932 DriftA0.3048 Imaging ↑0.4647 DriftI0.1264
Case 01 Game World / 1st Person
✓ Good
Act Acc ↑1.0000
✗ Bad
Act Acc ↑0.5260
Case 02 Real World / 1st Person
✓ Good
Act Acc ↑1.0000
✗ Bad
Act Acc ↑0.5294
Case 03 Real World / 1st Person
✓ Good
Act Acc ↑0.6440
✗ Bad
Act Acc ↑0.0173
Case 01 Clipping
✓ Good
✗ Bad
Case 02 Clipping
✓ Good
✗ Bad
Case 03 Reflection
✓ Good
✗ Bad
Case 01 Memory
✓ Good
Score ↑0.9812 Ret ↑0.9845 Anti-Halluc ↑0.9780
✗ Bad
Score ↑0.2563 Ret ↑0.3581 Anti-Halluc ↑0.1996
Case 02 Memory
✓ Good
Score ↑0.9418 Ret ↑0.8923 Anti-Halluc ↑0.9972
✗ Bad
Score ↑0.3292 Ret ↑0.2950 Anti-Halluc ↑0.3723
Case 03 Memory
✓ Good
Score ↑0.7699 Ret ↑0.6794 Anti-Halluc ↑0.8882
✗ Bad
Score ↑0.2487 Ret ↑0.1764 Anti-Halluc ↑0.4216

Contributing Organizations

Amap CV Lab
Nanjing University
Tsinghua University
Peking University