World2Motion

Turning Video World Models into
3D Human Motion Generators

Fangyuan Tu1,2,+·Xiangyue Zhang2,3·Yiyi Cai2,3·Yichen Peng4·Kunhang Li2,3
·Bo Zheng2·Zhixiang Wang2·Kaipeng Zhang2·Erwin Wu4·Haoran Xie1·Haiyang Liu2,3

1Japan Advanced Institute of Science and Technology2Alaya Lab3The University of Tokyo4Institute of Science Tokyo

+ Work done during internship at Alaya Lab.

01 / OVERVIEW

From a video world model.
Into the 3D Human Motion.

World2Motion turns the visual knowledge of a video world model into scene-aware 3D human motion. Given a single image and a text prompt, it generates the video and the corresponding motion together.

02 / METHOD

World2Motion architecture: initial pose estimation, video and motion branches, shared attention, and a shift-decoupled denoising schedule.
01

Joint video and
3D motion generation

We adapt Cosmos 3 to generate scene-aware 3D human motion and corresponding video in a single stage, conditioned on one image and a text prompt.

02

Hybrid video–motion
training data

We combine synthetic video–motion pairs with real videos paired with estimated 3D motion, bringing interaction supervision and diverse human movements into one dataset.

03

A shift-decoupled
noise schedule

Video and motion use different noise shifts with shared denoising progress. The same noise relationship is maintained during training and inference to improve temporal stability.

03 / COMPARISON

Quantitative comparison

Table 2. Quantitative results of motion generation methods. Best values are bold.
MethodFID ↓R@1 ↑MM-Dist ↓Contact-F1 ↑ISR ↑Jerk ↓(10 m/s³)FSR ↓(%)PoseError ↓(mm)Time ↓(s)
Text-to-motion
HY-Motion0.460.101.210.090.513.431.25—2.83
Kimodo0.660.051.300.030.2335.146.07—1.93
Go-to-Zero0.510.041.350.070.3617.503.42—2.57
Human–object interaction
HOIFHLI0.820.021.370.090.1210.626.28—13.94
LIGHT0.620.031.350.170.461.3711.60—25.88
Two-stage video generation and motion recovery
MiniMax-H3 + CameraHMR0.690.101.320.240.9137.226.41—523.67
Joint video–motion generation
CoMoVi0.710.121.250.090.6916.1024.4581.28270.45
World2Motion0.270.241.130.220.918.352.8067.50156.59

Best values are bold. PoseError measures pelvis-relative agreement with CameraHMR estimates from generated video. ISR measures interaction completion; it does not measure exact physical contact.

3.34×

Faster inference

vs. MiniMax-H3 + CameraHMR
156.59 s vs. 523.67 s

91%

Interaction success

Matches the two-stage baseline
on the evaluated benchmark

67.5mm

Video–motion PoseError

17.0% lower than CoMoVi
under the same evaluation

Qualitative comparison

HOIFHLI and LIGHT include their predicted object motion. In the other motion renderings, dashed objects show fixed scene context. Video-generating methods include the corresponding RGB output.

04 / RESULTS