AetherLabsAI/Video2World
Benchmarking Coding Agents for Interactive World Modeling from Embodied Videos - Rethinking Embodied Real2Sim from a Software Engineering Perspective
About AetherLabsAI/Video2World
AetherLabsAI/Video2World is an open-source project on GitHub, mainly written in Python. Benchmarking Coding Agents for Interactive World Modeling from Embodied Videos - Rethinking Embodied Real2Sim from a Software Engineering Perspective It currently holds 83 stars and 0 forks with 0 open issues, and was last pushed on an unknown date (repository created unknown).
Project Overview
AI Homed tracks it on the Today's Trending board, currently at rank #86 with 0 new stars today.
GitHub Repository Details
README
Video2World: Benchmarking Coding Agents for Interactive World Modeling from Embodied Videos
Jinzhou Tang1,2*, Zijun Zhang1*, Jing Yang1, Yuchen Yan1, Kun Zhou1†, Lingjun Mao1, Ruobing Han1, Jinglin Cao1, Wenpeng Xu1, Lukun He1, Minghao Fu1,2, Fan Feng1, Biwei Huang1,2
1Aether AI 2University of California, San Diego
*Equal contribution †Corresponding author and project lead
Building interactive simulators from real-world observations is a promising way to scale embodied data, but current pipelines still rely heavily on manual environment construction and calibration. Video2World asks whether frontier foundation models and coding agents can automate this process end to end. We formulate autonomous video-to-simulation as a software engineering task: an agent observes an embodied video, constructs the corresponding simulated environment and robot behavior, and iteratively refines the result through execution feedback.
Video2World comprises 222 reconstruction instances in 39 task families, built from 189 video clips of robot executions, egocentric human recordings and third-person human demonstrations. Reconstructed worlds are measured along geometric fidelity, dynamic fidelity and functional correctness, capturing spatial perception, physical reasoning and executable interaction.
Overview
Each instance pairs a source video with a target configuration (simulator, robot embodiment, control interface). The coding agent works in an isolated sandbox with the video, simulator APIs, documentation, assets and general-purpose coding tools; it receives no scene model, object trajectory, robot state or task specification. The submitted scene and robot behavior are executed by the evaluator from a fresh initialization, and every change of object state must arise from simulator dynamics. The rollout is compared with hidden, object-centric annotations:
- Geometry: interaction-weighted scene Chamfer distance, object shape Chamfer distance and size error;
- Dynamics: translation APE, rotation APE and translation RPE of the object motion produced by the submitted robot
- Functionality: task success and stage-wise task progress, judged on object-centric task requirements.
Installation
Scoring needs only the base package. Evaluation uses three simulator environments, pinned to the versions the reference results were produced with:
git clone https://github.com/AetherLabsAI/Video2World.git
cd Video2World
pip install -e . # scoring
ManiSkill / SAPIEN: FurnitureBench, DROID, hand, Push-T
pip install -e .[sim]
python -m mani_skill.utils.download_asset xarm6 # Robotiq gripper meshes
MuJoCo: rope routing, toy packing, cloth (a separate environment)
python -m venv .venv-twin && .venv-twin/bin/pip install -e .[twin]
v2w setup --twin-python .venv-twin/bin/python
Isaac Sim 5.1 + Isaac Lab: RoboDojo and in-house (with a RoboDojo checkout at commit 25691aa)
v2w setup --isaac-python /path/to/isaac/python --robodojo-source /path/to/RoboDojo
v2w setup # report what is configured
v2w setup stores these locations in v2w.local.json; --graphics-libs DIR adds a library directory for headless
Isaac rendering, and Isaac Sim needs GPUs with a working Vulkan device. ManiSkill pulls in the GUI build of OpenCV,
which needs libGL; on a headless server without it, install libgl1 or replace it with the headless build
(pip uninstall -y opencv-python && pip install --force-reinstall opencv-python-headless).
Data
Download the dataset into ./data (or elsewhere, then
v2w setup --data PATH):
huggingface-cli download AetherLabs-AI/Video2World --repo-type dataset --local-dir data
data/
manifest.json 222 instances with their metric contracts
samples// video.mp4 (agent input); gt_pkg/, hidden/, meta.json (evaluator only)
har//.json human-assisted reference results
sources/ assets/ source annotations, robot models and simulator assets
Usage
Score the human-assisted reference shipped with the data:
v2w har
Run a coding agent on the benchmark, evaluate its submissions and compute V2WScore. Agents are driven through their command-line interfaces (Claude Code, Codex, OpenCode), so any model those tools can reach can be evaluated, including self-hosted ones behind an OpenAI-compatible API:
# Claude Code
v2w run --name opus --agent claude --model claude-opus-5-5 --gpus 0,1 --workers 2
Codex
v2w run --name astra --agent codex --model gpt-6-astra --reasoning-effort high
any OpenAI-compatible endpoint (vLLM, SGLang, OpenRouter, DeepSeek, ...), driven by OpenCode
vllm serve Qwen/Qwen3-Coder --port 8000
v2w run --name qwen --base-url http://localhost:8000/v1 --model Qwen/Qwen3-Coder
OPENROUTER_API_KEY=... v2w run --name kimi --base-url https://openrouter.ai/api/v1 \
--api-key-env OPENROUTER_API_KEY --model moonshotai/kimi-k3
your own agent (see docs/custom_agent.md)
v2w run --name mine --agent-cmd "my_agent --brief {brief} --video {video} --out {out}"
packages produced elsewhere, laid out as //protocol.json
v2w eval --name submitted --packages /path/to/packages
v2w score --name opus # runs/opus/report/{summary.json,samples.json,samples.csv}
v2w list # task families and instances
With --base-url, OpenCode is used by default; --agent codex uses the Responses API of the endpoint instead. The
agent inspects the video through rendered frames, so the model should accept image input. To evaluate an agent of your
own, follow the tutorial in docs/custom_agent.md.
--family and --sample restrict a run, and every step is resumable. A submission is a
video2sim protocol package: protocol.json (scene, robot, camera frame, task), the action
stream, meshes and a short report. Agents may use only the video and the task brief of their profile
(v2w/agents/profiles/); the bundled driver stages exactly these into an isolated workspace and audits the agent log
for accesses to the hidden annotations. The full evaluation protocol is in docs/benchmark.md.
Repository structure
v2w/
cli.py runner.py scoring.py benchmark.py paths.py
agents/ coding-agent driver, task briefs and public task handouts
metrics/ scene / object geometry, trajectories, task success and progress
tracks/ evaluators: furniture (FurnitureBench, DROID, Push-T), hand, robodojo, inhouse, twins, cloth
config/ V2WScore configuration
video2sim/ package protocol, validator, simulation bridges and the agents' self-check tools
docs/ benchmark protocol, custom-agent tutorial
Citation
@article{tang2026video2world,
title = {Video2World: Benchmarking Coding Agents for Interactive World Modeling from Embodied Videos},
author = {Tang, Jinzhou and Zhang, Zijun and Yang, Jing and Yan, Yuchen and Zhou, Kun and Mao, Lingjun and
Han, Ruobing and Cao, Jinglin and Xu, Wenpeng and He, Lukun and Fu, Minghao and Feng, Fan and
Huang, Biwei},
journal = {arXiv preprint arXiv:2610.04432},
year = {2026}
}
Acknowledgements
Video2World builds on the following datasets, simulators, robot models and tools; we thank their authors for making them available.
- Video sources and annotations: FurnitureBench,
- Reconstructed twins: PhysTwin and the
- Simulators: ManiSkill 3 and SAPIEN,
- Robot models: Franka Panda and Robotiq 2F-85 from the ManiSkill assets,
- In-house household episodes: the AgiBot G1 with OmniPicker grippers and the household scenes of
- Coding agents: Claude Code, Codex
License
Released under the Apache License 2.0.