Vchitect/VBench

★ 1,778⑂ 0

[CVPR2024 Highlight] VBench - We Evaluate Video Generation

About Vchitect/VBench

Vchitect/VBench is an open-source project on GitHub, mainly written in Python. [CVPR2024 Highlight] VBench - We Evaluate Video Generation It currently holds 1,778 stars and 0 forks with 0 open issues, and was last pushed on an unknown date (repository created unknown).

Project Overview

AI Homed tracks it on the AI Video Projects board and on the AI AI Video Projects list.

GitHub Repository Details

Repository Vchitect/VBench · default branch - · size 0 KB · watchers 0 · source: GitHub REST API and repository README

README

vbench_logo

HuggingFace VBench Arena (View Generated Videos Here!) VBench-2.0 Arena (View Generated Videos Here!) Project Page Project Page Dataset Download PyPI Video Video Visitors

This repository provides unified implementations for the VBench series of works, supporting comprehensive evaluation of video generative models across a wide spectrum of capabilities and settings.

If your questions are not addressed in this README, please contact Ziqi Huang at ZIQI002 [at] e [dot] ntu [dot] edu [dot] sg.

Table of Contents

:mega: Overview

This repository provides unified implementations for the VBench series of works, supporting comprehensive evaluation of video generative models across a wide spectrum of capabilities and settings.

(1) VBench

TL;DR: Evaluating Video Generation — Benchmark • Evaluation Dimensions • Evaluation Methods • Human Alignment • Insights

VBench Paper (CVPR 2024) VBench: Comprehensive Benchmark Suite for Video Generative Models
Ziqi Huang, Yinan He, Jiashuo Yu, Fan Zhang, Chenyang Si, Yuming Jiang, Yuanhan Zhang, Tianxing Wu, Qingyang Jin, Nattapol Chanpaisit, Yaohui Wang, Xinyuan Chen, Limin Wang, Dahua Lin+, Yu Qiao+, Ziwei Liu+
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024
overall_structure

We propose VBench, a comprehensive benchmark suite for video generative models. We design a comprehensive and hierarchical Evaluation Dimension Suite to decompose "video generation quality" into multiple well-defined dimensions to facilitate fine-grained and objective evaluation. For each dimension and each content category, we carefully design a Prompt Suite as test cases, and sample Generated Videos from a set of video generation models. For each evaluation dimension, we specifically design an Evaluation Method Suite, which uses carefully crafted method or designated pipeline for automatic objective evaluation. We also conduct Human Preference Annotation for the generated videos for each dimension, and show that VBench evaluation results are well aligned with human perceptions. VBench can provide valuable insights from multiple perspectives.

Note: The code and README for the VBench components are located here, relative path: ..

@InProceedings{huang2023vbench,
    title={{VBench}: Comprehensive Benchmark Suite for Video Generative Models},
    author={Huang, Ziqi and He, Yinan and Yu, Jiashuo and Zhang, Fan and Si, Chenyang and Jiang, Yuming and Zhang, Yuanhan and Wu, Tianxing and Jin, Qingyang and Chanpaisit, Nattapol and Wang, Yaohui and Chen, Xinyuan and Wang, Limin and Lin, Dahua and Qiao, Yu and Liu, Ziwei},
    booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
    year={2024}
}

(2) VBench++

TL;DR: Extends VBench with (1) VBench-I2V for image-to-video, (2) VBench-Long for long videos, and (3) VBench-Trustworthiness covering fairness, bias, and safety.

VBench++ (TPAMI 2025) VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models
Ziqi Huang, Fan Zhang, Xiaojie Xu, Yinan He, Jiashuo Yu, Ziyue Dong, Qianli Ma, Nattapol Chanpaisit, Chenyang Si, Yuming Jiang, Yaohui Wang, Xinyuan Chen, Ying-Cong Chen, Limin Wang, Dahua Lin+, Yu Qiao+, Ziwei Liu+
IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2025
overall_structure

VBench++ supports a wide range of video generation tasks, including text-to-video and image-to-video, with an adaptive Image Suite for fair evaluation across different settings. It evaluates not only technical quality but also the trustworthiness of generative models, offering a comprehensive view of model performance. We continually incorporate more video generative models into VBench to inform the community about the evolving landscape of video generation.

Note: The code and README for the VBench++ components are located at:

*These modules belong to VBench++, not VBench, or VBench-2.0. However, to maintain backward compatibility for users who have already installed the repository, we preserve the original relative path names and provide this clarification here. title={{VBench++}: Comprehensive and Versatile Benchmark Suite for Video Generative Models}, author={Huang, Ziqi and Zhang, Fan and Xu, Xiaojie and He, Yinan and Yu, Jiashuo and Dong, Ziyue and Ma, Qianli and Chanpaisit, Nattapol and Si, Chenyang and Jiang, Yuming and Wang, Yaohui and Chen, Xinyuan and Chen, Ying-Cong and Wang, Limin and Lin, Dahua and Qiao, Yu and Liu, Ziwei}, journal={IEEE Transactions on Pattern Analysis and Machine Intelligence}, year={2025}, doi={10.1109/TPAMI.2025.3633890} }

(3) VBench-2.0

TL;DR: Extends VBench to evaluate intrinsic faithfulness — a key challenge for next-generation video generation models.

VBench-2.0 Report (arXiv) VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness
Dian Zheng, Ziqi Huang, Hongbo Liu, Kai Zou, Yinan He, Fan Zhang, Yuanhan Zhang, Jingwen He, Wei-Shi Zheng+, Yu Qiao+, Ziwei Liu+
overall_structure Overview of VBench-2.0. (a) Scope of VBench-2.0. Video generative models have progressed from achieving superficial faithfulness in fundamental technical aspects such as pixel fidelity and basic prompt adherence, to addressing more complex challenges associated with intrinsic faithfulness, including commonsense reasoning, physics-based realism, human motion, and creative composition. While VBench primarily assessed early-stage technical quality, VBench-2.0 expands the benchmarking framework to evaluate these advanced capabilities, ensuring a more comprehensive assessment of next-generation models. (b) Evaluation Dimension of VBench-2.0. VBench-2.0 introduces a structured evaluation suite comprising five broad categories and 18 fine-grained capability dimensions.

Note: The code and README for the VBench-2.0 components are located at link, relative path: VBench-2.0.

@article{zheng2025vbench2,
    title={{VBench-2.0}: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness},
    author={Zheng, Dian and Huang, Ziqi and Liu, Hongbo and Zou, Kai and He, Yinan and Zhang, Fan and Zhang, Yuanhan and He, Jingwen and Zheng, Wei-Shi and Qiao, Yu and Liu, Ziwei},
    journal={arXiv preprint arXiv:2503.21755},
    year={2025}
}

:fire: Updates

is documented HERE.

:mortar_board: Evaluation Results

See our leaderboard for the most updated ranking and numerical results (with models like Gen-3, Kling, Pika). HuggingFace

We visualize the evaluation results of the 12 most recent top-performing long video generation models across 16 VBench dimensions.

Additionally, we present radar charts separately for the evaluation results of open-source and closed-source models. The results are normalized per dimension for clearer comparisons.

:trophy: Leaderboard

See numeric values at our Leaderboard :1st_place_medal::2nd_place_medal::3rd_place_medal:

:film_projector: Model Info

See model info for video generation models we used for evaluation.

:hammer: Installation

Install with pip

pip install torch torchvision --index-url https://download.pytorch.org/whl/cu118 # or any other PyTorch version with CUDA<=12.1
pip install vbench

To evaluate some video generation ability aspects, you need to install detectron2 via:

   pip install detectron2@git+https://github.com/facebookresearch/detectron2.git
   
If there is an error during detectron2 installation, see here. Detectron2 is working only with CUDA 12.1 or 11.X.

Download VBench_full_info.json to your running directory to read the benchmark prompt suites.

Install with git clone

git clone https://github.com/Vchitect/VBench.git pip install torch torchvision --index-url https://download.pytorch.org/whl/cu118 # or other version with CUDA<=12.1 pip install VBench If there is an error during detectron2 installation, see here.

Usage

Use VBench to evaluate videos, and video generative models.

[New] Evaluate Your Own Videos

We support evaluating any video. Simply provide the path to the video file, or the path to the folder that contains your videos. There is no requirement on the videos' names.

To evaluate videos with customized input prompt, run our script with --mode=custom_input:

python evaluate.py \
    --dimension $DIMENSION \
    --videos_path /path/to/folder_or_video/ \
    --mode=custom_input
alternatively you can use our command:
vbench evaluate \
    --dimension $DIMENSION \
    --videos_path /path/to/folder_or_video/ \
    --mode=custom_input

To evaluate using multiple gpus, we can use the following commands:

torchrun --nproc_per_node=${GPUS} --standalone evaluate.py ...args...
or
vbench evaluate --ngpus=${GPUS} ...args...

Evaluation on the Standard Prompt Suite of VBench

Command Line
vbench evaluate --videos_path $VIDEO_PATH --dimension $DIMENSION
For example:
vbench evaluate --videos_path "sampled_videos/lavie/human_action" --dimension "human_action"
Python
from vbench import VBench
my_VBench = VBench(device, , )
my_VBench.evaluate(
    videos_path = <video_path>,
    name = ,
    dimension_list = [, , ...],
)
For example:
from vbench import VBench
my_VBench = VBench(device, "vbench/VBench_full_info.json", "evaluation_results")
my_VBench.evaluate(
    videos_path = "sampled_videos/lavie/human_action",
    name = "lavie_human_action",
    dimension_list = ["human_action"],
)

Evaluation of Different Content Categories

command line
vbench evaluate \
    --videos_path $VIDEO_PATH \
    --dimension $DIMENSION \
    --mode=vbench_category \
    --category=$CATEGORY
or ``` python evaluate.py \ --dimension $DIMENSION \ --videos_path /path/to/folder_

GitHub Stars & Activity

1,778Stars
0Forks
0Open issues
PythonLanguage

GitHub Popularity

GitHub stars1,778
Forks0
Open issues0
Primary languagePython
License-
Stars gained today0
Created-
Last pushed-

Trending History

Trending statusnot on today's boards

Related AI Projects

1

calesthio / OpenMontage

Python★ 59,376⑂ 0
2

ATH-MaaS / Pixelle-Video

Python★ 28,141⑂ 0
3

KlingAIResearch / LivePortrait

Python★ 19,050⑂ 0
4

Wan-Video / Wan2.2

Python★ 17,520⑂ 0
5

Zulko / moviepy

Python★ 14,897⑂ 0
6

zai-org / CogVideo

Python★ 13,018⑂ 0
7

Tencent-Hunyuan / HunyuanVideo

Python★ 12,527⑂ 0
8

HKUDS / ViMax

Python★ 12,393⑂ 0

More AI Rankings