PKU-YuanGroup/Helios

★ 2,147⑂ 0

Helios: Real Real-Time Long Video Generation Model

About PKU-YuanGroup/Helios

PKU-YuanGroup/Helios is an open-source project on GitHub, mainly written in Python. Helios: Real Real-Time Long Video Generation Model It currently holds 2,147 stars and 0 forks with 0 open issues, and was last pushed on an unknown date (repository created unknown).

Project Overview

AI Homed tracks it on the AI Video Projects board and on the AI AI Video Projects list.

GitHub Repository Details

Repository PKU-YuanGroup/Helios · default branch - · size 0 KB · watchers 0 · source: GitHub REST API and repository README

README

Helios: Real Real-Time Long Video Generation Model

⭐ 14B Real-Time Long Video Generation Model can be Cheaper, Faster but Keep Stronger than 1.3B ones ⭐

arXiv hf_paper Project Page hf_space HuggingFace ModelScope GitHub GitCode

Ascend Diffusers SGLang Diffusion vLLM-Omni

This repository is the official implementation of Helios, which is a breakthrough video generation model that achieves minute-scale, high-quality video synthesis at 19.5 FPS on a single H100 GPU (about 10 FPS on a single Ascend NPU) —without relying on conventional long video anti-drifting strategies or standard video acceleration techniques.


✨ Highlights

1. Without commonly used anti-drifting strategies (e.g., self-forcing, error-banks, keyframe sampling, or inverted sampling), Helios generates minute-scale videos with high quality and strong coherence.

2. Without standard acceleration techniques (e.g., KV-cache, causal masking, sparse/linear attention, TinyVAE, progressive noise schedules, hidden-state caching, or quantization), Helios achieves 19.5 FPS in end-to-end inference on a single H100 GPU.

3. We introduce optimizations that improve both training and inference throughput while reducing memory consumption, enabling image-diffusion-scale batch sizes during training while fitting up to four 14B models within 80 GB of GPU memory.

🎬 Video Demos

Demo Video of Helios or you can click here to get the video. Some best prompts are here.

📣 Latest News!!

🔥 Friendly Links

If your work has improved Helios and you would like more people to see it, please inform us.

⚙️ Requirements and Installation

Video Tutorial

If you prefer a step-by-step walkthrough, check out this community-made YouTube Tutorial. It covers local installation, 4K video generation, and how to run Helios on a consumer-grade PC, along with other practical usage tips.

Prepare Environment

# 0. Clone the repo
git clone --depth=1 https://github.com/PKU-YuanGroup/Helios.git
cd Helios

1. Create conda environment

conda create -n helios python=3.11.2 conda activate helios

2. Install PyTorch (adjust for your CUDA version)

CUDA 12.6

pip install torch==2.10.0 torchvision==0.25.0 torchaudio==2.10.0 --index-url https://download.pytorch.org/whl/cu126

CUDA 12.8

pip install torch==2.10.0 torchvision==0.25.0 torchaudio==2.10.0 --index-url https://download.pytorch.org/whl/cu128

CUDA 13.0

pip install torch==2.10.0 torchvision==0.25.0 torchaudio==2.10.0 --index-url https://download.pytorch.org/whl/cu130

3. Install dependencies

bash install.sh

Model Download

| Models | Download Link | Supports | Notes | |------------------|----------------------------------------------------------------------------------------------------------------------------------------------------------|-----------------------------------------------|---------------------------------------------------------------------------------------------| | Helios-Base | 🤗 Huggingface 🤖 ModelScope | T2V ✅ I2V ✅ V2V ✅ Interactive ✅ | Best Quality, with v-prediction, standard CFG and custom HeliosScheduler. | | Helios-Mid | 🤗 Huggingface 🤖 ModelScope | T2V ✅ I2V ✅ V2V ✅ Interactive ✅ | Intermediate Ckpt, with v-prediction, CFG-Zero* and custom HeliosScheduler. | | Helios-Distilled | 🤗 Huggingface 🤖 ModelScope | T2V ✅ I2V ✅ V2V ✅ Interactive ✅ | Best Efficiency, with x0-prediction and custom HeliosDMDScheduler. |

💡Note:
* All three models share the same architecture, but Helios-Mid and Helios-Distilled use a more aggressive multi-scale sampling pipeline to achieve better efficiency.
* Helios-Mid is an intermediate checkpoint generated in the process of distilling Helios-Base into Helios-Distilled, and may not meet expected quality.
* For Image-to-Video or Video-to-Video, since training is based on Text-to-Video, these two functions may be slightly inferior to Text-to-Video. You may enable is_skip_first_chunk if you find the first few chunks are static or imporve the value of image_noise_sigma_min, image_noise_sigma_max, video_noise_sigma_min, and video_noise_sigma_max.

Download models using huggingface-cli: ``` sh pip install "huggingface_hub[cli]" huggingface-cli download BestWishYSH/Helios-Base --local-dir BestWishYSH/Helios-Base huggingface-cli download BestWishYSH/Helios-Mid --local-dir BestWishYSH/Helios-Mid huggingface-cli download BestWishYSH/Helios-Distilled --local-dir BestWishYSH/Helios-Distilled


Download models using modelscope-cli:
sh pip install modelscope modelscope download BestWishYSH/Helios-Base --local_dir BestWishYSH/Helios-Base modelscope download BestWishYSH/Helios-Mid --local_dir BestWishYSH/Helios-Mid modelscope download BestWishYSH/Helios-Distilled --local_dir BestWishYSH/Helios-Distilled

🚀 Inference

Helios uses an autoregressive approach that generates 33 frames per chunk. For optimal performance, num_frames should be set to a multiple of 33. If a non-multiple value is provided, it will be automatically rounded up to the nearest multiple of 33.

Example frame counts for different video lengths:

| num_frames | Adjusted Frames | 24 FPS | 16 FPS | |------------|-----------------|--------|--------| | 1449 | 1452 (33×44) | ~60s (1min) | ~90s (1min 30s) | | 720 | 726 (33×22) | ~30s | ~45s | | 240 | 264 (33×8) | ~11s | ~16s | | 129 | 132 (33×4) | ~5.5s | ~8s | | 81 | 99 (33×3) | ~4s | ~6s |

Run the model

We provide inference scripts for all models covering text-to-video, image-to-video, and video-to-video in this directory.

bash cd scripts/inference

For Helios-Base

bash helios-base_t2v.sh bash helios-base_i2v.sh bash helios-base_v2v.sh

For Helios-Mid

bash helios-mid_t2v.sh bash helios-mid_i2v.sh bash helios-mid_v2v.sh

For Helios-Distilled

bash helios-distilled_t2v.sh bash helios-distilled_i2v.sh bash helios-distilled_v2v.sh

For Interactive

⚠️ This feature is still under development — results may not always meet expectations

cd scripts/inference/experiment_interactive

Sanity Check

Before trying your own inputs, we highly recommend going through the sanity check to find out if any hardware or software went wrong.

| Task | Helios-Base | Helios-Mid | Helios-Distilled | | ------- | -------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------- | | T2V | | | | | V2V | | | |

✨ Group Offloading to Save VRAM

Helios supports group offloading to significantly reduce VRAM consumption, allowing you to run on GPU with limited memory footprint. For more details on the underlying mechanics, please refer to the documentation.

The Helios model below requires ~6GB of VRAM.

Click to expand the code

bash CUDA_VISIBLE_DEVICES=0 python infer_helios.py \ --base_model_path "BestWishYsh/Helios-Distilled" \ --transformer_path "BestWishYsh/Helios-Distilled" \ --sample_type "t2v" \ --prompt "A vibrant tropical fish swimming gracefully among colorful coral reefs in a clear, turquoise ocean. The fish has bright blue and yellow scales with a small, distinctive orange spot on its side, its fins moving fluidly. The coral reefs are alive with a variety of marine life, including small schools of colorful fish and sea turtles gliding by. The water is crystal clear, allowing for a view of the sandy ocean floor below. The reef itself is adorned with a mix of hard and soft corals in shades of red, orange, and green. The photo captures the fish from a slightly elevated angle, emphasizing its lively movements and the vivid colors of its surroundings. A close-up shot with dynamic movement." \ --num_frames 240 \ --guidance_scale 1.0 \ --is_enable_stage2 \ --pyramid_num_inference_steps_list 2 2 2 \ --is_amplify_first_chunk \ --output_folder "./output_helios/helios-distilled" \ --enable_low_vram_mode \ --group_offloading_type "leaf_level"
  

✨ Context Parallelism on Multiple GPUs

Helios supports various Context Parallelism mechanisms, including Ulysses Attention, Ring Attention, Unified Attention, and Ulysses Anything Attention. For more details, please refer to the documentation.

For example, let's take Helios-Base with 4 GPUs.

Click to expand the code

bash CUDA_VISIBLE_DEVICES=0,1,2,3 torchrun --nproc_per_node 4 infer_helios.py \ --enable_parallelism \ # remember to enable this config --cp_backend "ulysses" \ # ["ring", "ulysses", "unified", "ulysses_anything"] --base_model_path "BestWishYsh/Helios-Base" \ --transformer_path "BestWishYsh/Helios-Base" \ --sample_type "t2v" \ --num_frames 99 \ --fps 24 \ --prompt "A vibrant tropical fish swimming gracefully among colorful coral reefs in a clear, turquoise ocean. The fish has bright blue and yellow scales with a small, distinctive orange spot on its side, its fins moving fluidly. The coral reefs are alive with a variety of marine life, including small schools of colorful fish and sea turtles gliding by. The water is crystal clear, allowing for a view of the sandy ocean floor below. The reef itself is adorned with a mix of hard and soft corals in shades of red, orange, and green. The photo captures the fish from a slightly elevated angle, emphasizing its lively movements and the vivid colors of its surroundings. A close-up shot with dynamic movement." \ --guidance_scale 5.0 \ --output_folder "./output_helios/helios-base"
  

✨ Diffusers Pipeline

Install diffusers from source:

bash pip install git+https://github.com/huggingface/diffusers.git

For example, let's take Helios-Distilled (Standard Pipeline).

Click to expand the code

bash import torch from diffusers import AutoModel, HeliosPyramidPipeline from diffusers.utils import export_to_video, load_video, load_image

vae = AutoModel.from_pretrained("BestWishYsh/Helios-Distilled", subfolder="vae", torch_dtype=torch.float32)

pipeline = HeliosPyramidPipeline.from_pretrained( "BestWishYsh/Helios-Distilled", vae=vae, torch_dtype=torch.bfloat16 ) pipeline.to("cuda")

negative_prompt = """ Bright tones, overexposed, static, blurred details, subtitles, style, works, paintings, images, static, overall gray, worst quality, low quality, JPEG compression residue, ugly, incomplete, extra fingers, poorly drawn hands, poorly drawn faces, deformed, disfigured, misshapen limbs, fused fingers, still picture, messy background, three legs, many people in the background, walking backwards """

# --- T2V --- prompt = """ A vibrant tropical fish swimming gracefully among colorful coral reefs in a clear, turquoise ocean. The fish has bright blue and yellow scales with a small, distinctive orange spot on its side, its fins moving fluidly. The coral reefs are alive with a variety of marine life, including small schools of colorful fish and sea turtles gliding by. The water is crystal clear, allowing for a view of the sandy ocean floor below. The reef itself is adorned with a mix of hard and soft corals in shades of red, orange, and green. The photo captures the fish from a slightly elevated angle, emphasizing its lively movements and the vivid colors of its surroundings. A close-up shot with dynamic movement. """

output = pipeline( prompt=prompt, negative_prompt=negative_prompt, num_frames=240, pyramid_num_inference_steps_list=[2, 2, 2], guidance_scale=1.0, is_amplify_first_chunk=True, generator=torch.Generator("cuda").manual_seed(42), ).frames[0] export_to_video(output, "helios_distilled_t2v_output.mp4", fps=24)

# --- I2V --- i2v_prompt = """ A towering emerald wave surges forward, its crest curling with raw power and energy. Sunlight glints off the translucent water, illuminating the intricate textures and deep green hues within the wave’s body. A thick spray erupts from the breaking crest, casting a misty veil that dances above the churning surface. As the perspective widens, the immense scale of the wave becomes apparent, revealing the restless expanse of the ocean stretching beyond. The scene captures the ocean’s untamed beauty and relentless force, with every droplet and ripple shimmering in the light. The dynamic motion and vivid colors evoke both awe and respect for nature’s might. """ image_path = "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/diffusers/helios/wave.jpg"

output = pipeline( prompt=i2v_prompt, negative_prompt=negative_prompt, image=load_image(image_path).resize((640, 384)), num_frames=240, pyramid_num_inference_steps_list=[2, 2, 2], guidance_scale=1.0, is_amplify_first_chunk=True, generator=torch.Generator("cuda").manual_seed(42), ).frames[0] export_to_video(output, "helios_distilled_i2v_output.mp4", fps=24)

# --- V2V --- v2v_prompt = """ A bright yellow Lamborghini Huracn Tecnica speeds along a curving mountain road, surrounded by lush green trees under a partly cloudy sky. The car's sleek design and vibrant color stand out against the natural backdrop, emphasizing its dynamic movement. The road curves gently, with a guardrail visible on one side, adding depth to the scene. The motion blur captures the sense of speed and energy, creating a thrilling and exhilarating atmosphere. A front-facing shot from a slightly elevated angle, highlighting the car's aggressive stance and the surrounding greenery. """ video_path = "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/diffusers/helios/car.mp4"

output = pipeline( prompt=v2v_prompt, negative_prompt=negative_prompt, video=load_video(video_path), num_frames=240, pyramid_num_inference_steps_list=[2, 2, 2], guidance_scale=1.0, is_amplify_first_chunk=True, generator=torch.Generator("cuda").manual_seed(42), ).frames[0] export_to_video(output, "helios_distilled_v2v_output.mp4", fps=24)


For example, let's take Helios-Distilled (Modular Pipeline).

Click to expand the code

bash import torch from diffusers import ModularPipeline, ClassifierFreeGuidance from diffusers.utils import export_to_video, load_image, load_video

mod_pipe = ModularPipeline.from_pretrained("BestWishYsh/Helios-Distilled") mod_pipe.load_components(torch_dtype=torch.bfloat16) mod_pipe.to("cuda")

# we need to upload guider to the model repo, so each checkpoint will be able to config their guidance d

GitHub Stars & Activity

2,147Stars
0Forks
0Open issues
PythonLanguage

GitHub Popularity

GitHub stars2,147
Forks0
Open issues0
Primary languagePython
License-
Stars gained today0
Created-
Last pushed-

Trending History

Trending statusnot on today's boards

Related AI Projects

1

calesthio / OpenMontage

Python★ 59,376⑂ 0
2

ATH-MaaS / Pixelle-Video

Python★ 28,141⑂ 0
3

KlingAIResearch / LivePortrait

Python★ 19,050⑂ 0
4

Wan-Video / Wan2.2

Python★ 17,520⑂ 0
5

Zulko / moviepy

Python★ 14,897⑂ 0
6

zai-org / CogVideo

Python★ 13,018⑂ 0
7

Tencent-Hunyuan / HunyuanVideo

Python★ 12,527⑂ 0
8

HKUDS / ViMax

Python★ 12,393⑂ 0

More AI Rankings