deepseek-ai/DeepGEMM

▲ 363 stars today★ 8,646⑂ 1,368

DeepGEMM: clean and efficient BLAS kernel library on GPU

About deepseek-ai/DeepGEMM

deepseek-ai/DeepGEMM is an open-source project on GitHub, mainly written in Cuda. DeepGEMM: clean and efficient BLAS kernel library on GPU It currently holds 8,646 stars and 1,368 forks with 0 open issues, and was last pushed on an unknown date (repository created unknown).

Project Overview

AI Homed tracks it on the Today's Trending board, currently at rank #15 with 363 new stars today.

GitHub Repository Details

Repository deepseek-ai/DeepGEMM · default branch - · size 0 KB · watchers 0 · source: GitHub REST API and repository README

README

DeepGEMM

DeepGEMM is a unified, high-performance tensor core kernel library that brings together the key computation primitives of modern large language models — GEMMs (FP8, FP4, BF16), fused MoE with overlapped communication (Mega MoE), MQA scoring for the lightning indexer, HyperConnection (HC), and more — into a single, cohesive CUDA codebase. All kernels are compiled at runtime through DeepJIT, requiring no CUDA compilation during installation.

DeepGEMM leverages some concepts from CUTLASS and CuTe, but avoids heavy reliance on their templates or algebras. The library is designed for simplicity, with only a limited number of core kernel functions, making it a clean and accessible resource for learning NVIDIA GPU kernel optimization techniques.

Despite its lightweight design, DeepGEMM's performance matches or exceeds expert-tuned libraries across various matrix shapes.

News

Quick start

Requirements

Development

# Submodule must be cloned
git clone --recursive git@github.com:deepseek-ai/DeepGEMM.git
cd DeepGEMM

Link some essential includes and build the C++ extension

cat develop.sh ./develop.sh

Installation

cat install.sh
./install.sh

Then, import deep_gemm in your Python project, and enjoy!

Interfaces

Notices

This library provides optimized GEMM kernels for NVIDIA GPUs with a naming convention: D = C + A @ B. The input shape layout is NT (non-transposed A, transposed B). While the SM90 implementation supports only the NT memory layout (row-major, col-major), the SM100 implementation supports all memory layouts (NT, TN, NN, TT). For example, fp8_gemm_nt will do a D = C + A @ B.T

For both architectures, the LHS scaling factor is required to have a TMA-aligned and transposed layout. And the data format for the scaling factor of SM90 and SM100 is different:

Please note that operations like input transposition or FP8 casting must be handled separately by the user, please implement or fuse them into prior kernels independently. While the library provides some simple PyTorch utility functions, these may result in slower performance, but our primary focus is on optimizing the GEMM kernels themselves.

Normal dense GEMMs (non-grouped)

To perform a basic non-grouped FP8 GEMM, call the fp8_gemm_{nt, nn, tn, tt} function. For more details, please refer to the function documentation.

Grouped GEMMs (contiguous layout)

Unlike traditional grouped GEMMs in CUTLASS, DeepGEMM groups only the M-axis, while N and K must remain fixed. This design is tailored for scenarios where experts in an MoE model share the same shape. For training forward passes or inference prefilling, where each expert may process a varying number of tokens, we concatenate these tokens into a single tensor, referred to as the "contiguous" layout. Note that each expert segment must be aligned to the GEMM M block size (get_mk_alignment_for_contiguous_layout()). For more information, please refer to the m_grouped_fp8_gemm_{nt, nn}_contiguous function documentation.

We also provide a K-axis-grouped API for MoE weight backward (with M and N must remain fixed), please refer to k_grouped_fp8_gemm_tn_contiguous for more information.

Grouped GEMMs (masked layout)

During the inference decoding phase, when CUDA graph is enabled and the CPU is unaware of the number of tokens each expert receives, we support masked grouped GEMMs. By providing a mask tensor, the kernel computes only the valid portions.

Use m_grouped_fp8_gemm_nt_masked for this purpose and consult the relevant documentation. An example usage is to use the output of low-latency kernels from DeepEP as input.

V3.2 MQA kernels for the indexer

The kernel family has two versions, non-paged (for prefilling) and paged (for decoding). Take the non-paged version fp8_fp4_mqa_logits as an example. Its main inputs are:

The output is compressed to [seq_len, max_seqlen_k]; row i stores its valid KV span starting at column zero. For each token i in q, it will iterate all tokens j from [cu_seq_len_k_start[i], cu_seq_len_k_end[i]), and calculate the corresponding compressed logit as:

kv_j = kv[0][j, :] * kv[1][j].unsqueeze(1)  # [head_dim]
out_ij = q[i, :, :] @ kv_j  # [num_heads]
out_ij = out_ij.relu() * weights[i, :]  # [num_heads]
out_ij = out_ij.sum()  # Scalar

For more details and the paged version fp8_fp4_paged_mqa_logits, please refer to tests/test_attention.py.

Mega MoE

Mega MoE fuses and overlaps EP dispatch, linear 1 and linear 2 (FP8xFP4 or FP8xFP8), SwiGLU, and EP combine into a single mega-kernel, overlapping NVLink communication and tensor core computation. It requires multi-process launch with symmetric memory. Usage:

# Allocate symmetric memory buffer

NOTES: requires PyTorch >= 2.9

buffer = deep_gemm.get_symm_buffer_for_mega_moe( group, num_experts, num_max_tokens_per_rank, num_topk, hidden, intermediate_hidden, mma_type='fp8xfp4', # Use 'fp8xfp8' for FP8 routed-expert weights )

Transform weights (FP4 or FP8 with UE8M0 SF) into the required layout

transformed_l1, transformed_l2 = deep_gemm.transform_weights_for_mega_moe(l1_weights, l2_weights)

(Optional) Localize weights into locality domains

transformed_l1 = (deep_gemm.localize(transformed_l1[0]), transformed_l1[1]) transformed_l2 = (deep_gemm.localize(transformed_l2[0]), transformed_l2[1]) deep_gemm.destroy_localizer()

Copy inputs into the buffer before each call

You may fuse these into previous kernels

buffer.x[:num_tokens].copy_(x_fp8) buffer.x_sf[:num_tokens].copy_(x_sf) buffer.topk_idx[:num_tokens].copy_(topk_idx) buffer.topk_weights[:num_tokens].copy_(topk_weights)

Run the fused mega MoE kernel

y = torch.empty((num_tokens, hidden), dtype=torch.bfloat16, device='cuda') deep_gemm.fp8_fp4_mega_moe(y, transformed_l1, transformed_l2, buffer)

For the full example with multi-process setup and benchmarking, please refer to tests/test_mega_moe.py.

Utilities

The library provides some utility functions besides the above kernels:

The library also provides some environment variables, which may be useful:

Each DG_JIT_* variable falls back to the corresponding global DJ_JIT_* variable when unset.

For additional examples and details, please refer to the test code or review the corresponding Python documentation.

Acknowledgement

DeepGEMM is inspired by the CUTLASS project. Thanks and respect to the developers!

License

This code repository is released under the MIT License.

Citation

@misc{deepgemm2025,
      title={DeepGEMM: clean and efficient BLAS kernel library on GPU}, 
      author={Chenggang Zhao and Zhean Xu and Liang Zhao and Jiashi Li and Chenhao Xu and Anyi Xu and Shengyu Liu and Kexing Zhou and Kuai Yu},
      year={2025},
      publisher = {GitHub},
      howpublished = {\url{https://github.com/deepseek-ai/DeepGEMM}},
}

GitHub Stars & Activity

8,646Stars
1,368Forks
0Open issues
CudaLanguage

GitHub Popularity

GitHub stars8,646
Forks1,368
Open issues0
Primary languageCuda
License-
Stars gained today363
Created-
Last pushed-

Trending History

Daily boardrank #15 · ▲ 363 stars

Related AI Projects

1

mattpocock / skills

Shell★ 277,947⑂ 23,273▲ 972 stars
→
2

affaan-m / ECC

JavaScript★ 274,182⑂ 40,904▲ 731 stars
→
3

NousResearch / hermes-agent

Python★ 251,657⑂ 0
→
4

deepseek-ai / deepseek-harness

TypeScript★ 244,544⑂ 0
→
5

n8n-io / n8n

TypeScript★ 206,762⑂ 0
→
6

firecrawl / firecrawl

TypeScript★ 189,134⑂ 0
→
7

Significant-Gravitas / AutoGPT

Python★ 187,666⑂ 0
→
8

ollama / ollama

Go★ 182,379⑂ 18,117▲ 137 stars
→

More AI Rankings