MikeTrendsTrends right now

⬢github Cuda · 8.7K ★ +131 since we first saw it · pushed 6 d ago · MIT

deepseek-ai/DeepGEMM

DeepGEMM: clean and efficient BLAS kernel library on GPU

DeepGEMM is an open-source CUDA library from DeepSeek providing high-performance tensor core kernels for LLM workloads: FP8, FP4, and BF16 GEMMs, fused MoE with overlapped communication (Mega MoE), and indexer scoring kernels. Kernels are compiled at runtime via DeepJIT, so no CUDA compilation is needed at install. It aims for a small, readable codebase while matching or beating expert-tuned libraries.

Why now: A recent update (2026.09.30) added locality domain optimizations and a DeepGEMM-Ascend port for Huawei Ascend hardware, alongside earlier additions like Mega MoE and FP8xFP4 GEMM.

Who it is for: GPU engineers and LLM inference/training developers who need fast, hackable tensor core kernels on NVIDIA (and now Ascend) hardware.

cudagpu-kernelsllmblasfp8moe

Open on GitHub →

Stars over our 31 snapshots: 8.5K to 8.7K, since 9 h ago.

Where people talked about it

API: https://socialmediatrends-api.osmike.com/v1/repos/deepseek-ai/DeepGEMM