vLLM Adds Day-0 Support for NVIDIA Vera Rubin NVL72, Claims 7.8x Throughput over GB200
vLLM announced support for NVIDIA Vera Rubin NVL72, claiming 7.8x throughput over GB200 NVL72. The platform offers daily container builds and day-0 model support for DeepSeek, Moonshot AI, Z.ai, and MiniMax, leveraging Blackwell kernel compatibility and Rubin-tuned kernels via FlashInfer 0.7.0.
Coverage timeline
vLLM BlogvLLM Team, Inferact, Red Hat, and NVIDIA
vLLM on NVIDIA Vera Rubin NVL72 ## vLLM now supports Vera Rubin NVL72! NVIDIA Vera Rubin is the next-generation platform built for agentic inference. Inferact, NVIDIA, Red Hat, and the vLLM community have been bringing vLLM up on Vera Rubin NVL72 since it was announced, and vLLM runs on Vera Rubin NVL72 today with daily container builds and support for models from DeepSeek, Moonshot AI, Z.ai, and MiniMax. This post is an early look at where things stand, and here are a few highlights from the work so far: * **Vera Rubin NVL72 hardware:** 5x the NVFP4 FLOPS, about 2.4x the HBM bandwidth and 1.7x the bidirectional NVLink bandwidth of GB200 NVL72, with 2-4x faster exponentials for softmax. * **Day-0 support:** Rubin builds on Blackwell's architecture family, so vLLM’s Blackwell kernels are compatible with Rubin. Thanks to this, vLLM already supports diverse models such as DeepSeek, Kimi, GLM, and MiniMax on Rubin. * **Rubin-tuned kernels:** Through FlashInfer 0.7.0, vLLM gets Rubin-tuned