Back to News

NVIDIA Vera Rubin NVL72 shows up to 7x token throughput per MW vs Blackwell on 1.6T DeepSeek model

#nvidia#inference#blackwell#vera-rubin

SemiAnalysis reports that NVIDIA's Vera Rubin NVL72 inference tests show up to 7x better token throughput per MW vs. Blackwell on a 1.6T DeepSeek model. This exceeds Jensen Huang's earlier claim of 3x for 1T-3T LLMs, suggesting the new platform may offer greater efficiency gains than officially stated.

Coverage timeline

  1. Techmeme

    Bryan Shan / SemiAnalysis : Vera Rubin NVL72 inference tests show up to 7x better token throughput per MW vs. Blackwell on a 1.6T DeepSeek model, above Huang's 3x claim for 1T-3T LLMs — Jensen Sandbagging Performance Again, 2x more Annual Profit Per GigaWatt, The More you Buy, The More you Earn, AgentX, InferenceX, Extreme Co-Design