Back to News

vLLM Announces vllm-metal for Apple Silicon with Concurrent Serving

#vllm#apple-silicon#serving#inference

vLLM announced vllm-metal, bringing its paged, continuously batched serving stack to Apple Silicon. The new offering features flatter time-to-first-token under concurrent agent load, batched multi-token prediction, and automatic M5 prefill acceleration.

Coverage timeline

  1. vLLM BlogRanran Haoran Zhang, Lik Xun Yuan, Chao Ju Chen, Eric Curtin, Michael Goin

    vllm-metal brings vLLM's paged, continuously batched serving stack to Apple Silicon, with flatter TTFT under concurrent agent load, batched MTP, and automatic M5 prefill acceleration.