Back to News

vLLM Adds Tenstorrent Plugin for Mesh-Optimized LLM Serving

#vllm#tenstorrent#llm-serving#mesh-architecture

vLLM announced an out-of-tree platform plugin enabling LLM serving on Tenstorrent accelerators, with optimizations tailored to the mesh architecture. Key features include phase-based scheduling, single-process data parallelism on Galaxy systems, on-device sampling with host fallback, and asynchronous decode overlap.

Coverage timeline

  1. vLLM BlogTenstorrent Team

    Tenstorrent accelerators join vLLM as an out-of-tree platform plugin, driven by mesh-architecture choices: phase-based scheduling, single-process data parallelism on Galaxy, on-device sampling with host fallback, and async decode overlap.