Back to News

Hetzner launches experimental LLM inference API with one Qwen model

156 points · 87 comments#llm#inference#hetzner#api

Hetzner has introduced Hetzner Inference, an experimental OpenAI-compatible API for LLM inference running on its own infrastructure. The service currently offers only the Qwen/Qwen3.6-35B-A3B-FP8 model and has no billing, SLA, or production guarantee. Hetzner aims to gauge user interest and test scalability before considering a full product launch.

Coverage timeline

  1. Hacker Newsjonas_scholz

    Hetzner is experimenting with LLM inference. That is not a sentence I expected to write, but I think it is pretty interesting :) Before anyone moves their production AI workloads to Hetzner: this is very much an **experiment**. There is no billing, no SLA, no production guarantee, and currently only one model. Hetzner says it wants to learn whether people actually want this, how the system scales, which features matter, and what kind of load it can handle. So this is not a finished product launch. It is Hetzner putting something early in front of users and seeing what happens. I really like that approach. ## What Is Hetzner Inference? Hetzner Inference is an OpenAI-compatible API running on Hetzner's own infrastructure. You create an API token in the Experiments dashboard, point an OpenAI client at Hetzner's base URL, and use it like most other inference APIs. Right now, the only available model is `Qwen/Qwen3.6-35B-A3B-FP8`. It is a 35-billion-parameter Mixture-of-Experts model with 3