Back to News

Tiny LLM Hits 21,000 Tokens/s on $250 FPGA in Live Demo

79 points · 30 comments#fpga#llm#inference#hardware

A Hacker News post showcases a live demo of a tiny language model running at 21,000 tokens per second on a $250 FPGA. The demonstration highlights the potential for low-cost, high-speed inference on edge hardware.

Coverage timeline

  1. Hacker Newsmikeayles