The short version
A public demo of extremely high-throughput language-model inference.
From the HackerLinks archive
A public demo of extremely high-throughput language-model inference.
The short version
A public demo of extremely high-throughput language-model inference.
Why it caught our attention
Multiple commenters singled out the experience as unusually compelling despite the small model.
Where it surfaced on Hacker News
“I freakin' love this demo. It feels magical.”
wxw · recommendation
Direct HN comment
“This is the coolest LLM thing I’ve seen since the original ChatGPT announcement a few years ago. IMO much more impressive than marginal gains of frontier models.”
appplication · recommendation
Direct HN comment
AMD acquires Taalas to boost inference performance by etching models in silicon