The short version
A 27B vision-language model that runs locally with long context.
From the HackerLinks archive
A 27B vision-language model that runs locally with long context.
The short version
A 27B vision-language model that runs locally with long context.
Why it caught our attention
Two RTX 5090 users reported excellent speed and strong real-world results.
Where it surfaced on Hacker News
“I'm now running NVFP4 quantized both weight and cache on my RTX5090 and getting excellent results: 264k cache allocated for pool, 10k tok/s prompt processing, 200 tok/s generation for single stream, or 801 tok/s generation for 8 concurrent streams.”
skolos · recommendation · first hand use
Direct HN comment
“For the kinds of things I use a local model for (legal document review), it’s just spectacular. It also has good vision support. I’ve been using 27B more and more over DS4.”
D13Fd · recommendation · first hand use
Direct HN comment
Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses