From the HackerLinks archive

Qwen3.8-27B

A 27B vision-language model that runs locally with long context.

At a glance:
First seen:2026-09-09
Last seen:2026-09-09
Times seen:1
Website:huggingface.co

The short version

A 27B vision-language model that runs locally with long context.

Why it caught our attention

Two RTX 5090 users reported excellent speed and strong real-world results.

Where it surfaced on Hacker News

2026-09-09

I'm now running NVFP4 quantized both weight and cache on my RTX5090 and getting excellent results: 264k cache allocated for pool, 10k tok/s prompt processing, 200 tok/s generation for single stream, or 801 tok/s generation for 8 concurrent streams.

skolos · recommendation · first hand use

Direct HN comment

For the kinds of things I use a local model for (legal document review), it’s just spectacular. It also has good vision support. I’ve been using 27B more and more over DS4.

D13Fd · recommendation · first hand use

Direct HN comment

Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses