The short version
128B model people weighed on local speed, cost, and quantization tradeoffs.
From the HackerLinks archive
128B model people weighed on local speed, cost, and quantization tradeoffs.
The short version
128B model people weighed on local speed, cost, and quantization tradeoffs.
Why it caught our attention
It triggered serious debate about what local LLMs can realistically do.
Where it surfaced on Hacker News
Editorial paraphrase
HN compared its Pareto efficiency, local VRAM needs, and token throughput against Sonnet and other frontier models.
Original threadMistral Medium 3.5