The short version
A speculative-decoding drafter collection for accelerating supported local language models.
From the HackerLinks archive
A speculative-decoding drafter collection for accelerating supported local language models.
The short version
A speculative-decoding drafter collection for accelerating supported local language models.
Why it caught our attention
A local-model user reported a large throughput increase with DFlash enabled.
Where it surfaced on Hacker News
“So yeah base might be 30tps (I used iq4) but mtp or dflash help a lot and should be used when checking what is useful and what is not for running models as it is not fare to judge without them.”
alex7o · recommendation · first hand use
Direct HN comment
M5 Ultra Mac Studio Review