The short version
Sparse-attention model chasing million-token context on less compute.
From the HackerLinks archive
Sparse-attention model chasing million-token context on less compute.
The short version
Sparse-attention model chasing million-token context on less compute.
Why it caught our attention
Commenters praised the tech and highlighted its large compute and speed gains.
Where it surfaced on Hacker News
Editorial paraphrase
One commenter said they loved the tech; another pointed to the 64.5x compute reduction and 56x speedup claim.
Original threadSubQ 1.1 Small