The short version
Cursor's coding-model evaluation suite for comparing practical agent performance.
From the HackerLinks archive
Cursor's coding-model evaluation suite for comparing practical agent performance.
The short version
Cursor's coding-model evaluation suite for comparing practical agent performance.
Why it caught our attention
A commenter said its rankings consistently matched their subjective model evaluations.
Where it surfaced on Hacker News
“Their bench has always been one that most-fit my mental model of how good each of these models are.”
jjcm · recommendation · evaluated
Direct HN comment
Artificial Analysis Intelligence Index v4.2
Also surfaced in this discussion
Artificial Analysis Intelligence Index v4.2
2026-09-06