The short version
Open benchmark for deterministic structured outputs across text, image, and audio.
From the HackerLinks archive
Open benchmark for deterministic structured outputs across text, image, and audio.
The short version
Open benchmark for deterministic structured outputs across text, image, and audio.
Why it caught our attention
It targets a real agent failure mode: valid JSON with wrong values.
Where it surfaced on Hacker News
Editorial paraphrase
Commenters focused on value accuracy, modality-specific rankings, and the need to measure structured hallucinations.
Original threadShow HN: A new benchmark for testing LLMs for deterministic outputs