Terminal-Bench is blind to destruction
Official grading scored 83 of 83 runs as passes after off-task user assets were wiped. The audit produced an executable frame gate and upstream patches for Harbor and Inspect.
I audit the inference from benchmark result to capability claim.
I find the hidden assumption, construct the test that can break it, and return evidence someone else can replay.
Official grading scored 83 of 83 runs as passes after off-task user assets were wiped. The audit produced an executable frame gate and upstream patches for Harbor and Inspect.
The audit proved a 15% floor of tasks whose graded value no reading of the materials determines, and shipped the cases, grading receipts, and regrade path.
Four of 113 answer keys failed their own verifiers. All four climbed in the later re-grade while the pooled rate stayed flat, and the revised v1.1, re-audited, mostly held.
Twenty-one tasks carry a verified recall-only witness, so the zero-percent floor across nine models measured recall of published algorithms rather than reconstruction.
Repeated audits expose the same boundary failure in different forms: the grader certifies less than the benchmark claims.
How to Audit a Benchmark gives the full operational checklist. Assurance at the Boundary collects eight public-side audits and their receipts.
About · GitHub · june@june.kim