Research

I audit the inference from benchmark result to capability claim.

I find the hidden assumption, construct the test that can break it, and return evidence someone else can replay.

Selected findings

Terminal-Bench is blind to destruction

Official grading scored 83 of 83 runs as passes after off-task user assets were wiped. The audit produced an executable frame gate and upstream patches for Harbor and Inspect.

ProgramBench measures recall

Twenty-one tasks carry a verified recall-only witness, so the zero-percent floor across nine models measured recall of published algorithms rather than reconstruction.

FrontierCode, MirrorCode, and τ²-bench

Repeated audits expose the same boundary failure in different forms: the grader certifies less than the benchmark claims.

Method

  1. Name the capability claim and the conclusion drawn from its score.
  2. Read the task, environment, and grader as one measurement instrument.
  3. Construct the smallest counterexample that separates the claim from the check.
  4. Publish pinned artifacts, observed outcomes, and a regrade command.
  5. Disclose the finding and send the fix upstream.

How to Audit a Benchmark gives the full operational checklist. Assurance at the Boundary collects eight public-side audits and their receipts.

Research mechanisms

Papers

About · GitHub · june@june.kim