June Kim · Research Engineer, AI Agent Systems
I turn ambiguous questions about agent behavior into environments, experiments, graders, and tools.
I build the system, run it against real workflows, find where its feedback becomes unreliable, and turn the result into a better product or protocol. I bring 10+ years shipping across Google, Loom, and startups. Upstream, I've landed fixes in Rust, Go, C++, and Python projects. I'd rather hand you something you can verify than something you have to believe.
Case studies
- Two benchmarks changed after my audits. I built and ran public construct-validity audits of frontier coding benchmarks (ProgramBench, SWE-bench Pro, DeepSWE, and Terminal-Bench) and extracted determinacy, a reusable auditing tool. DeepSWE shipped a re-graded v1.1 eighteen days after my audit; exactly the four tasks I flagged climbed while the pooled rate held flat. OpenAI retracted its SWE-bench Pro recommendation twenty-nine days after my audit went up on Scale's tracker. I claim the dates, not causation; the dates sit on other people's servers. What the comparison is actually about.
- Receipts anyone can re-run. Every audit ships versioned receipts to Zenodo with a regrade script: Terminal-Bench audit, SWE-bench Pro run, Assurance at the Boundary. Anyone can re-run the grading and disagree with me on the record. The fixes go upstream too: harbor#2266 implements the Terminal-Bench frame gate as an opt-in observational check (CI green, open), and inspect_ai#4462 does the same for Inspect. When I say a benchmark is broken, I send the patch.
- Engineering system, merged by strangers. My agentic contribution pipeline turns issue discovery, patch generation, testing, and maintainer feedback into the loop that produced 100 pull requests merged across 81 external repositories in 2026 (GitHub search, July 26), including hyper, TiDB, Servo, Enzyme, wild, and flux. A maintainer pressing merge is a cost paid by someone with no reason to flatter me. Re-run the count.
- Research mechanism. The Hypothesis Graph is a harness-layer data structure for coding agents: testable-claim nodes and refutation-condition edges that make an agent's reasoning inspectable and reusable. It's the harness behind the pipeline above.
Roles I fit
- Research engineering for agent environments, evaluation infrastructure, and graders.
- Applied AI work where an ambiguous question about agent behavior has to become a reproducible experiment, then a shipped system.
- Agent reliability: diagnosing broken feedback loops, misleading success signals, and failures that pass current checks.
More: Verifiable Knowledge, Slop Slope, and the theory layer.
Resume & Contact
Resume page · PDF · Markdown
june@june.kim · LinkedIn · GitHub
Based in Vancouver, Canada (Pacific time). My LinkedIn says Bay Area for visibility.