Guided tour
Stress-testing agents for failure boundaries
Thesis. A better aggregate score is not a graduation decision. Pre-register gates for scope and false completion. Treat the harness as an experimental factor, including whether the agent can call an external verifier.
Exploratory
Local + live probes
Not production validation
Runnable demo
Start the tour
Thirteen steps. Method and labels first, then the flagship rejection, the essay, the verify-tool ablation, the other measurement stories, then try the method, limits, and resume language.
Flagship
H1 rejection
34/36 beats 28/36, still not promoted. False completion gate.
Essay
The harness that scored better
Same story in longform, with the verify-tool sequel.
New
Verify tool ablation
Haiku + local coder flip; abstainer control.
Runnable
Try the method
Public demo split + selftest in verified-done.
All stories
01
H1 rejection
Pre-registered honesty gate.
AI safety / eval
02
Multi-path coding
Multi-file is the breakpoint.
Local coding agents
03
Task-family transfer
Coding winner fails constrained prose.
Eval design
04
RAG routes
Use-case routes, not one Elo.
RAG / retrieval
05
Agent-as-evaluator
Skill edge matrix + thin smoke receipts.
Prompt / skills
06
Verify tool ablation
Harness × model honesty (Findings A–D).
Agent reliability
Method cards
- Public harness descriptions — structure without fixtures
- Harness isolation protocol
Downloads
- H1 rejection one-pager (PDF)
- Multi-path hard screen one-pager (PDF)
- Task-family transfer one-pager (PDF)
Related public work
- verified-done — runnable honesty demo
- claim-audit-lab — claim support vs nomination
- apparatus-contracts — handoff contracts
- research-scaffold-harness — scaffold vs unsupported claims
- evidence-bundler — evidence-bundle pipeline