Guided tour
Stress-testing agents for failure boundaries
Start the tour
Twelve steps. Method and labels first, then the flagship rejection, the essay, the verify-tool ablation, the other measurement stories, the runnable method, and limitations.
H1 rejection
34/36 beats 28/36, still not promoted. Honesty gate; H0 could not emit that metric.
The harness that scored better
Same story in longform, with the verify-tool sequel.
Verify tool ablation
Haiku + local coder flip; model-dependence measured, not an abstainer control.
Try the method
Public demo split + selftest in verified-done.
All stories
H1 rejection
Pre-registered honesty gate.
AI safety / eval
Multi-path coding
Multi-file is the breakpoint.
Local coding agents
Task-family transfer
Coding winner fails constrained prose.
Eval design
RAG routes
Use-case routes, not one Elo.
RAG / retrieval
Agent-as-evaluator
Skill edge matrix + thin smoke receipts.
Prompt / skills
Verify tool ablation
Harness × model honesty (Findings A, D, E, F). Finding C withdrawn.
Agent reliability
Method cards
- Public harness descriptions — structure without fixtures
- Harness isolation protocol
Downloads
- H1 rejection one-pager (PDF)
- Multi-path hard screen one-pager (PDF)
- Task-family transfer one-pager (PDF)
Related public work
- verified-done — runnable honesty demo
- claim-audit-lab — claim support vs nomination
- apparatus-contracts — handoff contracts
- research-scaffold-harness — scaffold vs unsupported claims
- evidence-bundler — evidence-bundle pipeline