THE HINGE
Cheap to Make, Expensive to Check
Beneath every promise of leverage sits one asymmetry: cheap output is useful only to the extent that a relevant check can reject a plausible error.
- Generation can collapse before review does
- Verification may still consume expert time
- Price both costs before offloading
THE GATE
The Generation-Verification Gap
- Saad-Falcon et al., arXiv 2506.18203v2 (2025): a correct candidate can be generated but not selected
- The paper's LM judges and reward models are weak, imperfect verifiers, not ground truth
- Operator rule: offload only when a relevant check is reliable and materially cheaper than remaking
The paper defines a model-selection failure, not this article's business-cost inequality. The practical bridge is to price verification separately and reject another fluent answer as proof.
THE LAST THIRTY
The 70% Problem
The first stretch was never the whole job. The 70% problem is a cost-shape heuristic: the residue matters because checking it can consume the expertise the draft appeared to save.
- First 70%: scaffolding, patterns, boilerplate
- Last 30%: edge cases, security, judgment
- Osmani, The 70% Problem (2024): illustrative split, not a measured ratio
REVIEW IS THE WORK
Checking is not what comes after the work. It is the work.
Final State editorial rule
Not a rubber stamp but the actual labour — and when checking nears making, oversight becomes re-work.
WHERE ORACLES LIVE
Verifiable Space
- AlphaEvolve (Google DeepMind, 2025): automated evaluators score proposed programs
- GNoME (Merchant et al., Nature 624, 2023): DFT checks predicted computational stability, not usefulness or synthesizability
- AlphaProteo (Google DeepMind, 2024): wet-lab assays test binding for specified targets
- Each check is bounded to the property it actually measures
AI can search widely inside verifiable space, but a passed check supports only the property tested. Computational stability is not a useful material; binding is not a medicine.
TWO STUDIES, TWO SETTINGS
Two AI-Assistance Studies, Two Settings
- NBER Working Paper 31161 (2023), conversational assistant, 5,179 support agents: issues resolved per hour +14% average, +34% for novice and lower-skilled workers
- METR early-2025 RCT, primarily Cursor Pro with Claude 3.5/3.7 Sonnet, 16 experienced open-source developers and 246 tasks: completion time +19%; participants believed it fell 20%
- METR, February 2026: the old result no longer reflects current tools; its later experiment could not yield a reliable current estimate
These studies do not isolate verifier cost or explain each other's result. The bounded editorial synthesis is that checking burden varies by workflow and must be measured locally: felt speed is not measured speed, so every offload needs its own baseline and check.
THE SHELF
The Director-or-Limb Shelf
This shelf prices verification, not capability. The reverse-centaur risk begins when approval remains visible but no affordable check can change the result.
- Estimate make and check costs task by task
- A short check counts only when it tests the failure that matters
- When review approaches remake cost, the claimed saving is unproven
BUILD THE CHECK
If No Verifier Exists, Build One
- 01Scheel, Schijen & Lakens, Adv. Methods Pract. Psychol. Sci. 4(2), 2021: 44% positive Registered Reports versus a 96% literature estimate; comparison, not a causal effect
- 02Seal the source, threshold, and stop rule before seeing the output
- 03Use a fresh-model attack as a challenge, not independent assurance; reserve that label for an outside reviewer
Adapt the Three Lines Model without pretending one person has created formal independence: precommit first, challenge second, and schedule a genuinely outside review for consequential work.
DIRECTOR OR LIMB
The Question That Decides Your Role
If you cannot point to the source, threshold, and stop rule, the generation-verification gap is still open. Do not count generated output as a saving until the check closes it.
- Before offloading, name the output, source, threshold, owner, and stop rule
- Test the verifier against plausible failures before trusting it
- No relevant, affordable check: build one, narrow the task, or keep the work human