Skip to content
Final StateCheap to Make, Expensive to Check
VOL. I  ·  NODE 005▢  ATLAS

THE HINGE

Cheap to Make, Expensive to Check

Cost-curve chart showing making work getting cheaper while checking remains costly.The conceptual figure separates falling generation cost from task-specific verification cost; the widening gap shows why output volume alone does not establish a defensible offload.GENERATION FELL FIRSTCOST(CHECK): TASK-SPECIFICBEFOREAFTERUNPRICED GAPCOST(MAKE)PRICE THE CHECK SEPARATELY

Beneath every promise of leverage sits one asymmetry: cheap output is useful only to the extent that a relevant check can reject a plausible error.

  • Generation can collapse before review does
  • Verification may still consume expert time
  • Price both costs before offloading

THE GATE

The Generation-Verification Gap

Offload gate requiring a relevant, reliable, separate, and cheaper check.Four rows test whether the verifier targets the important failure, has a known miss rate, is more than model self-agreement, and costs less than remaking the work.THE OFFLOAD GATECOST(CHECK) < COST(REMAKE)RELEVANTTESTS THE FAILURERELIABLEKNOWN MISS RATESEPARATENOT SELF-AGREEMENTCHEAPERCHECK < REMAKEALL FOURANY ONE FAILSOFFLOADBUILD / NARROW / KEEPCHEAP ALONE IS NOT ENOUGH
  • Saad-Falcon et al., arXiv 2506.18203v2 (2025): a correct candidate can be generated but not selected
  • The paper's LM judges and reward models are weak, imperfect verifiers, not ground truth
  • Operator rule: offload only when a relevant check is reliable and materially cheaper than remaking

The paper defines a model-selection failure, not this article's business-cost inequality. The practical bridge is to price verification separately and reject another fluent answer as proof.

THE LAST THIRTY

The 70% Problem

Effort curve with a fast first seventy percent and a steep last thirty percent.The curve treats seventy/thirty as an illustrative work shape: scaffolding and boilerplate come quickly, while edge cases, correctness, security, and judgment concentrate at the end.THE COSTLY RESIDUE COMES LASTFAST 70%SCAFFOLDINGLAST 30%EDGE CASESILLUSTRATIVE, NOT MEASURED
Illustrative work shape, not a measured ratio.Addy Osmani, The 70% Problem, 2024

The first stretch was never the whole job. The 70% problem is a cost-shape heuristic: the residue matters because checking it can consume the expertise the draft appeared to save.

  • First 70%: scaffolding, patterns, boilerplate
  • Last 30%: edge cases, security, judgment
  • Osmani, The 70% Problem (2024): illustrative split, not a measured ratio

REVIEW IS THE WORK

Checking is not what comes after the work. It is the work.

Final State editorial rule

Not a rubber stamp but the actual labour — and when checking nears making, oversight becomes re-work.

WHERE ORACLES LIVE

Verifiable Space

Search space with candidate outputs screened by mathematical, physical, and laboratory verifiers.The figure shows AI replacing search inside verifiable spaces while external checks reject proposals; checkable does not mean automatically true.THREE CHECKS, THREE BOUNDSALPHAEVOLVE '25CHECK: AUTO EVALUATORPASSES: PROGRAM SCOREGNOME / NATURE '23CHECK: DFT CALCULATIONPASSES: PREDICTED STABILITYALPHAPROTEO '24CHECK: WET-LAB ASSAYPASSES: TARGET BINDINGA PASS SUPPORTS ONLYTHE PROPERTY TESTEDSTABLE != USEFUL; BINDING != MEDICINE
  • AlphaEvolve (Google DeepMind, 2025): automated evaluators score proposed programs
  • GNoME (Merchant et al., Nature 624, 2023): DFT checks predicted computational stability, not usefulness or synthesizability
  • AlphaProteo (Google DeepMind, 2024): wet-lab assays test binding for specified targets
  • Each check is bounded to the property it actually measures

AI can search widely inside verifiable space, but a passed check supports only the property tested. Computational stability is not a useful material; binding is not a medicine.

TWO STUDIES, TWO SETTINGS

Two AI-Assistance Studies, Two Settings

Two-column exhibit comparing a 2023 support-agent field study with METR's bounded early-2025 developer trial.The left column reports the NBER support-agent field study. The right reports METR's bounded early-2025 developer trial and marks METR's 2026 warning that the result is historical, not current.TWO STUDIES, TWO SETTINGSPEOPLE / TASKS / TOOLS / DATES DIFFERNBER W31161 / 20235,179SUPPORT AGENTS+14%AVG+34%NOVICEMETR / EARLY 202516 DEVS246 TASKSFELT +20%MEASURED -19%METR 2026: RESULT IS HISTORICALNOT A CURRENT ESTIMATENOT A CAUSAL COMPARISON
Different populations, tasks, tools, and dates; juxtaposition is not a causal comparison.NBER Working Paper 31161 (2023); METR (July 2025, updated February 2026)
  • NBER Working Paper 31161 (2023), conversational assistant, 5,179 support agents: issues resolved per hour +14% average, +34% for novice and lower-skilled workers
  • METR early-2025 RCT, primarily Cursor Pro with Claude 3.5/3.7 Sonnet, 16 experienced open-source developers and 246 tasks: completion time +19%; participants believed it fell 20%
  • METR, February 2026: the old result no longer reflects current tools; its later experiment could not yield a reliable current estimate

These studies do not isolate verifier cost or explain each other's result. The bounded editorial synthesis is that checking burden varies by workflow and must be measured locally: felt speed is not measured speed, so every offload needs its own baseline and check.

THE SHELF

The Director-or-Limb Shelf

Task shelf showing make and check bars for tasks above and below a flip line.Illustrative rows compare make or remake cost with check cost and mark the point where review approaches reconstruction of the task.THE DIRECTOR-OR-LIMB SHELFMAKE / REMAKECHECKCHECK ~= REMAKETEMPLATE EMAILBANK TIE-OUTSECURITY EDGETASTE / VALIDITYA SHORT CHECK MUST TESTTHE FAILURE THAT MATTERSMEASURE COSTS LOCALLY
Illustrative task-cost comparison; estimate with local time, error, and review data.Editorial model

This shelf prices verification, not capability. The reverse-centaur risk begins when approval remains visible but no affordable check can change the result.

  • Estimate make and check costs task by task
  • A short check counts only when it tests the failure that matters
  • When review approaches remake cost, the claimed saving is unproven

BUILD THE CHECK

If No Verifier Exists, Build One

  1. 01Scheel, Schijen & Lakens, Adv. Methods Pract. Psychol. Sci. 4(2), 2021: 44% positive Registered Reports versus a 96% literature estimate; comparison, not a causal effect
  2. 02Seal the source, threshold, and stop rule before seeing the output
  3. 03Use a fresh-model attack as a challenge, not independent assurance; reserve that label for an outside reviewer

Adapt the Three Lines Model without pretending one person has created formal independence: precommit first, challenge second, and schedule a genuinely outside review for consequential work.

Three-part verification discipline separating precommitment, model challenge, and outside assurance.The figure separates precommitment, an adversarial model challenge that is not independent assurance, and a periodic outside human review.REBUILD THE THREE LINESPRECOMMITSEALSOURCETHRESHOLDCHALLENGEATTACKMODELNOT ASSURANCEASSUREOUTSIDEHUMANREVIEWOUTSIDE REVIEW SUPPLIES ASSURANCEDISCIPLINE, NOT FORMAL COMPLIANCE

DIRECTOR OR LIMB

The Question That Decides Your Role

Gate asking for output, source of truth, threshold, owner, and stop rule before offload.The closing figure separates work the operator can direct from work where no cheap independent verifier exists and the seat must stay.BEFORE YOU OFFLOADNAME THE CONTRACTOUTPUTSOURCE OF TRUTHTHRESHOLDOWNERSTOP RULEFAILURE TESTCATCHES WHAT MATTERS?YES: OFFLOADNO: BUILD / NARROW / KEEPCHECKED OUTPUT IS THE GAIN

If you cannot point to the source, threshold, and stop rule, the generation-verification gap is still open. Do not count generated output as a saving until the check closes it.

  • Before offloading, name the output, source, threshold, owner, and stop rule
  • Test the verifier against plausible failures before trusting it
  • No relevant, affordable check: build one, narrow the task, or keep the work human
Read the transcript

01 · THE HINGE

The quarterly report arrives before the analyst has finished opening the source files. It is polished, complete, and cheap to produce. Knowing whether its claims survive those files still consumes the analyst's attention. Generation collapsed first. Verification did not necessarily follow. That split, cheap to make and dear to check, is the quiet hinge the whole era turns on.

02 · THE GATE

A 2025 paper by Jon Saad-Falcon and colleagues gives the generation-verification gap a precise model-level meaning. A correct candidate may be generated, yet an imperfect verifier fails to select it. Their language-model judges and reward models are weak verifiers, not ground truth, and verification can dominate inference cost. That paper does not prove this article's business inequality. It sharpens the warning behind it. Price the check separately, and hand a task over only when that check is relevant, reliable, and materially cheaper than remaking the work.

03 · THE LAST THIRTY

Addy Osmani named a useful work shape in The 70% Problem in 2024. Seventy and thirty are a heuristic, not a measured universal ratio. The fast stretch is scaffolding, familiar patterns, and boilerplate. Then the line turns toward edge cases, security, correctness, and judgment. The exact share changes by task. The economic point survives: a quick draft does not price the residue, and the residue is often where expert checking begins.

04 · REVIEW IS THE WORK

Reviewing the machine's output gets booked as a free action, a glance laid on top of the real labour. It is the costly step itself. When checking a thing costs nearly as much as making it, the human in the loop stops being oversight and becomes unpaid re-work at machine speed. We keep entering the expensive part in the ledger as nothing.

05 · WHERE ORACLES LIVE

Three research systems show the generate-and-test pattern, with three different bounds. Google DeepMind's 2025 AlphaEvolve proposes programs and automated evaluators score them. Merchant and colleagues' 2023 GNoME work filters crystal candidates and uses density-functional calculations to check predicted computational stability, not usefulness or synthesizability. Google DeepMind's 2024 AlphaProteo designs binders and wet-lab assays test binding for specified targets, not whether a binder is a medicine. A passed check supports the property tested, and nothing larger.

06 · TWO STUDIES, TWO SETTINGS

Set two bounded AI-assistance studies beside each other, without treating unlike interventions as a controlled comparison. In the 2023 NBER working paper Generative AI at Work, Brynjolfsson, Li, and Raymond studied five thousand one hundred and seventy-nine support agents using a conversational assistant. Issues resolved per hour rose fourteen percent on average and thirty-four percent for novice and lower-skilled workers. METR's early-2025 trial covered sixteen experienced open-source developers and two hundred and forty-six tasks in familiar repositories, primarily using Cursor Pro with Claude three point five or three point seven Sonnet. With AI allowed, completion time rose nineteen percent, while participants believed it had fallen twenty. METR now says that historical result no longer reflects current tools, and its later experiment could not produce a reliable current estimate. Neither study isolates verifier cost or explains the other's result. The bounded editorial inference is narrower: checking burden differs by workflow, so every claim of speed needs its own baseline and a locally priced check.

07 · Advertisement · 11x AI Growth Workers

Pipeline is not one task. It is a chain of small handoffs: find the buyer, research the account, respond fast, qualify cleanly, follow up on time. 11x gives sales, marketing, and RevOps teams digital workers for that background motion. The system keeps work moving, while people focus on the conversations that deserve them.

08 · THE SHELF

Stand every task on a shelf with two bars, one for making it and one for checking it. Then ask whether the short bar tests the failure that matters. A bank tie-out may be cheap and decisive for a reconciliation. A passing unit test may still miss a security edge. Higher on the shelf, review starts to require the same expertise and reconstruction as making. At that point the claimed saving is unproven. The shelf prices verification, not capability, task by task.

09 · BUILD THE CHECK

When no cheap verifier exists, the task is not human forever. Build the check before you collapse the work. Adapt the Three Lines Model without claiming that one person has manufactured formal independence. First, seal the source, threshold, and stop rule before seeing the output. Second, use a fresh model to attack the work, while remembering that another model is a challenge, not independent assurance. Third, bring an outside reviewer through on a fixed clock for consequential work. Scheel, Schijen, and Lakens reported forty-four percent positive Registered Reports in psychology, compared with an earlier ninety-six percent estimate for the standard literature. That is a comparison, not a causal estimate. Its practical value here is precommitment: define success before the result can bargain with you.

10 · DIRECTOR OR LIMB

The closing test is smaller than it sounds. Before offloading, name five things: the output, its source of truth, the threshold, the owner, and the stop rule. Then test the check against a plausible failure. If the check catches what matters at a cost well below remaking the work, the saving is defensible. If it does not, build a better verifier, narrow the task, or keep the work human. Cheap output is not the gain. Checked output is.

01 / 10 · THE HINGE0:00 / 7:06