Skip to content
Final StateThe Jagged Frontier: The Edge You Can't See Until You're Past It
VOL. I  ·  NODE 113▢  ATLAS

TWO TASKS, ONE AFTERNOON

The Jagged Frontier: The Edge You Can't See Until You're Past It

A jagged frontier is the uneven edge of model capability: tasks that look equally hard to a person can sit on opposite sides.

NOT A FENCE, A COASTLINE

The boundary is not a line

Capability boundary drawn as a jagged coastline rather than a straight fence.The figure shows bays of failure reaching into apparently solid ground, so task appearance does not reliably reveal the edge.THE EDGE YOU EXPECTEQUALLY HARD TO THE EYEINSIDE THE FRONTIEROUTSIDE FRONTIEREASYHARDAPPARENT DIFFICULTY
  • A fence: easy inside, hard outside, failures announced
  • A coastline: bays of failure reach deep inside ground that looks solid
  • Two equally hard-looking tasks can land on opposite sides

The difficulty a person can see is not the boundary that matters. The edge is jagged, and it does not signal where it turns.

SEVEN HUNDRED AND FIFTY-EIGHT

Inside the line, the gains are real

Bar chart of BCG study gains on 18 inside-frontier tasks: quality up about 40%, speed up 25.1%, and tasks completed up 12.2%.The chart covers the study's 18 realistic inside-frontier consulting tasks, not the selected outside-frontier failure task.758 CONSULTANTS · 18 REALISTIC TASKSBOSTON CONSULTING GROUP · GPT-4 · 2023+40%QUALITY25.1%FASTER+12.2%TASKS DONEVS CONSULTANTS WITHOUT AI
Source: Dell'Acqua, McFowland, Mollick et al., HBS Working Paper 24-013, 2023; preregistered experiment with 758 BCG consultants.

This is the half the headlines carried, and it is true. Give the machine well-shaped work and it lifts a whole team — the wins you can actually measure.

  • 758 consultants; three randomized conditions
  • Inside-frontier quality about +40%, work 25.1% faster
  • 12.2% more tasks completed

THE PLANTED TASK

One task, a step past the edge

Selected outside-frontier task where AI users were 19 percent less likely to produce a correct solution.The exhibit uses a normalized comparison, not invented absolute rates: relative likelihood of a correct answer is 1.00 without AI and 0.81 with AI.RELATIVE LIKELIHOOD CORRECT1.00WITHOUT AI0.81WITH AI19% LESSLIKELYASLEEP ATTHE WHEEL
Source: Dell'Acqua, McFowland, Mollick et al., HBS Working Paper 24-013, 2023; one selected outside-frontier task.
  • Reconcile quantitative data against interview evidence
  • The selected task sat outside the measured GPT-4 frontier
  • With AI: 19% less likely to produce a correct solution

The paper describes this over-reliance as falling asleep at the wheel: fluent output can suppress the checking the task still needs.

Capability can fall away while the confident tone holds.

How sure the output sounds need not reveal which side of the boundary a task occupies. Fluency is not an external check.

CENTAURS AND CYBORGS

Two reported habits near the frontier

Reported centaur and cyborg work habits near the frontier, both emphasizing verification.The comparison scopes the labels to the BCG study's reported high-performer habits: split tasks or intertwine work, but keep checking because neither reveals the edge.REPORTED CENTAURREPORTED CYBORGHUMANAISPLIT TASKSKEEP CHECKINGINTERTWINEVERIFY EACH PASSREPORTED HABITSNOT A MAP OF THE EDGE
  • Centaur: split subtasks by what the person can safely verify
  • Cyborg: work back-and-forth, checking each pass
  • In the BCG study, these described high performers; they are not proof the edge is known

The point is discipline under uncertainty, not a mascot. Not to be confused with the reverse centaur — the arrangement flipped, with the machine holding the plan and the human reduced to its hands.

AN UNEVEN EDGE

AI capabilities cover an expanding, but uneven, set of knowledge work we call a jagged technological frontier.

Dell'Acqua, McFowland, Mollick et al., "Navigating the Jagged Technological Frontier," 2023

The paper that named the shape. The load-bearing word is uneven: the reach grows with every release, but never into a clean circle you could fence.

WHAT THE STUDY DOES NOT MAP

A result is not a diagnosis

Evidence boundary separating the study's demonstrated performance result from OOD status and untested causal explanations.The supporting figure prevents the outside-frontier result from being recast as proof of OOD or a demonstrated account of why participants failed.DEMONSTRATEDSELECTED TASKWITH AI19% LOWERLIKELIHOOD OF ACORRECT SOLUTIONNOT ESTABLISHEDOOD STATUSA SINGLE CAUSEJUDGING HARDERTHAN DOINGRESULT ≠ OOD DIAGNOSISOR CAUSAL ACCOUNTVALIDATE EXTERNALLY
  • Demonstrated: lower correctness with AI on one selected outside-frontier task
  • Not demonstrated: that the task was out of distribution for GPT-4
  • Not established by this result: the causal mechanism or that judging was harder than doing

The experiment establishes the performance contrast, not an OOD diagnosis or a single cause. Treat the mechanism as open and require a cheap external check.

THE COAST THAT MOVES

Each release can redraw the coast

  • An upgrade can expand the frontier — and still leave pockets to retest
  • Last year's safe handoff is only a hypothesis this year
  • Diagnostic: what external check would catch a confident miss here?

Treat each release as a new map to test, then return to the value of an honest "I don't know".

Read the transcript

01 · TWO TASKS, ONE AFTERNOON

It is the same afternoon. The same operator, the same machine, two jobs across the desk an hour apart. The first comes back sharper than the work would have been alone, and the day feels lighter for it. The second looks no harder. A memo, a read of a market, a call on a number. It comes back just as quick, just as clean, just as sure of itself. Only this one is wrong. Nothing on the page says so. The voice that got the first one right is the voice that got this one wrong, and between them it never changed its tone.

02 · NOT A FENCE, A COASTLINE

The trouble starts with what we expect. We expect a tool to have a smooth edge. Easy things it does, hard things it fails, and the failures announce themselves. This tool does not work that way. Its edge is jagged. A coastline, not a fence. Two tasks that look equally hard to a person can sit on opposite sides of it. One is inland, well inside what the machine can do. The other is a single step into the water, and there the machine drowns without a splash. From where you stand, the two look identical. The difficulty you can see is simply not the boundary that decides.

03 · SEVEN HUNDRED AND FIFTY-EIGHT

None of this is a fable. In 2023, Harvard Business School and Boston Consulting Group ran a preregistered experiment with seven hundred and fifty-eight consultants. After a baseline, participants entered one of three randomized conditions: no AI, GPT-4, or GPT-4 with a prompt-engineering overview. Across eighteen realistic tasks selected to sit inside the model's frontier, the gains were not small. Those using AI completed twelve point two percent more tasks, worked twenty-five point one percent faster, and produced work rated about forty percent higher in quality. This is the part the headlines carried, and it is true. But the experiment also included a task selected to sit outside the frontier. That is the result worth remembering.

04 · THE PLANTED TASK

The researchers selected one task a single pace past the edge. To the eye, it looked no harder than the rest: reconcile quantitative data against interview evidence and recommend which brand deserved investment. The machine took the bait and walked its users off the cliff. Consultants using AI were nineteen percent less likely to produce a correct solution than consultants without AI. The paper's phrase for the over-reliance is unkind and accurate: falling asleep at the wheel. Fluent output can look finished enough to suppress the checking the task still needs.

05 · THE VOICE HOLDS

Here is the danger in a single line. Capability can fall away while the confident tone holds. Step across the edge, and the machine may sound as certain as it did on a task it handled well. No tremor, no hedge, no tell you can safely treat as a verifier. How sure the output sounds need not reveal which shore you are on. Fluency is not an external check. The frontier is hard to see from the inside because the voice need not bend where the ground gives way.

06 · CENTAURS AND CYBORGS

And yet some people worked that ground more safely. The consultants who did best were not simply the ones who leaned hardest on the machine. The study described two habits among high performers. The centaur draws a clean line down the middle: the person keeps the work they can judge and verify, the machine takes the work it appears to do well, and the seam stays visible. The cyborg blends the two but keeps checking, trading lines back and forth, catching slips as they appear. The evidence is narrower than the metaphor: in this experiment, these were successful patterns for keeping person and model from merging into unchecked automation. What they share is a person who never assumes which side of the edge a task is sitting on.

07 · Advertisement · 11x AI Growth Workers

Pipeline is not one task. It is a chain of small handoffs: find the buyer, research the account, respond fast, qualify cleanly, follow up on time. 11x gives sales, marketing, and RevOps teams digital workers for that background motion. The system keeps work moving, while people focus on the conversations that deserve them.

08 · AN UNEVEN EDGE

The paper gave the shape its name. Dell'Acqua, Mollick and seven others, spread across business schools and the Boston Consulting Group, called it a jagged technological frontier. Listen to the word carrying the weight. Not expanding, though it is. Uneven. The reach of these models grows with every version, but it never grows into a tidy circle you could draw a fence around. It grows the way a coastline grows: in fingers and inlets, with bays of failure reaching far inside ground that looks perfectly solid. That is why no one hands you the map. The edge is real, and it is ragged, and it will not hold still.

09 · WHAT THE STUDY DOES NOT MAP

Keep the evidence boundary visible. The experiment showed that participants using AI were nineteen percent less likely to solve one selected outside-frontier task correctly. It did not demonstrate that this task was out of distribution relative to GPT-4's training or deployment distribution. It also did not isolate a single causal mechanism, or establish that judging the answer was harder than doing the work. Those may be hypotheses for later tests, not findings to import into this result. What the result supports is operational caution: validate the task externally instead of treating fluency as a capability signal.

10 · THE COAST THAT MOVES

One last thing, and it is the cruel one. The frontier moves. A new model can redraw the coastline, often pushing it outward, but sometimes changing the shape in ways that make an old handoff unsafe. What you safely handed the machine in the spring should be retested by autumn, because the same flat, confident voice may hide a new failure. You should not count on a release note for that. So the frontier is not a fact you learn once. It is a thing you re-earn every time the tool changes. Before you hand it over again, ask one operator's question: what external check would catch a confident miss here?

01 / 10 · TWO TASKS, ONE AFTERNOON0:00 / 7:20