01 · THE WATCHMAN'S ALARM
A watchman keeps a detector with what sounds like a tight spec: it catches the real alarm, and on quiet nights it cries wolf only one percent of the time. That sounds like a tool you would trust with your life. Listen to what it does to him instead. The nights stack up. Alert, alert, alert, and each one, checked, turns out to be nothing. A fox. A gust. A shadow. He runs them all down anyway, because the one time he doesn't is the time it matters. Except it never seems to matter. Slowly the alerts stop meaning danger and start meaning noise. And on the night the wolf finally comes, its alarm looks exactly like the false alarms that came before. His hand is already halfway to the switch.
02 · RARITY MOVES THE GOALPOSTS
Here is why the good detector fails him. Ninety-nine percent specificity means only one false alarm in a hundred quiet cases. That sounds like the whole story, but it is not. When the event you are hunting is rare, the number that decides the system is not a blended accuracy score. It is the false-positive rate at the real base rate. Think about the arithmetic underneath. There is a vast pile of non-events, and a sliver of real ones. Even a tiny false-positive rate, applied to that vast pile, throws off a mountain of false alarms, while the true events, being rare, stay a sliver. The detector can be useful and still bury you if the base rate is low enough. This is not just a bad model. It is the mathematics of rarity, and a summary accuracy figure can hide it.
03 · NINETY-NINE OF A HUNDRED
Let us do it slowly, with numbers you can hold. Say the event strikes once in ten thousand. Give the detector perfect recall and ninety-nine percent specificity. It catches the one real event, and misfires on one percent of the nine thousand nine hundred and ninety-nine non-events. That produces about one hundred false alarms. So the alert pile holds about one hundred and one flags, and just one is real. Roughly ninety-nine percent of the alerts are wrong, and precision is about one percent. Recall says you caught the event. Precision says the operator now has to find it in the noise.
04 · THE PERSONALITY SKETCH
And this is not only a machine problem. It is an old human one. In the early nineteen-seventies, two psychologists, Daniel Kahneman and Amos Tversky, ran a now-famous test. They handed people a short personality sketch, said to be drawn at random from a hundred professionals. One group was told the hundred held seventy engineers and thirty lawyers. Another was told the reverse. That mix, the base rate, should swing a probability-aware guess. It barely moved the needle. People judged by how much the sketch sounded like an engineer or a lawyer, and all but ignored the odds. The founders of behavioural economics had caught the human mind doing a version of what the watchman does.
05 · INSENSITIVITY TO PRIOR PROBABILITY
They gave the effect a careful name: insensitivity to prior probability of outcomes. Say it slowly, because every word earns its place. Insensitivity. Not ignorance; we can state the base rate if asked, we simply do not feel its weight. Prior probability. The odds before any evidence arrives, the seventy in a hundred. Of outcomes. The thing you are trying to predict. Later it got a blunter name: base-rate neglect. The vivid detail in front of you, the story that fits, crowds out the quiet number that should have anchored the whole judgment. And a fluent, confident AI flag can be exactly that kind of vivid detail. So a machine flag does not automatically cure the bias. In the wrong design, it supplies it at scale.
06 · THE INTRUSION DETECTOR
Security researchers ran straight into this wall. In 2000, Stefan Axelsson formalised it for intrusion detection: when the attack is rare, the variable that can bind the whole system is the false-positive rate. If that rate is too high, the operator drowns. But drowning is only the first harm. The second is quieter. Live with false alarms long enough and a team can stop treating them as alarms at all. Diane Vaughan's term normalization of deviance describes the related organizational drift: repeated departures become the background people learn to accept. Applied here, the flood gets muted, and the one real signal can be muted with it.
07 · Advertisement · Bubble AI App Builder
Some ideas do not need another document before they become testable. With Bubble AI, you describe the app you want, and Bubble creates a working starting point: interface, data, and logic you can inspect. From there, you refine visually, connect AI models and services, and turn the first version into something real enough to use.
08 · THE ONE YOU MUTED
A threshold is not simply right or wrong. It selects an operating point on the detector's ROC tradeoff. Raise the threshold and the false-positive rate may fall, but recall usually falls as more true events are missed. Lower it and recall may improve while the alert queue grows. At the deployed base rate, precision follows from that operating point and the event prevalence. Choose the threshold from the cost of false alarms, the cost of misses, capacity to investigate, and required precision. In some systems a threshold can meet those requirements; in others no available point can.
09 · WHAT SURVIVES THE FLOOD
Conjunction can help, but only under conditions. If channels make sufficiently independent errors, requiring several signals to agree may lower the joint false-positive rate. The same requirement can miss true events that appear in only one channel, so recall may fall. Correlated channels can simply repeat the same error and provide little gain. Measure both sides at the deployed base rate. Then use the surviving flags for triage: narrate the evidence, state a counter-argument, and run a cheap real-world probe before acting. The conjunction is a filter with a tradeoff, not a truth machine.
10 · THE FIRST ENEMY
So this is the first enemy of every rare-event check you build. Rarity floods the signal until the truth is a needle in a field of needles. And it does not travel alone. Its partner is fluency, the trouble of judging work in a field you cannot command, where confidence reads as skill. Rarity buries the true signal. Fluency talks you into the false one. Between them, they can make the verifier you were counting on fail without a sound. Which returns us to the one honest reading a machine can offer at the edge. Not a louder alarm. A quieter, rarer word: I don't know. That word is one thing the flood cannot easily fake.