Change your mind by the right amount


How much should one piece of evidence change your mind?

Here’s the version of that question that broke my intuition the first time I met it. There’s a test for some disease, and it’s 99% accurate. You take it. It comes back positive. How worried should you be?

Most people say 99%. To be honest, my gut said something like that too. But for a disease that’s rare enough, the real answer is about 9%. Not ninety. Nine.1

That gap, between the 99% your gut screams and the 9% that’s actually true, is the most useful idea I know for thinking clearly. It has a name that makes it sound like a homework problem (Bayes’ rule), and it almost always gets taught as a formula you plug numbers into. Forget the formula. The formula is the least interesting part. What sits underneath it is a way of changing your mind that, once you see it, you can’t un-see.

Two of the best explanations of this live on YouTube, and they pair up almost perfectly. Julia Galef’s “A Visual Guide to Bayesian Thinking” shows you the discipline: how the good version of this reasoning actually runs.2 Veritasium’s “The Bayesian Trap” shows you where it bites: how the same reasoning, done a little too well, can quietly wall you in.1 Watch them back to back and you get the whole thing. This post is me trying to put both halves in one place, from the ground up.

A belief is a number

Start here, because everything else is a consequence of it.

Most of us carry beliefs as if they were switches. True or false. I’m right or I’m wrong. He likes me or he doesn’t. The first move in Bayesian thinking is to stop doing that and treat a belief as a dial instead: a number between “no way” and “definitely,” somewhere in the middle most of the time. “I’m about 70% sure it’ll rain.”

Sounds trivial. It isn’t. Once a belief is a number, “changing your mind” stops being a mood and becomes arithmetic. The question is no longer “did this prove me right or wrong,” it’s “by exactly how much should this nudge the dial?” A little, if the evidence is weak. A lot, if it’s strong. Bayesian thinking is just the rules for how far the dial moves.

That’s the whole subject. The rest is figuring out what “the right amount” actually is.

Two things, multiplied

Every honest update needs exactly two ingredients, and if you skip either one you get a predictable kind of wrong.

The first is your starting point: how likely was this before the new evidence showed up? How common is the thing in general? This is the boring one, and it’s the one people forget.

The second is how much the evidence discriminates: is this clue genuinely more expected in the world where you’re right than in the world where you’re wrong? This is the subtle one, and it’s the one people overrate.

You multiply them. That’s it. Your new belief is the starting point, scaled by how lopsided the evidence is.

Galef has a lovely way of seeing this as a picture, and it’s worth stealing. She poses a puzzle: you meet a shy, softly spoken guy named Tom on a big university campus. Is he more likely to be a maths PhD student or a business student? Almost everyone says maths. Shyness just fits the mathematician stereotype better, and honestly, that instinct isn’t wrong… shyness probably is more common among maths PhDs. Galef’s version descends straight from a classic Kahneman and Tversky experiment where people did exactly the same thing with a personality sketch, guessing the field that fit the vibe and ignoring how many people were actually in each field.3

Here’s the picture. Draw a rectangle for all the students. Split it into two columns by how many of each there are. There are maybe ten times as many business students as maths PhDs, so the business column is wide and the maths column is a thin sliver. Now shade, in each column, the fraction who come across as shy. The maths sliver gets a tall shaded band (lots of them seem shy). The wide business column gets a short one (fewer of them do).

Tom is shy, so he’s somewhere in the shaded area. The question is which shaded piece he’s more likely in. And the short-but-wide business band has more area than the tall-but-thin maths one. Not by a little. About twice as much.2

      maths (few)        business (many, ~10x)
     +-----------+   +-----------------------------+
shy  |###########|   |#############################|  <- short band,
     |###########|   +-----------------------------+     but very wide
     |###########|   |                             |
not  |###########|   |         not shy             |
shy  +-----------+   |                             |
                     +-----------------------------+
     tall + thin          short + wide
     = small area         = bigger area  -> Tom is likelier here

Your gut only looks at the heights (“shy… tall maths band… mathematician!”). Bayes makes you look at the areas, and area is width times height. That’s why you multiply. It was never a formula to memorise. It’s just “how much of the rectangle is this.” The column widths are your starting point. The band heights are how well the evidence fits. The shaded area is the answer.

And notice what’s doing the damage: the widths. Ignore how common each type is, mentally set the columns equal, and shyness really would point to maths. The only thing holding the answer down is the base rate. Which brings us to the failure that has its own name.

The 99% test that’s mostly wrong

Back to that medical test. Same rectangle, scarier numbers.

The disease hits about 1 in 1,000 people. The test is 99% accurate: it catches 99% of sick people, and wrongly flags only 1% of healthy people. You test positive. Count it out, no formula needed:

1,000 people, disease in 1 of them
     1 actually sick    -> test flags them        ->   1 true positive
   999 healthy          -> 1% wrongly flagged      -> ~10 false alarms
                                                       ------------------
   11 people test positive, but only 1 is really sick  ->  1 in 11  ~  9%

The test is fine. Nothing is broken. Your gut just grabbed the accuracy (“99%!”) and threw away the base rate (“but it’s rare”). In a thousand people, the one real case gets swamped by ten false alarms, because there are so many more healthy people to draw a stray 1% from.1

This error is common enough to have a name: base rate neglect.4 And once you’re watching for it, you see it everywhere. Somebody snoops around your house and you jump to “burglar,” forgetting that honest people vastly outnumber burglars. A striking coincidence feels like destiny, when coincidences are just common. The vivid, specific detail hijacks your attention and the boring background number, the one that should anchor the whole estimate, quietly falls out of the calculation.

The fix is a reflex: before you weigh the evidence, ask how common the thing was to begin with. Doctors put it well. When you hear hoofbeats, think horses, not zebras.

The one question that kills bad evidence

The base rate is the ingredient people drop. The other ingredient, the lopsidedness of the evidence, is the one they get fooled by. And there’s a single question that fixes it.

Galef tells a story about a colleague, call him Bob, who she suspected was quietly competitive with a coworker, Alice. One day Bob complains about Alice missing a deadline, and Galef catches herself thinking, “See? Proof he’s jealous of her.” Then she stops and asks the question that matters: if Bob weren’t jealous of Alice, would he still be complaining about the missed deadline?

Of course he would. People complain about missed deadlines all the time. It has nothing to do with jealousy. The complaint felt like evidence because it fit the theory… but it fits the opposite theory almost exactly as well. So it should barely move the dial at all.2

This is the whole reason the medical test had two accuracy numbers instead of one. “99% accurate” is meaningless on its own. What matters is the gap between “catches 99% of sick people” and “false-alarms 1% of healthy people.” A clue is only worth something to the extent it’s more likely when you’re right than when you’re wrong. Line up three clues for the same suspicion and the difference is stark:

  • A clue that’s 80% likely if you’re right and 10% if you’re wrong: that’s an eight-to-one hammer. Update hard.
  • A clue that’s 45% likely if you’re right and 35% if you’re wrong: a faint nudge. Barely bother.
  • A clue that’s 40% either way: worthless. It “happened,” it “fits,” and it tells you exactly nothing.

That last case is the important one, because it’s the one we lie to ourselves about. We go looking for things that confirm the theory we already like, find them (there are always some), and feel more sure. But evidence that would show up whether or not you’re right is not evidence. The only honest test is Galef’s: if I were wrong, would I still be seeing this? If the answer is yes, drop it, however good it feels.

You don’t flip, you drift

One more piece before the trap, and it’s the one that keeps you honest in the other direction.

Beliefs rarely flip in a single step, and they shouldn’t. Galef talks about being convinced, when she first moved to Berkeley, that meditation was basically fake. Then a smart, thoughtful friend told her it had genuinely made him happier. Easy to explain that away (placebo, or something else in his life changed). But she had to admit: a smart friend reporting real benefit is at least a little more likely in the world where meditation works than in the world where it doesn’t. Not proof. A nudge. And if the nudges keep coming, all pointing the same way, at some point the honest thing is to have quietly changed your mind without ever having had a single dramatic conversion.2

That’s the rhythm. Snowflakes of evidence, each nearly weightless, accumulating until the branch bends. You don’t wait for one knockout proof, and you don’t ignore the drift either. You keep turning the dial by small amounts, and you let it add up. Which, not coincidentally, is exactly how Bayes’ rule was meant to be used: over and over, each new fact refining the last estimate, never quite certain, always updating. It’s also how your spam filter works, if you want a concrete example. Every word in an email nudges a running probability, and Paul Graham’s 2002 write-up got that idea working well enough to flag mail above 0.9 and miss fewer than 5 spams in 1,000, with zero false positives.5 Not bad for a minister’s abandoned side note.

A quick word on where this came from

The idea is older and stranger than you’d guess. Thomas Bayes, an 18th-century English minister, worked it out and then… didn’t publish it. Apparently he didn’t think it was a big deal. It sat in his papers until after he died in 1761, when his friend Richard Price dug it out, cleaned it up, and read it to the Royal Society in 1763.6

The story goes that Bayes framed it with a thought experiment about a ball thrown onto a flat table: with his back turned, he could keep refining his guess of where it landed by being told whether each new throw landed left or right, in front or behind. Never certain, steadily less wrong. That picture, of an estimate that sharpens as evidence arrives, is the whole thing in miniature.

Funny detail: Bayes never actually wrote down the formula everyone now attaches to his name. The clean, general version came from Pierre-Simon Laplace a decade later, working independently and apparently unaware of Bayes at all. As the historian Stephen Stigler put it, the modern theory really comes from Laplace, but the name belongs to Bayes.7 Bayes gets the naming rights for being first. Laplace did the work. Ah well… that’s science for you.

The trap

Everything so far says humans under-update. We forget base rates, we cling to theories, we mistake comfortable clues for real ones. So here’s the twist that makes Veritasium’s video worth its title: you can also be too good at this.

Think about what updating does when the evidence keeps saying the same thing. You get rejected, again. The pay stays low, again. Some door stays shut, again. Being a good Bayesian, you update: this is just how it is. And you keep updating, run after run, until you’ve converged on near-certainty that this is simply the way the world works and there’s nothing you can do about it.1

Here’s the problem, and it falls right out of the rectangle. Your starting point is a column width. What happens to a column with zero width? Its area is zero. Forever. No matter how tall the evidence band on top of it gets, zero times anything is still zero. Once you’ve concluded some outcome is flatly impossible, you’ve set its width to zero, and Bayes’ rule will now faithfully multiply every future scrap of evidence by that zero and leave your belief exactly where it was. The maths locks the door behind you.

"that's impossible for me"  ->  you stop trying it
     ->  so you generate no evidence either way
     ->  so nothing ever updates
     ->  so it stays impossible   (a very tidy self-fulfilling prophecy)

And this is the real hole in the whole framework: Bayes’ rule tells you how to update a belief, but it says nothing about where your starting point should come from. Two people can look at the identical evidence, one with a 100% prior and one with a 0% prior, and neither will ever budge. As Nate Silver points out, there’s not much use debating someone who’s already at 100% or 0%, because no evidence in the universe can move them.8 The prior is the input Bayes can’t give you.

So what saves you? Not more updating. The thing the formula quietly forgets is that your own actions are part of the evidence stream. If you never try, you never get the data that could prove you wrong. The escape from the trap isn’t cleverer arithmetic, it’s going out and generating the evidence on purpose: doing the thing you’ve decided is hopeless, precisely to see what comes back. A wrong prior of zero is the one bug Bayes can’t fix on its own. You have to feed it data it would never have produced by sitting still. Read properly, a good understanding of Bayes doesn’t say “accept the odds.” It says: experiment.

So what do you actually do

You’re never going to run this as a calculation at a dinner table, and you don’t need to. Stripped down, it’s a short checklist you can actually carry around:

  1. Start from how common it is. Before you weigh anything, ask how likely this was in the first place. The rare thing stays rare unless the evidence is overwhelming. Skip this step and you’re the person who’s 99% sure of a 9% diagnosis.
  2. Ask what you’d see if you were wrong. Then compare it to what you’re actually seeing. If they look about the same, your evidence is weak, no matter how much it flatters your theory. This one question does more work than anything else on the list.
  3. Move in small steps. Let weak evidence nudge you a little and strong evidence move you a lot, and let it accumulate. Don’t demand a knockout, don’t ignore the drift.
  4. Never round yourself to zero or one. The moment a belief hits “impossible” or “certain,” it stops responding to reality. And when you catch yourself stuck near certain about something that matters, don’t collect the same clue again. Go get a different one, by acting differently and watching what happens.

That’s Bayesian thinking, and I’ve come to think of it less as a piece of maths and more as a posture: hold your beliefs as dials rather than switches, move them by amounts the evidence actually justifies, and stay suspicious of any dial you’ve quietly pinned to the end of the scale. To be honest, I’m not sure I always manage it. The pull toward switches is strong, and the trap is comfortable in a way that’s hard to notice from the inside. But the reframe has changed how I argue with myself, which is most of what I wanted from it. A change for the better, and one I’m still working at.

References

Footnotes

  1. Veritasium (Derek Muller), “The Bayesian Trap” (2017). Source of the medical-test example (roughly 9% after one positive test, about 91% after a second independent one), the historical retelling, and the “trap” argument that over-internalising Bayesian updating can become a self-fulfilling prophecy. 2 3 4

  2. Julia Galef, “A Visual Guide to Bayesian Thinking” (2015). The Tom, Bob-and-Alice, and meditation examples, and the rectangle/area way of seeing Bayes, are all from this video. Galef co-founded the Center for Applied Rationality and later wrote The Scout Mindset, whose thesis is essentially principle two: ask what you’d expect to see if you were wrong. 2 3 4

  3. Daniel Kahneman & Amos Tversky, “On the Psychology of Prediction,” Psychological Review 80(4), 1973. The original “Tom W.” study on the representativeness heuristic. The paper’s version contrasts computer science with humanities; Galef adapts it to maths versus business, and the ~10:1 figures are her illustrative numbers, not Kahneman and Tversky’s.

  4. Base rate fallacy. The general name for judging a conditional probability from the specific evidence while ignoring the underlying prevalence; the false-positive paradox in the medical test is a special case.

  5. Paul Graham, “A Plan for Spam” (2002). The canonical engineering use of Bayes: reframing spam filtering as a per-user statistical problem. Graham trained on corpora of roughly 4,000 spam and 4,000 legitimate messages, flagged mail scoring above 0.9, and reported missing fewer than 5 in 1,000 spams with zero false positives.

  6. Thomas Bayes & Richard Price, “An Essay towards solving a Problem in the Doctrine of Chances” (1763). Bayes died in 1761; Price found the manuscript among his papers, edited it, and communicated it to the Royal Society, where it was read on 23 December 1763.

  7. Stephen M. Stigler, “Laplace’s 1774 Memoir on Inverse Probability,” Statistical Science 1(3), 1986. Laplace derived the general form of the rule independently in 1774, apparently unaware of Bayes. Stigler’s histories are the standard reference for the “Laplace did the theory, Bayes got the name” account.

  8. Nate Silver, The Signal and the Noise (Penguin Press, 2012). The point that a debate between someone with a 100% prior and someone with a 0% prior is unwinnable, because no evidence can shift either, is made in his chapter on Bayesian reasoning.