Probability

The door that was not opened

Three doors, one prize, a host who opens a losing door and offers a swap. Switching wins two times in three, and the reason is not about doors — it is about what the host was allowed to do.
16 min read 6 figures Small cases lie

Worth reading first: Bayes' theorem is a picture of a square.

Three doors. A prize behind one. A contestant picks a door, the host — who knows where the prize is — opens one of the other two to reveal no prize, and offers a swap.

Switching wins two times in three.

Three doors, as areasStaying wins 33.3% of the time and switching wins 66.7%, because the host's choice is constrained by what the host can see, so opening a door rules a region out without moving any boundary.the first pick was right — staying winsthe prize is behind door 2 — switching winsthe prize is behind door 3 — switching winsthe host knowsstay: 33.3%switch: 66.7%3 equally likely worlds,3 of them surviveno boundary moved:a region was ruled out
Fig. 1 The unit square split by where the prize is, before the host does anything. One band in three is the case where the first pick was already right; the other two are the cases where switching wins. The generator enumerates the equally likely worlds and checks both rates against the arithmetic.

This is the most argued-about elementary probability problem there is, and the arguing is instructive: the correct answer was published in 1990 and drew ten thousand letters of objection, about a thousand of them from people with doctorates.

The picture is a square that was drawn before the host moved

The whole content is in the order of events.

The prize’s location was fixed before anything happened, and the contestant’s first pick was made with no information. So at the moment of picking, the square divides into three equal bands: the prize is behind the chosen door, or behind one of the other two.

That division is a fact about the setup. Nothing the host does afterwards can change where the prize is, and nothing the host does can change the chance that the first pick was right — that number was determined when the pick was made, and it is one in three.

So after the host opens a door, the chance that staying wins is still one in three, and therefore the chance that switching wins is two in three. There is no calculation. The argument is that a probability fixed before the evidence arrived cannot be revised by evidence that could not have distinguished the cases.

Why the host’s constraint is the whole thing

The step everyone skips is the one that matters: the host was going to open a losing door no matter what.

That is the hypothesis, and it means the host’s action carries no information about the contestant’s own door. Whatever was behind it, a losing door among the other two existed and was opened. The event “the host opened a losing door” has probability one, and an event of probability one cannot update anything.

What the action does do is concentrate the other two-thirds. Before, the two unchosen doors held two-thirds between them, split evenly. After, one of them is known to be empty, so the whole two-thirds sits on the remaining one.

Put differently: the host has ruled out a region without redrawing any boundary. The one-third band is untouched; the two-thirds region has been relabelled as belonging entirely to a single door.

This is exactly the structure of Bayes’ theorem drawn as a square, which is the rung below. There, a test result cut the square and the posterior was one shaded area over the total shaded area. Here the host’s action cuts nothing — that is the unusual feature, and it is why the answer feels wrong. Readers expect evidence to move a boundary, and this evidence does not.

Bayes' theorem as two rectanglesA unit square split by how common the condition is (1.0%) and then by how the test behaves. Of everyone who tests positive, the fraction who have it is 16.7%.has itdoes nottests positivetrue, and positive: 0.99%false, and positive: 4.95%so of the positives,16.7% really have it
Fig. 2 The rung below, for comparison. A medical test splits the square twice — once by how common the condition is and once by how the test behaves — and the answer is a ratio of two shaded regions. That is what an update normally looks like, and it is what a reader arrives at the door problem expecting.

The contrast is the point. In the test figure the evidence genuinely partitions the population, and both cells that survive have positive area, so the posterior is a genuine ratio. At the doors the evidence partitions nothing, and the entire arithmetic is the observation that one band was never touched.

The same story with a different host

Now change one thing. The host does not know where the prize is, opens one of the other two doors at random, and it happens to be empty.

The same story, told by a host who does not knowStaying wins 50.0% of the time and switching wins 50.0%, because the host's choice is unconstrained, and the two worlds in which the prize was revealed have been thrown away.stay winsstay winsprize revealedswitch winsswitch winsprize revealedthe host does notstay: 50.0%switch: 50.0%6 equally likely worlds,4 of them survivetwo worlds were discarded,and that is what moved the answer
Fig. 3 The six equally likely worlds when the host chooses blindly: three prize positions, two doors that can be opened. Two of them reveal the prize and are struck out. Of the four that survive, two favour staying and two favour switching.

Now switching wins half the time.

Everything the contestant sees is identical. Same doors, same open door, same empty room behind it, same offer. And the answer changed, because what changed is not the evidence but the process that could have generated it.

The reason is visible in the figure. When the host is ignorant, some of the worlds in which the first pick was wrong end with the prize being revealed, and those worlds are removed by the conditioning. Removing worlds in which the first pick was wrong raises the chance that the first pick was right, from one-third to one-half. The informed host removes no such worlds, because he never reveals the prize in any of them.

That comparison is the actual lesson of the problem, and it is much more useful than the answer. A probability depends on the mechanism that produced the observation, not only on the observation. Two situations that look identical from inside can have different answers, and no amount of staring at the evidence recovers the difference — it has to be known.

Where the intuition goes wrong

The universal wrong answer is that the two remaining doors are equally likely, so switching does not matter.

The error has a precise name: it treats the two survivors as symmetric when the process that produced them was not. One door was chosen by the contestant, at random, from three. The other survived a filter applied by someone who knew the answer and was obliged to avoid the prize. Those doors did not arrive at the final round by the same route, and there is no reason for them to be equally likely.

The instinct being violated is that ignorance implies uniformity — that when two options remain and nothing more is known, they must be even. That instinct is right when the two options are genuinely interchangeable and wrong the moment they have different histories. It is the same error as reading a positive test result as a probability of illness without the base rate, and it is the same error as being surprised by a shared birthday because twenty-three seems small next to three hundred and sixty-five — in each case an intuition about symmetry is applied to a situation that is not symmetric.

Making it obvious

The argument that convinces almost everybody is to change the number of doors.

The same game, with more doorsWith d doors the host opens all but two, and switching wins (d − 1)/d of the time. Three doors is the smallest game in which the effect exists at all, which is why it is the one that looks like a paradox.5101520253000.20.40.60.81doorschance of winning3 doors: 67%10 doors: 90%30 doors: 97%staying
Fig. 4 Switching’s chance against the number of doors, where the host opens all but two. At three doors it is two-thirds; at thirty it is twenty-nine thirtieths. Three is the smallest game in which the effect exists at all, which is why it is the one that looks like a paradox.

With a hundred doors, the contestant picks one — a one per cent chance of being right — and the host, who knows, opens ninety-eight empty doors, leaving the contestant’s original pick and one other.

Almost nobody hesitates at that version. The original pick was a wild guess with a one in a hundred chance; the surviving door has been selected by someone who knew, from ninety-nine candidates. Switching wins ninety-nine times in a hundred.

The same game, with more doorsWith d doors the host opens all but two, and switching wins (d − 1)/d of the time. Three doors is the smallest game in which the effect exists at all, which is why it is the one that looks like a paradox.468101200.20.40.60.81doorschance of winning3 doors: 67%12 doors: 92%staying
Fig. 5 The same curve over the range where the argument is actually had. Between three and twelve doors the two lines separate quickly, and by ten the case is no longer disputed by anyone — the disagreement is confined to the leftmost point of a curve that nobody disputes anywhere else.

Nothing structural changed between the two versions. The three-door case is the same argument with the effect at its smallest — a factor of two rather than ninety-nine — and small effects are where intuition is least reliable. That is a theme this collection returns to: the smallest instance of a phenomenon is the one that is remembered, quoted and argued about, and it is systematically the worst one to reason from. That is a general property worth carrying: an argument that is disputed at its smallest instance is often not disputed at all once the instance is enlarged, and enlarging it is a cheaper move than arguing.

What it costs to be sure

The problem is small enough to settle by enumeration, and it was — repeatedly, because the argument would not stop.

There are three equally likely prize positions and the contestant’s choice can be fixed at door one without loss of generality, so there are three cases. Staying wins in one. That is the whole computation, and it takes a table with three rows.

It was also settled by simulation, famously and en masse, after the 1990 dispute. Simulation is the right tool for a claim about a process because it forces the process to be written down: a great many of the objectors, asked to code it, discovered that they could not write the host’s behaviour without deciding whether the host knew — which is the question. That is the same virtue a generator has on this site, where a figure that must state its parameters before it can be drawn cannot smuggle in an assumption the prose left vague.

That is the practical value of a simulation over an argument here. An argument can be conducted with the ambiguity intact; a program cannot. Whoever writes if host_knows has already found the crux.

The cost of a simulation is the usual one: the estimate’s error falls as 1/N1/\sqrt{N}, so distinguishing two-thirds from one-half needs a few hundred runs and distinguishing two-thirds from 0.660.66 needs a few hundred thousand. That is fine here and it is a bad way to settle anything whose answer is close to another answer, which is why the enumeration is better in every respect except persuasiveness. It is the same square-root cost that governs every estimate made by sampling, and the same reason a simulation is a poor instrument for a precise number and an excellent one for a disputed qualitative claim.

The 1990 dispute is worth one more sentence because of what settled it. Not the argument, which had been made correctly and rejected; not the enumeration, which fits in three rows. What settled it was several hundred school classes running the experiment with paper cups and reporting the results — a physical simulation, run by people with no stake in the answer, at a scale where the two-thirds is unmistakable. The correct reasoning had been available throughout and was not persuasive; a count of outcomes was.

There is also a counting route that avoids both, and it is the shortest argument in the essay. The three prize positions are the three cases; staying wins in exactly the one where the first pick was right; so staying wins one time in three and switching wins the rest. That is a counting argument of the kind that settles a question without computing anything, and its brevity is a fair measure of how much of the difficulty is in the statement rather than in the mathematics.

The same test, at every base rateA test with 99% sensitivity and 95% specificity, applied to populations in which the condition is more or less common. The chance that a positive result is real is a property of the population as much as of the test.00.20.40.60.8100.20.40.60.81how common the condition ischance a positive is real0.1% → 2%1% → 17%10% → 69%50% → 95%
Fig. 6 The rung below again, showing what a genuine dependence on a prior looks like: the same test at every base rate, with the answer sweeping from nothing to everything. The door problem has no such curve, because its prior is fixed at one-third by the rules of the game and nothing in the game varies it.

That absence is worth naming, because it explains why the problem is a puzzle rather than a technique. The base-rate figure is useful precisely because the base rate varies between situations, so a reader has to learn to ask what it is. The door problem’s prior is nailed down by the format, so there is nothing to ask and nothing to estimate — which leaves only the protocol, and the protocol is the part the telling omits.

Where the statement needs to be nailed down

The problem as usually told is genuinely ambiguous, and the objectors were not simply wrong — they were answering a different question, one the statement permitted.

Three assumptions are needed and none is always stated:

The host always opens a door. If the host may decline — say, only offering the swap when the contestant has picked correctly — then switching is a disaster. Nothing in the usual telling rules this out, and this is the reading under which the objectors’ answer is closer to right. The actual game show, incidentally, worked this way some of the time: the host was under no obligation, which means the televised game and the problem named after it are not the same game.

The host always reveals a losing door. This is the ignorance variant above, and it changes the answer to a half.

The host chooses uniformly when both remaining doors are empty. If the host has a bias — always opening the leftmost when free to choose — then which door was opened carries extra information, and the probability of winning by switching becomes either 1/21/2 or 11 depending on the case, averaging to 2/32/3. The average is unchanged and the individual answers are not, which is a small and genuinely useful warning: a correct average can conceal that no individual case has that value, in the same way that the most even distribution is the one that forces the conclusion and the average distribution says nothing.

Once all three are stated the answer is two-thirds and there is nothing to argue about. Before they are stated, the argument is about the problem rather than about probability, and a great deal of the heat came from participants who had each silently supplied a different one.

The general lesson is worth more than the specific answer: in a conditional probability problem, the protocol is part of the data. A description of what was observed is not enough; what is needed is a description of what would have been observed in the cases that did not happen.

What the picture cannot show

The first figure shows the square before the host acts, and the claim is about what happens after. The reason the picture is drawn that way is the argument — the boundary does not move — but a reader who wants to see the update sees nothing happen, which is unsatisfying and is the honest state of affairs.

The second figure has the opposite problem. It shows the six worlds and strikes two out, which makes the conditioning visible, and that visibility is available only because the ignorant host’s version has something to strike. The two figures are drawn differently on purpose, and the difference between them is the content; a single format for both would have made one of them a lie.

Neither figure can show the protocol, which is the thing the whole essay says is decisive. A protocol is a description of counterfactuals — what the host would have done in cases that did not occur — and there is nothing in a picture of what happened that can carry it.

The ladder from here

Rungs on this anchor: the base-rate square, below this one. The boy-or-girl paradox, where the same ambiguity appears with no host at all and is harder to resolve. The three prisoners problem, which is this one in different clothes and predates it by decades. Simpson’s paradox, where aggregating groups reverses a comparison. The sleeping beauty problem, where the protocol ambiguity has never been settled. Likelihood ratios as the general form of what the host’s constraint does. And the two-envelope problem, where a naive argument gives an answer that cannot be right and locating the error is genuinely subtle.

The observation was not the evidence

The single sentence worth taking away is that the contestant’s information is not what the contestant saw.

An open door with nothing behind it is the observation. The evidence is that a host who knew the answer and was obliged to open an empty door did so — and that is a statement about rules rather than about doors. It is the same distinction Bayes’ square makes between a test result and a test’s characteristics: the result is what arrives, and the characteristics are what make it mean anything. Change the rules and the same open door means something else. Remove the rules from the telling and the problem has no answer.

Probability is often introduced as a way of reasoning about uncertain events, and this problem is a good argument that it is really a way of reasoning about uncertain processes. The event here is not in dispute by anybody. All the disagreement, and the entire factor of two, lives in what the process was allowed to do.

That reading also explains why the problem has outlived its game show by decades. As a puzzle about doors it is small and settled. As a demonstration that an observation’s meaning is fixed by the rule that generated it, it is the shortest available statement of something that goes wrong constantly and expensively elsewhere: in survey design, where who was asked determines what an answer means; in medical testing, where the population screened determines what a positive result implies; and in any analysis of data that was collected for some other purpose, where the collection rule is usually unrecorded and always decisive.

The doors are a memorable wrapper on the smallest possible instance of that, which is exactly what a good elementary problem should be.