Probability

The envelope that always looks better

Two envelopes, one holding twice as much as the other. Open one, see an amount, and reason that the other holds double or half with equal chance — so switching gains a quarter on average. By symmetry the same argument says switch back. The step that fails is not the arithmetic; it is the claim that double and half are equally likely whatever amount is seen, which no honest prior allows — and there is one prior under which the other envelope really does look better at every amount.
17 min read 6 figures The same thing twiceSmall cases lie

Worth reading first: Evidence measured in decibans · The door that was not opened.

Two envelopes each contain money, one exactly twice as much as the other. One is handed over at random; it is opened, and it holds some amount xx. The offer is to switch.

The argument for switching is short. The other envelope holds either 2x2x or x/2x/2. Each is equally likely, since the envelope in hand was chosen at random. So the other envelope’s expected value is

12⋅2x+12⋅x2=54x,\tfrac12 \cdot 2x + \tfrac12 \cdot \tfrac{x}{2} = \tfrac54 x,

a quarter more than what is in hand. Switch. But the argument never used the value of xx, and it applies equally to the other envelope, so having switched it says switch back, and so on for ever — and before any envelope is opened, it says that each is worth more than the other. Something is wrong, and it is not the arithmetic. Nor is it the symmetry: the two envelopes really are interchangeable before either is opened, and no argument that ends by preferring one of them unopened can be sound. The fault has to lie in a premise, and there is only one premise with a probability in it.

This is the problem the sleeping-beauty essay pointed to as the one where an argument that looks exactly like the others produces a conclusion that cannot be right. In the earlier puzzles, stating the protocol exactly settled the answer. Here the protocol is perfectly clear. What has to be found is the step at which a probability was quietly assumed that no probability distribution can supply.

A prior on the amounts

The phrase “equally likely” hides a question: equally likely under what? Before the envelopes were filled, someone chose the amounts, and whatever process they used is a probability distribution on the possible pairs. The reasoning after seeing xx is Bayes’ theorem: the chance that xx is the smaller amount depends on how likely the pair (x,2x)(x, 2x) was, compared with the pair (x/2,x)(x/2, x).

Take the simplest honest version. The pair is (2n,2n+1)(2^n, 2^{n+1}), with nn chosen uniformly from 0 to 5 — six possible pairs, from (1,2)(1, 2) up to (32,64)(32, 64). Now compute, for each amount that could be seen, what switching is really worth.

Switching envelopes when the largest amount is 64. A bar chart of the probability-weighted gain from switching at each amount that might be seen, 1 to 64: small positive bars and one large negative bar, adding to zero.
Fig. 1 The amounts are 2n2^n and 2n+12^{n+1} with n equally likely from 0 to 5. For each amount that might be seen, the average gain from switching, weighted by how likely that amount is to be seen. At the middle amounts switching gains a quarter of what is in hand, exactly as the naive argument says. At the largest, 64, the other envelope is certainly smaller, and that one loss cancels every gain: the bars add to exactly zero.

For every middle amount — 2, 4, 8, 16, 32 — the two pairs that could have produced it are equally likely, so double and half really are equally likely, and switching really does gain a quarter on average. The naive argument is correct at those amounts. Seeing 1, the other must be 2, and switching gains even more. But seeing 64, the other must be 32: there is no pair (64,128)(64, 128). At that single amount switching loses half of a large sum, and the figure shows that the loss exactly cancels everything gained elsewhere. The weighted gains add to zero, as they must, since keeping and switching are symmetric before anything is seen.

So the naive argument’s error is local and precise. It is true that double and half are equally likely at most amounts. It is false at the largest amount, and the largest amount, rare as it is, carries the most money.

The same question every earlier puzzle asked

The mistake in the naive argument has the same shape as the mistakes in the earlier puzzles, which is worth making explicit. Seeing xx is evidence, and Bayes’ theorem needs two numbers to use it: how likely it was to see xx if the envelope in hand is the smaller, and how likely if it is the larger. The first is the chance that the pair was (x,2x)(x, 2x), the second the chance that it was (x/2,x)(x/2, x). “Equally likely” says that those two numbers are equal — that seeing xx is no evidence at all about which envelope is in hand.

Stated like that, the claim is plainly not free. It is a claim about how the amounts were chosen, and different choices make it true at different amounts and false at others. In the bounded example the evidence is neutral at the middle amounts and decisive at the ends: seeing 1 proves the envelope is the smaller, seeing 64 proves it is the larger. The naive argument treated every amount as a middle amount, which is exactly the error of treating the host’s opened door as if it carried no information about the host.

No distribution makes them always equal

Could a different prior make double and half equally likely at every amount? That would require the pair (x,2x)(x, 2x) to be exactly as likely as (x/2,x)(x/2, x) for every xx — so every pair (2n,2n+1)(2^n, 2^{n+1}) equally likely, for every whole number nn, positive and negative. There are infinitely many such pairs, and equal probabilities on infinitely many outcomes cannot add up to one. No probability distribution does what the naive argument assumes. The argument silently used a “uniform distribution on all powers of two”, which does not exist.

That is one resolution, and it is the one most accounts stop at. It can be seen growing in the bounded example: as the range of pairs is widened, the naive argument becomes correct at more and more amounts, and the whole of the correction is pushed onto the single largest one.

Switching envelopes when the largest amount is 512. A bar chart of the probability-weighted gain from switching at each amount that might be seen, 1 to 512: small positive bars and one large negative bar, adding to zero.
Fig. 2 The same calculation with n equally likely from 0 to 8, so the largest amount is 512. Switching gains a quarter at every amount from 2 to 256, and the gains, weighted by how likely each amount is, grow towards the top; the single loss at 512 is larger still, and the bars again add to exactly zero.

However far the range is extended, the loss at the top is exactly as large as all the gains below it put together, and it is carried by an amount that is seen only once in every 2(N+1)2(N+1) openings. A player who never happens to see the top amount will find switching profitable every time, and the whole account is settled in the rare round when it appears. It is not quite the end, because there is a genuine distribution under which the conclusion of the argument survives.

A prior where switching always looks better

In 1995 John Broome gave a prior under which the other envelope is worth more on average whatever amount is seen. The pair is again (2n,2n+1)(2^n, 2^{n+1}), but with nn chosen with probability 13(23)n\tfrac13 \left(\tfrac23\right)^n for n=0,1,2,…n = 0, 1, 2, \dots — a perfectly good distribution, whose probabilities add to one.

Broome's envelopes: the other one always looks better. Bars of the expected amount in the other envelope divided by the amount seen, for amounts 1 to 512 under Broome's prior: 2 at 1, and 1.1 everywhere else, all above the line at 1.
Fig. 3 Broome’s prior. Seeing 1, the other envelope surely holds 2. Seeing any other amount, it holds double with chance 2/5 and half with chance 3/5, and is worth 1.1 times the amount seen on average. Every bar is above the break-even line: whatever is seen, switching looks better — and by symmetry, so does switching back.

Seeing an amount 2k2^k with k≥1k \ge 1, the two pairs that could have produced it are (2k−1,2k)(2^{k-1}, 2^k), of probability proportional to (2/3)k−1(2/3)^{k-1}, and (2k,2k+1)(2^k, 2^{k+1}), proportional to (2/3)k(2/3)^k. The ratio is 3:23 : 2, so the other envelope is half with probability 3/53/5 and double with probability 2/52/5 — no longer equal, and weighted towards half. And still

35⋅x2+25⋅2x=1110 x.\tfrac35 \cdot \tfrac{x}{2} + \tfrac25 \cdot 2x = \tfrac{11}{10}\,x.

Double is less likely than half, but it is worth four times as much, and the prior does not fall off fast enough to compensate. The likelihood ratio the opened envelope supplies — in the language of odds, the evidence for “this is the smaller” — is the same at every amount, 2:32 : 3, a steady −1.8-1.8 decibans. Seeing the amount always makes “smaller” a little less likely, and never enough to make switching a bad bet. So under this prior, for every single amount the envelope might hold, switching is worth ten per cent more on average — and before anything is opened the two envelopes are symmetric. The paradox has come back with an honest distribution.

Where the expected value went

The escape is that under Broome’s prior the expected amount in an envelope is infinite.

The average amount, summed under two priors. Partial sums of the expected amount on a logarithmic axis for 20 pairs: under Broome's prior they grow geometrically, under a thinner prior they converge to 1.500.
Fig. 4 The expected amount in an envelope, added up pair by pair. Under Broome’s prior the n-th pair contributes (1/3)(4/3)n(1/3)(4/3)^n times 3/2, and the partial sums (orange) grow by a third at each step, past 470 after twenty pairs and without bound. Under a prior with probabilities (1/2)(1/4)n(1/2)(1/4)^n the same sums (blue) settle at 1.5. Only a prior with a finite average gives switching something finite to be compared against.

Each pair contributes its probability times its average amount: 13(23)n×32⋅2n=12(43)n\tfrac13 (\tfrac23)^n \times \tfrac32 \cdot 2^n = \tfrac12 (\tfrac43)^n, and a geometric series of ratio 4/34/3 diverges — the square that half, a quarter, an eighth fit into has no counterpart when each term is a third bigger than the last, and the comparison test settles it at once. The statement “the other envelope is worth 1.1 times this one, on average, whatever this one holds” then compares two infinite quantities. It is true amount by amount, and adding it up over all amounts gives ∞=1.1×∞\infty = 1.1 \times \infty, which says nothing.

The step that fails is the one that turns “better given every xx” into “better overall”. For finite expectations that step is the law of total expectation and is always valid. With infinite ones it is not, for much the same reason that rearranging a series that does not converge absolutely can give any answer: the order of summation, here “first condition on xx, then average”, changes the result when the total is infinite.

An average that never arrives

Simulation makes the infinite expectation concrete.

Running averages that never settle. Two erratic curves of average winnings per round against rounds played, keeping and switching, under Broome's prior; both keep jumping upward and neither converges.
Fig. 5 A hundred thousand rounds under Broome’s prior, and the running average winnings of always keeping the first envelope (blue) and always switching (orange), on a logarithmic axis of rounds. Neither settles. Every so often a very large pair turns up and lifts both averages at once, and whichever strategy happened to hold the larger envelope that round pulls ahead — until the next large pair.

The running averages behave like those of a distribution with no mean: long quiet stretches, then a jump when a pair far out in the tail appears, then a slow decay back until the next one. In this run the switching strategy happens to be ahead — a matter of which envelope was handed over in the few enormous rounds — and with a different seed the order is as likely to be reversed. There is no number that either average converges to, so there is no sense in which switching is “worth 10% more” in the long run. The conditional expectations are finite and the unconditional one is not, and the paradox lives in that gap.

The same behaviour appears in the older St Petersburg game, which pays 2n2^n with probability 2−n2^{-n} and has infinite expected value: simulated averages of its payouts climb in jumps for ever, roughly like the logarithm of the number of plays, and never settle on a price. Broome’s envelopes are a St Petersburg game played twice with a symmetry between the two plays, and the symmetry is what turns an unbounded average into a contradiction.

A rule that really does beat a coin toss

There is a surprising coda. Although no rule gains money on average — the symmetry forbids it — there are rules that end with the larger envelope more often than half the time, for every pair of amounts.

A random threshold that beats a coin toss. Bars for pairs (2ⁿ, 2ⁿ⁺¹), n = 0 to 12, of the probability of ending with the larger envelope under a random-threshold switching rule, all above one half.
Fig. 6 Pick a random threshold, here lying between 2m2^m and 2m+12^{m+1} with probability 0.3·0.7ᵐ, and switch exactly when the amount seen is below it. For each possible pair, the chance of ending with the larger envelope, computed exactly: every bar is above one half, by half the chance that the threshold falls between the two amounts — 65% for the smallest pair, falling towards 50% for large ones.

The rule is Thomas Cover’s, and it has the air of a trick until the reason is seen. Choose a threshold at random, independently of the envelopes, from any distribution that puts some probability between every pair of possible amounts. Switch if the amount seen is below the threshold, keep it otherwise. If the threshold falls below both amounts, the rule always keeps, and wins half the time; if above both, it always switches, and wins half the time. But if the threshold falls between the two amounts, the rule keeps the larger and switches away from the smaller — it wins for certain. So the chance of ending with the larger envelope is one half plus half the chance that the threshold lands between them, which is strictly more than one half for every pair.

This does not contradict the symmetry. The rule gains in probability of holding the larger envelope, not in expected money, and its edge is largest for small amounts, where the threshold is most likely to fall between them, and vanishes for large ones. It uses the amount seen as information, which the naive argument claimed to do and did not.

Thresholds are the natural shape for such rules, for the same reason they are in the problem of when to stop looking: a single number seen once can only be compared with something, and the best a rule that knows nothing else can do is compare it with a cut-off. There the cut-off is set by how many candidates have been seen; here it is drawn at random, precisely because nothing is known about the scale of the amounts, and a random cut-off is the only kind guaranteed to fall between the two amounts with some positive chance, whatever they are.

What the figures leave out

Every prior drawn is a choice. The bounded prior, Broome’s prior and the thin-tailed prior are three examples. The paradox is about what can go wrong in general, and the figures show it going wrong in the two standard ways — a boundary that the naive argument ignores, and an infinite expectation that makes conditional comparisons meaningless — without claiming these are the only ones.

The simulation is one run. A hundred thousand rounds under a distribution with infinite mean shows the averages failing to settle, and a different seed would show different jumps at different times. What does not change from seed to seed is the failure to settle, which is a theorem rather than an observation.

Only powers of two are drawn. Every prior here puts the amounts at 2n2^n and 2n+12^{n+1}, which keeps the arithmetic exact. With a continuous prior on the smaller amount the same analysis goes through with densities in place of probabilities: switching gains wherever the density falls off slowly enough and loses where it falls fast, and the two balance exactly whenever the expected amount is finite.

The threshold rule’s edge is exact but small. For large amounts it approaches a half, and for any particular pair it depends on the threshold distribution chosen. The figure computes it exactly for one choice.

Still open: how to reason when expectations diverge

The mathematics of the problem is settled: with a finite expected amount, switching is neutral once averaged over what might be seen; with an infinite one, conditional expectations can all favour switching and mean nothing. What is not settled is how a decision-maker should behave when faced with a prior like Broome’s — whether to accept that expected value is the wrong guide when it is infinite, and if so what should replace it. The same question is raised by the St Petersburg game, which pays 2n2^n with probability 2−n2^{-n} and has infinite expected value yet is worth only a few units to anyone asked to pay for it. Proposals include bounded utility, which makes every expectation finite; ignoring outcomes of sufficiently small probability; and comparisons that look at the difference between the two envelopes rather than their separate values, which is finite even when each value is not. None has become the agreed answer, and philosophers continue to publish variants of the envelopes designed to defeat each proposal in turn.

There is a mathematical version of the question too: which ways of comparing random amounts with infinite means are consistent — agree with expectation whenever it is finite, respect symmetry, and do not contradict themselves on examples like Broome’s. Several such extensions of expected value have been proposed and shown to fail on cleverly built distributions, and whether any reasonable one exists is not known.

An argument that was true at every amount but one

The naive argument’s arithmetic is right. Its premise — double and half equally likely at every amount — is true at almost every amount under any reasonable prior, and false at exactly the amounts that carry the most weight. Under an unbounded prior with a finite average, those amounts are rare and far out, and the loss there balances the gains; under Broome’s prior they never come, and the balance is paid instead by an infinite expectation. Either way the resolution is the same as in every earlier puzzle of this kind: write down the prior, compute what the evidence of the opened envelope actually says, and the paradox becomes a bookkeeping error with an address.

The address differs from puzzle to puzzle. In the door that was not opened it was the host’s rule for choosing a door; in the sleeping-beauty problem it was which events were being counted; here it is the prior on the amounts, and specifically its tail — the rare, large pairs that the naive argument never pictured. In every case the fix is to ask what the evidence would have looked like under each hypothesis. The envelopes add one lesson the others did not need: when the amounts at stake have no average, no answer phrased as an average can be trusted, however correct each of its steps.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A dashed tag is an object no other essay names yet.

Bayes' theoremConditional probabilityExpectationRandom walkSymmetry