Probability

One coin, counted by runs and by wakings

Beauty is put to sleep and a fair coin is tossed. Heads, she is woken once; tails, twice, with the first waking erased from her memory. Each time she wakes she is asked how likely heads is. One half, say some; one third, say others; and unlike every earlier puzzle of this kind, stating the protocol exactly does not end the argument.

Worth reading first: Two children and the sentence about one of them · The door that was not opened.

On Sunday evening Beauty is put to sleep, and a fair coin is tossed. If it lands heads, she is woken on Monday, interviewed, and put back to sleep until the experiment ends. If it lands tails, she is woken on Monday, interviewed, given a drug that erases the memory of that waking, and woken again on Tuesday and interviewed again. She knows all of this in advance. At each interview she is asked one question: how likely is it that the coin landed heads?

Nothing she experiences on waking distinguishes the three possible wakings — Monday after heads, Monday after tails, Tuesday after tails. She learns nothing she did not know on Sunday, except that she is awake. On Sunday the answer was one half.

One experiment, measured by runs and by awakenings. Two unit squares for the same coin and the same schedule: the same experiment weighed two ways: by runs, heads keeps half the square; by awakenings, heads is one of 3 equal slices.
Fig. 1 The same experiment drawn twice. On the left the square is weighted by runs of the experiment: heads takes half, and the tails half is shared between its two wakings. On the right it is weighted by wakings: each of the three possible wakings takes an equal third. Heads is one half of the first square and one third of the second.

The problem was circulated by Arnold Zuboff, named by Robert Stalnaker, and brought into print by Adam Elga in 2000, who argued for one third. David Lewis replied the next year arguing for one half. Two decades and a very large literature later, both answers have careful defenders, and a third position has joined them.

The two earlier essays in this series found that a conditional probability depends on the process that produced the evidence — a host’s rules, a parent’s habits of speech — and that stating the process settles the answer. Here the process is stated completely, in the first paragraph. What remains undetermined is something the earlier puzzles never had to decide: what, exactly, is being counted.

Two squares for one experiment

Both answers can be drawn, and drawing them shows that they are not computations that disagree but measures that differ.

The halfer’s square is weighted by runs of the experiment. Each run has one coin toss, half the runs are heads and half tails, and a heads run contains one waking while a tails run contains two. Waking tells Beauty nothing about the coin, because she was certain to wake at least once whichever way it fell. The event Beauty is awake has probability one under both hypotheses, and — exactly as with the informed host’s open door — an event certain under every hypothesis cannot move any of them. So heads keeps its half, and the tails half is shared between Monday and Tuesday, a quarter each.

The thirder’s square is weighted by wakings. Over many runs there are, on average, one and a half wakings per run: one from heads runs and two from tails runs. A third of all wakings follow heads. Beauty, on waking, is at one of those wakings and cannot tell which, so she should assign each an equal share, and heads gets one third.

The two squares have the same cells and the same labels. They differ in the areas given to the cells, and the areas are what a probability is. The halfer divides the square by the coin first and then, within tails, by day. The thirder divides it by waking and gives each equal weight. Both are internally consistent. The argument is over which one is Beauty’s credence when she wakes.

This is where the puzzle departs from the two-child problem. There, the ambiguity was in the English: the sentence did not say how it was produced, and supplying a process removed it. Here there is no missing process. There is a missing answer to the question of what a credence is a proportion of.

The same coins, counted twice

The difference between the two measures is not hidden in a subtle model. It shows up the moment the experiment is run many times and the results are tallied.

The same coins, counted by run and by awakening. Two running shares from one simulated sequence of experiments: 1500 runs of the experiment: 766 came up heads, a share of 0.511; they produced 2234 awakenings, of which 766 were heads, a share of 0.343.
Fig. 2 Fifteen hundred simulated runs of the experiment, all from one sequence of fair coins. The upper curve is the running share of runs that were heads, which settles near one half. The lower curve is the running share of wakings that followed heads, which settles near one third. The coins are the same coins; only the denominator differs.

In the simulation — an estimate from random trials, as every such tally is — roughly half the runs are heads, as a fair coin requires. Those runs produce one waking each, and the tails runs two each, so about a third of all wakings follow heads. Both curves settle, as the law of large numbers guarantees any running share of independent trials must, and they settle in different places because they are averages over different things.

Nobody disputes either curve. The halfer agrees that a third of wakings follow heads, and the thirder agrees that half of runs are heads. What they dispute is which of these frequencies Beauty’s credence at a waking should match. The thirder says she is at a waking, so her credence should be calibrated to wakings: if she answered “one third” at every waking of a long series, she would be right about the frequency of heads among the occasions on which she was asked. The halfer says the coin was tossed once per run, her evidence on waking is the evidence she had on Sunday, and her credence in a fair coin’s outcome should be the chance of that outcome.

The two curves are also a small instance of a familiar technique. Converting a share of runs into a share of wakings means giving every run a weight equal to the number of wakings it contains — one for heads, two for tails — and averaging with those weights. Reweighting one set of samples to estimate an average over a different population is exactly what importance sampling does, and the weights here are the ratio between the two populations. The thirder’s measure is the halfer’s measure reweighted by wakings, and the halfer’s is the thirder’s reweighted by one over wakings; each can be computed from the other’s data, which is why the data cannot choose between them.

The simulation also shows why neither side can be refuted by experiment. Every observable frequency in the experiment is agreed; the curves are what they are whichever answer is right. The question is which curve is the answer to how likely is it that the coin landed heads? — and that is a question about the meaning of the words, not about coins.

What each answer is worth at the betting table

One way to give a credence consequences is to ask what bets it makes fair. Suppose that at each waking Beauty may buy a ticket that pays 1 if the coin landed heads. What price should she pay?

Two ways to sell a bet on heads. Expected profit per experiment of buying the ticket at each price: a ticket that pays 1 if the coin fell heads, sold once per experiment or once at every awakening; the first breaks even at 1/2 and the second at 1/3.
Fig. 3 Expected profit per run of the experiment from buying a ticket that pays 1 on heads, against its price. If a ticket is sold once per run, the profit line crosses zero at a price of one half. If a ticket is sold at every waking, a tails run sells two losing tickets, and the profit crosses zero at one third.

The answer depends on the betting scheme, and the dependence is exact. If the ticket is offered at every waking and Beauty buys at every waking, then a heads run earns one ticket’s payout and a tails run pays for two losing tickets. The expected profit per run at price c is ½(1 − c) − ½·2c, which is zero at c = 1/3 — the price at which the bet is a fair game, neither side gaining on average. A thirder’s credence prices the ticket correctly under that scheme, and a halfer who pays one half at each waking loses on average a quarter per run.

If instead only one ticket is sold per run — say, only the purchase made on the last waking counts — the expected profit is ½ − c, zero at one half, and it is the halfer’s credence that prices the bet fairly.

Christopher Hitchcock argued in 2004 that the per-waking scheme is the natural one, since Beauty cannot know at a waking whether it is her only chance to bet, and that a halfer is therefore exposed to a sure loss. Halfers reply that the loss comes from the scheme counting tails twice, not from any error in credence — that a bet placed twice on the same coin is a bet on tails at double stakes, and pricing it by the coin’s chance is correct for the coin and wrong only for the doubled stake. Each reply is coherent. What the figure establishes is the thing both agree on: betting consequences do not decide between the credences, because each credence is exactly the fair price under one scheme.

Told that it is Monday

The sharpest disagreement arises when Beauty learns something. Suppose that after she gives her answer, she is told: today is Monday.

Sleeping Beauty, told that it is Monday. Two unit squares for the same coin and the same schedule: told that it is Monday, every later awakening is crossed out; the per-run square then gives heads 2/3 and the per-awakening square gives 1/2.
Fig. 4 Both squares after Beauty learns that it is Monday, with the tails-Tuesday waking crossed out. On the left, what survives is the heads half and the tails-Monday quarter, so heads has two thirds of what is left. On the right, heads-Monday and tails-Monday are equal thirds, and heads has one half.

For the thirder this is straightforward. Of three equally likely wakings, the Tuesday one is ruled out, leaving heads-Monday and tails-Monday, still equal, so heads is one half. And that is exactly what the thirder should want, because on Monday the two branches are indistinguishable up to that point — indeed Elga’s argument notes that the coin could be tossed on Monday night, after the interview, in which case on Monday it has not yet been tossed and its chance of heads is one half. Elga then ran the argument backwards: if heads-Monday and tails-Monday must be equal, and tails-Monday and tails-Tuesday must be equal because Beauty cannot tell them apart, then all three are equal, and heads before the news is one third.

For the halfer the news is awkward. Starting from heads one half and tails-Monday one quarter, crossing out tails-Tuesday leaves heads with two thirds of the surviving area. Lewis accepted this: learning that it is Monday should make Beauty believe, at two to one, that a fair coin — possibly not yet tossed — will land or has landed heads. That consequence is widely regarded as the most serious cost of the halfer position, and it is why a third camp, the double halfers, hold one half both before and after the news, at the price of giving up ordinary conditioning on the information that it is Monday.

The picture shows why no one can have both. Starting from one half, the only way to end at one half after crossing out a region that belongs entirely to tails is to cross out nothing — to refuse the update. Starting from a third, the update lands at one half automatically. The three camps are the three ways of trading one intuition against another, and the square makes the trade exact.

Many wakings on tails

Nick Bostrom’s extreme version sharpens the thirder’s side. Change the protocol so that tails brings not two wakings but many — five, or a million — with memory erased after each.

Sleeping Beauty with 5 awakenings on tails. Two unit squares for the same coin and the same schedule: the same experiment with 5 tails awakenings weighed two ways: by runs, heads keeps half the square; by awakenings, heads is one of 6 equal slices.
Fig. 5 The same two squares with five wakings on tails. By runs, heads still holds half the square and the tails half is sliced five ways. By wakings, heads is one slice of six. The halfer’s credence on waking is unchanged at one half; the thirder’s has fallen to one sixth.

The thirder’s answer falls as 1/(n + 1): with a million wakings on tails, a waking Beauty should be almost certain the coin landed tails, since almost all wakings are tails wakings. The halfer’s answer stays at one half however large n is, because the coin was still fair and waking was still certain.

More awakenings on tails. Credence in heads at each number of tails awakenings under the two accounts: with n awakenings on tails, counting awakenings gives heads 1/(n + 1); counting runs gives 1/2 before any news and n/(n + 1) once told it is the first day.
Fig. 6 Credence in heads against the number of wakings on tails. Weighted by wakings it falls as 1/(n + 1). Weighted by runs it stays at one half on waking, and rises to n/(n + 1) once Beauty learns that this is the first waking. At one waking all three agree at one half.

The same figure shows the halfer’s cost growing just as fast. A halfer told this is the first waking must update to n/(n + 1), which for a million wakings is near certainty that the coin landed heads — on the basis of learning which day it is, in an experiment where the first waking happens under either outcome. Each extreme makes one side’s position look untenable, and the curves are mirror images of each other about one half: the thirder’s credence before the news and the halfer’s after it add to exactly one.

The extreme cases have been used to argue both ways, which is itself informative. Bostrom used a closely related structure, in which a hypothesis producing many more observers is favoured by the thirder’s reasoning, to show that the thirder’s principle has startling consequences in cosmology — the presumptuous philosopher, who is certain on philosophical grounds that the universe is vast. So each side has an extreme case in which its reasoning seems absurd, and both are the same figure read from opposite ends.

Why the square has two weightings

Every earlier problem in this series had a single natural sample space. Families were the objects in the two-child problem; worlds of prize and host choice were the objects at the doors. A probability was a share of those objects, and the only question was which of them the evidence left standing.

Sleeping Beauty has two kinds of object and no rule for choosing between them. A run is an outcome of the coin together with the schedule it causes. A waking is a moment at which Beauty exists, is awake and is asking the question. Ordinary probability has never had to distinguish these, because in ordinary situations each run contains exactly one moment at which the question is asked. The problem is constructed so that the number of askings depends on the outcome being asked about, and that is precisely what makes the two weightings disagree.

So the puzzle is not about Bayes’ theorem. Both sides apply it correctly to their own square, and the theorem drawn as areas gives each side the answer it expects. The puzzle is about self-locating belief — belief not about what the world is like but about where in it one is — and whether such belief should be measured over worlds or over places in them. That question has no analogue in coin tossing, and the coin was only ever the simplest possible thing to be uncertain about while being uncertain about the day.

Its consequences reach well beyond the laboratory. Arguments about how to reason as one observer among many, whether in cosmology, in estimates of how long humanity will last, or in the fine-tuning of physical constants, all require choosing a weighting, and the choice made in the bedroom is the choice made there.

What the squares cannot show

Which weighting is credence. The two squares are drawn at equal size and in equal detail, because the figures cannot adjudicate. Every computation in this essay is agreed by all parties; each figure checks its answer against the corresponding formula, and those checks hold for both measures.

Memory. The drug is essential and nothing in the pictures depicts it. Without it, Beauty on Tuesday would remember Monday and know where she was, and the two squares would coincide at every waking she could not tell apart — which is none. The squares show the result of the erasure, not the reason it matters, which is that it makes two moments indistinguishable from the inside.

Whether the question has one answer. It is a live position that how likely is heads? asked at a waking is ambiguous between two well-defined quantities, and that both answers are correct answers to different questions. That position dissolves the puzzle rather than solving it, and much of the literature resists it, on the grounds that Beauty must bet, act and plan with one credence at a time.

Where the reasoning goes next

The natural continuation is the general form that all these problems share. Every update in this series has been a prior multiplied by a likelihood ratio — the host’s constraint, the parent’s habit, the chance of waking — and writing the update in odds, where it is a single multiplication, makes the common structure visible and the arithmetic trivial.

Beyond that lie problems where the evidence is a count rather than a sentence: sequential tests where each posterior becomes the next prior, and questions of what a test is worth in bits of information rather than in percentages. And there is the two-envelope problem, where an argument that looks exactly like the ones here produces a conclusion that cannot be right — the case where the task is not to choose between two measures, but to find the step at which a single measure was quietly abandoned.

Two measures, one coin

The coin in Sleeping Beauty is fair, the protocol is fully stated, and every frequency in the experiment is agreed. Half of runs are heads; a third of wakings follow heads. A ticket bought once per run is fair at one half and a ticket bought at every waking is fair at one third. Told that it is Monday, a thirder moves to one half and a halfer to two thirds, and in the extreme versions each side’s answer runs to an extreme that the other side can point to.

What is not agreed is whether Beauty’s credence on waking is a share of runs or a share of wakings. Earlier puzzles of this kind were settled by stating how the evidence was produced. This one states it in full and is still not settled, because the disagreement was never about the evidence. It is about what an uncertain person, at one moment among several that look the same from the inside, is uncertain over.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A dashed tag is an object no other essay names yet.

AreaBayes' theoremConditional probabilityCounting argumentExpectationLaw of large numbersLikelihoodSample space