Concept

Conditional probability

The chance of one event given that another is known to have happened. It is the quantity every updating rule is written in terms of, and it is computed as a ratio of two probabilities on the same space.

Named by 16 essays across 2 fields — each of them below, with the objects they name alongside it.

Bayes' theorem as two rectangles. A unit square split by how common the condition is (1.0%) and then by how the test behaves. Of everyone who tests positive, the fraction who have it is 16.7%.

Bayes' theorem is a picture of a square

A test that is 99% accurate returns a positive result. The chance it is right can easily be under one in five, and the reason is visible the moment the population is drawn as a square rather than described as a formula.

probability · Bayes
Three doors, as areas. Staying wins 33.3% of the time and switching wins 66.7%, because the host's choice is constrained by what the host can see, so opening a door rules a region out without moving any boundary.

The door that was not opened

Three doors, one prize, a host who opens a losing door and offers a swap. Switching wins two times in three, and the reason is not about doors — it is about what the host was allowed to do.

probability · Bayes
5,040 orders, 7 thresholds, one best rule. For each number of candidates passed over, the share of the 5,040 possible arrival orders in which the rule ends up with the best of the 7. The count is exhaustive.

When to stop looking

Candidates arrive one at a time in a random order. Each must be accepted or rejected on the spot, with no going back and no way to know what is still to come. The best possible rule is to look at about a third of them and then take the first one that beats everything seen — and it works about a third of the time, however many there are.

probability · Optimal stopping
A signal both can see, and neither wants to disobey. A two-by-two game with a distribution over its four cells, drawn as the weight on each. Obeying the recommendation is a best reply for both choosers, and the pair collects 21/2 between them.

A signal both can see

Two choosers who randomise privately can reach a set of outcomes that is smaller, and worse, than the set they reach when a device draws one cell and whispers each of them their half of it. Nothing is enforced and nobody is bound, and the arrangement is stable anyway.

applied · Equilibrium
Two equilibria, and two tests that disagree. The row chooser's expected payoff from each option against the column chooser's behaviour, for a joint effort worth more than a safe one. The lines cross at 0.750, which is the mixed equilibrium and the boundary between the two basins.

Two equilibria and no way to choose

A game can have two states nobody wants to leave, one paying more than the other, and the definition of an equilibrium has nothing to say about which happens. The two standard tie-breakers disagree, and the one that wins is usually the worse.

applied · Equilibrium
58.6% at 46 candidates, against 37% without the values. The chance of ending with the best candidate when the values are shown, against the number of candidates, for 10 sizes. It falls towards 0.5802 rather than towards 1/e.

When the numbers are shown

The secretary rule wins a third of the time and cannot do better, because it is told only who is ahead. Show the actual values and say where they came from, and the same problem is won three times in five — by a standard that falls as the end approaches.

probability · Optimal stopping
About the fourth-best, whatever the size of the field. The smallest expected rank achievable by an online rule, against the number of candidates, for 10 sizes. It rises to 3.8516 at 2500 candidates and its limit is 3.8695.

Giving up on the best

The secretary rule treats landing the second-best exactly as badly as landing the worst, which is a strange thing to want. Ask instead for the smallest average rank and the answer is about the fourth-best candidate — whatever the size of the field, and whether it is ten or ten million.

probability · Optimal stopping
An online rule taking nine tenths of what an oracle takes. The share of the oracle's expected maximum secured by the best single threshold, and by the threshold at the median of the maximum, for 8 field sizes of independent uniform values.

Half of what an oracle takes

Compare an online rule not against the best it could have done but against a rule that has seen every value in advance. One fixed threshold secures half of what the oracle collects, whatever the distributions are — and there is an example on which half is all there is.

probability · Optimal stopping
Add the odds from the end: start at event 5 of 12. The odds of 12 independent events with chances 1/k — the secretary problem, the sum of the odds from the end reaching one at event 5, and the chance of stopping on the last success from each starting point, highest at 39.6%.

Add the odds from the end

Watch a sequence of independent events and try to stop exactly on the last one that happens. Add up the odds of the events from the end backwards until the total reaches one, and stop at the first success from there. That rule is the best possible for any probabilities whatever, and the secretary problem is the special case in which the chances are one over the position.

probability · Optimal stopping
The band two equilibria occupy, and the point a noisy reading leaves. A line of values of the payoff parameter with three regions marked — staying out dominant, both actions equilibria, investing dominant — and a single threshold inside the middle region.

The reading that is almost right

Every account of simultaneous choice so far has assumed the payoffs are known to both choosers and known to be known. Replace that with each chooser seeing a private reading off by a little, and a band of equilibria closes to a single point — so the assumption nobody states decides the answer.

applied · Equilibrium
How long 4 equally likely patterns take to appear. A bar for each of 4 patterns of 3 coin tosses giving the expected number of tosses before it first appears, with the lengths at which each pattern overlaps itself listed.

Two patterns, one chance, different waits

HTH and HTT are equally likely in any given window of three tosses. Waiting for HTH takes ten tosses on average and waiting for HTT takes eight, and the difference is not about probability at all — it is about what a failed attempt leaves behind.

probability · Expectation
Two children, and at least one is a boy. Four equally likely families drawn as quarters of a square: the question “is at least one a boy?” rules out only the girl–girl family, and leaves three equal quarters; the chance of two boys is 33.3%.

Two children and the sentence about one of them

A family has two children and at least one is a boy. The chance that both are boys is one in three — or one in two, or anything from one in three to certainty — and every one of those answers is right for some way the sentence could have come to be said. There is no host and no door, and the protocol is still the whole problem.

probability · Bayes
One experiment, measured by runs and by awakenings. Two unit squares for the same coin and the same schedule: the same experiment weighed two ways: by runs, heads keeps half the square; by awakenings, heads is one of 3 equal slices.

One coin, counted by runs and by wakings

Beauty is put to sleep and a fair coin is tossed. Heads, she is woken once; tails, twice, with the first waking erased from her memory. Each time she wakes she is asked how likely heads is. One half, say some; one third, say others; and unlike every earlier puzzle of this kind, stating the protocol exactly does not end the argument.

probability · Bayes
How much the first player can guarantee, as the coin's bias moves. A plot of the first player's best guaranteed winning chance in Penney's game against the probability of heads, a third on a fair coin and rising past one half only when the coin is heavily biased.

A coin that lets the first player win

On a fair coin the second player in Penney's game always has a better pattern than the first, and the first can hold them to no worse than two to one. Bend the coin and every overlap is paid for in the letters it uses: the replies change, the first player's share swings between a third and a half, and past a heads chance of 1/∛2 the first player simply names HHH and wins.

probability · Expectation
Evidence as steps on a scale of decibans. A horizontal decibel scale of odds with a starting dot at the prior and one arrow per piece of evidence, ending at a posterior probability of 44.2%.

Evidence measured in decibans

Write a probability as odds and take the logarithm, and every piece of evidence becomes a length. A positive result on a good test is thirteen decibans; a negative one is minus twenty. Lay the lengths end to end from the prior and the posterior is where they stop, in any order. The rule fails in exactly one way — when two pieces of evidence share a cause — and Turing built a code-breaking method on the arithmetic.

probability · Bayes
Switching envelopes when the largest amount is 64. A bar chart of the probability-weighted gain from switching at each amount that might be seen, 1 to 64: small positive bars and one large negative bar, adding to zero.

The envelope that always looks better

Two envelopes, one holding twice as much as the other. Open one, see an amount, and reason that the other holds double or half with equal chance — so switching gains a quarter on average. By symmetry the same argument says switch back. The step that fails is not the arithmetic; it is the claim that double and half are equally likely whatever amount is seen, which no honest prior allows — and there is one prior under which the other envelope really does look better at every amount.

probability · Bayes

Named alongside it

The objects these essays reach for when they reach for this one.

ExpectationBayes' theoremCounting argumentOptimal stoppingThreshold ruleAreaDecision procedureIrrevocable decisionSample spaceBackward inductionBest replyCorrelated equilibrium

All concepts