Applied

Patience instead of a contract

Commitment had to assume an announcement binds. Play the same game again tomorrow and the assumption is unnecessary — the future does the binding. What it costs is that nearly every outcome becomes an equilibrium, so a theory that could not choose between two now cannot choose between infinitely many.
17 min read 7 figures The same thing twiceOne point away

Worth reading first: Worth more for being seen first · Two equilibria and no way to choose.

The essay on moving first ends on an assumption it cannot discharge. A leader who announces a mixture gains from the announcement only if the announcement binds, and nothing in the game makes it bind — a leader free to revise after seeing the reply would revise, a follower who knows that does not believe the announcement, and the whole construction falls back to the simultaneous game.

The usual repairs are external: a contract, a public randomising device, a bridge burnt behind the army. There is one repair that is internal, and it is the subject of this page. Play the same game again tomorrow.

The payoffs repetition makes available in the prisoner's dilemma. A plot of the two choosers' average payoffs, with the stage game's four cells marked, their convex hull drawn, the two minmax values shown as lines, and the region above both shaded.
Fig. 1 The prisoner’s dilemma as a picture of payoffs rather than of choices. The four dots are its four cells; the dashed outline is every average a long run of play can produce; the two dotted lines are what each chooser can be held to whatever they do; and the shaded region is what patience makes available. The one equilibrium of the game played once sits in the corner.

The single equilibrium of that game pays (1,1)(1, 1) and the shaded region reaches (3,3)(3, 3) and beyond. Repetition converts a game with one bad answer into a game with an enormous number of answers, most of them good. That is the folk theorem, and the second half of this essay is about why it is usually described as an embarrassment rather than as a solution.

The stage game, and why it has nothing to offer

Best replies in the prisoner's dilemma. A bimatrix with every best reply marked on both sides and every cell that is a best reply for both boxed as a pure equilibrium. One such cell was found.
Fig. 2 The game played once, with every best reply marked on both sides. Defecting beats cooperating in both columns — 4 against 3, and 1 against 0 — so it is dominant; the same holds by symmetry for the columns; and the one cell carrying both marks pays each chooser 1 when they could have had 3.

Nothing in that matrix is in doubt. D is better than C whatever the other does, so both choose D, both receive 1, and the cell paying 3 to each is reachable by nothing the definition of an equilibrium permits. It is the cleanest failure in the subject: not a selection problem, not a coordination problem, just a unique answer that both parties would pay to avoid.

The announcement device works on it. A leader who could bind itself to cooperate would not be believed, because cooperating is dominated — but neither chooser here is a leader, and there is nothing to announce.

What a chooser can be held to

Repetition introduces a quantity that the one-shot game does not need, and getting it right is where the theorem’s content lives.

Ask how badly one chooser can be treated if the other sets out to punish. The punisher chooses first, knowing that the punished will then best-reply, so the number is

v  =  minwhat the punisher does  maxwhat the punished then does  what the punished receives.\underline{v} \;=\; \min_{\text{what the punisher does}} \; \max_{\text{what the punished then does}} \; \text{what the punished receives.}

That is the minmax, and it is a bound nobody can be pushed below: whatever the punisher is doing, best-replying to it earns at least v\underline{v}. So no equilibrium of a repeated game can pay a chooser less than that on average, because a chooser earning less would deviate and best-reply for ever. That is the same guarantee both sides name in a game of pure conflict, asked of one side of a game that is not zero-sum — where the two guarantees no longer add to nothing and the gap between them is the region this essay is about.

The minimisation is over mixtures, not over actions, and the difference is not a refinement.

The value of a 2×2 zero-sum game, named from both sides. The row chooser's expected payoff against each column as a line over the mixing probability, with the lower envelope and its maximum, beside the same construction from the column chooser's side. Both give 0.
Fig. 3 Matching pennies, solved the way a zero-sum game is solved: each chooser’s expected payoff against each of the other’s pure actions is a straight line, the lower envelope of the two lines is the guarantee, and its peak is the value. The value is nought, reached by mixing evenly — and no pure choice by the punisher holds the other below 1.

In matching pennies a punisher restricted to pure actions cannot hold the other below 11, since whichever side the punisher shows, the other can match it. A punisher allowed to mix holds the other to 00. The minmax over pure actions is 1 and over mixtures is 0, and a folk theorem stated with the first number would claim that a region above (1,1)(1,1) is sustainable in a game whose payoffs never add to more than nought — which is to say it would claim an empty region and be vacuously false.

That is why every figure here computes the minmax by minimising over a mixing probability, as an exact rational, with the punished chooser’s best reply drawn as the upper envelope of two straight lines. The minimum of such an envelope sits at an endpoint or at the crossing, so three candidates settle it exactly and no search is needed — and each figure then checks, on a lattice of twenty-one mixtures, that no other mixture of the punisher’s does better.

The region, and what fills it

The hull in the hero figure is the set of payoff pairs a long run can average to. Its corners are the stage game’s four cells, and every point between them is reachable by playing the cells in the right proportions — a repeated game’s average payoff is a weighted average of the cells visited, and any weights can be arranged by alternating. The hull is therefore a convex set by construction rather than by hypothesis, which is what makes the clipping in the figures exact rather than approximate.

Clip that hull to the payoffs strictly above both minmax numbers and what is left is the folk-theorem region:

every feasible average payoff strictly better than the minmax for both choosers is the average payoff of some equilibrium of the repeated game, once the choosers are patient enough.

For two choosers that statement is exact and needs no further hypothesis; with three or more it needs the feasible set to be genuinely multi-dimensional, because punishing one chooser must be arrangeable without punishing the punishers equally, and a degenerate payoff space can make that impossible.

The payoffs repetition makes available in chicken. A plot of the two choosers' average payoffs, with the stage game's four cells marked, their convex hull drawn, the two minmax values shown as lines, and the region above both shaded.
Fig. 4 The same construction on chicken, where the two equilibria of the game played once pay 7 and 2 in one order or the other. The minmax pair is (2,2)(2,2), the region above it is most of the hull, and its best total is 12 against the 9 the one-shot equilibria manage. Every asymmetric split of that 12 is in the region too.

The chicken figure is where the theorem’s reach becomes uncomfortable. Both one-shot equilibria are inside the region, so repetition does not remove them. The efficient pair at (6,6)(6,6) is inside. So is (6.5,5.5)(6.5, 5.5), and (3,9)(3, 9), and every other point of a two-dimensional set. Repetition does not select. It adds.

The payoffs repetition makes available in the stag hunt. A plot of the two choosers' average payoffs, with the stage game's four cells marked, their convex hull drawn, the two minmax values shown as lines, and the region above both shaded.
Fig. 5 The stag hunt, where the news is different again. Both of its one-shot equilibria pay each chooser at least the minmax of 3, so both survive; the efficient pair at (4,4)(4,4) was already an equilibrium; and the region repetition adds is the thin sliver of unequal splits between them. Here the folk theorem contributes almost nothing, because the one-shot game was not the problem.

Three games, three completely different amounts of new ground. The prisoner’s dilemma gains everything, chicken gains a great deal, the stag hunt gains almost nothing — and the quantity that predicts which is how far the one-shot equilibria sit above the minmax. A game whose equilibrium already pays the minmax has the whole region to gain; a game whose equilibria are already efficient has nothing.

How patient is patient enough

The theorem says once the choosers are patient enough and that phrase hides an exact number, which is worth extracting because it is the one quantity in this essay that a person could plausibly estimate about a real situation.

Take the simplest strategy that supports cooperation: cooperate, and if the other ever defects, defect for ever afterwards. That is the grim trigger. Whether it is an equilibrium is an inequality between two streams of payoff.

The patience an agreement needs in the prisoner's dilemma. Two curves against the discount factor: the total value of keeping an agreement for ever and the total value of breaking it once and being punished afterwards, crossing at one point.
Fig. 6 Two totals against the discount factor: keeping the agreement for ever at 3 a period, and breaking it once for 4 and then living in the punishment at 1 a period. Below a discount factor of a third, breaking pays; above it, keeping pays; and at exactly a third the two streams are worth the same.

Write RR for what the agreement pays each period, TT for what breaking it pays once, PP for the punishment, and δ\delta for how much the next period is worth against this one. Keeping the agreement is worth R/(1δ)R/(1-\delta). Breaking it is worth TT now and then PP for ever, which is T+δP/(1δ)T + \delta P/(1-\delta). The two are equal at

δ=TRTP,\delta^\ast = \frac{T - R}{T - P},

and above it the threat is enough. In the prisoner’s dilemma of these figures that is (43)/(41)=1/3(4-3)/(4-1) = 1/3 — so a pair who value tomorrow at a third of today can hold the agreement, which is a very low bar.

Two things about that formula are worth carrying. It depends only on the ratios between the three payoffs, so scaling the game changes nothing. And it rises with the temptation and falls with the severity of the punishment, which is the statement that an agreement is easier to hold when breaking it gains little and when the consequence is bad — obvious in words, and here as a closed form whose two differences are exactly the two intuitions.

The patience an agreement needs in chicken. Two curves against the discount factor: the total value of keeping an agreement for ever and the total value of breaking it once and being punished afterwards, crossing at one point.
Fig. 7 The same computation on chicken, where breaking the agreement is worth 7 against a jointly-agreed 6 and the punishment pays 2. The critical discount factor is a seventh — patience matters even less here, because the gain from deviating is one against a punishment that costs six.

The discount factor has two readings and they are not the same. It can mean impatience, with δ\delta near one describing choosers who value the future nearly as much as the present. It can equally mean probability of continuing: if the relationship ends after each period with some chance, then δ\delta is the chance it does not, and the arithmetic is identical. The second reading is the one that makes the theorem applicable, because it says the condition is about whether there is a next time rather than about anybody’s temperament.

What has actually been repaired

The essay opened with a gap — an announcement that binds nothing — and it is worth checking precisely what has replaced it.

The binding is now a best reply. Under grim trigger, a chooser who has been cooperating and considers defecting computes the two streams above and finds defection worse. Nothing external enforces anything; the constraint is the other side’s strategy, and the other side’s strategy is optimal against this one. That is the whole content of the repair, and it is genuinely different from a contract: no third party, no enforcement, no observability beyond seeing what happened last period.

What is not repaired is the threat’s credibility after the fact. Grim trigger says punish for ever, and once a defection has happened, carrying out the punishment costs the punisher too — both are then earning 1 instead of 3. A chooser who has already deviated might reasonably propose starting again, and a punisher who would accept is a punisher whose threat was never credible.

That objection is exactly right and it has an exact answer. Requiring the strategies to be optimal at every point of the game, including after deviations nobody expected is what subgame perfection asks, and grim trigger in the prisoner’s dilemma passes it — because permanent defection is an equilibrium of the game from any point onward, so carrying out the punishment is a best reply to being punished. In games where the minmax punishment is not itself an equilibrium the threat has to be more elaborate, and the folk theorem still holds: the construction punishes for a finite number of periods and then rewards the punishers for having done it. The region does not shrink.

The name, and what it is confessing

The result is called the folk theorem because it was known and used before anybody published a proof of it — in the 1950s it was common property among game theorists with no attributable source, which is what the word means in this context. Aumann and Shapley, and Rubinstein, gave versions for the undiscounted case around 1976; Fudenberg and Maskin’s 1986 paper is the one that settled the discounted case and the subgame-perfect version together.

The name has a second reading that the subject has come to prefer. A theorem saying almost anything is an equilibrium is not a prediction; it is the announcement that the concept has stopped predicting. Two equilibria and no way to choose was a difficulty; a continuum of them is the same difficulty with the arithmetic run to its conclusion.

Read that way the folk theorem is the strongest available argument that equilibrium is not by itself a theory of behaviour. It settles what is consistent — which is a real question, and the answer is a region — and it says nothing about what happens. Everything that has been built on top of it is an attempt to say more: bargaining solutions that pick a point of the region by axioms, evolutionary arguments that ask which points are reached by a population whose shares grow with how well they do, and learning models that ask what choosers converge to when they are not assumed to know the answer in advance. A public signal that recommends does something similar with a much weaker device and reaches a much smaller set, and the comparison is instructive: a correlating device enlarges the set of outcomes by a polytope that can be computed in one linear program, while repetition enlarges it to nearly everything and computes nothing.

Where repetition is doing the work in practice

Three cases where the mechanism is visibly the one above, and the third is the one worth arguing with.

Long relationships between firms. A supplier who could cut quality once for a gain does not, because the relationship is worth more than the gain — which is the inequality of this essay with δ\delta read as the chance of continuing. Contracts exist and are famously incomplete; what holds the part nobody wrote down is the arithmetic here. It is worth noticing that this is the same repair as the planner who closes a road before anybody chooses a route, arrived at from the opposite end: there an outside party removed an option, here the parties remove it from each other by what they will do next.

The tournament that made the mechanism famous. Axelrod’s round-robin of prisoner’s-dilemma strategies in 1980 was won by tit-for-tat, four lines long, which cooperates and then copies. Its success is often described as a discovery about cooperation and is better described as a demonstration of this theorem’s hypothesis: the tournament repeated the game two hundred times, and in a repeated game a strategy that rewards cooperation and punishes defection is a best reply to itself.

And the case against reading too much into it. A finitely repeated prisoner’s dilemma with a known end has a unique equilibrium, and it is defection in every period — work backwards from the last, where defection is dominant, and the induction eats the whole game. So the entire region depends on the horizon being unbounded or uncertain, and a known last period destroys it completely. That discontinuity is severe: two hundred periods of certain repetition gives full defection, and two hundred periods with a small chance of a two-hundred-and-first gives the folk theorem back. Whether that says something about behaviour or something about the model is a question the model cannot answer.

Two sources of the same region, and a warning about counting them

There is a second construction that reaches a region of this shape and it is worth separating from the first, because they are often run together and only one of them is about repetition.

The folk theorem’s region comes from strategies that condition on history. Nothing is agreed in advance; each chooser’s rule says what to do given what has happened, and the equilibrium condition is checked at every history.

A correlated equilibrium’s region comes from a device that draws a cell and whispers, and it is a set of distributions over the cells of the game played once. Both sets contain the game’s Nash equilibria, both are convex, and both can be drawn as regions in the same payoff plane — so a figure of one looks like a figure of the other.

They are not the same set and the containment runs one way. Every correlated equilibrium’s payoff is feasible and pays each chooser at least their minmax, so it sits inside the folk region; the reverse fails badly, since the folk region of the prisoner’s dilemma reaches (3,3)(3,3) and its only correlated equilibrium is (1,1)(1,1) — a device that whispers cannot make a dominated action attractive, and no amount of correlation repairs a dominant strategy. What repetition adds that correlation cannot is a future to lose.

The warning is about arithmetic rather than about concepts. A region drawn in a payoff plane is compatible with several completely different constructions behind it, and the constructions differ in what they assume about enforcement, about observation and about time. Two theories that produce the same picture are not the same theory, which is the same caution three functions sharing a name earn in a different corner of this collection.

A region is not a strategy

The region is drawn closed and the theorem is about its interior. A payoff exactly equal to a chooser’s minmax is not sustainable — a chooser held to exactly what they can guarantee is indifferent to deviating, and the construction needs a strict margin to pay for the punishment threat. The shading runs to the two dotted lines and the claim does not.

Nothing in any figure is a strategy. Every point in the region is the average payoff of some equilibrium, and the equilibrium is a rule saying what to do after every history of play — an object with no picture, indexed by every finite sequence of past cells. What is drawn is the arithmetic of the outcome, and the construction that reaches it is prose.

And the patience threshold is drawn for one strategy. The critical discount factor in the two discount figures is grim trigger’s. A different supporting strategy gives a different threshold, usually higher, since a milder punishment needs more patience to deter the same temptation; the folk theorem’s own construction is not grim trigger and its threshold is not the number drawn. What the figure shows is that a threshold exists and where it comes from.

Still open: what a region is a prediction of

Repetition makes an external commitment unnecessary and leaves a set where a point was wanted. Two directions lead out.

One narrows the set by asking more of the strategies — that they be simple, that they be robust to mistakes, that they not require remembering everything. A single mistaken defection under grim trigger destroys the agreement for ever, which is a poor property for a rule meant to survive in a noisy world, and the strategies that survive noise are a much smaller family than the theorem permits.

The other narrows it by taking away the assumption that has been carried silently since the first of these games: that the payoff matrix is known to everybody and known to be known. A slightly noisy reading of it removes exactly that, and the effect is the opposite of this one — a continuum of equilibria collapses to a single point, and it does so because of a perturbation too small to matter.

What repetition is and is not

The habit worth taking is to distinguish two things a model can do when it has too few answers.

Adding structure to an underdetermined problem usually narrows it. Repetition is the case where adding structure widens it — and widens it so far that the concept being used stops distinguishing anything. That is worth recognising because it looks like progress while it is happening: the prisoner’s dilemma’s bad answer goes away, cooperation becomes explicable, and the theory appears to be working. What has happened is that the set of consistent answers has grown to include the one that was wanted, along with everything else.

A concept that can accommodate any outcome has not explained the outcome that occurred. The test to apply is the one this essay’s three figures apply: ask what the concept rules out, and if the answer is a set with an interior, the work of choosing inside it has not been started.

What links here

Computed from the collection, not written here: the essays that point at this one.

Reads more easily once this is understood

Essays that name this one as worth reading first.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A dashed tag is an object no other essay names yet.

Best replyCommitmentConvexityEquilibrium selectionMinimaxMixed strategyNash equilibriumZero-sum game