Applied

A signal both can see

Two choosers who randomise privately can reach a set of outcomes that is smaller, and worse, than the set they reach when a device draws one cell and whispers each of them their half of it. Nothing is enforced and nobody is bound, and the arrangement is stable anyway.
15 min read 6 figures Decided by exhaustionOne point away

Worth reading first: The value from both sides · The road that makes everyone later.

Two choosers face a matrix in which each would rather the other gave way. If both press on, both get nothing. If one gives way, the one who pressed on collects seven and the one who yielded collects two. If both give way, each collects six.

There are two equilibria in pure choices — one for each way of deciding who yields, found by the same best-reply search the rung below runs over every cell — and each pays nine between the pair. There is a third in private mixtures, and it pays 28/328/3, which is worse than both. Every arrangement the definition of an equilibrium reaches on this matrix is worse than one of the cells nobody can reach, and the cell in question, where both give way, pays twelve.

A signal both can see, and neither wants to disobey. A two-by-two game with a distribution over its four cells, drawn as the weight on each. Obeying the recommendation is a best reply for both choosers, and the pair collects 21/2 between them.
Fig. 1 A device draws one of the four cells from the distribution shown and tells each chooser only their own half of it. Told to press on, a chooser knows the other was told to yield; told to yield, a chooser faces an even split. Obeying is a best reply in every case, the four conditions hold exactly, and the pair collects 21/221/2 — more than any equilibrium without the device.

Nothing about that arrangement is enforced. Nobody is bound to obey, no contract exists, and each chooser sees only their own recommendation. It is stable for the same reason a Nash equilibrium is: nobody can do better by disobeying alone.

What the device changes

A mixed equilibrium is a pair of independent lotteries. The row chooser flips their coin, the column chooser flips theirs, and the two coins have nothing to do with each other — so the distribution over the four cells is a product, and the products form a curved surface inside the set of all distributions.

A correlated equilibrium is any distribution over the four cells satisfying an incentive condition. The distributions are no longer required to be products, so the set is everything the condition allows, and the condition is that on hearing a recommendation, neither chooser prefers to do something else.

Best replies in a game with no pure equilibrium. A bimatrix with every best reply marked on both sides and every cell that is a best reply for both boxed as a pure equilibrium. No such cell was found.
Fig. 2 The reason the enlargement matters. This matrix has no pure equilibrium at all — every cell is marked and none carries both marks — so before mixtures were allowed there was nothing to talk about. Correlation is the same kind of enlargement made a second time, and it buys more than the first one did.

The word correlated is doing exactly one job. The recommendations are allowed to depend on each other, and that dependence is the whole of what the device supplies. It supplies no enforcement, no communication between the choosers, and no information about the payoffs that either lacked.

What a recommendation tells a chooser

The incentive condition is a statement about conditional beliefs, and it is worth writing out because the arithmetic is short and the conclusion is not obvious.

Suppose the distribution puts weight 14\tfrac14 on press–yield, 14\tfrac14 on yield–press, and 12\tfrac12 on yield–yield, with nothing on the crash.

Told to press on, the row chooser knows the drawn cell was press–yield, since that is the only cell in which they are told to press. So the other has certainly been told to yield, pressing on pays seven and yielding pays six, and pressing on is right.

Told to yield, the row chooser knows the cell was yield–press or yield–yield, with weights 14\tfrac14 and 12\tfrac12 — so, conditionally, one chance in three that the other presses and two in three that they yield. Yielding pays 132+236=143\tfrac13\cdot 2 + \tfrac23\cdot 6 = \tfrac{14}{3}. Pressing on pays 130+237=143\tfrac13\cdot 0 + \tfrac23\cdot 7 = \tfrac{14}{3}. Equal — so yielding is a best reply, and the condition holds.

The column chooser’s two conditions are the same by symmetry. Four inequalities, all satisfied, and the arrangement is an equilibrium.

The expected total is 149+149+1212=212\tfrac14\cdot 9 + \tfrac14\cdot 9 + \tfrac12\cdot 12 = \tfrac{21}{2}, against nine for either pure equilibrium and 28/328/3 for the mixed one.

Why the mixture cannot do it

The reason no private randomising reaches this is worth isolating, because it explains what correlation is buying.

Under independent mixtures the crash has probability pqp\cdot q, where pp and qq are the two chances of pressing on. To make the crash rare, both pp and qq must be small; but if both are small, both are yielding most of the time and each has an overwhelming reason to press on instead. The equilibrium condition ties pp and qq to values that make the other indifferent, and at those values the crash has probability 19\tfrac19.

The device removes the crash entirely by putting no weight on it — which is not a choice available to two independent coins, since a product distribution has weight on a cell whenever it has weight on the corresponding row and column. Independence forces every combination that either party plays to occur, and the whole gain here is from a combination that occurs never.

The value of a 2×3 zero-sum game, named from both sides. The row chooser's expected payoff against each column as a line over the mixing probability, with the lower envelope and its maximum, beside the same construction from the column chooser's side. Both give 19/15.
Fig. 3 The machinery of private mixing, on a zero-sum matrix where it works perfectly: the guarantee each side can secure is the envelope of straight lines, and the two guarantees meet. Nothing in that construction has any way of putting nought on a cell both sides play with positive probability.

Found by exhaustion, not by argument

The distribution above was not derived. It is the best point of a lattice, found by testing every distribution whose four weights are twelfths — all 455455 of them — against the four conditions and keeping the feasible one with the largest total.

A signal both can see, and neither wants to disobey. A two-by-two game with a distribution over its four cells, drawn as the weight on each. Obeying the recommendation is a best reply for both choosers, and the pair collects 10 between them.
Fig. 4 The distribution most often quoted for this game: equal weight on the three cells other than the crash. It is a correlated equilibrium — the conditions are checked and hold — and it pays 1010, which is less than the 21/221/2 the search finds. The tidy answer is not the best one.

That the search beats the quotable answer is worth a moment. Putting a third on each of the three non-crash cells is the arrangement every textbook uses, it is symmetric and memorable, and it leaves half a unit on the table. The better one puts more weight on the cell where both yield and less on the two where somebody wins, which is the direction a reader would guess and not the amounts.

The search is what makes the claim checkable. An argument that a particular distribution is optimal would need the linear program’s dual; a sweep of every distribution on a lattice needs nothing but arithmetic, and it establishes optimality over the lattice rather than over the whole polytope — which is a weaker statement, stated as such.

What the enlargement is worth, in one number

The previous rung measured the gap between an equilibrium and the best achievable outcome and gave it a name — the price of anarchy, which on its network was 144/119144/119. The same ratio can be asked here, and asking it says what the device is worth.

On the chicken matrix the best total any arrangement could reach is 1212, at the cell where both give way, and no equilibrium of any kind reaches it. Without a device the best is 99; with one it is 21/221/2. So a public coin closes 38\tfrac{3}{8} of the gap between the best equilibrium and the best outcome, and closes none of the rest.

That fraction is the honest summary and it is worth stating in both directions. The device is not a repair. It does not reach the cooperative cell, it does not remove the conflict, and it does not make either chooser better off than the pure equilibrium they prefer — the row chooser gets 21/421/4 under the correlated arrangement and 77 at the pure equilibrium where the other yields. What it does is make an arrangement neither can veto better than the arrangements that were previously available, which is a claim about the set of stable outcomes rather than about anybody’s welfare.

And the gap it closes varies with the game. On the battle matrix, where the conflict is about which of two agreements to reach rather than about who yields, the device closes nothing at all: the best correlated total equals the best pure one, because the loss there comes from failing to agree rather than from a combination independence forces. Correlation buys exactly what independence was costing, and on a game where independence costs nothing it buys nothing.

The set is a polytope, and that is the structural fact

The conditions are linear in the distribution: each says one weighted sum of the weights is at least another. So the set of correlated equilibria is defined by finitely many linear inequalities together with the requirement that the weights are non-negative and sum to one, which makes it a polytope.

The vertex that maximises 3x₁ + 4x₂. A two-variable linear program's feasible region, drawn from the exact intersection of every pair of its constraints, with the objective's contour lines and the optimal vertex marked.
Fig. 5 A linear program in two variables: a region cut out by inequalities, and the corner at which a linear objective is largest. The set of correlated equilibria has exactly this shape in a higher-dimensional space, and the consequence is that maximising anything linear over it — a total payoff, one chooser’s payoff, the chance of avoiding a particular cell — is a problem of this kind.

That is a much stronger structural statement than anything available for Nash equilibria, and the contrast is the reason the notion is taken seriously.

Finding a Nash equilibrium is a fixed-point problem. Existence is proved by a theorem about a map of a region into itself, the proof is not a construction, and the set of equilibria is generally not convex — the average of two Nash equilibria is usually not one. That non-convexity is not a technicality: it is why there is no sensible notion of “the” equilibrium of a game, and why selection is a subject rather than a calculation.

Finding a correlated equilibrium is a linear program. Existence needs no fixed point at all, since every Nash equilibrium is a correlated one and one of those exists; the set is convex by construction, so the average of two correlated equilibria is one; and the whole apparatus of duality applies, so the best point comes with a certificate that it is the best.

The gap between those two situations is large and is not a matter of degree. It is the difference between a problem with a curve in it and a problem without one.

Where every Nash equilibrium sits inside it

The containment is easy to check and worth checking, because it is what makes the enlargement an enlargement rather than a different subject.

Take any pair of independent mixtures that is a Nash equilibrium and read it as a distribution over cells. Told to play a particular row, the chooser learns nothing at all about the column, since the two lotteries are independent — so the conditional belief is the unconditional one, and the incentive condition reduces to the Nash condition, which holds by assumption.

A signal both can see, and neither wants to disobey. A two-by-two game with a distribution over its four cells, drawn as the weight on each. Obeying the recommendation is a best reply for both choosers, and the pair collects 5 between them.
Fig. 6 The same construction on a game where the two want to agree and disagree about what on. Its two pure equilibria are correlated equilibria, the search confirms it, and here the best correlated outcome ties with the better pure one — the enlargement buys nothing on every game.

So the set of correlated equilibria contains every Nash equilibrium, and the containment is generally strict. On the battle game above it is not strict in the quantity being maximised: the enlarged set is larger and its best total is the same. An enlargement that always helps would be suspicious, and the honest description is that correlation buys something exactly when the loss comes from combinations that independence cannot avoid.

What the device has to be

Nothing above requires an actual device, and it is worth being precise about what is and is not needed, because the assumption is doing real work and is easy to smuggle.

A public randomising event, observed differently. That is all. A traffic light is the standard illustration and it is a good one: it produces correlated recommendations, neither driver sees the other’s, and obeying is a best reply because disobeying while the other obeys is expensive.

Not enforcement. The light does not stop anybody. What stops them is the arithmetic of the conditional belief, and the whole content of the notion is that the arithmetic suffices.

Not communication. The choosers never speak. If they could, they would face a different and larger problem, in which promises can be exchanged and the question of whether a promise binds arrives — which is the territory of a division nobody can walk away from, where an agreement is assessed by what a group could get by leaving rather than by what one participant could get by deviating.

And not a fair device. The distribution need not treat the two symmetrically. A device favouring the row chooser is still a correlated equilibrium provided the column chooser’s two conditions hold, and the set of correlated equilibria of the chicken game includes points far more favourable to one side than either pure equilibrium.

That last is where the notion gets uncomfortable, and it is the honest limitation. The set is large, the definition selects nothing within it, and the question of which correlated equilibrium happens is no more settled than the question of which Nash equilibrium does — which is the next rung’s subject, asked of a smaller set.

Where it needs care

The recommendation must be private. If each chooser sees the whole draw, the row chooser told to yield knows exactly what the column chooser was told, the conditional belief collapses to a point, and the conditions change. The arrangement above survives that change on this matrix and does not in general.

The device’s distribution must be common knowledge. Both choosers reason about conditional probabilities, and those need the distribution. A device whose behaviour one party does not know is not this construction; it is a game with incomplete information, which is a different and much harder object.

The conditions are inequalities on the drawn distribution, not on the payoffs. Changing a payoff changes which distributions are feasible, and the change is not monotone: raising the reward for pressing on shrinks the set, and raising it enough leaves only the pure equilibria.

And the search is over a lattice. Twelfths were enough to find a better point than the quoted one; a finer lattice could find better still, and the exact optimum is the solution of a linear program rather than of a sweep. The figure says what it tested and does not claim more.

Aumann, and a definition that arrived sideways

Robert Aumann introduced correlated equilibrium in 1974, in a paper whose subject was not this at all. He was asking what happens when the notion of a mixed strategy is taken seriously as a belief rather than as a physical randomisation, and the correlated version fell out as the thing that survives when the independence assumption is dropped.

That is worth registering because the independence assumption had never been argued for. It is present in Nash’s definition because a mixed strategy is defined as one chooser’s lottery, and a pair of lotteries is a product because that is what a pair of separate objects is. Nobody chose independence; it arrived with the notation.

Aumann returned to the notion in 1987 with a much stronger claim: that players who are Bayesian rational and share a common prior must be playing a correlated equilibrium, whatever they believe about each other. On that reading it is not a generalisation of Nash’s notion at all but the correct one, with Nash’s being the special case where beliefs happen to be independent. A definition that started as a technical relaxation became the primary object, which is a fair description of how several of the notions in this field arrived.

What the pictures cannot show

The figures draw a distribution over four cells and the four conditional comparisons. What they cannot draw is the set — the polytope of all feasible distributions — because it lives in the three-dimensional space of distributions over four cells and there is no room on a page for it beside the game it belongs to.

That absence hides the structural fact this essay is mostly about. The reader is shown one point and told it was the best of 455455; a picture of the polytope with that point at a corner would be the honest illustration of “this is a linear program”, and it would be a picture of an object with no payoffs visible in it.

Nor can any figure show what makes the arrangement stable. Stability is a statement about what each chooser would do on hearing something they were not told, and a drawing of what happened has no room for what did not.

The ladder from here

Rungs above: why a best-reply search terminates on a congestion game, which is the structural fact that makes the equilibria of the previous rung’s network computable at all. Two equilibria and no way to choose, where the selection question is asked of the smaller set. Communication equilibria, where the choosers may talk to the device as well as listen. Coarse correlated equilibrium, which weakens the condition to a comparison made before the recommendation arrives and is what the standard learning rules converge to. And the price of anarchy over the correlated set, which is bounded by the same 4/34/3 the previous rung quoted and for a reason that is easier rather than harder.

An assumption that arrived with the notation

The habit worth carrying is about where the independence went unexamined.

Nash’s definition needs each chooser’s behaviour to be a lottery, and two lotteries held by two people are independent unless something makes them otherwise. That is not an assumption anybody defended; it is what “two separate objects” means, and the notation for a pair of mixed strategies has no slot in which a dependence could be written.

So the assumption was invisible, and removing it required rewriting the notation before the question could even be posed. A distribution over cells has a slot for dependence; a pair of distributions over rows and columns does not. What Aumann changed first was the object, and the theory followed.

That is a recurring shape. When a generalisation seems unavailable, the obstruction is often the representation rather than the mathematics — and the test is whether the more general object is easier, which here it is: a polytope rather than a fixed point, a linear program rather than a search.