Applied

Worth more for being seen first

Moving first sounds like a disadvantage, since the other side gets to see the move and answer it. When the move is a mixture that is announced and believed, it is never a disadvantage, it is worth exactly nothing in a game of pure conflict, and in other games it is worth more than any equilibrium — sometimes by announcing an action that would never be played in secret.
16 min read 6 figures The same thing twiceOne point away

Worth reading first: The value from both sides · Two equilibria and no way to choose.

In every game on the pages before this one the choosers move at the same instant. Neither sees what the other did; each chooses against a belief. That is the defining feature of an equilibrium, and it hides a question that the definition never raises: what if one of them could move first, in public, and be believed?

The obvious answer is that moving first is a disadvantage. The other side sees the move and picks the best reply to it, which is exactly the information a secret move withholds. In rock, paper, scissors, a player forced to show a hand before the other chooses loses every time.

The obvious answer is wrong whenever the move being announced is a mixture — a probability of playing each action, with the dice rolled afterwards.

Announcing a mixture is worth 5/3 more than any equilibrium. The leader's payoff against the probability it announces for its first action, for a leader with a dominant action that is better off not being seen to play it, with the follower's reply switching where the follower is indifferent. The best announcement is worth 11/3; the best equilibrium of the simultaneous game is worth 2.
Fig. 1 A leader who announces the probability of playing U, and a follower who sees the announcement and replies with whichever column pays the follower more. The thick line is what the leader receives, switching where the follower’s preference switches. Announcing U with probability 2/32/3 is worth 11/311/3; the only equilibrium of the simultaneous game is worth 22.

Announcing a mixture is worth 5/35/3 more than the only equilibrium the game has. And the announcement that achieves it puts weight on an action the leader would never play in a simultaneous game.

A game with a dominant action

The game in that figure is small and completely transparent. The leader chooses U or D, the follower chooses L or R, and the payoffs are these.

Best replies in a game whose leader has a dominant action. A bimatrix with every best reply marked on both sides and every cell that is a best reply for both boxed as a pure equilibrium. One such cell was found.
Fig. 2 The payoffs, with every best reply marked on both sides. U beats D for the leader in both columns — 2 against 1, and 4 against 3 — so U is dominant. Against U the follower prefers L, and the one cell that is a best reply for both is UL.

In the simultaneous game there is nothing to discuss. U is better for the leader whatever the follower does, so the leader plays U; the follower, knowing that, plays L; the equilibrium is UL, the leader receives 2, and iterated elimination reaches the same cell without ever considering beliefs.

Now let the leader announce first. Announcing pure U produces exactly the same outcome, 2, because the follower answers L as before. But announcing pure D changes the follower’s reply: against D the follower prefers R, 2 to 0, and the leader receives 3. Committing to the dominated action is worth more than playing the dominant one. The leader gains nothing from D directly — D is worse in every column — and gains a whole point by what the announcement does to the follower.

The mixture does better still. Against an announced probability pp of U, the follower’s L is worth pp and R is worth 2(1p)2(1-p), so the follower prefers R whenever pp is below 2/32/3. Against R the leader receives 4p+3(1p)=3+p4p + 3(1-p) = 3 + p, which rises with pp — so the leader wants pp as large as possible while still keeping the follower on R. That is p=2/3p = 2/3, where the leader receives 3+2/3=11/33 + 2/3 = 11/3.

The tie, and why it does not matter

At exactly p=2/3p = 2/3 the follower is indifferent between L and R, and the figure resolves the tie in the leader’s favour. That convention needs a sentence of defence, because the whole gain appears to rest on it.

It does not. At p=0.66p = 0.66 the follower strictly prefers R and the leader receives 3.663.66; at p=0.666p = 0.666 it receives 3.6663.666. Every announcement a little below 2/32/3 gets a follower who strictly prefers R, and the leader’s payoff approaches 11/311/3 as closely as wanted. The tie-breaking rule is the tidy statement of a limit, not a gift from the follower.

What the convention does do is make the optimum a single point rather than a supremum, and it matches how these problems are solved in practice: the leader announces a probability very slightly inside the region where the follower’s reply is the one it wants, and the difference is below any precision that matters.

The gain comes from the discontinuity. The follower’s reply jumps at 2/32/3 from R to L, and the leader’s payoff jumps with it from 11/311/3 down to 5/35/3. A leader who can place itself just on the favourable side of a jump in the other side’s behaviour is extracting value from the jump — and a simultaneous mover, who does not know which side of it the other believes it is on, cannot.

In a game of pure conflict, it is worth nothing

Now take a game in which every gain for one side is a loss for the other.

Moving first in a zero-sum game is worth nothing. The leader's payoff against the probability it announces for its first action, for matching pennies, where every gain is the other's loss, with the follower's reply switching where the follower is indifferent. The best announcement is worth 0; the best equilibrium of the simultaneous game is worth 0.
Fig. 3 Matching pennies: the leader wins 1 if the two coins match and loses 1 if they differ. The follower, seeing the announced probability of heads, always picks the side that mismatches, so the leader’s payoff is the lower of the two lines. Its peak is at a half, where it is worth 00 — exactly what the game’s only equilibrium is worth.

The shape is completely different. The follower’s interests are the leader’s reversed, so the follower’s best reply is always the column that is worst for the leader, and the thick line is the lower envelope of the two payoff lines. The best the leader can do is make the two lines equal, which is a half, and there the leader receives nought.

That is exactly the value of the game — the number both sides name when they move simultaneously.

The value of a 2×2 zero-sum game, named from both sides. The row chooser's expected payoff against each column as a line over the mixing probability, with the lower envelope and its maximum, beside the same construction from the column chooser's side. Both give 0.
Fig. 4 The same game drawn the way a simultaneous zero-sum game is solved: each chooser’s expected payoff against each of the other’s pure actions, the lower envelope, and its peak. From both sides the guarantee is 00, reached by mixing evenly.

The two pictures are the same picture. The commitment figure’s thick line is the lower envelope from the minimax construction, because in a zero-sum game a follower who best-replies is a follower who minimises the leader’s payoff. So the value of committing is the maximum of the lower envelope — the most the leader can guarantee — and the minimax theorem says that equals what the other side can hold the leader to when it commits instead.

Von Neumann’s theorem is a statement that order does not matter. Whoever announces first, the payoff is the value; announcing gains nothing and loses nothing. That is the content of the theorem restated as a claim about timing, and it is the one class of games in which the intuition that moving first is a handicap is not merely wrong but exactly balanced.

Never worse than an equilibrium

The two examples are the extremes, and there is a clean statement covering everything between them: a leader who commits to a mixture receives at least as much as in any equilibrium of the simultaneous game.

The argument is one line. Take any equilibrium, and let the leader announce its equilibrium mixture. The follower’s equilibrium reply is a best reply to that mixture, so the follower is willing to play it; and if the tie is broken in the leader’s favour, the follower plays something at least as good for the leader. So the leader can always reproduce any equilibrium’s payoff by announcement, and may do better. The statement has no empty cases, because a finite game in which mixtures are allowed always has an equilibrium to reproduce — which is also why the comparison can be made in every game at all, rather than only in those where somebody has found one.

Without the ability to announce a mixture the statement fails. Committing to a pure action can be worse than the simultaneous game — in matching pennies, announcing heads loses 1 for certain — and that is where the intuition about moving first as a handicap comes from. The handicap belongs to pure commitment. A mixed commitment keeps the randomness the other side cannot exploit and adds the information the other side can, and the information only ever helps the announcer.

In the stag hunt, it chooses

The third case is the one two equilibria and no way to choose left open, and commitment settles it.

Moving first decides which equilibrium happens. The leader's payoff against the probability it announces for its first action, for a joint effort worth more than a safe one, with the follower's reply switching where the follower is indifferent. The best announcement is worth 4; the best equilibrium of the simultaneous game is worth 4.
Fig. 5 The stag hunt with a leader. Announcing the joint effort for certain is worth 44, which is the better of the game’s two pure equilibria — the other is worth 33, and so is the mixed one. Moving first gains nothing over the better equilibrium and 11 over the worse, and it decides which one happens.

The leader announces that it will hunt stag, for certain. The follower, told that, prefers stag to hare, 4 to 3, and the joint effort happens. The value of moving first here is not a gain over the best equilibrium; it is the removal of the doubt that made the worse equilibrium’s basin three times the size of the better one’s.

Two equilibria, and two tests that disagree. The row chooser's expected payoff from each option against the column chooser's behaviour, for a joint effort worth more than a safe one. The lines cross at 0.750, which is the mixed equilibrium and the boundary between the two basins.
Fig. 6 The same game without a leader, as a chooser facing an unknown opponent sees it: the joint effort is the better reply only when the other is believed to commit with probability above 0.7500.750, which is a quarter of the possible beliefs.

Put the two figures side by side and the selection problem dissolves in a specific way. In the simultaneous game, a chooser needs to believe the other will commit with probability at least three-quarters, and nothing in the game supplies that belief. With a leader, the belief is supplied — it is certainty — by the one act the simultaneous game forbids. The announcement is not a signal of intent. It is a change in the order of moves, and the order of moves is what the equilibrium definition takes as given.

That also explains why the device of a public signal could reach the better outcome and could not guarantee it. A signal recommends; a commitment acts. The follower in the stag hunt is not asked to trust that the leader will do what it said; the leader has already done it.

A road closed on purpose

The same structure — a restriction imposed first beating freedom exercised simultaneously — is the resolution of the road that makes everyone later.

In that network, adding a link that costs nothing to use made every traveller strictly slower, because each traveller, choosing a route against everybody else’s, found the new link individually worthwhile, and the equilibrium they settled into was worse for all of them than the one without it. No traveller can escape by acting alone: a single driver who ignores the free link while everyone else uses it simply arrives later than they do.

A planner who closes the link before anybody chooses a route is a leader in exactly the sense of this essay. The planner moves first, publicly and bindingly; the travellers are followers who best-reply to the network they are given; and the equilibrium they reach on the smaller network is better for every one of them. Removing an option from everybody, in advance, does what no traveller could do by giving it up alone.

The difference from the games above is worth stating, because it is what makes the network example uncontroversial. There, the leader’s gain came partly at the follower’s expense — in the first game the follower receives less when the leader announces 2/32/3 than in the simultaneous equilibrium. Here, the planner’s payoff is the travellers’ own total delay, so the leader’s interest and the followers’ coincide, and commitment is a pure improvement. That is the case in which a leader is most obviously welcome, and it is the case in which the simultaneous game’s failure is most plainly a failure of timing rather than of anybody’s preferences.

What makes an announcement believable

Everything above assumes the announcement is binding, and that is the substantive assumption rather than a technicality.

A leader who could revise after seeing the follower’s reply would. In the first game, having announced 2/32/3 and seen the follower choose R, the leader would rather play U for certain, earning 4 instead of the 11/311/3 average. A follower who knows the leader can revise does not believe the announcement, and the whole construction collapses back to the simultaneous game. So the gain requires something that makes revision impossible or costly: a contract, a public randomising device, a reputation that repeated play would lose, or a physical arrangement that removes the choice.

Schelling’s work on the strategy of conflict is largely a catalogue of such arrangements — the general who burns the bridges behind the army, the negotiator who cannot be reached, the government that passes a law requiring it to retaliate — and the point of each is to make an announcement credible by destroying the option to go back on it. Reducing one’s own options can increase one’s payoff, which is the paradox the leader figure draws in numbers.

In the mixed case there is a further requirement: the randomisation itself must be verifiable, or at least not manipulable after the fact. Otherwise a leader who announced 2/32/3 could quietly play U more often, and the follower, anticipating that, would not reply as the figure assumes.

Where commitment is used

The mixed version of this idea runs security operations. When a limited number of patrols must cover more places than they can watch at once, a fixed schedule is exploitable — an adversary observes it and strikes where the patrol is not — and a secret schedule is unnecessary, because the adversary will observe it anyway over time. The defender is therefore a leader whether it likes it or not, and the right move is to choose the probabilities with which each place is covered, publish them in effect by using them, and randomise each day’s assignment from them.

Systems built on exactly this calculation have scheduled checkpoints and canine patrols at Los Angeles International Airport since 2007, air marshals on flights, and coast guard patrols in American ports. The optimisation they solve is the one in the figure — maximise the leader’s payoff over announced mixtures, given a follower who best-replies — with many more actions and a linear program in place of three candidate points.

The same structure appears in pricing, where a firm that posts a price is a leader and a customer who decides whether to buy is a follower, and in the posted thresholds of the prophet inequality, where a single announced price secures a guaranteed share of what an all-seeing allocator could collect.

Where the account needs care

The follower must know the payoffs and best-reply. A follower who is uncertain about what the leader will receive, or who replies imperfectly, gives a different optimum, and the knife-edge at 2/32/3 is replaced by a smooth trade-off. The clean gain of 5/35/3 belongs to a perfectly informed, perfectly rational follower.

Ties are broken for the leader. As argued above this is a limit rather than a gift, but in games where the follower’s indifference is not a single point — a follower indifferent over a range of announcements — the difference between breaking ties for and against the leader can be a genuine difference in the answer, and the literature distinguishes strong and weak versions of the solution for that reason.

Two leaders is not a leader. If both choosers can commit, the question becomes who commits first, which is itself a game with its own equilibria, and the advantage of leadership becomes something to race for.

And more actions need more than three candidates. In a two-by-two game the best announcement is at an endpoint or at the follower’s single point of indifference. With more actions the follower’s reply regions are polygons in a simplex, the leader’s payoff is linear on each, and the optimum is found by solving one linear program per follower action — which is how the security systems compute their schedules, and which duality makes tractable.

Von Stackelberg’s firms

The leader–follower structure is named after Heinrich von Stackelberg, whose 1934 book on market forms analysed two firms choosing output where one moves first. In his model the leader produces more than it would in the simultaneous game, the follower adjusts down, and the leader’s profit rises — a pure commitment that works because output, unlike a coin toss, is visible and costly to reverse.

The mixed version is much more recent. That a leader can always do at least as well as in any equilibrium by committing to a mixture was stated precisely by von Stengel and Zamir in the 2000s, along with examples like the one at the top of this page, where the commitment puts weight on a dominated action. The computational side — finding the optimal commitment in large games — came out of work on security scheduling in the same decade.

The two versions answer different questions. Von Stackelberg asked what a firm gains by being first; the modern question is what a defender gains by being predictable in its probabilities while unpredictable in its actions. The arithmetic is the same shape and the conclusions point in opposite directions: the firm gains by being aggressive, and the defender gains by being honest about how random it is.

What a line of payoffs cannot show

The figures draw expected payoffs, and a single play of an announced mixture produces one cell. A leader who announces 2/32/3 and then rolls U receives 4 or 2 depending on the follower’s reply, and nothing on the page is the outcome of a play. The claim is about the average, and the average is only what happens if the game is played often or the leader is risk-neutral.

The pictures also show a follower who never errs. A follower who replies R with probability 0.990.99 rather than certainty moves the leader’s best announcement away from 2/32/3, and no figure here shows how far.

And the enforcement is invisible. The difference between the leader figure and the simultaneous game is a single assumption — that the announcement binds — and the assumption has no picture. Everything the figures show is conditional on it, and in practice it is the part that is hard.

Still open: commitment in sequence

A single announcement against a single follower is the simplest case. When the game is repeated, reputation can substitute for a binding contract, and the question becomes which announcements a leader can make credible by the threat of losing future trust — the territory of the folk theorems, where a population’s dynamics and a pair’s reputations are different routes to the same selection problem. Leadership with several followers who also interact, and bargaining in which each side’s commitment is itself strategic, are the two directions in which the analysis above stops being a one-dimensional picture and becomes a problem of computation.

Giving away the option to deviate

The habit is about what an option is worth to the person who holds it.

In a single decision an extra option can only help: at worst it goes unused. Among choosers who anticipate each other, an extra option can hurt, because the others anticipate its use. The leader in the first game is better off without the ability to change its mind, and the defender scheduling patrols is better off publishing its probabilities than keeping them secret.

When others respond to what a chooser can do, removing an option is a move, and sometimes the best one available. The test is to ask what the other side would do if they knew the option were gone — and whether that change is worth more than the option.