The best average is the worst bet
Worth reading first: A wait that ends and has no average · A growth rate no step contains.
A coin lands heads with probability 0.6. At each toss a bettor may stake any part of their wealth on heads at even money: win and the stake is doubled, lose and it is gone. The offer is a thousand tosses. How much should be staked each time?
The arithmetic of expectation gives a clear answer. A stake of wins with probability 0.6 and loses with probability 0.4, so its expected gain is — positive, and larger the larger the stake. To maximise expected wealth, bet everything on every toss. After tosses the expected wealth is , which after a thousand tosses is a number with eighty digits.
The bettor who stakes everything doubles their wealth at each head and loses all of it at the first tail. On the sequence in the figure the first tail comes at the tenth toss; the chance that it has not come by the thousandth is , about . Every bettor who follows the strategy that maximises the expectation is ruined, with probability one, and the expectation is still . All of it sits on the single sequence of a thousand heads.
This is the sharpest form of a lesson that runs through the study of expectation: an average can be infinite while the typical value is small, and here an average grows without bound while the typical value is nought. The right quantity to maximise, for a bettor who bets repeatedly, is not the expected wealth but the expected logarithm of wealth, and the fraction that does it — John Kelly’s of 1956 — is , a fifth of current wealth.
Wealth is a product, so its logarithm is a sum
Bet a fixed fraction of current wealth on every toss. A head multiplies wealth by , a tail by , and after tosses with heads the wealth is
That is a product of independent random factors, and its logarithm is a sum of independent random terms. The law of large numbers applies to the sum: converges, with probability one, to the expected value of a single term,
So almost every run grows like , and — the expected logarithm, not the logarithm of the expectation — is the growth rate the bettor actually experiences. It is the same mechanism as in a sequence that adds or subtracts by the toss of a coin: products of random factors grow at a rate set by the average of their logarithms, and that rate is contained in no single step.
The growth rate is zero at , rises to a maximum, and falls. Setting its derivative to nought, , gives , the edge itself: bet the fraction of wealth equal to the amount by which the chance of winning exceeds the chance of losing. At that is 0.2, and the growth rate there is 0.0201 a toss — wealth multiplies by about 1.0203 per toss in the long run, doubling roughly every thirty-five tosses.
Beyond the maximum the growth falls quickly. At it is zero, and above that it is negative: a bettor who has the edge on every single toss, and bets more than 39 per cent of their wealth each time, sees that wealth shrink towards nothing with probability one. The realised rates over twenty thousand tosses lie on the curve within a few thousandths. Nothing about this depends on luck; it is the law of large numbers applied to logarithms.
The average and the typical part company
The expected wealth and the typical wealth are both exact functions of the fraction, and they disagree about everything.
The expectation is , increasing in : after 200 tosses it is about at , at and at Kelly’s 0.2. The median, computed exactly from the median number of heads, runs the other way. At Kelly’s fraction it is about after 200 tosses, at it has fallen to , and at it is nought from the second toss on.
The gap is the one the envelope paradox exploits and St Petersburg’s game made famous. An expectation over a product is dominated by its luckiest outcomes, because a product amplifies them. At , a run of heads multiplies wealth by 1.6 a toss and a run of tails by 0.4; the expectation counts the long runs of heads at their full multiplied value, and they are rare enough that the typical run never sees one.
Which one to maximise is not a matter of mathematics alone. A bettor offered a single toss, or one who is risk-neutral about a small fraction of their wealth, may reasonably maximise the expectation. A bettor who will bet again and again with the proceeds, and cares about where they will be rather than about an average over imaginary copies of themselves, is maximising the typical outcome, and that is the logarithm.
Kelly is eventually ahead of every rival
The growth rate is a statement about the limit. Over a finite run, a bolder fraction can be ahead.
Run Kelly against a rival on the same tosses, and record how often Kelly is richer. After ten tosses it is ahead on about 62 to 65 per cent of runs — only a little more than half, because a bolder bettor gains more from a lucky start and a more timid one loses less from an unlucky one. After a hundred tosses it is ahead on 68 to 75 per cent; after a thousand, on 94 to 99 per cent; after five thousand, on every run in the sample.
Leo Breiman proved in 1961 that this is general: for any strategy that is essentially different from Kelly’s — whose long-run growth rate is lower — the ratio of Kelly’s wealth to the rival’s tends to infinity with probability one, and Kelly’s strategy reaches any fixed target wealth in the shortest expected time. The rivals closest to Kelly take longest to fall behind, because their growth rates are closest: 0.1 and 0.3 sit symmetrically on either side of 0.2 and are overtaken at the same pace, 0.35 a little faster.
Why the pace is so slow follows from the growth rates. Over tosses the difference between the logarithms of Kelly’s wealth and a rival’s is a sum of independent terms with mean and spread of order . Against the rival at 0.3 the mean difference is about 0.0054 a toss and the spread of a single term about 0.105, so the accumulated mean exceeds twice the accumulated spread only after about 1,500 tosses — which is where the figure’s curves close in on one. A rival that is close to Kelly is beaten in the end with certainty, but the end can be very far off, and over any horizon a person actually bets for, a fraction near Kelly’s is as good as Kelly’s.
That is the sense in which Kelly’s fraction is best. It does not maximise any finite-horizon expectation, and it is not ahead on every run. It is the strategy that, given enough tosses, is ahead of any other with probability approaching one.
The price is the size of the dips
Kelly’s strategy grows fastest, and it is a rough ride.
Over 2,000 runs with a smaller edge, and Kelly’s fraction 0.1, wealth fell to half its starting value at some point on 49 per cent of runs, to a fifth on 20 per cent and to a tenth on 9.7 per cent. The chance of ever falling to a fraction of the start is about , and that is a theorem in the continuous-time limit, independent of the size of the edge: a Kelly bettor falls to half their stake about half the time.
Halving the fraction changes this dramatically. Half Kelly’s growth rate is three quarters of Kelly’s — the growth curve is flat at its top, so a small step back from the maximum costs little — but its chance of ever falling to is about : 11.5 per cent of runs fell to half, 0.55 per cent to a fifth, and one run in two thousand to a tenth. In general, staking times Kelly’s fraction gives a growth rate of times Kelly’s and a chance of about of falling to . Twice Kelly has a growth rate of nought, and a chance of one of eventually falling as far as anyone cares to name.
This is why people who use the rule in practice — Edward Thorp at the blackjack table and later in his fund, and many after him — tend to bet a fraction of Kelly. The other reason is that the edge is never known exactly. Overestimate and the fraction is too large, and the growth curve falls faster on the high side than on the low; betting a fraction of the estimated Kelly stake is insurance against the estimate.
A growth rate that is an information rate
Kelly did not publish his rule as a betting system. His paper of 1956 was called A new interpretation of information rate, and it was about a gambler receiving tips over a noisy channel.
At Kelly’s fraction the growth rate is
where is the coin’s entropy. Read the coin as a tip that is right with probability , and is exactly the capacity of the channel carrying it — the rate at which the tip conveys information about the outcome. A perfect tip, , conveys per toss and doubles the wealth each time; a useless one, , conveys nothing and earns nothing.
For a small edge the growth rate is about half the square of the edge: at it is 0.0050 a toss, and at — a one per cent edge — it is 0.0002. A gambler with a one per cent edge, betting optimally, doubles their money in about 3,500 tosses. Evidence measured in decibans adds up the same way, as logarithms of likelihood ratios, and that is not a coincidence. With several outcomes and any odds, Kelly’s rule says to divide wealth among the outcomes in proportion to one’s own probabilities for them, and the growth rate is then the amount by which one’s probabilities beat the odds the house has posted, measured in the same logarithmic units as evidence — so a bettor’s growth is a score of how much better their model of the coin is than the house’s.
Unequal odds, and a demon that profits from noise
The even-money coin is the simplest case of a general rule. If a bet pays to 1 and wins with probability , the growth rate of staking a fraction is , and its maximum is at
the expected gain per unit staked divided by the payout — edge over odds, as the rule is usually remembered. A bet with no edge gets nothing; a long shot with a small edge gets a small stake even though its expected gain per unit is large, because the payout that makes it attractive also makes the losses frequent. When several outcomes are available at once — a horse race — the rule becomes: spread the wealth across the horses in proportion to one’s own probabilities for them, and keep back whatever the track’s take makes unprofitable.
The same calculation produces a result that looks like a paradox. Take an asset that, each period, either doubles or halves with equal probability. Its expected value grows by a quarter a period, but its median goes nowhere: the growth rate of holding it is . Now hold half of one’s wealth in the asset and half in cash, and rebalance to half and half after every period. Each period the wealth is multiplied by or , and the growth rate is — positive, from combining an asset that does not grow with cash that does not grow. Claude Shannon is said to have lectured on it, and it is sometimes called Shannon’s demon. It is the same arithmetic as two losing games that win when alternated: a fixed fraction sells after rises and buys after falls, and that converts volatility into growth whenever the logarithm, rather than the mean, is what compounds.
Ruin, and what ruin is not
A fixed fraction can never be completely ruined: is positive, so wealth stays positive after any sequence. The ruin in this essay is wealth tending to nought, not reaching it. That is different from the classical gambler’s ruin, where a bettor staking a fixed amount hits nought in finite time — and for a fixed amount on a favourable coin, the ruin probability starting from capital is , positive however large the capital.
Kelly’s proportional betting escapes that by shrinking the stake as the wealth shrinks. It cannot be wiped out, and it pays for the safety by betting less after losses, which is also what makes its growth rate a property of the logarithm. The all-in bettor of the first figure is the one case in which a proportional bettor can hit nought, because .
What the simulations show and what they assume
Every figure here is for the simplest case — a single even-money bet repeated with a known probability. Kelly’s rule extends to unequal odds, where the fraction is the edge divided by the odds, to several simultaneous bets, where it becomes a concave optimisation over portfolios, and to continuous time, where it becomes the growth-optimal portfolio of a stock and a bond; in each case it maximises the expected logarithm. The drawdown law is exact for the continuous-time version, and the discrete runs in the figure match it to within sampling error.
The thing the figures cannot settle is whether maximising the logarithm is what a particular person should do. Paul Samuelson argued for decades that it is not, that a bettor with a different attitude to risk should maximise a different function of wealth, and he was right that nothing forces the logarithm on someone who cares about a fixed horizon. What the figures do show is what the logarithm buys — the fastest growth of the typical outcome, eventual dominance of every other fixed fraction — and what the expectation costs.
Still open: how much to trust an estimate
How should a bettor size their stakes when the probability is estimated rather than known? Betting the Kelly fraction for an estimated edge overbets whenever the estimate is high, and the growth curve punishes overbetting more than underbetting. Fractional Kelly is the practical answer, but which fraction is optimal depends on the estimate’s uncertainty in a way that has precise answers only in special models; for a general sequence of bets with estimated, changing edges, the right generalisation of Kelly’s rule — one that is optimal for the outcomes a bettor actually faces rather than for a model of them — is not settled, and the theory of universal portfolios, which competes with the best fixed portfolio in hindsight, gives guarantees only up to a factor polynomial in the number of rounds.
Maximise the logarithm
When wealth is reinvested, it is a product of random factors, and its logarithm is a sum that the law of large numbers controls. The fraction that maximises the expected logarithm, Kelly’s , grows the typical bettor faster than any other and is eventually ahead of every rival, while the strategy that maximises the expectation — stake everything — is ruined with certainty. The growth it buys is , an information rate, and its price is a fall to half the starting wealth about half the time, which half Kelly reduces to an eighth.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The heuristic that cannot be a proof — both name expectation, logarithm, random walk
- A threshold no average can see — both name expectation, random walk
- A walk that beats trying everything — both name expectation, random walk
- Almost none of the roots are real — both name expectation, logarithm
- Fair bits from an unfair coin — both name entropy, expectation
- No single input can move it far — both name expectation, random walk
Named objects
A dashed tag is an object no other essay names yet.
EntropyExpectationGamblers ruinLaw of large numbersLogarithmMedianRandom walk