Dynamics

A dimension from the stretching rates

An attractor has no construction rule, so its dimension has to be counted — that was the rung below's argument for defining dimension by counting at all. Kaplan and Yorke's formula computes it instead, from two numbers that describe the map and never look at the set.

Worth reading first: A dimension that is not a whole number · A carpet with two dimensions.

The rung below argued that the reason to define dimension by counting is that counting works on sets with no rule behind them — the attractor of a map being the standing example, since nothing says how many copies of itself it holds or at what scale.

Kaplan and Yorke’s formula computes the dimension anyway, from two numbers that are properties of the map.

A dimension of 1.2576, from two stretching rates. The running averages of the Hénon map's two Lyapunov exponents, settling at 0.4177 and -1.6217. Kaplan and Yorke's formula turns them into a dimension of 1.2576 without counting a single box.
Fig. 1 The running averages of the Hénon map’s two Lyapunov exponents over two hundred thousand steps, settling at 0.41890.4189 and 1.6229-1.6229. Their sum is 1.2040-1.2040, which is the logarithm of the map’s area factor exactly — a check the arithmetic has to pass. The formula then gives a dimension of 1.25811.2581, with no boxes counted anywhere.

The counted value for the same attractor is about 1.241.24. The two agree to about the accuracy either can claim, and neither is a computation of the other.

What the two numbers are

Start two orbits at nearby points and watch the gap between them. It grows — and the growth is what a difference too small to draw becomes — and how fast it grows is the Lyapunov exponent — the average rate at which the logarithm of the separation increases per step.

In two dimensions there are two such rates, because a small disc of starting points is carried to a small ellipse, and an ellipse has two axes. The larger exponent λ1\lambda_1 is the rate along the direction that stretches most; the smaller λ2\lambda_2 is the rate along the direction that is left over once the first is accounted for.

Computing them means carrying a pair of vectors along the orbit under the map’s derivative and re-orthogonalising at every step, keeping a running total of how much each was scaled. That is Gram–Schmidt run inside a dynamical system, and it is necessary because both vectors would otherwise swing round onto the most-stretched direction and report the same exponent twice.

The check the figure makes is on the sum. The map’s derivative has determinant b-b, exactly, at every point — the Hénon map contracts area by the same factor everywhere — so the two exponents must add to logb\log|b|. That is a number known in advance, it is log0.3=1.2040\log 0.3 = -1.2040, and the measured sum agrees to four decimal places.

That check has real teeth. A Lyapunov exponent computed wrongly is a plausible number with nothing to compare it against, and the sum is the only quantity in the calculation that is known independently.

The formula

D=1+λ1λ2D = 1 + \frac{\lambda_1}{|\lambda_2|}

when λ1>0>λ2\lambda_1 > 0 > \lambda_2 and λ1+λ2<0\lambda_1 + \lambda_2 < 0, which is the situation for the Hénon map and for every dissipative map with one unstable direction.

The reasoning behind it is a covering argument and it is short. Follow a small square of starting points for kk steps. It becomes a long thin strip: stretched by ekλ1e^{k\lambda_1} along one direction and squeezed by ekλ2e^{k\lambda_2} along the other. Covering that strip with squares of side ekλ2e^{k\lambda_2} — the strip’s width — takes about ek(λ1λ2)e^{k(\lambda_1 - \lambda_2)} of them.

So at scale ε=ekλ2\varepsilon = e^{k\lambda_2} the count is ε(λ1λ2)/λ2\varepsilon^{-(\lambda_1-\lambda_2)/|\lambda_2|}, and the exponent is

λ1λ2λ2=λ1λ2+1,\frac{\lambda_1 - \lambda_2}{|\lambda_2|} = \frac{\lambda_1}{|\lambda_2|} + 1,

which is the formula. The dimension is one for the direction that survives, plus a fraction for how much of the stretched direction the squeezing fails to eliminate.

Read that way the formula is not mysterious at all. It says the attractor is a curve — dimension one — thickened by however much the stretching outruns the contraction, and the thickening is measured by a ratio of rates.

The two limiting cases confirm the reading. If the contraction is overwhelming — λ2|\lambda_2| enormous beside λ1\lambda_1 — the ratio goes to nought and the dimension goes to one, which says an attractor whose transverse direction collapses instantly is a curve. If the two rates approach each other, the ratio goes to one and the dimension to two, which says a map that barely contracts leaves something area-filling. Every value between one and two is achieved by some ratio, and the formula’s whole content is that the intermediate cases interpolate the way the covering argument says.

It also makes the self-affine structure of the previous rung explicit. A blob carried kk steps is a rectangle ekλ1e^{k\lambda_1} long and ekλ2e^{k\lambda_2} wide, so its aspect ratio is ek(λ1λ2)e^{k(\lambda_1-\lambda_2)} and grows without bound — exactly the situation the carpet was built to exhibit, arriving here from a map rather than from a grid. The attractor is a self-affine object whose two contraction rates are eλ1e^{\lambda_1} and eλ2e^{\lambda_2}, and everything the carpets showed about such objects applies, including that its box and Hausdorff dimensions have no reason to agree.

the Hénon attractor, after 5 steps. the Hénon attractor drawn from its own rule: the orbit of x ↦ 1 − 1.4x² + y, y ↦ 0.3x, which obeys no self-similar rule.
Fig. 2 Thirty thousand points of one orbit. The bands are the stretched-and-folded strip after many steps, and the fine structure between them is what the ratio of the two exponents is measuring.
Box counts against box size, on logarithmic axes. For each set, the logarithm of the number of occupied boxes plotted against the logarithm of one over the box size, with a straight line fitted and its slope reported.
Fig. 3 The counting method on the same attractor, beside a set with a rule. The attractor’s counts give about 1.241.24 over five box sizes, and the range cannot be widened — a finer grid counts boxes the sampled orbit happened to miss, and the estimate would fall for a reason that is about the sample rather than about the set.

What it is not

The formula’s status is worth being precise about, because it is usually quoted as though it were a theorem and it is a conjecture.

Kaplan and Yorke proposed it in 1979 on exactly the covering argument above, and the argument is a heuristic: it treats the stretching and squeezing as if they happened at uniform rates, when the exponents are only long-run averages and the local rates vary from point to point along the orbit.

What has been proved is an inequality. Ledrappier and Young, and others, showed that the formula is an upper bound for the dimension of the attractor — the true value can be smaller, and it is smaller for maps constructed to make it so. Equality is known for some families and is expected for typical ones, and there is no general theorem.

So the agreement in the numbers at the top of this page is evidence and not confirmation. 1.25811.2581 from the formula against about 1.241.24 from counting is consistent with equality and equally consistent with a small genuine gap, since the counted value is itself an estimate over five box sizes.

Which dimension it computes

There is a further imprecision worth separating out, because there are now three notions in play and the formula is about a fourth.

The formula estimates the dimension of the natural measure on the attractor — the distribution an orbit spends its time in — rather than the dimension of the attractor as a set. Those differ. A measure’s dimension asks how the measure of a small ball scales; a set’s asks how many balls are needed to cover it. A set can be larger than the part of it that carries the measure.

In practice the distinction is invisible and in principle it is not, and this is the case where it matters: an attractor may contain points an orbit visits with vanishing frequency, and box counting sees them while the measure does not.

That is also why the counted value in the rung below and the formula’s value here can differ without either being wrong. The counting sampled a hundred and twenty thousand points of one orbit, so it saw the parts the measure lives on — which is the measure’s dimension, approached from a set-counting method that had no way of finding the rest.

Infinite below 1.2619, nought above it. The total of the s-th powers of the diameters in the natural cover of the Koch curve, plotted against s for 4 depths. Every curve passes through one at s = 1.2619 and they separate either side of it.
Fig. 4 The measure-based definition on a set with a rule, where all the notions coincide. The whole difficulty of the attractor is that none of this machinery is available: there is no cover to write down and no exponent at which anything pivots that anybody can compute.

Why the method exists at all

Both halves of the calculation are worth putting against the alternative, because the case for the formula is a practical one.

Counting boxes on an attractor needs an orbit. A long one, since the count at a fine scale is an undercount whenever a box the attractor meets contains no sampled point, and the undercount grows as the boxes shrink. The rung below stopped at the scale where that would start to bite, and the range of scales left was five or six box sizes wide — narrow for a definition about a limit.

Computing exponents needs an orbit too, and uses it differently. Every step contributes to the running average, so a long orbit gives a better estimate rather than a finer scale, and the convergence is visible in the figure: the two curves settle and stay settled. There is no smallest scale to run out of.

And the exponents are useful for other reasons anyway. They are what says whether a system is chaotic at all, they set the horizon past which a prediction is worthless, and they are computed as a matter of course. The dimension arriving as a by-product costs nothing.

There is a fourth difference and it is the one that decides the matter in higher dimensions. Box counting’s cost grows with the dimension of the space and the exponents’ does not. Covering a set in ten dimensions with boxes of side ε\varepsilon means a grid with ε10\varepsilon^{-10} cells, and the number of sampled points needed to occupy a useful fraction of the occupied ones grows the same way; the estimate is unusable past three or four dimensions and hopeless past six. Carrying ten vectors along an orbit and re-orthogonalising them costs a hundred multiplications a step, whatever the attractor looks like.

So for the systems anybody cares about — flows in many variables, discretisations of partial differential equations, where the state space has thousands of dimensions and the attractor is expected to have a modest one — counting is not merely worse, it is unavailable, and the formula is the only method there is. Its status as a conjecture is a great deal easier to live with in that light: the alternative is not a rigorous number but no number.

The Lorenz attractor is the standing example, at about 2.062.06, and that figure has never been produced by counting.

A dimension of 1.2033, from two stretching rates. The running averages of the Hénon map's two Lyapunov exponents, settling at 0.3073 and -1.5112. Kaplan and Yorke's formula turns them into a dimension of 1.2033 without counting a single box.
Fig. 5 The same map at a different parameter. Both exponents move, their sum is still exactly the logarithm of the area factor because that does not depend on the parameter, and the dimension comes out at 1.20331.2033 instead.

That the sum is unchanged while both exponents move is the sharpest thing the figures show. The area contraction is fixed by the map’s form and the split between the two directions is not, so the dimension varies with the parameter while the constraint on the pair does not.

What the two exponents say apart from the dimension

The pair is worth reading on its own, because between them they describe the whole of what the map does to a small blob of starting points and the dimension is only one thing they encode.

The positive one is a horizon. An uncertainty δ\delta in the starting point grows like δekλ1\delta e^{k\lambda_1}, so it reaches size one after about log(1/δ)/λ1\log(1/\delta)/\lambda_1 steps. At λ1=0.4189\lambda_1 = 0.4189 that is about 2.42.4 steps per factor of ee — so knowing the start to one part in a million buys about thirty-three steps of prediction, and knowing it to one part in 101510^{15} buys about eighty-two. Each extra decimal place of the initial measurement buys the same fixed number of steps, which is the arithmetic behind a closer start buying only time.

The sum is a volume. Two exponents adding to log0.3\log 0.3 says a blob’s area shrinks by a factor of 0.30.3 per step, without exception and without averaging, because the determinant of this map’s derivative is constant. After thirty steps a region has 0.3300.3^{30} of its area, which is about 2×10162 \times 10^{-16} — so an attractor that started as a filled region has no area at all in the limit, and the set the orbit visits must be something thinner.

And the two together say what kind of thinner. Area going to nought forbids dimension two; the positive exponent forbids dimension one, since a curve carried by a map with a stretching direction does not stay a curve of finite length. So the dimension is strictly between, before any formula is quoted, and the formula’s job is only to say where.

That is worth separating out because it is the part that is certain. That the dimension lies strictly between one and two follows from the signs of the two exponents alone, and needs neither the covering argument nor the conjecture.

Where the account needs care

The exponents must be ordered and the signs must be right. The formula’s general form finds the largest jj for which λ1++λj0\lambda_1 + \cdots + \lambda_j \ge 0 and interpolates; the two-dimensional version above is the case j=1j = 1, and it needs λ1>0\lambda_1 > 0 and λ1+λ2<0\lambda_1 + \lambda_2 < 0. A map with two positive exponents, or two negative ones, needs a different case of the formula and gets a different kind of answer.

The averages are over one orbit. They are the right numbers only if the orbit is typical for the natural measure, which holds for almost every starting point with respect to that measure and can fail for a badly chosen one. Discarding an initial run of steps, as the figures do, is the standard precaution and is not a proof.

Re-orthogonalisation is not optional. Without it both carried vectors align with the most-stretched direction within a few dozen steps and the second exponent comes out equal to the first, which is a wrong answer with no symptom.

And the parameter must be one with an attractor. The Hénon map at some parameters has orbits that escape to infinity, and the calculation then reports the exponents of a divergence rather than of an attractor. The figure’s guard is on the parameter range.

Kaplan, Yorke, and a conjecture that stayed one

Kaplan and Yorke published the formula in 1979, in a paper about chaotic behaviour in difference equations, and stated it as a conjecture. Forty-odd years later that is still its status in general, which is unusual for a formula this widely used.

The reason the situation is stable is that the inequality half is enough for what the formula is used for. A practitioner wanting to know whether an attractor is nearly two-dimensional or nearly one-dimensional gets an answer that is right to the precision the data supports, and whether the last decimal is an upper bound or an equality does not change any conclusion.

A conjecture that is an inequality in the useful direction is not a gap anybody is troubled by, which is worth registering as a genuine mode of mathematical life rather than a failure. The situation is the reverse of the dissection piece counts, where the missing half is the useful one.

What the pictures cannot show

The exponents are limits and the figure draws running averages over two hundred thousand steps. That the curves have flattened is evidence the limits exist and is not a proof, and a map with slow convergence would produce a figure that looked exactly as settled at a value that was still moving.

The attractor’s dimension is not drawn at all. What is drawn is two numbers and an arithmetic combination of them, and the connection to the set in the orbit figure is the covering argument, which is prose.

And the sum check — the one assertion with real force — is a check on the arithmetic rather than on the formula. Two exponents adding to the right total says they were computed correctly; whether 1+λ1/λ21 + \lambda_1/|\lambda_2| is the dimension is the conjecture, and no figure here bears on it.

The ladder from here

Rungs above: the Ledrappier–Young formula, which expresses the measure’s dimension exactly in terms of the exponents and the entropy and is a theorem rather than a conjecture. Multifractal analysis, where the local scaling rate varies across the attractor and a whole spectrum of dimensions replaces the single number. The dimension of the Lorenz attractor, which is about 2.062.06 and is known only by these methods. Pesin’s formula, relating the entropy to the positive exponents. And the numerical difficulty of computing exponents for a flow rather than a map, where the zero exponent along the trajectory has to be found and discarded.

Measuring the machine instead of the output

The move is the one worth carrying and it is the opposite of the rung below’s.

That rung argued that a definition working only on sets built by a rule is a definition of the rule, and that the reason to count boxes is that counting can be done on anything. This rung goes back to the machine — but to a different property of it. Not the construction rule, which the attractor does not have, but the stretching rates, which every map has whether or not it builds anything.

The general form is: when the object resists measurement, measure the process that produces it. The attractor is hard to cover and easy to generate; the exponents are properties of the generating step and are estimated by running it. Everything difficult about the set has been traded for something easy about the map.

The trade has a cost and it is the one the middle of this essay is about. The set and the process are not the same object, so a quantity computed from the process is a statement about the process — and turning it into a statement about the set is where the conjecture lives. A measurement of the machine is evidence about the output and never identical with it, and every method of this shape carries the same gap.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A dashed tag is an object no other essay names yet.

AttractorBox dimensionDissipationGram schmidtIterated function systemLyapunov exponentScalingSelf affinitySensitive dependence