Dynamics

A dimension for every rate of crowding

Spread a unit of mass over an interval by splitting it unevenly, again and again, and the result covers the whole interval while crowding almost all of its weight onto a set of smaller dimension. Every rate of crowding picks out its own set of points with its own dimension, and the whole family is read off one curve.

Worth reading first: Infinite on one side and nought on the other · A dimension from the stretching rates.

Kaplan and Yorke’s formula ended on a distinction that most of the time is invisible: the dimension of an attractor as a set is one number, and the dimension of the distribution an orbit spends its time in is another. A set can be larger than the part of it that carries the weight.

The simplest object on which that distinction becomes a whole subject is not an attractor at all. It is a unit of mass spread over an interval by a rule that splits it unevenly.

A mass split 10 times, 0.3 to the left and 0.7 to the right. A self-similar measure on the unit interval: the mass is split 10 times, 0.3 of each piece's share going left and 0.7 right, and each of the 1024 pieces is drawn as a bar as tall as its mass. The heaviest carries 0.0282 and the lightest 5.9 × 10⁻⁶.
Fig. 1 A unit of mass on the interval, split in half ten times, with the left piece of every split taking 0.3 of what it is given and the right 0.7. Each of the 1,024 pieces is drawn as tall as its mass. The heaviest, at the right-hand end, carries 0.0282 and the lightest, at the left, 0.0000059 — 4,784 times less.

Every point of the interval gets some mass, so the set the mass lives on is the whole interval, of dimension one. But the mass is not spread evenly across it, and it is not spread unevenly in any simple way either: the picture has spikes inside spikes at every scale. The way to describe such a thing is not one dimension but a curve of them — one for each rate at which mass crowds into a shrinking neighbourhood — and the curve is computable exactly.

How fast the mass in a small interval vanishes

Take a point xx and the piece of length 2n2^{-n} that contains it after nn splits. The piece’s mass is 0.3j×0.7nj0.3^j \times 0.7^{n-j}, where jj is the number of times the point fell in a left half, and jj is simply the number of noughts among the first nn binary digits of xx.

For an evenly spread mass the piece would carry 2n2^{-n}, its own length. Here it carries its length raised to a power:

mass=lengthα,α=jln(1/0.3)+(nj)ln(1/0.7)nln2.\text{mass} = \text{length}^{\alpha}, \qquad \alpha = \frac{j\ln(1/0.3) + (n-j)\ln(1/0.7)}{n \ln 2}.

The exponent α\alpha depends only on the proportion of noughts in the point’s binary expansion. A point whose digits are all noughts sits where the mass is thinnest, with α=ln(1/0.3)/ln2=1.737\alpha = \ln(1/0.3)/\ln 2 = 1.737 — its neighbourhoods lose mass faster than their length. A point whose digits are all ones sits where it is thickest, with α=ln(1/0.7)/ln2=0.515\alpha = \ln(1/0.7)/\ln 2 = 0.515 — its neighbourhoods keep mass longer than length would allow. Every exponent between is the exponent of some proportion.

That turns the question into one about digits, and digits have been counted before. The average of many coin tosses settles: almost every number, picked at random along the interval, has noughts and ones in equal proportion, and so almost every point in the sense of length has α=1.126\alpha = 1.126. But almost every point in the sense of the mass is different. A point picked according to the mass falls left with chance 0.3 at every split, so its proportion of noughts settles at 0.3, and its exponent at

0.3ln(1/0.3)+0.7ln(1/0.7)ln2=0.881.\frac{0.3\ln(1/0.3) + 0.7\ln(1/0.7)}{\ln 2} = 0.881.

The two notions of “typical” disagree, and each is certain in its own sense. The mass is spread over every point of the interval and concentrated, all but a vanishing fraction of it, on points whose digits are 30 per cent noughts — and those points, as the next sections show, form a set of dimension 0.881.

Counting the pieces that share a mass

After nn splits there are n+1n+1 possible masses, and the number of pieces carrying the mass with jj left choices is (nj)\binom{n}{j} — the number of ways to place jj noughts among nn digits. The figure checks that count at every one of its 1,024 pieces.

That is enough to define a dimension for each exponent. Pieces of length ε=2n\varepsilon = 2^{-n} that share one exponent number (nj)\binom{n}{j}, and a set that needs about εf\varepsilon^{-f} boxes of side ε\varepsilon has dimension ff. So the points whose neighbourhoods scale with exponent α\alpha make up a set whose dimension, at finite nn, is

fn(α)=ln(nj)nln2.f_n(\alpha) = \frac{\ln \binom{n}{j}}{n \ln 2}.

As nn grows that approaches the entropy of the proportion, tlnt(1t)ln(1t)-t\ln t - (1-t)\ln(1-t) with t=j/nt = j/n, divided by ln2\ln 2. The limit is exact, and it is not only a box count. Abram Besicovitch in 1934 and Harold Eggleston in 1949 proved that the set of numbers whose binary digits have noughts in proportion tt has Hausdorff dimension equal to that entropy over ln2\ln 2 — so the dimension of each level set is the one the finest covers see, not just the one a grid sees.

A spectrum of dimensions from 0.51 to 1.74. The multifractal spectrum of a self-similar measure on the unit interval with shares 0.3 and 0.7: the dimension f(α) of the set of points where the mass scales with exponent α. Its top is 1.0000 and it touches the diagonal at 0.8813.
Fig. 2 The spectrum of the measure: for each exponent α\alpha from 0.515 to 1.737, the dimension f(α)f(\alpha) of the points where the mass scales like that. Its top, 1 at α=1.126\alpha = 1.126, is the whole interval; it touches the dashed diagonal at 0.881, the dimension of the set that carries the mass. The dots are the counts after 10, 40 and 160 splits, which sit below the curve by at most 0.202, 0.075 and 0.025.

Two points on that curve say most of what the section above said in words. The top of the curve is the typical point for length: the most numerous exponent, 1.126, belongs to a set of full dimension one, because almost every number has half its digits noughts. The point where the curve touches the diagonal is the typical point for mass: exponent 0.881 and dimension 0.881. Everywhere else the curve lies strictly below the diagonal, and that is not a coincidence of this example. A set of dimension ff can hold mass that scales like εα\varepsilon^\alpha per box only if fαf \le \alpha; where the two are equal, the set is exactly big enough to carry all the mass.

The finite counts approach the curve from below, as they must — a binomial coefficient is smaller than its entropy estimate by a factor that grows like n\sqrt n — and the gap shrinks roughly like lnn/n\ln n / n.

Almost all the mass on almost none of the pieces

The distance between the top of the curve and its point of contact with the diagonal can be turned into a count, and the count is startling.

After ten splits, the 120 pieces with exactly three left choices — about one piece in nine — carry 120×0.33×0.77=0.267120 \times 0.3^3 \times 0.7^7 = 0.267 of the mass, more than any other group. As the number of splits grows the groups near a proportion of 0.3 take more and more of the mass, because the proportion of left choices in a mass-weighted piece obeys the law of large numbers, and the groups away from it take less. After a thousand splits the pieces whose proportion lies within a few per cent of 0.3 carry all but a sliver of the mass. How many of them are there? About 20.881×10002^{0.881 \times 1000}, out of 210002^{1000} pieces in all.

The mass-carrying pieces are a fraction 21192^{-119} of the pieces, and that fraction goes to nought exponentially as the splitting continues. So the measure is spread over the whole interval in the sense that no piece is empty, and it lives on almost none of the interval in the sense that matters for where it actually is. The dimension 0.881 is exactly the exponent of that shrinking fraction: the carrying set needs ε0.881\varepsilon^{-0.881} boxes where the whole interval needs ε1\varepsilon^{-1}.

This is the same count that underlies data compression. A long string of symbols from a source that emits one symbol with probability 0.3 and the other with 0.7 is, with overwhelming likelihood, one of about 20.881n2^{0.881 n} typical strings rather than one of all 2n2^n, so it can be written down in 0.881 bits per symbol. The information dimension of the measure and the entropy of the source are the same number for the same reason: both count the strings that actually occur. The name “information dimension” is that coincidence, recorded.

Measuring it without knowing the rule

The curve above was computed from the construction. An orbit or an experiment does not come with one, and the practical question is how to get the spectrum from samples.

The method is box counting with the boxes weighted. Cover the samples with boxes of side ε\varepsilon, record the share of the points in each box, and add up those shares raised to a power qq. At q=0q = 0 every occupied box counts once, which is ordinary box counting. At q=1q = 1 the shares add to one whatever ε\varepsilon is. At q=2q = 2 the sum is the chance that two points picked independently land in the same box. Large positive qq is dominated by the heaviest boxes and negative qq by the lightest.

Mass-weighted box counts for 4 exponents, against the formula. Points sampled from a lopsided self-similar measure, binned at shrinking box sizes, with the logarithm of the sum of each box's share raised to the power q plotted for q = −1, 0, 2, 4. The fitted slopes are 2.254, 1.000, −0.788, −2.022, against 2.252, 1.000, −0.786, −2.010 from the formula.
Fig. 3 Four hundred thousand points sampled from the measure by the random-iteration rule — shrink towards the left half with probability 0.3 — and binned at eight box sizes. For each qq the logarithm of the sum of shares to the power qq falls on a straight line. The slopes are 2.254, 1.000, −0.788 and −2.022 for q=1,0,2,4q = -1, 0, 2, 4, against 2.252, 1.000, −0.786 and −2.010 from the formula.

For this measure the sum after nn splits is exactly (0.3q+0.7q)n(0.3^q + 0.7^q)^n, so on these axes each line has slope

β(q)=ln(0.3q+0.7q)ln2,\beta(q) = \frac{\ln(0.3^q + 0.7^q)}{\ln 2},

and the figure fits the sampled counts and requires the fitted slope to match it. The samples were never told the formula; the random-iteration rule produces points and the grid counts them. The agreement to three decimal places for positive qq is the check that the sampled points really are distributed the way the construction says, and the looser agreement at q=1q = -1 is honest: negative powers are dominated by the emptiest boxes, where a few points more or less move the sum.

The slopes also give the family of numbers the literature calls generalised dimensions, β(q)/(1q)\beta(q)/(1-q): 1 at q=0q = 0, 0.881 in the limit at q=1q = 1, 0.786 at q=2q = 2 and 0.670 at q=4q = 4. They fall as qq rises, because higher powers look ever more exclusively at where the mass is densest.

The curve and the counts are one object

The slopes β(q)\beta(q) and the curve f(α)f(\alpha) look like two unrelated descriptions — one a family of power laws for weighted sums, the other a family of dimensions for sets of points. They are the same information, and the translation between them is short.

Break the weighted sum up by exponent. There are about εf(α)\varepsilon^{-f(\alpha)} boxes with exponent α\alpha, each with share about εα\varepsilon^{\alpha}, so they contribute εqαf(α)\varepsilon^{q\alpha - f(\alpha)}. As ε\varepsilon shrinks the sum is dominated by whichever α\alpha makes that exponent smallest, so

β(q)=maxα(f(α)qα),f(α)=minq(qα+β(q)).\beta(q) = \max_\alpha \big( f(\alpha) - q\alpha \big), \qquad f(\alpha) = \min_q \big( q\alpha + \beta(q) \big).

That pair is the Legendre transform, the operation that describes a convex curve by its tangent lines instead of its points. Each qq is the slope of one tangent to the spectrum, and β(q)\beta(q) is where that tangent meets the vertical axis. The tangent of slope nought is the horizontal line through the top, so β(0)\beta(0) is the largest dimension present, the support’s. The tangent of slope one passes through the origin, because β(1)=0\beta(1) = 0 — the shares add to one — and that is why the diagonal, the line of slope one through the origin, touches the curve: its point of contact is the exponent that carries the whole mass.

The figure of the spectrum checks the transform both ways. It computes the minimum of qα+β(q)q\alpha + \beta(q) by brute force over six thousand values of qq and requires it to match the parametric curve, and at several finite nn it solves for the qq whose tilted proportion 0.3q/(0.3q+0.7q)0.3^q/(0.3^q + 0.7^q) equals j/nj/n and requires the curve at that qq to be the limit of the binomial count.

The same rule on a set with gaps

Nothing above needed the mass to live on an interval. Split the Cantor set’s mass the same way — the left third of every piece taking a quarter and the right third three quarters — and the support is itself a fractal.

A mass split 7 times, 0.25 to the left and 0.75 to the right. A self-similar measure on the Cantor set: the mass is split 7 times, 0.25 of each piece's share going left and 0.75 right, and each of the 128 pieces is drawn as a bar as tall as its mass. The heaviest carries 0.1335 and the lightest 6.1 × 10⁻⁵.
Fig. 4 The same construction on the middle-thirds Cantor set, split seven times, with shares 0.25 and 0.75. The 128 pieces sit in the Cantor set’s intervals and the skyline drops to nothing across its gaps. The heaviest carries 0.1335 and the lightest 0.000061, 2,187 times less.

The pieces now have length 3n3^{-n}, so every ln2\ln 2 in the formulas becomes ln3\ln 3 and the whole spectrum shrinks by the factor ln2/ln3\ln 2/\ln 3.

A spectrum of dimensions from 0.26 to 1.26. The multifractal spectrum of a self-similar measure on the Cantor set with shares 0.25 and 0.75: the dimension f(α) of the set of points where the mass scales with exponent α. Its top is 0.6309 and it touches the diagonal at 0.5119.
Fig. 5 The spectrum of that measure. Its exponents run from 0.262 to 1.262, its top is 0.6309 — the dimension of the Cantor set itself — and it touches the diagonal at 0.5119, the dimension of the part of the Cantor set that carries the mass. The counts after 10, 40 and 160 splits sit below it by at most 0.128, 0.047 and 0.016.

The top of this curve is the number the Cantor set’s own count gave, ln2/ln3=0.6309\ln 2/\ln 3 = 0.6309, and it has to be: at q=0q = 0 the weighted count forgets the weights. What the spectrum adds is everything the set-counting could not see. Inside a set of dimension 0.63 the mass concentrates on a subset of dimension 0.51, and around that sits a continuum of thinner and thicker subsets, each measured.

When the shares are equal, the curve collapses

The unevenness is doing all the work, and the way to see that is to remove it.

Equal shares, and a spectrum that is one point. The multifractal spectrum of a self-similar measure on the unit interval with shares 0.5 and 0.5: the dimension f(α) of the set of points where the mass scales with exponent α. With equal shares it is the single point 1.0000.
Fig. 6 The same construction with the mass split evenly, half to each side. Every piece after any number of splits carries exactly the same mass, every point has the same exponent, 1, and the spectrum is the single point at 1 — which is also where it meets the diagonal. The dashed curve behind it is the spectrum for shares of 0.3 and 0.7, drawn for comparison.

With equal shares the measure is just length, every point is typical in both senses at once, and the top of the curve and its point of contact with the diagonal are the same point. A measure whose spectrum is a single point has one dimension; a measure whose spectrum is a curve has infinitely many, and the width of the curve is how uneven it is. For shares of 0.3 and 0.7 the exponents span 1.22; for 0.1 and 0.9 they would span 3.17, and as one share goes to nought the lightest exponent goes to infinity.

Where the spectrum came from

The pieces of this subject arrived separately and were assembled late.

The digit-counting result is the oldest: Besicovitch’s and Eggleston’s dimensions of the sets of numbers with a given frequency of digits, in the 1930s and 1940s, were theorems about numbers with no measure in sight. Alfréd Rényi defined a one-parameter family of entropies in 1961, of which Shannon’s is the case q=1q = 1, and those entropies are the logarithms of exactly the weighted sums in the moment figure. Benoit Mandelbrot built measures by repeated uneven splitting in 1974, as models of how energy is shared out among eddies in a turbulent fluid.

The generalised dimensions — one for each qq — were introduced for strange attractors by H. G. E. Hentschel and Itamar Procaccia in 1983. Uriel Frisch and Giorgio Parisi, working on turbulence, described the same structure in 1985 as a spectrum of local exponents, and the paper that gave the curve its usual name, f(α)f(\alpha), and showed it to be the Legendre transform of the moments, was by Thomas Halsey, Mogens Jensen, Leo Kadanoff, Itamar Procaccia and Boris Shraiman in 1986.

That paper’s point was practical. A spectrum can be estimated from the weighted box counts of a sampled orbit, as the moment figure does, without ever identifying the sets it describes, and then compared across systems. Two attractors with the same box dimension can have visibly different spectra, and the difference is a fingerprint of how their dynamics share out time. The rigorous theory for measures without an exact construction came afterwards and is still incomplete, which is the subject of the last section but one.

Where the formalism needs its hypotheses

The construction is exactly self-similar. Every formula here comes from the moments being an exact power, (pq+(1p)q)n(p^q + (1-p)^q)^n, and that comes from the same split being repeated identically in every piece. For a measure with no such rule the Legendre relation between the counts and the level sets is a conjecture to be checked, not a theorem, and there are measures for which it fails — the spectrum computed from the counts is then only an upper envelope of the true one.

The grid lines up with the construction. Boxes of side 2k2^{-k} on the interval, or 3k3^{-k} on the Cantor set, fall exactly on the pieces. A grid of another size cuts pieces in two and blurs the exponents at the finest scales, which is harmless in the limit and visible in any finite count.

Negative powers need many samples. The sum at q<0q < 0 is dominated by the emptiest occupied boxes, and a box whose expected count is a handful of points is estimated badly. The figure keeps its range to eight box sizes for that reason — at eight splits the lightest box still expects 26 of the 400,000 points — and allows its q=1q = -1 slope a looser tolerance than the others.

And the level sets are not separate regions. The points with exponent 0.881 are dense in the interval, and so are the points with every other exponent; each level set is a dust spread through all the others. The dimension f(α)f(\alpha) measures how much of the interval each dust is, not where it is.

None of the sets the spectrum describes can be drawn

The measure figure is ten splits of an infinite construction and its skyline is bounded by the pixel. The spikes at every scale continue below the resolution, and the drawing is evidence of the pattern rather than a picture of the limit.

The spectrum is drawn as a curve, and none of the sets it describes can be drawn at all. A set of numbers whose binary digits are 30 per cent noughts has no picture: it is dense, uncountable and of length nought, and its dimension 0.881 is a theorem about covers rather than something visible. The dots are finite counts of pieces, which are the nearest thing to the sets a figure can show.

And the moment figure measures slopes over eight box sizes, the same narrow range the attractor’s count had to settle for. Its agreement with the formula is evidence about the sample, and the exactness belongs to the construction.

Still open: whether the classic Hénon attractor has a spectrum at all

The natural next object is the measure an orbit of a chaotic map spends its time in, and the standard example is the one the stretching-rates essay used, the Hénon map at a=1.4a = 1.4, b=0.3b = 0.3. Its orbit is spread over a banded set, its mass is visibly uneven across the bands, and numerical multifractal analysis produces a smooth spectrum for it with a plausible top and a plausible point of contact.

None of that is proved, and the gap is not a technicality. It has never been shown that the Hénon map at those classical parameters has a strange attractor at all. Michael Benedicks and Lennart Carleson proved in 1991 that strange attractors exist for a set of parameters of positive measure near b=0b = 0, and Benedicks and Lai-Sang Young that those attractors carry a natural measure of the kind this essay needs; whether a=1.4a = 1.4, b=0.3b = 0.3 is one of those parameters, or whether the orbit everybody draws is a very long transient on its way to a stable periodic cycle, is open. The spectrum computed from the counts is a measurement of something whose existence is not established, which is a sharper version of the warning the stretching-rates essay ended on.

A measure is not described by where it is

The habit is about what a single number hides.

The measure in the first figure lives on a set of dimension one, and so does length. By the measure of its support the two are identical, and by every measurement that counts occupied boxes they are indistinguishable. The difference is entirely in how the mass is shared out, and a description that records only where the mass is cannot see it. The spectrum is what a description looks like when it records how much is where, at every scale at once: the top of the curve is the old dimension, and everything below it is information the old dimension threw away.

The same move applies well beyond fractals. Two distributions with the same range, two networks with the same nodes, two images with the same pixels lit — any comparison that asks only about the support will call them equal. When the object carries weights, ask how the weights scale, not just where they are, and expect the answer to be a curve rather than a number.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A dashed tag is an object no other essay names yet.

Binomial coefficientBox dimensionCantor setHausdorff dimensionIterated function systemLaw of large numbersMeasureScalingSelf-similarity