A dimension for every rate of crowding
Worth reading first: Infinite on one side and nought on the other · A dimension from the stretching rates.
Kaplan and Yorke’s formula ended on a distinction that most of the time is invisible: the dimension of an attractor as a set is one number, and the dimension of the distribution an orbit spends its time in is another. A set can be larger than the part of it that carries the weight.
The simplest object on which that distinction becomes a whole subject is not an attractor at all. It is a unit of mass spread over an interval by a rule that splits it unevenly.
Every point of the interval gets some mass, so the set the mass lives on is the whole interval, of dimension one. But the mass is not spread evenly across it, and it is not spread unevenly in any simple way either: the picture has spikes inside spikes at every scale. The way to describe such a thing is not one dimension but a curve of them — one for each rate at which mass crowds into a shrinking neighbourhood — and the curve is computable exactly.
How fast the mass in a small interval vanishes
Take a point and the piece of length that contains it after splits. The piece’s mass is , where is the number of times the point fell in a left half, and is simply the number of noughts among the first binary digits of .
For an evenly spread mass the piece would carry , its own length. Here it carries its length raised to a power:
The exponent depends only on the proportion of noughts in the point’s binary expansion. A point whose digits are all noughts sits where the mass is thinnest, with — its neighbourhoods lose mass faster than their length. A point whose digits are all ones sits where it is thickest, with — its neighbourhoods keep mass longer than length would allow. Every exponent between is the exponent of some proportion.
That turns the question into one about digits, and digits have been counted before. The average of many coin tosses settles: almost every number, picked at random along the interval, has noughts and ones in equal proportion, and so almost every point in the sense of length has . But almost every point in the sense of the mass is different. A point picked according to the mass falls left with chance 0.3 at every split, so its proportion of noughts settles at 0.3, and its exponent at
The two notions of “typical” disagree, and each is certain in its own sense. The mass is spread over every point of the interval and concentrated, all but a vanishing fraction of it, on points whose digits are 30 per cent noughts — and those points, as the next sections show, form a set of dimension 0.881.
Counting the pieces that share a mass
After splits there are possible masses, and the number of pieces carrying the mass with left choices is — the number of ways to place noughts among digits. The figure checks that count at every one of its 1,024 pieces.
That is enough to define a dimension for each exponent. Pieces of length that share one exponent number , and a set that needs about boxes of side has dimension . So the points whose neighbourhoods scale with exponent make up a set whose dimension, at finite , is
As grows that approaches the entropy of the proportion, with , divided by . The limit is exact, and it is not only a box count. Abram Besicovitch in 1934 and Harold Eggleston in 1949 proved that the set of numbers whose binary digits have noughts in proportion has Hausdorff dimension equal to that entropy over — so the dimension of each level set is the one the finest covers see, not just the one a grid sees.
Two points on that curve say most of what the section above said in words. The top of the curve is the typical point for length: the most numerous exponent, 1.126, belongs to a set of full dimension one, because almost every number has half its digits noughts. The point where the curve touches the diagonal is the typical point for mass: exponent 0.881 and dimension 0.881. Everywhere else the curve lies strictly below the diagonal, and that is not a coincidence of this example. A set of dimension can hold mass that scales like per box only if ; where the two are equal, the set is exactly big enough to carry all the mass.
The finite counts approach the curve from below, as they must — a binomial coefficient is smaller than its entropy estimate by a factor that grows like — and the gap shrinks roughly like .
Almost all the mass on almost none of the pieces
The distance between the top of the curve and its point of contact with the diagonal can be turned into a count, and the count is startling.
After ten splits, the 120 pieces with exactly three left choices — about one piece in nine — carry of the mass, more than any other group. As the number of splits grows the groups near a proportion of 0.3 take more and more of the mass, because the proportion of left choices in a mass-weighted piece obeys the law of large numbers, and the groups away from it take less. After a thousand splits the pieces whose proportion lies within a few per cent of 0.3 carry all but a sliver of the mass. How many of them are there? About , out of pieces in all.
The mass-carrying pieces are a fraction of the pieces, and that fraction goes to nought exponentially as the splitting continues. So the measure is spread over the whole interval in the sense that no piece is empty, and it lives on almost none of the interval in the sense that matters for where it actually is. The dimension 0.881 is exactly the exponent of that shrinking fraction: the carrying set needs boxes where the whole interval needs .
This is the same count that underlies data compression. A long string of symbols from a source that emits one symbol with probability 0.3 and the other with 0.7 is, with overwhelming likelihood, one of about typical strings rather than one of all , so it can be written down in 0.881 bits per symbol. The information dimension of the measure and the entropy of the source are the same number for the same reason: both count the strings that actually occur. The name “information dimension” is that coincidence, recorded.
Measuring it without knowing the rule
The curve above was computed from the construction. An orbit or an experiment does not come with one, and the practical question is how to get the spectrum from samples.
The method is box counting with the boxes weighted. Cover the samples with boxes of side , record the share of the points in each box, and add up those shares raised to a power . At every occupied box counts once, which is ordinary box counting. At the shares add to one whatever is. At the sum is the chance that two points picked independently land in the same box. Large positive is dominated by the heaviest boxes and negative by the lightest.
For this measure the sum after splits is exactly , so on these axes each line has slope
and the figure fits the sampled counts and requires the fitted slope to match it. The samples were never told the formula; the random-iteration rule produces points and the grid counts them. The agreement to three decimal places for positive is the check that the sampled points really are distributed the way the construction says, and the looser agreement at is honest: negative powers are dominated by the emptiest boxes, where a few points more or less move the sum.
The slopes also give the family of numbers the literature calls generalised dimensions, : 1 at , 0.881 in the limit at , 0.786 at and 0.670 at . They fall as rises, because higher powers look ever more exclusively at where the mass is densest.
The curve and the counts are one object
The slopes and the curve look like two unrelated descriptions — one a family of power laws for weighted sums, the other a family of dimensions for sets of points. They are the same information, and the translation between them is short.
Break the weighted sum up by exponent. There are about boxes with exponent , each with share about , so they contribute . As shrinks the sum is dominated by whichever makes that exponent smallest, so
That pair is the Legendre transform, the operation that describes a convex curve by its tangent lines instead of its points. Each is the slope of one tangent to the spectrum, and is where that tangent meets the vertical axis. The tangent of slope nought is the horizontal line through the top, so is the largest dimension present, the support’s. The tangent of slope one passes through the origin, because — the shares add to one — and that is why the diagonal, the line of slope one through the origin, touches the curve: its point of contact is the exponent that carries the whole mass.
The figure of the spectrum checks the transform both ways. It computes the minimum of by brute force over six thousand values of and requires it to match the parametric curve, and at several finite it solves for the whose tilted proportion equals and requires the curve at that to be the limit of the binomial count.
The same rule on a set with gaps
Nothing above needed the mass to live on an interval. Split the Cantor set’s mass the same way — the left third of every piece taking a quarter and the right third three quarters — and the support is itself a fractal.
The pieces now have length , so every in the formulas becomes and the whole spectrum shrinks by the factor .
The top of this curve is the number the Cantor set’s own count gave, , and it has to be: at the weighted count forgets the weights. What the spectrum adds is everything the set-counting could not see. Inside a set of dimension 0.63 the mass concentrates on a subset of dimension 0.51, and around that sits a continuum of thinner and thicker subsets, each measured.
When the shares are equal, the curve collapses
The unevenness is doing all the work, and the way to see that is to remove it.
With equal shares the measure is just length, every point is typical in both senses at once, and the top of the curve and its point of contact with the diagonal are the same point. A measure whose spectrum is a single point has one dimension; a measure whose spectrum is a curve has infinitely many, and the width of the curve is how uneven it is. For shares of 0.3 and 0.7 the exponents span 1.22; for 0.1 and 0.9 they would span 3.17, and as one share goes to nought the lightest exponent goes to infinity.
Where the spectrum came from
The pieces of this subject arrived separately and were assembled late.
The digit-counting result is the oldest: Besicovitch’s and Eggleston’s dimensions of the sets of numbers with a given frequency of digits, in the 1930s and 1940s, were theorems about numbers with no measure in sight. Alfréd Rényi defined a one-parameter family of entropies in 1961, of which Shannon’s is the case , and those entropies are the logarithms of exactly the weighted sums in the moment figure. Benoit Mandelbrot built measures by repeated uneven splitting in 1974, as models of how energy is shared out among eddies in a turbulent fluid.
The generalised dimensions — one for each — were introduced for strange attractors by H. G. E. Hentschel and Itamar Procaccia in 1983. Uriel Frisch and Giorgio Parisi, working on turbulence, described the same structure in 1985 as a spectrum of local exponents, and the paper that gave the curve its usual name, , and showed it to be the Legendre transform of the moments, was by Thomas Halsey, Mogens Jensen, Leo Kadanoff, Itamar Procaccia and Boris Shraiman in 1986.
That paper’s point was practical. A spectrum can be estimated from the weighted box counts of a sampled orbit, as the moment figure does, without ever identifying the sets it describes, and then compared across systems. Two attractors with the same box dimension can have visibly different spectra, and the difference is a fingerprint of how their dynamics share out time. The rigorous theory for measures without an exact construction came afterwards and is still incomplete, which is the subject of the last section but one.
Where the formalism needs its hypotheses
The construction is exactly self-similar. Every formula here comes from the moments being an exact power, , and that comes from the same split being repeated identically in every piece. For a measure with no such rule the Legendre relation between the counts and the level sets is a conjecture to be checked, not a theorem, and there are measures for which it fails — the spectrum computed from the counts is then only an upper envelope of the true one.
The grid lines up with the construction. Boxes of side on the interval, or on the Cantor set, fall exactly on the pieces. A grid of another size cuts pieces in two and blurs the exponents at the finest scales, which is harmless in the limit and visible in any finite count.
Negative powers need many samples. The sum at is dominated by the emptiest occupied boxes, and a box whose expected count is a handful of points is estimated badly. The figure keeps its range to eight box sizes for that reason — at eight splits the lightest box still expects 26 of the 400,000 points — and allows its slope a looser tolerance than the others.
And the level sets are not separate regions. The points with exponent 0.881 are dense in the interval, and so are the points with every other exponent; each level set is a dust spread through all the others. The dimension measures how much of the interval each dust is, not where it is.
None of the sets the spectrum describes can be drawn
The measure figure is ten splits of an infinite construction and its skyline is bounded by the pixel. The spikes at every scale continue below the resolution, and the drawing is evidence of the pattern rather than a picture of the limit.
The spectrum is drawn as a curve, and none of the sets it describes can be drawn at all. A set of numbers whose binary digits are 30 per cent noughts has no picture: it is dense, uncountable and of length nought, and its dimension 0.881 is a theorem about covers rather than something visible. The dots are finite counts of pieces, which are the nearest thing to the sets a figure can show.
And the moment figure measures slopes over eight box sizes, the same narrow range the attractor’s count had to settle for. Its agreement with the formula is evidence about the sample, and the exactness belongs to the construction.
Still open: whether the classic Hénon attractor has a spectrum at all
The natural next object is the measure an orbit of a chaotic map spends its time in, and the standard example is the one the stretching-rates essay used, the Hénon map at , . Its orbit is spread over a banded set, its mass is visibly uneven across the bands, and numerical multifractal analysis produces a smooth spectrum for it with a plausible top and a plausible point of contact.
None of that is proved, and the gap is not a technicality. It has never been shown that the Hénon map at those classical parameters has a strange attractor at all. Michael Benedicks and Lennart Carleson proved in 1991 that strange attractors exist for a set of parameters of positive measure near , and Benedicks and Lai-Sang Young that those attractors carry a natural measure of the kind this essay needs; whether , is one of those parameters, or whether the orbit everybody draws is a very long transient on its way to a stable periodic cycle, is open. The spectrum computed from the counts is a measurement of something whose existence is not established, which is a sharper version of the warning the stretching-rates essay ended on.
A measure is not described by where it is
The habit is about what a single number hides.
The measure in the first figure lives on a set of dimension one, and so does length. By the measure of its support the two are identical, and by every measurement that counts occupied boxes they are indistinguishable. The difference is entirely in how the mass is shared out, and a description that records only where the mass is cannot see it. The spectrum is what a description looks like when it records how much is where, at every scale at once: the top of the curve is the old dimension, and everything below it is information the old dimension threw away.
The same move applies well beyond fractals. Two distributions with the same range, two networks with the same nodes, two images with the same pixels lit — any comparison that asks only about the support will call them equal. When the object carries weights, ask how the weights scale, not just where they are, and expect the answer to be a curve rather than a number.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A carpet with two dimensions — both name box dimension, hausdorff dimension, iterated function system, measure, scaling, self-similarity
- The room a jagged graph takes up — both name box dimension, hausdorff dimension, scaling, self-similarity
- A staircase with no steps — both name cantor set, measure, self-similarity
- Almost none of it left, and still uncountably many — both name cantor set, measure, self-similarity
- No interval in it, and length to spare — both name cantor set, measure, self-similarity
- A constant that does not care which map — both name scaling, self-similarity
Named objects
A dashed tag is an object no other essay names yet.
Binomial coefficientBox dimensionCantor setHausdorff dimensionIterated function systemLaw of large numbersMeasureScalingSelf-similarity