Two numbers in one jagged record
Worth reading first: The room a jagged graph takes up · A dimension that is not a whole number.
The room a jagged graph takes up counted the boxes needed by graphs that came with rules — Takagi’s sums of raised midpoints and Lévy’s random construction of Brownian motion — and found dimensions between one and two that matched the rules to two decimal places. It ended on the harder case. A record of something measured — a river’s level, a temperature, the height of ground along a line — has a finest scale set by how often it was sampled, a coarsest set by how long it runs, and no rule to say what happens outside that range.
It also ended on a claim. For fractional Brownian motion, the graph’s dimension is , where is the Hurst exponent that says how strongly the path’s rises persist; so the dimension of a record, that essay said, is a statement about its correlations, and a record of dimension below one and a half is one whose rises tend to be followed by rises.
For a record, both halves of that need qualifying. The dimension of a finite record is an estimate, and the most familiar way of estimating it — counting boxes — is the least reliable, for a reason as mundane as the choice of units. And the tie between dimension and correlation is a property of one model, not of records: a stationary random record can be rough and have long memory, or smooth and forget quickly, in any combination. Roughness is read at the finest scales of a record and memory at the coarsest, and nothing forces the two readings to agree.
Four records, and what each end of them shows
The four records are stationary Gaussian series: each value is normally distributed, and the correlation between two values depends only on how far apart they are. That correlation is the whole description, and these four use a family introduced by Tilmann Gneiting and Martin Schlather in 2004, called the Cauchy class:
The two parameters act at opposite ends. For small separations the correlation is plus smaller terms, so two nearby values differ by an amount whose square grows like : sets how rough the record is up close. For large separations the correlation decays like : sets how long the record remembers. A small means correlations that fade so slowly they never sum to anything finite. With the formula is , the shape of the Cauchy density, which is where the family’s name comes from; the general member keeps that density’s tail, falling off as a power rather than exponentially, and a power-law tail is exactly what lets correlations persist.
The close-ups in the figure are the roughness and the whole records are the memory. The two records with are smooth over four units whatever their ; the two with are jagged whatever theirs. At full length, the records with drift in swings lasting hundreds of units, and the ones with look like evenly packed noise. Each record carries both numbers, and each number is visible only at its own end of the range. None of the four is a walk built step by step and scaled down; each is a stationary series whose statistics are the same at every point along it, which is the natural model for a record that is not drifting anywhere in particular.
The box count reads low
Before separating the two numbers, it is worth knowing how to measure one. The graph of a record is a set of points in the plane, and the box count of the essay on dimension applies to it directly: cover the graph with a grid of squares of side , count the squares it touches, and read the dimension off how the count grows as shrinks. That was how a curve with a corner at every point was first measured, and how the Takagi graphs were. For a graph sampled at points, the smallest usable square is one sample wide.
The other estimators in common use do not count anything in the plane. The variogram is the average squared difference between values a given lag apart, . For a path whose increments over a lag grow like , the variogram grows like , the dimension is , and so the dimension can be read from the variogram at lags one and two alone:
The madogram does the same with average absolute differences in place of squared ones.
On fractional Brownian paths, whose true dimension is known exactly, the variogram and madogram are right: their estimates lie on the diagonal across the whole range. The box count is wrong in a systematic way. It is close for the smoothest paths and reads lower and lower as the paths get rougher, until a path of true dimension 1.9 is reported as 1.6.
The reason is the finest scale. At a column one sample wide, a box count sees the path’s rise and fall between two neighbouring samples and nothing of what happens between them, and for a rough path what happens between samples is most of the path. The count uses the finest columns to fix the slope, and the finest columns are exactly the ones where the sampling has thrown the detail away. Tilmann Gneiting, Hana Ševčíková and Donald Percival compared the estimators in 2012 and recommended the variogram and madogram over the box count for exactly this family of reasons.
A count that changes with the units
The box count has a second problem, and it is more embarrassing. A box is a square, and a square in the plane of a graph has one side measured in time and the other in whatever the record measures — metres, degrees, pounds. The number of boxes a column needs depends on the ratio of the two units, and that ratio is arbitrary.
Squash the path flat and each column’s rise and fall fits inside a single box, so the count is one box per column, the count grows exactly like , and the graph is reported as a line of dimension one. Stretch it tall and each column needs a stack of boxes; the count climbs, and it levels off at 1.355 — below the true 1.5, because at the finest columns the sampling has flattened the path in the way the previous figure showed. Between the extremes, every value from 1.0 to 1.36 can be had from the same data by choosing the units.
The variogram does not have this problem at all. Multiplying the record by a constant multiplies every squared difference by the square of that constant, and the ratio of two variograms — which is all the estimate uses — does not change. The Hausdorff measure has the same indifference, infinite on one side of the dimension and nought on the other whatever the units; it is a count on a fixed grid that brings the units in. The estimate is the same to every digit at every stretch. A dimension is supposed to be a property of the shape, and a rule that returns different shapes for a river measured in metres and in feet is measuring the ruler.
The same trouble appeared, in a cleaner form, in a carpet with two dimensions: stretch one direction more than the other and the box count and the Hausdorff dimension part company, because a square box is the wrong shape for a set whose two directions scale differently. The graph of a record is always such a set — time and value scale by different factors — and the variogram sidesteps the problem by never forming a box. For a graph built from a formula none of this bites, because the formula supplies the detail below every grid and the count can be pushed down scale after scale until the units no longer matter. For a record there is no below.
Roughness at one end, memory at the other
With a reliable estimator of dimension in hand, the second question becomes answerable. The Hurst exponent measures memory, and it is read at the opposite end of the record: split the record into blocks of samples, average each block, and see how quickly the averages settle as grows. For independent samples the variance of a block average falls like — the wobble of an average shrinks like one over the square root of the count. For a record with long memory the samples in a block are correlated, the block’s average is less informative than its size suggests, and the variance falls more slowly, like with between one half and one. For the Cauchy class, when .
The two panels are the two ends of the same pair of records. On the left, at lags of a few samples, the two records’ variograms are parallel lines of the same slope — their heights differ because also sets how quickly correlation drops at first, but their slopes are both about one, and the slope is all the dimension reads. The measured dimensions are and . On the right, at blocks of thousands of samples, the lines part. The short-memory record’s block averages settle at close to the rate independent samples would, and the long-memory record’s settle slowly. The measured Hurst exponents are and .
Harold Hurst found the effect in 1951, in the long records of the Nile’s annual floods: the river’s good and bad years came in runs longer than independent years would produce, with an exponent of about , and the same held for rainfall, tree rings and lake sediments. Independent years would have given one half, since a sum of independent terms spreads like the square root of their number — the law behind the bell curve built from coin flips — and a reservoir sized on that law would run dry more often than planned. Benoit Mandelbrot and James Wallis called it the Joseph effect, after the seven years of plenty and seven of famine, and fractional Brownian motion was the model Mandelbrot built to reproduce it.
A grid, not a line
Measured over a grid of settings, the two numbers vary independently. Dimension changes with alone, to within four hundredths in every row; the Hurst exponent changes with alone, more loosely, missing its target by up to . The dashed line is , which is where every fractional Brownian path lies and where the jagged-graph essay placed every record. Only one of the twelve settings sits on it.
That is the content of Gneiting and Schlather’s paper, whose title is its theorem: stochastic models that separate fractal dimension and the Hurst effect. Self-affinity is the single assumption that ties them. A self-affine path looks statistically the same after the horizontal scale is multiplied by and the vertical by , at every , so the same governs how it looks up close and how it behaves over its whole length. A record that is not self-affine — a record whose small-scale and large-scale behaviour come from different physics — has no reason to have , and fitting a single self-affine model to it reports one number where there are two.
The claim at the top of this essay, that a dimension below one and a half means persistent rises, holds for records that are self-affine. The Cauchy class is the demonstration that the hypothesis is doing the work: a record whose rises persist over years can be as rough over minutes as white noise summed, and a record with no memory past a week can be as smooth as a spline.
A record with a different dimension at every scale
Self-affinity can fail in a quieter way, too. A record can be made of two processes added together, one dominating at small scales and the other at large ones, and then even the dimension on its own is not a single number.
At the shortest lags the rough component dominates the differences, and the record reads as dimension . At the longest lags the smooth one does, and it reads as about . In between, at every lag, the estimate takes an intermediate value, and there is no range of lags over which it is flat. A reported dimension of such a record is a report of which lags were used.
This is not an artificial case. A measured profile of ground is roughened by rocks at one scale and shaped by drainage at another; a temperature record has daily noise on top of seasonal and climatic swings; an instrument adds its own noise at its finest resolution. Lewis Fry Richardson’s measurements of coastlines, which Mandelbrot reinterpreted in 1967 as the first fractal dimensions, gave straight lines on logarithmic axes over the range of ruler lengths he could use — and it is the straightness over a stated range that was the finding, not the slope alone.
How long a record must be for each number
The two numbers also differ in how much data they need. The dimension is read from differences between neighbouring samples, and a record of samples has about of those, so its estimate is pinned early: from a thousand samples, with a spread of . The Hurst exponent is read from blocks of thousands of samples, and a record has only a handful of those. The shortest records’ blocks are only a few units long, where the record is merely smooth rather than persistent, and the estimate there is — a record that is too short for its memory reports its smoothness instead. Only at a quarter of a million samples does the estimate come down to , still with a spread of .
Memory lives at the scales a record has fewest of. It is the number whose evidence is thinnest and the one most often wanted, because it is the one that decides whether a run of bad years is a warning or a coincidence.
What the figures cannot show
Every record here is Gaussian, drawn exactly from its covariance by circulant embedding: the correlations are laid round a circle twice the record’s length, the Fourier transform turns them into independent variances, and independent normal draws with those variances are transformed back. The method is exact whenever the transformed correlations are non-negative, and every figure checks that they are. Real records are rarely Gaussian, and heavy-tailed jumps affect the variogram and the madogram differently; that is one reason the madogram is kept in the comparison. A record whose mass is spread unevenly across scales, like the measures that crowd at every rate, needs a whole spectrum of exponents rather than one, and neither estimator is designed for that.
The Hurst estimates use aggregated variances over blocks of a fixed range of sizes, and the grid’s misses of up to come from the range rather than from chance: the correlations of the Cauchy class approach their final power law only when is large compared with one, and blocks of sixteen to 256 units are not far enough out for the smallest . Other estimators of memory — rescaled range, detrended fluctuation analysis, the slope of the spectrum near zero frequency — have their own biases and the same shortage of long blocks.
And every estimate is of a model’s parameter. That a real record is well described by any member of the Cauchy class, or by fractional Brownian motion, is itself a claim the data can only support over the range it covers.
Still open: telling a crossover from a curve
A dimension estimated at lags and changes with for a record with a crossover, and it also changes with for a record that is simply not yet in its asymptotic regime — the Cauchy class does it at its larger lags, as its correlations bend from one law to another. Deciding from a single finite record whether there is a range of lags over which the estimate is genuinely constant, and so whether the record has a dimension at all over that range, is a statistical question without a settled answer. The estimates at neighbouring lags are strongly correlated with each other, and a record short enough to be practical has too few independent pieces at large lags to test flatness against a slow curve.
The same issue bites harder for memory. Whether a hydrological or climatic record has true long-range dependence — correlations decaying like a power, forever — or merely correlations that decay slowly over the range observed and then stop, is a question on which Hurst’s own data has been argued over since the 1970s, and it matters, because the two predict very different frequencies for long droughts. A record cannot contain its own infinite future, and the part of a model that concerns the scales beyond the record is the part the record cannot test.
Two readings, not one
The lesson of the jagged-graph essay was that roughness is a rate — how the size of the detail shrinks from one scale to the next — rather than an amount. A record adds a second rate, read at the other end: how slowly its averages settle. For a path built by one rule at every scale the two are the same number seen twice. For a record, they are two measurements, each confined to its own end of the record’s range, each needing a different amount of data, and each estimated reliably only by a method that does not depend on the units.
So a dimension reported for a record should come with three things beside it: the estimator, the range of scales over which it was read, and, if the record is to be called persistent or anti-persistent, a separate measurement of its memory at the scales where memory lives.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A dimension from the stretching rates — both name box dimension, scaling, self affinity
- An average that never settles — both name scaling, variance
- Drawn without putting back — both name correlation, variance
- How fast the bell arrives — both name scaling, variance
- The shape that averaging leaves alone — both name scaling, variance
Named objects
A dashed tag is an object no other essay names yet.
Box dimensionCorrelationFractal dimensionFractional brownian motionHurst exponentScalingSelf affinityVarianceVariogram