Two children and the sentence about one of them
Worth reading first: The door that was not opened.
A family has two children. At least one of them is a boy. What is the chance that both are boys?
The answer usually given is one in three, and the argument for it is short. Two children, each a boy or a girl with equal chance and independently, make four equally likely families: boy–boy, boy–girl, girl–boy, girl–girl, listing the older child first. The sentence rules out girl–girl and nothing else. Of the three families left, one is boy–boy.
Most people’s first answer is one in two, on the grounds that the other child is a boy or a girl and the sexes of siblings are independent. That answer is also right, for a different sentence. The problem was put this way by Martin Gardner in 1959, alongside a companion — the older child is a girl; what is the chance both are girls? — whose answer is one in two, and in a later column he conceded that the first version was ambiguous: the answer depends on how the information was obtained, and the question does not say.
This is the same lesson as the door that was not opened, with the host removed. At the doors, a person with rules decided what to reveal, and the rules were what decided the answer. Here there is nobody visibly choosing anything, and the ambiguity is harder to see because of it. But somebody said the sentence, or answered a question, or was met in the street, and each of those is a different process that ends in the same words.
The square before anything is said
Everything starts from the square. Four families, each a quarter, is the whole of the prior: before any sentence, the chance of two boys is one in four, of two girls one in four, and of one of each one in two, because a mixed family can arise in two orders.
That last point is where the arithmetic hides. A family with a boy and a girl is twice as likely as a family with two boys, not equally likely, because it occupies two quarters of the square. Anyone who reasons “there are three kinds of family — two boys, two girls, one of each” and treats them as equally likely has made the error that Bayes’ theorem drawn as a square exists to prevent: counting kinds instead of measuring areas. The kinds are not equally likely and the square says so.
The square is also a model rather than a fact about families. Real births run at roughly 105 boys to 100 girls, and siblings’ sexes are very slightly correlated. Neither changes anything below by more than a percentage point, and both are set aside: the puzzle is about what a sentence rules out, and it is clearest where the prior is simplest.
With the prior drawn, every version of the problem has the same shape. Some process produces a piece of evidence; the evidence crosses out the parts of the square in which that process could not have produced it; and the answer is the share of what survives that is boy–boy. The versions differ only in which parts are crossed out.
Naming a child changes the answer
Take the sentence the older child is a boy. It rules out the whole bottom half of the square: both families in which the older child is a girl.
Two quarters are left, boy–boy and boy–girl, and they are equally likely. The answer is one in two, which is the naive answer, now correct.
The difference from the first version is that at least one is a boy does not say which child. It rules out one quarter; the older is a boy rules out two. The mixed family survives the first sentence in both of its orders and the second in only one of them, so the first sentence leaves the mixed families at two to one against boy–boy, and the second leaves them level.
That is the whole mechanism, and it generalises. Any sentence that picks out a particular child — the older, the taller, the one whose name comes first alphabetically, the one who answered the door — and says that child is a boy leaves the chance of two boys at one half, because it cuts the square along a line that treats the two orders separately. Only a sentence that refuses to identify the child, that says some child is a boy without saying which, crosses out the girl–girl quarter alone.
So the one-third answer requires a specific and slightly unusual kind of evidence: a guarantee about the family as a whole, with no child in particular behind it. A census clerk asking does this household include a boy? produces that evidence, and so does a school that invites only families with at least one son. Meeting a boy in the street does not.
Meeting one child is naming one
Suppose the family is visited, one of the two children is met at random, and it is a boy. The sentence reporting this — at least one of them is a boy — is true. But the process was not the census clerk’s.
Every family now splits in two, according to which child happened to be met. The eighths in which a girl was met are crossed out. In the boy–boy family, both halves survive, because whichever child was met was a boy. In each mixed family, only one half survives. So boy–boy keeps its full quarter and each mixed family keeps an eighth, and the survivors are two eighths of boy–boy against two eighths of mixed. One in two.
This is the same structure as the host who opened a door at random. The random host threw away worlds in which he would have revealed the prize; the random meeting throws away worlds in which a girl would have been met. In both cases the family’s composition affects how likely the observed evidence was, and that likelihood is what redistributes the probability. A boy–boy family is twice as likely as a mixed family to produce a boy on the doorstep, and Bayes’ theorem converts that factor of two into exactly the difference between one in three and one in two.
Put in the language of likelihoods: under the census clerk’s question, every family with a boy produces a “yes” with probability one, so the evidence does not distinguish among them and the prior ratio of one to two survives. Under the random meeting, a boy–boy family produces a boy with probability one and a mixed family with probability one half, so the ratio of one to two is multiplied by two to one, and becomes even.
The same families, two ways of finding out
The two processes can be run side by side on the same families, which settles any suspicion that the difference lies in the setup rather than in the questions.
Every family in the simulation was produced the same way, and each was put through both procedures. Of the three thousand, about three quarters answer yes to the census question, and of those about a third have two boys. About half produce a boy when one child is chosen at random, and of those about half have two boys. Both running shares wobble early and then settle, as a running average of independent trials must, and they settle in different places.
The two sets of kept families overlap heavily. Every boy–boy family is in both. The mixed families are in the first set always and in the second only when the boy happened to be chosen — half the time. That is the entire difference, visible in the counts: the second procedure keeps roughly half as many mixed families as the first, and the same number of boy–boy families.
A simulation — an estimate built from random trials — proves nothing that the square has not already shown, but it does one thing the square cannot: it makes it impossible to argue that the problem is about words. The families are the same families. The only thing that differs between the two curves is what was done to find out about them.
When a parent chooses what to say
The version that most people meet in conversation is a third process, and it is the one where the answer is genuinely undetermined. A parent of two says, unprompted: one of them is a boy.
What that sentence rules out depends on what the parent would have said otherwise. A parent of two boys can only say one of them is a boy. A parent of two girls can only say one of them is a girl. A parent of one of each can say either, and chooses — by habit, by which child is on their mind, by which child is standing beside them.
Call the chance that a mixed family mentions its boy q. The boy–boy quarter survives in full, and the two mixed quarters survive in the proportion q, so the chance of two boys is one quarter divided by one quarter plus one half times q, which is 1/(1 + 2q).
At q = 1 that is one in three: the parent always volunteers the boy, so the sentence carries no more information than the census clerk’s answer. At q = 1/2 it is one in two: the parent picks a child to talk about at random, which is meeting a child at random. And at q = 0 it is one — a parent who always mentions a girl when there is one has just revealed that there is not.
So the conversational version has no answer until the parent’s habits are specified, and every answer from one third to certainty is available. The figure also makes a less obvious point. The answer is never below one third, and it never falls below the prior of one quarter, because the sentence one of them is a boy is always at least as likely from a boy–boy family as from a mixed one. Evidence that a hypothesis makes at least as likely as its alternatives can only raise that hypothesis. How far it raises it is the likelihood ratio, which here runs from 1 to infinity as q runs from 1 to 0.
A boy born on a Tuesday
In 2010, at a gathering held in Gardner’s memory, the puzzle designer Gary Foshee posed a version that looked like a joke: a parent of two children says that one of them is a boy born on a Tuesday; what is the chance that both are boys? The weekday looks irrelevant. It is not, and the answer is 13/27.
The count is in the grid. Each child is one of fourteen kinds, a sex and a weekday, so there are 196 equally likely families. The families with at least one Tuesday boy are the ones whose older child is a Tuesday boy — a row of fourteen — together with those whose younger child is — a column of fourteen — less the one family in which both are, counted twice — the smallest case of inclusion and exclusion. That is 27. Among them, the boy–boy families are the ones where the other child is also a boy: seven in the row, seven in the column, less the shared corner, which is 13.
The reason the weekday matters is that it nearly identifies a child. At least one boy is satisfied in a boy–boy family by either child, and that double coverage is what held the answer at one third. At least one Tuesday boy is satisfied in a boy–boy family by either child only if both were born on Tuesday, which is rare; almost always exactly one child fits, and the sentence behaves like a sentence naming that child. So the answer moves almost all the way from one third to one half. It stops at 13/27 rather than 1/2 because of the corner cell, the one family in which two Tuesday boys both fit and are counted once.
The census clerk’s protocol is being assumed here — the family is kept because at least one child is a Tuesday boy, however that came to be known. Under a random-meeting protocol the weekday changes nothing, and the answer stays at one half. The whole disagreement about Foshee’s puzzle, which was considerable, is the disagreement about Gardner’s puzzle one level down.
How much detail makes a child
The weekday is one example of a general effect, and the effect has a formula.
If the confirmed boy has a detail that a fraction p of boys share, the same row-and-column count gives the chance of two boys as (2 − p)/(4 − p). At p = 1 the detail is shared by everyone and the formula gives 1/3. At p = 1/7 it gives 13/27. At p = 1/365, a birthday, it gives 729/1459, which is 0.4997. The curve rises smoothly from one third to one half as the detail becomes rarer.
The formula is the reason the weekday result, which sounds like a paradox, is really the most natural thing in the problem. The one-third answer is not the “default” answer that a weekday perturbs. It is the extreme case of a family of answers, the one in which the confirming detail is so common that it cannot tell the two children apart. Every rarer detail — a name, a birthday, a hair colour, a school — tells them apart a little, and the answer drifts toward the one-half that naming a child outright would give.
This also makes plain why the effect is so hard to see in words. A boy and a boy born on a Tuesday sound like the same claim with a decoration. As evidence about a family, they are different events with different areas, and the second has a smaller area in a mixed family than twice the area in a boy–boy family — which is exactly what the count of 13 against 27 records. Twenty-three people turns on a similar count, where a detail everyone ignores — how many pairs a group contains — decides the answer.
Where the words do not settle it
Two claims need separating, because the literature on this puzzle has repeatedly run them together.
The mathematics is not ambiguous. For any fully specified process — census, random meeting, a parent with habit q, a census for Tuesday boys — the answer is determined, and the square computes it. There is no dispute about the arithmetic anywhere in this essay. Every figure’s answer is counted off the areas and checked against the corresponding formula.
The English is ambiguous. “A family has two children, and at least one is a boy” does not say which process produced the knowledge, and a reader must supply one. Readers who hear a guarantee about the family supply the census and get one third. Readers who hear a report about a child in front of someone supply a random meeting and get one half. Neither is misreading; they are completing an incomplete problem differently. Maya Bar-Hillel and Ruma Falk’s 1982 paper on conditional-probability teasers made exactly this point: small changes of wording change which process a reader supplies, and so change the answer the reader will defend.
What should settle it is the question the door essay closed on: what would have been said in the cases that did not happen? If the girl–girl family would have produced some different statement, and the mixed families would have produced exactly this one, then the answer is one third. If the mixed families would have produced this statement only some of the time, the answer is higher. The problem as usually posed says nothing about the cases that did not happen, and that silence is the entire difficulty.
The practical rule is the one the medical test taught in a different setting: before interpreting a piece of evidence, ask how often it would appear under each hypothesis. A positive test from a well person, a losing door from an informed host, a boy mentioned by a parent — each is evidence only in proportion to how differently the hypotheses would produce it.
What the square cannot show
The prior itself. Every figure draws four equal quarters. The true proportions differ slightly from equal, and families with more children, or with deliberate stopping rules — keep having children until a boy — produce different priors again. With a stopping rule the families no longer fill a square at all, and the waiting-time arguments are the right tool rather than a table of four cells.
How a real sentence was produced. The mention figure makes the dependence on q exact, but no figure can measure q for an actual speaker. It is a fact about a person rather than about families, and it is precisely the fact that the puzzle’s wording leaves out.
Why intuition resists. The square shows which cells survive; it does not explain why three cells feel like they should be “two boys, two girls, mixed” rather than “boy–boy, boy–girl, girl–boy”. That resistance is psychological, and it is well documented, but it is outside what areas can depict.
The problems this one opens
The same structure — evidence whose likelihood depends on the answer, produced by a process the telling omits — reappears in problems where the omission cannot be repaired by asking.
The sharpest is the Sleeping Beauty problem, where the question is not how was the sentence produced but how many times is the question being asked, and where the two natural ways of counting give one half and one third. There, unlike here, specifying the protocol completely does not end the argument, because the protocol is fully specified in the statement and the disagreement survives it.
Nearer to hand are the three prisoners problem, which is the door problem told about a pardon; the general form of all of these as a likelihood ratio multiplying prior odds, which is the arithmetic this essay’s figures kept performing; and the question of what happens to these answers when the families are not a square — larger families, stopping rules, correlated sexes — where the same method applies and the numbers stop being tidy.
The sentence is not the evidence
At least one is a boy has one meaning as a sentence and several as evidence. As a sentence it rules out the girl–girl family. As evidence it rules out whatever the process that produced it would not have produced — one quarter of the square if a clerk asked the question, one quarter and half of the mixed families if a child was met, a variable share if a parent chose what to mention.
The answers are one third, one half, and anything from one third to one. Foshee’s Tuesday boy lands at 13/27 because a rare enough detail almost names a child, and naming a child gives one half. None of these is a paradox once the square is drawn and the process is stated. The difficulty is entirely that the usual telling states the sentence and not the process, and a sentence alone cannot say which parts of the square it came from.
What links here
Computed from the collection, not written here: the essays that point at this one.
Reads more easily once this is understood
Essays that name this one as worth reading first.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Area by counting dots — both name area, counting argument
- Giving up on the best — both name conditional probability, counting argument
- Nobody gets their own hat — both name counting argument, sample space
- One point in every big enough shape — both name area, counting argument
- The colouring nobody has ever seen — both name counting argument, independence
- When to stop looking — both name conditional probability, counting argument
Named objects
A dashed tag is an object no other essay names yet.
AreaBayes' theoremConditional probabilityCounting argumentIndependenceLaw of large numbersLikelihoodSample space