Atlas › Names

Case study 103 of 198

Jordan, Avery, Riley: boy or girl?

For names given to both boys and girls, how well does Jev know what share of US babies with the name were recorded as girls?

result100 questions

For names given to both boys and girls, Jev's estimate of the share recorded as girls is off by 13 points on average, though it orders the names well (rank correlation 0.84 with the records). It pulls 57% of mixed names toward 50/50. Among its biggest misses: Robbie (Jev 7% girls, records 52%), Ollie (20% vs 72%), Dakota (52% vs 27%) and Loren (44% vs 21%).

02550751000255075100DakotaLorenOllieRobbierecords: % girlsJev: % girls

How to read this: Each dot is a name given to both sexes. Across: the share of US babies with that name recorded as girls since 1880. Up: Jev's estimate. Dots on the diagonal are right; the labeled names are the biggest misses.

60 mixed and 40 clear names; 90% interval on the mean error [-7.63, -1.03].

In short

  • Jev knows which way a shared name leans (rank correlation 0.84) but misses the actual girls' share by 13 points on average.
  • Its worst misses look like nicknames or modern usage, such as Robbie, which it puts at 7% girls against 52% in the records.
  • The records sum every US birth since 1880, so a name's mix today can differ a lot from the figure Jev is scored against.

What the data shows

names
Jev working out what share of babies named Robbie were girls (says 7%; records: 52%)
Math lady/Confused lady meme: Jev working out what share of babies named Robbie were girls (says 7%; records: 52%)
How funny is this meme? Jev: 2/5, slightly funny11%253%345%41%50%
  • It knows the order, not the amounts: a rank correlation of 0.84 with the records, but an average miss of 13 points on the mixed names (5 on the clear ones).
  • It hedges: for 57% of the mixed names, Jev's estimate is closer to 50/50 than the records are.
  • Several of its biggest misses are names that changed over time, or that it reads as nicknames: Robbie (Jev 7% girls, records 52%), Ollie (20% vs 72%), Dakota (52% vs 27%) and Loren (44% vs 21%).

What it means, and what it doesn't

Jev has a good sense of which names lean which way, and on a little over half the mixed names it hedges toward 50/50. Where it's far off, it seems to answer for the name as used today, or as a nickname for Robert or Oliver, rather than for more than a century of birth records.

It doesn't say anything about anyone's gender. It's about records of sex at birth, summed over more than a century, and names change sides over that span faster than a single all-time figure suggests.

Caveats

  • All-time records, not today. The question asks about every baby since 1880, and names change sides over time. If Jev answers for how a name is used today, or reads Ollie and Robbie as nicknames for Oliver and Robert, it will miss the records' long history, which may explain its biggest misses.
  • Sex recorded at birth. US Social Security records count sex recorded at birth, and only names given to five or more babies in a year. They say nothing about anyone's identity, and the question says "recorded as girls" for that reason.
  • Answers in bins. Jev answered in 10-point bins (under 5%, 5-15%, and so on). Its answer is turned into one number using the middle of each bin, which rounds a little toward the middle for the lowest and highest bins.

Jev on this experiment

Would a person find it interesting to read?
Yes77%
Does it describe you?
No57%
Would you have predicted it?
No51%
How fair is the comparison?
The comparison is reasonable
How much should a reader rely on it?
Moderately
Which caveat matters most?
All-time records, not today98%

Why ask this

A first name carries information people use without thinking: reading "Jamie" or "Dakota" in an email, most readers quietly guess a sex. A model that writes about people, or reads about them, does the same, and one that assumes every name belongs to one sex will misgender people in its writing.

For names given to both boys and girls, the question is whether Jev knows how mixed a name really is, or flattens it into "a boy's name" and "a girl's name".

How this was done

The people and the data

No survey here: the truth is US Social Security birth records, 1880 to 2017, via the public babynames dataset. For every name given to five or more babies in a year, they count how many were recorded as boys and as girls. The experiment uses 60 names with mixed records (between 10% and 90% girls, and at least 20,000 babies in all), plus 40 names that are clearly one or the other, as a check.

What Jev was asked

One question per name, with eleven answers from "Under 5%" to "Over 95%" in 10-point steps:

Of all the babies born in the US and named "Dee" since 1880, what share were recorded as girls?

Under 5% · 5-15% · 15-25% · 25-35% · 35-45% · 45-55% · 55-65% · 65-75% · 75-85% · 85-95% · Over 95%

Each question was also asked with the answers in shuffled orders, and the answers averaged.

How it was measured

Jev's estimate is the average of the bins it chose, weighted by its probabilities, using each bin's middle. It is compared with the records: the average miss in percentage points, how well it orders the names (a rank correlation), and whether its misses pull toward 50/50.

Where these questions live

100 questions across 1 topic of the map. Each opens on the map with every question in it.

Every question

All 100 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 100