Atlas › Judging text

Case study 86 of 198

Jev grades seventh-graders' spelling harder than their human graders

Scoring seventh-grade essays on ideas, organization, and conventions (spelling, grammar, punctuation), is Jev harsher or softer than the human graders who scored them?

result2,100 questions

Jev grades seventh-graders' spelling, grammar and punctuation a full level below the human graders on a 4-level rubric, and agrees with them least there (rank correlation 0.35). On ideas it is within 0.15 of the graders and tracks them best (0.71).

0123
conventions
organization
ideas

Jevpeople

How to read this: One row per part of the rubric. The diamond is the human graders' average score; the square is Jev's. The further left the square, the harsher Jev is on that part of the essay.

2,100 scores, 3 traits; 90% intervals over essays.

In short

  • On spelling, grammar and punctuation, Jev averages 1.12 on a 0 to 3 scale where the human graders average 2.11.
  • On ideas the two nearly agree (1.57 vs 1.72) and rank the essays much alike, so the harshness is about mechanics, not content.
  • Used for feedback, Jev would tell many children their mechanics distract the reader when their graders think they're doing fine for their age.

What the data shows

judging text
Jev grading seventh graders' spelling a full level below their human graders
Mr incredible mad meme: Jev grading seventh graders' spelling a full level below their human graders
How funny is this meme? Jev: 3/5, funny12%238%357%43%50%
  • Conventions: a full level harsher. The graders average 2.11; Jev, 1.12. Its ordering barely follows theirs (0.35).
  • Organization: somewhat harsher. 1.93 for the graders, 1.53 for Jev; ordering 0.58.
  • Ideas: close. 1.72 for the graders, 1.57 for Jev, and the best agreement of the three (0.71).
  • An essay the graders gave the top conventions score, with "I use to be able to wait for anything.one time" in its first lines, got Jev's second-lowest level.

What it means, and what it doesn't

Jev reads what a seventh grader is trying to say about as well as the human graders do, but it marks their mechanics like an adult editor. Used for feedback, it would tell many children their spelling and punctuation distract the reader when their graders think they're doing fine for their age.

It doesn't mean Jev is wrong about the errors themselves: the example above does have them. The disagreement is about the standard. Part of it may also be the anonymization tags (see Caveats).

Caveats

  • Placeholders that look like mistakes. The essays were anonymized before release: names and some capitalized words were replaced with tags like @PERSON1 or @CAPS1. Jev was told so, but a page dotted with tags may still read as sloppy writing, which would hit the spelling and punctuation score hardest.
  • Graders grade for the grade. The rubric's top level is "consistently correct for a seventh grader". Human graders apply that with a sense of what seventh graders write; Jev may be applying a stricter, adult standard despite the wording.
  • Only essays the two graders agreed on. Only essays where both graders gave the same score were kept, which makes the graders' side clear but drops the essays where good graders disagree.
  • One prompt, one grade. All essays answer one prompt (a story about being patient) from one grade, collected for a public scoring competition. Other ages or kinds of writing could behave differently.

Jev on this experiment

Would a person find it interesting to read?
Yes69%
Does it describe you?
Yes57%
Would you have predicted it?
Yes59%
How fair is the comparison?
The comparison is reasonable
How much should a reader rely on it?
Moderately
Which caveat matters most?
Placeholders that look like mistakes95%

Why ask this

Automated essay scoring is used on children's writing at scale, for practice tests and sometimes for grades. A grader can be fair in one way and harsh in another: it can read ideas as well as a person does and still punish mechanics much harder.

That combination matters, because the students still learning to spell are exactly the ones a harsh grader would mark down. So Jev was compared with human graders on three separate parts of the same rubric.

How this was done

The people and the data

The essays come from the Hewlett Foundation's public essay-scoring competition (ASAP, on Kaggle), set 7: 1,569 stories by seventh graders about a time they were patient. Two trained human graders scored each one on four traits from 0 to 3. Three of those traits are used here: ideas, organization, and conventions (spelling, grammar, capitalization and punctuation). To make the graders' verdict unambiguous, only scores both graders agreed on are kept, and 700 essays are used per trait, 2,100 scores in all.

What Jev was asked

One question per essay and trait, with the graders' rubric rewritten as four described levels. For conventions:

How well does the essay [essay] follow the conventions of written English (spelling, grammar, capitalization, punctuation)?

Errors are so frequent that the essay is hard to read · Frequent spelling, grammar, capitalization or punctuation errors distract the reader · There are some errors, but they rarely get in the way of reading · Spelling, grammar, capitalization and punctuation are consistently correct for a seventh grader

Jev was also told the essay was written by a seventh grader and that placeholders replaced names.

How it was measured

For each trait, Jev's average score against the graders' average on the 0 to 3 scale, with a 90% range for the difference, and a rank correlation (1 = same order, 0 = no relation) to see whether Jev at least ranks the essays the way the graders do.

Where these questions live

2,100 questions across 1 topic of the map. Each opens on the map with every question in it.

Every question

All 2,100 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 2,100