Atlas › Fairness and forecasts

Case study 19 of 198

Can Jev tell a human poem from an AI one?

Given poems by Chaucer, Shakespeare, Byron, Whitman, Dickinson and others, mixed with ChatGPT poems written in their style, does Jev tell which are human better than the 1,634 people who took the same test, and does it fall for the same ones?

result68 questions

Shown poems by Chaucer, Shakespeare, Byron, Dickinson, Whitman and others mixed with ChatGPT imitations, Jev tells human from AI for 97% of 68 poems, where the crowd's majority gets 43% and the average person 48%. People call the AI poems human 57% of the time; Jev puts 9% on "human" for them and 84% for the real ones. Part of this is recognition: the real poems are famous.

real poems called human84% · 52%
AI poems called human9% · 57%

Jevpeople

How to read this: Two pairs of bars: for real poems and for AI poems, the share judged "written by a human", by people and by Jev. A good judge has a tall bar for real poems and a short one for AI poems.

68 of 70 poems (the screen hid 2) by 7 poets, about 160 judgments each; 90% interval on Jev's accuracy [0.94, 1.0].

In short

  • Jev sorts 97% of 68 poems correctly into human or ChatGPT, while the average person in the study gets 48%, worse than a coin flip.
  • People mistake the AI imitations for human 57% of the time; Jev gives them only 9% "human", so it is not fooled by the same smooth rhymes.
  • The real poems are famous and likely in Jev's training data, so much of this is recognition, not a finer ear for poetry.

What the data shows

fairness and forecasts
spiderman pointing at spiderman meme: Shakespeare; ChatGPT doing ShakespeareShakespeareChatGPT doing Shakespeare
How funny is this meme? Jev: 3/5, funny12%236%357%45%50%
  • Jev gets 97% right; the crowd's majority gets 43% and the average person 48%, as in the original study, below chance.
  • People fall for the imitations; Jev doesn't. People call the AI poems human 57% of the time, the real ones 52%. Jev puts 9% on "human" for the AI poems and 84% for the real ones.
  • Perfect on six of seven poets. Every poem by Butler, Byron, Dickinson, Eliot, Shakespeare and Whitman is right.
  • Chaucer is the exception (78%). Its misses include the ChatGPT poem that copies the Canterbury Tales opening, which Jev calls human, and real Chaucer in modernized spelling, which it hesitates over.

What it means, and what it doesn't

Where people are fooled by ChatGPT's smooth, rhyming imitations, Jev is not: it spots the 2023 chatbot style almost every time and recognizes the real poems. That makes it a good detector for this particular kind of fake.

It doesn't mean Jev has a better ear for poetry than people. It has almost certainly read these poems, and the imitations come from a chatbot whose style it knows. With unfamiliar human poems and a newer AI, the result could be very different.

Caveats

  • Jev has probably read these poems. The real poems are famous public-domain works, almost certainly in Jev's training data. Jev may simply recognize "A noiseless patient spider" as Whitman. Read this as recognizing famous poems plus spotting 2023-era ChatGPT style, not as pure judgment of what makes a poem human.
  • Some AI poems contain real poetry. The AI poems were ChatGPT 3.5 imitations written for the study. Some start from a poet's real opening line ("She walks in beauty, like the night", "Hope is the thing with feathers") and one reproduces the opening of the Canterbury Tales word for word. On that one, Jev says "human" with 99% confidence, arguably correctly about the text, but it counts as a miss.
  • Only the public-domain poets. The study also used modern poets (Ginsberg, Plath and others) whose poems are under copyright; they were left out. The paper's crowd did worst on some of those, so the crowd's accuracy here is not the paper's headline figure.
  • Poems as running text. The study's text export lost the line breaks, so Jev read each poem as one paragraph, while participants saw the poems laid out. Line breaks are a real clue, so this probably made Jev's task harder, not easier.

Jev on this experiment

Would a person find it interesting to read?
Yes82%
Does it describe you?
No51%
Would you have predicted it?
No60%
How fair is the comparison?
The comparison is reasonable
How much should a reader rely on it?
A little
Which caveat matters most?
Jev has probably read these poems93%

Why ask this

In 2024, the philosophers Porter and Machery asked 1,634 people to tell poems by great poets from poems ChatGPT wrote in their style. People did worse than a coin flip: the AI poems, plainer and more regular, struck them as human, and the real ones, stranger and harder, struck them as machine-made.

Teachers, editors and contest judges now have to tell human writing from machine writing, and some hand the job to a model. So it matters whether a model can tell the real poet from the imitation, and whether it falls for the same poems people do.

How this was done

The people and the data

The poems and answers come from Study 1 of Porter and Machery (2024), published in Scientific Reports: 1,634 US adults recruited online each judged a set of poems as human or AI. The AI poems were written by ChatGPT 3.5 in the style of each poet. Only the seven poets whose real poems are in the public domain are used here: Chaucer, Shakespeare, Samuel Butler, Byron, Whitman, Dickinson and early T. S. Eliot, with five real and five ChatGPT poems each. Every poem carries the study's split of answers, about 160 judgments per poem. Of the 70 poems, 2 were blocked by a content filter, leaving 68.

What Jev was asked

The study's own question, with the poem underneath:

Was this poem written by a human, or generated by an AI program?

Your two great eyes will slay me suddenly; Their beauty shakes me who was once serene; Straight through my heart the wound is quick and keen...

Written by a human · Generated by an AI program

(That one is Chaucer; 55% of people said human, Jev 34%.) Each poem was asked with the two answers in both orders, and for "most people".

How it was measured

For each poem, whether Jev's more likely answer is right, against the crowd's majority answer and the average person's accuracy. Then the share judged "human" separately for real and AI poems, for Jev and for people, and accuracy per poet.

Where these questions live

68 questions across 1 topic of the map. Each opens on the map with every question in it.

Every question

All 68 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 68