What the data shows

- Familiarity: close agreement (0.82). Jev knows which pairings are stock phrases.
- Aptness: loose agreement (0.38; 90% interval 0.29 to 0.45).
- It finds almost everything apt. Metaphors average 4.2 for Jev against 3.4 for people, and literal phrases 4.3 against 3.7.
- Neither side separates metaphors much from literal phrases. Both rate literal phrases slightly higher.
- Where it's stingiest relative to people: "lazy eye", "fragrant shadow" and "lonely oval" fall furthest below people's rankings.
What it means, and what it doesn't
Jev is a generous reader of figurative language: it rarely calls a pairing strained. Ask it whether an image in your draft works and it will usually say yes. Its sense of which phrases are familiar is much sharper than its sense of which ones are good.
It doesn't mean Jev has no taste in images: it ranks them loosely like people. But the level descriptions written for this project sound approving in the middle (see Caveats), which could inflate its averages, though not its ranking.