What the data shows

Jev's arousal ratings track people's pleasantness in reverse (-0.43) about as strongly as they track people's own arousal ratings (0.37). For people, the dials are moderately linked the other way: pleasant words lean stirring (+0.37). For Jev, the two collapse into one: its own arousal and its own pleasantness ratings run opposite each other at -0.50.
The words it gets most wrong show the pattern plainly:
- Pleasant words it calls calming that people find stirring: cuddle, loving, snuggle, caress, warmth, beach. People put "cuddle" at 6.6 of 9; Jev put it at the very bottom, "sleepy or soothing".
- Unpleasant words it calls intensely stirring that people rate calm to middling: puke, misery, awful, amputate, anguish, deluge. People rated "misery" about 3.3 of 9 for arousal: sad, but not jolting. Jev put it near the top.
What it means, and what it doesn't
When Jev says something is "intense" or "exciting", it's partly telling you it's unpleasant. That matters for any task that asks a model to rate energy separately from mood: music, ads, customer messages, even alarm levels.
It doesn't mean Jev can't tell pleasant from unpleasant. Its pleasantness ratings match people's closely (see "Jev knows how pleasant a word is, not how exciting"). The confusion is specific to the second dial. Part of it may also come from the project's own answer wording (see Caveats); rerunning the arousal questions with neutral examples is the obvious next test.