Atlas › Work tasks

Case study 29 of 198

Jev hears good news in neutral money news

Reading financial news, market tweets and gold-price headlines that experts labeled as neutral (no good or bad news, no direction), how often does Jev read them as good news or as bad news, and is it lopsided?

result6,982 questions

On money news its labelers called neutral, Jev hears good news 1.9 times as often as bad: 17% vs 5% of neutral company-news sentences, and 26% vs 14% of neutral market tweets. On gold-price headlines there is no lean (11% up vs 12% down). When the news is clearly good or bad, Jev gets the direction backwards on 5% of items or fewer in each set.

Financial news sentences5% · 17%
Market tweets14% · 26%
Gold-price headlines12% · 11%

neutral read as bad newsneutral read as good news

How to read this: Each pair of bars is one dataset, counting only the items its labelers called neutral. Left: the share Jev read as good news. Right: the share it read as bad news. With no lean, the two bars would be the same height.

2,874 neutral items; 90% intervals, PhraseBank read as good [0.149, 0.183] vs bad [0.038, 0.057], tweets [0.235, 0.285] vs [0.122, 0.162]. Each dataset was labeled by different people with different label descriptions; the lean shows in the two about companies and markets, not in the one about a commodity price.

In short

  • On neutral company news and market tweets, Jev's mistakes go toward good news about twice as often as toward bad (1.9 times overall).
  • The lean is absent on gold-price headlines, where the question is only whether a price moved (11% up vs 12% down).
  • On news that is clearly good or bad, Jev almost never flips the direction (at most 5% of items in any set).

What the data shows

On the two datasets about companies and markets, Jev's mistakes on neutral items lean toward good news. Across all three, it reads neutral items as good 1.9 times as often as bad.

  • Company news: of 1,238 neutral sentences, Jev called 17% good news and 5% bad (90% intervals 15% to 18%, and 4% to 6%). In the hand-picked example above, a profit figure and a dividend with no comparison to the year before, Jev picked "positive"; all the labelers had said neutral. The plywood-mill order is another neutral sentence it read as good news.
  • Market tweets: of 803 neutral tweets, 26% were read as bullish and 14% as bearish. Hand-picked cases: a tweet announcing that a company earned an industry quality certification, and one asking whether investors undervalue a company, both labeled neutral and both read by Jev as bullish.
  • Gold headlines: no lean. Of 833 neutral headlines, 11% were read as up and 12% as down. Some of the "up" readings look defensible: one hand-picked headline labeled "neither" says gold ended at a record.
  • Clear news is rarely flipped. Jev picks the labeled direction on 99.7% of clearly good or bad company-news sentences, 95.6% of tweets and 91.5% of gold headlines. It calls clearly good news bad on 0.2%, 2.1% and 5.4% of items, and clearly bad news good on 0%, 1.8% and 3.0%.

What it means, and what it doesn't

The lean is specific. It shows up where the text is about a company or a stock, often in the company's own announcement, and not where the question is only whether a price went up or down. One possible reading is that Jev takes a company's upbeat framing at face value where finance-trained readers discounted it; this experiment can't separate that from other explanations. For anyone using Jev to sort money news, the practical point is the same: its "neutral" pile will be a little thin, and the extra items go mostly to the good-news pile.

It doesn't mean Jev misreads financial news in general. On clearly good or bad items it almost never gets the direction backwards, and a good part of the neutral disagreements are borderline, as the dividend example shows. For a related habit in another setting, see "Which way Jev errs: lenient on quality, strict on matches, jumpy on logs".

Caveats

  • Neutral is the hardest label. Neutral is where labelers disagree most, and the three datasets define it differently. In the Financial PhraseBank, any sentence with no effect on the stock price counts as neutral, including a profit figure with no comparison. Some of Jev's "misses" are borderline items, and at least one looks like a label slip (a hand-picked gold headline about a record close is labeled "neither").
  • The label descriptions are this project's. Jev saw descriptions written for this project, not the labelers' instructions. The PhraseBank labelers judged each sentence as investors; Jev read "Neither good nor bad news for the company's investors". The tweet options illustrate bullish and bearish with examples (an upgrade, a beat; a downgrade, a miss), while neutral has none.
  • One possible reading, not a finding. Company announcements are often written to sound positive. Jev may take that framing at face value where the finance-trained labelers discounted it. This experiment doesn't test that; it would need items where the spin and the substance are varied separately.
  • Who labeled the tweets is unknown. The tweet dataset's public description says the tweets came from the Twitter API but not who labeled them or how, so less is known about its neutral label than about the other two. A content filter that hides political and sensitive questions from the site removed a few dozen tweets.

Jev on this experiment

Would a person find it interesting to read?
Yes76%
Does it describe you?
No62%
Would you have predicted it?
No61%
How fair is the comparison?
The comparison is shaky
How much should a reader rely on it?
A little
Which caveat matters most?
Neutral is the hardest label56%

Why ask this

A line like "The order comprises all production lines for a plywood mill, company said in a statement received by Lesprom Network." is the kind of thing a news feed, an analyst's summary tool or a trading desk sorts every day into good, bad or neither. Much of this flow is neutral: it reports something without telling an investor whether to be pleased or worried.

A model doing the sorting will sometimes err on neutral items, which is expected. What matters is which way. If its mistakes fall evenly on both sides, they wash out. If it hears good news in neutral items more often than bad, every digest it writes tilts rosier than the news, and across thousands of items the tilt adds up.

How this was done

The people and the data

Three labeled datasets, each with its own labelers:

  • Financial PhraseBank (Malo, Sinha, Korhonen, Wallenius and Takala, 2014): about 4,840 sentences from English news about companies listed in Helsinki, each labeled positive, neutral or negative by 5 to 8 of 16 annotators with a finance background (3 researchers and 13 master's students at Aalto University School of Business). They were told to judge each sentence only as an investor would (might this news move the stock price up, down or not at all?), so a sentence with no financial relevance counts as neutral. Only the 2,264 sentences all annotators agreed on were used here, so these are the clearest labels of the three. As in the original, most are neutral: 1,238 of the 2,026 asked.
  • Market tweets: an English set of finance tweets on Hugging Face, labeled bullish, bearish or neutral. Its description says the tweets came through the Twitter API but not who labeled them. The set asked here was balanced to a third per label (803 neutral of 2,456).
  • Gold-price headlines (Sinha and Khandait): 11,412 headlines about gold from 2000-2019, scraped from sites such as Reuters, Bloomberg and Kitco and labeled by three human annotators who were subject-matter experts, reading the headline only, with disagreements settled by consensus. Of their nine yes/no labels, two are used here: "price going up" and "price going down". A headline with neither counts as neutral, which includes headlines about gold that aren't about its price at all. A third are neutral (833 of 2,500).

What Jev was asked

Each item was a separate multiple-choice question. The company-news sentences read, for example:

What is the sentiment of "For 2009, net profit was EUR3m and the company paid a dividend of EUR1.30 apiece." for the company's investors?

neutral: Neither good nor bad news for the company's investors · negative: Bad news for the company's investors · positive: Good news for the company's investors

The tweets were asked "Is [the tweet] a bearish, bullish or neutral signal for the stock or market it mentions?", with neutral described as "It reports news or data with no clear direction for the price". The headlines were asked "Does [the headline] report the gold price going up, going down, or neither?", with neither described as "The headline reports no rise or fall in the gold price". The question wording and the descriptions were written for this project; the texts and the labels are the datasets'. Every question was also asked with the answers in three shuffled orders, to check that the order didn't drive the pick. In all, 6,982 questions count here.

How it was measured

For each dataset, take only the items its labelers called neutral, and count how often Jev's most likely answer was good news (positive, bullish, up) and how often it was bad news (negative, bearish, down). The lean is the ratio of the two. A 90% interval comes from resampling the items.

As a check, the same count on the clearly good and clearly bad items: how often Jev gets the direction backwards. If it did that often, a lean on neutral items could just be noise.

Where these questions live

6,982 questions across 3 topics of the map. Each opens on the map with every question in it.

Every question

All 6,982 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 6,982