Atlas › Work tasks

Case study 120 of 198

When Jev misroutes, the right answer is usually next door

When Jev sends a customer message to the wrong intent, how wrong is it: a neighbor of the right intent, or somewhere else entirely, and does a longer list of intents make it worse?

result32,430 questions

When Jev sends a customer message to the wrong place, the right answer is its second choice 51% of the time, and the typical miss is a close sibling: "order a physical card" sent to "get a physical card". A longer menu barely hurts: with 77 banking intents it's right 81% of the time, and across 13 routing datasets the number of options and the share right are barely related (rank correlation 0.10).

snips intents (7)57% · 97%
multiwoz domain (7)26% · 98%
math routing (7)67% · 75%
dolly tasks (7)38% · 60%
abcd flows (9)51% · 82%
cfpb complaints (9)49% · 68%
airline complaints (9)49% · 76%
clinc150 (11)71% · 90%
skill select (11)65% · 87%
massive en (18)54% · 90%
sgd dialogue (24)56% · 89%
hwu64 intents (64)53% · 88%
banking77 (77)47% · 81%

misses where the answer was 2nd choiceright

How to read this: Each row is a routing dataset, ordered by how many options Jev chose from. The grey bar is how often it picked the right one; the magenta bar is, among its misses, how often the right one was its second choice.

32,430 routing questions, 5,952 misses, 13 datasets with 7-77 options.

In short

  • Across 13 routing datasets and 32,430 messages, when Jev picks the wrong intent, the right one is its second choice 51% of the time.
  • Menu length barely matters (rank correlation 0.10); overlapping categories do the damage, with Dolly instruction types at 60% right from only 7 options.
  • A router built on Jev should use its second choice (show two options, or ask when the top two are close) rather than trust the top pick alone.

What the data shows

  • The right answer is usually close. Among 5,952 misses, the right intent was Jev's second choice 51% of the time. For CLINC150 it's 71%, for math topics 67%.
  • Misses are siblings. In banking, the most common miss is "get a physical card" sent to "change PIN", then "order a physical card" sent to "get a physical card". In support flows, "manage account" sent to "subscription inquiry"; in airline complaints, a complaint about a flight attendant sent to "customer service".
  • Menu length barely matters. With 7 options Jev ranges from 60% right (instruction types) to 98% (travel domains); with 77 banking intents it's right 81% of the time. Menu length and accuracy are barely related (0.10).
  • Fuzzy menus are the problem. The weakest datasets are the ones with overlapping categories: instruction types (60%, "brainstorming" sent to "open question"), consumer complaints (68%), math topics (75%, "intermediate algebra" sent to "algebra").

What it means, and what it doesn't

For routing, Jev's mistakes are mostly the cheap kind: the neighbor of the right answer, often ranked second. A router built on it should use the second choice (show two options, or ask a clarifying question when the top two are close) rather than trusting the top pick alone.

Menu length showed little relation to accuracy here, while the fuzziest menus did worst. Cleaning up intents that mean nearly the same thing will do more than shortening the list.

Caveats

  • Some intents overlap by design. Many "misses" are between intents a person would also hesitate over. "Where is the card PIN?" is labeled "get a physical card"; Jev said "change PIN". The datasets' labels treat these as wrong, so the share right is a floor.
  • Clean research messages. Most of these datasets use short, tidy messages collected or written for research. Real customer messages are longer, messier and often ask two things at once.
  • Different menus, different difficulty. Some datasets have crisp, separate intents (travel domains, music vs weather); others blur (task types like "brainstorming" vs "open question"). Comparing datasets mixes menu length with how distinct the intents are.

Jev on this experiment

Would a person find it interesting to read?
Yes60%
Does it describe you?
Yes53%
Would you have predicted it?
Yes54%
How fair is the comparison?
The comparison is reasonable
How much should a reader rely on it?
Moderately
Which caveat matters most?
Some intents overlap by design68%

Why ask this

Routing is one of the most common jobs a model like Jev does: read a customer's message and send it to the right team, flow or tool. It's the task TypeSafe itself leads with.

What matters isn't only how often the router is right, but how it's wrong. A message sent to a neighboring intent ("order a card" vs "get a card") costs little; one sent to an unrelated team costs a lot. And if the right answer is usually the router's second choice, a simple fallback (show both, or ask) recovers most misses.

How this was done

The people and the data

The project used 13 public routing datasets, 32,430 messages in all: bank customer queries (BANKING77, 77 intents), voice assistant commands (SNIPS, HWU64, MASSIVE, CLINC150), travel and service dialogues (MultiWOZ, Schema-Guided Dialogue), consumer complaints to the US Consumer Financial Protection Bureau, airline complaints on Twitter, customer-service chats for a fictional clothing store (ABCD), picking the right tool for an AI assistant (MetaTool), competition-math topics (MATH), and instruction types from the Dolly dataset. The menus range from 7 to 77 options.

Each message's right intent comes from its dataset: labeled by crowd workers (the airline tweets), chosen by the person who wrote the message for it (the Dolly instructions, the ABCD chats), or the product a complaint was filed under.

What Jev was asked

Each message was one pick-one question, with every intent written out with a one-line description. For example:

What is the customer asking the bank about in this message?

Message: "Where is the card PIN?"

Options (77): age limit: the minimum age to open or use an account · change PIN: changing the card PIN · ATM support: which ATMs the card works at · … · get a physical card …

How it was measured

For each dataset: the share of messages Jev routes right; among its misses, the share where the right intent was its second choice; and the most common confusions. Across the 13 datasets, the analysis compares menu length with the share right (rank correlation: 1 means longer menus always do better, -1 always worse, 0 no relation).

Where these questions live

32,430 questions across 15 topics of the map. Each opens on the map with every question in it.

Every question

All 32,430 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 32,430