Atlas › Work tasks

Case study 89 of 198

Jev reads what code says, not what it does

Given a function, can Jev tell whether its docstring or commit message describes it, and can it tell whether it contains a security bug or needs a reviewer's comment?

result9,000 questions

Jev reads what code says, not what it does. It tells whether a docstring describes a function 93% of the time, and whether a commit message describes a change 93%. But whether a function has a security bug, it gets right 54% of the time, a coin flip: it misses 48% of the vulnerable functions and flags 43% of the fixed ones. Whether a change needs a reviewer's comment: 60%.

docstring fits the function93%
commit message fits the diff93%
change needs a review comment60%
function has a security bug54%

50% = a coin flip

How to read this: Each bar is one code task, showing how often Jev got it right. Every set is half yes and half no, so a bar near 50% means Jev can't tell the two apart.

9,000 questions, each set balanced; 90% interval on the security-bug share [0.526, 0.563].

In short

  • Jev reliably checks whether a docstring or commit message matches the code (93% each), but spotting a security bug in a C function is a coin flip (54%).
  • On security it isn't biased one way: it misses 48% of vulnerable functions and flags 43% of fixed ones, so its verdicts carry almost no signal.
  • It is a useful checker of stale comments and mislabeled commits, not a vulnerability scanner.

What the data shows

work tasks
Scooby doo mask reveal meme: a well-documented function; let's see what it really does; a well-documented function; a buffer overflowa well-documented functionlet's see what it really doesa well-documented functiona buffer overflow
How funny is this meme? Jev: 3/5, funny10%216%374%410%50%
  • Matching text to code: strong. Docstrings 93% (it misses 12% of true pairs and almost never accepts a wrong docstring, 1%); commit messages 93%.
  • Security bugs: chance. 54% right: it misses 48% of vulnerable functions and flags 43% of fixed ones. It isn't leaning one way; it simply can't separate them.
  • Review comments: weak. 60% right on whether a change drew a reviewer's comment, missing 37% and flagging 42%.

What it means, and what it doesn't

Jev is a good reader of code: it can check whether documentation and commit messages match what's there, which is useful for catching stale comments and mislabeled commits. It is not a vulnerability scanner. Given one function at a time, it has no real signal on whether that function is exploitable.

It doesn't mean Jev understands nothing about code behavior. These tasks are hard for everyone at the function level, and the security labels are noisy. But a code-review tool built on Jev should keep security and "does this need review" judgments with people or dedicated tools.

Caveats

  • The security labels are noisy. A function counts as "vulnerable" if a vulnerability-fixing commit touched it, and "fixed" after the fix. That's a known noisy label: some "vulnerable" functions may not contain the flaw themselves. Part of Jev's coin flip is the label's.
  • Functions out of context. Each C function from FFmpeg and QEMU is shown alone. Many memory-safety bugs depend on how the function is called or how buffers are sized elsewhere, which Jev can't see. A person reviewing one function would face the same limit.
  • Review comments are one reviewer's call. "Needs a review comment" means a human reviewer happened to comment on that part of a real pull request. Reviewers skip things and comment on style. Published models built for this task reach roughly 70-73%.
  • Mismatches were easy. The wrong docstrings and commit messages come from other functions or commits in the same project, not near-misses written to fool. That makes the matching tasks easier than subtle drift between code and comments.

Jev on this experiment

Would a person find it interesting to read?
Yes75%
Does it describe you?
Yes66%
Would you have predicted it?
Yes61%
How fair is the comparison?
The comparison is shaky
How much should a reader rely on it?
Moderately
Which caveat matters most?
The security labels are noisy95%

Why ask this

Code review asks two different skills. One is reading: does this comment describe this function, does this commit message match this change? The other is reasoning about behavior: can this function overflow a buffer, does this change need a second look?

A model strong at the first and weak at the second would look very helpful in code review while missing exactly what matters.

How this was done

The people and the data

Four public datasets, 9,000 questions, each balanced half yes and half no:

  • Docstring vs function (CodeSearchNet, Python, JavaScript, Java and Go): the function's own docstring, or one from another function in the same project.
  • Commit message vs change (CommitBench): a commit's own message, or another commit's from the same project.
  • Security bugs (Devign, from the CodeXGLUE suite): C functions from FFmpeg and QEMU, before and after a vulnerability fix.
  • Needs a review comment (CodeReviewer, from Microsoft): changes from real GitHub pull requests, labeled by whether a human reviewer commented on them.

What Jev was asked

Each item was a yes/no question over real code. For example:

Does the function contain a security vulnerability or memory-safety bug? (a C function from QEMU or FFmpeg follows)

Does the docstring accurately describe what the code does?

Docstring: "Assert node is a text node."

Code: function text(node) { unist(node); assert.strictEqual('children' in node, false, …); assert.ok('value' in node, …) }

How it was measured

For each task, the share Jev got right, the share of real cases it missed, and the share of clean cases it wrongly flagged, with a range showing how much it could vary by chance. Because every set is half yes and half no, 50% is what guessing would get.

Where these questions live

9,000 questions across 4 topics of the map. Each opens on the map with every question in it.

Every question

All 9,000 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 9,000