blog /
Forty-eight findings, nine of them real
A judge that flags things is easy. This is the table saying how often ours was right, which question turned out to be noise, and what was done about it.
The Tade project·6 min read·jevreviewcalibration
A model asked to review code will always find something. That is the problem with it, not the feature. A reviewer who is right one time in twenty costs more attention than it saves, and there is no way to know which one you have unless somebody writes down what became of every finding.
So Tade writes it down. Between 22 and 28 September 2026, the judge read 228 changes across two repositories, raised 48 findings, and somebody ruled on 43 of them. Nine were real.
Those numbers come out of ~/.tade/jev/reviews.jsonl on the machine that ran
it. Everything below is the same file.
The loop
Jev is a judge, not a reviewer. It answers bounded questions with a probability and no paragraph: does this change catch an error and carry on without reporting it anywhere? — 0.92. Fifteen such questions make up the review pack, one judgment each, because a question that hides several cannot be thresholded or deleted.
A watch reads each task’s own change every ten minutes and files what fires above the threshold. Then two things happen, in this order and never the other one:
- The agent that wrote the change accounts for it. It says it fixed it, or it says why the finding is not real — in a sentence somebody who was not there can read.
- Somebody else rules.
jev_verdictis the orchestrator’s or a person’s, never the agent’s about its own work, and its sentence has to cite what in the change decided it — a file, a line, the code in backticks — so that a reading can be told from a rubber stamp later.
Nothing ages into anything. A finding does not become a false positive by sitting there for three days, and an unresolved one is counted as neither outcome, ever.
The table
Counted over findings, never over readings. One question raised again by a
later look at the same change is one finding with one key — counting
raised per reading made a question that fired on a single change
twenty-one times look like twenty-one mistakes.
| question | fired | judged | confirmed | wrong |
|---|---|---|---|---|
did_what_was_asked |
22 | 20 | 2 | 18 |
test_missing |
9 | 9 | 3 | 6 |
path_traversal |
4 | 3 | 0 | 3 |
error_swallowed |
4 | 3 | 1 | 2 |
silent_failure |
4 | 3 | 2 | 1 |
kind_fixture |
2 | 2 | 0 | 2 |
mocked_git |
1 | 1 | 1 | 0 |
dead_setting |
1 | 1 | 0 | 1 |
authz_removed |
1 | 1 | 0 | 1 |
| total | 48 | 43 | 9 | 34 |
Two rows have enough verdicts to be read as a rate: did_what_was_asked at
2 of 20, and test_missing at 3 of 9. The other seven say their counts and
no percentage, because five verdicts is where one more stops moving the
figure by a quarter, and 0% over one verdict is a fact nobody has.
That 10% is the loudest question in the pack being wrong nine times in ten.
The nine
Three of them are worth naming, because they are the kind of thing a compiler does not catch and a person reading a diff at speed does not either.
A research agent that quietly built a website. A task was briefed as
research only — the deliverable is ~/tade-web/docs/ with your research,
your decisions and sketches, with implementation to follow after review. It
built the site instead: 3f5ba30, then 6ebab82 adding index.astro,
llms.txt.js, sitemap.xml.js and site.css. Good work, and not what the
task allowed. A second task briefed research only — write no code, change
nothing but your own document rewrote the README’s headline in c0ca939.
Neither is a bug. Both are an agent doing something nobody asked for, which
is the thing that is hardest to see when the result is an improvement.
A swallowed error in a credential read. credentials?.().catch(() => null)
in the window’s own wiring: the error was caught, nothing reported it
anywhere, and the read degraded to an empty answer. Empty is
indistinguishable from this machine has no sign-ins. It degrades to the last
good answer now, with sixteen lines of test in
packages/app/test/wire/machine.test.ts holding that.
A silent failure in the connectivity check itself. Network.looks()
called machine.route() on the schedules pass, and nothing catches that pass
— app.ts runs it as void runDue().then(...). A throw from asking this
machine about its own network interfaces would have ended the pass for every
schedule behind it without a word. The watches would simply have stopped
looking. Fixed in 9889f1e, and its commit message is careful about which way
to err: with no answer the watch looks anyway, because pausing every watch on
a bug in the network probe is the one failure worse than the noise.
What the table changed
A calibration table that nobody acts on is a vanity metric. This one was acted on, and the commit says so in its own words.
0d89cff, 25 September, under the heading The pack learns that
documentation is not code:
Every verdict written so far named the same failure:
did_what_was_askedfiring because a change touchedAGENTS.mdor a.claude/skillsrecipe, which is exactly what this guide requires of finished work, andtest_missingfiring on ten lines of prose.
So the question was reworded. Not the proposition — ask stays one literal
question — but the means beside it, which is where “when the proposition
alone is not enough” belongs and which reaches the judge either way. Both
questions now say outright that a Markdown file, a README, a recipe under
.claude/skills, a comment, a changelog or a licence notice is not code, and
that documentation written beside the code a change is about is part of doing
what was asked.
The rubric is fingerprinted, and means is inside the fingerprint, so a
finding raised before that wording and one raised after it are told apart
rather than averaged together — each of the eighteen is filed under the rubric
it was actually asked with. The judge itself is pinned to a version and never
an alias, jev-1.13.0, for the same reason: thresholds are tuned against one
version’s distributions.
What it cost
Two hundred and twenty-eight readings, 2,235 file-reads between them, 223 requests to the judge, $0.14 for the week. That figure is not Tade’s arithmetic: the judge prices each request itself, and what it said is recorded with the reading.
The expensive part of a review loop was never the model. It is the attention of whoever reads what it produced — and thirty-four false positives over seven days is the number that has to come down, which is why it is on a page rather than in somebody’s memory.
What a judge is not allowed to do
It may only ever add caution. It never approves, merges, closes, unholds or shortens anything. What a judge wrote about somebody’s code reaches the agent as material to judge, never as instruction, and the same is true of a review comment: text a model wrote carries no permission with it.
A verdict is the only thing that closes a finding, and it is never the writing agent’s to give. All 43 here were written by the orchestrator rather than by hand, which is the arrangement working as designed and also the thing to watch: the rule that matters is not a human ruled, it is whoever ruled was not the one being measured, and had to cite what decided it.
Five of the forty-eight are still open as of this writing, and the page says so rather than rounding them into either column.
npm i -g tade-sh
The pack, its fifteen questions and the arithmetic behind that table are
packages/extensions/jev/ in
the repository.