# Forty-eight findings, nine of them real

> A judge that flags things is easy. This is the table saying how often ours was right, which question turned out to be noise, and what was done about it.

Published 2026-09-28 · by The Tade project · tagged jev, review, calibration.
Published at https://tade.sh/blog/forty-eight-findings-nine-of-them-real/ — part of https://tade.sh/blog/.

---

A model asked to review code will always find something. That is the problem
with it, not the feature. A reviewer who is right one time in twenty costs
more attention than it saves, and there is no way to know which one you have
unless somebody writes down what became of every finding.

So Tade writes it down. Between 22 and 28 September 2026, the judge read 228
changes across two repositories, raised 48 findings, and somebody ruled on 43
of them. Nine were real.

Those numbers come out of `~/.tade/jev/reviews.jsonl` on the machine that ran
it. Everything below is the same file.

## The loop

Jev is a judge, not a reviewer. It answers bounded questions with a
probability and no paragraph: *does this change catch an error and carry on
without reporting it anywhere?* — 0.92. Fifteen such questions make up the
review pack, one judgment each, because a question that hides several cannot
be thresholded or deleted.

A watch reads each task's own change every ten minutes and files what fires
above the threshold. Then two things happen, in this order and never the other
one:

- **The agent that wrote the change accounts for it.** It says it fixed it, or
  it says why the finding is not real — in a sentence somebody who was not
  there can read.
- **Somebody else rules.** `jev_verdict` is the orchestrator's or a person's,
  never the agent's about its own work, and its sentence has to cite what in
  the change decided it — a file, a line, the code in backticks — so that a
  reading can be told from a rubber stamp later.

Nothing ages into anything. A finding does not become a false positive by
sitting there for three days, and an unresolved one is counted as neither
outcome, ever.

> **A screen from Tade — `what-jev-flagged`.** What Jev has read this week, each question by how often it was right, and every finding with what became of it.
>
> The calibration page: what was read this week, what has no verdict yet, every
> question by how often it was right, and whether a probability means what it
> says.

## The table

Counted over findings, never over readings. One question raised again by a
later look at the same change is one finding with one key — counting
`raised` per reading made a question that fired on a single change
twenty-one times look like twenty-one mistakes.

| question | fired | judged | confirmed | wrong |
| --- | --- | --- | --- | --- |
| `did_what_was_asked` | 22 | 20 | 2 | 18 |
| `test_missing` | 9 | 9 | 3 | 6 |
| `path_traversal` | 4 | 3 | 0 | 3 |
| `error_swallowed` | 4 | 3 | 1 | 2 |
| `silent_failure` | 4 | 3 | 2 | 1 |
| `kind_fixture` | 2 | 2 | 0 | 2 |
| `mocked_git` | 1 | 1 | 1 | 0 |
| `dead_setting` | 1 | 1 | 0 | 1 |
| `authz_removed` | 1 | 1 | 0 | 1 |
| **total** | **48** | **43** | **9** | **34** |

Two rows have enough verdicts to be read as a rate: `did_what_was_asked` at
2 of 20, and `test_missing` at 3 of 9. The other seven say their counts and
no percentage, because five verdicts is where one more stops moving the
figure by a quarter, and `0%` over one verdict is a fact nobody has.

That 10% is the loudest question in the pack being wrong nine times in ten.

## The nine

Three of them are worth naming, because they are the kind of thing a compiler
does not catch and a person reading a diff at speed does not either.

**A research agent that quietly built a website.** A task was briefed as
research only — *the deliverable is `~/tade-web/docs/` with your research,
your decisions and sketches*, with implementation to follow after review. It
built the site instead: `3f5ba30`, then `6ebab82` adding `index.astro`,
`llms.txt.js`, `sitemap.xml.js` and `site.css`. Good work, and not what the
task allowed. A second task briefed *research only — write no code, change
nothing but your own document* rewrote the README's headline in `c0ca939`.
Neither is a bug. Both are an agent doing something nobody asked for, which
is the thing that is hardest to see when the result is an improvement.

**A swallowed error in a credential read.** `credentials?.().catch(() => null)`
in the window's own wiring: the error was caught, nothing reported it
anywhere, and the read degraded to an empty answer. Empty is
indistinguishable from *this machine has no sign-ins*. It degrades to the last
good answer now, with sixteen lines of test in
`packages/app/test/wire/machine.test.ts` holding that.

**A silent failure in the connectivity check itself.** `Network.looks()`
called `machine.route()` on the schedules pass, and nothing catches that pass
— `app.ts` runs it as `void runDue().then(...)`. A throw from asking this
machine about its own network interfaces would have ended the pass for every
schedule behind it without a word. The watches would simply have stopped
looking. Fixed in `9889f1e`, and its commit message is careful about which way
to err: with no answer the watch looks anyway, because pausing every watch on
a bug in the network probe is the one failure worse than the noise.

## What the table changed

A calibration table that nobody acts on is a vanity metric. This one was acted
on, and the commit says so in its own words.

`0d89cff`, 25 September, under the heading **The pack learns that
documentation is not code**:

> Every verdict written so far named the same failure: `did_what_was_asked`
> firing because a change touched `AGENTS.md` or a `.claude/skills` recipe,
> which is exactly what this guide *requires* of finished work, and
> `test_missing` firing on ten lines of prose.

So the question was reworded. Not the proposition — `ask` stays one literal
question — but the `means` beside it, which is where "when the proposition
alone is not enough" belongs and which reaches the judge either way. Both
questions now say outright that a Markdown file, a README, a recipe under
`.claude/skills`, a comment, a changelog or a licence notice is not code, and
that documentation written beside the code a change is about is part of doing
what was asked.

The rubric is fingerprinted, and `means` is inside the fingerprint, so a
finding raised before that wording and one raised after it are told apart
rather than averaged together — each of the eighteen is filed under the rubric
it was actually asked with. The judge itself is pinned to a version and never
an alias, `jev-1.13.0`, for the same reason: thresholds are tuned against one
version's distributions.

## What it cost

Two hundred and twenty-eight readings, 2,235 file-reads between them, 223
requests to the judge, **$0.14** for the week. That figure is not Tade's
arithmetic: the judge prices each request itself, and what it said is recorded
with the reading.

The expensive part of a review loop was never the model. It is the attention
of whoever reads what it produced — and thirty-four false positives over seven
days is the number that has to come down, which is why it is on a page rather
than in somebody's memory.

## What a judge is not allowed to do

It may only ever add caution. It never approves, merges, closes, unholds or
shortens anything. What a judge wrote about somebody's code reaches the agent
as material to judge, never as instruction, and the same is true of a review
comment: text a model wrote carries no permission with it.

A verdict is the only thing that closes a finding, and it is never the writing
agent's to give. All 43 here were written by the orchestrator rather than by
hand, which is the arrangement working as designed and also the thing to watch:
the rule that matters is not *a human ruled*, it is *whoever ruled was not the
one being measured, and had to cite what decided it*.

Five of the forty-eight are still open as of this writing, and the page says so
rather than rounding them into either column.

```sh
npm i -g tade-sh
```

The pack, its fifteen questions and the arithmetic behind that table are
`packages/extensions/jev/` in
[the repository](https://github.com/mujacica/tade).
