# From a red main to green in thirty-six minutes

> On 28 September ubuntu-latest went red on Tade's own main at 08:15. A watch found it at 08:20, an agent removed the race, and CI was green again by 08:51.

Published 2026-09-28 · by The Tade project · tagged ci, watches, checks, agents.
Published at https://tade.sh/blog/a-red-main-fixed-before-anybody-noticed/ — part of https://tade.sh/blog/.

---

A red `main` is the failure nobody owns. It is not on anybody's review, no
robot comments on it, and the person who broke it has moved on to the next
thing. It gets noticed when somebody else pulls and their own tests fail.

So Tade watches it. The whole of 28 September 2026, in UTC, from GitHub's own
check runs and `~/.tade/events.jsonl`:

| | |
| --- | --- |
| 08:15:13 | `check (ubuntu-latest)` reports **failure** on `8b71105` |
| 08:20:48 | the watch `review.branch-checks` finds it, and starts one agent |
| 08:43:35 | the agent commits `a1c4b6f` — 2 files, +29 −2 — and pushes |
| 08:44:12 | CI starts on the new commit |
| 08:51:32 | `check (ubuntu-latest)` reports **success** |
| 08:51:52 | the agent says what it did and stops |

Thirty-six minutes from red to green, and nobody went looking. What follows is
each row of that, and where to check it.

## The half of CI that has nothing to hang off

Tade already watched CI on the reviews somebody opened. That watch reads a
forge and asks about pull requests. A project whose work goes straight to its
base branch opens none of those, so the run that decides whether the branch
everybody else pulls is broken was watched by nobody.

`review.branch-checks` is that run, watched. It looks every ten minutes, and
it stands itself up wherever there is a forge rather than waiting to be
offered — a red `main` is everybody's, and nobody should have to notice it by
hand.

It asks git before it asks anybody else. A commit this machine has and the
remote has not is not a failure; it is CI that has not run yet, and `rev-list`
answers that for nothing. The whole look costs one request, and none at all
where there is nothing pushed.

## What the agent was told

The watch does not fix anything. It writes a finding and puts one agent on it,
and the words it hands over are fixed in the source rather than composed by a
model. Verbatim, from `packages/extensions/review/src/branch.ts`:

```
CI is failing on the branch this project is on. Reproduce it here before you
change anything — what failed, the commit it failed on and the tail of each
log are in your context file. A test that fails because the code disproved
its premise is fixed by correcting the premise, never by deleting the test;
a race is fixed by removing the race, never by loosening the assertion that
caught it. Then fix the cause and push. Never force-push, never revert
somebody else's commit, and never merge anything: if the fix is not yours to
make, say so and stop.
```

Three clauses of that are refusals, and they are there because the failure
mode of an agent let loose on a red branch is not a bad fix. It is a green
branch that proves nothing: a deleted test, a loosened assertion, a revert of
somebody else's morning.

The tail of each failing log goes into the agent's context file before it
starts, at most three of them, read from the forge. So the agent begins with
what CI printed rather than with a guess about what CI might have printed.

> **A screen from Tade — `a-shell-beside-the-agent`.** A shell open beside the agent, in the same worktree, with a divider you can drag.
>
> Reproducing is a shell in the same worktree as the agent, beside it, with a
> divider you can drag — not a second window somebody has to keep in step.

## Why the bug was hard

The failing test was `packages/drivers/pty/test/memory.test.ts`, and it
asserted something narrow: a lane that closes gives its screen back. A
terminal at two hundred columns and ten thousand lines of scrollback is tens
of megabytes of typed arrays, so a lane nobody can ever read again holding
onto them is a real leak.

It was red on ubuntu, green on macOS, and red only sometimes. That shape is
the one that gets a test deleted.

The cause is in xterm. `dispose()` does not drop what has been written to the
emulator and not yet parsed, and a queued chunk holds the whole screen behind
it: xterm parses in the background off a chain of
`setTimeout(() => this._innerWrite())`, and that closure reaches the write
buffer, its parse action, the input handler and every `Uint32Array` of the
buffer. Letting go of Tade's own reference then lets go of nothing. The
comment the fix left in `packages/drivers/pty/src/index.ts` puts a number on
it: **1KB still queued kept 24MB of screen alive**, until the backlog
happened to drain.

The machine-dependence falls straight out of that. `screen()` settles the
emulator for up to 50ms before it answers, and this machine always drained
inside that window. A GitHub runner slow enough to still be behind did not.

## What the fix was, and what it was not

Two changes, and the second is the one that matters.

`release()` now calls `reset()` before `dispose()`, so what the leftover
parse holds is a fresh, empty buffer rather than the full one.

And the test no longer waits for the lane to go quiet. It closes the lane
from **inside an output listener**: the driver writes to the emulator before
it tells anybody, so inside a listener there is always a chunk queued, on any
machine. The commit's own words: *without the fix it fails every run, with
the same numbers CI reported.*

That is the difference between a flake that was made to stop flaking and a
race that was removed. The assertion — `after` under half of `full` — is
untouched.

## What it cost

The agent ran the project's own checks through Tade three times before it
finished — `checks_run`, not a shell, so each run is recorded against the
commit it ran on instead of being described in a message. Then it asked the
judge about its own range, answered what came back, and committed.

> **A screen from Tade — `what-an-agent-has-done`.** The ACTIONS tab: the commits this agent made, what is not committed, the review it is out for, and each of the project’s own checks with what it ran, how long it took and what it counted — with the tests red.
>
> What an agent did, read back rather than remembered: its own commits, and each
> check with the command it ran and what it counted — red at the commit in hand,
> with the failing tail a click away.

What it spent, from the journal:

```console
$ tade spend --days 7
  tade/check-ubuntu-latest-failing-on-main-say    —   9.8M tokens   1h 35m
```

Nine point eight million tokens, and an hour and thirty-five minutes of lane
time of which about twenty-four minutes were a model actually inside a turn —
the lane was still open when that was read, so both figures are floors rather
than totals.

The dash is not zero and it is not missing. The run was on a Claude Code
subscription, which bills a flat fee, so there is no per-token cost to report
and Tade reports none. That is the subject of [its own
post](/blog/three-kinds-of-dollar-each-one-marked/).

## What it saved, honestly

Not engineering time on the fix. Somebody would have had to write that either
way, and an agent wrote it under instructions a person set.

What it saved is the interval — thirty-six minutes between CI saying red and
CI saying green, five and a half of which were the watch not having looked yet.
In a week where `tests` ran 131 times on this machine and failed 52 of them,
nobody had to notice this one, and nobody pulled a broken `main` in between.

The claim is small and every row of it is checkable. The watch, its prompt and
its refusals are in `packages/extensions/review/src/branch.ts`. The fix is
`a1c4b6f`. The two CI conclusions are what GitHub returns for `8b71105` and
`a1c4b6f`, and the rest is `~/.tade/events.jsonl` on the machine that ran it.

```sh
npm i -g tade-sh
```

Tade is MIT, and the source is
[on GitHub](https://github.com/mujacica/tade).
