blog /
From a red main to green in thirty-six minutes
On 28 September ubuntu-latest went red on Tade's own main at 08:15. A watch found it at 08:20, an agent removed the race, and CI was green again by 08:51.
A red main is the failure nobody owns. It is not on anybody’s review, no
robot comments on it, and the person who broke it has moved on to the next
thing. It gets noticed when somebody else pulls and their own tests fail.
So Tade watches it. The whole of 28 September 2026, in UTC, from GitHub’s own
check runs and ~/.tade/events.jsonl:
| 08:15:13 | check (ubuntu-latest) reports failure on 8b71105 |
| 08:20:48 | the watch review.branch-checks finds it, and starts one agent |
| 08:43:35 | the agent commits a1c4b6f — 2 files, +29 −2 — and pushes |
| 08:44:12 | CI starts on the new commit |
| 08:51:32 | check (ubuntu-latest) reports success |
| 08:51:52 | the agent says what it did and stops |
Thirty-six minutes from red to green, and nobody went looking. What follows is each row of that, and where to check it.
The half of CI that has nothing to hang off
Tade already watched CI on the reviews somebody opened. That watch reads a forge and asks about pull requests. A project whose work goes straight to its base branch opens none of those, so the run that decides whether the branch everybody else pulls is broken was watched by nobody.
review.branch-checks is that run, watched. It looks every ten minutes, and
it stands itself up wherever there is a forge rather than waiting to be
offered — a red main is everybody’s, and nobody should have to notice it by
hand.
It asks git before it asks anybody else. A commit this machine has and the
remote has not is not a failure; it is CI that has not run yet, and rev-list
answers that for nothing. The whole look costs one request, and none at all
where there is nothing pushed.
What the agent was told
The watch does not fix anything. It writes a finding and puts one agent on it,
and the words it hands over are fixed in the source rather than composed by a
model. Verbatim, from packages/extensions/review/src/branch.ts:
CI is failing on the branch this project is on. Reproduce it here before you
change anything — what failed, the commit it failed on and the tail of each
log are in your context file. A test that fails because the code disproved
its premise is fixed by correcting the premise, never by deleting the test;
a race is fixed by removing the race, never by loosening the assertion that
caught it. Then fix the cause and push. Never force-push, never revert
somebody else's commit, and never merge anything: if the fix is not yours to
make, say so and stop.
Three clauses of that are refusals, and they are there because the failure mode of an agent let loose on a red branch is not a bad fix. It is a green branch that proves nothing: a deleted test, a loosened assertion, a revert of somebody else’s morning.
The tail of each failing log goes into the agent’s context file before it starts, at most three of them, read from the forge. So the agent begins with what CI printed rather than with a guess about what CI might have printed.
Why the bug was hard
The failing test was packages/drivers/pty/test/memory.test.ts, and it
asserted something narrow: a lane that closes gives its screen back. A
terminal at two hundred columns and ten thousand lines of scrollback is tens
of megabytes of typed arrays, so a lane nobody can ever read again holding
onto them is a real leak.
It was red on ubuntu, green on macOS, and red only sometimes. That shape is the one that gets a test deleted.
The cause is in xterm. dispose() does not drop what has been written to the
emulator and not yet parsed, and a queued chunk holds the whole screen behind
it: xterm parses in the background off a chain of
setTimeout(() => this._innerWrite()), and that closure reaches the write
buffer, its parse action, the input handler and every Uint32Array of the
buffer. Letting go of Tade’s own reference then lets go of nothing. The
comment the fix left in packages/drivers/pty/src/index.ts puts a number on
it: 1KB still queued kept 24MB of screen alive, until the backlog
happened to drain.
The machine-dependence falls straight out of that. screen() settles the
emulator for up to 50ms before it answers, and this machine always drained
inside that window. A GitHub runner slow enough to still be behind did not.
What the fix was, and what it was not
Two changes, and the second is the one that matters.
release() now calls reset() before dispose(), so what the leftover
parse holds is a fresh, empty buffer rather than the full one.
And the test no longer waits for the lane to go quiet. It closes the lane from inside an output listener: the driver writes to the emulator before it tells anybody, so inside a listener there is always a chunk queued, on any machine. The commit’s own words: without the fix it fails every run, with the same numbers CI reported.
That is the difference between a flake that was made to stop flaking and a
race that was removed. The assertion — after under half of full — is
untouched.
What it cost
The agent ran the project’s own checks through Tade three times before it
finished — checks_run, not a shell, so each run is recorded against the
commit it ran on instead of being described in a message. Then it asked the
judge about its own range, answered what came back, and committed.
What it spent, from the journal:
$ tade spend --days 7
tade/check-ubuntu-latest-failing-on-main-say — 9.8M tokens 1h 35m
Nine point eight million tokens, and an hour and thirty-five minutes of lane time of which about twenty-four minutes were a model actually inside a turn — the lane was still open when that was read, so both figures are floors rather than totals.
The dash is not zero and it is not missing. The run was on a Claude Code subscription, which bills a flat fee, so there is no per-token cost to report and Tade reports none. That is the subject of its own post.
What it saved, honestly
Not engineering time on the fix. Somebody would have had to write that either way, and an agent wrote it under instructions a person set.
What it saved is the interval — thirty-six minutes between CI saying red and
CI saying green, five and a half of which were the watch not having looked yet.
In a week where tests ran 131 times on this machine and failed 52 of them,
nobody had to notice this one, and nobody pulled a broken main in between.
The claim is small and every row of it is checkable. The watch, its prompt and
its refusals are in packages/extensions/review/src/branch.ts. The fix is
a1c4b6f. The two CI conclusions are what GitHub returns for 8b71105 and
a1c4b6f, and the rest is ~/.tade/events.jsonl on the machine that ran it.
npm i -g tade-sh
Tade is MIT, and the source is on GitHub.