Long unattended agent runs: where keel stops them, and why
A coding agent left to run for hours rarely fails by crashing. It fails by carrying on: past a question nobody answered, past a check that could not run, into a merge. So the useful question about an unattended run is not whether it finished. It is whether, when it could not finish correctly, it stopped at a place someone will look, with the reason written down.
keel is a workflow backbone for that. Your agent host runs /keel:ship on a GitHub issue and does the work; the hosts are Claude Code, Codex, Cursor (partial) and Antigravity. At each fixed step the keel CLI decides whether the run may go on. This page maps six practices for long runs to what keel does for each, then shows one run of keel's own that stopped.
Six practices, and what keel does for each
| Practice for long agent runs | What keel does |
|---|---|
| Define "done" first | An issue with no deliverable or no acceptance criteria comes back needs-input with questions, and /keel:ship stops before it cuts a branch (intake.py, ship Step 0) |
| Name the stops | A live run declares its consent scopes (filesystem, git, github, …) before it acts; merges go only through keel merge, inside the merge window (an audited --hotfix is the one bypass) |
| Keep state in a file | A resumable checkpoint (keel checkpoint, keel resume) and an append-only run ledger (.keel/state/run-ledger.jsonl) |
| Fan out, then check | /keel:regression runs parallel reviewers and keeps low-confidence findings as review-only instead of filing them |
| Review before a person does | Review verdicts are pinned to the head SHA; reviewers can run on another vendor, and knobs.evidence_require_distinct_vendors makes distinct reviewer vendors a requirement |
| Read what's blocked first | /keel:morning puts the cross-session deferrals at the top of the briefing |
These map to the long-run advice in Anthropic's Getting the most out of Opus 5.5 guide; read it there. They hold for any agent. keel is independent and not affiliated with Anthropic.
One run that stopped: four pull requests and a missing verdict
On 25 August 2026 keel's own repository had four pull requests at the merge step. keel sorts changes into risk tiers, and tier 3 covers its release toolchain and its core modules. At the time a tier-3 merge required three reviewer verdicts and a verdict from the cross-vendor jury, a panel run of three agent CLIs that costs money every time it runs.
The four were two dependency pin bumps (#919, #920), a whole-tree reformat (#967) and new parsing logic for untrusted input (#958). The pre-merge evidence check refused each one and named what was missing. For the idna bump, before its reviewers had posted:
$ keel evidence-verify projects/keel.yaml --root . --pr 920 --phase pre-merge
required : 4
missing : review-verdict-1, review-verdict-2, review-verdict-3, jury-verdict
Once the reviewer verdicts were in, two of the four still stopped on one item: jury-verdict. Nothing merged.
The agent driving the run had two ways past the gate: a one-flag edit to the workflow, or the evidence-waiver label. It took neither. Removing a check that blocks the agent's own pull requests is a decision about what tier 3 means, so it opened #965, "DECISION NEEDED", with four options, what each would cost, and a table of the four pull requests.
It was settled the same day. The jury became advisory at tier 3, and the three reviewer verdicts stayed required (#968). A jury stays the tool for changes where independent readings can differ, and #958 was one of those. .github/workflows/keel-ship.yml still passes --no-jury at that call site.
What made it a safe stop: the tool named the reason (jury-verdict), so nobody had to read a log to find it. It stopped at the merge, the one step whose effect is hard to take back. And the way on was a recorded decision, not a waiver that turns into a habit.
What the record also shows
Stopping is not the whole story, and keel's own history says where it fell short:
- A bot reverted a review mid-gate. Fixes pushed onto a bot's branch were wiped by the bot's next push: 222 deletions, including the tests that would have caught the regression. keel did not stop that; the fixes were re-landed in #1126. The rule that a bot's branch is read-only came after (#1127).
- A gate seat caught a fix passing on a ref nobody observed. The first review round on the fix for #1184 blocked it: the change reported a
passagainstorigin/<base>while it had read the local branch. It was reworked before it merged (#1191). - A closed issue is not a verified issue. An audit reverted fourteen closed fixes. At the time their issues were closed, three could be reverted, wholly or in half, without a test failing (#1289). keel now asks every fix for a test that fails as an assertion when that one change is reverted.
Trying it
pipx install keel-workflow
keel setup --root .
# then, in your agent host:
/keel:ship <issue>
The Quickstart covers the first issue, and install.md covers each host. On Cursor a local checkout registers one skill, so use the marketplace route for the commands (#1332).