Skip to content

feat(api): scale-out phase B, runs that outlive their API process - #669

Merged
jonwiggins merged 25 commits into
mainfrom
feat/scale-out-runs
Oct 11, 2026
Merged

jonwiggins merged 25 commits into
mainfrom
feat/scale-out-runs

Conversation

@jonwiggins

@jonwiggins jonwiggins commented Oct 10, 2026 •

Copy link
Copy Markdown
Owner

Scale-out phase B (docs/plans/scale-out.md §1–2): agent runs outlive the API process that started them. Every run kind starts under the run protocol from phase A and is streamed through one attach loop (services/run-attachment.ts), and any instance can pick a run up where its rows stop: no duplicate line, no gap, cost added once.

Merged with origin/main (C1 #668, C2 #667): C1's claim transactions and C2's chat-turn leases are kept as they are; the attach sweep is listed in REPEAT_SCHEDULERS.

The shared machinery

  • Hold / lease. The claiming instance takes attached_by / attach_lease_until (60 s, renewed every 20 s with WHERE attached_by = me). A renewal that matches nothing severs the stream at once and the worker returns. On SIGTERM, detachAll severs every stream and sets every lease to now() before the workers close.
  • Start. One update records the start: exec_state='started', exec_pid, consumed_bytes=0, the lease. It is a CAS on the holder, so a run that was queued again meanwhile gets its new agent stopped instead.
  • Stream (streamRun). Attaches from byte 0. Lines that end at or before consumed_bytes only feed the in-memory derivations: the PR tool-call tracker, the session id, the terminal event (which sends the stdin EOF, on a replay too), the result/auth digest and a command's exit receipt. Every other line becomes a log row. Each chunk's rows and the new offset are committed in one transaction (appendRunLogsBatch, a CAS on the holder and the old offset), at line ends, counted in bytes.
  • No allLogs. Each agent adapter now reads its result line by line (createResultParser; parseResult feeds a whole output through it, so the existing tests still apply). inferExitCode and auth-failure detection also work per line (run-output-digest.ts).
  • Settle. settleRun writes exec_state='exited' and the run's spend in one statement, only while this instance holds the run, and only once. If an instance dies after settling but before the run's own transition, the next attach finishes the run without adding the spend again.
  • Attach jobs. runAttachJob is one driver for every kind: claim by CAS (exit when the claim is lost), then either the kind's pod-death path when the pod is gone or the kind's finish.
  • Any instance acts on a run. Cancel looks up the pod and run dir in the database, then calls killRun, for tasks, Job runs (fail hook), reviews (cancel) and agents (archive, or failed by the reconciler). A mid-run message to a task is appended to the agent's stdin by the route. The Redis task-message bus is gone.
  • Run directories. The attached worker removes the run directory once the run is terminal and consumed. The exec-stream path is removed: the inline exec scripts, the EPIPE watchdog and the exit-trap run-home removal. A fresh start in a run directory stops a stale supervisor still running there, so each run directory has one agent.

Per worker

Worker Changes
task-worker.ts (Repo Tasks) Starts with startTaskInRepoPod. The task stays provisioning until the start is recorded, so running always means an agent some instance can attach to; a cancel during setup stops the new agent. Finishes with streamRun and the digest, then settleRun, then the existing PR detection and transitions. Handles attach-task jobs, which move a still-provisioning task to running when the instance that started it died first. A failure in the start exec now fails the run instead of taking the provisioning retry.
workflow-worker.ts (Job runs) Same shape. A run cancelled during setup has its agent stopped. A command Job's exit line is a replayed derivation. Handles attach-workflow-run.
persistent-agent-worker.ts (turns) Drains the turn's messages and moves the agent to running only after the start is recorded (beginTurn), so a turn that never started leaves its messages pending. The turn's spend is added to the agent's total inside the settle transaction. A turn finishing on a paused or archived agent leaves its state alone; the old code moved it back to idle. Handles attach-persistent-agent-turn.
pr-review-worker.ts (review runs) Records the agent type before the start. A run is running only after its start is recorded. A cancelled review is never flipped to failed by its stopped run. Handles attach-pr-review-run.

Reconciler

  • The snapshot carries execution: exec_state, attached_by, attach_lease_until. That is the row's own for tasks and Job runs, the open turn's for agents, and the active run's for reviews.
  • reattach: the run is running/provisioning, its agent started (or its exit is half recorded), and the lease has expired. This enqueues attach__<run>__<expired lease ms>. The executor checks the version under CAS without writing anything; the job is removed when done.
  • requeueRun: the run never started, its lease expired, and the protocol had held it. Each kind's own retry handles it: tasks and review runs go back to queued; a Job run goes through failed to queued without spending a retry; an agent's turn is halted and the agent goes back to idle.
  • A stall counts only while the run is attached.
  • Runs from before the protocol are handled the way boot used to handle them: a task goes to needs attention; a Job run or agent fails as uncertain.
  • The attach sweep runs every 30 s and enqueues reconcile for expired attachments.
  • Boot fails nothing. It re-enqueues queued work, repairs slot counts, sweeps expired attachments and kicks a resync.
  • OPTIO_ATTACH_LEASE_MS, OPTIO_ATTACH_RENEW_MS and OPTIO_ATTACH_SWEEP_INTERVAL_MS can override the lease, renewal and sweep intervals; the e2e suite sets 6 s / 2 s / 1 s.

E2E matrix (apps/api/e2e/scale-out-runs.e2e.test.ts, startApiCluster)

Every case checks the final state, that the log rows are exactly the agent's tick 1…N lines (each once, with no gap), and that the cost was added once.

# Case Result
1 One server, a Repo Task running, SIGKILL, a new server boots ✅ pr_opened, cost once
2 The attached server is SIGKILLed while two others are up ✅ re-attached in under 10 s (the lease is 6 s), completes
3 Three servers, 10 runs (6 tasks on 2 repos, 2 Job runs, 2 turns), the busiest server killed, DB polled every 250 ms ✅ all complete; global ≤ 4 and per-repo ≤ 2 every time
4 Cancel through a server that is not attached (task, Job run, agent archive) ✅ each agent exits 143 long before N
5 A mid-run message through a server that is not attached ✅ the task's agent echoes it once, and it is marked delivered
9 SIGTERM to the attached server ✅ another server holds the run in under 5 s
10 Fake destroy of the pod mid-run ✅ failed with the pod's reason
11 A Job run and an agent turn through cases 1, 4 and 5 ✅ 1 and 4 for both. For 5, the agent takes its message through the inbox and the next turn; Job runs have no message API
12 A boot with runs in flight ✅ nothing fails, nothing is taken over

Runs: 3 green on the branch before the merge, 1 inside the full e2e tier and 3 more green on the merged tree (about 100 s each). The first-ever run found three bugs, all fixed: a reconciler agent archive or failure did not stop the turn's agent, a lost task's error message was not recorded, and the test was too strict about a legitimate requeue.

Tiers (on the merged tree)

  • pnpm turbo test: 13/13 packages green; API 189 files
  • pnpm --filter @optio/api test:integration: 55 files, 455 tests green. This includes the new run-attachment.int.test.ts: two instances, as two module graphs, race the claim 25 times and exactly one wins each time. It also covers the live lease, the loser's log batch rolling back, spend settled once, and the queued reset.
  • pnpm --filter @optio/api test:e2e: 26 files, 120 tests green. The scale-out file was also run 3 more times, all green.
  • pnpm format:check and pnpm turbo typecheck: green. Shared types under packages/shared/src/types are unchanged, so no Swift or Kotlin regeneration was needed.
  • run-protocol.linux.test.ts, which is skipped on macOS, was run in a node:22-bookworm container: 6/6 green, including a new test for a fresh start stopping a stale supervisor. Its 200 KB-line case depends on the locale (${#l} counts characters), so it passes under C.UTF-8 and fails under POSIX, on main as well.

Plan corrections

  • running means the agent started. A task or review run stays provisioning, and an agent stays provisioning for its turn, until the start is recorded. Otherwise a running run with no agent would be ambiguous.
  • Every transition back to queued clears the attachment columns, so a new attempt never inherits the last attempt's expired lease or exited state.
  • The attach job id is attach__<run>__<expired lease ms>, not attach:<run>:<attempt>. BullMQ rejects : in custom ids except in a legacy three-part form it plans to drop, and runKey already uses __.
  • The executor applies reattach under a version CAS that writes nothing, so a worker finishing the run at the same moment never loses its own CAS.
  • Job runs have no mid-run message API. An agent's messages reach it through its inbox and the next turn, so case 5 covers Repo Tasks, and agents only through the inbox. An agent's "cancel" is its archive.

Left as is

  • The resync kick and the attach sweep only enqueue reconcile passes; they do not take a poller lease, because overlapping passes are harmless.
  • A task's start exec that dies mid-setup is requeued. If its supervisor did launch, the next start stops it; for pooled runs, the start guard refuses with 75 and the attempt fails as uncertain.

Review fixes

Fifteen findings from the code review, each with a test. Commits: b9ceb58c (protocol scripts), 8618bd87 (digest), d61d4995 (late messages), cb7f34c8 + 2a38646e + 1337008e (attach loop and workers), e5ea0a18 (e2e), a75b6839 (docs).

A design change under several fixes: attached_by names the hold, not just the instance (<instance>/<nonce>). Every column write (renew, start, log batch, settle, close) is a CAS on the hold's own mark. Without it, a force-redo picked up by the same instance let the old attempt's job renew, append to and settle the new attempt's row, since both jobs matched attached_by = me.

# Finding Fix Test
1 Force-redo skipped the attempt reset One helper, resetRunAttempt (run-log-service), for every path back to queued: a task's transition, its force-redo (reset, log delete and event now in one transaction), a Job run's retry and requeue, a review run's requeue. It locks and reads the row, clears the attachment columns in the same transaction, and after commit TERMs the last attempt's agent if it was started, scoped to its pid (stopPreviousAttempt). int: queue again, force-redo, Job run requeue each signal pid 4242/6262; e2e 13
2, 3, 14 Agents orphaned on error paths One rule in run-attachment, endRunJob, called first in every worker's finally and in runAttachJob. If the job started or claimed an agent and ends without finishing the run or leaving it for another instance (detached, shutdown), the agent gets TERM (pid-scoped) unless its exit is recorded or another hold streams it under a live lease. Then the lease and the slot are given back. The per-kind signals after a failed recordRunStarted or running transition are gone. User cancel paths are unchanged. unit: agentLeftBehind decisions, signal before release, reset/finished/detached; int: real rows
4 Same-instance double claim The claim CAS takes only an expired or closed lease, whoever held it (the attached_by = me clause is gone). runAttachJob checks the in-process registry (holdsRun) first. int: a second claim by the same instance loses; unit: attach job skips without a claim
5 Claiming a finished run The claim CAS requires an active run (provisioning/running, or a turn not finished). An exited run that is still active can be claimed; it finishes as already, with no spend added. int: five terminal states never claimed, the row untouched; a failed Job run; a half-settled run finishes without re-adding cost
6 Review cancel during setup The worker checks the parent review before the agent launches, between the start and recording it, and right after. If cancelled, it marks the run cancelled and endRunJob stops the agent. A review cancel also moves its queued and provisioning runs to cancelled. A run completes, and its results are written, only while the run and its review are still active (transitionRunFromActive). e2e 14 (slow start + cancel: run cancelled, agent exit 143, no summary, verdict or agent_drafted); fails without the fix
7 detachAll released settling holds Holds have a phase: setup → streaming → settling (once the stream ended in an exit or loss). SIGTERM lets setup/streaming go at once and waits up to 20 s (OPTIO_DETACH_SETTLE_WAIT_MS) for settling ones, which keep renewing. A run let go while its job was still starting it gives the fresh lease up again. unit: streaming released, settling awaited, bound; e2e 16 (another server never holds it); fails without the fix
8 Stale supervisor on a fresh start TERM → up to 5 s → KILL → up to 5 s. If it is still alive (zombies don't count), the start fails without touching anything. The old attempt's files are moved to <runDir>.prev-<epoch>-<pid> and removed in the background; files are never truncated in place. The agent script and supervisor are written only after that. The fake runtime mirrors this and follows an attachment's output by descriptor. Linux (node:22-bookworm, 10/10): TERM-ignoring supervisor KILLed, straggler's late write not in the new output, an unkillable one fails the start with files intact; unit script checks
9 Delivered after EOF The deliver script, holding a lock on the run's stdin, refuses when stdin.ndjson holds the sentinel: it prints __OPTIO_STDIN_CLOSED__ → RunStdinClosedError. The delivery marks the message "The agent already finished its turn; the message was not delivered". A replayed EOF is not an error. Linux; fake; unit (script, error, delivery service, replayed EOF)
10 Re-attach parsed by the live definition A Job run re-attach uses the run row's agent_type (null means command). A turn records agentRuntime in its wake_payload (no migration), and the re-attach reads it, falling back to the agent's runtime. e2e 15 (agent → command edit, holder SIGKILLed: completes, contiguous ticks); fails without the fix
11 Digest exit-code rule Before this PR, tasks and review runs read the exit code from the output alone (first result event, then fatal raw lines; the exit code was never consulted). Now a seen result event is authoritative, then a failure signal, then the process's exit code. Job runs and turns keep the process's exit code, as before. unit: result vs exit 1/2/143, no result, fatal line, codex, no-inference kinds
12 Requeue reported as taken Settle and stream classify a loss: taken (same agent, another hold) or reset (exec state cleared, or another pid). On reset the worker gives its slot back, and endRunJob stops the old agent. unit (stream: reset vs taken); int (settle after requeue, after a new attempt started)
13 stderr lost removeRunDirectory reads the last 8 KB of stderr.log and removes the directory in one exec, logging the tail as a warning. A failed, lost or no-output run reads the tail before its transition (readRunStderr) and appends it to the error (withStderr). unit; e2e 17 ([[mock:stderr:…]] reaches the task's error)
15 removeRunDirectory guard isRemovableRunDir: exactly /home/agent/optio/runs/<segment>, segment ^[A-Za-z0-9_-][A-Za-z0-9_.-]*$, not ./... unit: 12 rejects

Also: optio_run_pid never returns pid 0/1, because kill -- -1 would signal the whole pod.

Disagreements and judgment calls

  • 2/3/14, applied literally to attach jobs. An attach job that throws mid-stream now stops the agent. Before, the run stayed for the next attach. The next attach then finishes it as a failure (exit 143). This matches the pre-protocol behaviour, where a worker error failed the task, and avoids an agent nobody streams. A transient DB blip can now fail a run that would otherwise have survived.
  • 13, one exec vs two. Reading and removal share one exec, as decided. A failure message needs the tail before the run's transition, and removing the directory before the transition would make a crash in between unrecoverable. So failed, lost or silent runs pay one extra short read exec.
  • 11, agents without result events (Codex, Copilot, OpenCode, Gemini, OpenClaw). The old task path ignored the process exit code entirely. I kept "failure signal, else exit code" for them: a non-zero exit with no error pattern now fails the run where it used to succeed.
  • 6, queued runs too. A review cancel also cancels queued runs, not only provisioning ones (the worker would cancel them on pickup anyway).
  • Hold tokens (above) go beyond the review's text. Item 4's clause removal alone doesn't stop one instance's old and new attempt jobs from passing for each other.

Tiers after the fixes

  • pnpm turbo test: 13/13 packages green. API: 190 files, 3096 tests. container-runtime: 179 tests, plus 10 skipped (the Linux file).
  • pnpm --filter @optio/api test:integration: 55 files, 462 tests green. run-attachment.int.test.ts has 13 tests, 7 of them new.
  • pnpm --filter @optio/api test:e2e: 26 files, 125 tests green.
  • scale-out-runs.e2e.test.ts (now 13 cases, a third cluster for 13–17): 4 green runs on the final tree (one inside the full tier, then 3 more in a row), plus earlier runs while building. Cases 14, 15 and 16 were also checked against the code with their fix reverted: each fails. Case 16 needed a holder sampler, because the conditional review-run transitions alone hide the second finisher.
  • run-protocol.linux.test.ts in node:22-bookworm (C.UTF-8): 10/10 green, 4 of them new.
  • pnpm format:check, pnpm turbo typecheck: green. No shared types changed, so no Swift or Kotlin regeneration was needed.

Review fixes, round 2 (c77f6119, 032981e6, e0dad4c4, 7683cb13, a0a082ef, bc3e6003)

Streaming errors detach instead of killing. Round 1 applied item 2's rule to every error. Now endRunJob and the workers split errors by where the job failed:

  • Before streaming (setup, the start script, the start's own state changes): the rule is unchanged. The agent the job started is stopped and the run fails as before.
  • Once the agent streams or settles (a DB error writing logs, a publish error, a parser exception): onStreamError gives the lease up (attach_lease_until = now()). It does not signal the agent or fail the run. The reconciler re-attaches.
  • The cap. Attach jobs that fail one attempt are counted under attachFailures in the row's existing jsonb (no migration): a task's or Job run's metadata, a review run's metadata, a turn's wake_payload. The counter is a CAS on the hold, and the column's other keys are kept. It is cleared when the attempt's start is recorded and when an attach job commits progress; the starting job's own failure does not count. The third failed attach in a row fails the run the kind's way (onAttachFailed: task attach_failed, Job run recoveryRequired so it is not auto-retried, turn halted and agent failed, review run and review failed) with "Attaching to the run failed 3 times in a row; last error: …". endRunJob then stops the agent.
  • Unchanged: user cancel and the reset path.
  • Tests:
    • unit: the decision, and the attach job's detach and give-up outcomes
    • integration: the counter per hold, its other keys preserved, then detach, detach, fail on the third
    • e2e 18: the starting job's log write fails once mid-stream (OPTIO_TEST_STREAM_FAULT_FIRST on a [[mock:mark:10:FAULT-ONCE]] line). The run is re-attached and reaches pr_opened with contiguous logs, the line stored once, cost once, the same agent pid and exit 0. Checked to fail under the old kill rule.
    • e2e 19: every log write fails. The Job run fails after 3 attaches with the message above, attachFailures = 3, and the agent (still ticking) exits 143.

Item 11 for agents without a result event. Codex, Copilot, OpenCode, Gemini and OpenClaw are back to exactly the pre-PR rule: their failure signals decide, and the exit code is not consulted. The Claude (and Cursor) result-event rule stays. A unit test pins the old Codex behaviour across exit codes 0, 1, 3 and 143 against inferExitCode.

CI flake in case 3 (e0dad4c4). CI failed case 3 at d2058147 (a turn halted error) and at 1337008e (a create answered 500). Both logs show the same thing: the repo-cleanup health check cleaned the scale-b repo pod up as crashed. Cause: every cluster in the file shares one database but had its own fake-runtime directory, so the ready pod rows an earlier cluster left behind pointed at containers that no longer existed. Locally a worker's pod pick usually met them first and dropped them quietly. On CI the 5 s health check sometimes won: it ran the crash path (the stubbed kubectl delete networkpolicy error in the log) and marked the repo's running tasks with that pod's error. The file's clusters now share one fake-runtime directory, as a deployment's pods outlive its API servers. The scale-out file's create assertions now print the response body, and a failure prints 80 log lines per server, so a recurrence is diagnosable.

Hermeticity. Every e2e server was already off any real cluster: kubectl and helm are failing stubs, KUBECONFIG points at a read-only fake API, and the developer's credentials are stripped. GitHub was the gap: the dummy GITHUB_TOKEN the tiers seed was sent to the real api.github.com (PR detection, the reconciler's PR reads). The hermetic env now preloads a fetch guard (NODE_OPTIONS --import). A request to github.com, githubusercontent.com or gitlab.com is answered in-process with the same 401 and noted on stderr. A unit test covers it, and so does a hermetic-server e2e case: a seeded token's GitHub call is answered locally.

Case 3's agent assertion. Running the file again surfaced the error-turn symptom locally in 1 of 4 runs, with the turns printed: the server case 3 kills was still setting an agent's turn up, so the reconciler halted that never-started turn ("went away before its agent started; queued again", no cost, no logs) and the next turn ran to natural. That is recovery working. Case 3 now waits for the turn that ran to its end and checks that any turn before it is exactly that requeue.

A pod row abandoned mid-creation (bc3e6003). CI then failed case 3 once more: an agent never finished a turn in 120 s. The agent pool waited on a pod row left provisioning by a server that died between inserting the row and creating the pod. Every turn waited 120 s and failed, so the agent ended failed. pickPod (repo and Job pods) already took such rows over after 10 minutes. Both now share one rule, provisioningAbandoned: a bare pod's row still provisioning past OPTIO_STALE_PROVISIONING_MS (10 min; a live creator marks it ready or failed within the runtime's 120 s). A waiter (waitForProvisioning) waits only until the row would count as abandoned, then removes it and makes the pod. Tests:

  • unit: the agent pool's three outcomes
  • integration: pickPod takes over a row the moment it goes stale; waitForProvisioning
  • e2e 20: an idle agent's pod row is set back to provisioning, and its next turn runs within seconds. With the old code it fails with CI's symptom.

The scale-out clusters run with a 15 s stale age.

Tiers on the final tree (bc3e6003):

  • pnpm turbo test: 13/13 packages (API 190 files, 3105 tests).
  • Integration: 55 files, 465 tests.
  • E2E tier: 26 files, 129 tests.
  • scale-out-runs.e2e.test.ts (16 cases): 3 green runs in a row after the tier, and 6 more of the three-server block while hunting the flake.
  • Format and typecheck green.
  • CI fully green at bc3e6003 (run 38106457189), E2E included.

Every adapter gets createResultParser(): push(line) as the output streams,
finish(exitCode) once the agent exited. parseResult(exitCode, logs) feeds a
whole output through it, so the existing tests pin the semantics (the PR
URL is the first match, the cost the last; no pattern spans a line). A
worker no longer has to hold a run's whole output to read its result.
…in the fake agent

A fresh start of a run stops a previous attempt's supervisor still alive in
the same run directory before it truncates the output (its starter went
away before recording it): one run directory, one agent. The fake runtime
does the same. The fake agent plays [[mock:ticks:N:MS]] — N numbered text
events after init — so a consumer that changed hands mid-run can prove it
stored each line once. Tests for both, on Linux too.
The world snapshot carries a run's execution — exec_state, attached_by,
attach_lease_until (the row's own; the open turn's for an agent, the active
run's for a review) — and every kind's pure decisions gain two actions:
reattach (running or provisioning, the agent started or its exit half
recorded, nobody holds it) and requeueRun (the instance setting the run up
went away before its agent started). A stall counts only while a run is
attached; a run that predates the protocol is settled the way a boot used
to. Unit tests for every branch.
…-once offset

run-attachment.ts: the instance that claims a run holds its attachment (a
60 s lease renewed every 20 s, WHERE attached_by = me; a renewal that
matches nothing severs the stream). streamRun attaches from byte 0, feeds
every line to the in-memory derivations, and stores the lines past
consumed_bytes as log rows committed with the new offset in one
transaction (appendRunLogsBatch: a CAS on the holder and the old offset),
at line ends, counted in bytes. settleRun records the exit and the run's
spend in one statement, once. runAttachJob claims an expired attachment by
CAS for any kind; detachAll gives every lease up at shutdown. The result,
exit code and auth failure are read line by line (run-output-digest.ts).
Unit tests for offsets, replay and the lease; an integration test of two
instances racing the claim.
Repo Tasks, Job runs, persistent-agent turns and PR-review runs start
through the start script (the first Claude stream-json message is the run's
initial stdin), record the start in one update (exec_state started,
exec_pid, consumed_bytes 0, the lease: a CAS on the holder, so a run queued
again meanwhile stops the agent instead), and finish through the shared
attach loop: no allLogs, the agent's turn ended by the stdin EOF line, an
exited agent settled through the kind's existing completion, a lost one
through its pod-death path, a detached stream left for the next instance.
Each kind takes attach jobs (attach-task, attach-workflow-run,
attach-persistent-agent-turn, attach-pr-review-run) through the same
driver. A task or review run is running (an agent running a turn) only once
its agent started; a start-exec failure is the run's, not provisioning's.
A transition back to queued clears the attachment for the new attempt.

Any instance acts on a run: cancel (a task, a Job run, a review, an
archived or reconciler-failed agent) signals the agent's process group
found from the database; a mid-run message to a task is appended to its
agent's stdin by the route; the Redis message bus is gone. The run queues
use lockDuration 60 s / stalledInterval 30 s.

The exec-stream run path (execTaskInRepoPod, execRunInPod, execTurnInPod,
the inline exec scripts, the EPIPE watchdog, the exit-trap run-home
removal) is removed; the attached worker removes a run's directory once it
is terminal and consumed.
The snapshot builder reads each run's execution, and the executor applies
reattach (a version check that writes nothing, then the kind's attach job,
id attach__<run>__<expired lease ms>, removed when done) and requeueRun
(the kind's own retry: a task or review run to queued, a Job run through
failed to queued without spending a retry, a turn halted and its agent back
to idle with its messages still pending). An attach sweep (every 30 s, one
upserted job scheduler) gives every run whose attachment expired a
reconcile pass. On SIGTERM an instance gives its runs up before its
workers close. Boot no longer fails or flags runs in flight: it re-enqueues
queued work, repairs slot counts, sweeps expired attachments and kicks a
resync. Lease, renew and sweep intervals take OPTIO_ATTACH_* overrides.
…them

Real API servers on one database and one fake-runtime directory
(api-cluster.ts), with a 6 s lease and a 1 s sweep: SIGKILL of a run's only
server and a new one booting; the attached server killed with two others up
(another attaches within 10 s); ten runs across kinds with one server killed
and the limits polled from the database every 250 ms; cancel and a mid-run
message through a server not attached to the run; SIGTERM handing a run
over in under 5 s; the pod destroyed mid-run; a Job run and an agent turn
through kill, cancel and message; a boot with runs in flight. Every case
checks the final state, the log rows (each tick once, contiguous) and the
cost added once.
docs/reconciliation.md explains how runs are held, streamed, settled and
handed over, and the reattach / requeueRun decisions. The plan records
phase B as built and what building it found.
The cancellation tests imported the whole attachment module through the
real signalRun, which made them slow enough to time out under a loaded
test run. They now mock signalRun, and run-attachment.test.ts covers its
pod lookup and runtime call.
The attach sweep is a scheduler REPEAT_SCHEDULERS knows (scheduleRepeat),
so boot's cleanup keeps it. C1's claim transactions and C2's chat-turn
leases stay as they are.
…s the old attempt aside

A stale supervisor in the run directory gets TERM, up to 5 s, KILL, up to
5 s more; one that survives fails the start with nothing touched. The old
attempt's files are moved to a sibling directory (removed in the
background) instead of being truncated in place, so a straggler writing
through an open descriptor never lands in the new attempt's output.

The deliver script refuses a line once the run's stdin holds the EOF
sentinel (RunStdinClosedError), under a lock so a line and the sentinel
never race. killRun can name the attempt's pid, so a newer attempt in the
same directory is never hit. The fake runtime mirrors all of it, follows
an attachment's output by descriptor, and takes [[mock:slow-start:MS]].
Before the run protocol, tasks and review runs read their exit code from
the output alone: the first Claude-style result event decided (is_error),
and a fatal raw line or an agent's failure signal otherwise. The digest
took the process's exit code over a successful result, so a turn that
reported success and then exited non-zero failed. Now a seen result event
is authoritative, a failure signal fails the run, and only an output that
says neither leaves it to the exit code. Job runs and agent turns keep the
process's exit code, as they always did.
The deliver script refuses a line once the run's stdin holds the EOF
sentinel; the route's delivery marks the message with "The agent already
finished its turn; the message was not delivered" instead of delivered.
- Holds: attached_by names the hold (<instance>/<nonce>) and every column
  write is a CAS on it, so one instance's old-attempt and new-attempt jobs
  are never mistaken for each other. A claim takes only an expired (or
  closed) lease of a run still active (provisioning/running; a turn not
  finished): a run already terminal is never claimed, and an attach job
  exits at once when a job of this process still holds the run.
- One rule stops an agent its job left behind (endRunJob): a job that
  started or claimed an agent and ends without finishing the run or
  leaving it to another instance (an error halfway, a cancel or requeue
  under it) TERMs that agent, scoped to its supervisor's pid, before it
  gives the lease or the pod slot back. The per-kind kill hooks it made
  redundant go; this covers the persistent-agent and review-run error
  paths that left agents running.
- Every path back to queued (a task's transition and its force-redo, a Job
  run's retry or requeue, a review run's requeue) goes through
  resetRunAttempt: the attachment columns cleared in the same transaction,
  the last attempt's started agent TERMed (pid-scoped) after the commit.
  A force-redo's reset, log removal and event are one transaction.
- A settle or stream that lost its run says taken (another hold has the
  same agent) or reset (queued again, or a new attempt started); a reset
  job gives its pod slot back.
- Shutdown lets runs being set up or streamed go at once and waits (up to
  20 s) for runs past their stream to finish settling, their leases kept.
- A review cancelled while its run is set up: the worker checks the review
  before the agent launches, before its start is recorded and right after,
  cancels the run and stops the agent; a review cancel also cancels its
  queued and provisioning runs; a run only completes, and its results are
  only written, while it and its review are still active.
- A re-attach reads the output by what the attempt ran: a Job run's
  recorded agent type, a turn's runtime (now recorded in its wake payload).
- Before a run directory is removed, the tail of the agent's stderr is read
  in the same exec and logged; a failed, lost or silent run's error message
  carries it. removeRunDirectory accepts only runs/<one plain segment>.
…tling

New cluster cases: a force-redo of a running task stops the old agent and
keeps only the new attempt's logs; a review cancelled while its run is set
up cancels the run, stops the agent and writes nothing; a Job edited from
an agent to a command mid-run is still read as the agent on re-attach;
SIGTERM while a review run settles finishes it on that server, once (no
other server ever holds it); a failed run's error carries its stderr tail.
The fake runtime takes [[mock:stderr:TEXT]] and answers a stderr-tail exec.
…s its lease up

The start, recorded after shutdown released the hold, takes a fresh lease;
endRunJob gives it up again for a released hold, so another instance
attaches within a sweep instead of after a lease period.
For tasks and review runs, Codex, Copilot, OpenCode, Gemini and OpenClaw
were judged by their failure signals alone before the run protocol: the
process's exit code was never consulted. The digest consulted it when no
signal fired; it no longer does. Claude and Cursor keep the result-event
rule.
The point of the run protocol is that a run survives trouble on the API
side. endRunJob's rule now depends on where the job failed: an error
while the run is set up (setup, the start script, the start's own state
changes) keeps the rule (the agent it started is stopped, the run failed
as before); an error once the agent streams or settles (a database error
writing the logs, a publish, a parser) gives the lease up without
signalling the agent and leaves the run for the reconciler to re-attach
(onStreamError). Attach jobs that fail one attempt are counted in the run
row's existing jsonb (attachFailures; cleared when the attempt starts and
when an attach job makes progress; the starting job's failure does not
count): the third in a row fails the run with the last error, the kind's
way, and its agent is stopped through the rule. User cancel and the reset
path are unchanged.

Tests: unit (the decision, the attach job's two outcomes), integration
(the counter per hold, preserving the column's other keys; detach, then
fail on the third), e2e 18 (a log write fails once: re-attached, same
agent, contiguous logs, cost once) and 19 (fails every time: failed after
3 attaches, agent stopped). The fake runtime takes [[mock:mark:N:TEXT]];
OPTIO_TEST_STREAM_FAULT[_FIRST] are the e2e tier's fault knobs.
…s never reach GitHub

CI flaked in case 3 (ten runs, three servers): every cluster in the file
works on one database but had a fake-runtime directory of its own, so the
pod rows an earlier cluster left ready (scale-a, scale-b, its Jobs' and
agents' pods) pointed at containers that no longer existed. Whether a
worker's pod pick or the 5 s health check met them first was timing: on
CI the health check sometimes won, cleaned the scale-b pod up as crashed
(the stubbed kubectl's network-policy error in the log) and marked the
repo's running tasks with that pod's error. The clusters now share one
directory, as a deployment's pods outlive its API servers.

The hermetic env (every startApiServer) also preloads a fetch guard: a
request to github.com, githubusercontent.com or gitlab.com is answered in
the process with the 401 a dummy token gets, and noted on stderr, so the
dummy GITHUB_TOKEN the tiers seed never leaves the machine (kubectl was
already a failing stub and the Kubernetes API a read-only fake). A unit
test and the hermetic-server e2e prove it. The scale-out file's create
assertions now print the response body, and a failure prints 80 log lines
per server.
…died

When the server case 3 kills was still setting an agent's turn up, the
reconciler halts that turn (it never ran: no cost, no logs) and the
agent's next turn does the work; the case read the halted turn and
failed, as it once did for a task requeued the same way. It now waits for
the turn that ran to its end and checks that any turn before it is exactly
that recovery. Reproduced locally (1 in 4 runs) with the turns printed.
A server that dies between a pod's row and the pod leaves the row
provisioning forever. pickPod took such a row over after 10 minutes, but a
persistent agent's pod acquisition waited on it every turn (120 s, then
failing the turn) and never recovered: the agent ended failed. CI's case 3
hit it when the server it kills was creating a new agent's pod. Both now
share one rule (provisioningAbandoned: a bare pod's row still provisioning
past OPTIO_STALE_PROVISIONING_MS, 10 min; a live creator marks it within
the runtime's 120 s), and a waiter waits only until the row would count as
abandoned, then removes it and makes the pod (waitForProvisioning).

Tests: unit (the agent pool's three outcomes), integration (pickPod takes
over a row the moment it goes stale; waitForProvisioning), e2e 20 (an
idle agent's pod row set back to provisioning: its next turn runs within
seconds; fails without the fix with CI's symptom). The scale-out e2e
clusters run with a 15 s stale age.
@jonwiggins
jonwiggins merged commit 2d51d24 into main Oct 11, 2026
13 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant