Repository navigation
feat(api): scale-out phase B, runs that outlive their API process - #669
Merged
Merged
Conversation
Every adapter gets createResultParser(): push(line) as the output streams, finish(exitCode) once the agent exited. parseResult(exitCode, logs) feeds a whole output through it, so the existing tests pin the semantics (the PR URL is the first match, the cost the last; no pattern spans a line). A worker no longer has to hold a run's whole output to read its result.
…in the fake agent A fresh start of a run stops a previous attempt's supervisor still alive in the same run directory before it truncates the output (its starter went away before recording it): one run directory, one agent. The fake runtime does the same. The fake agent plays [[mock:ticks:N:MS]] — N numbered text events after init — so a consumer that changed hands mid-run can prove it stored each line once. Tests for both, on Linux too.
The world snapshot carries a run's execution — exec_state, attached_by, attach_lease_until (the row's own; the open turn's for an agent, the active run's for a review) — and every kind's pure decisions gain two actions: reattach (running or provisioning, the agent started or its exit half recorded, nobody holds it) and requeueRun (the instance setting the run up went away before its agent started). A stall counts only while a run is attached; a run that predates the protocol is settled the way a boot used to. Unit tests for every branch.
…-once offset run-attachment.ts: the instance that claims a run holds its attachment (a 60 s lease renewed every 20 s, WHERE attached_by = me; a renewal that matches nothing severs the stream). streamRun attaches from byte 0, feeds every line to the in-memory derivations, and stores the lines past consumed_bytes as log rows committed with the new offset in one transaction (appendRunLogsBatch: a CAS on the holder and the old offset), at line ends, counted in bytes. settleRun records the exit and the run's spend in one statement, once. runAttachJob claims an expired attachment by CAS for any kind; detachAll gives every lease up at shutdown. The result, exit code and auth failure are read line by line (run-output-digest.ts). Unit tests for offsets, replay and the lease; an integration test of two instances racing the claim.
Repo Tasks, Job runs, persistent-agent turns and PR-review runs start through the start script (the first Claude stream-json message is the run's initial stdin), record the start in one update (exec_state started, exec_pid, consumed_bytes 0, the lease: a CAS on the holder, so a run queued again meanwhile stops the agent instead), and finish through the shared attach loop: no allLogs, the agent's turn ended by the stdin EOF line, an exited agent settled through the kind's existing completion, a lost one through its pod-death path, a detached stream left for the next instance. Each kind takes attach jobs (attach-task, attach-workflow-run, attach-persistent-agent-turn, attach-pr-review-run) through the same driver. A task or review run is running (an agent running a turn) only once its agent started; a start-exec failure is the run's, not provisioning's. A transition back to queued clears the attachment for the new attempt. Any instance acts on a run: cancel (a task, a Job run, a review, an archived or reconciler-failed agent) signals the agent's process group found from the database; a mid-run message to a task is appended to its agent's stdin by the route; the Redis message bus is gone. The run queues use lockDuration 60 s / stalledInterval 30 s. The exec-stream run path (execTaskInRepoPod, execRunInPod, execTurnInPod, the inline exec scripts, the EPIPE watchdog, the exit-trap run-home removal) is removed; the attached worker removes a run's directory once it is terminal and consumed.
The snapshot builder reads each run's execution, and the executor applies reattach (a version check that writes nothing, then the kind's attach job, id attach__<run>__<expired lease ms>, removed when done) and requeueRun (the kind's own retry: a task or review run to queued, a Job run through failed to queued without spending a retry, a turn halted and its agent back to idle with its messages still pending). An attach sweep (every 30 s, one upserted job scheduler) gives every run whose attachment expired a reconcile pass. On SIGTERM an instance gives its runs up before its workers close. Boot no longer fails or flags runs in flight: it re-enqueues queued work, repairs slot counts, sweeps expired attachments and kicks a resync. Lease, renew and sweep intervals take OPTIO_ATTACH_* overrides.
…them Real API servers on one database and one fake-runtime directory (api-cluster.ts), with a 6 s lease and a 1 s sweep: SIGKILL of a run's only server and a new one booting; the attached server killed with two others up (another attaches within 10 s); ten runs across kinds with one server killed and the limits polled from the database every 250 ms; cancel and a mid-run message through a server not attached to the run; SIGTERM handing a run over in under 5 s; the pod destroyed mid-run; a Job run and an agent turn through kill, cancel and message; a boot with runs in flight. Every case checks the final state, the log rows (each tick once, contiguous) and the cost added once.
docs/reconciliation.md explains how runs are held, streamed, settled and handed over, and the reattach / requeueRun decisions. The plan records phase B as built and what building it found.
The cancellation tests imported the whole attachment module through the real signalRun, which made them slow enough to time out under a loaded test run. They now mock signalRun, and run-attachment.test.ts covers its pod lookup and runtime call.
The attach sweep is a scheduler REPEAT_SCHEDULERS knows (scheduleRepeat), so boot's cleanup keeps it. C1's claim transactions and C2's chat-turn leases stay as they are.
…s the old attempt aside A stale supervisor in the run directory gets TERM, up to 5 s, KILL, up to 5 s more; one that survives fails the start with nothing touched. The old attempt's files are moved to a sibling directory (removed in the background) instead of being truncated in place, so a straggler writing through an open descriptor never lands in the new attempt's output. The deliver script refuses a line once the run's stdin holds the EOF sentinel (RunStdinClosedError), under a lock so a line and the sentinel never race. killRun can name the attempt's pid, so a newer attempt in the same directory is never hit. The fake runtime mirrors all of it, follows an attachment's output by descriptor, and takes [[mock:slow-start:MS]].
Before the run protocol, tasks and review runs read their exit code from the output alone: the first Claude-style result event decided (is_error), and a fatal raw line or an agent's failure signal otherwise. The digest took the process's exit code over a successful result, so a turn that reported success and then exited non-zero failed. Now a seen result event is authoritative, a failure signal fails the run, and only an output that says neither leaves it to the exit code. Job runs and agent turns keep the process's exit code, as they always did.
The deliver script refuses a line once the run's stdin holds the EOF sentinel; the route's delivery marks the message with "The agent already finished its turn; the message was not delivered" instead of delivered.
- Holds: attached_by names the hold (<instance>/<nonce>) and every column write is a CAS on it, so one instance's old-attempt and new-attempt jobs are never mistaken for each other. A claim takes only an expired (or closed) lease of a run still active (provisioning/running; a turn not finished): a run already terminal is never claimed, and an attach job exits at once when a job of this process still holds the run. - One rule stops an agent its job left behind (endRunJob): a job that started or claimed an agent and ends without finishing the run or leaving it to another instance (an error halfway, a cancel or requeue under it) TERMs that agent, scoped to its supervisor's pid, before it gives the lease or the pod slot back. The per-kind kill hooks it made redundant go; this covers the persistent-agent and review-run error paths that left agents running. - Every path back to queued (a task's transition and its force-redo, a Job run's retry or requeue, a review run's requeue) goes through resetRunAttempt: the attachment columns cleared in the same transaction, the last attempt's started agent TERMed (pid-scoped) after the commit. A force-redo's reset, log removal and event are one transaction. - A settle or stream that lost its run says taken (another hold has the same agent) or reset (queued again, or a new attempt started); a reset job gives its pod slot back. - Shutdown lets runs being set up or streamed go at once and waits (up to 20 s) for runs past their stream to finish settling, their leases kept. - A review cancelled while its run is set up: the worker checks the review before the agent launches, before its start is recorded and right after, cancels the run and stops the agent; a review cancel also cancels its queued and provisioning runs; a run only completes, and its results are only written, while it and its review are still active. - A re-attach reads the output by what the attempt ran: a Job run's recorded agent type, a turn's runtime (now recorded in its wake payload). - Before a run directory is removed, the tail of the agent's stderr is read in the same exec and logged; a failed, lost or silent run's error message carries it. removeRunDirectory accepts only runs/<one plain segment>.
…tling New cluster cases: a force-redo of a running task stops the old agent and keeps only the new attempt's logs; a review cancelled while its run is set up cancels the run, stops the agent and writes nothing; a Job edited from an agent to a command mid-run is still read as the agent on re-attach; SIGTERM while a review run settles finishes it on that server, once (no other server ever holds it); a failed run's error carries its stderr tail. The fake runtime takes [[mock:stderr:TEXT]] and answers a stderr-tail exec.
…s its lease up The start, recorded after shutdown released the hold, takes a fresh lease; endRunJob gives it up again for a released hold, so another instance attaches within a sweep instead of after a lease period.
For tasks and review runs, Codex, Copilot, OpenCode, Gemini and OpenClaw were judged by their failure signals alone before the run protocol: the process's exit code was never consulted. The digest consulted it when no signal fired; it no longer does. Claude and Cursor keep the result-event rule.
The point of the run protocol is that a run survives trouble on the API side. endRunJob's rule now depends on where the job failed: an error while the run is set up (setup, the start script, the start's own state changes) keeps the rule (the agent it started is stopped, the run failed as before); an error once the agent streams or settles (a database error writing the logs, a publish, a parser) gives the lease up without signalling the agent and leaves the run for the reconciler to re-attach (onStreamError). Attach jobs that fail one attempt are counted in the run row's existing jsonb (attachFailures; cleared when the attempt starts and when an attach job makes progress; the starting job's failure does not count): the third in a row fails the run with the last error, the kind's way, and its agent is stopped through the rule. User cancel and the reset path are unchanged. Tests: unit (the decision, the attach job's two outcomes), integration (the counter per hold, preserving the column's other keys; detach, then fail on the third), e2e 18 (a log write fails once: re-attached, same agent, contiguous logs, cost once) and 19 (fails every time: failed after 3 attaches, agent stopped). The fake runtime takes [[mock:mark:N:TEXT]]; OPTIO_TEST_STREAM_FAULT[_FIRST] are the e2e tier's fault knobs.
…s never reach GitHub CI flaked in case 3 (ten runs, three servers): every cluster in the file works on one database but had a fake-runtime directory of its own, so the pod rows an earlier cluster left ready (scale-a, scale-b, its Jobs' and agents' pods) pointed at containers that no longer existed. Whether a worker's pod pick or the 5 s health check met them first was timing: on CI the health check sometimes won, cleaned the scale-b pod up as crashed (the stubbed kubectl's network-policy error in the log) and marked the repo's running tasks with that pod's error. The clusters now share one directory, as a deployment's pods outlive its API servers. The hermetic env (every startApiServer) also preloads a fetch guard: a request to github.com, githubusercontent.com or gitlab.com is answered in the process with the 401 a dummy token gets, and noted on stderr, so the dummy GITHUB_TOKEN the tiers seed never leaves the machine (kubectl was already a failing stub and the Kubernetes API a read-only fake). A unit test and the hermetic-server e2e prove it. The scale-out file's create assertions now print the response body, and a failure prints 80 log lines per server.
…died When the server case 3 kills was still setting an agent's turn up, the reconciler halts that turn (it never ran: no cost, no logs) and the agent's next turn does the work; the case read the halted turn and failed, as it once did for a task requeued the same way. It now waits for the turn that ran to its end and checks that any turn before it is exactly that recovery. Reproduced locally (1 in 4 runs) with the turns printed.
A server that dies between a pod's row and the pod leaves the row provisioning forever. pickPod took such a row over after 10 minutes, but a persistent agent's pod acquisition waited on it every turn (120 s, then failing the turn) and never recovered: the agent ended failed. CI's case 3 hit it when the server it kills was creating a new agent's pod. Both now share one rule (provisioningAbandoned: a bare pod's row still provisioning past OPTIO_STALE_PROVISIONING_MS, 10 min; a live creator marks it within the runtime's 120 s), and a waiter waits only until the row would count as abandoned, then removes it and makes the pod (waitForProvisioning). Tests: unit (the agent pool's three outcomes), integration (pickPod takes over a row the moment it goes stale; waitForProvisioning), e2e 20 (an idle agent's pod row set back to provisioning: its next turn runs within seconds; fails without the fix with CI's symptom). The scale-out e2e clusters run with a 15 s stale age.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Scale-out phase B (
docs/plans/scale-out.md§1–2): agent runs outlive the API process that started them. Every run kind starts under the run protocol from phase A and is streamed through one attach loop (services/run-attachment.ts), and any instance can pick a run up where its rows stop: no duplicate line, no gap, cost added once.Merged with
origin/main(C1 #668, C2 #667): C1's claim transactions and C2's chat-turn leases are kept as they are; the attach sweep is listed inREPEAT_SCHEDULERS.The shared machinery
attached_by/attach_lease_until(60 s, renewed every 20 s withWHERE attached_by = me). A renewal that matches nothing severs the stream at once and the worker returns. On SIGTERM,detachAllsevers every stream and sets every lease tonow()before the workers close.exec_state='started',exec_pid,consumed_bytes=0, the lease. It is a CAS on the holder, so a run that was queued again meanwhile gets its new agent stopped instead.streamRun). Attaches from byte 0. Lines that end at or beforeconsumed_bytesonly feed the in-memory derivations: the PR tool-call tracker, the session id, the terminal event (which sends the stdin EOF, on a replay too), the result/auth digest and a command's exit receipt. Every other line becomes a log row. Each chunk's rows and the new offset are committed in one transaction (appendRunLogsBatch, a CAS on the holder and the old offset), at line ends, counted in bytes.allLogs. Each agent adapter now reads its result line by line (createResultParser;parseResultfeeds a whole output through it, so the existing tests still apply).inferExitCodeand auth-failure detection also work per line (run-output-digest.ts).settleRunwritesexec_state='exited'and the run's spend in one statement, only while this instance holds the run, and only once. If an instance dies after settling but before the run's own transition, the next attach finishes the run without adding the spend again.runAttachJobis one driver for every kind: claim by CAS (exit when the claim is lost), then either the kind's pod-death path when the pod is gone or the kind's finish.killRun, for tasks, Job runs (fail hook), reviews (cancel) and agents (archive, or failed by the reconciler). A mid-run message to a task is appended to the agent's stdin by the route. The Redis task-message bus is gone.Per worker
task-worker.ts(Repo Tasks)startTaskInRepoPod. The task staysprovisioninguntil the start is recorded, sorunningalways means an agent some instance can attach to; a cancel during setup stops the new agent. Finishes withstreamRunand the digest, thensettleRun, then the existing PR detection and transitions. Handlesattach-taskjobs, which move a still-provisioningtask torunningwhen the instance that started it died first. A failure in the start exec now fails the run instead of taking the provisioning retry.workflow-worker.ts(Job runs)attach-workflow-run.persistent-agent-worker.ts(turns)runningonly after the start is recorded (beginTurn), so a turn that never started leaves its messages pending. The turn's spend is added to the agent's total inside the settle transaction. A turn finishing on a paused or archived agent leaves its state alone; the old code moved it back toidle. Handlesattach-persistent-agent-turn.pr-review-worker.ts(review runs)runningonly after its start is recorded. A cancelled review is never flipped tofailedby its stopped run. Handlesattach-pr-review-run.Reconciler
execution:exec_state,attached_by,attach_lease_until. That is the row's own for tasks and Job runs, the open turn's for agents, and the active run's for reviews.reattach: the run isrunning/provisioning, its agent started (or its exit is half recorded), and the lease has expired. This enqueuesattach__<run>__<expired lease ms>. The executor checks the version under CAS without writing anything; the job is removed when done.requeueRun: the run never started, its lease expired, and the protocol had held it. Each kind's own retry handles it: tasks and review runs go back toqueued; a Job run goes throughfailedtoqueuedwithout spending a retry; an agent's turn is halted and the agent goes back toidle.OPTIO_ATTACH_LEASE_MS,OPTIO_ATTACH_RENEW_MSandOPTIO_ATTACH_SWEEP_INTERVAL_MScan override the lease, renewal and sweep intervals; the e2e suite sets 6 s / 2 s / 1 s.E2E matrix (
apps/api/e2e/scale-out-runs.e2e.test.ts,startApiCluster)Every case checks the final state, that the log rows are exactly the agent's
tick 1…Nlines (each once, with no gap), and that the cost was added once.pr_opened, cost oncedestroyof the pod mid-runfailedwith the pod's reasonRuns: 3 green on the branch before the merge, 1 inside the full e2e tier and 3 more green on the merged tree (about 100 s each). The first-ever run found three bugs, all fixed: a reconciler agent archive or failure did not stop the turn's agent, a lost task's error message was not recorded, and the test was too strict about a legitimate requeue.
Tiers (on the merged tree)
pnpm turbo test: 13/13 packages green; API 189 filespnpm --filter @optio/api test:integration: 55 files, 455 tests green. This includes the newrun-attachment.int.test.ts: two instances, as two module graphs, race the claim 25 times and exactly one wins each time. It also covers the live lease, the loser's log batch rolling back, spend settled once, and thequeuedreset.pnpm --filter @optio/api test:e2e: 26 files, 120 tests green. The scale-out file was also run 3 more times, all green.pnpm format:checkandpnpm turbo typecheck: green. Shared types underpackages/shared/src/typesare unchanged, so no Swift or Kotlin regeneration was needed.run-protocol.linux.test.ts, which is skipped on macOS, was run in anode:22-bookwormcontainer: 6/6 green, including a new test for a fresh start stopping a stale supervisor. Its 200 KB-line case depends on the locale (${#l}counts characters), so it passes underC.UTF-8and fails under POSIX, on main as well.Plan corrections
runningmeans the agent started. A task or review run staysprovisioning, and an agent staysprovisioningfor its turn, until the start is recorded. Otherwise arunningrun with no agent would be ambiguous.queuedclears the attachment columns, so a new attempt never inherits the last attempt's expired lease orexitedstate.attach__<run>__<expired lease ms>, notattach:<run>:<attempt>. BullMQ rejects:in custom ids except in a legacy three-part form it plans to drop, andrunKeyalready uses__.reattachunder a version CAS that writes nothing, so a worker finishing the run at the same moment never loses its own CAS.Left as is
Review fixes
Fifteen findings from the code review, each with a test. Commits:
b9ceb58c(protocol scripts),8618bd87(digest),d61d4995(late messages),cb7f34c8+2a38646e+1337008e(attach loop and workers),e5ea0a18(e2e),a75b6839(docs).A design change under several fixes:
attached_bynames the hold, not just the instance (<instance>/<nonce>). Every column write (renew, start, log batch, settle, close) is a CAS on the hold's own mark. Without it, a force-redo picked up by the same instance let the old attempt's job renew, append to and settle the new attempt's row, since both jobs matchedattached_by = me.resetRunAttempt(run-log-service), for every path back toqueued: a task's transition, its force-redo (reset, log delete and event now in one transaction), a Job run's retry and requeue, a review run's requeue. It locks and reads the row, clears the attachment columns in the same transaction, and after commit TERMs the last attempt's agent if it was started, scoped to its pid (stopPreviousAttempt).endRunJob, called first in every worker'sfinallyand inrunAttachJob. If the job started or claimed an agent and ends without finishing the run or leaving it for another instance (detached, shutdown), the agent gets TERM (pid-scoped) unless its exit is recorded or another hold streams it under a live lease. Then the lease and the slot are given back. The per-kind signals after a failedrecordRunStartedorrunningtransition are gone. User cancel paths are unchanged.agentLeftBehinddecisions, signal before release, reset/finished/detached; int: real rowsattached_by = meclause is gone).runAttachJobchecks the in-process registry (holdsRun) first.provisioning/running, or a turn not finished). Anexitedrun that is still active can be claimed; it finishes asalready, with no spend added.endRunJobstops the agent. A review cancel also moves its queued and provisioning runs to cancelled. A run completes, and its results are written, only while the run and its review are still active (transitionRunFromActive).agent_drafted); fails without the fixdetachAllreleased settling holdssetup→streaming→settling(once the stream ended in an exit or loss). SIGTERM letssetup/streaminggo at once and waits up to 20 s (OPTIO_DETACH_SETTLE_WAIT_MS) forsettlingones, which keep renewing. A run let go while its job was still starting it gives the fresh lease up again.<runDir>.prev-<epoch>-<pid>and removed in the background; files are never truncated in place. The agent script and supervisor are written only after that. The fake runtime mirrors this and follows an attachment's output by descriptor.stdin.ndjsonholds the sentinel: it prints__OPTIO_STDIN_CLOSED__→RunStdinClosedError. The delivery marks the message "The agent already finished its turn; the message was not delivered". A replayed EOF is not an error.agent_type(null means command). A turn recordsagentRuntimein itswake_payload(no migration), and the re-attach reads it, falling back to the agent's runtime.taken(same agent, another hold) orreset(exec state cleared, or another pid). Onresetthe worker gives its slot back, andendRunJobstops the old agent.removeRunDirectoryreads the last 8 KB ofstderr.logand removes the directory in one exec, logging the tail as a warning. A failed, lost or no-output run reads the tail before its transition (readRunStderr) and appends it to the error (withStderr).[[mock:stderr:…]]reaches the task's error)removeRunDirectoryguardisRemovableRunDir: exactly/home/agent/optio/runs/<segment>, segment^[A-Za-z0-9_-][A-Za-z0-9_.-]*$, not./...Also:
optio_run_pidnever returns pid 0/1, becausekill -- -1would signal the whole pod.Disagreements and judgment calls
Tiers after the fixes
pnpm turbo test: 13/13 packages green. API: 190 files, 3096 tests. container-runtime: 179 tests, plus 10 skipped (the Linux file).pnpm --filter @optio/api test:integration: 55 files, 462 tests green.run-attachment.int.test.tshas 13 tests, 7 of them new.pnpm --filter @optio/api test:e2e: 26 files, 125 tests green.scale-out-runs.e2e.test.ts(now 13 cases, a third cluster for 13–17): 4 green runs on the final tree (one inside the full tier, then 3 more in a row), plus earlier runs while building. Cases 14, 15 and 16 were also checked against the code with their fix reverted: each fails. Case 16 needed a holder sampler, because the conditional review-run transitions alone hide the second finisher.run-protocol.linux.test.tsinnode:22-bookworm(C.UTF-8): 10/10 green, 4 of them new.pnpm format:check,pnpm turbo typecheck: green. No shared types changed, so no Swift or Kotlin regeneration was needed.Review fixes, round 2 (
c77f6119,032981e6,e0dad4c4,7683cb13,a0a082ef,bc3e6003)Streaming errors detach instead of killing. Round 1 applied item 2's rule to every error. Now
endRunJoband the workers split errors by where the job failed:onStreamErrorgives the lease up (attach_lease_until = now()). It does not signal the agent or fail the run. The reconciler re-attaches.attachFailuresin the row's existing jsonb (no migration): a task's or Job run'smetadata, a review run'smetadata, a turn'swake_payload. The counter is a CAS on the hold, and the column's other keys are kept. It is cleared when the attempt's start is recorded and when an attach job commits progress; the starting job's own failure does not count. The third failed attach in a row fails the run the kind's way (onAttachFailed: taskattach_failed, Job runrecoveryRequiredso it is not auto-retried, turn halted and agent failed, review run and review failed) with "Attaching to the run failed 3 times in a row; last error: …".endRunJobthen stops the agent.OPTIO_TEST_STREAM_FAULT_FIRSTon a[[mock:mark:10:FAULT-ONCE]]line). The run is re-attached and reachespr_openedwith contiguous logs, the line stored once, cost once, the same agent pid and exit 0. Checked to fail under the old kill rule.attachFailures= 3, and the agent (still ticking) exits 143.Item 11 for agents without a result event. Codex, Copilot, OpenCode, Gemini and OpenClaw are back to exactly the pre-PR rule: their failure signals decide, and the exit code is not consulted. The Claude (and Cursor) result-event rule stays. A unit test pins the old Codex behaviour across exit codes 0, 1, 3 and 143 against
inferExitCode.CI flake in case 3 (
e0dad4c4). CI failed case 3 atd2058147(a turn haltederror) and at1337008e(a create answered 500). Both logs show the same thing: the repo-cleanup health check cleaned thescale-brepo pod up ascrashed. Cause: every cluster in the file shares one database but had its own fake-runtime directory, so thereadypod rows an earlier cluster left behind pointed at containers that no longer existed. Locally a worker's pod pick usually met them first and dropped them quietly. On CI the 5 s health check sometimes won: it ran the crash path (the stubbedkubectl delete networkpolicyerror in the log) and marked the repo's running tasks with that pod's error. The file's clusters now share one fake-runtime directory, as a deployment's pods outlive its API servers. The scale-out file's create assertions now print the response body, and a failure prints 80 log lines per server, so a recurrence is diagnosable.Hermeticity. Every e2e server was already off any real cluster:
kubectlandhelmare failing stubs,KUBECONFIGpoints at a read-only fake API, and the developer's credentials are stripped. GitHub was the gap: the dummyGITHUB_TOKENthe tiers seed was sent to the real api.github.com (PR detection, the reconciler's PR reads). The hermetic env now preloads afetchguard (NODE_OPTIONS --import). A request to github.com, githubusercontent.com or gitlab.com is answered in-process with the same 401 and noted on stderr. A unit test covers it, and so does a hermetic-server e2e case: a seeded token's GitHub call is answered locally.Case 3's agent assertion. Running the file again surfaced the
error-turn symptom locally in 1 of 4 runs, with the turns printed: the server case 3 kills was still setting an agent's turn up, so the reconciler halted that never-started turn ("went away before its agent started; queued again", no cost, no logs) and the next turn ran tonatural. That is recovery working. Case 3 now waits for the turn that ran to its end and checks that any turn before it is exactly that requeue.A pod row abandoned mid-creation (
bc3e6003). CI then failed case 3 once more: an agent never finished a turn in 120 s. The agent pool waited on a pod row leftprovisioningby a server that died between inserting the row and creating the pod. Every turn waited 120 s and failed, so the agent endedfailed.pickPod(repo and Job pods) already took such rows over after 10 minutes. Both now share one rule,provisioningAbandoned: a bare pod's row stillprovisioningpastOPTIO_STALE_PROVISIONING_MS(10 min; a live creator marks it ready or failed within the runtime's 120 s). A waiter (waitForProvisioning) waits only until the row would count as abandoned, then removes it and makes the pod. Tests:pickPodtakes over a row the moment it goes stale;waitForProvisioningprovisioning, and its next turn runs within seconds. With the old code it fails with CI's symptom.The scale-out clusters run with a 15 s stale age.
Tiers on the final tree (
bc3e6003):pnpm turbo test: 13/13 packages (API 190 files, 3105 tests).scale-out-runs.e2e.test.ts(16 cases): 3 green runs in a row after the tier, and 6 more of the three-server block while hunting the flake.bc3e6003(run 38106457189), E2E included.