Repository navigation
feat: scale-out phase A — instance ids, the schema, leases, the run protocol, a multi-server e2e harness - #666
Merged
Merged
Conversation
… leases, tokens, claims
… no exec holds the agent
… a process of its own
… the workers still use
…ec, FIFO closed by the sentinel
…runs; prefer the active review
…inished run homes swept
…erver exposes kill(signal)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Phase A of
docs/plans/scale-out.md(in this PR): the foundation for API pods that share one database and for runs that outlive the pod that started them. Nothing user-facing changes yet; the workers still run agents through the classic exec path. B (re-attachable runs), C1 (coordination) and C2 (sessions / glance / skills) build on this and touch no schema.What landed
services/instance.ts:INSTANCE_ID=OPTIO_INSTANCE_ID(the chart sets it to the pod's name) + a suffix new at every boot, so a restarted pod is a new instance. Logged at boot;GET /api/healthreturnsinstanceId.1792010000_scale_out): the attachment columnsexec_state,exec_pid,consumed_bytes,attached_by,attach_lease_untilontasks(throughrunColumns(), sorepo_tasksandworkflow_runscarry them),pr_review_runsandpersistent_agent_turns; tablesleases,ws_upgrade_tokens,inbound_webhook_deliveries,glance_state,installed_skill_files,ticket_sync_claims; a unique partial indexpr_reviews_active_pr_url_key(older active duplicates cancelled first).migrate-safe.tsnow leaves out any later-added run column when it remakes the views between historical migrations (one list instead of therecovery_requiredspecial case). Integration test proves columns, views, indexes and constraint behavior.services/lease-service.ts:acquireLease(insert-or-update-where-expired, returning a row iff acquired),renewLease,releaseLease,leaseHolder,seizeLease(newest-wins),sweepLeases, andwithLease(key, fn, opts)that renews on an interval, abortsfn's signal and rejects withLeaseLostErrorwhen the lease is taken away. Integration-tested with two contending instances.ContainerRuntime(packages/container-runtime/src/run-protocol.ts, re-exported byutils/pod-env.ts):startRun,attachRun,deliverStdin,killRun. A start script ends withsuperviseAgent(runDir, agentLines), which writesagent.shandsupervisor.shinto the run home, primesstdin.ndjsonfrom$OPTIO_RUN_STDIN, launches the supervisor withsetsid … &, and prints__OPTIO_RUN_STARTED__:<pid>. The supervisor runs the agent with stdin fed from the file through a FIFO until__OPTIO_STDIN_EOF__, appends stdout tooutput.ndjson, and writesexit(143 on TERM). Attach istail -c +N --pid=<pid> -f, with the exit marker on stderr so stdout is pure file bytes. Kubernetes and Docker share one exec-based implementation; Docker's non-TTY exec now demuxes stderr. Verified for real onubuntu:24.04(start → attach → deliver → EOF → exit 7; attach from an offset; kill TERM → 143 with no leftover processes; SIGKILLed supervisor → lost; zero exit marks the env warm);run-protocol.linux.test.tsdoes the same under vitest on Linux (skips on macOS).buildTaskStartScript(repo-pool-service.ts; the task setup is now onetaskSetupLinesshared with the unchangedbuildTaskExecScript, whichexecTaskInRepoPodstill uses) andbuildPooledStartScript(pod-env.ts). No trap, no EPIPE watchdog, no run-home removal; the env-ready marker moves into the supervisor.FakeContainerRuntime({ dir })/OPTIO_FAKE_RUNTIME_DIRkeeps containers on disk;startRunspawnsfake-agent.mjsdetached, which plays every existing[[mock:*]]directive (plus[[mock:echo-stdin]]) intooutput.ndjsonwith aseqon every event, readsstdin.ndjson, waits for the EOF sentinel after its result like claude (exit 97 if none comes withinOPTIO_FAKE_EOF_TIMEOUT_MS), and writesexit. Kill-style utility execs anddestroyreach protocol runs. In-memory mode is unchanged for the existing tests.test-utils/e2e/api-cluster.ts:startApiCluster({ size, fakeRuntimeDir?, env? })→{ servers, fakeRuntimeDir, live(), client(i), any(), kill(i, signal?), add(), stop() }.api-cluster.e2e.test.ts: two servers boot together on one DB, a Job made through A is read through B and its run completes, A is SIGKILLed and B keeps serving and running new work, a third server joins.What B / C1 / C2 build on
startRun(handle, { runId, runDir, script, initialStdin }) → { pid, output };attachRun(handle, { runId, runDir, fromByte }) → { output: Readable, exit: Promise<RunExit>, close() }withRunExit = { kind: "exited", code } | { kind: "lost" } | { kind: "detached", reason };deliverStdin(handle, { runId, runDir, line });killRun(handle, { runId, runDir, signal: "TERM" | "KILL" | "INT" | "HUP" }) → boolean.runDirisrunHome(runId).podDir.execState,execPid,consumedBytes,attachedBy,attachLeaseUntilontasks/workRuns/workflowRuns/prReviewRuns/persistentAgentTurns;leases,wsUpgradeTokens,inboundWebhookDeliveries,glanceState,installedSkillFiles,ticketSyncClaims.STDIN_EOF_SENTINELis what a worker appends where it calledstdin.end().Corrections to the plan, found while building
ticket_external_idis written by five paths besides the sync, so a unique index on tasks would reject work a person starts on purpose: the sync getsticket_sync_claimsinstead.webhook_deliveriesalready exists (the outbound webhooks' log): the inbound dedupe table isinbound_webhook_deliveries.pod-env.tsre-exports them.RunExithas a third kind,detached, for an attach whose exec dropped without a marker.Tiers
pnpm turbo test: shared 748, container-runtime 162 (+3 Linux-only skipped here), api 3021 — green.pnpm --filter @optio/api test:integration: 50 files / 405 tests — green.pnpm --filter @optio/api test:e2e: 23 files / 102 tests — green (incl. the cluster smoke).pnpm format:check,pnpm turbo typecheck(13/13),helm lint— green.Review fixes (15 items)
trap on_term …and after launching the agent and the feeder; the start script waits for that file (≤10 s, 50 ms polls) before__OPTIO_RUN_STARTED__. fake-agent.mjs writes its pid/identity file only afterprocess.on("SIGTERM"), and the fake'sstartRunwaits for it. CI's red test passes (3× locally).<pid> <starttime>(field 22 of/proc/$$/stat, read past the(comm)withsed 's/.*) //'then field 20). One shell helper,optio_run_pid(RUN_ALIVE_HELPER), shared by attach, kill and the pooled start guard, trusts a pid only when the live process's start time matches; a mismatch is "not running" (attach → LOST after what was written, kill →none, start guard → proceed). Unit-tested as text, exercised in the Linux vitest and the Docker check with a stale pid file.tiniadded toimages/base.Dockerfile;scripts/repo-init.shand both pooled init scripts (workflow-pool-service.ts,persistent-agent-pool-service.ts, viaKEEP_POD_ALIVEin pod-env.ts) end withexec tini -- sleep infinity, falling back toexec sleep infinity. The Docker check runs under tini as pid 1 and asserts no<defunct>after three runs, a TERM kill and a SIGKILL. The plan notes the images must be rebuilt (the fake is unaffected).tail … | ( while …; done > "$d/stdin.pipe" ): only the loop holds the FIFO, sobreakcloses the agent's stdin at once. Docker check: the run ends 999 ms after the sentinel (bounded bytail --pid's 1 s poll).launchPrReview. The migration's CTE also cancels the queued/provisioning/runningpr_review_runsof the reviews it cancels.findReviewByUrl(new, used bylaunchPrReview) orders active-state rows first, then newestupdated_at; integration-tested with a cancelled duplicate both newer and older than the active one.exitafter stdout.RunAttachment.exitsettles only after the marker and the stdout end (writer'sfinish, reader'send, orclose()which abandons the output); documented on the type; unit test with stdout still flowing after the marker.startRunPrelude(bytes)(exec 3<&0 0</dev/null; export OPTIO_RUN_STDIN_BYTES=<n>) +head -c "$OPTIO_RUN_STDIN_BYTES" <&3 > stdin.ndjsonin the start script; the runtime writes exactlyBuffer.byteLengthbytes to the exec's stdin and closes it.deliverStdinishead -c <n> >> stdin.ndjsonthe same way.OPTIO_RUN_STDINandOPTIO_STDIN_LINEare gone. Docker check and the Linux vitest: a 300 KB first message and a 200 KB delivered line (multi-byte UTF-8 in the vitest) arrive byte-exact,stdin.ndjsoncompared withcmp.__OPTIO_STDIN_DELIVERED__after the append; stderr only goes into the error message on failure (the runtime package has no logger; the API's debug log is phase B's worker).normalizeFromByte:Number.isFinite ? Math.max(0, Math.floor) : 0.viewSourceColumns()extracts the quoted source columns from each view statement ("col" AS "alias"→col),withoutColumns()rebuilds the select list without those absent frominformation_schema;LATER_RUN_VIEW_COLUMNSis gone. Unit-tested against the realRUN_VIEWS_SQL; the fresh-install path runs in every integration/e2e boot. No manual step remains, so no CLAUDE.md note.buildTaskStartScriptprunes the-wtworktree and runsgit worktree pruneright after the worktree setup. The plan (§1 "Where things live") and both builders' docs state that the attached worker removes the run directory when the run is terminal andconsumed_bytesequals the file size (phase B). Backstop now:run-home-sweep-service.ts, called from the repo-cleanup worker's health cycle, lists/home/agent/optio/runsin every ready pod (with each output file's size) and removes directories of runs that are terminal and either non-protocol, drained, or terminal for over an hour; unknown directories are left alone.fake-tape.mjs(+fake-tape.d.mts) holds the markers, directive helpers,planCommandRunandplanAgentRun; both fake.ts's exec session and fake-agent.mjs play the plan.ApiServerHandle.kill(signal)(group kill, SIGKILL after 15 s for other signals);stop()iskill("SIGTERM")with 10 s; api-cluster uses it — onekillGroup./api/healthreturnsinstanceIdonly whenreq.useris set (schemaoptional); route test covers both; the e2e asserts the public probe carries nothing identifying.INSTANCE_STARTED_ATdropped;instance.tsimported statically in index.ts.Disagreed with nothing. Two notes: the fake runtime verifies a pid's identity through
/procon Linux and falls back to "pid alive" on macOS (no/proc), which is where the fake's own tests run; and item 7's prelude also points fd 0 at/dev/null, so no setup command can eat the stdin bytes beforeheaddoes.Verification after the fixes: Docker check on
ubuntu:24.04+ tini (ALL PASS: pid 1 is tini; start/attach/deliver/EOF exit 7; offset replay; TERM → 143 with no leftovers; SIGKILL → lost; stale pid → killnone/ attach LOST; 300 KB + 200 KB byte-exact; no zombies after every phase). Tiers: unit (shared 748, runtime 169 + 5 Linux-only skipped here, api 3027), integration 51 files / 409, e2e 23 files / 102, format, typecheck 13/13, helm lint — all green.