Repository navigation
[Discussion] Condition-driven waiting and wake-up for Sessions #5920
Replies: 1 comment
Research appendix: event sources, transport trade-offs, and validationThis appendix preserves candidate use cases, reuse points, external precedents, and failure analysis that would make the main discussion too long. It is not a commitment to implement every source or a claim that these cases have already been tested in Maka. Prepared on 2026-10-01. Code observations come from targeted investigation during the discussion, with key foundations rechecked at checkout A. Later sources: prioritize real workflows, not adapter countThe main proposal prioritizes ShellRun, CI, PR/review, and file readiness. The following are deferred candidates, not claims of low usage. Maintainers with frequent workflows may reasonably move them ahead of a currently planned adapter.
Two rules apply throughout:
Real notification examples without a public inbound endpoint
The first two are local or locally observable events; the third demonstrates remote push without inbound reachability. All still require current-state verification and missed-event reconciliation. B. Why GitHub Webhooks are not an initial prerequisiteA Webhook is an event transport, not a mandatory product or tool called “webhooker,” and not a CI-only mechanism. PRs, reviews, and deployments can use it, but their changes can also be discovered through APIs. B.1 Reachability and operationsGitHub must deliver to a reachable HTTPS endpoint. A user's machine may sit behind NAT or a corporate firewall, lack a stable domain, and disconnect, sleep, or shut down.
A relay can retain events while the client is offline. It cannot by itself wake a powered-off computer; local model execution still requires Host to be available. B.2 Read access does not imply permission to install a WebhookA typical open-source contributor can push to a personal fork and read upstream public PR/CI state, but cannot configure upstream Webhooks or install a GitHub App there. Webhook-only support would therefore exclude a common contribution workflow. GitHub Apps improve integration management but still depend on repository/organization approval. Polling is likely a lasting fallback, not just a temporary shortcut. B.3 Event delivery is not condition satisfactionA workflow completion event may refer to an old SHA or attempt, an optional check, or only one required check. Notifications may also be duplicated, reordered, or missed while offline. Use: Signature validation, replay protection, routing, load/payload limits, and credential-revocation handling are still required. Payload text must not become model instructions. B.4 Suggested evolutionStart with Host polling to validate waiting and continuation. Later add “Webhook-triggered verification plus low-frequency reconciliation” behind the same adapter. Tune intervals against quotas, task phase, and acceptable latency rather than polling every subscription aggressively. Describe this honestly as polling-based detection with event-driven Session continuation, not as converting the provider API into push. C. Existing code to reuse—and limits to that reuseThese are entry-point observations, not claims that a generic waiting tool already exists.
Code entry points:
These links track a moving main branch. Recheck the implementation revision before coding. In particular, historical comments that continuation is entirely volatile or ScheduledTask always rewrites its entire table should not be reused as current facts. D. External precedents: useful patterns, not wholesale replacementsD.1 Google ADK: LongRunningFunctionToolThe official documentation describes a model invoking a tool that starts a long operation, returns an initial result such as an operation ID, and pauses/ends the agent run. The application can later supply intermediate or final responses to continue. Useful lessons:
This does not supply Maka's storage, Goal accounting, permissions, or restart policy automatically. Whether to send progress back to the model remains an application decision; unchanged observations should not trigger repeated model work. D.2 LangGraph: interrupts, checkpoints, and resumeThe official interrupts documentation supports pausing nodes/tools, persisting through a checkpointer, and resuming the same thread with Important lessons:
Maka need not adopt node replay. Existing execution ledgers and admission should instead prevent “resume the wait” from becoming “rerun arbitrary commands preceding registration.” D.3 Temporal: Signals, conditions, and durable schedulingThe official message-passing documentation shows workflows receiving Signals, updating state, waiting on conditions, and continuing. The useful precedents are the separation of intake, state evaluation, and execution, plus explicit identity, timeout, and recovery rules. Temporal is not a model tool; this proposal does not recommend importing or rebuilding a workflow platform of that scale. Common lesson: the tool is an entry point; the runtime and persistence protocol provide suspension and continuation. This research does not rank framework performance/correctness or infer support for arbitrary external predicates from their examples. E. Alternatives rejected or deferred
“New conditions over general-purpose data” does not mean arbitrary scripts. Versioned templates with bounded parameters can express new instances, such as a declared file set, required fields, and a completion marker. Do not silently replace an unsupported requirement with an approximate signal. F. Multiple conditions and event storms“One effective wait initially” must not mean “only one observable object”:
Semantics to specify:
One-shot subscriptions simplify the effective-successor bound. They do not limit queries while a condition never becomes true; TTL and observation budgets remain necessary. G. Fault injection and measurementG.1 Suggested test matrix
G.2 Establish a baseline before claiming numbersCompare four paths: model-scheduled checks, a single polling script, existing Goal waiting, and the proposed mechanism. Use the same resource with controlled event timing and record:
“Checking every 30 seconds for 20 minutes is about 40 queries” is an accounting illustration, not a benchmark. Both a script and a Host watcher can avoid per-query model calls; their main difference remains execution lifetime, recovery, and control. 中文说明(点击展开)调研附录:事件来源、传输取舍与验证计划本附录保留正文之外的候选场景、代码复用点、业界参考和故障分析,供讨论者调整优先级。它不是对所有来源的实现承诺,也不表示这些场景已在 Maka 中做过实验。 整理日期:2026-10-01。代码依据为本次讨论中的路径核查,关键基础在 checkout A. 中长期来源:按真实价值调整,而不是按名称扩张短期正文优先 ShellRun、CI、PR/review、文件就绪。下列场景暂列后续,并非断言其使用频率低;如果维护者有高频工作流,它们完全可能优先于某个短期适配器。
两点适用于所有来源:
不依赖公网入口的真实通知示范
前两项主要是本地或本机可见事件,第三项证明远端推送也不一定要求本机入站网络。通知仍需要当前状态核验与漏事件校准。 B. GitHub Webhook 为什么不是首期必要条件这里的 Webhook 是事件传输方式,不是必须采购或安装的名为 “webhooker” 的工具,也不是 CI 专属功能。PR、review、部署等都可采用它,但也能通过 API 查询发现变化。 B.1 可达性与运维GitHub 要向一个可访问的 HTTPS 地址发送事件,而用户机器可能在 NAT、企业防火墙之后,没有固定域名,并会断网、休眠或关闭。
中转可以保存离线事件,不能自动把关机的电脑唤醒。用户的 Host 不在线时,模型仍不能在本机执行。 B.2 权限不对称:能看 CI,不代表能装 Webhook普通开源贡献者可以向个人 fork push,并读取上游公共 PR/CI;但通常不能为上游仓库配置 Webhook 或安装 GitHub App。 因此,仅支持 Webhook 会排除一部分最常见的贡献流程。GitHub App 可以改善权限和接入管理,但仍依赖组织/仓库许可。轮询 fallback 很可能需要长期保留。 B.3 事件到达不等于条件已满足一次 workflow 完成通知可能对应旧 SHA、旧 attempt、optional check,或只是 required checks 中的一部分。通知也可能重复、乱序或在离线期间丢失。 建议采用: 还需要签名验证、重放防护、正确路由、负载与 payload 限制,以及凭据撤销后的处理。不能让通知携带的文本直接成为模型指令。 B.4 推荐演进首期采用 Host 轮询,先证明等待与执行恢复协议。之后在相同适配器下加入“Webhook 及时触发核验+低频校准兜底”。具体间隔应依据配额、任务阶段和可接受延迟测量,而非固定地高频查询所有订阅。 这应表述为“轮询检测+事件驱动的 Session 恢复”,不能宣称外部 API 已变成推送。 C. 当前代码可以复用什么,不能假定什么下列是入口级核查,不是声称已有通用等待工具。
代码入口:
这些链接指向会变化的 main。实施前应按当时 revision 重新核验。尤其不能照搬旧评论中“continuation 完全不持久化”或“ScheduledTask 当前总是全表重写”等历史判断。 D. 业界机制:相似点与不能照搬的部分D.1 Google ADK:LongRunningFunctionTool官方文档描述了模型调用工具启动长任务、返回 operation ID 等初始结果、agent run 暂停/结束,再由应用提交中间或最终结果继续的方式。 可借鉴:
不能据此推导出框架已经替 Maka 解决存储、Goal 预算、权限和重启恢复。是否、何时提交进度给模型,仍是应用决策;不应每次无变化都回传。 D.2 LangGraph:interrupt、checkpoint 与 resume官方 interrupts 文档支持在节点和工具内暂停,以持久化 checkpointer 和相同 thread ID 恢复,并使用 关键经验:
Maka 不必复制节点重放模型;反而应利用现有执行日志与接纳语义,避免把“恢复等待”变成重跑登记前的任意命令。 D.3 Temporal:Signal、条件等待与持久化调度官方 message-passing 文档展示了向工作流发送 Signal、更新状态、等待条件及继续执行。 值得学习的是接收、状态判断、执行三个层次,以及消息身份、超时和恢复规则。Temporal 不是模型工具;不建议为当前需求直接引入或重建同等规模的通用工作流平台。 共同结论:模型工具只是入口,暂停/恢复能力主要由运行时和持久化协议提供。 本调研没有建立跨框架性能或正确性排名,也没有把官方例子等同于它们支持任意外部条件。 E. 拒绝或推迟的方案,以及原因
“通用数据上的新条件”不等于上述任意脚本:可通过有版本的固定模板和有界参数支持新组合,例如声明文件清单、所需字段和明确完成标记。不能表达的条件不应被近似信号悄悄替换。 F. 多条件与事件风暴的设计检查不把“首版一个有效等待”误解为只能观察一个对象:
需要固定的语义:
初期 one-shot 可把“一个等待的有效后继”限制得更清楚,但它不限制条件一直未满足时的查询次数,所以仍需 TTL 与观测预算。 G. 故障注入与测量计划G.1 建议测试矩阵
G.2 先建立基线,不预设数字比较四条路径:模型逐次查询、单工具脚本轮询、既有 Goal waiting、新机制。使用相同对象与可控的事件时间,记录:
“每 30 秒检查 20 分钟约 40 次查询”只能说明计数差异,不是性能测量。脚本循环和 Host watcher 都可能不逐次调用模型;两者的主要区别仍是执行生命周期、恢复和治理。 Preparation note: drafted with assistance from Maka. Framework observations are based on the linked official documentation; repository findings are targeted code inspection, not production validation or benchmark results. |
Uh oh!
There was an error while loading. Please reload this page.
I would like to discuss a concrete implementation path for condition-driven waiting and wake-up: replace model-driven status checks and long-held tool calls with structured waits owned by Runtime Host, then continue the original work through existing execution admission when an actionable result becomes available.
This is not a new workflow engine or a replacement for Goal, Session, Turn, or recovery authority. The near-term target is a common mechanism with ShellRun, GitHub CI, PR/review state, and external-file readiness adapters. The interfaces and PR boundaries below are proposals, not implemented APIs or an approved design.
1. Motivation: distinguish three ways of waiting
A common development task is to change code, push, and wait for CI. In one such execution, the agent first repeatedly scheduled
sleep → query GitHub, then switched to a polling loop inside a single script. Both approaches kept the same execution Turn open.Maka also has a separate path: after a Turn settles, the Goal evaluator can return
waiting, and an exponential timer schedules another execution to check again. The current interval starts at 5 seconds and is capped at 5 minutes; the timer itself does not verify CI.By Host observation with condition-driven continuation, I mean handing the wait to background code that does not depend on model reasoning: the agent declares a resource and a condition, Host saves that declaration, and the current execution can end. Host then receives notifications or reads source state periodically. Only a matching result, timeout, or actionable exception requests a follow-up under existing execution rules.
For example, when waiting for CI on a specific commit, background code distinguishes “still running” from “finished”; the model need not return periodically to ask. When CI finishes, Host delivers the commit identity, result, and log references to the original Session so the model can analyze them or verify completion. This continues the original task rather than keeping a tool call open throughout the wait.
The benefits are fewer unnecessary model calls, releasing an execution slot held only for waiting, and consistent cancellation, deadlines, recovery, and result delivery.
More concretely:
Compared with repeated model-scheduled checks, the main gain is less unnecessary reasoning and context processing. Compared with a single polling script, it is primarily releasing execution occupancy, durable waiting, and reliable continuation.
There is also a separate accounting issue: check Turns settled through the normal waiting path can consume Goal work iterations. It is the Turn—not each HTTP request—that counts. Work iterations and wait lifetime should be separate, but removing the former bound must coincide with a TTL, observation limits, and a wake bound. We must not introduce an unbounded intermediate behavior.
2. Proposed architecture and model interaction
The core needs three responsibilities, not three independent execution systems:
Model-facing entry point
Provide a high-level tool, provisionally
WaitForEvent. The model should not have to compose “create watcher, save continuation, suspend Session” correctly.Illustrative request:
{ "source": "shell_run", "resource": {"taskRef": "maka://runtime/background-tasks/task-123"}, "condition": {"type": "terminal", "version": 1}, "timeoutSeconds": 3600 }The model proposes a resource and condition. Host binds Session, Turn, Goal identity, and permissions from trusted call context. If the condition already holds, the result can be delivered immediately; otherwise Host registers the wait and yields the current execution. Yielding must be a runtime control operation, not merely text asking the model to stop.
The initial registration acknowledgement is separate from the eventual result. On delivery, the model receives the original objective, wait reason, resource version, and bounded evidence references. Timeout, failure, cancellation, or invalidation may also require handling. A wake is not proof that the overall goal succeeded.
Host-side capabilities to add and connect
This section describes changes in Runtime Host and the state it exposes to clients. Adapters are its integration layer for individual sources, not another execution backend.
pendingContinuationand execution identity rather than maintaining a competing running/pending authority.3. Correctness and lifecycle requirements
“Any success / all success / any failure” is not a complete generic condition language. PR closure can satisfy a condition without representing business success; collecting test results may require all tasks to reach a terminal state. Start with a few deterministic templates: one task becoming terminal, all conditions in a declared set being satisfied, all tasks settling, and CI-specific fail-fast aggregation. Do not expose arbitrary expressions.
4. Scope and restart boundaries
Existing mechanisms to cover or preserve
whenIdleA ShellRun is a command managed by Maka's process manager with a task record. Built-in foreground and background Bash use this path, but helper processes inside other tools, remote jobs, and external user terminals are not automatically equivalent wait targets. Background tasks are the initial focus; not every command completion should invoke the model.
Condition categories in the near-term scope
The model may propose parameters for a data condition; Host validates them and runs deterministic checks. This does not permit arbitrary model-generated observer code. “No directory changes for five minutes” proves a stability interval, not that every expected upload finished. Reject or clarify conditions that cannot be expressed or verified.
Preserve current product behavior after interruption
This change does not redefine when continuation or rerunning requires approval. It restores wait facts, checks source state, and hands valid results to existing Session/Goal control and recovery:
orphaned, unreachable, or missing a result.Recovering a wait is not recovering a process or granting retry authority. Automatic replay of side-effecting operations is not an implicit promise of this optimization.
5. Security and authorization
Validate four separate capabilities: registering a wait, observing a source, creating a follow-up, and performing concrete actions.
userAuthorized: true. Host supplies trusted identity and policy context. Reuse the existing confirmation mechanism where confirmation is required.analyze-only: verification, analysis, and reporting only. If adopted, enforce this boundary through actual capabilities, not prompting alone.This does not alter the general permission mode of existing Goals. It prevents a new wake entry point from bypassing it.
6. Architecture acceptance criteria
The first user-visible slice should satisfy at least:
Compare the same real scenario before and after implementation: model calls, tokens, API requests, detection-to-admission latency, missed/duplicate deliveries, and user intervention. The research does not establish numerical performance gains; baseline measurements and implementation tests must do that.
7. Near-term plan: three core PRs and four adapter PRs
Split the core around three independently testable questions: how to retain a wait, how to decide it has a result, and how to return that result to the original task. The first three PRs establish the protocol with test sources. The fourth proves a user-visible ShellRun flow before adding external adapters.
PR 1: contracts and persistence—retain what we are waiting for
Define and persist the resource, condition version, task ownership, deadline, cancellation, and delivery identity. Handle duplicate registration, stale updates, Session cleanup, and schema migration so later components read one authoritative record rather than maintaining separate in-memory copies.
PR 2: Host observation and condition management—observe without model work
Introduce the adapter interface and wait manager for initial reads, notifications, background polling, TTL, limits, cancellation, and startup reconciliation. Adapters return explicit condition outcomes and the manager persists them; this layer does not start the model directly.
PR 3: execution bridge and Goal integration—admit one effective follow-up
Connect wait registration, yielding the current execution, blocking unconditioned Goal continuation, and result delivery. Reuse existing continuation and root-Turn admission, carry typed provenance, and check current Goal validity, permissions, and Session availability.
Integrate normal budget accounting and settlement, including replacement limits when idle waits stop consuming work iterations. Reconcile crash windows such as a saved result whose execution has not been admitted; do not create another background execution queue.
PR 4: ShellRun—ship the first usable long-task wait
Let the model register waits on background ShellRuns it is authorized to observe, consuming reliable terminal facts and result references. Completion, failure, cancellation, and
orphanedstate need explicit outcomes. If the task finished before registration, read the existing record rather than waiting for another notification.Expose the tool, minimal status display, and cancellation so users can actually “run tests and continue handling the result when finished.” Keep resource-cleanup callbacks in their existing role and do not turn every command completion into a model wake.
PR 5: GitHub CI—apply the same protocol to a remote condition
Add explicitly authorized GitHub connections and API observation bound to repository, head SHA, and run/attempt, with defined run-terminal or check-set conditions. Handle success, failure, cancellation, throttling, credential loss, and offline reconciliation without letting old commits trigger current work.
PR 6: PR/review—reuse GitHub integration for collaboration conditions
Build on existing authentication, observation, and identity handling to support merge, closure, review arrival, and explicit approval conditions. Registration binds the target PR and applicable revision, identity, or role requirements; the adapter verifies them.
PR 7: file readiness—deterministic conditions for externally delivered data
Combine filesystem notifications with current-state scans for completion markers, explicit manifests, and a small set of bounded checks—for example, declared files with matching checksums or an explicit CSV set containing required columns. The model supplies template parameters, not arbitrary observer scripts.
PRs 1–3 separate responsibilities, not authorities. PRs 3 and 4 may merge if that makes review easier; keep the series close enough to avoid long-lived unused abstractions. Reuse execution persistence and outbox ownership rather than adding an independent wake database or a second execution queue.
Deduplication, cancellation, deadlines, permissions, and restart safety ship with the first usable slice. They are not post-release hardening for a version that can lose wakes or wait indefinitely. Accounting changes may land at integration, but replacement bounds must become effective with them.
Goal is a reasonable first execution consumer. If ordinary Session follow-ups must also ship initially, define their ownership and admission tests explicitly rather than assuming coverage. Additional sources, transports, and implementation precedents are in the appendix.
中文说明(点击展开)
我想讨论一条条件驱动等待与唤醒的具体实施路线:将“等待外部条件”从模型反复查询或长时间挂起的工具调用,转为 Runtime Host 管理的结构化等待;有值得处理的新结果时,再经现有执行入口继续原任务。这不是另建一套工作流引擎,也不替换现有 Goal、Session、Turn 或恢复权威。短期希望完成共同机制,并接入 ShellRun、GitHub CI、PR/review 状态和外部文件就绪。以下接口和 PR 划分是讨论建议,尚未实现或形成维护共识。
1. 为什么改:三种等待方式的区别
一个实际开发流程是:修改代码、push,然后等待 CI。在一次这样的执行中,先多次由模型安排
sleep → 查询 GitHub,后来改为一个脚本内部轮询;两种方式都发生在同一个尚未结束的 Turn 内。Maka 也已有另一条路径:Turn 结束后,Goal evaluator 判断
waiting,退避定时器再安排一次执行检查条件。当前退避为 5 秒起、最高间隔 5 分钟;定时器本身并不核验 CI。这里提出的“Host 观测+条件驱动恢复”,是把等待交给不依赖模型推理的后台程序:Agent 登记“等待哪个对象、什么条件成立”,Host 保存这份记录,当前执行即可结束。之后由 Host 接收通知或定期读取状态,只有条件满足、超时或出现需要处理的异常时,才按现有规则安排后续执行。
例如等待指定提交的 CI,后台程序负责辨认“仍在运行”与“已结束”;模型不必每隔一段时间回来询问。CI 结束后,Host 将对应提交、结果和日志引用交给原 Session,由模型继续分析或核验任务是否完成。这里恢复的是原任务的后续执行,不是把一个长时间未返回的工具调用继续挂着。
主要收益是减少无效模型调用、释放纯等待占用的执行位置,以及统一等待的取消、期限、恢复和结果交付。
具体来说:
相对模型反复查询,收益主要体现在减少无效推理与上下文处理;相对单个脚本内部轮询,收益更集中在释放执行占用、持久化等待和可靠接续。
还有一个独立问题:当前按正常 waiting 路径结算的检查 Turn 仍可能消耗 Goal 工作轮次。这里计入轮次的是 Turn,不是每次 HTTP 请求。应将工作轮次与等待期限分开,但取消旧的轮次消耗时,必须同时提供 TTL、观测限制和唤醒次数上限,不能产生无界等待。
2. 建议的架构与模型交互
核心应有三个职责,而不是三个独立的执行系统:
模型入口
提供一个高层等待工具,例如暂称
WaitForEvent。不要求模型组合“创建 watcher、保存 continuation、暂停 Session”等底层操作。示意请求:
{ "source": "shell_run", "resource": {"taskRef": "maka://runtime/background-tasks/task-123"}, "condition": {"type": "terminal", "version": 1}, "timeoutSeconds": 3600 }模型提出等待对象与条件;Host 从可信调用上下文绑定 Session、Turn、Goal 身份和权限。条件已满足时可以直接交付结果;否则登记等待并让出当前执行。让出必须是运行时控制行为,不能仅靠工具文本要求模型“不要继续”。
“已登记”等初始回执与未来的最终结果分离。结果到达后,模型获得原目标、等待原因、资源版本和有界证据引用,继续在原权限范围内处理。超时、失败、取消或条件失效也可以是需要处理的结果;唤醒不等于目标成功。
Host 后台需要增加和衔接的能力
这一部分描述 Runtime Host 的后台改动,以及它向客户端提供的状态;适配器是其中面向具体事件来源的接入层,不是另一个执行后台。
pendingContinuation和执行身份;不能为了方便再维护另一套“正在运行/等待执行”的权威状态。3. 正确性与生命周期注意事项
多条件不宜只用“any 成功/all 成功/any 失败”概括所有来源。PR 关闭是条件满足,不一定是业务成功;汇总测试还需要“全部进入终态”。短期可提供少量确定模板,例如单任务终态、明确集合全部满足、全部终态,以及 CI 的“任一 required check 失败,或全部通过”,不开放任意表达式。
4. 修改范围与中断恢复边界
覆盖什么,不覆盖什么
whenIdleShellRun 指经 Maka 受管进程管理器执行、具有任务记录的命令。内置 Bash 的前台与后台都走这条路径,但其他工具内部的进程、远端 job、用户外部终端不自动成为同类等待对象。首期重点是后台任务;不应把所有命令完成都转换成模型唤醒。
短期支持的条件类别
通用数据条件可以由模型提出参数,再由 Host 校验并执行确定性检查;不是允许模型提交任意观察代码。“目录安静五分钟”只能证明稳定期,不能擅自等同于“所有文件已经上传”。无法表达或验证的条件应明确拒绝或请求澄清。
中断后沿用现有产品逻辑
本次不重新定义是否自动续跑或重跑的产品策略。 新机制负责恢复等待事实、核验来源状态,并交给现有 Session/Goal 控制与恢复路径:
orphaned、失联或未记录结果而自动重派旧命令。恢复等待不等于恢复进程,也不等于授予新的重试权限。自动重试有副作用的操作不属于这个优化的隐含承诺。
5. 安全与权限
需要分别核验:登记等待、读取来源、创建后续执行、执行具体动作。
userAuthorized: true等参数自证授权。身份与权限来自 Host 上下文;确需确认时复用现有可信交互机制。analyze-only(仅核验、分析和汇报)作为保守范围;若采用,应由实际能力限制执行,而非只靠提示词。这不改变已有 Goal 的一般权限模式,而是明确:外部唤醒入口不能成为绕过它的新通道。
6. 架构验收条件
首个用户可用版本至少通过以下验证:
按同一真实案例比较模型调用、token、API 请求量、检测到接纳的延迟、重复/漏恢复以及用户介入次数。当前调研没有证明具体性能收益数字;这些应由基线和实现后的测试给出。
7. 短期开发计划:三个核心 PR,四个来源 PR
核心部分按三个可分别验证的问题拆分:怎样保存等待、怎样判定等待有了结果、怎样将结果交回原任务。前三项先用测试来源建立协议,第四项用 ShellRun 形成第一个用户可用闭环,再增加外部适配器。
PR 1:契约与持久化——重启后仍知道在等什么
定义并保存等待对象、条件版本、所属任务、截止时间、取消状态及交付身份。处理重复登记、旧记录更新、Session 清理和 schema 迁移,使后续组件能够读取同一份权威事实,而不是各自保存一份内存状态。
PR 2:Host 观测与条件管理——没有新结果时由后台继续观察
建立来源适配器接口和等待管理器,处理首次状态读取、通知接收、后台轮询、TTL、限流、取消及启动后校准。适配器返回明确的条件结果,管理器将其持久化;这一层不直接启动模型。
PR 3:执行桥接与 Goal 整合——结果到来后只安排一次有效后续执行
把登记等待、让出当前执行、阻止无条件 Goal 续跑,以及结果交付连接起来。复用已有 continuation 和 root-Turn admission,携带明确来源身份,并检查 Goal 是否仍有效、权限是否仍允许、Session 是否繁忙。
同时接入正常预算和结算路径,明确等待不占用工作轮次时的替代限制。对“结果已保存但执行尚未接纳”等崩溃窗口做恢复核验,而不是另建一个后台执行队列。
PR 4:ShellRun——交付第一个可用的长任务等待流程
允许模型对自己有权访问的后台 ShellRun 登记等待,消费任务可靠终态和结果引用。正常完成、失败、取消或
orphaned都有明确结果;任务早于登记结束时,直接读取已有记录,不等待下一次通知。这一项同时开放等待工具、最小状态展示和取消入口,让用户能够实际体验“运行测试,结束后继续处理结果”。现有资源清理回调保持原职责,不将所有命令结束都自动变成唤醒。
PR 5:GitHub CI——用相同协议处理远端条件
增加明确授权的 GitHub 连接与 API 观测,绑定 repo、head SHA、run/attempt,支持约定的 run 终态或检查集合条件。处理成功、失败、取消、限流、凭据失效和离线后的补查,避免旧提交结果误触发当前任务。
PR 6:PR/review——复用 GitHub 接入,增加协作条件
在已有认证、轮询和资源身份基础上,增加 merged、closed、review 到达及明确审批条件。登记时绑定目标 PR 和适用的版本、身份或角色要求,再由适配器核验是否满足。
PR 7:文件就绪——支持外部程序交付的确定性数据条件
使用文件系统通知与当前状态扫描,支持完成标记、明确 manifest,以及少量有界校验模板,例如声明的文件齐全且校验通过,或指定 CSV 集合具备所需列。模型可以填写模板参数,但不能提交任意观察脚本。
PR 1–3 是依职责拆分,不是另建三套权威。PR 3 与 PR 4 可按规模合并;各项可紧邻评审,不应留下长期无人使用的抽象。现有执行存储和 outbox 应复用,避免独立 wake 数据库或第二条执行队列。
去重、取消、期限、权限和恢复安全必须随首个用户闭环交付。 不能先开放一个会漏唤醒或无限等待的版本,再把这些列为上线后增强。等待计量调整可以落在整合阶段,但与替代边界必须同时生效。
首个执行消费者建议从 Goal 开始;如果普通 Session follow-up 也必须首期开放,应明确其所有权和接纳测试,而不是自动假定已被覆盖。更多来源、传输方式和可复用经验列在附录。
Preparation note: this draft combines discussion, code-path inspection, and official framework documentation, with assistance from Maka. It is input for review, not a claim of implementation, completed benchmarking, or design approval.
All reactions