Repository navigation
fix(explore): preserve durable evidence links beyond the context budget - #5974
Conversation
Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>
Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>
loopx-agent
left a comment
There was a problem hiding this comment.
Reviewer: model_agent; gpt-6.1-sol; OpenAI; runtime_reported; reasoning_effort=xhigh
Approval conclusion (author-owned PR; GitHub blocks formal self-approval)
动机
持续记录 Explore 证据的 agent,以及在 Goal 成果页复核这些证据的用户,需要在第九次记录后仍保留此前任务关联。以前同一任务超过八条显式关联时,继续记录会失败;读取已存在的更长关联又会静默截断,后面的反驳可能变成看不见的证据。这个限制把紧凑上下文的展示成本施加到了持久数据上。本 PR 让有效关联持续保留,同时继续限制每次展示的大小。实测 File、SQLite 的第九次 CLI 写回、重放和页面关联均可读,省略部分可展开。本次关闭的是这条保存与找回链路,不宣称模型必然采用证据或长期研究结论变得更准确。
改动思路
复用既有 TypeScript Todo 工作要求 owner,保留显式关联的稳定顺序、去重和合法 id 校验,去掉八条持久容量上限。Python 读适配器保持完整合法关联,不新增决策 owner。Explore 完整审计先解析所有关联并分类风险;turn-context 只展示八条 id、三条明细和省略数量,展开命令指向同一份 Todo 审计。现有成果页复用原生读取 API,无需用户输入一遍已知关联,也没有增加授权步骤、能力开关或后台自动运行。
具体改动
关键代码讲解
normalizeTodoWorkRequirements移除关联数量拒绝,现有格式校验和稳定去重保留。normalize_explore_result_node_refs用 set 辅助去重,取消读回的提前 break;旧的九条上限拒绝测试相应改为十二条持续保留。branchContext在展示层截取 requested/unknown refs 并报告省略数量,不改变完整 audit 的 hazards、status counts、诊断模式或零 score delta。projectExploreTurnContext在紧凑化之前从完整关联挑选已链接结果,后面的反驳不会被前八条展示预算排除。harness.plan_command从 worker 分组方案改为todo-branch-plan --agent-id,展开当前可见任务的同一份证据审计。重复读取仍受原有 Goal、agent 和文件来源约束。- 新持久场景经过普通 CLI claim、真实 File/SQLite authority、hard lease、refresh-state、重放和冷读。外国 owner 拒绝时证据日志不变;第九条记录及全部历史保留,反驳 hazard 可见,重放不重复写入。
- API 场景核验第九条关联、已归档任务、四十一条结果的分页、跨 Goal cursor/Origin 拒绝、损坏源失败和恢复。现有 React 成果页按 canonical API 返回的关联展示,不另造数量上限。双语 README 明确持久关联不再限制八条,以及紧凑视图的预算和完整读取方式。
此前 README 明确写有八条容量,不能把本次变更说成始终未有限制。这里是有意扩展旧契约,同时修复读适配器静默丢关联;当前评审请求的持续摄取与完整负面证据结果,是这次容量调整的依据。新 README 属于受审变更,不用于自证正确。既有权限、显式关联、零诊断分数和只读展示合同保持。
对主干的风险
独立检查当前十个文件完整差异及未改动的 Explore audit、结果日志、Todo 变更/读取和 React 消费者。相同第九次捕获场景在不可变基线的 File、SQLite 上均失败,在当前 head 均成功;100 项 Python、15 项 TypeScript 测试通过,控制面 typecheck 通过。原生 premerge 的 5 项直接检查、19 项风险检查通过,包含语义、Todo 生命周期与公开边界;未查询远端 CI。
关闭能力的完整 turn-context、启动分发及旧合法形状在同一合成输入下 base/head 完全一致,并断言不会读取 Explore 日志。显式增加九条以上关联是有意放宽既有 API,即使仅为保存历史也生效;不会隐式激活 Graph/Harness。无关新结果、显示省略和分页不清除完整风险;invalid id、相互冲突的 clear/append、外国 owner、缺 lease、过期分页和损坏来源仍拒绝。
另从当前 head 构建正式 Chat bundle,用隔离原生服务器与合成 Goal 在 Ego Lite 验证“选择 Goal→成果→第九条关联→下一页”。整页能同时看见 Goal、refuted 状态、适用范围及关联任务;损坏证据时清空旧结果,恢复后点击刷新找回关联。没有 mock API、启动模型、改动活动 Goal 或安装候选到正在运行的 App。首次浏览等待了错误的“文件”标签,已按实际“成果”入口纠正;属于探针选择错误。bundle 构建保留既有大 chunk 提示,未当成失败或为它改预算。未验证安装升级、实际模型采纳、无限增长的数据规模或 PostgreSQL;本 PR 没有更改存储 provider/transport,不属于 authority-store 重构。
风险是持久关联可增长,完整审计成本随明确关联数增加;预算应约束投影,而非静默删数据。测试含四十条关联和省略计数;这不是无限规模容量或长程净耗时证明。无法读取来源时保持失败并提供恢复,不制造“没有反驳”的空白成功结果。
我的整体评价
APPROVE,head a1f6b8e。 这是一条完整且可回滚的现有能力修复:agent 可以继续记录第九条结果,后续冷读、同任务审计和用户成果页都能找到它;重复 refresh 不叠加日志,负面证据不被显示预算抹掉。对长程效果是保留已有承诺与反驳,对体验是免除清理旧关联或重录的补救动作;上下文维持有界,有结构上的效率收益。完整持久扫描有成本,模型是否减少重复实验、实际净耗时改善仍未实测。
未来改造检查已应用:用同一个 typed Todo owner 和原生 audit,Python set 去重替代重复线性检查,不引入重复关联账本、新状态或通用框架。关联 id 是显式意图;omitted 字段是可派生投影,不能代表缺乏证据。当前“先验证首个有用结果与真实消费者”的经验适用,本次用第九条 CLI 写回、完整审计和打包页面核验;历史案例裁决未继承,经验召回也不证明本次质量提升。评审结论不授予合并权,不关闭整个 Explore 路线图。
English verdict: APPROVE - a1f6b8e: durable links survive the ninth capture and replay while compact views preserve complete hazard classification and same-Todo expansion. Independent File/SQLite base-to-head regression, 100 Python tests, 15 TS tests, typecheck, 5 direct/19 risk checks, full disabled parity, and packaged UI pagination/failure/recovery passed. Model adoption and large-scale net efficiency remain untested.
A valid ninth Explore result attachment was rejected because the Todo mutation owner treated a context-display budget as a persistent link limit. The Python read codec could also silently drop associations beyond eight. Preserve all valid distinct node IDs in the existing Todo contract, while keeping the compact Explore turn view bounded.
todo-branch-plan, matching the Todo audit being summarized, rather than the different worker-lane grouping.Validation: the real ninth-capture CLI regression failed on the base for both canonical File and SQLite, then passed after the fix. The related Python suite passed 199 tests; final cold-command and HTTP readback checks passed 3 tests after the command correction. TypeScript owner tests passed 40 tests, and the final control-plane typecheck passed. Risk-based premerge completed 19 checks plus 5 direct checks with no failures or skips; the 10-file public-boundary scan was clean.
Coverage includes ninth-link preservation, exact refresh replay, wrong-owner denial, retained refuted evidence, result pagination, 40-link compact/full audit readback, and the existing frontend HTTP API association/recovery path. No frontend code changed; packaged visual interaction and PostgreSQL were not rerun. This changes the existing link codec/projection, not the authority provider or schema. Frozen benchmark workers were not modified and no benchmark speedup is claimed.
Placement/refactor: reuse the typed Todo mutation owner and Explore projection; the Python codec remains a read adapter. Replace quadratic list deduplication with an order-preserving set. No new capability, storage abstraction or parallel decision owner. This closes the bounded evidence-retention defect, not the broader research/adoption acceptance.
Maintainer merge required after exact-head independent review.