The loop refuses in one voice, and a persona loads by fit - #222
Merged
Merged
Conversation
…ng on refusals Eight version declarations moved in lockstep: pyproject.toml, __init__.py, package.json, package-lock.json (2 sites), .claude-plugin/plugin.json, the ENGINE stamp in all four engine twins, and the metadata version in all three skill trees. The dogfood bundle's index.md engine stamp follows, and ENGINE_MD5 plus the SKILL.md sha256 prose pin are re-aimed. CHANGELOG records what the two milestones actually found, including the two items carried into 3.7 knowingly rather than quietly: eleven verbs still answering a refusal with `False`, and the SKILL.md prose pin still held in two files. Suite: 1439 passed, 7 skipped, 0 failed. The tag and the npm/PyPI publish are human-owned and NOT done here.
…op` is the 27th verb
3.6.0 made a 40-task roadmap legible: every row carries its title, the headline counts
the unauthored ones. It still could not answer what a reviewer actually asks — is this a
queue or a graveyard? A task the plan is working toward and a task the plan walked away
from both read `[scaffold]`.
An unauthored task now reports which plan wants it:
yowo — … · 2 queued · 1 abandoned · 1 adrift (add todo)
· cli-bounded-memory [queued] Bounded CLI memory
· legacy-shim [abandoned] Legacy config shim
· stray-idea [adrift] Something I jotted down
DERIVED, never stored (R:NEWFIELD). The plan that queued a task IS its milestone, so the
answer already exists in two places that cannot disagree — the task's `milestone:` and
that milestone's `status:`. A third copy would drift out of step with the thing it
describes. `archived` counts as closed alongside `done`; a `milestone:` naming nothing
is `adrift`, identical to naming none, and a read verb never raises on the typo.
`milestone-done` is where a graveyard comes from. It tallied exit criteria and never
looked at its member tasks, so a milestone closed and whatever it queued stopped being
anybody's, with nothing said. It now refuses, names every unauthored member, and offers
all three real exits — author it, drop it, or re-home it under a live milestone. A
refusal naming one fix does not present a choice, it applies pressure.
`add drop <slug> --reason "<why>"` is the 27th verb. `dropped` was already a status the
engine READ in three places and no verb could WRITE — vocabulary living only in the
reader (R:DEADWORD). Withdrawing work meant deleting a file or leaving it to rot as a
scaffold, and neither leaves a reason. The reason is required for the same purpose the
whole task serves. A `done` task is refused: that verdict was gated against a receipt,
and `reopen` is the verb that revisits it without erasing it.
The verb ripples into every registry that enumerates the set — CLI WIRED lists, three
skill trees, two READMEs, the book command reference, four count pins. E6 records that
cost, and a check now enumerates the registries so the ripple is finite rather than
discovered one CI failure at a time.
Four checks pinned the literal word `scaffold`. Rule kept, expired premise replaced:
each now asserts membership in `SCAFFOLD_KINDS` and that the word is not `direction`.
`loop.md`'s Gather step promised one word and now promises three — each registered and
proven from driven stdout, including re-detecting the sentence after a rewrite dropped
it below the claim detector's verb list.
SKILL.md lands at exactly 176/176 lines and 13258/13258 bytes: the new cookbook line and
the wired-surface entry are funded by compressing the Engine paragraph, the init bullet
and three cookbook comments — not by moving a human-set pin.
Suite: 1444 passed, 7 skipped, 0 failed.
Records the three surfaces that closed the last gap from the yowo PR #16 review, plus the three lessons they filed, drained and bound as decisions: - every status the engine reads has a verb that writes it - a close inspects what it holds - a count of pending work splits by the property the reader will act on Suite: 1444 passed, 7 skipped.
…rientation report
On this bundle `add status` printed three rows — PROJECT, `index`, and an ARCHIVED
milestone — while withholding 112 nodes, and ended in `next: add new task <slug>`,
which is not a command. It was tuned for a bundle mid-flight and degraded at both ends.
Before / after, same bundle:
AIDD-Book — … · 295 nodes AIDD-Book — … · 296 nodes
· PROJECT [—] · status-answers… [build] Task status answers…
· v3-final-collateral [archived] … 3 Persona · 1 Project · 5 Spec · 1 manifest
· index [—] carrying no state — not listed (`--all`)
next: add new task <slug> last: run a-plan-says-what-it-wants · 2026-09-10
next: add brief status-answers-what-needs-me
What was wrong, and what each fix was:
- ordering was by node TYPE, so an archived milestone outranked every open task. Rows now
sort by ATTENTION_RANK — verify · build · direction · queued · abandoned · adrift — with
an unrecognised beat FIRST, because that is exactly when a person should look.
- `archived` was missing from the answered set, so the deadest state in the engine was the
one work row a finished bundle showed.
- the stateless-row rule was a TYPE LIST naming Spec and Persona, so Project and the
manifest kept their exemption from the rule written to remove them. It is a predicate now.
- `status --all` printed "(`--all` for done nodes)" — a hint advising the flag already in
force, with no other route to those rows. The cap applies to the bare report only, and its
hint names `add status --all`, driven and proven to produce exactly the withheld rows.
- `BEAT_NEXT["build"]` handed back `<test cmd>` on EVERY frozen task. A notary cannot know a
project's test command — but it is given one on every `run`, so `run` now remembers it on
`index.md` and the hint replays what actually worked here.
- rows wrapped at 108 columns and the type column never lined up, because `[{beat}]` is
variable-width. Every column is fixed now and the title takes what is left, at 100.
- an empty board printed nothing, indistinguishable from a broken read. It says
`nothing needs you — N answered, M carrying no state`, and `--all` says how many rows are
coming before they scroll.
New: a `last:` line naming the most recent recorded act and its node — absent, never guessed,
when nothing has happened. A resume point that omits the last session is not a resume point.
Six checks re-aimed (rule kept, expired premise replaced), and one skill claim REPAIRED
rather than culled: `seed.md` promised a stateless node "shows `[—]`", which stopped being
true when those rows were collapsed. Both prose repairs had to be re-detected — the claim
guard reads line by line and needs its rendering verb beside the backticked command.
Suite: 1449 passed, 7 skipped.
The gate refused the PASS: M5 (an empty board says so) was asserted by test_the_board_shows_only_what_needs_a_decision and named by no covers: entry. The check was always proving it; the binding was missing.
Three lessons filed, drained and bound as decisions: - orientation answers what needs a decision; answered or stateless nodes are context - a hint carrying a slot the ENGINE could have filled is a defect; a slot only the human can fill is guidance - a rule about a property is a predicate on that property, never a list of the types that happened to have it
…e 3 that stay False Two checks in one branch were GREEN only because `freeze` refused with `False`, and `False is not None`. Eleven verbs still answered that way. Each was self-consistent, so none looked wrong on its own — the caller crossing two verbs was the one who paid. The sweep is NOT "every False becomes None". The question is whether the verb DID its work or ANSWERED a question: - 37 refusals now return `None`, including two in `done` that a 2-tuple survey could not see because `done` returns 3-tuples. - `fresh` keeps `False` for stale — a receipt that observed code which has since changed is the ANSWER a caller asked for — and gains `None` for "freshness cannot be established". Three states, and the third is the point: collapsing them would report a clean `stale` over a question the engine could not ask. - `render_card` keeps `False` for "card is current". - six verbs answer an empty COLLECTION from a query that ran (`deltas`, `search`, `interview`, `todo`, `locate`, `neighborhood`). That is an answer too. My own caller sweep wrongly tightened 10 of those to `is None` before the suite caught it. Callers: 57 `assert not <name>` tightened to `is None` where the name came from a refusal verb, and 32 `is False` assertions converted. Both were done by AST — matching the name to the `add.<verb>()` it was bound from — never by text, because `assert not hits` about a list of grep matches has nothing to do with this contract and flagging it would train the reader to ignore the guard. The message pin names its RANGE, not the working tree — this repo's own bound decision from two tasks ago. It asserts LOSSES only: a reword shows as one loss and one gain, so losses still catch every edit, while a message added by a later rung on a stacked branch is not this check's business. Proved red by rewording one refusal. Suite: 1454 passed, 7 skipped.
…nly edges `uncovered_obligations` bound filled edges and probed assumptions and said in its own docstring that Musts and Rejects stayed out `until the cost of widening is MEASURED rather than estimated`. This measures it and widens the rung. The measurement, run over this bundle at direction: 0 of 105 nodes carrying RULES have an uncovered Must or Reject, and 21 carry no RULES at all. Nothing in the corpus is newly refused — because the GATE already refuses an uncovered Must, so nothing could have shipped with one. What the widening buys is the TIMING: the author learns at the freeze, holding the pen, instead of at the gate after the whole build, when the fix is still one line of authoring. The rung is now one line — every referent minus everything a `covers:` names — so freeze and gate read one definition of an obligation rather than two. - retire test_the_rung_does_not_widen_to_musts: it held a real line until the measurement existed, and now asserts against its own condition being met - drop the E/A prefix filter in test_the_rung_and_the_gate_agree — a filter that outlives its reason is what makes two definitions of one thing - re-aim two fixtures whose premise expired: the `1 uncovered` count now covers its Reject so the count is again about the single edge, and the refreeze fixture rewords a check instead of deleting one Q33 (a count check is pinned to its fixture's shape) and M46 (a rung held narrow until measured is a debt with a due date) filed. author: Tin Dang
Both tasks gated PASS and both exits held on the sweep: no check anywhere asserts a refusal by truthiness where identity is meant (enumerated by test_no_caller_is_blind_to_the_change), and the two changes are separable in history — 94dd3b5 moves the shape, f47bd74 widens the rung. Deltas drained, each bound to a decision: M45 — None is the refusal; a verb that ANSWERS keeps its boolean M46 — a narrowing held "until measured" is retired in place when it is Q32 — survey call sites by CALLABLE, never by the arity the first few share Q33 — when the producer widens, re-aim the fixture, not the expected number author: Tin Dang
Three entries: the 37-site normalisation to `None`, the suite sweep for truthiness assertions on a refusal, and the widened `R:UNCOVERED` rung with the measurement that unblocked it (0 of 105 nodes). The "carried into 3.7" bullet about eleven verbs still answering `False` STAYS, with a resolved marker appended. It was true when it was written, and the append-only guard on this file exists precisely so a dated record is not brought into line with the present — the marker says what happened without falsifying what was said. author: Tin Dang
Personas loaded only when a subagent was spawned. Measured over this bundle:
15 of 190 lifecycle nodes carry a lens, and 3 of 40 milestones — the one lane
intake.md already says must load one.
The mandate was never missing. It is well written and lives in
agents/add-worker.md §2 ("Become the persona FIRST — before any task-specific
instruction"), a file loaded ONLY on a spawn. The three beat guides mentioned a
persona zero, zero and once, and that once is the security R:NOCOVERAGE
paragraph. So the lens was a function of whether the human spawned an agent.
- direction.md, build.md and verify.md each carry the selector now, appended to
the paragraph that already marks the beat's entry (GROUND, the brief, and the
trust sentence): the project roster on `flow:` + `task-kinds:`, the teacher
index second, proceed if neither fits, `add advise` to record it
- SKILL.md stops calling the LOAD opt-in. `opt-in` is true of the ROSTER (a
bundle may have none) and false of the load (if a persona fits, it loads)
- the four copies of the selector — three guides plus add-worker.md §2 — are
ENUMERATED and pinned to agree on both keys and the tier order, so a fifth
surface cannot state a different rule and no copy drifts alone
Funded, not re-pinned: +7 lines of instruction paid by -7 compressed from
streams.md, so the surface holds at 1492 and SKILL.md at exactly 176/176 lines
and 13258/13258 bytes. The sha256 prose pin is re-aimed with its reason.
Both checks that were green at red-first were proved red by withholding their
subject — one mutation (the selector added to explore.md) fires both, each for
its own reason.
author: Tin Dang
`brief` emitted `<persona>` only when the node ALREADY carried `persona:`/`advised_by:`. The only verb that stamps one is `add advise`, and no next-hint on the normal path names it — the todo row refuses a second verb by design (A12). So a lens could reach a node only through an act nothing in the loop ever asked for, and 175 of this bundle's 190 lifecycle nodes carried none. The no-lens branch now lists the Persona nodes whose `flow:` names this beat's surface and whose `task-kinds:` covers the node's `kind:`, and names `add advise` as the verb that records a pick. The circle closes at the one surface that is already the agent's instructions. The NO-EXEC floor is the point, not an obstacle: - the engine emits the FITTING SET and stops — ranking, choosing and loading stay the orchestrating agent's judgment - sorted by slug, because an alphabetical list is visibly not a ranking and any other order would be a preference the engine has no basis for - frontmatter only; a candidate's body never opens, so a brief cannot grow by the size of a persona it did not pick - an absent `kind:` skips the kind gate — it is optional on a Task, and gating on it would leave the roster dark for most of a bundle - a roster with no fitting entry emits what it always did, byte for byte `LENS_SURFACE`/`LENS_FALLBACK` restate `add-worker.md` §2's mapping rather than inventing a second one, so the sequential path cannot route a node to a different lens than a delegated one would. Both checks that were green at red-first were proved red by mutation: leaking a candidate's body fires R:BODYLEAK, and offering candidates beside a recorded lens fires A2. Four twins mirrored, ENGINE_MD5 re-aimed (engine_pin.py has twins too). author: Tin Dang
Both tasks gated PASS. The selector now reads the same on the sequential path
as it does inside a spawn, and an unlensed brief names who could fit.
Deltas drained, each bound to a decision:
M47 — an instruction only exists on the paths that LOAD the file holding it;
enumerate the copies in a check rather than trusting one statement
S13 — name what is optional; a word covering both a capability and the data
it reads will be taken as covering the capability
CHANGELOG records both entries under 3.6.0.
author: Tin Dang
CI went red on a check this repo wrote two tasks ago. `git merge-base HEAD origin/main` returns 128 on `actions/checkout`'s depth-1 clone, and the check read "cannot establish a baseline" as "the claim is false" — passing locally and failing CI for a reason unrelated to what it asserts. Behind that one failure was a class. Nine call sites in the suite compare the working tree to a git ref, and every one of them is satisfied by `git commit`. The repo had already bound the decision (Q28) and built a shape guard for it — over `git diff` only, because that is the spelling the two retired guards used. the-message-pin-is-a-content-pin The baseline is now a sha256 over the sorted (verb, message) pairs the check already extracts, with the aiming task and reason on its line. No git, so it runs identically on a shallow clone; and it keeps holding after this branch merges, which a merge-base containing the change could not. A third rung refuses the easy fix: no skip, no try, no early return, no subprocess. a-head-guard-declares-its-lifetime The shape guard now enumerates `git show` and `git merge-base` alongside `git diff`. Two new rungs: a ref a fresh shallow checkout lacks is refused, and every function comparing the working tree to HEAD must say in its own source that it is a live-editing tripwire — it fires while the edit is made and is inert once committed. That is a legitimate lifetime; the defect was leaving it unlabelled so a green CI read as the claim having held. All seven keep their assertions; a rung refuses annotating a guard into an empty one. one-home-for-the-prose-pin Carried out of the 3.6.0 cut. The SKILL.md sha256 was a literal in test_surface.py while test_skill_reads_the_graph.py recovered it by regexing that file's source. Both pins are now `skill_budget.PROSE_PINS`, imported. The claim is driven, not restated: the owning guard is pointed at a mutated tree and must still raise. Also: `feature-builder` seeded as the fourth persona — the roster had no lens for building a user-facing behaviour, so every `kind: feature` task got no candidate; 155 of 155 tasks now fit one. And `specs/domain.md` is authored, clearing the last doctor warning: it still held the scaffold placeholder in `## Decisions that bind`, which `bind_sections()` fed into every brief. Narrowed while here: `test_the_addition_was_funded` pinned skill_budget.py byte-for-byte, which is broader than its rule — adding an unrelated constant reported a budget bump that never happened. 1475 passed, 7 skipped in tests/ · 8 in tooling/. Every new check proved red by withholding its subject. author: Tin Dang
Collaborator
Author
|
Pushed
Also:
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Supersedes #221 (and #220, #219) — same linear stack, all their commits are contained here. Merging this closes the 3.6.0 line; the tag is still not cut and npm still serves 3.5.0.
Two milestones, both closed on checked exit criteria.
refuse-in-one-voice— the loop refuses in one voice, and earlyTwo checks in an earlier branch were green only because
freezerefused withFalseandFalse is not None. The shape was not one rung's habit.Noneis the refusal, bundle-wide — 37 return sites, enumerated by a check so a 28th verb cannot answer differently. Three stayFalseon purpose and say why:freshdistinguishes stale from cannot establish,render_cardanswers the card is current. Six verbs whose empty collection is an answer are exempt too. Messages and everycli.pyexit code are byte-identical.FalseandNonealike. Swept, and the shape is now refused suite-wide by an AST check narrowly aimed at that binding.freezedemands a check for every Must and Reject, not only edges — theR:UNCOVEREDrung held Musts out until the cost was measured. Measured: 0 of 105 nodes carrying RULES have an uncovered Must or Reject, 21 carry no RULES at all. Nothing is newly refused — the gate already refused it — so the widening buys the timing: the author is asked while still holding the pen. The rung is one expression now, so freeze and gate read one definition of an obligation.personas-load-by-fit— a persona loads by fit, on the path that does not spawnReported from use: personas were adopted only inside a spawned subagent. Measured over this bundle: 15 of 190 lifecycle nodes carried a lens, and 3 of 40 milestones — the one lane the skill already says must load one.
The mandate was never missing. It lives in
agents/add-worker.md§2 ("Become the persona FIRST"), a file loaded only on a spawn.phases/direction.mdandbuild.mdmentioned a persona zero times;verify.mdonce, inside the securityR:NOCOVERAGEparagraph.flow:+task-kinds:, teacher index second, proceed if neither fits,add adviseto record it. The four copies (three guides +add-worker.md§2) are enumerated and pinned to agree on both keys and the tier order.opt-inis true of the roster (a bundle may have none) and false of the load. A roster-less bundle still behaves exactly as before.briefnames the roster entries that fit.briefemitted a persona only when one was already recorded, and the only verb that records one isadd advise— which no hint on the normal path names, because thetodorow refuses a second verb by design. The lens could reach a node only through an act nothing ever asked for.The NO-EXEC floor is the point, not an obstacle: the engine emits the fitting set and stops, sorted by slug so the list is visibly not a ranking, reading frontmatter only so a brief cannot grow by the size of a persona it did not pick, and byte-identical to before when nothing fits.
Budgets
Funded, never re-pinned. +7 lines of instruction paid by 7 compressed out of
streams.md: the skill surface holds at 1492/1500 and SKILL.md at exactly 176/176 lines, 13258/13258 bytes. The sha256 prose pin is re-aimed with its reason.Verification
1467 passed, 7 skippedintests/and8 passedintooling/— both roots, what CI runs. Every check that was green at red-first was proved red by withholding its subject: one mutation adding the selector toexplore.mdfires two of them, and leaking a candidate's body or offering candidates beside a recorded lens fires the other two.Four engine twins mirrored and
ENGINE_MD5re-aimed — includingengine_pin.py's own twins, which is whattest_all_four_twins_and_both_pins_agreecaught.Decisions bound
Noneis the refusal; a verb that ANSWERS keeps its booleanKnown and carried
kind: feature, sopersona_candidatesreturns empty there. That is a roster gap, now visible instead of silent.