Skip to content

[Bug]: lead lane always fails with SESSION_POPULATION_BINDING_MISMATCH on Claude while sessionEvidence succeeds #164

Description

@nth2k8wjng-ux

Summary

On the Claude host, harness evidence-bundle can never return a usable bundle whenever the workspace has Claude sessions to analyze. The lead lane is unconditionally downgraded to unavailable with SESSION_POPULATION_BINDING_MISMATCH, which forces status: "failed" in both quick and normal depth. The Better Harness skill is then required to stop, so no report can ever be produced.

The same workspace analyzed with harness analyze directly succeeds, which shows the lead analyzer itself is healthy and the defect is in the bundle's population-consistency reconciliation.

Relation to existing issue #34 — this is a distinct variant

This is not a duplicate of #34 (SESSION_POPULATION_BINDING_MISMATCH when reviewing from inside the target workspace), which was closed on 2026-08-01 with no linked PR. The two reports differ in the failing lane, the divergence size, and the version:

#34 (closed) This report
sessionEvidence lane unavailable available
lead lane unavailable unavailable
Count divergence 1 (130 vs 131) 2–3 (18 vs 20/21)
Version 0.4.0 0.7.0-alpha1
Platform macOS arm64 Windows 10
Fix PR none linked —

Because sessionEvidence stays available here, the check in session-evidence.mjs (eligibleCount !== population.binding.eligible.count) evidently passes in this workspace — the failure occurs one stage later, in the lead-lane reconciliation. That is a different code path from the one #34 exercised, and it still reproduces on the current release.

The documented workaround is not reachable

#34 suggested running the review from a different working directory as an unverified workaround. That is not possible on the CLI, because --cwd is validated against the workspace:

$ node scripts/better-harness.mjs harness evidence-bundle \
    --platform claude --workspace <home> --cwd <other-dir> --depth quick --format json
harness evidence-bundle failed: cwd must resolve inside workspace

So a user with a non-empty session store has no supported way to obtain a bundle on the Claude host — neither by changing cwd nor by narrowing the window (see the window matrix below).

Environment

Item Value
Package @qoder-ai/better-harness 0.7.0-alpha1
Node v24.15.0 (engines >=22.20.0 <25.0.0 — satisfied)
OS Windows 10 Pro 19045
Install source Claude Code plugin marketplace (QoderAI/better-harness)
Platform flag --platform claude
Workspace a home directory; not a git repository
~/.claude/projects 2 slug directories

Steps to reproduce

  1. Use a workspace whose Claude session store is non-empty, on a machine where the workspace is not a git repository:

    node scripts/better-harness.mjs harness evidence-bundle \
      --platform claude \
      --workspace <home> \
      --cwd <home> \
      --language en \
      --depth quick \
      --format json
  2. Observe status: "failed" and lead.status: "unavailable".

  3. For contrast, run the standalone analyzer against the same workspace:

    node scripts/better-harness.mjs harness analyze \
      --platform claude --workspace <home> --cwd <home> --format json
  4. Observe that it succeeds and reports a complete lead contract.

Reproduced at three different windows, so it is not window-dependent:

depth window eligible (bundle) lead lane bundle status
quick 7d (explicit) 7 unavailable failed
normal 30d (explicit) 18 unavailable failed
normal default (unset) 18 unavailable failed

Control case — a workspace with zero eligible sessions:

workspace eligible lead lane bundle status
other directory 0 available partial

So the failure is gated on eligible > 0.

Expected vs actual result

Expected: harness evidence-bundle returns a bundle whose lead lane is available (as harness analyze does for the same input), so the skill can proceed to reconciliation.

Actual: lead is replaced by unavailableLane(..., SESSION_POPULATION_BINDING_MISMATCH) and status becomes failed, blocking every downstream step.

Evidence

bundle.diagnostics.sessionPopulationBinding (default-window run):

{
  "status": "conflict",
  "population": {
    "eligible": { "count": 18, "fingerprint": "431e788f61a0a9eb" }
  },
  "sessionSelection": {
    "strategy": "all-eligible",
    "selected": { "count": 18, "fingerprint": "431e788f61a0a9eb" }
  },
  "leadSelection": null,
  "episodes": {
    "comparison": "not-comparable-selection-or-policy",
    "sessionTaskEpisodes": 105,
    "leadProjectedEpisodes": 0,
    "leadRetainedEpisodes": 0,
    "leadZeroSignalDiscardedEpisodes": 0
  },
  "errorCodes": ["SESSION_POPULATION_BINDING_MISMATCH"]
}

Lane states for the same run:

sessionEvidence : available
projectHarness  : unavailable | GIT_COMMAND_FAILED
agentCustomize  : available
lead            : unavailable | SESSION_POPULATION_BINDING_MISMATCH

Standalone analyzer against the same workspace — succeeds, and its public selection counts disagree with the bundle's population binding:

{
  "summaryFacts.evidenceBoundary.manifest.selection": {
    "strategy": "all-eligible",
    "eligibleCount": 21,
    "analyzedCount": 21,
    "confidence": "High"
  }
}

With an explicit 30-day window pinned to match the bundle run, harness analyze still succeeds with eligibleCount: 20 — while the bundle's population.binding.eligible.count for the same window is 18.

Note also that the standalone harness analyze JSON has no top-level sessionBinding key at all (keys: kind, schemaVersion, evidence, summaryFacts).

Suspected cause

populationDiagnostics() reads the lead binding from lead?.data?.sessionBinding:

scripts/harness-analysis/evidence-bundle/index.mjs:74

const leadBinding = lead?.data?.sessionBinding ?? null;

It then reconciles public counts against the frozen population:

scripts/harness-analysis/evidence-bundle/index.mjs:87-90

if (Number(leadSelection.eligibleCount ?? -1) !== population.binding.eligible.count
  || Number(leadSelection.analyzedCount ?? -1) !== leadBinding?.selection?.selected?.count) {
  errors.push("lead public counts do not match its population binding");
}

When the lead data carries no sessionBinding (or no summaryFacts.evidenceBoundary.manifest.selection), both comparisons collapse to -1 !== count, which is always true for a non-empty population — hence the conflict is unconditional whenever eligible > 0. The conflict then destroys the lead lane itself:

scripts/harness-analysis/evidence-bundle/index.mjs:151-157

const sessionPopulationBinding = populationDiagnostics(sessionPopulation, sessionEvidence, lead);
if (sessionPopulationBinding.status === "conflict") {
  lead = unavailableLane("lead-analyzer", Object.assign(
    new Error("Session population binding mismatch"),
    { code: "SESSION_POPULATION_BINDING_MISMATCH" },
  ));
}

Two things look worth checking:

  1. Population-source divergence. collectSessionPopulation and the lead analyzer appear to select sessions under different criteria for the same window (18 vs 20/21). The reconciliation exists to catch exactly this, but here it is the only thing that trips — and it takes the whole report down with it.

  2. sessionBinding contract. collectLead passes sessionPopulation into analyzeHarnessEvidence and returns its payload verbatim (index.mjs:34-61), while populationDiagnostics expects a data.sessionBinding sub-object. Worth confirming which shape the analyzer emits on the claude platform when sessionPopulation is injected.

Secondary observation (may be by design)

In normal depth a non-git workspace also fails via projectHarness: GIT_COMMAND_FAILED, because normal treats any unavailable lane as blocking (index.mjs:164). Users who run Claude Code from a non-versioned directory (a global ~/.claude harness home, for example) can therefore never obtain a normal report even after the lead defect is fixed. If that is intended, a clearer precondition error early in evidence-bundle would be friendlier than a late failed bundle.

Related: using a home directory as --workspace produces a partial topology — inventoryMode: "filesystem-fallback" scanned 50000 files of 50001 (truncated: true), yielding 108 members made up largely of plugin-cache directories. That is bounded-scan behavior rather than a defect, but it makes the intended workspace choice for a "global harness review" unclear.

Impact

Blocks the Claude host entirely for any user whose session store is non-empty — report generation is impossible, so the whole skill short-circuits at Step 1.

Activity

  1. phodal commented on Sep 11, 2026

    @phodal
    Member

    Welcome to PR @nth2k8wjng-ux

  2. 1339190177 commented on Sep 23, 2026

    @1339190177
    Contributor

    认领留言草稿(发在 #164)

    Hi, I'd like to help with this one. I dug into the lead lane path on current
    main (post-alpha2) with synthetic Claude homes and found a diagnosability
    defect that almost certainly wraps whatever is really failing on your machine:

    When the lead analyzer throws for its own reason, populationDiagnostics
    reconciles lead-side counts from absent data (lead.data missing → every
    count compares as -1), reports conflict, and collectEvidenceBundle then
    replaces the lead lane with a generic SESSION_POPULATION_BINDING_MISMATCH
    envelope
    — the original error code and message are discarded. So
    "unconditionally fails with SESSION_POPULATION_BINDING_MISMATCH" is what a
    failed lead looks like from the outside, regardless of the real cause.

    Two observations that may matter for the remaining Windows-specific failure:

    1. The bundle's population discovery runs with the frozen topology /
      analysisScope (strict workspace qualification) while standalone
      harness analyze runs without them (lenient). That explains the
      18 vs 20/21 divergence between the two commands without any in-bundle
      count disagreement.
    2. In my synthetic repros (Linux, git and non-git workspaces, multi-slug
      homes, root-candidate cwd sessions, live-session exclusion), the frozen
      population, session lane, and lead always reconcile — so I could not
      reproduce a genuine binding mismatch on Linux. The real lead failure on
      your machine is likely Windows-specific (path flavor / drive-letter
      handling in workspace matching or transcript reads).

    I've opened #187, which stops the masking: an unobserved lead lane now
    keeps its own failure code on the lane, binding diagnostics record
    leadObserved instead of manufacturing lead errors, and genuine conflicts
    additionally expose the concrete reconciliation error strings. The fail-closed
    contract is unchanged. With this change, re-running harness evidence-bundle
    on the affected workspace should reveal the actual error behind #164.

    Happy to iterate on the follow-up once the real code is visible.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions