Skip to content

[Typed evaluation P2] Optimize executable decision programs with independent final assessment #1299

Description

@drewstone

Parent: tangle-network/agent-eval#768. Roadmap item 14; M8. Owner: Runtime improve/strategy evolution with Eval comparison. Depends on #1296/#1297/#1298 and agent-eval#770/#771.

  • Use existing improve() and strategy evolution to vary questions, evidence renderers, thresholds, models and branch policies, not only prose prompts.
  • Execute exact candidate programs and retained state under comparable resources; account for search, evaluations, retrieval and failed/incomplete attempts.
  • Re-score saved distributions for deterministic mapping changes, but rerun changed execution paths before claiming downstream improvement.
  • Keep final cases and final judge feedback outside adaptive search; retain candidate/version lineage and require explicit promotion of the exact assessed program.

Acceptance: an executable policy beats an appropriate frozen baseline on untouched tasks with uncertainty and full costs reported. Work on working-evaluator quality is assessed separately from optimizing a worker against it. No new optimizer facade or automatic self-modification authority; a higher development judge score alone is not success.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions