Skip to content

fix: preserve Luna context-limit error code - #284

Open
deep1725 wants to merge 1 commit into
agentcontrol:mainfrom
deep1725:fix/luna-context-limit-code
Open

deep1725 wants to merge 1 commit into
agentcontrol:mainfrom
deep1725:fix/luna-context-limit-code

Conversation

@deep1725

@deep1725 deep1725 commented Oct 6, 2026 •

Copy link
Copy Markdown

Summary

  • Preserve metadata.error_code: context_limit for the Luna model server's explicit maximum-length rejection.
  • Keep the server's existing error-text redaction intact. Generic runtime failures are not classified as context limits.
  • Document the code so clients can distinguish unsupported inputs from transient evaluator errors.

Scope

  • User-facing/API changes: one safe metadata value on confirmed context-limit failures.
  • Internal changes: classify the existing RuntimeError produced by a failed scorer response.
  • Out of scope: changing the model context window or retrying/truncating input.

Risk and Rollout

  • Risk level: low; successful results and other failures are unchanged.
  • Rollout requires publishing the updated Galileo evaluator and rebuilding consumers that pin its version. Roll back by restoring the previous package version.

Testing

  • make lint and make typecheck passed for the Galileo evaluator using public PyPI.
  • No tests added or run locally; automated CI results will be reported separately.

Checklist

  • Updated documentation for the metadata code.
  • Identified evaluator package release and consumer image update as required follow-up work.

Related PRs

Repo PR Purpose Status
agentcontrol/agent-control #284 Safe Luna context-limit metadata Open
rungalileo/ml-deploy #159 Explicit exclusions and valid latency evidence Open
rungalileo/orbit #2681 Opt-in H100 128K ceiling and exact 100K GPU probe Draft
rungalileo/galileo-infra-automation #1140 H100 validation Helm context and served-context check Draft

Context validation rollout

Release the Wizard image that permits the opt-in H100 128K ceiling.
Validate the production Gemma bundle with the exact 100K GPU probe before enabling the Infra configuration.
Confirm the selected chart uses that image and the live model catalog advertises 131,072 tokens.
Then dispatch a fresh Aegis evaluation with the fixed worker and evaluator.
The two context PRs remain drafts until live GPU validation succeeds.
Historical evaluation attempts keep their original evidence.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant