Skip to content

Use GPT-6 Luna as default scoring model - #230

Merged
Ankur Goyal (ankrgyl) merged 1 commit into
mainfrom
ankur/default-scorer-gpt-6-luna
Oct 1, 2026
Merged

Ankur Goyal (ankrgyl) merged 1 commit into
mainfrom
ankur/default-scorer-gpt-6-luna

Conversation

@ankrgyl

@ankrgyl Ankur Goyal (ankrgyl) commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • update AutoEvals' exported DEFAULT_MODEL constant to gpt-6-luna in TypeScript and Python
  • keep the SDK runtime fallback (getDefaultModel / get_default_model) unchanged at gpt-5-mini
  • add TypeScript and Python regression coverage for the shared scorer-model constant

Testing

  • pnpm run test --run js/llm.test.ts
  • OPENAI_API_KEY=test-key .venv/bin/python -m pytest -c pyproject.toml py/autoevals/test_llm.py -k default_scorer_model
  • pre-commit run --all-files
  • pnpm run build

@github-actions

github-actions Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

Braintrust eval report

Autoevals (HEAD-1790893223)

Score Average Improvements Regressions
NumericDiff 78.5% (+33pp) 84 🟢 3 🔴
Time_to_first_token 7.92tok (+2.74tok) 4 🟢 116 🔴
Llm_calls 1.55 (+0) - -
Tool_calls 0 (+0) - -
Errors 0 (-0.45) 99 🟢 -
Llm_errors 0 (+0) - -
Tool_errors 0 (+0) - -
Prompt_tokens 512.02tok (+244.68tok) - 114 🔴
Prompt_cached_tokens 0tok (+0tok) - -
Prompt_cache_creation_tokens 0tok (+0tok) - -
Prompt_cache_creation_5m_tokens 0tok (+0tok) - -
Prompt_cache_creation_1h_tokens 0tok (+0tok) - -
Completion_tokens 461.31tok (+422.58tok) - 219 🔴
Completion_reasoning_tokens 343.27tok (+340.64tok) - 219 🔴
Completion_accepted_prediction_tokens 0tok - -
Completion_rejected_prediction_tokens 0tok - -
Completion_audio_tokens 0tok (+0tok) - -
Total_tokens 973.34tok (+667.25tok) - 219 🔴
Estimated_cost 0$ (+0$) - 125 🔴
Duration 7.93s (+3.48s) 4 🟢 215 🔴
Llm_duration 8.75s (+3.53s) 6 🟢 213 🔴

@ankrgyl
Ankur Goyal (ankrgyl) force-pushed the ankur/default-scorer-gpt-6-luna branch from 8bde5a4 to 53f6168 Compare October 1, 2026 22:19
@ankrgyl
Ankur Goyal (ankrgyl) merged commit 38228ee into main Oct 1, 2026
4 of 15 checks passed
@github-actions

github-actions Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

Braintrust eval report

Autoevals (main-1790894503)

Score Average Improvements Regressions
NumericDiff 77.4% (-1pp) 5 🟢 11 🔴
Time_to_first_token 8.98tok (+1.06tok) 58 🟢 161 🔴
Llm_calls 1.55 (+0) - -
Tool_calls 0 (+0) - -
Errors 0 (+0) - -
Llm_errors 0 (+0) - -
Tool_errors 0 (+0) - -
Prompt_tokens 514.63tok (+2.61tok) 23 🟢 30 🔴
Prompt_cached_tokens 0tok (+0tok) - -
Prompt_cache_creation_tokens 0tok (+0tok) - -
Prompt_cache_creation_5m_tokens 0tok (+0tok) - -
Prompt_cache_creation_1h_tokens 0tok (+0tok) - -
Completion_tokens 464.69tok (+3.38tok) 113 🟢 97 🔴
Completion_reasoning_tokens 347.93tok (+4.65tok) 94 🟢 79 🔴
Completion_accepted_prediction_tokens 0tok (+0tok) - -
Completion_rejected_prediction_tokens 0tok (+0tok) - -
Completion_audio_tokens 0tok (+0tok) - -
Total_tokens 979.32tok (+5.99tok) 112 🟢 100 🔴
Estimated_cost 0$ (+0$) 82 🟢 61 🔴
Duration 8.98s (+1.06s) 58 🟢 161 🔴
Llm_duration 9.83s (+1.08s) 58 🟢 161 🔴

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants