Skip to content

feat: add voice task success scorer - #241

Draft
kev (kevchoi) wants to merge 2 commits into
mainfrom
kevchoi/20261002-voice-task-success
Draft

kev (kevchoi) wants to merge 2 commits into
mainfrom
kevchoi/20261002-voice-task-success

Conversation

@kevchoi

@kevchoi kev (kevchoi) commented Oct 7, 2026 •

Copy link
Copy Markdown

AI-created or modified and not human-reviewed in its current form; treat this artifact as provisional. Updated 2026-10-07.

Replaces #232, whose branch was renamed to kevchoi/20261002-voice-task-success. Bottom of a stack: #241 ← #238 ← #239.

Summary

Adds VoiceTaskSuccess, an LLM judge for whether a voice agent completed the caller's request, judged from the whole call (instructions, tool calls, tool results).

  • templates/voice_task_success.yaml: A/B/C rubric over {{thread_with_system}}.
  • JS and Python VoiceTaskSuccess: picks the conversation from a thread_with_system argument, else trace.getThread() / trace.get_thread(). With no conversation it returns a null score without calling the model. Otherwise it passes the messages to the template as a JSON string, so tool calls and results are included.
  • js/manifest.ts entry, so it's available as a built-in scorer.
  • SCORERS.md, README.md, AGENTS.md.
  • Tests in JS and Python: the trace's conversation is sent as JSON, and a missing conversation skips the score.

Testing

  • Build, tsc --noEmit (one pre-existing error in js/render-messages.test.ts), vitest, pre-commit. Tests that need an OpenAI key fail the same way on main.
  • Ran JS and Python (sync and async) against a live model for each source: thread_with_system, trace, no conversation (null), empty trace (null), and .partial.
  • JS JSON.stringify and Python json.dumps produce identical text for the same conversation. The judge cites tool names and arguments (a transfer of $1000 to Bob when the caller asked for $10 to Alice scores 0).

🤖 Generated with Claude Code

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@kevchoi
kev (kevchoi) added this pull request to stack #242 October 7, 2026 14:50
@github-actions

github-actions Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

Braintrust eval report

Autoevals (HEAD-1791387778)

Score Average Improvements Regressions
NumericDiff 77.9% (-2pp) 3 🟢 10 🔴
Time_to_first_token 8.44tok (-0.94tok) 157 🟢 62 🔴
Llm_calls 1.55 (+0) - -
Tool_calls 0 (+0) - -
Errors 0 (+0) - -
Llm_errors 0 (+0) - -
Tool_errors 0 (+0) - -
Prompt_tokens 513.89tok (-1.49tok) 30 🟢 26 🔴
Prompt_cached_tokens 0tok (+0tok) - -
Prompt_cache_creation_tokens 0tok (+0tok) - -
Prompt_cache_creation_5m_tokens 0tok (+0tok) - -
Prompt_cache_creation_1h_tokens 0tok (+0tok) - -
Completion_tokens 466.72tok (-18.48tok) 117 🟢 97 🔴
Completion_reasoning_tokens 347.64tok (-20.65tok) 99 🟢 77 🔴
Completion_accepted_prediction_tokens 0tok (+0tok) - -
Completion_rejected_prediction_tokens 0tok (+0tok) - -
Completion_audio_tokens 0tok (+0tok) - -
Total_tokens 980.61tok (-19.97tok) 125 🟢 89 🔴
Estimated_cost 0$ (0$) 78 🟢 51 🔴
Duration 8.44s (-0.94s) 157 🟢 62 🔴
Llm_duration 9.09s (-1.13s) 161 🟢 58 🔴

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant