Skip to content

Add SpeechClarity, TurnTaking, and VoiceTaskSuccess voice scorers - #229

Closed
kev (kevchoi) wants to merge 2 commits into
mainfrom
kevchoi/20261001-voice-scorers
Closed

kev (kevchoi) wants to merge 2 commits into
mainfrom
kevchoi/20261001-voice-scorers

Conversation

@kevchoi

Copy link
Copy Markdown

Summary

Adds three voice scorers. Two listen to the call recording itself, not a transcript:

  • SpeechClarity (audio): how easily a phone listener can understand every word the agent says.
  • TurnTaking (audio): whether the agent talks over the caller, cuts them off, or ignores interruptions. Response latency is ignored.
  • VoiceTaskSuccess (transcript): whether the agent correctly completed the caller's request, judged from {{thread_with_system}} (instructions, tool calls, tool results).

Grades are A = 1, B = 0.5, C = 0.

How

  • Each scorer is a template plus one manifest entry, under a new "Voice" group. No new classes.
  • Any LLM classifier now accepts an optional audio argument in OpenAI's input_audio shape ({data, format}). It's sent as a user message after the prompt, in both TS and Python.
  • Templates can set requires_audio: true. Those scorers fail fast without audio instead of judging nothing.
  • The audio templates default to gpt-audio. The gpt-5 models can't take audio input. OpenAI audio input accepts only WAV and MP3.

Testing

  • New mocked tests in TS (SpeechClarity sends audio to gpt-audio; audio is required) and Python (TurnTaking request shape, VoiceTaskSuccess renders the system prompt, audio is required).
  • The full suites show the same failures as main without a key (25 Python, 8 JS: missing credentials or litellm).
  • Not yet run against a real audio model; that needs an OpenAI key.

Draft. The product wiring is in the companion braintrust monorepo PR.

🤖 Generated with Claude Code

kev (kevchoi) and others added 2 commits October 1, 2026 00:47
LLM-as-a-judge scorers that listen to a voice call recording instead of a
transcript. Any LLM classifier now accepts an OpenAI input_audio `audio`
arg, sent as a user message after the prompt. The voice templates default
to gpt-audio.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
SpeechClarity and TurnTaking use the agreed prompts, and their templates
declare requires_audio so they fail fast without a recording.
VoiceTaskSuccess judges the call from the trace's thread_with_system.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

Braintrust eval report

Autoevals (HEAD-1790828103)

Score Average Improvements Regressions
NumericDiff 78.6% (0pp) 6 🟢 8 🔴
Time_to_first_token 7.66tok (-0.88tok) 154 🟢 65 🔴
Llm_calls 1.55 (+0) - -
Tool_calls 0 (+0) - -
Errors 0 (+0) - -
Llm_errors 0 (+0) - -
Tool_errors 0 (+0) - -
Prompt_tokens 514.26tok (-1.49tok) 25 🟢 21 🔴
Prompt_cached_tokens 0tok (+0tok) - -
Prompt_cache_creation_tokens 0tok (+0tok) - -
Prompt_cache_creation_5m_tokens 0tok (+0tok) - -
Prompt_cache_creation_1h_tokens 0tok (+0tok) - -
Completion_tokens 471.45tok (+5.78tok) 97 🟢 107 🔴
Completion_reasoning_tokens 356.36tok (+6.69tok) 76 🟢 86 🔴
Completion_accepted_prediction_tokens 0tok (+0tok) - -
Completion_rejected_prediction_tokens 0tok (+0tok) - -
Completion_audio_tokens 0tok (+0tok) - -
Total_tokens 985.7tok (+4.29tok) 101 🟢 104 🔴
Estimated_cost 0$ (+0$) 64 🟢 78 🔴
Duration 7.66s (-0.88s) 154 🟢 65 🔴
Llm_duration 8.38s (-0.94s) 156 🟢 63 🔴

@kevchoi
kev (kevchoi) deleted the kevchoi/20261001-voice-scorers branch October 2, 2026 22:13
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant