Skip to content

feat: add speech clarity audio scorer - #239

Draft
kev (kevchoi) wants to merge 4 commits into
kevchoi/20261002-voice-timing-scorersfrom
kevchoi/20261007-speech-clarity
Draft

kev (kevchoi) wants to merge 4 commits into
kevchoi/20261002-voice-timing-scorersfrom
kevchoi/20261007-speech-clarity

Conversation

@kevchoi

@kevchoi kev (kevchoi) commented Oct 7, 2026 •

Copy link
Copy Markdown

AI-created or modified and not human-reviewed in its current form; treat this artifact as provisional. Updated 2026-10-07.

Adds SpeechClarity, an LLM judge that listens to a voice call recording and rates how easily the agent's speech can be understood: garbled or cut-off words, distortion, dropouts, echo and noise. Scores are A = 1, B = 0.5 and C = 0. This PR carries over only the speech clarity part of #228. TurnTaking, VoiceTaskSuccess (#241) and the thread changes stay out.

  • templates/speech_clarity.yaml: the prompt, with model: gemini-3.8-flash and a new audio: input.audio field.
  • Templates take an optional audio path (JS LLMClassifierFromTemplate, Python LLMClassifier). The value there must be {data: bytes, content_type: "audio/..."} or a list of them, such as a recording split into chunks. Each is sent unchanged as its own base64 file part after the prompt text, in list order; nothing is joined. If the value is missing or the list is empty, the score is skipped (null). Chunk boundaries aren't marked in the prompt, so a word split across two chunks may be judged as cut off.
  • js/manifest.ts: lists SpeechClarity as a built-in under "LLM-as-a-Judge". Built-ins get audio as an attachment reference, so it needs the API server to pass the bytes (braintrustdata/braintrust#21432, in the matching braintrust stack). Merge the stacks together.
  • SCORERS.md documents SpeechClarity and the audio: field; README.md and AGENTS.md list it.

Only Gemini reads audio this way. Tests mock the API, so the scorer hasn't been run against a real model or calibrated against human ratings.

Testing

  • pnpm run test -- js/llm.test.ts: passes.
  • uv run --extra dev --extra scipy pytest py/autoevals/test_llm.py -k speech: passes. Five unrelated tests in that file need an OpenAI key and fail without one.
  • tsc --noEmit reports one error in js/render-messages.test.ts. This PR doesn't touch that file.

🤖 Generated with Claude Code

@github-actions

github-actions Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

Braintrust eval report

Autoevals (HEAD-1791394521)

Score Average Improvements Regressions
NumericDiff 78.4% (0pp) 9 🟢 10 🔴
Time_to_first_token 8.77tok (+2.17tok) 15 🟢 204 🔴
Llm_calls 1.55 (+0) - -
Tool_calls 0 (+0) - -
Errors 0 (+0) - -
Llm_errors 0 (+0) - -
Tool_errors 0 (+0) - -
Prompt_tokens 516.87tok (+0.37tok) 22 🟢 23 🔴
Prompt_cached_tokens 0tok (+0tok) - -
Prompt_cache_creation_tokens 0tok (+0tok) - -
Prompt_cache_creation_5m_tokens 0tok (+0tok) - -
Prompt_cache_creation_1h_tokens 0tok (+0tok) - -
Completion_tokens 460.81tok (-12.29tok) 112 🟢 99 🔴
Completion_reasoning_tokens 345.89tok (-13.38tok) 89 🟢 77 🔴
Completion_accepted_prediction_tokens 0tok (+0tok) - -
Completion_rejected_prediction_tokens 0tok (+0tok) - -
Completion_audio_tokens 0tok (+0tok) - -
Total_tokens 977.68tok (-11.92tok) 110 🟢 101 🔴
Estimated_cost 0$ (0$) 72 🟢 60 🔴
Duration 8.78s (+2.17s) 15 🟢 204 🔴
Llm_duration 9.63s (+2.52s) 13 🟢 206 🔴

@kevchoi
kev (kevchoi) changed the base branch from main to kevchoi/20261002-voice-timing-scorers October 7, 2026 14:35
@kevchoi
kev (kevchoi) force-pushed the kevchoi/20261007-speech-clarity branch from 07fbf1c to 8c665cc Compare October 7, 2026 14:35
@kevchoi
kev (kevchoi) added this pull request to stack #240 October 7, 2026 14:35
@kevchoi
kev (kevchoi) removed this pull request from stack #240 October 7, 2026 14:50
@kevchoi
kev (kevchoi) added this pull request to stack #242 October 7, 2026 14:50
kev (kevchoi) and others added 3 commits October 7, 2026 11:41
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@kevchoi
kev (kevchoi) force-pushed the kevchoi/20261007-speech-clarity branch from 1697885 to 58729a8 Compare October 7, 2026 15:42
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant