Repository navigation
feat: add speech clarity audio scorer - #239
Draft
kev (kevchoi) wants to merge 4 commits into
Draft
kev (kevchoi) wants to merge 4 commits into
kev (kevchoi) wants to merge 4 commits into
Conversation
Braintrust eval report
|
kev (kevchoi)
changed the base branch from
main
to
kevchoi/20261002-voice-timing-scorers
October 7, 2026 14:35
kev (kevchoi)
force-pushed
the
kevchoi/20261007-speech-clarity
branch
from
October 7, 2026 14:35
07fbf1c to
8c665cc
Compare
kev (kevchoi)
added this pull request to stack #240
October 7, 2026 14:35
kev (kevchoi)
removed this pull request from stack #240
October 7, 2026 14:50
kev (kevchoi)
added this pull request to stack #242
October 7, 2026 14:50
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
kev (kevchoi)
force-pushed
the
kevchoi/20261007-speech-clarity
branch
from
October 7, 2026 15:42
1697885 to
58729a8
Compare
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds
SpeechClarity, an LLM judge that listens to a voice call recording and rates how easily the agent's speech can be understood: garbled or cut-off words, distortion, dropouts, echo and noise. Scores are A = 1, B = 0.5 and C = 0. This PR carries over only the speech clarity part of #228. TurnTaking, VoiceTaskSuccess (#241) and the thread changes stay out.templates/speech_clarity.yaml: the prompt, withmodel: gemini-3.8-flashand a newaudio: input.audiofield.audiopath (JSLLMClassifierFromTemplate, PythonLLMClassifier). The value there must be{data: bytes, content_type: "audio/..."}or a list of them, such as a recording split into chunks. Each is sent unchanged as its own base64filepart after the prompt text, in list order; nothing is joined. If the value is missing or the list is empty, the score is skipped (null). Chunk boundaries aren't marked in the prompt, so a word split across two chunks may be judged as cut off.js/manifest.ts: listsSpeechClarityas a built-in under "LLM-as-a-Judge". Built-ins get audio as an attachment reference, so it needs the API server to pass the bytes (braintrustdata/braintrust#21432, in the matching braintrust stack). Merge the stacks together.SCORERS.mddocumentsSpeechClarityand theaudio:field;README.mdandAGENTS.mdlist it.Only Gemini reads audio this way. Tests mock the API, so the scorer hasn't been run against a real model or calibrated against human ratings.
Testing
pnpm run test -- js/llm.test.ts: passes.uv run --extra dev --extra scipy pytest py/autoevals/test_llm.py -k speech: passes. Five unrelated tests in that file need an OpenAI key and fail without one.tsc --noEmitreports one error injs/render-messages.test.ts. This PR doesn't touch that file.🤖 Generated with Claude Code