Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 6 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -334,6 +334,12 @@ Eval(
- Translation
- Fine-tuned binary classifiers

### Voice evaluations

- Speech clarity
- Turn-taking
- Voice task success

### RAG evaluations

- Context precision
Expand Down
51 changes: 51 additions & 0 deletions SCORERS.md
Original file line number Diff line number Diff line change
Expand Up @@ -194,6 +194,57 @@ Evaluates translation quality.

---

## Voice scorers

`SpeechClarity` and `TurnTaking` listen to a voice call recording, not a transcript. They default to `gpt-audio`, which accepts WAV and MP3 audio. `VoiceTaskSuccess` reads the call's transcript from the trace.

### SpeechClarity

Evaluates how easily a listener on a phone call can understand every word the agent says: garbled or cut-off words, distortion, dropouts, echo, and background noise.

**Parameters:**

- `audio` (object, required): The recording as OpenAI `input_audio`: base64 `data` and a `format` of `"wav"` or `"mp3"`
- `model` (string, optional): An audio-capable model to use

**Score Range:** 0-1

- `1.0` = Clear and understandable without effort
- `0.5` = Audible flaws, but easy to understand
- `0.0` = Important words are unclear or lost

### TurnTaking

Evaluates how the agent handles turn-taking: talking over the caller, cutting them off, or not stopping when interrupted. Response latency is ignored.

**Parameters:**

- `audio` (object, required): The recording as OpenAI `input_audio`: base64 `data` and a `format` of `"wav"` or `"mp3"`
- `model` (string, optional): An audio-capable model to use

**Score Range:** 0-1

- `1.0` = Smooth turn-taking
- `0.5` = Brief overlaps or a slow yield, but no caller words are lost
- `0.0` = Caller words are lost

### VoiceTaskSuccess

Evaluates whether the agent correctly completed the caller's request, from the trace's `{{thread_with_system}}`: its instructions, tool calls, tool results, and replies.

**Parameters:**

- `trace` (Trace, required): The voice call's trace
- `model` (string, optional): Model to use

**Score Range:** 0-1

- `1.0` = Completed correctly, and everything told to the caller is supported
- `0.5` = Completed, with minor problems that did not change the outcome
- `0.0` = Not completed, a wrong or unauthorized action, or false or unsupported claims

---

## RAG (Retrieval-Augmented Generation) scorers

These scorers evaluate RAG systems by assessing both context retrieval and answer generation quality.
Expand Down
58 changes: 58 additions & 0 deletions js/llm.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,7 @@ import {
buildClassificationTools,
LLMClassifierFromTemplate,
OpenAIClassifier,
SpeechClarity,
templateUsesThreadVariables,
} from "../js/llm";
import {
Expand Down Expand Up @@ -74,6 +75,63 @@ describe("LLM Tests", () => {
).toBe(true);
});

test("SpeechClarity sends the audio to an audio model", async () => {
let body: any;
server.use(
http.post(
"https://api.openai.com/v1/chat/completions",
async ({ request }) => {
body = await request.json();
return HttpResponse.json({
id: "chatcmpl-test",
object: "chat.completion",
created: 0,
model: body.model,
choices: [
{
index: 0,
finish_reason: "tool_calls",
message: {
role: "assistant",
content: null,
tool_calls: [
{
id: "call_test",
type: "function",
function: {
name: "select_choice",
arguments: JSON.stringify({
reasons: "Clear.",
choice: "A",
}),
},
},
],
},
},
],
});
},
),
);

const audio = { data: "UklGRg==", format: "wav" as const };
const score = await SpeechClarity({ output: "", audio });

expect(score.score).toBe(1);
expect(body.model).toBe("gpt-audio");
expect(body.messages.at(-1)).toEqual({
role: "user",
content: [{ type: "input_audio", input_audio: audio }],
});
});

test("SpeechClarity requires audio", async () => {
await expect(SpeechClarity({ output: "" } as any)).rejects.toThrow(
"SpeechClarity needs the call recording",
);
});

test("openai classifier should evaluate titles", async () => {
let callCount = -1;
server.use(
Expand Down
38 changes: 38 additions & 0 deletions js/llm.ts
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,7 @@ import {
} from "./oai";
import { ModelGradedSpec, templates } from "./templates";
import {
ChatCompletionContentPartInputAudio,
ChatCompletionMessage,
ChatCompletionMessageParam,
ChatCompletionTool,
Expand Down Expand Up @@ -143,6 +144,7 @@ export type OpenAIClassifierArgs<RenderArgs> = {
choiceScores: Record<string, number>;
classificationTools: ChatCompletionTool[];
cache?: ChatCache;
audio?: ChatCompletionContentPartInputAudio.InputAudio;
} & LLMArgs &
RenderArgs;

Expand Down Expand Up @@ -175,6 +177,7 @@ export async function OpenAIClassifier<RenderArgs, Output>(
reasoningBudget,
useResponsesApi,
cache,
audio,
...remainingRenderArgs
} = remaining;

Expand Down Expand Up @@ -212,6 +215,12 @@ export async function OpenAIClassifier<RenderArgs, Output>(
};

const messages = renderMessages(messagesArg, renderArgs);
if (audio) {
messages.push({
role: "user",
content: [{ type: "input_audio", input_audio: audio }],
});
}

const resp = await cachedChatCompletion(
{
Expand Down Expand Up @@ -306,6 +315,7 @@ export function LLMClassifierFromTemplate<RenderArgs>({
reasoningEnabled,
reasoningBudget,
useResponsesApi,
requiresAudio,
}: {
name: string;
promptTemplate: string;
Expand All @@ -318,6 +328,7 @@ export function LLMClassifierFromTemplate<RenderArgs>({
reasoningEnabled?: boolean;
reasoningBudget?: number;
useResponsesApi?: boolean;
requiresAudio?: boolean;
}): Scorer<string, LLMClassifierArgs<RenderArgs>> {
const choiceStrings = Object.keys(choiceScores);
const ret = async (
Expand Down Expand Up @@ -376,6 +387,10 @@ export function LLMClassifierFromTemplate<RenderArgs>({
useCoT,
};

if (requiresAudio && !classifierArgs.audio) {
throw new Error(`${name} needs the call recording as \`audio\``);
}

return await OpenAIClassifier(classifierArgs);
};
Object.defineProperty(ret, "name", {
Expand All @@ -398,6 +413,7 @@ export function LLMClassifierFromSpec<RenderArgs>(
useCoT: spec.use_cot,
temperature: spec.temperature,
maxTokens: spec.max_tokens,
requiresAudio: spec.requires_audio,
});
}

Expand Down Expand Up @@ -492,3 +508,25 @@ export const Translation = buildLLMClassifier<{
language: string;
input: string;
}>("Translation", "translation");

/**
* Test whether the agent in a voice call `audio` recording speaks clearly.
*/
export const SpeechClarity = buildLLMClassifier<{
audio: ChatCompletionContentPartInputAudio.InputAudio;
}>("SpeechClarity", "speech_clarity");

/**
* Test whether the agent in a voice call `audio` recording takes turns smoothly.
*/
export const TurnTaking = buildLLMClassifier<{
audio: ChatCompletionContentPartInputAudio.InputAudio;
}>("TurnTaking", "turn_taking");

/**
* Test whether a voice agent completed the caller's request, from the trace's thread.
*/
export const VoiceTaskSuccess = buildLLMClassifier<{}>(
"VoiceTaskSuccess",
"voice_task_success",
);
26 changes: 26 additions & 0 deletions js/manifest.ts
Original file line number Diff line number Diff line change
Expand Up @@ -6,9 +6,12 @@ import {
Humor,
Possible,
Security,
SpeechClarity,
Sql,
Summary,
Translation,
TurnTaking,
VoiceTaskSuccess,
} from "./llm";
import { NumericDiff } from "./number";
import { EmbeddingSimilarity, Levenshtein } from "./string";
Expand Down Expand Up @@ -103,6 +106,29 @@ export const Evaluators: {
},
],
},
{
label: "Voice",
methods: [
{
method: SpeechClarity,
description:
"Test whether the agent in a voice call recording speaks clearly.",
template: templates.speech_clarity,
},
{
method: TurnTaking,
description:
"Test whether the agent in a voice call recording takes turns smoothly.",
template: templates.turn_taking,
},
{
method: VoiceTaskSuccess,
description:
"Test whether a voice agent correctly completed the caller's request, from the trace's thread.",
template: templates.voice_task_success,
},
],
},
{
label: "RAG",
methods: [
Expand Down
7 changes: 7 additions & 0 deletions js/templates.ts
Original file line number Diff line number Diff line change
Expand Up @@ -7,9 +7,12 @@ import factuality from "../templates/factuality.yaml";
import humor from "../templates/humor.yaml";
import possible from "../templates/possible.yaml";
import security from "../templates/security.yaml";
import speech_clarity from "../templates/speech_clarity.yaml";
import sql from "../templates/sql.yaml";
import summary from "../templates/summary.yaml";
import translation from "../templates/translation.yaml";
import turn_taking from "../templates/turn_taking.yaml";
import voice_task_success from "../templates/voice_task_success.yaml";

export const modelGradedSpecSchema = z.object({
prompt: z.string(),
Expand All @@ -18,6 +21,7 @@ export const modelGradedSpecSchema = z.object({
use_cot: z.boolean().optional(),
temperature: z.number().optional(),
max_tokens: z.number().optional(),
requires_audio: z.boolean().optional(),
});

export type ModelGradedSpec = z.infer<typeof modelGradedSpecSchema>;
Expand All @@ -29,9 +33,12 @@ const templateStrings = {
humor,
possible,
security,
speech_clarity,
sql,
summary,
translation,
turn_taking,
voice_task_success,
} as const;

// eslint-disable-next-line @typescript-eslint/consistent-type-assertions
Expand Down
Loading
Loading