Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 2 additions & 1 deletion AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -54,7 +54,7 @@ The library maintains parallel implementations in TypeScript (`js/`) and Python

### Key Modules (both languages)

- `llm.ts` / `llm.py` - LLM-as-a-judge scorers (Factuality, Battle, ClosedQA, Humor, Security, Sql, Summary, Translation, VoiceTaskSuccess)
- `llm.ts` / `llm.py` - LLM-as-a-judge scorers (Factuality, Battle, ClosedQA, Humor, Security, Sql, Summary, Translation, VoiceTaskSuccess, SpeechClarity)
- `ragas.ts` / `ragas.py` - RAG evaluation metrics (ContextRelevancy, Faithfulness, AnswerRelevancy, etc.)
- `string.ts` / `string.py` - Text similarity (Levenshtein, EmbeddingSimilarity)
- `json.ts` / `json.py` - JSON validation and diff
Expand All @@ -69,6 +69,7 @@ YAML templates in `templates/` define LLM classifier prompts. Templates use Must
- Prompt rendering with chain-of-thought (CoT) suffix
- Tool-based response parsing via `select_choice` function
- Score mapping from choice letters to numeric scores
- An optional `audio:` path (e.g. `input.audio`) to `{data, content_type}` bytes, sent as a file part; missing audio gives a `null` score

### Python Scorer Pattern

Expand Down
1 change: 1 addition & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -333,6 +333,7 @@ Eval(
- SQL
- Translation
- Voice task success
- Speech clarity
- Fine-tuned binary classifiers

### RAG evaluations
Expand Down
29 changes: 29 additions & 0 deletions SCORERS.md
Original file line number Diff line number Diff line change
Expand Up @@ -211,6 +211,35 @@ The judge sees the messages as JSON, including tool calls and results. With no c
- `0.5` = Completed with minor problems
- `0.0` = Not completed, or wrong or unsupported

### SpeechClarity

Evaluates how clearly the agent speaks, from a recording of the whole call.

**Parameters:**

- `input.audio` (object or list): The recording, as `{ data, content_type }`, where `data` is the bytes (`Uint8Array` in TypeScript, `bytes` in Python) and `content_type` is an `audio/*` type such as `audio/ogg`. A list of these, such as a recording split into chunks, is sent as separate files in list order.
- `model` (string, optional): Model to use (default: `gemini-3.8-flash`). It must accept audio as a Chat Completions `file` part.

Without `input.audio`, or with an empty list, the score is `null`.

**Score Range:** 0-1

- `1.0` = Clear and understandable without effort
- `0.5` = Audible flaws, but easy to understand
- `0.0` = Important words unclear or lost

**Example:**

```python
from pathlib import Path
from autoevals import SpeechClarity

audio = {"data": Path("call.ogg").read_bytes(), "content_type": "audio/ogg"}
result = SpeechClarity().eval(input={"audio": audio}, output=None)
```

To build your own audio judge, set `audio` to the recording's path in the scorer's arguments: `audio: input.audio` in a template, `audio="input.audio"` for `LLMClassifier`, or `audio: "input.audio"` for `LLMClassifierFromTemplate`.

---

## RAG (Retrieval-Augmented Generation) scorers
Expand Down
61 changes: 61 additions & 0 deletions js/llm.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -9,6 +9,7 @@ import {
DEFAULT_MODEL,
LLMClassifierFromTemplate,
OpenAIClassifier,
SpeechClarity,
templateUsesThreadVariables,
VoiceTaskSuccess,
} from "../js/llm";
Expand Down Expand Up @@ -724,3 +725,63 @@ describe("VoiceTaskSuccess", () => {
expect(score.score).toBeNull();
});
});

describe("SpeechClarity", () => {
test("sends each audio as a file to gemini-3.8-flash through chat completions", async () => {
let body: any;
server.use(
http.post(
"https://api.openai.com/v1/chat/completions",
async ({ request }) => {
body = await request.json();
return HttpResponse.json({
choices: [
{
message: {
role: "assistant",
tool_calls: [
{
id: "call_test",
type: "function",
function: {
name: "select_choice",
arguments: '{"reasons":"Clear.","choice":"A"}',
},
},
],
},
},
],
});
},
),
);

const score = await SpeechClarity({
input: {
audio: [
{ data: new Uint8Array([1, 2, 3]), content_type: "audio/ogg" },
{ data: new Uint8Array([4, 5, 6]), content_type: "audio/ogg" },
],
},
output: undefined,
openAiApiKey: "test-api-key",
});

expect(score.score).toBe(1);
expect(body.model).toBe("gemini-3.8-flash");
expect(body.messages[0].content[0].type).toBe("text");
expect(
body.messages[0].content.slice(1).map((p: any) => p.file.file_data),
).toEqual(["data:audio/ogg;base64,AQID", "data:audio/ogg;base64,BAUG"]);
});

test("skips when there is no audio", async () => {
const score = await SpeechClarity({
input: { text: "hello" },
output: undefined,
openAiApiKey: "test-api-key",
});
expect(score.score).toBeNull();
});
});
73 changes: 71 additions & 2 deletions js/llm.ts
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,7 @@ import {
} from "./oai";
import { ModelGradedSpec, templates } from "./templates";
import {
ChatCompletionContentPart,
ChatCompletionMessage,
ChatCompletionMessageParam,
ChatCompletionTool,
Expand Down Expand Up @@ -62,6 +63,42 @@ function filterSystemMessagesFromThread(thread: unknown[]): unknown[] {
});
}

function getPath(args: unknown, path: string): unknown {
let value = args;
for (const key of path.split(".")) {
value =
value && typeof value === "object" && !Array.isArray(value)
? Reflect.get(value, key)
: undefined;
}
return value;
}

export type Audio = { data: Uint8Array; content_type: string };

function audioPart(audio: unknown): ChatCompletionContentPart.File {
const data = Reflect.get(Object(audio), "data");
if (!(data instanceof Uint8Array)) {
throw new TypeError(
"Audio must be an object with `data` bytes and a `content_type`",
);
}
const contentType = String(Reflect.get(Object(audio), "content_type") ?? "")
.split(";")[0]
.trim()
.toLowerCase();
if (!contentType.startsWith("audio/")) {
throw new Error(
`Audio must have an audio/* content type, got "${contentType}"`,
);
}
const base64 = Buffer.from(data).toString("base64");
return {
type: "file",
file: { file_data: `data:${contentType};base64,${base64}` },
};
}

const NO_COT_SUFFIX =
"Answer the question by calling `select_choice` with a single choice from {{__choices}}.";

Expand Down Expand Up @@ -142,6 +179,7 @@ export type OpenAIClassifierArgs<RenderArgs> = {
messages: ChatCompletionMessageParam[];
choiceScores: Record<string, number>;
classificationTools: ChatCompletionTool[];
audioFiles?: ChatCompletionContentPart.File[];
cache?: ChatCache;
} & LLMArgs &
RenderArgs;
Expand Down Expand Up @@ -174,6 +212,7 @@ export async function OpenAIClassifier<RenderArgs, Output>(
reasoningEnabled,
reasoningBudget,
useResponsesApi,
audioFiles,
cache,
...remainingRenderArgs
} = remaining;
Expand Down Expand Up @@ -212,6 +251,13 @@ export async function OpenAIClassifier<RenderArgs, Output>(
};

const messages = renderMessages(messagesArg, renderArgs);
if (audioFiles) {
const last = messages[messages.length - 1];
messages[messages.length - 1] = {
role: "user",
content: [{ type: "text", text: String(last.content) }, ...audioFiles],
};
}

const resp = await cachedChatCompletion(
{
Expand Down Expand Up @@ -306,6 +352,7 @@ export function LLMClassifierFromTemplate<RenderArgs>({
reasoningEnabled,
reasoningBudget,
useResponsesApi,
audio,
}: {
name: string;
promptTemplate: string;
Expand All @@ -318,11 +365,23 @@ export function LLMClassifierFromTemplate<RenderArgs>({
reasoningEnabled?: boolean;
reasoningBudget?: number;
useResponsesApi?: boolean;
audio?: string;
}): Scorer<string, LLMClassifierArgs<RenderArgs>> {
const choiceStrings = Object.keys(choiceScores);
const ret = async (
runtimeArgs: ScorerArgs<string, LLMClassifierArgs<RenderArgs>>,
) => {
const audioValue = audio ? getPath(runtimeArgs, audio) : undefined;
const audioList =
audioValue == null
? []
: Array.isArray(audioValue)
? audioValue
: [audioValue];
if (audio && audioList.length === 0) {
return { name, score: null };
}

const useCoT = runtimeArgs.useCoT ?? useCoTArg ?? true;
// Use runtime model > template model > configured default model
const model = runtimeArgs.model ?? modelArg ?? getDefaultModel();
Expand Down Expand Up @@ -374,6 +433,7 @@ export function LLMClassifierFromTemplate<RenderArgs>({
// Since the logic is a bit funky for computing this, include
// it at the end to prevent overrides
useCoT,
audioFiles: audio ? audioList.map(audioPart) : undefined,
};

return await OpenAIClassifier(classifierArgs);
Expand All @@ -398,6 +458,7 @@ export function LLMClassifierFromSpec<RenderArgs>(
useCoT: spec.use_cot,
temperature: spec.temperature,
maxTokens: spec.max_tokens,
audio: spec.audio,
});
}

Expand All @@ -409,10 +470,10 @@ export function LLMClassifierFromSpecFile<RenderArgs>(
return LLMClassifierFromSpec(name, doc);
}

function buildLLMClassifier<RenderArgs>(
function buildLLMClassifier<RenderArgs, Output = string>(
name: string,
templateName: keyof typeof templates,
): ScorerWithPartial<string, LLMClassifierArgs<RenderArgs>> {
): ScorerWithPartial<Output, LLMClassifierArgs<RenderArgs>> {
if (!(templateName in templates)) {
throw new Error(`Model template ${name} not found`);
}
Expand Down Expand Up @@ -521,3 +582,11 @@ export const VoiceTaskSuccess = makePartial<
thread_with_system: JSON.stringify(messages),
});
}, "VoiceTaskSuccess");

/**
* Test how clearly an agent speaks, from a recording of the whole conversation (`input.audio`), given as one audio or a list of chunks in recording order.
*/
export const SpeechClarity = buildLLMClassifier<
{ input: { audio?: Audio | Audio[]; [key: string]: unknown } },
unknown
>("SpeechClarity", "speech_clarity");
7 changes: 7 additions & 0 deletions js/manifest.ts
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,7 @@ import {
Humor,
Possible,
Security,
SpeechClarity,
Sql,
Summary,
Translation,
Expand Down Expand Up @@ -84,6 +85,12 @@ export const Evaluators: {
description: "Test whether an output is malicious.",
template: templates.security,
},
{
method: SpeechClarity,
description:
"Test how clearly an agent's speech can be understood, from a recording of the whole conversation (`input.audio`).",
template: templates.speech_clarity,
},
{
method: Sql,
description:
Expand Down
3 changes: 3 additions & 0 deletions js/templates.ts
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,7 @@ import factuality from "../templates/factuality.yaml";
import humor from "../templates/humor.yaml";
import possible from "../templates/possible.yaml";
import security from "../templates/security.yaml";
import speech_clarity from "../templates/speech_clarity.yaml";
import sql from "../templates/sql.yaml";
import summary from "../templates/summary.yaml";
import translation from "../templates/translation.yaml";
Expand All @@ -19,6 +20,7 @@ export const modelGradedSpecSchema = z.object({
use_cot: z.boolean().optional(),
temperature: z.number().optional(),
max_tokens: z.number().optional(),
audio: z.string().optional(),
});

export type ModelGradedSpec = z.infer<typeof modelGradedSpecSchema>;
Expand All @@ -30,6 +32,7 @@ const templateStrings = {
humor,
possible,
security,
speech_clarity,
sql,
summary,
translation,
Expand Down
Loading
Loading