Evals¶
Eval tooling for prompts, agents, and orchestrators: build a regression
suite of test scenarios once, then grade a candidate system prompt or
instruction against it before promoting it into that subject's live
base_system_prompt or instruction,
all without touching the subject's actual configuration or taking a live
call.
A dataset is a named collection of cases for one subject, either
an agent or an orchestrator. Each case is a scenario (what a caller might
say or ask) plus expected criteria (what a good response should do). A
run takes a candidate prompt or instruction, generates a real
response to every case's scenario with it, then grades each response
pass or fail against that case's expected criteria using a second,
separate LLM-as-judge call. For an orchestrator subject, "generates a real
response" builds and drives a real ADK agent tree with the candidate
instruction, tools and sub-orchestrators included, rather than a bare
one-shot completion.
Endpoint family¶
An orchestrator subject uses the identical endpoint family, cases and
runs included, with /api/v1/orchestrators/{orchestrator_id}/eval-datasets
in place of /api/v1/agents/{agent_id}/eval-datasets throughout:
GET /api/v1/agents/{agent_id}/eval-datasets
POST /api/v1/agents/{agent_id}/eval-datasets
GET /api/v1/agents/{agent_id}/eval-datasets/{dataset_id}
DELETE /api/v1/agents/{agent_id}/eval-datasets/{dataset_id}
GET /api/v1/agents/{agent_id}/eval-datasets/{dataset_id}/cases
POST /api/v1/agents/{agent_id}/eval-datasets/{dataset_id}/cases
DELETE /api/v1/agents/{agent_id}/eval-datasets/{dataset_id}/cases/{case_id}
GET /api/v1/agents/{agent_id}/eval-datasets/{dataset_id}/runs
POST /api/v1/agents/{agent_id}/eval-datasets/{dataset_id}/runs
GET /api/v1/agents/{agent_id}/eval-datasets/{dataset_id}/runs/{run_id}
GET /api/v1/orchestrators/{orchestrator_id}/eval-datasets
POST /api/v1/orchestrators/{orchestrator_id}/eval-datasets
GET /api/v1/orchestrators/{orchestrator_id}/eval-datasets/{dataset_id}
DELETE /api/v1/orchestrators/{orchestrator_id}/eval-datasets/{dataset_id}
GET /api/v1/orchestrators/{orchestrator_id}/eval-datasets/{dataset_id}/cases
POST /api/v1/orchestrators/{orchestrator_id}/eval-datasets/{dataset_id}/cases
DELETE /api/v1/orchestrators/{orchestrator_id}/eval-datasets/{dataset_id}/cases/{case_id}
GET /api/v1/orchestrators/{orchestrator_id}/eval-datasets/{dataset_id}/runs
POST /api/v1/orchestrators/{orchestrator_id}/eval-datasets/{dataset_id}/runs
GET /api/v1/orchestrators/{orchestrator_id}/eval-datasets/{dataset_id}/runs/{run_id}
Authentication and roles¶
Every endpoint takes an admin session token (see
Authentication). Reading (GET)
is available to every role, including Viewer; every mutating action
requires Owner, Admin, or Editor.
Run lifecycle¶
stateDiagram-v2
[*] --> pending: create run
pending --> running: worker claims it
running --> completed: every case scored
running --> failed: run-level error
completed --> [*]
failed --> [*]
Creating a run only enqueues it. A background worker (polling every few
seconds, the same shape as the campaign pacing tick and
retention worker) claims one pending run at a time and processes every
case in it sequentially. Poll GET .../runs/{run_id} until status
leaves pending or running. A failed run means the run itself
errored, for example the Gemini API was unreachable, not that a case
scored badly. A per-case failure instead sets that one result's
error_message and leaves verdict empty, while the rest of the run
continues.
Create a dataset¶
| Field | Type | Required | Description |
|---|---|---|---|
name |
string | Yes | 1-120 characters. |
{
"id": 1,
"agent_id": 1,
"orchestrator_id": null,
"name": "Hours FAQ",
"created_at": "2026-01-15T10:00:00Z"
}
agent_id and orchestrator_id are set by which endpoint family created
the dataset, exactly one of them is ever non-null: an orchestrator-scoped
dataset instead returns "agent_id": null, "orchestrator_id": 1.
Add a case¶
| Field | Type | Required | Description |
|---|---|---|---|
scenario |
string | Yes | What the caller says or asks. |
expected_criteria |
string | Yes | What a good response should do. |
curl -X POST https://api.your-domain.com/api/v1/agents/1/eval-datasets/1/cases \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $ADMIN_TOKEN" \
-d '{
"scenario": "A customer asks what your business hours are.",
"expected_criteria": "States specific opening and closing hours or days of the week."
}'
Criteria no response can pass¶
A case runs against the subject's real tools, so criteria that depend on something the case cannot control end up grading themselves rather than the prompt. Four shapes to avoid:
- Predicting what a lookup finds. "...and the lookup returns no such ticket" fails whenever the tool actually finds it, and punishes a response for honestly reporting the real result. Describe the request and let the criteria be about handling whatever comes back. The judge is instructed to void a criteria clause the real tool result contradicts, but such a case still tests nothing useful.
- Expecting a literal the scenario never gave. "replies with the provided update text" cannot be met when no text appears in the scenario. Put the value in the scenario, or expect the agent to ask for it.
- Demanding exact wording or ordering. "lists the six features in that order" fails a response covering all six that groups them sensibly. The judge grades on substance, so write criteria about coverage.
- Describing two turns. A run generates one response with nobody to answer back, so "asks for their email address, saves it, and sends a recap" cannot be met: the saving needs a value the asking has not returned yet. Require the asking, or put the address in the scenario.
Run against a candidate prompt¶
| Field | Type | Required | Description |
|---|---|---|---|
candidate_system_prompt |
string | Yes | The prompt or instruction to test, not saved to the subject. |
For an orchestrator-scoped run, candidate_system_prompt holds the
candidate instruction, the field name is unchanged since it's the
same EvalRun schema either way.
Returns 422 if the dataset has no cases yet.
{
"id": 1,
"dataset_id": 1,
"candidate_system_prompt": "You are a helpful receptionist. We are open 9-5 Mon-Fri.",
"status": "pending",
"error_message": null,
"created_at": "2026-01-15T10:05:00Z"
}
Check a run's results¶
{
"id": 1,
"dataset_id": 1,
"candidate_system_prompt": "You are a helpful receptionist. We are open 9-5 Mon-Fri.",
"status": "completed",
"error_message": null,
"created_at": "2026-01-15T10:05:00Z",
"results": [
{
"id": 1,
"run_id": 1,
"case_id": 1,
"actual_response": "We're open Monday through Friday, 9am to 5pm.",
"verdict": "pass",
"judge_reasoning": "States the specific hours and days as required.",
"error_message": null
}
]
}
Next¶
- Agents: where a winning candidate prompt gets promoted
into
base_system_prompt. - Orchestrators: where a winning candidate instruction
gets promoted into
instruction. - Writing effective instructions: what makes a good system prompt to begin with.