Skip to content
MeridFlow AiFlow v8.x • self-hosted

Evals

Eval tooling for prompts, agents, and orchestrators: build a regression suite of test scenarios once, then grade a candidate system prompt or instruction against it before promoting it into that subject's live base_system_prompt or instruction, all without touching the subject's actual configuration or taking a live call.

A dataset is a named collection of cases for one subject, either an agent or an orchestrator. Each case is a scenario (what a caller might say or ask) plus expected criteria (what a good response should do). A run takes a candidate prompt or instruction, generates a real response to every case's scenario with it, then grades each response pass or fail against that case's expected criteria using a second, separate LLM-as-judge call. For an orchestrator subject, "generates a real response" builds and drives a real ADK agent tree with the candidate instruction, tools and sub-orchestrators included, rather than a bare one-shot completion.

Endpoint family

An orchestrator subject uses the identical endpoint family, cases and runs included, with /api/v1/orchestrators/{orchestrator_id}/eval-datasets in place of /api/v1/agents/{agent_id}/eval-datasets throughout:

GET    /api/v1/agents/{agent_id}/eval-datasets
POST   /api/v1/agents/{agent_id}/eval-datasets
GET    /api/v1/agents/{agent_id}/eval-datasets/{dataset_id}
DELETE /api/v1/agents/{agent_id}/eval-datasets/{dataset_id}
GET    /api/v1/agents/{agent_id}/eval-datasets/{dataset_id}/cases
POST   /api/v1/agents/{agent_id}/eval-datasets/{dataset_id}/cases
DELETE /api/v1/agents/{agent_id}/eval-datasets/{dataset_id}/cases/{case_id}
GET    /api/v1/agents/{agent_id}/eval-datasets/{dataset_id}/runs
POST   /api/v1/agents/{agent_id}/eval-datasets/{dataset_id}/runs
GET    /api/v1/agents/{agent_id}/eval-datasets/{dataset_id}/runs/{run_id}

GET    /api/v1/orchestrators/{orchestrator_id}/eval-datasets
POST   /api/v1/orchestrators/{orchestrator_id}/eval-datasets
GET    /api/v1/orchestrators/{orchestrator_id}/eval-datasets/{dataset_id}
DELETE /api/v1/orchestrators/{orchestrator_id}/eval-datasets/{dataset_id}
GET    /api/v1/orchestrators/{orchestrator_id}/eval-datasets/{dataset_id}/cases
POST   /api/v1/orchestrators/{orchestrator_id}/eval-datasets/{dataset_id}/cases
DELETE /api/v1/orchestrators/{orchestrator_id}/eval-datasets/{dataset_id}/cases/{case_id}
GET    /api/v1/orchestrators/{orchestrator_id}/eval-datasets/{dataset_id}/runs
POST   /api/v1/orchestrators/{orchestrator_id}/eval-datasets/{dataset_id}/runs
GET    /api/v1/orchestrators/{orchestrator_id}/eval-datasets/{dataset_id}/runs/{run_id}

Authentication and roles

Every endpoint takes an admin session token (see Authentication). Reading (GET) is available to every role, including Viewer; every mutating action requires Owner, Admin, or Editor.

Run lifecycle

stateDiagram-v2
    [*] --> pending: create run
    pending --> running: worker claims it
    running --> completed: every case scored
    running --> failed: run-level error
    completed --> [*]
    failed --> [*]

Creating a run only enqueues it. A background worker (polling every few seconds, the same shape as the campaign pacing tick and retention worker) claims one pending run at a time and processes every case in it sequentially. Poll GET .../runs/{run_id} until status leaves pending or running. A failed run means the run itself errored, for example the Gemini API was unreachable, not that a case scored badly. A per-case failure instead sets that one result's error_message and leaves verdict empty, while the rest of the run continues.

Create a dataset

POST /api/v1/agents/{agent_id}/eval-datasets
Field Type Required Description
name string Yes 1-120 characters.
curl -X POST https://api.your-domain.com/api/v1/agents/1/eval-datasets \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $ADMIN_TOKEN" \
  -d '{"name": "Hours FAQ"}'
{
  "id": 1,
  "agent_id": 1,
  "orchestrator_id": null,
  "name": "Hours FAQ",
  "created_at": "2026-01-15T10:00:00Z"
}

agent_id and orchestrator_id are set by which endpoint family created the dataset, exactly one of them is ever non-null: an orchestrator-scoped dataset instead returns "agent_id": null, "orchestrator_id": 1.

Add a case

POST /api/v1/agents/{agent_id}/eval-datasets/{dataset_id}/cases
Field Type Required Description
scenario string Yes What the caller says or asks.
expected_criteria string Yes What a good response should do.
curl -X POST https://api.your-domain.com/api/v1/agents/1/eval-datasets/1/cases \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $ADMIN_TOKEN" \
  -d '{
    "scenario": "A customer asks what your business hours are.",
    "expected_criteria": "States specific opening and closing hours or days of the week."
  }'

Criteria no response can pass

A case runs against the subject's real tools, so criteria that depend on something the case cannot control end up grading themselves rather than the prompt. Four shapes to avoid:

  • Predicting what a lookup finds. "...and the lookup returns no such ticket" fails whenever the tool actually finds it, and punishes a response for honestly reporting the real result. Describe the request and let the criteria be about handling whatever comes back. The judge is instructed to void a criteria clause the real tool result contradicts, but such a case still tests nothing useful.
  • Expecting a literal the scenario never gave. "replies with the provided update text" cannot be met when no text appears in the scenario. Put the value in the scenario, or expect the agent to ask for it.
  • Demanding exact wording or ordering. "lists the six features in that order" fails a response covering all six that groups them sensibly. The judge grades on substance, so write criteria about coverage.
  • Describing two turns. A run generates one response with nobody to answer back, so "asks for their email address, saves it, and sends a recap" cannot be met: the saving needs a value the asking has not returned yet. Require the asking, or put the address in the scenario.

Run against a candidate prompt

POST /api/v1/agents/{agent_id}/eval-datasets/{dataset_id}/runs
Field Type Required Description
candidate_system_prompt string Yes The prompt or instruction to test, not saved to the subject.

For an orchestrator-scoped run, candidate_system_prompt holds the candidate instruction, the field name is unchanged since it's the same EvalRun schema either way.

Returns 422 if the dataset has no cases yet.

curl -X POST https://api.your-domain.com/api/v1/agents/1/eval-datasets/1/runs \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $ADMIN_TOKEN" \
  -d '{"candidate_system_prompt": "You are a helpful receptionist. We are open 9-5 Mon-Fri."}'
{
  "id": 1,
  "dataset_id": 1,
  "candidate_system_prompt": "You are a helpful receptionist. We are open 9-5 Mon-Fri.",
  "status": "pending",
  "error_message": null,
  "created_at": "2026-01-15T10:05:00Z"
}

Check a run's results

GET /api/v1/agents/{agent_id}/eval-datasets/{dataset_id}/runs/{run_id}
{
  "id": 1,
  "dataset_id": 1,
  "candidate_system_prompt": "You are a helpful receptionist. We are open 9-5 Mon-Fri.",
  "status": "completed",
  "error_message": null,
  "created_at": "2026-01-15T10:05:00Z",
  "results": [
    {
      "id": 1,
      "run_id": 1,
      "case_id": 1,
      "actual_response": "We're open Monday through Friday, 9am to 5pm.",
      "verdict": "pass",
      "judge_reasoning": "States the specific hours and days as required.",
      "error_message": null
    }
  ]
}

Next

  • Agents: where a winning candidate prompt gets promoted into base_system_prompt.
  • Orchestrators: where a winning candidate instruction gets promoted into instruction.
  • Writing effective instructions: what makes a good system prompt to begin with.