Proving it works
Grade a change before a caller ever hears it.
Build a dataset of real scenarios and the criteria a good answer should meet, then run a candidate prompt against it and watch a live call while it happens.
What the capability actually does.
Test cases from real questions
A dataset holds the scenarios worth checking and the criteria a good response should meet, built from questions people actually ask rather than ones nobody asks.
Run a candidate prompt before it ships
Point a proposed system prompt at the dataset and run it. The result is graded against your own criteria, not a subjective read of one transcript, and the candidate runs on the same model class as a real task rather than a cheaper stand-in.
Orchestrators get the same treatment
The same dataset, run, and comparison flow applies to an Orchestrator, where the candidate is executed as a real agent tree rather than a single generation.
Watch a call while it happens
Live call monitoring streams a conversation in progress rather than only its transcript afterward, for the period a new agent is still earning trust.
The screens this happens on.

A prompt change is a hypothesis
A dataset holds the scenarios worth checking and the criteria a good answer should meet. A candidate prompt runs against all of them and is graded against those criteria, on the same model class a real task would use, so the result reflects what would actually happen rather than a cheaper approximation of it.
- The same flow covers Orchestrators, where the candidate runs as a real agent tree
What it deliberately does not do.
- Evals and live call monitoring are off on Demo and Basic. The 14 day trial unlocks both temporarily; a permanent licence needs Pro or Enterprise.
- An eval grades what a candidate prompt actually produced against your dataset. It does not predict how a change will perform on a case nobody wrote.
- Grading is a model judging a model. It is a great deal better than reading one transcript and guessing, and it is not a proof.