Evaluations
Evaluations answer the question every prompt change raises: did that help, and what did it break? You write cases, run them against an agent, and every answer gets a verdict from one or more scorers, each with the reason it passed or failed.
Open Monitor → Evaluations.
Concepts
Section titled “Concepts”| Term | Meaning |
|---|---|
| Dataset | A named set of cases, the scorers that judge them, and a rule for what tools may do during a run. |
| Case | One question for the agent, plus what a good outcome looks like. |
| Scorer | A check applied to each answer. Free scorers are deterministic; Costs tokens scorers ask a model to judge. |
| Run | One execution of a dataset against one agent. |
| Pass rate | The share of scored cases that passed. A case passes only when every scorer that ran on it passed. |
Create a dataset
Section titled “Create a dataset”-
Select New dataset. Give it a Name and a Description of what it protects against.
-
Under Tools during a run, choose what the agent’s tools may do. An evaluation runs the agent for real, so this matters.
Option What happens Mock (recommended) Tools return the stub declared on each case. Nothing leaves the platform, so a fifty-case run sends no mail and files no tickets. Knowledge-base search still runs. Allowlist Only the tools you list under Tools allowed to run for real execute. Every other tool returns a refusal the agent can see. Real Every tool runs for real, with real consequences, once per case and repeat. Only sensible for read-only agents. -
Under Scorers, add the checks you want. Each has a weight (how much it counts toward the case’s score) and a Configuration (JSON) box. See Scorer reference.
-
Leave Run judges even when a deterministic check already failed off unless you want the full picture on every case. By default, a case that already failed a free check skips the paid ones.
-
Select Create dataset.
Add cases
Section titled “Add cases”On the dataset page, select Add case.
| Field | What it does |
|---|---|
| Name | A short label, for example “Refund window is 30 days”. |
| What the user asks | The message sent to the agent. |
| A good answer (optional) | A reference answer. Needed only by the correctness scorer; the wording doesn’t have to match. |
| Tools it should call (optional) | Comma-separated tool names. Catches an agent that answered from memory instead of looking anything up. Used by tool_trajectory. |
| Tags (optional) | Comma-separated labels such as policy, billing. Results are broken down by tag, which shows which kind of question the agent is bad at. |
| Tool stubs (optional, JSON) | Shown when tools are mocked. The answer each tool returns during the run, for example {"get_order": {"id": "ORD-42", "status": "shipped"}}. A tool without a stub tells the agent it was mocked. |
Import from conversations
Section titled “Import from conversations”Import from conversations turns the 25 most recent web chat answers that users marked with a thumbs-down into cases. Each imported case has the user’s question and no reference answer: it is a question to answer better, not a wrong answer to reproduce. Importing twice does not create duplicates.
Only web chat feedback can be imported.
Run a dataset
Section titled “Run a dataset”-
On the dataset page, under Run this dataset, choose the Agent.
-
Set Repeats (1 to 20). Running each case several times shows whether a result is stable or luck. Repeats multiply the cost.
-
Select Run. The run is queued and its page opens; progress updates as cases finish.
A run can hold up to 2,000 executions (cases × repeats). It is refused before it starts if the workspace’s budget for the agent is exhausted. Select Cancel run to stop a run; cases already running finish.
To evaluate a draft before publishing it, use Evaluate this draft against on the agent’s form. See Drafts and publishing.
Read a run
Section titled “Read a run”The run page answers four questions.
How did it do? The Pass rate, the number of cases passed, failed and errored, the total cost and the average latency. A run started from the dataset page is compared with the dataset’s previous completed run, and vs baseline shows the change in points.
What changed? Against the baseline, What changed lists the cases that regressed and the ones that were fixed. An aggregate pass rate can hide three fixes and three new failures; this does not. If the agent itself was edited between the two runs, a warning says the difference can’t be attributed to the change under test alone.
Where is the problem?
- Case × scorer: a grid of every verdict. A whole column in red points to a miscalibrated scorer; a whole row, to a broken case.
- By scorer and By tag: pass rates per check and per kind of question.
- Latency: the distribution of response times.
- Stability: with repeats, the cases that passed on some repeats and not others.
Why did this case fail? Open a case from the Cases list (filter by all, failed or passed) to see each scorer’s verdict and reason, the tools called, the agent’s answer, and its worklog.
The Evaluations overview shows the workspace’s pass rate, runs, cases evaluated and evaluation spend over the last 7, 30 or 90 days, a Quality over time chart per agent, and Open regressions: agents whose newest run scored worse than the one before. Each dataset page adds its own History, Cost per run and Duration per run.
Scorer reference
Section titled “Scorer reference”Free scorers run first. Scorers that cost tokens use the judge model and pass when their score is at least threshold (default 0.7).
| Scorer | Kind | Checks | Configuration |
|---|---|---|---|
contains |
Free | Expected strings appear in the answer. | values (list), match (all or any, default all), case_sensitive (default false) |
not_contains |
Free | No forbidden string appears, such as leaked instructions or banned phrasing. | values, case_sensitive |
regex |
Free | The answer matches a pattern, such as an order number or a date format. | pattern, ignore_case (default true), should_match (default true) |
json_schema |
Free | The answer is valid JSON and, if given, matches a schema. | schema |
tool_trajectory |
Free | The agent called the expected tools. | expected (defaults to the case’s tools), mode: subset (default, extras allowed), exact, or ordered |
latency_budget |
Free | The answer arrived in time. | max_ms |
cost_budget |
Free | The answer cost less than a limit. | max_usd |
no_error |
Free | The turn finished, no tool call failed, and the answer isn’t empty. | None |
correctness |
Costs tokens | The answer agrees with the case’s reference answer. | threshold |
groundedness |
Costs tokens | Every claim is supported by what the agent retrieved from its knowledge bases. | threshold |
rubric |
Costs tokens | A criterion you write in plain language. | criterion, threshold |
tone_policy |
Costs tokens | The answer stays within a persona and policy. | policy (defaults to the agent’s system prompt), threshold |
A scorer with nothing to check passes and says so: correctness on a case with no reference answer, groundedness when nothing was retrieved.
Set up the judge model
Section titled “Set up the judge model”Scorers that cost tokens need a judge. Set EVAL_JUDGE_PROVIDER (for example gemini, openai or anthropic) and EVAL_JUDGE_MODEL in Workspace settings.
Without them, those scorers report the case as unscored rather than failed. The same happens if the judge call itself fails, so an outage never turns a good answer red. Judge tokens are billed to the workspace and counted in Evaluation spend.
Score real traffic
Section titled “Score real traffic”Datasets tell you whether an agent still passes the cases someone thought to write. Online sampling scores a share of real web chat answers as they happen, using only scorers that need no reference answer.
Configure it in Workspace settings:
| Setting | Default | What it does |
|---|---|---|
EVAL_ONLINE_ENABLED |
false |
Turns sampling on. |
EVAL_ONLINE_SAMPLE_PERCENT |
5 |
Share of recent conversations scored, 0–100. |
EVAL_ONLINE_DAILY_CAP |
50 |
Maximum conversations scored per day, however much traffic grows. |
EVAL_ONLINE_SCORERS |
groundedness |
Comma-separated scorers. Use groundedness, tone_policy and no_error; budgets take a limit, as in latency_budget:8000 or cost_budget:0.01. |
Sampling reads the stored answer and the sources it used; it never runs the agent again. Results appear as the Real traffic figure on the overview, as a dashed line on Quality over time, and under an automatically created Production sampling dataset. They are kept apart from dataset results, so real traffic never inflates or hides a dataset’s pass rate.
Fix-and-evaluate with the Agent Builder
Section titled “Fix-and-evaluate with the Agent Builder”The Agent Builder can create a dataset, run it, read the failures, try a revised system prompt without saving it, and apply the change only if the score improves, all within one conversation.
Permissions
Section titled “Permissions”| Permission | Allows |
|---|---|
evaluations.view |
See datasets, runs and results. |
evaluations.manage |
Create and edit datasets and cases, and import cases. |
evaluations.run |
Start and cancel runs, including evaluating a draft. |
See Users and roles.
