Skip to content

Evaluations

Evaluations answer the question every prompt change raises: did that help, and what did it break? You write cases, run them against an agent, and every answer gets a verdict from one or more scorers, each with the reason it passed or failed.

Open Monitor → Evaluations.

Term Meaning
Dataset A named set of cases, the scorers that judge them, and a rule for what tools may do during a run.
Case One question for the agent, plus what a good outcome looks like.
Scorer A check applied to each answer. Free scorers are deterministic; Costs tokens scorers ask a model to judge.
Run One execution of a dataset against one agent.
Pass rate The share of scored cases that passed. A case passes only when every scorer that ran on it passed.
  1. Select New dataset. Give it a Name and a Description of what it protects against.

  2. Under Tools during a run, choose what the agent’s tools may do. An evaluation runs the agent for real, so this matters.

    Option What happens
    Mock (recommended) Tools return the stub declared on each case. Nothing leaves the platform, so a fifty-case run sends no mail and files no tickets. Knowledge-base search still runs.
    Allowlist Only the tools you list under Tools allowed to run for real execute. Every other tool returns a refusal the agent can see.
    Real Every tool runs for real, with real consequences, once per case and repeat. Only sensible for read-only agents.
  3. Under Scorers, add the checks you want. Each has a weight (how much it counts toward the case’s score) and a Configuration (JSON) box. See Scorer reference.

  4. Leave Run judges even when a deterministic check already failed off unless you want the full picture on every case. By default, a case that already failed a free check skips the paid ones.

  5. Select Create dataset.

On the dataset page, select Add case.

Field What it does
Name A short label, for example “Refund window is 30 days”.
What the user asks The message sent to the agent.
A good answer (optional) A reference answer. Needed only by the correctness scorer; the wording doesn’t have to match.
Tools it should call (optional) Comma-separated tool names. Catches an agent that answered from memory instead of looking anything up. Used by tool_trajectory.
Tags (optional) Comma-separated labels such as policy, billing. Results are broken down by tag, which shows which kind of question the agent is bad at.
Tool stubs (optional, JSON) Shown when tools are mocked. The answer each tool returns during the run, for example {"get_order": {"id": "ORD-42", "status": "shipped"}}. A tool without a stub tells the agent it was mocked.

Import from conversations turns the 25 most recent web chat answers that users marked with a thumbs-down into cases. Each imported case has the user’s question and no reference answer: it is a question to answer better, not a wrong answer to reproduce. Importing twice does not create duplicates.

Only web chat feedback can be imported.

  1. On the dataset page, under Run this dataset, choose the Agent.

  2. Set Repeats (1 to 20). Running each case several times shows whether a result is stable or luck. Repeats multiply the cost.

  3. Select Run. The run is queued and its page opens; progress updates as cases finish.

A run can hold up to 2,000 executions (cases × repeats). It is refused before it starts if the workspace’s budget for the agent is exhausted. Select Cancel run to stop a run; cases already running finish.

To evaluate a draft before publishing it, use Evaluate this draft against on the agent’s form. See Drafts and publishing.

The run page answers four questions.

How did it do? The Pass rate, the number of cases passed, failed and errored, the total cost and the average latency. A run started from the dataset page is compared with the dataset’s previous completed run, and vs baseline shows the change in points.

What changed? Against the baseline, What changed lists the cases that regressed and the ones that were fixed. An aggregate pass rate can hide three fixes and three new failures; this does not. If the agent itself was edited between the two runs, a warning says the difference can’t be attributed to the change under test alone.

Where is the problem?

  • Case × scorer: a grid of every verdict. A whole column in red points to a miscalibrated scorer; a whole row, to a broken case.
  • By scorer and By tag: pass rates per check and per kind of question.
  • Latency: the distribution of response times.
  • Stability: with repeats, the cases that passed on some repeats and not others.

Why did this case fail? Open a case from the Cases list (filter by all, failed or passed) to see each scorer’s verdict and reason, the tools called, the agent’s answer, and its worklog.

The Evaluations overview shows the workspace’s pass rate, runs, cases evaluated and evaluation spend over the last 7, 30 or 90 days, a Quality over time chart per agent, and Open regressions: agents whose newest run scored worse than the one before. Each dataset page adds its own History, Cost per run and Duration per run.

Free scorers run first. Scorers that cost tokens use the judge model and pass when their score is at least threshold (default 0.7).

Scorer Kind Checks Configuration
contains Free Expected strings appear in the answer. values (list), match (all or any, default all), case_sensitive (default false)
not_contains Free No forbidden string appears, such as leaked instructions or banned phrasing. values, case_sensitive
regex Free The answer matches a pattern, such as an order number or a date format. pattern, ignore_case (default true), should_match (default true)
json_schema Free The answer is valid JSON and, if given, matches a schema. schema
tool_trajectory Free The agent called the expected tools. expected (defaults to the case’s tools), mode: subset (default, extras allowed), exact, or ordered
latency_budget Free The answer arrived in time. max_ms
cost_budget Free The answer cost less than a limit. max_usd
no_error Free The turn finished, no tool call failed, and the answer isn’t empty. None
correctness Costs tokens The answer agrees with the case’s reference answer. threshold
groundedness Costs tokens Every claim is supported by what the agent retrieved from its knowledge bases. threshold
rubric Costs tokens A criterion you write in plain language. criterion, threshold
tone_policy Costs tokens The answer stays within a persona and policy. policy (defaults to the agent’s system prompt), threshold

A scorer with nothing to check passes and says so: correctness on a case with no reference answer, groundedness when nothing was retrieved.

Scorers that cost tokens need a judge. Set EVAL_JUDGE_PROVIDER (for example gemini, openai or anthropic) and EVAL_JUDGE_MODEL in Workspace settings.

Without them, those scorers report the case as unscored rather than failed. The same happens if the judge call itself fails, so an outage never turns a good answer red. Judge tokens are billed to the workspace and counted in Evaluation spend.

Datasets tell you whether an agent still passes the cases someone thought to write. Online sampling scores a share of real web chat answers as they happen, using only scorers that need no reference answer.

Configure it in Workspace settings:

Setting Default What it does
EVAL_ONLINE_ENABLED false Turns sampling on.
EVAL_ONLINE_SAMPLE_PERCENT 5 Share of recent conversations scored, 0–100.
EVAL_ONLINE_DAILY_CAP 50 Maximum conversations scored per day, however much traffic grows.
EVAL_ONLINE_SCORERS groundedness Comma-separated scorers. Use groundedness, tone_policy and no_error; budgets take a limit, as in latency_budget:8000 or cost_budget:0.01.

Sampling reads the stored answer and the sources it used; it never runs the agent again. Results appear as the Real traffic figure on the overview, as a dashed line on Quality over time, and under an automatically created Production sampling dataset. They are kept apart from dataset results, so real traffic never inflates or hides a dataset’s pass rate.

The Agent Builder can create a dataset, run it, read the failures, try a revised system prompt without saving it, and apply the change only if the score improves, all within one conversation.

Permission Allows
evaluations.view See datasets, runs and results.
evaluations.manage Create and edit datasets and cases, and import cases.
evaluations.run Start and cancel runs, including evaluating a draft.

See Users and roles.