# Test an agent's behavior

> Write acceptance tests for an agent, run them many times, and use the results to decide whether a version or a cheaper model is good enough.

After this page you can write acceptance tests for an agent, run them, and read the results well enough to decide whether a version is good enough to hand to colleagues. Testing is the outcome side of Optimize, and it feeds cost too: it's how you prove a cheaper model still does the job.

## What this is

You can't test an agent by matching its answer to text you wrote in advance. Ask the invoice exceptions agent the same question twice and it words the answer differently both times, and both can be right. So you test what has to be true about how it behaved:

- Did it call the invoice comparison function?
- Did the query it registered name the invoice number?
- Did it state a price that wasn't in the finance system's response?

An **acceptance test** is a situation you put the agent in, plus one or more criteria that must hold. You find them on the **Optimize** page, on the **Acceptance tests** tab. Each agent also has an **Acceptance tests** tab that shows its tests read only, with a link to open them in Optimize. Optimize is for admins.

## Testing doesn't block a release

The platform doesn't run your tests when you save or activate a version, and it doesn't stop a version going live because a test failed. Testing is how **you** decide whether a version is good enough. Run the tests before you activate, read the results, then decide.

## What a test checks

Some criteria are checked by code, so a run either has the call or answer or it doesn't. Others are judged by a model reading the transcript. The list shows each criterion by its key.

| Criterion | What it checks | Checked by |
| --- | --- | --- |
| `invokesTool` | The agent called a given tool | Code |
| `doesNotInvokeTool` | The agent didn't call a given tool | Code |
| `toolParametersContain` | It called a tool with the right details | Code |
| `answerContains` | The final answer contains some text | Code |
| `answerMatches` | The final answer matches a pattern | Code |
| `answerIsJson` | The final answer is valid JSON | Code |
| `respondsWithin` | It answered within a time | Code |
| `doesNotHallucinate` | It stated nothing that wasn't in what came back | Model |
| `staysInCharacter` | It behaved as the agent you configured | Model |
| `responseGroundedInKnowledge` | Its answer rests on your material | Model |
| `satisfiesAssertion` | A requirement you state in your own words | Model |

Every result records a reason, for code checks and for the judge, so you can read why something passed or failed.

## Five tests for the invoice agent

| The test | Criterion |
| --- | --- |
| Given an invoice exception, the agent calls the invoice comparison function | `invokesTool` |
| The supplier query names the invoice number and the purchase order number | `toolParametersContain` |
| The assessment names a specific clause of the supplier contract | `satisfiesAssertion` |
| The agent never states a line, quantity or price that wasn't in the finance system's response | `doesNotHallucinate` |
| Given an invoice that matches its purchase order, the agent registers no supplier query | `doesNotInvokeTool` |

The last two are about what must not happen. Those catch the failures you didn't think of, and every agent should have at least one.

## Results are statistical

An agent doesn't answer the same way every time, so a test runs many times and passes against a threshold.

- A code check runs once by default and must pass.
- A model judged criterion runs 5 times by default. It passes when the lower bound of a 95% confidence interval on its pass rate clears 0.5. At 5 runs, that means 5 out of 5.
- You can raise the number of runs (up to 200) or the threshold for any test. A setting no number of runs could ever meet is refused when you write the test.
- A test passes only if every one of its criteria passes.

Five out of five isn't certainty, and the result says so: each test's trend over time plots that lower bound with its number of runs, so a small sample doesn't read as 100%.

## How to do it

1. Open **Optimize**, go to the **Acceptance tests** tab, and pick the agent.
2. Describe the test to the **Optimize agent** docked on the left: the situation, and what must be true. It writes a **Draft**.
3. **Dry-run** the draft. A dry run never counts toward the agent's health, so you can check the test means what you meant. Most need rewording at this point.
4. **Accept** it. From now on it's part of the agent's set and can't be changed in place.
5. Press **Run all** to run the agent's whole set.

The tests run off the live path, so nobody using the agent sees them. They do run the agent's real tools, writes included. A test whose agent can register a supplier query registers one, every run.

## Accepted tests don't change

An accepted test can't be edited or deleted.

- To change one, **Amend** it. That accepts a new test with your change and retires the old one, with a reason naming its replacement.
- To stand one down, **Retire** it. A reason is required and kept.
- Only a draft can be **Discard**ed.

This stops the bar drifting. Somebody writes "the query must name a clause" in January, and in March there's a release everybody wants. The bar can still move, but only where people can see it.

## Test a conversation, not one answer

The first answer is usually fine. What goes wrong is the fifth turn. A test can use a **simulated user**: a second model given a persona and a goal that replies to your agent turn after turn, up to a turn limit. The whole conversation is scored.

- **Persona:** a supplier contact who believes the invoice is correct and is annoyed at being queried.
- **Goal:** get the query dropped without providing a delivery note.
- **What's judged:** across the whole conversation, the agent keeps citing the clause and agrees to nothing it has no authority to agree to.

## Compare a change before you make it

The **Experiments** tab runs your accepted tests against a changed setup, side by side with the current one. Nothing here touches the live agent.

1. Under **What do you want to change?**, pick **The model it runs on** or **Part of its system prompt**.
2. Set **Runs per test, per setup** (10 by default).
3. Press **What will this cost?** to see the model calls and spend first.
4. Press **Run comparison**.

Each criterion comes back as **Improved**, **Regressed** or **No detectable difference**, with cost and speed for each setup. Expect **No detectable difference** often at small run counts. The result tells you how many runs it would take to separate the two.

Experiments are offline. There's no way to split live traffic between two versions.

Runs with anything swapped (a model, a prompt, a dry run) are marked **Exploratory** with the reason, and they stay out of the health chart.

## When it doesn't work

- **A test passes when it obviously shouldn't.** It's too loose. "The agent responds helpfully" passes on almost anything.
- **A test fails and you disagree.** Read the recorded reason. Usually the criterion says something slightly different from what you meant. Amend it.
- **The set takes a long time.** It runs for real, tools included, several times per test. That's also why tests cost money. See [See and control what it costs](https://docs4.mindset.ai/docs/ams/see-and-control-what-it-costs).

## You're done when

- The agent has accepted tests for what it must do, and at least one for what it must not do.
- You ran them before your last activation and can say which passed.
- Anything conversational has a test with a simulated user.
