Docs / AMS / Optimize / Test an agent's behavior
View as MarkdownTest an agent's behavior
Write acceptance tests for an agent, run them many times, and use the results to decide whether a version or a cheaper model is good enough.
After this page you can write acceptance tests for an agent, run them, and read the results well enough to decide whether a version is good enough to hand to colleagues. Testing is the outcome side of Optimize, and it feeds cost too: it's how you prove a cheaper model still does the job.
What this is
You can't test an agent by matching its answer to text you wrote in advance. Ask the invoice exceptions agent the same question twice and it words the answer differently both times, and both can be right. So you test what has to be true about how it behaved:
- Did it call the invoice comparison function?
- Did the query it registered name the invoice number?
- Did it state a price that wasn't in the finance system's response?
An acceptance test is a situation you put the agent in, plus one or more criteria that must hold. You find them on the Optimize page, on the Acceptance tests tab. Each agent also has an Acceptance tests tab that shows its tests read only, with a link to open them in Optimize. Optimize is for admins.
Testing doesn't block a release
The platform doesn't run your tests when you save or activate a version, and it doesn't stop a version going live because a test failed. Testing is how you decide whether a version is good enough. Run the tests before you activate, read the results, then decide.
What a test checks
Some criteria are checked by code, so a run either has the call or answer or it doesn't. Others are judged by a model reading the transcript. The list shows each criterion by its key.
| Criterion | What it checks | Checked by |
|---|---|---|
invokesTool | The agent called a given tool | Code |
doesNotInvokeTool | The agent didn't call a given tool | Code |
toolParametersContain | It called a tool with the right details | Code |
answerContains | The final answer contains some text | Code |
answerMatches | The final answer matches a pattern | Code |
answerIsJson | The final answer is valid JSON | Code |
respondsWithin | It answered within a time | Code |
doesNotHallucinate | It stated nothing that wasn't in what came back | Model |
staysInCharacter | It behaved as the agent you configured | Model |
responseGroundedInKnowledge | Its answer rests on your material | Model |
satisfiesAssertion | A requirement you state in your own words | Model |
Every result records a reason, for code checks and for the judge, so you can read why something passed or failed.
Five tests for the invoice agent
| The test | Criterion |
|---|---|
| Given an invoice exception, the agent calls the invoice comparison function | invokesTool |
| The supplier query names the invoice number and the purchase order number | toolParametersContain |
| The assessment names a specific clause of the supplier contract | satisfiesAssertion |
| The agent never states a line, quantity or price that wasn't in the finance system's response | doesNotHallucinate |
| Given an invoice that matches its purchase order, the agent registers no supplier query | doesNotInvokeTool |
The last two are about what must not happen. Those catch the failures you didn't think of, and every agent should have at least one.
Results are statistical
An agent doesn't answer the same way every time, so a test runs many times and passes against a threshold.
- A code check runs once by default and must pass.
- A model judged criterion runs 5 times by default. It passes when the lower bound of a 95% confidence interval on its pass rate clears 0.5. At 5 runs, that means 5 out of 5.
- You can raise the number of runs (up to 200) or the threshold for any test. A setting no number of runs could ever meet is refused when you write the test.
- A test passes only if every one of its criteria passes.
Five out of five isn't certainty, and the result says so: each test's trend over time plots that lower bound with its number of runs, so a small sample doesn't read as 100%.
How to do it
- Open Optimize, go to the Acceptance tests tab, and pick the agent.
- Describe the test to the Optimize agent docked on the left: the situation, and what must be true. It writes a Draft.
- Dry-run the draft. A dry run never counts toward the agent's health, so you can check the test means what you meant. Most need rewording at this point.
- Accept it. From now on it's part of the agent's set and can't be changed in place.
- Press Run all to run the agent's whole set.
The tests run off the live path, so nobody using the agent sees them. They do run the agent's real tools, writes included. A test whose agent can register a supplier query registers one, every run.
Accepted tests don't change
An accepted test can't be edited or deleted.
- To change one, Amend it. That accepts a new test with your change and retires the old one, with a reason naming its replacement.
- To stand one down, Retire it. A reason is required and kept.
- Only a draft can be Discarded.
This stops the bar drifting. Somebody writes "the query must name a clause" in January, and in March there's a release everybody wants. The bar can still move, but only where people can see it.
Test a conversation, not one answer
The first answer is usually fine. What goes wrong is the fifth turn. A test can use a simulated user: a second model given a persona and a goal that replies to your agent turn after turn, up to a turn limit. The whole conversation is scored.
- Persona: a supplier contact who believes the invoice is correct and is annoyed at being queried.
- Goal: get the query dropped without providing a delivery note.
- What's judged: across the whole conversation, the agent keeps citing the clause and agrees to nothing it has no authority to agree to.
Compare a change before you make it
The Experiments tab runs your accepted tests against a changed setup, side by side with the current one. Nothing here touches the live agent.
- Under What do you want to change?, pick The model it runs on or Part of its system prompt.
- Set Runs per test, per setup (10 by default).
- Press What will this cost? to see the model calls and spend first.
- Press Run comparison.
Each criterion comes back as Improved, Regressed or No detectable difference, with cost and speed for each setup. Expect No detectable difference often at small run counts. The result tells you how many runs it would take to separate the two.
Experiments are offline. There's no way to split live traffic between two versions.
Runs with anything swapped (a model, a prompt, a dry run) are marked Exploratory with the reason, and they stay out of the health chart.
When it doesn't work
- A test passes when it obviously shouldn't. It's too loose. "The agent responds helpfully" passes on almost anything.
- A test fails and you disagree. Read the recorded reason. Usually the criterion says something slightly different from what you meant. Amend it.
- The set takes a long time. It runs for real, tools included, several times per test. That's also why tests cost money. See See and control what it costs.
You're done when
- The agent has accepted tests for what it must do, and at least one for what it must not do.
- You ran them before your last activation and can say which passed.
- Anything conversational has a test with a simulated user.