Docs / AMS / Optimize / See and control what it costs
View as MarkdownSee and control what it costs
See what your agents spend by agent, function, model, provider and user, and bring it down without making the agent worse.
After this page you can say what your agents cost and why, and you know how to make them cost less without making them worse. This is the cost side of Optimize.
What this is
The Cost tab of Monitor, the default tab. Every model call is recorded with the tokens the provider reported, priced at that model's rates when it's recorded. Six figures across the top:
| Figure | What it tells you |
|---|---|
| Total spend | What the window you're looking at cost |
| Metered calls | How many model calls were made |
| Avg cost / call | Total divided by calls. The number to watch after a change |
| Avg input tokens / call | How much is sent each time. Usually the number that's too high |
| Tokens in | With the share that came from cache |
| Tokens out | How much the model generated |
Pick the Time window (last 7, 30 or 90 days, or 12 months). Spend over time charts it by day, week or month.
Break it down
Switch between the breakdowns. Each answers a different question.
| Breakdown | The question it answers |
|---|---|
| By agent | Which piece of work costs the most. Start here |
| By function | Whether a function makes more model calls than you expected |
| By model | What you pay for capability you may not need |
| By provider | The split across model providers |
| By user | Who's using it. One person generating most of the spend is a conversation about how they use it |
| By interaction | Conversations, triggered runs and test runs, so you can see what your own testing costs |
For the invoice exceptions agent, by agent gives you its total. By function might show the comparison function's model step running far more than once per invoice. By model might show the assess phase's judged gate running on your largest model.
The cache figure
Model providers charge less for input they've seen recently. For an agent with a long system prompt, that's most of what it sends. If Tokens in shows a low share from cache for an agent that runs constantly, something changes at the start of every request and nothing can be reused.
The main lever: prove a cheaper model still passes
The largest saving in most orgs, and you can measure it.
- Open Optimize → Experiments.
- Under What do you want to change?, pick The model it runs on and choose a cheaper model under Compare against.
- Press What will this cost?, then Run comparison.
- Read the result per criterion: Improved, Regressed or No detectable difference, with cost and speed for each setup.
If the invoice agent shows no regression on a smaller model, the smaller model is the better choice, and you have the run to show anyone who asks. The experiment changes nothing. To switch, change the model on the agent, save a version and activate it. See Test an agent's behavior.
Bringing it down
In the order that usually pays:
- Use a cheaper model, proved as above.
- Turn fixed work into a function. Anything with one right answer costs less as steps than as reasoning, and it stops varying. Comparing invoice lines to PO lines is arithmetic. See Build a function.
- Use machine checks instead of judged gates. A machine check tests what was collected and costs nothing. A judged gate is another model call every time the phase is checked. Keep judged gates where the condition is about quality.
- Return less. An operation that returns a whole invoice record when the function needs four fields sends the rest into the model, and it shows in Avg input tokens / call.
- Schedule for the rate the data changes. An agent that runs hourly against data that updates daily costs 24 times what it needs to.
What caps spend today
There's no spend cap you can set in AMS today. A model connection's settings have no budget field that Mindset enforces. What bounds spend is the run limits, which you can't change:
| Bound | Default |
|---|---|
| Tokens in one agent turn | 400,000 |
| Tokens in one scheduled firing | 1.2 million |
| Output tokens per model call | The Size cap (max output tokens per call) you set on the model connection |
A turn that reaches its token limit stops and says so. See Limits and run behavior. Before you leave a scheduled agent alone for a month, check its cost here after the first few days.
Things to be aware of
- Cost is per environment. A busy test environment is real spend.
- Tests cost money. They run the agent for real, several times per test, and the judge and the simulated user are model calls too.
- Read the user breakdown before the model breakdown. Usage patterns explain more of the total than model choice does.
- Mindset's own builder agents, such as the Agent Builder, are recorded against Mindset when they run on Mindset's model key. If your org has them use your own key, they're recorded against you.
- The figure is priced from token counts at list rates. Your provider's invoice is the final word.
You're done when
- You can name your three most expensive agents and say why each is where it is.
- You've run one model comparison and either changed the model or can say why not.
- You know what each scheduled agent costs per week.