m4Mindset docs

Docs / AMS / Optimize / See and control what it costs

View as Markdown

See and control what it costs

See what your agents spend by agent, function, model, provider and user, and bring it down without making the agent worse.

After this page you can say what your agents cost and why, and you know how to make them cost less without making them worse. This is the cost side of Optimize.

What this is

The Cost tab of Monitor, the default tab. Every model call is recorded with the tokens the provider reported, priced at that model's rates when it's recorded. Six figures across the top:

FigureWhat it tells you
Total spendWhat the window you're looking at cost
Metered callsHow many model calls were made
Avg cost / callTotal divided by calls. The number to watch after a change
Avg input tokens / callHow much is sent each time. Usually the number that's too high
Tokens inWith the share that came from cache
Tokens outHow much the model generated

Pick the Time window (last 7, 30 or 90 days, or 12 months). Spend over time charts it by day, week or month.

Break it down

Switch between the breakdowns. Each answers a different question.

BreakdownThe question it answers
By agentWhich piece of work costs the most. Start here
By functionWhether a function makes more model calls than you expected
By modelWhat you pay for capability you may not need
By providerThe split across model providers
By userWho's using it. One person generating most of the spend is a conversation about how they use it
By interactionConversations, triggered runs and test runs, so you can see what your own testing costs

For the invoice exceptions agent, by agent gives you its total. By function might show the comparison function's model step running far more than once per invoice. By model might show the assess phase's judged gate running on your largest model.

The cache figure

Model providers charge less for input they've seen recently. For an agent with a long system prompt, that's most of what it sends. If Tokens in shows a low share from cache for an agent that runs constantly, something changes at the start of every request and nothing can be reused.

The main lever: prove a cheaper model still passes

The largest saving in most orgs, and you can measure it.

  1. Open Optimize → Experiments.
  2. Under What do you want to change?, pick The model it runs on and choose a cheaper model under Compare against.
  3. Press What will this cost?, then Run comparison.
  4. Read the result per criterion: Improved, Regressed or No detectable difference, with cost and speed for each setup.

If the invoice agent shows no regression on a smaller model, the smaller model is the better choice, and you have the run to show anyone who asks. The experiment changes nothing. To switch, change the model on the agent, save a version and activate it. See Test an agent's behavior.

Bringing it down

In the order that usually pays:

  • Use a cheaper model, proved as above.
  • Turn fixed work into a function. Anything with one right answer costs less as steps than as reasoning, and it stops varying. Comparing invoice lines to PO lines is arithmetic. See Build a function.
  • Use machine checks instead of judged gates. A machine check tests what was collected and costs nothing. A judged gate is another model call every time the phase is checked. Keep judged gates where the condition is about quality.
  • Return less. An operation that returns a whole invoice record when the function needs four fields sends the rest into the model, and it shows in Avg input tokens / call.
  • Schedule for the rate the data changes. An agent that runs hourly against data that updates daily costs 24 times what it needs to.

What caps spend today

There's no spend cap you can set in AMS today. A model connection's settings have no budget field that Mindset enforces. What bounds spend is the run limits, which you can't change:

BoundDefault
Tokens in one agent turn400,000
Tokens in one scheduled firing1.2 million
Output tokens per model callThe Size cap (max output tokens per call) you set on the model connection

A turn that reaches its token limit stops and says so. See Limits and run behavior. Before you leave a scheduled agent alone for a month, check its cost here after the first few days.

Things to be aware of

  • Cost is per environment. A busy test environment is real spend.
  • Tests cost money. They run the agent for real, several times per test, and the judge and the simulated user are model calls too.
  • Read the user breakdown before the model breakdown. Usage patterns explain more of the total than model choice does.
  • Mindset's own builder agents, such as the Agent Builder, are recorded against Mindset when they run on Mindset's model key. If your org has them use your own key, they're recorded against you.
  • The figure is priced from token counts at list rates. Your provider's invoice is the final word.

You're done when

  • You can name your three most expensive agents and say why each is where it is.
  • You've run one model comparison and either changed the model or can say why not.
  • You know what each scheduled agent costs per week.