m4Mindset docs

Docs / AMS / Run / Limits and run behavior

View as Markdown

Limits and run behavior

Size a job before you build it, and know what happens when a run hits a limit, waits for a person, or stops part way.

After this page you can size a job against the limits a run works inside, and you know what happens when a run reaches one. The limits serve the cost and security sides of Optimize: they cap how much a mistake can spend and how far it can reach.

The limits

These are the defaults. Mindset sets them for the whole deployment. You can't change them per org or per agent.

An agent turn

A turn is one go at the work: the agent is asked, it calls what it needs, and it answers.

LimitDefaultWhat it bounds
Time, when the caller waits for one reply90 secondsAn API call that holds open until the answer is ready
Time, for streamed or delegated work240 secondsA turn streamed back as it runs, or one another agent asked for
Model roundsAbout 99How many times the agent can think and call tools in one turn
Tokens400,000What one turn can spend on the model
Repeated callsStops at 5The same call with the same input and nothing new learned. The agent gets a warning at 3

A scheduled run

A trigger on a schedule fires with nobody watching, so it gets its own bounds over the whole firing.

LimitDefault
Time for the whole firing15 minutes
Turns12, each up to 240 seconds
Tokens for the whole firing1.2 million

A function run

A function run has its own clock and its own budget of outside calls.

LimitDefaultWhat it bounds
Time, when an agent or script calls it5 minutesThe whole function run
Calls to outside systems1,000 per runEvery operation or model call the function makes. Each item in a list counts once
Items in a list a step repeats over1,000A longer list is refused before any call goes out
Items in flight at once8How many of those calls run at the same time

Delegation and MCP

LimitDefault
Calls from one run to other agents8, shared across the whole chain
A call to a tool on an MCP server120 seconds

When an agent hands work to another agent and the answer takes longer than 30 seconds, the caller gets a task to check back on and the delegated work keeps going to its own 240 second limit.

Size the invoice agent against them

Count what one invoice costs. The invoice exceptions agent works through four phases:

PhaseWhat it doesRough cost
1. GatherFetches the invoice and the purchase order2 operation calls
2. CompareCalls the invoice comparison function1 function run
3. AssessSearches the supplier contracts knowledge base2 or 3 searches
4. PrepareRegisters the supplier query1 write operation

One invoice fits easily inside one turn's 99 rounds and well inside 240 seconds.

Now try 4,000 invoices in one run. Time is the limit you hit first. Four phases of model work per invoice, thousands of times over, won't finish in 15 minutes, and the function that repeats over a list stops at 1,000 items. No setting makes this fit, so change the unit of work.

Small work per item: repeat across a list. A function step can repeat once per item, eight at a time. A first pass that takes a list of invoice numbers and fetches each PO number is one call per item, so a list of a few hundred fits. An item that fails is recorded as failed in the results, and the rest carry on.

Bigger work per item: one run per item. Run the agent once per exception. Your finance system can call Mindset when an invoice fails its PO match. Four thousand exceptions become four thousand short runs. Each one has its own record, and each one can be re-run on its own.

If you batch, size the batch against whichever limit you hit first, and size it for a slow morning. If the job only fits when every system answers at once, it doesn't fit.

When a run reaches a limit

A turn that reaches a limit stops where it is and says which limit it reached: out of rounds, out of time, out of tokens, or repeating itself. It hasn't failed. The run keeps everything it did, and it can carry on:

  • In a chat, the person sends another message and the agent picks up from there.
  • In a run nobody is chatting in (a schedule, a webhook or an API call) that follows a script, Mindset gives the run another turn later, and it continues from the phase it was in.
  • A scheduled firing that runs out of turns, time or tokens is recorded as parked, with the reason.

A function run that reaches its time or call limit ends with an error naming the limit.

When a run waits for a person

A script phase can ask a person a question (for example with Slack buttons) and wait for the answer. While it waits, the run is parked. It holds no turn and spends nothing. When the answer comes, the run resumes from that phase. A run nobody is chatting in resumes on its own. A run a person started in chat resumes when that person next sends a message.

  • The author sets how long to wait, at least 5 minutes, and which phase to go to if nobody answers.
  • A run that is still parked after 31 days is closed as abandoned.

When Mindset itself has a problem

If the server driving a run stops, the run isn't lost. Its hold on the run expires after about 5 minutes and Mindset picks it up again, continuing from where it stopped. If driving it keeps failing, Mindset backs off and retries. After an hour, or 10 failures in a row, it closes the run as abandoned.

Retries inside a function

A function step can retry a failed call. Retries are off by default: a step makes one attempt unless you set more, up to 10. You choose which kinds of failure to retry, such as rate limits, timeouts and network failures. A request the other system rejects isn't worth retrying, because it will be rejected again. You can also give the step its own timeout per attempt.

Retries don't count against the 1,000 call budget, and the function's 5 minute limit doesn't stretch to fit them. When time runs out, attempts in flight stop.

Calling twice doesn't run twice

When your own software starts a conversation with an agent, it can pass an idempotency key. Sending the same key again returns the first conversation and starts nothing new, so a caller that isn't sure its request landed can safely send it again. Each slot on a schedule fires once, for the same reason.

Build so a re-run is safe

A write happens when the agent calls it, as long as the operation is enabled. So a run that stops part way may already have changed something.

  • Make steps repeatable where you can. Running one twice should have no extra effect.
  • Put what you can't undo last. The invoice agent's first three phases only read. Phase four is the only one that writes, so a run that stops in phase three has changed nothing.
  • Read the record before you re-run. Check what happened shows how far the run got and what it called.

You're done when

  • You can state one run's worst case: how long it takes and how much it calls.
  • Both sit well inside the limits, and you know which limit you'd hit first.
  • For each write your agent makes, you know whether doing it twice would matter, and the ones that would happen last.