Docs / AMS / Run / Limits and run behavior
View as MarkdownLimits and run behavior
Size a job before you build it, and know what happens when a run hits a limit, waits for a person, or stops part way.
After this page you can size a job against the limits a run works inside, and you know what happens when a run reaches one. The limits serve the cost and security sides of Optimize: they cap how much a mistake can spend and how far it can reach.
The limits
These are the defaults. Mindset sets them for the whole deployment. You can't change them per org or per agent.
An agent turn
A turn is one go at the work: the agent is asked, it calls what it needs, and it answers.
| Limit | Default | What it bounds |
|---|---|---|
| Time, when the caller waits for one reply | 90 seconds | An API call that holds open until the answer is ready |
| Time, for streamed or delegated work | 240 seconds | A turn streamed back as it runs, or one another agent asked for |
| Model rounds | About 99 | How many times the agent can think and call tools in one turn |
| Tokens | 400,000 | What one turn can spend on the model |
| Repeated calls | Stops at 5 | The same call with the same input and nothing new learned. The agent gets a warning at 3 |
A scheduled run
A trigger on a schedule fires with nobody watching, so it gets its own bounds over the whole firing.
| Limit | Default |
|---|---|
| Time for the whole firing | 15 minutes |
| Turns | 12, each up to 240 seconds |
| Tokens for the whole firing | 1.2 million |
A function run
A function run has its own clock and its own budget of outside calls.
| Limit | Default | What it bounds |
|---|---|---|
| Time, when an agent or script calls it | 5 minutes | The whole function run |
| Calls to outside systems | 1,000 per run | Every operation or model call the function makes. Each item in a list counts once |
| Items in a list a step repeats over | 1,000 | A longer list is refused before any call goes out |
| Items in flight at once | 8 | How many of those calls run at the same time |
Delegation and MCP
| Limit | Default |
|---|---|
| Calls from one run to other agents | 8, shared across the whole chain |
| A call to a tool on an MCP server | 120 seconds |
When an agent hands work to another agent and the answer takes longer than 30 seconds, the caller gets a task to check back on and the delegated work keeps going to its own 240 second limit.
Size the invoice agent against them
Count what one invoice costs. The invoice exceptions agent works through four phases:
| Phase | What it does | Rough cost |
|---|---|---|
| 1. Gather | Fetches the invoice and the purchase order | 2 operation calls |
| 2. Compare | Calls the invoice comparison function | 1 function run |
| 3. Assess | Searches the supplier contracts knowledge base | 2 or 3 searches |
| 4. Prepare | Registers the supplier query | 1 write operation |
One invoice fits easily inside one turn's 99 rounds and well inside 240 seconds.
Now try 4,000 invoices in one run. Time is the limit you hit first. Four phases of model work per invoice, thousands of times over, won't finish in 15 minutes, and the function that repeats over a list stops at 1,000 items. No setting makes this fit, so change the unit of work.
Small work per item: repeat across a list. A function step can repeat once per item, eight at a time. A first pass that takes a list of invoice numbers and fetches each PO number is one call per item, so a list of a few hundred fits. An item that fails is recorded as failed in the results, and the rest carry on.
Bigger work per item: one run per item. Run the agent once per exception. Your finance system can call Mindset when an invoice fails its PO match. Four thousand exceptions become four thousand short runs. Each one has its own record, and each one can be re-run on its own.
If you batch, size the batch against whichever limit you hit first, and size it for a slow morning. If the job only fits when every system answers at once, it doesn't fit.
When a run reaches a limit
A turn that reaches a limit stops where it is and says which limit it reached: out of rounds, out of time, out of tokens, or repeating itself. It hasn't failed. The run keeps everything it did, and it can carry on:
- In a chat, the person sends another message and the agent picks up from there.
- In a run nobody is chatting in (a schedule, a webhook or an API call) that follows a script, Mindset gives the run another turn later, and it continues from the phase it was in.
- A scheduled firing that runs out of turns, time or tokens is recorded as parked, with the reason.
A function run that reaches its time or call limit ends with an error naming the limit.
When a run waits for a person
A script phase can ask a person a question (for example with Slack buttons) and wait for the answer. While it waits, the run is parked. It holds no turn and spends nothing. When the answer comes, the run resumes from that phase. A run nobody is chatting in resumes on its own. A run a person started in chat resumes when that person next sends a message.
- The author sets how long to wait, at least 5 minutes, and which phase to go to if nobody answers.
- A run that is still parked after 31 days is closed as abandoned.
When Mindset itself has a problem
If the server driving a run stops, the run isn't lost. Its hold on the run expires after about 5 minutes and Mindset picks it up again, continuing from where it stopped. If driving it keeps failing, Mindset backs off and retries. After an hour, or 10 failures in a row, it closes the run as abandoned.
Retries inside a function
A function step can retry a failed call. Retries are off by default: a step makes one attempt unless you set more, up to 10. You choose which kinds of failure to retry, such as rate limits, timeouts and network failures. A request the other system rejects isn't worth retrying, because it will be rejected again. You can also give the step its own timeout per attempt.
Retries don't count against the 1,000 call budget, and the function's 5 minute limit doesn't stretch to fit them. When time runs out, attempts in flight stop.
Calling twice doesn't run twice
When your own software starts a conversation with an agent, it can pass an idempotency key. Sending the same key again returns the first conversation and starts nothing new, so a caller that isn't sure its request landed can safely send it again. Each slot on a schedule fires once, for the same reason.
Build so a re-run is safe
A write happens when the agent calls it, as long as the operation is enabled. So a run that stops part way may already have changed something.
- Make steps repeatable where you can. Running one twice should have no extra effect.
- Put what you can't undo last. The invoice agent's first three phases only read. Phase four is the only one that writes, so a run that stops in phase three has changed nothing.
- Read the record before you re-run. Check what happened shows how far the run got and what it called.
You're done when
- You can state one run's worst case: how long it takes and how much it calls.
- Both sit well inside the limits, and you know which limit you'd hit first.
- For each write your agent makes, you know whether doing it twice would matter, and the ones that would happen last.