m4Mindset docs

Docs / AMS / Optimize / Diagnose and improve an agent

View as Markdown

Diagnose and improve an agent

Work out why an agent didn't do what you expected, fix it in the Agent Builder, and check the fix before you make it live.

After this page you can work out why an agent didn't do what you expected, make the fix, and check it before it goes live. This is the outcome side of Optimize: making an agent right more often. To find out what happened first, see Check what happened.

Where the fix gets made

Almost every fix is made by talking to the Agent Builder, docked on the left of every tab of an agent. Describe the change in plain words and it makes it. Open the tab the problem lives on first, because that's what the builder is looking at. On Script, "the assess phase shouldn't pass unless the answer names a contract clause" changes that phase. On Resources, "take the write operation off the early phases" changes that.

An agent's tabs are Overview, Chat, Script, Triggering, Resources, System Prompt, Settings, Acceptance tests, Versions & Availability and Embed.

Use Orca to create something you don't have yet: a new agent, function or connection. It proposes a plan as a card with Approve plan and Don't build this, and stops until you choose. An agent that exists and behaves badly is the Agent Builder's job.

Edits don't go live until you activate them

Every change to an agent (its script, system prompt, settings, resources and the timing of its schedules) goes into one draft. Creating or deleting a schedule is the exception: that takes effect at once.

  1. Press Save as version. That creates a version and puts nothing live.
  2. Run the agent's acceptance tests against it, or try it on the Chat tab.
  3. On Versions & Availability, Activate the version.

A conversation already under way moves to the newly active version at its next turn. A run that follows a script finishes on the version it started with.

Work out which of eight things happened

Start at the top. Most reports of "the agent is broken" are settled in the first three.

#What you seeUsually
1No output at allIt never ran
2It ran and nothing changed in your systemA write was refused, or the run is waiting on a person
3It stopped part wayA phase's exit condition didn't hold, or the run reached a limit
4It finished and the answer is wrongA condition can be met without the work being done
5It did something too earlyIt had that operation in that phase
6An error naming an outside systemA connection failed, usually a credential
7Plausible output built on wrong dataIt reached the wrong account or read the wrong fields
8Right most times, wrong sometimesToo much left to the model

1. It never ran

Look for the agent in Monitor → Resources. No run means nothing started it, and the agent isn't what to fix. Go to its Triggering tab and check the schedule is enabled, or that the webhook source or API call is reaching it. See Making it run.

The one people miss is the schedule that half happened. Fourteen expected firings and nine actual isn't healthy, and the last run will look fine. Ask the Monitor agent for the run count over a window.

2. It ran and nothing changed

Two causes.

  • The write was refused. If your org has turned off automatic enabling of writes, a new write operation waits for a person to approve it, and calls to it are refused until then. Approve it on the connection's Operations tab.
  • The run is parked. A script phase with an $ask posted a question to Slack and is waiting for an answer. Check the channel.

3. It stopped part way

Open the session in Log → Sessions and read Steps.

When a phase's exit condition doesn't hold, Mindset gives the agent a gate status every turn: each condition marked as met or not, the judge's reason for any judged condition that failed, and the instruction never to claim a phase is complete when it isn't. That status usually names your problem in plain words.

Usually the condition is worded badly. If the assess phase requires a clause number and your contracts knowledge base has none, no answer will ever meet it. Fix it on the Script tab: tell the Agent Builder what the condition should say, or what the phase is missing.

If the turn stopped on a limit (out of time, rounds or tokens), see Limits and run behavior.

4. It finished and the answer is wrong

Every condition passed, and the output is still wrong. The conditions can be met without the work.

  • A loose judged gate. "A recommendation exists" passes on anything. "A recommendation naming a specific clause of the supplier contract" doesn't. Judged gates are strict: only a clear yes passes, and the reason is recorded.
  • A phase that skipped its work. If the assess phase is granted the contracts knowledge base and never searches it, the answer sounds just as confident. Check granted versus used in Check what happened.

Tighten the condition so meeting it needs the work, then re-run the same invoice and read the judge's reason.

5. It did something too early

It registered the supplier query before it compared anything. There's one cause: it had that operation in that phase. A phase can reach only what it declares. On the Script tab, take the write operation off every phase except the last.

6. A connection failed

Every operation on one connection fails at once while everything else works. Nearly always the credential expired or was rotated.

Replace the credential on the connection's Settings tab. For some connection types that means Set this up again, which checks the new credential works before it goes live. Then run one read on its own before you re-run the agent: Data preview for a Google Sheet, Query console for a database. Credentials never reach the agent or the model, so nothing else needs changing.

If one operation fails with an authorization error while the others work, the credential is valid but scoped too narrowly for that operation.

7. The data is wrong and nothing errored

The output is plausible and about the wrong invoice. Check in this order:

  • The connection points at the wrong account. A credential aimed at the wrong account returns valid data. Run one read and check a value you can verify elsewhere, such as a PO total.
  • A function reads fields that don't exist. If an operation's output shape came from documentation rather than a real response, the comparison function finds no lines and reports no differences on every invoice. Open the function's Preview tab, run it on INV-4471, and watch which step returns nothing.

8. Right most times, wrong sometimes

Something was left to the model that shouldn't have been.

  • The script does too little. One phase that says "handle the exception" lets the model take a different route each time. Split it into phases with exit conditions.
  • It can't reach what it needs. An agent missing the payment terms won't stop and ask. Compare granted versus used for a run that went wrong and one that went right.
  • The instruction reads two ways. "Recent invoices" means this month on one run and this quarter on the next. Say the number.
  • Nothing checks the output. Add a judged gate to the phase where quality matters.

When a prompt needs a definition, move it into a function

The system prompt says to query any material discrepancy. "Material" is a real rule in your business: over $500, or over 5% of the PO value. In the prompt, a model applies it by reading a sentence, and a $502 difference on a $90,000 order goes one way on one run and the other way on the next.

Move it into the invoice comparison function. Have its last step return whether the total difference is over the threshold. Then the threshold is a number in one versioned place, the phase condition becomes a machine check, and every run applies the same rule.

If a phrase in your prompt would need a definition before somebody else could apply it the same way every time, that definition belongs in a function.

After any change, run the tests

  1. Open Optimize → Acceptance tests, pick the agent, and press Run all.
  2. Read the failures before the passes.
  3. If no test covers the problem you fixed, add one now, while you remember the exact situation.

The platform doesn't block activation on these results. You decide. See Test an agent's behavior.

Things to be aware of

  • To stop something right now, revoke the operation on the connection. Every call checks it, including calls in a conversation already under way.
  • A write happens when it's called. Before you re-run a run that stopped part way, check what it already changed.
  • Change one thing at a time, or you won't know which fix worked.

When it doesn't work

  • You changed the wording and nothing changed. The wording wasn't the cause. Go back through the eight, and check granted versus used.
  • It passes your tests and fails in real use. Your tests are cleaner than reality. Turn the invoice that actually failed into a test.
  • It got worse. Roll back: on Versions & Availability, activate the earlier version. Functions and widgets keep versions the same way. See Change something that's already live.

You're done when

  • You can name which of the eight it was before you changed anything.
  • The fix was made in the Agent Builder, saved as a version, tested, then activated.
  • Anything that needed a consistent definition is a function's output, not a phrase in a prompt.
  • A test now covers the thing that went wrong.