m4Mindset docs

Docs / AMS / Test and go live / Test before it goes live

View as Markdown

Test before it goes live

Run a script's or function's test parameters, agree the writes a test makes, and get past the go-live gate on activation, publish and rollback.

After this page you can prove that every step of a script or function works before people meet it, agree the writes a test makes, and put the version live without meeting a refusal you didn't expect.

What this is

Mindset keeps a tested status on every operation: the named calls an agent makes to a connected system. An operation is Tested when a successful call to it is on record since its last lapse. Otherwise it is Untested.

The status matters at one moment: going live. Mindset refuses to put a script or function version in front of real users while it calls an untested operation. This is the go-live gate. A test run of the script or function makes its operations tested, and the gate then lets the version through with nothing to accept.

You test with test parameters: named sets of inputs kept beside a script or function. Each set has a Run button.

Where test parameters live

ForWhere
A scriptThe agent's Script tab, under the script, in "Test parameters kept beside this script"
A functionThe function's Preview tab, under the Diagram run area

Each set shows its name, a badge (Never run, Verified with a date, or Not verified with a date) and Run. Changing the sets never creates a new version of the script or function.

The page has no form for adding a set. When the list is empty it says "No test parameters yet. Ask the builder to keep some". Tell the builder docked beside the script or function which cases to keep. For the invoice exceptions script, ask for three:

  • An invoice that matches its purchase order, so Gather skips straight to Close.
  • An invoice with a price difference the supplier contract allows.
  • An invoice with a difference the contract doesn't allow, so Prepare registers a supplier query.

Orca keeps test parameters for what it builds, and you can also create them from Claude through the admin MCP.

Run a set

  1. Open the Script tab (or the function's Preview tab).
  2. Press Run on a set.
  3. If the run would write to a connected system, a dialog opens first. See Agree the writes a test makes.
  4. Read the result.

A run uses the real logic on the set's inputs. Reads go to your real systems. A function's run uses its newest version, so you can test a version before you publish it.

The result lists every step of the script or function:

  • Exercised, with the scenario, the inputs, when it ran and what came back.
  • Not exercised, with the reason: Branch not taken, Not reached, Refused, Write not agreed or Failed.

The verdict reads Verified: every step was exercised only when every step was. Otherwise it reads Not verified and names each step that wasn't. Whether the expectations written into the scenarios held is reported on a separate line.

To verify all four phases of the invoice script, you need cases that take each branch. The matching invoice exercises Gather and Close. Only the disallowed difference reaches Prepare.

Agree the writes a test makes

A test run that writes needs a go-ahead from an org admin before it starts. Before anything runs, Mindset works out from the script or function itself every operation the run can call, and whether each call reads or writes. A run with no writes needs no go-ahead.

When the run would write, the dialog This test writes to connected systems lists each call ("Writes 'register_supplier_query' on finance system") and says "any write not listed here is refused". Press Go ahead and the run goes end to end. A write that isn't on the list is refused when it tries to leave.

Things to know about a go-ahead:

  • It covers one exact list. A later run with the identical list reuses it. If the script changes so that the list changes, the dialog asks again, saying what the test would do has changed.
  • Only an org admin can give it, in person or through their own agent acting on their instruction. An org API key can't.
  • It's on the record. Each go-ahead is stored with who gave it, how, and the list it covered, and written to the audit log.
  • The writes are real. The test registers a real supplier query. Write the test inputs so it lands somewhere safe, such as a test supplier record.

An operation can also hold its own test parameters, separate from any script. On a connection's Operations tab, open Test parameters on an operation to set its Inputs, its Write target and a Safe value for the write target. These are used when the operation is tested on its own. They are never substituted into a script's or function's test run.

Chat on a version that isn't active

The same rule applies on the agent's Chat tab. When you talk to a saved version that isn't active, and that version can write, the tab says "This saved version can write to finance system. Its writes are refused until you agree them." Press Agree writes and confirm. Its chat turns then perform the agreed writes for real. The active version is never asked. See Agent versions.

Tested status

An operation becomes tested when a successful call to it is recorded. Any successful call counts, from a test run or from real use. A call that returns nothing also counts as a success.

The status shows as a Tested or Untested badge on each operation, on the connection's Operations tab or Tools tab. Hover over it to read why. Monitor → Resources lists "Enabled operations not yet tested", and marks the ones a saved script or function calls.

What makes it lapse

Tested status lapses back to untested only when one of four things changes:

ChangeThe reason shown
The operation is registered again with a different method, path or parameters"The operation was re-registered with a different call shape (method, path or parameters)."
Its connection is pointed at a different base URL"Its connection was pointed at a different base URL."
A different credential is bound to its connection (a new key, header, binding or OAuth grant). A routine automatic token refresh doesn't count"A different credential was bound to its connection."
A successful call breaks the output shape the operation had established"A call to it returned output that broke the shape it had established."

A lapse never happens on a timer. It stops nothing either: an agent already calling the operation keeps calling it. The next successful call makes it tested again. A lapse matters only the next time a version that calls the operation tries to go live. For what counts as breaking an output shape, see When a connected system changes.

Each environment keeps its own

Tested status is kept per environment. A test run in Sandbox doesn't make the Production copy of register_supplier_query tested. When you rebuild the invoice agent in Production, the gate stops you again until a test run there succeeds or an admin goes live untested.

The go-live gate

The gate checks three acts:

ActWhereRefused when
Activating an agent versionVersions & Availability tabThe version puts a new script, or a new text of the script, live, and the script's own steps call an untested operation. Rolling back to an earlier version is an activation and is checked the same way
Publishing a function versionThe function's Versions tab, PublishThe version's steps call an untested operation
Rolling back a functionThe function's Versions tab, Roll back to thisThe same

Saving is never gated. An activation that changes only the model, the prompt wording or a schedule is never refused, because it changes nothing the script runs. The gate also leaves out tools given directly to the agent's model on its Resources tab, and calls whose connection is chosen while the run goes.

A script that calls a function doesn't meet the gate for that function's operations. The function met the gate when its version was published.

When the gate refuses, a notice titled "1 operation has never succeeded" (or "N operations have never succeeded") lists each untested operation and its connection, and says a test run makes them tested. Run your test parameters, then try again.

Go live untested

If you need the version live before a test can succeed, an org admin can accept the risk:

  1. Tick Go live with these untested. This is recorded with your name, and they stay untested until a real call succeeds.
  2. Optionally say why in Why (optional).
  3. Press Go live untested.

The override covers exactly the operations the notice listed. It's written to the audit log with who accepted it, through which surface, which operations and the reason, in the same step as the activation. The operations stay Untested until a real call succeeds. Members who aren't org admins, and org API keys, can't override.

Re-runs when a system changes

When an operation's output breaks its established shape, Mindset re-runs the stored test parameters of every script and function whose live version calls it. A background sweep picks these up every two minutes. Each set then shows one line, such as "Re-ran by itself after 'get_invoice' on finance system stopped returning 'due_date' as it used to."

The re-run is nobody's action, so it can't give a go-ahead. Its writes run only under an earlier go-ahead for the identical list. Any other write is refused and shows as Write not agreed. Saving or activating a version doesn't start a re-run.

How this differs from behavior tests

Mindset has a second kind of test. Test an agent's behavior covers it.

Test parametersBehavior tests
What it provesEvery step of a script or function runs, and every operation it calls worksThe agent behaves as you want across many conversations
WhereThe Script tab, or a function's Preview → DiagramOptimize, and the agent's Acceptance tests tab
WritesReal, within the list an admin agreedRefused before they leave, every one
ResultVerified or Not verified, step by stepPass or fail against a threshold over several runs
Blocks going liveYes, through tested status and the go-live gateNo. Nothing runs them on save or activation
Runs again on its ownWhen an operation it calls changes shapeOn a schedule you set

What can go wrong

  • "N operations have never succeeded" on an operation that worked last week. Something lapsed it, such as a rotated key. Hover over its Untested badge to see which change, then run your test parameters.
  • Verified never appears. A step wasn't exercised. Read its reason. Branch not taken usually means you need another case that takes that branch.
  • A step reads Write not agreed. The run was a re-run with no go-ahead for its list, or the list changed. Press Run yourself and give the go-ahead.
  • The gate stops you in Production after you tested in Sandbox. Tested status is per environment. Test there too.

You're done when

  • Each script and function has test parameters that take every branch, and each set reads Verified.
  • You gave a go-ahead for a test that writes, and checked where its writes landed.
  • You activated a version with no untested operations, so the gate let it through.
  • You know who in your org may press Go live untested, and that it goes on the record.