Why shadow mode exists
Test cases show an agent can handle the situations someone thought to write down. Shadow mode shows what it does with the situations nobody wrote down, on the real data, at the real volume, next to the people who know the work. It is the evidence a team needs before it lets an agent act, and the evidence IT needs before it approves.
It is also how a team discovers that a rule was ambiguous. When the agent and the person disagree, someone reads why, and the process description or the rule gets fixed at the source rather than patched in the agent.
What happens to writes in shadow mode
Nothing real is touched. A write in a shadow run is simulated: the trace records the email that would have been sent or the record that would have been created, marked as simulated, so the person can read it without anyone receiving it. Simulated writes never appear in the approval queue, because there is nothing to approve. The queue is for real writes from live runs.
When the agreement bar is met and IT approves, the agent is promoted. Promotion is versioned and reversible: a previous version can always be restored.
The terms, as the product defines them
- Shadow mode
The domain agent runs alongside humans without acting, to prove itself.
In shadow, the domain agent proposes what it would do on real cases while a human makes the actual call. Tacit measures how often they agree. A domain agent must reach a high agreement bar (default 90%) before it can be promoted to act on its own.
- Promote
Move a verified domain agent to production so it can run.
Promotion makes a domain agent version the live, production one. It requires passing evals + shadow and a governance/IT approval. Promotion is append-only and reversible — you can always roll back to a previous version.
- Earn autonomy
A domain agent acts on its own only after it proves it agrees with you.
The core safety idea: nothing acts autonomously by default. A domain agent drafts → is reviewed → shadows → and only then, with evidence and a human approval, earns the right to act. Trust is earned, not assumed.
Questions people ask about this
- The Results tab shows green checks — does that mean the checks/evals passed?
- No — those are two different things. The Results tab (“What this agent did”) is a LOG of every time the agent ran. A green check there means a run finished and produced an output — not that it was correct and not that the checks passed. Rows tagged “Trial” ran with writes simulated, so nothing real was touched. The actual checks/evals live under “Getting it ready” on the Overview tab; that’s where you’ll see pass/fail. So a run can finish (green in Results) while a check still fails (red under Getting it ready) — “it ran” is not the same as “it ran correctly.”
- My trial ended and I see “Free trial expired” — what happens now?
- Your 7-day trial ended without a payment method on file, so nothing was charged and the workspace closed. Nothing is deleted: processes, domain agents, connections and knowledge are all kept. The owner picks Team or Business on that screen and pays through Stripe Checkout, or schedules a call for Enterprise; the workspace reopens the moment the payment goes through and any deployed agents resume. Team members who are not the owner see the same screen with the owner’s email to ask. If you added a card during the trial instead, the plan started by itself on day 7 and you never see this screen — the header badge reads the plan name rather than “Free trial”.
- What’s the difference between “checks” and a “trial” — aren’t both just simulations?
- Both run on MOCKED tools (nothing real is touched), but they check different things. CHECKS (evals) auto-grade CORRECTNESS: the system runs the agent on sample situations with known-good answers and scores whether its output is right — you need ≥80% to send to IT. A TRIAL is a simulated dry-run where its PROPOSED decision is compared to what a human would do (agreement) — pre-launch it’s just a preview you can watch, and it is NOT required to send to IT (only checks + connections are). The REAL domain agent trial happens AFTER go-live: the agent works alongside a person who approves each draft in the queue, and it earns unattended autonomy once it agrees with real human decisions ≥90% over ≥10 cases. WHERE to run the mocked trial: in the BUSINESS view, click “Try it safely” at the top of the agent; in the IT view, open “Checks & activity → 2 · Domain Agent trial”. IMPORTANT: the IT view also has a “3 · Simulate (live)” tab — that one runs against your LIVE connections (real writes can hit Salesforce) and is the last check before Deploy, so it is NOT the mocked trial; use Domain Agent trial for a no-side-effects run. Neither checks nor the pre-send trial touches production — live connections only happen once IT approves and it’s deployed.
Related
See it on one of your own processes. Free for the whole product for a trial period, no card needed to start, every write held for your approval.