Why derived beats hand-written
Hand-written tests check what their author remembered. Cases derived from the graph check every declared rule and every required output, including the exceptions a rule carries ("unless the input supplies it"), because those live in the same structure. When the process changes, the cases change with it and the old passing result is marked stale.
Each verdict is a single sample of a model’s judgement, so the platform reports repeatability alongside the score rather than pretending one run is the truth, and a failing case explains what happened and what to change: a rule to reword, an example to dismiss, a system to connect.
Evals, then shadow, then approval
Evals are the first gate: an agent cannot be submitted to IT without passing them. Shadow mode is the second, on real cases. IT approval is the third. Each produces evidence a person can read, and none of them can be skipped by an agent that is confident.
The terms, as the product defines them
- Eval
An automated test that checks a domain agent makes good decisions.
A suite of test cases — real-ish inputs with expected outcomes derived from your process and corpus — run against a domain agent. Each case is graded (pass/fail) so you can see, with real numbers, whether the domain agent is reliable before it goes live.
- Compile
Turn your description into a structured process graph.
When you compile, Tacit reads everything you described or uploaded and produces a structured graph (the BPMN diagram) — the steps, the order, the decision points. A "critic" pass then flags gaps (missing steps, unclear rules) for you to resolve.
- Graph / BPMN
The visual map of a process — steps and how they connect.
The diagram of your process: boxes (steps) connected by arrows (flow), branching on conditions. Steps are color-coded by who does them (human / system / domain agent / hybrid). You can edit the graph directly on the canvas.
Questions people ask about this
- The checks didn’t pass — what do I do?
- Each failing check now opens with “What happened and what to do”: the platform reads the run itself and tells you in plain words — for example “the rule guard blocked the send under rule 3, so the agent re-queried instead of reusing the supplied count” or “rule 11 and this check disagree on the SKU count scope” — with the buttons to act: apply a suggested rewording of the rule, keep the rule and dismiss the example, disable the rule, tell the agent the rule, re-run, or view the run. It is the same card in the IT Checks & activity tab and in the business view. If an enabled rule changed after the checks ran, a banner says the results may be out of date and offers a re-run. Below the card: open the agent; under “Getting it ready” you’ll see each example the agent disagreed with. A failed check is a DISAGREEMENT, not a verdict: the panel states the one question to answer (“Your agent and this example disagree about record count — which one is right?”) and shows the two sides field by field, highlighting only what actually differs. Fields both sides agree on are collapsed, so you are reading the argument rather than hunting for it. Type the missing rule in plain words (e.g. “Deny refunds past 30 days unless the item arrived damaged”) and click Save & re-check, or upload the SOP that covers it. We update the agent and re-run the checks automatically — no technical editing. If an example is just unrealistic (the agent was actually right), click “This example is wrong” to stop it counting against the agent. Not sure of the rule? Loop in a colleague before sending to IT. Note: for some agents the automatic checks can only verify the OUTPUT SHAPE (not that it does the job) — in that case we won’t let you send it to IT on shape-checks alone. Click “Try it safely” with a real example first (or add a real check); once you’ve genuinely tried it, submit unlocks. Note: the panel separates failures you can fix (a policy/wording gap — “tell us the rule”) from purely technical ones (an output-SHAPE or system-contract check) — the latter never ask you to write a rule; just re-run the checks, and if one persists your IT admin can adjust the agent’s output format.
- The Results tab shows green checks — does that mean the checks/evals passed?
- No — those are two different things. The Results tab (“What this agent did”) is a LOG of every time the agent ran. A green check there means a run finished and produced an output — not that it was correct and not that the checks passed. Rows tagged “Trial” ran with writes simulated, so nothing real was touched. The actual checks/evals live under “Getting it ready” on the Overview tab; that’s where you’ll see pass/fail. So a run can finish (green in Results) while a check still fails (red under Getting it ready) — “it ran” is not the same as “it ran correctly.”
- In what order do I do this — build, evals, submit, connect, go live, autonomy?
- The business builds and validates; IT connects and approves — a clean handoff. The order: (1) BUSINESS builds the agents and runs evals (≥80% pass) — that’s all it takes to submit; you do NOT need tools connected first. (2) BUSINESS clicks “Submit the whole process for IT review”. (3) IT picks it up: connects the required tools and binds the API fields in Integrations. (4) IT approves — approval is the go-live gate, so it can’t approve until the tools are connected, and nothing ever runs unconnected. (5) It deploys live in domain agent mode. Then the domain agent trial / agreement begins: it works alongside a human who approves each draft in the queue, and earns the right to act UNATTENDED once it agrees with real human decisions ≥90% over ≥10 real cases. IMPORTANT: the “domain agent trial” agreement is NOT a pre-launch step — the platform never fakes “what a human decided,” so the score stays empty until it’s live and people review its drafts. Two separate gates: Evals + IT approval (with tools connected) make it LIVE (human-in-the-loop); agreement with real decisions earns AUTONOMY (unattended).
- What’s the difference between “checks” and a “trial” — aren’t both just simulations?
- Both run on MOCKED tools (nothing real is touched), but they check different things. CHECKS (evals) auto-grade CORRECTNESS: the system runs the agent on sample situations with known-good answers and scores whether its output is right — you need ≥80% to send to IT. A TRIAL is a simulated dry-run where its PROPOSED decision is compared to what a human would do (agreement) — pre-launch it’s just a preview you can watch, and it is NOT required to send to IT (only checks + connections are). The REAL domain agent trial happens AFTER go-live: the agent works alongside a person who approves each draft in the queue, and it earns unattended autonomy once it agrees with real human decisions ≥90% over ≥10 cases. WHERE to run the mocked trial: in the BUSINESS view, click “Try it safely” at the top of the agent; in the IT view, open “Checks & activity → 2 · Domain Agent trial”. IMPORTANT: the IT view also has a “3 · Simulate (live)” tab — that one runs against your LIVE connections (real writes can hit Salesforce) and is the last check before Deploy, so it is NOT the mocked trial; use Domain Agent trial for a no-side-effects run. Neither checks nor the pre-send trial touches production — live connections only happen once IT approves and it’s deployed.
- Do I have to test my agent or run checks before it can go live?
- No — Tacit runs the checks for you, automatically. As soon as you finish describing it (or change it), we quietly run the quality checks and a safe trial in the background. The agent’s Overview shows a simple tracker — Built → Checking it works → Approved by IT → Live. When the checks pass and any needed platforms are connected, the agent points you to “Submit the whole process for review” — agents go to IT together as one process, not one at a time. Your IT team gives the final OK before anything goes live — you never have to run evals or trials by hand.
Related
See it on one of your own processes. Free for the whole product for a trial period, no card needed to start, every write held for your approval.