Play two weeks of your rules forward before any of it happens. Test Run never calls your AI provider and never changes a task.
Two weeks in about ten seconds
Test run
No risk. Nothing goes to your AI provider, nothing changes in your tasks.
Day 1 of 5
Plan a weekend trip to Lisbon › Shortlist a few things to see
about 3.1k tokens
Plan a weekend trip to Lisbon › Save an offline map and phrasebook
about 1.8k tokens
2
Actions so far
~5k
Est. tokens
0
You'd approve
Press Play, or step a day at a time
No risk, and that is a promise about the code
Test Run calls the same rule engine the real thing calls, with the same guardrails, and stops there. It never sends anything to your AI provider, never writes an artifact, and never changes a task. The estimates it shows you are arithmetic on a price table, not a bill.
Pick what it runs against
My tasksuses a snapshot of your real account, which is the honest answer to "what would this do to me". Sample project uses a made-up project so an empty account can still try a rule out.
Play it, or walk it
Play advances a day at a time on its own. Step moves one day. Step by step moves one decision at a time and highlights the rule on the canvas that produced it, which is the fastest way to find the node you drew wrong.
Read the day table
Each row is one decision. It tells you:
A day where nothing matches says so. An empty timeline is usually a scope that is narrower than you meant.
Look at the totals, not the rows
The footer is the part that changes minds: how many actions over the two weeks, how many of them you would have to approve, and the estimated tokens. If you would be approving twenty things a week, the rule is too broad, and you have learned that for free.
Try it on one crumb
Simulation cannot tell you whether the draft is any good. This button picks one real step and runs the action for real, in preview: it uses your key once, and nothing is applied to the task. It asks you to confirm first, because it is the one part of this screen that spends anything.
It needs a paid plan. Simulation itself does not.
Test Run and shadow mode are different tools
Test Run answers "what would this do" in ten seconds, against a snapshot, with made-up days. Shadow mode answers the same question over three real days, as things actually happen, and writes it into Activity.
Use Test Run while you are drawing. Use shadow before you trust it. They are not substitutes for each other, which is why saving a rule set puts it in shadow regardless of how good its simulation looked.