Flags and evals
| Command | Purpose |
|---|---|
vatio flags | Every open flag as markdown: the conversation, each tool call of that turn with your backend's response, and what the reply should have been |
vatio flags --save [PATH] | The same, written to vatio-flags.md or PATH to hand to a coding agent |
vatio flags --json | The open flags as JSON rows |
vatio eval [--env NAME] [--agent SLUG] | Replay every eval case against an environment, wait for the judge, print the report; exits 1 if a case fails |
vatio eval --no-wait | Start the runs and print their ids |
vatio eval --show RUN_ID | One run's report |
vatio eval defaults to the branch's environment on a branch and to preview elsewhere, and waits up to 900 seconds (--timeout N). A case that passes on live and fails on the environment you ran is listed first, as a regression.
Each flag with a note is an eval case. Replays call the environment's tools for real. The whole loop, and what resolves a flag, is in Improve your agent from flags.
