Improve your agent from flags
The people reading your agent's conversations know when a reply is wrong. In the inbox, a supervisor flags that reply and writes what the agent should have said. Vatio turns that note into work your coding agent can do on its own:
- A digest.
vatio flagsprints every open flag as one markdown document written for a coding agent: the conversation up to the reply, every tool the agent called on that turn with what your backend answered, and the supervisor's note. - An eval case per flag. The visitor's turns become the input and the note becomes the expected outcome.
vatio evalreplays the cases against any environment and a judge compares each new reply with the note. - Resolution on evidence. When live passes a flag's case, the flag is marked fixed in the inbox, with the version that fixed it. Nobody has to say it is done.
Your coding agent has your repository and your backend in front of it, which Vatio does not. That is why the fix is its job: the cause is often a tool response missing a field, not the prompt.
The loop
vatio flags --save # vatio-flags.md: hand it to your coding agent
git switch -c fix/flags-semana
# ... your agent fixes vatio.yml, knowledge, tools or the backend they call ...
vatio push # → environment fix-flags-semana
vatio eval # replays every case there; exits 1 if one fails
gh pr create # with GitHub connected, the PR runs the evals tooMerge when the evals pass. The publish replays every case on live, and the flags whose cases now pass are marked fixed.
A prompt that works with any coding agent:
Read the output of
vatio flags. Group the flags by cause, fix each cause where it lives —vatio.yml, knowledge, a tool, or the backend a tool calls — on a branch, prove it withvatio eval, and open a pull request.
The digest itself carries the same instructions at the top, so an agent handed only the file knows what to do with it.
Let your coding agent find it on its own
So you do not have to remember the command, two pointers tell a coding agent that vatio flags is where improving the agent starts:
AGENTS.md.vatio initwrites a short Vatio section into the workspace'sAGENTS.md, or appends one to an existing file that does not mention Vatio. Most coding agents read it without installing anything.The
vatio-improveskill, for Claude Code:/plugin marketplace add vatio-ai/skills /plugin install vatio@vatioClaude Code then uses it when you ask it to improve or fix your agent, or as
/vatio:vatio-improve. For another coding agent, copyskills/vatio-improve/SKILL.mdto wherever it reads skills from.
Both only point at the CLI. The instructions come from the platform each time vatio flags runs, so neither goes out of date.
What a flag carries
| In the digest | Where it comes from |
|---|---|
| Agent, environment, deployed version, channel | The chat the reply was in |
| The conversation up to the flagged reply | The last 10 messages before it |
| What the agent did on that turn | Each tool call: its arguments and the response your backend returned; knowledge lookups; tokens |
| What it should have said | The supervisor's note |
| Eval case | The case Vatio created from the flag |
A flag without a note is not in the digest and has no eval case: "this was wrong" without the right answer is not something an agent can be fixed from. Editing the note reopens a fixed flag; removing the flag removes its case.
Eval runs
vatio eval replays every eval case of the workspace against one environment, matching cases to that environment's agents by slug. Each case runs in a fresh chat: the visitor's turns (up to the last 8) are sent one by one and the agent answers each. A judge then reads the replay and the expected outcome and answers pass or fail, with one sentence of reasoning.
| Default environment | The branch's, on a branch; otherwise preview. --env NAME to choose |
| Exit code | 1 when any case fails or errors, so it can gate CI |
| Regressions | A case that passes on live and fails on this environment is listed first |
| Console | Every run, with the judge's reasoning, on the agent's eval runs page. A person's pass or fail there overrides the judge |
| Judge | A model from a different family than your agent's, so it never grades its own kind of answer |
| Replay chats | Not in the inbox and not counted as conversations, live included: they are Vatio testing the agent, not visitors |
A replay is a real agent turn, so a case can pass once and fail the next without anything changing. Run it again before you trust a single failure.
Evals call your tools for real
A replay is a real agent turn, so the tools it calls hit your backend with the environment's secrets. Point a branch at a sandbox backend with a branch-only secret if a tool writes. See Secrets.
Runs start on their own in two places, besides vatio eval:
- After every publish to live, if the workspace has cases. This is the run that resolves flags, and the baseline regressions are measured against.
- On every pull request preview, with GitHub connected. The result is commented on the pull request. See GitHub.
You can also add cases by hand on the agent's eval cases page in the console.
The daily digest
Every morning at 8:00 (America/Santiago), each workspace with new flags gets:
- A mail to its owner and developers, with how many flags arrived and the command to hand them to a coding agent.
- A GitHub issue, with GitHub connected: one issue labelled
vatio-flagsthat holds the digest, updated every day and closed when every flag is fixed. See GitHub.
API
| Endpoint | Returns |
|---|---|
GET /api/developer/v1/:workspace/flags.md | The digest, as vatio flags prints it |
GET /api/developer/v1/:workspace/flags | The open flags as JSON |
POST /api/developer/v1/:workspace/evals/runs | Starts a run; body { "environment": "fix-pagos", "agent": "main" }, both optional |
GET /api/developer/v1/:workspace/evals/runs/:id | One run's progress: cases, done, passed, failed, finished |
GET /api/developer/v1/:workspace/evals/runs/:id.md | One run's report |
All of them take the same token as the CLI.
