How it works
Rerunner is designed to test-drive every change to your coding agent's handbook before it's accepted. It will give the agent the same real tasks twice, once with the old version and once with the new, several times each, because an agent doesn't do exactly the same thing every time. Then it will tell you, task by task: same, worse, better, or not sure yet. Think of testing an edited pilot checklist in a simulator before real flights.
The rest of this page is the technical detail. Rerunner is designed to ask one question of every pull request that touches your agent's setup: will the agent behave differently after this change? It will answer in three layers, cheapest first, and block a merge only on evidence that doesn't depend on chance. In the report, same, worse, better and not sure yet are PASS REGRESSED IMPROVED and INCONCLUSIVE. It's designed to run in your CI; other terms are in the glossary.
1. Context diff
What the agent will see differently. No model calls, same answer every time.
Before any agent runs, Rerunner is designed to work out what the agent will load on the base branch and on the pull request: the chain of instruction files and everything they import, rules and the paths that trigger them, skills, MCP servers and their tool descriptions, and settings such as the pinned model. Then it will show the difference, file by file, including when each file loads.
| Context | main | PR #482 |
|---|---|---|
| CLAUDE.md | Loads at session start | Loads at session startChanged: +3 −9 lines |
| docs/testing.md | Imported by CLAUDE.md:8 | MissingRenamed to docs/guides/testing.md; the import still points at the old path |
| .claude/rules/api.md | When the agent reads or edits src/api/**/*.ts |
Unchanged |
| .claude/skills/release | Not present | AddedDescription listed in every session; body loads when used |
| .mcp.json | 1 server, docs | Unchanged |
| Model | Pinned in settings | Unchanged |
Deterministic checks
- Broken imports
An
@pathimport or a linked file that doesn't exist on the pull request branch. In our testing with Claude Code 2.1.293, a broken import produced no warning.- Missing files
Scripts, docs and directories named in the context that were moved or deleted.
- Hidden Unicode
Zero-width and bidirectional control characters that can hide instructions from a human reviewer.
- Likely secrets
Keys and tokens pasted into context, which every agent session would read and send to its model provider.
- Invalid settings
Settings files and
.mcp.jsonthat don't parse or don't match their schema, and rule frontmatter that doesn't parse, which turns a scoped rule into an always-on one.- Size limits per agent
Context bigger than a given agent will load. Codex stops adding AGENTS.md content at 32 KiB by default; Claude Code skips a CLAUDE.md file over 4 MiB and warns when files pass its recommended length. Each agent you run will be checked against its own limits.
A failure in any of these is designed to block the merge, because it gives the same answer on every run.
Planned: checks that need a model. Some problems, like two instruction files that contradict each other, need judgment. These checks will use the Claude API, and every finding must quote the files verbatim; a finding whose quote isn't found in the file is dropped. They are reported, never blocking.
2. Regression run
The same tasks, on both sides, several times each.
Rerunner will check out the base branch and the pull request side by side and run each task in your suite on both: same agent version, same pinned model, same settings, same starting files. The only difference between the two sides is the change under review. Each run will start in a fresh, disposable environment.
Agents aren't deterministic. The same task can pass on one run and fail on the next. So each task will run several times per side (three by default), and Rerunner will compare how often it passed, not a single result. Each run will be scored by the task's own check: a command, such as your tests, that exits 0 on success.
For every run, Rerunner will also record which instruction files loaded and why, which tools the agent called, and the run's turns, tokens and cost. With Claude Code those come straight from the agent; see Built on Claude.
Verdicts
Four answers, including “we can't tell yet”.
- PASS
No difference: the pull request passes as often as base. If a task fails on both sides, the report says so, because the task or its check probably needs attention.
- REGRESSED
The pull request passes less often than base, and a one-sided Fisher exact test (a standard way to check whether a difference in counts could be chance) on the pass counts gives p < 0.10.
- IMPROVED
The pull request passes more often than base, by the same standard.
- INCONCLUSIVE
The results differ, but not clearly enough to say the change caused it. More runs usually settle it. A task where the agent errored on every run, on both sides, is also inconclusive.
Two worked examples
The test asks: if the change made no difference, how likely is a split at least this lopsided? Take all the runs of a task together, count the passes, and ask how often chance alone would put that few of them on the pull request's side.
7 of the 12 runs passed. If the change made no difference, the chance that at most 1 of those 7 passes falls on the pull request's side is:
C(6,1)·C(6,6) / C(12,7) = 6 / 792 ≈ 0.008
p ≈ 0.008, below 0.10REGRESSED
6 of the 12 runs passed. The chance that 2 or fewer of those 6 passes fall on the pull request's side is:
(1 + 36 + 225) / C(12,6) = 262 / 924 ≈ 0.28
p ≈ 0.28, not below 0.10INCONCLUSIVE
Why p < 0.10, and what three runs can tell you
Statistical results never block a merge, so the threshold only decides what gets flagged for a human. With the default three runs per side, only a complete flip clears it: 3 of 3 against 0 of 3 gives p = 0.05, while 3 of 3 against 1 of 3 gives p = 0.20 and is reported as inconclusive. For finer distinctions, raise the run count for the tasks that matter most. Cost grows with it.
What blocks a merge
Only failures that would happen on every run.
- A context check fails: a broken import, hidden Unicode, a likely secret, an invalid settings or
.mcp.jsonfile, or context over an agent's size limit. - An expected file stops loading: a file listed in the task's
expect_loadedloaded in every base run and in no pull request run. - A forbidden file loads: a file listed in the task's
must_not_loadloads in a pull request run.
Differences in pass rates are statistical. Rerunner is designed to report them with their counts, mark them clearly in the pull request comment, and leave the decision to your reviewers. We'd rather say “we can't tell yet” than fail a build on noise.
3. Root cause Roadmap
Which edit did it?
A pull request often changes several pieces of context at once: a paragraph in CLAUDE.md, a skill, an MCP server's configuration. When a task regresses, you need to know which one did it.
Root cause analysis will restore each changed piece to its base version, one at a time, and re-run the regressed task. If restoring one piece brings the old behavior back, that piece is the likely cause, and the report will say so. If no single piece does, the report will say that too.
| Restored to base | Re-run | Result |
|---|---|---|
| CLAUDE.md §Auth | 5/6 | Likely cause |
| docs/testing.md rename | 1/6 | No effect |
| release skill | 1/6 | No effect |
Suites
A short TOML file in your repository. Draft format: field names may change before the first stable release.
Each task has the prompt the agent receives and a check that decides whether it succeeded. Write tasks from work your agent really does: small, recent, and with a check you already trust.
# rerunner.toml (draft format)
model = "sonnet" # pin the model so both sides are comparable
runs = 3 # runs per task, on base and on the pull request
max_budget_usd = 0.50 # cost guard per run
[[task]]
id = "account-export"
prompt = """
Add GET /account/export. It returns the signed-in user's data as JSON.
Follow the API conventions in this repository.
"""
check = "npm test -- account && npm run lint:auth" # exit code 0 means the task passed
expect_loaded = ["CLAUDE.md", ".claude/rules/api.md"]
timeout_s = 900
max_turns = 20
[[task]]
id = "invoice-pdf"
prompt = "Add a PDF download for invoices at GET /invoices/:id/pdf, with tests."
check = "npm test -- invoices"
must_not_load = [".claude/rules/legacy-*.md"] # retired rules must never load again
| Field | Where | What it does |
|---|---|---|
| model | suite | The model both sides run with. Pinning it keeps the comparison fair. |
| runs | suite | Runs per task, per side. Defaults to 3. |
| max_budget_usd | suite | A cost guard for each run. |
| [[task]] | suite | One block per task. |
| id | task | The task's name in reports. |
| prompt | task | What the agent is asked to do. |
| check | task | A shell command that decides success. Exit code 0 means the run passed. |
| expect_loaded | task | Instruction files that must load. One that loads on every base run and on no pull request run blocks the merge. |
| must_not_load | task | Files that must never load. One that loads on the pull request blocks the merge. |
| timeout_s | task | Time limit for each run, in seconds. |
| max_turns | task | Limit on the agent's turns in each run. |
Suites mined from your pull requests Roadmap
Writing tasks by hand is the slowest part. Rerunner will propose tasks from your own merged pull requests: the prompt from what the pull request did, the check from the tests it added, kept only if the check fails before the merge and passes after it. You review every proposed task before it joins the suite.
In your CI
A check designed to run where your code already gets tested, on GitHub or GitLab.
Rerunner is designed to run as a check in your own CI, on your runners. It will run on pull requests that change the agent's context: CLAUDE.md and AGENTS.md files, .claude/, skills, .mcp.json, settings, the pinned model, or the suite itself. Pull requests that don't touch any of these won't trigger a run, so they cost nothing.
The report will arrive as one pull request comment, updated on every push, with a pass or fail status you can make required.
It will never run automatically on pull requests from forks: a fork can change what the agent executes, so a maintainer starts those runs by hand, after review, in a sandbox.
Every run will call your model provider under your account, so cost grows with tasks × runs × two sides. Give Rerunner a dedicated API key with a spending limit, and run it on a fresh runner with no production secrets. The security page has the details.
Agents
Rerunner starts with Claude Code, the agent where we can see exactly which instruction files loaded and why. Codex, Cursor, Gemini CLI and OpenCode are on the roadmap; for them we'll rely on each tool's documented loading rules and its headless output.
It's vendor-neutral. We don't make a coding agent, and the same suite is meant to run against whichever agents your team uses.
Test-drive your agent's changes before they ship.
We're looking for a few teams to try Rerunner on one real repository and tell us where it's wrong.
Agent regression report
mainat3f2c1a09be71d4Context diff
CLAUDE.md+3 −9 linesdocs/testing.mdrenamed todocs/guides/testing.md.claude/skills/release/SKILL.mdaddedCLAUDE.md:8imports@docs/testing.md, which no longer exists on this branch. This pull request renamed it todocs/guides/testing.md. In our testing with Claude Code 2.1.293, a missing import produced no warning..mcp.jsonare valid. Context is within Claude Code's size limits.Regression run
Merge blocked by one deterministic check: the missing import.
account-exportregressed (6 of 6 against 1 of 6, one-sided Fisher exact test p ≈ 0.008). That difference is statistical, so it's reported for your reviewers, not enforced.date-parsingis inconclusive at six runs per side.