Rerunner Early access

How it works

Rerunner is designed to test-drive every change to your coding agent's handbook before it's accepted. It will give the agent the same real tasks twice, once with the old version and once with the new, several times each, because an agent doesn't do exactly the same thing every time. Then it will tell you, task by task: same, worse, better, or not sure yet. Think of testing an edited pilot checklist in a simulator before real flights.

The rest of this page is the technical detail. Rerunner is designed to ask one question of every pull request that touches your agent's setup: will the agent behave differently after this change? It will answer in three layers, cheapest first, and block a merge only on evidence that doesn't depend on chance. In the report, same, worse, better and not sure yet are PASS REGRESSED IMPROVED and INCONCLUSIVE. It's designed to run in your CI; other terms are in the glossary.

1. Context diff

What the agent will see differently. No model calls, same answer every time.

Before any agent runs, Rerunner is designed to work out what the agent will load on the base branch and on the pull request: the chain of instruction files and everything they import, rules and the paths that trigger them, skills, MCP servers and their tool descriptions, and settings such as the pinned model. Then it will show the difference, file by file, including when each file loads.

Context diff main against PR #482, Claude Code
ContextmainPR #482
CLAUDE.md Loads at session start Loads at session startChanged: +3 −9 lines
docs/testing.md Imported by CLAUDE.md:8 MissingRenamed to docs/guides/testing.md; the import still points at the old path
.claude/rules/api.md When the agent reads or edits src/api/**/*.ts Unchanged
.claude/skills/release Not present AddedDescription listed in every session; body loads when used
.mcp.json 1 server, docs Unchanged
Model Pinned in settings Unchanged
Illustration — not captured output.

Deterministic checks

Broken imports

An @path import or a linked file that doesn't exist on the pull request branch. In our testing with Claude Code 2.1.293, a broken import produced no warning.

Missing files

Scripts, docs and directories named in the context that were moved or deleted.

Hidden Unicode

Zero-width and bidirectional control characters that can hide instructions from a human reviewer.

Likely secrets

Keys and tokens pasted into context, which every agent session would read and send to its model provider.

Invalid settings

Settings files and .mcp.json that don't parse or don't match their schema, and rule frontmatter that doesn't parse, which turns a scoped rule into an always-on one.

Size limits per agent

Context bigger than a given agent will load. Codex stops adding AGENTS.md content at 32 KiB by default; Claude Code skips a CLAUDE.md file over 4 MiB and warns when files pass its recommended length. Each agent you run will be checked against its own limits.

A failure in any of these is designed to block the merge, because it gives the same answer on every run.

Planned: checks that need a model. Some problems, like two instruction files that contradict each other, need judgment. These checks will use the Claude API, and every finding must quote the files verbatim; a finding whose quote isn't found in the file is dropped. They are reported, never blocking.

2. Regression run

The same tasks, on both sides, several times each.

Rerunner will check out the base branch and the pull request side by side and run each task in your suite on both: same agent version, same pinned model, same settings, same starting files. The only difference between the two sides is the change under review. Each run will start in a fresh, disposable environment.

Agents aren't deterministic. The same task can pass on one run and fail on the next. So each task will run several times per side (three by default), and Rerunner will compare how often it passed, not a single result. Each run will be scored by the task's own check: a command, such as your tests, that exits 0 on success.

For every run, Rerunner will also record which instruction files loaded and why, which tools the agent called, and the run's turns, tokens and cost. With Claude Code those come straight from the agent; see Built on Claude.

Verdicts

Four answers, including “we can't tell yet”.

PASS

No difference: the pull request passes as often as base. If a task fails on both sides, the report says so, because the task or its check probably needs attention.

REGRESSED

The pull request passes less often than base, and a one-sided Fisher exact test (a standard way to check whether a difference in counts could be chance) on the pass counts gives p < 0.10.

IMPROVED

The pull request passes more often than base, by the same standard.

INCONCLUSIVE

The results differ, but not clearly enough to say the change caused it. More runs usually settle it. A task where the agent errored on every run, on both sides, is also inconclusive.

Two worked examples

The test asks: if the change made no difference, how likely is a split at least this lopsided? Take all the runs of a task together, count the passes, and ask how often chance alone would put that few of them on the pull request's side.

account-export six runs per side
main 6/6
PR 1/6

7 of the 12 runs passed. If the change made no difference, the chance that at most 1 of those 7 passes falls on the pull request's side is:

C(6,1)·C(6,6) / C(12,7) = 6 / 792 ≈ 0.008

p ≈ 0.008, below 0.10REGRESSED

Illustration — not captured output. Illustrative counts.
date-parsing six runs per side
main 4/6
PR 2/6

6 of the 12 runs passed. The chance that 2 or fewer of those 6 passes fall on the pull request's side is:

(1 + 36 + 225) / C(12,6) = 262 / 924 ≈ 0.28

p ≈ 0.28, not below 0.10INCONCLUSIVE

Illustration — not captured output. Illustrative counts.

Why p < 0.10, and what three runs can tell you

Statistical results never block a merge, so the threshold only decides what gets flagged for a human. With the default three runs per side, only a complete flip clears it: 3 of 3 against 0 of 3 gives p = 0.05, while 3 of 3 against 1 of 3 gives p = 0.20 and is reported as inconclusive. For finer distinctions, raise the run count for the tasks that matter most. Cost grows with it.

What blocks a merge

Only failures that would happen on every run.

  • A context check fails: a broken import, hidden Unicode, a likely secret, an invalid settings or .mcp.json file, or context over an agent's size limit.
  • An expected file stops loading: a file listed in the task's expect_loaded loaded in every base run and in no pull request run.
  • A forbidden file loads: a file listed in the task's must_not_load loads in a pull request run.

Differences in pass rates are statistical. Rerunner is designed to report them with their counts, mark them clearly in the pull request comment, and leave the decision to your reviewers. We'd rather say “we can't tell yet” than fail a build on noise.

3. Root cause Roadmap

Which edit did it?

A pull request often changes several pieces of context at once: a paragraph in CLAUDE.md, a skill, an MCP server's configuration. When a task regresses, you need to know which one did it.

Root cause analysis will restore each changed piece to its base version, one at a time, and re-run the regressed task. If restoring one piece brings the old behavior back, that piece is the likely cause, and the report will say so. If no single piece does, the report will say that too.

Root cause account-export, PR #482
Restored to baseRe-runResult
CLAUDE.md §Auth5/6Likely cause
docs/testing.md rename1/6No effect
release skill1/6No effect
Illustration of a planned feature — not captured output.

Suites

A short TOML file in your repository. Draft format: field names may change before the first stable release.

Each task has the prompt the agent receives and a check that decides whether it succeeded. Write tasks from work your agent really does: small, recent, and with a check you already trust.

# rerunner.toml (draft format)

model = "sonnet"        # pin the model so both sides are comparable
runs = 3                # runs per task, on base and on the pull request
max_budget_usd = 0.50   # cost guard per run

[[task]]
id = "account-export"
prompt = """
Add GET /account/export. It returns the signed-in user's data as JSON.
Follow the API conventions in this repository.
"""
check = "npm test -- account && npm run lint:auth"   # exit code 0 means the task passed
expect_loaded = ["CLAUDE.md", ".claude/rules/api.md"]
timeout_s = 900
max_turns = 20

[[task]]
id = "invoice-pdf"
prompt = "Add a PDF download for invoices at GET /invoices/:id/pdf, with tests."
check = "npm test -- invoices"
must_not_load = [".claude/rules/legacy-*.md"]   # retired rules must never load again
Draft format. An example suite, not a captured file.
FieldWhereWhat it does
modelsuiteThe model both sides run with. Pinning it keeps the comparison fair.
runssuiteRuns per task, per side. Defaults to 3.
max_budget_usdsuiteA cost guard for each run.
[[task]]suiteOne block per task.
idtaskThe task's name in reports.
prompttaskWhat the agent is asked to do.
checktaskA shell command that decides success. Exit code 0 means the run passed.
expect_loadedtaskInstruction files that must load. One that loads on every base run and on no pull request run blocks the merge.
must_not_loadtaskFiles that must never load. One that loads on the pull request blocks the merge.
timeout_staskTime limit for each run, in seconds.
max_turnstaskLimit on the agent's turns in each run.

Suites mined from your pull requests Roadmap

Writing tasks by hand is the slowest part. Rerunner will propose tasks from your own merged pull requests: the prompt from what the pull request did, the check from the tests it added, kept only if the check fails before the merge and passes after it. You review every proposed task before it joins the suite.

In your CI

A check designed to run where your code already gets tested, on GitHub or GitLab.

Rerunner is designed to run as a check in your own CI, on your runners. It will run on pull requests that change the agent's context: CLAUDE.md and AGENTS.md files, .claude/, skills, .mcp.json, settings, the pinned model, or the suite itself. Pull requests that don't touch any of these won't trigger a run, so they cost nothing.

The report will arrive as one pull request comment, updated on every push, with a pass or fail status you can make required.

Rerunner check on pull request #482

Agent regression report

Base
main at 3f2c1a0
Head
9be71d4
Agent
Claude Code, pinned model
Runs
6 per task, per side

Context diff

  • CLAUDE.md +3 −9 lines
  • docs/testing.md renamed to docs/guides/testing.md
  • .claude/skills/release/SKILL.md added
  • Blocks merge. CLAUDE.md:8 imports @docs/testing.md, which no longer exists on this branch. This pull request renamed it to docs/guides/testing.md. In our testing with Claude Code 2.1.293, a missing import produced no warning.
  • No hidden Unicode or likely secrets. Settings and .mcp.json are valid. Context is within Claude Code's size limits.

Regression run

TaskmainPRVerdict
account-export6/61/6REGRESSED
invoice-pdf6/66/6PASS
changelog-entry2/66/6IMPROVED
date-parsing4/62/6INCONCLUSIVE

Merge blocked by one deterministic check: the missing import. account-export regressed (6 of 6 against 1 of 6, one-sided Fisher exact test p ≈ 0.008). That difference is statistical, so it's reported for your reviewers, not enforced. date-parsing is inconclusive at six runs per side.

Illustration — not captured output. With root cause analysis (roadmap), the comment would also name the edit behind the regression.

It will never run automatically on pull requests from forks: a fork can change what the agent executes, so a maintainer starts those runs by hand, after review, in a sandbox.

Every run will call your model provider under your account, so cost grows with tasks × runs × two sides. Give Rerunner a dedicated API key with a spending limit, and run it on a fresh runner with no production secrets. The security page has the details.

Agents

Rerunner starts with Claude Code, the agent where we can see exactly which instruction files loaded and why. Codex, Cursor, Gemini CLI and OpenCode are on the roadmap; for them we'll rely on each tool's documented loading rules and its headless output.

It's vendor-neutral. We don't make a coding agent, and the same suite is meant to run against whichever agents your team uses.

Test-drive your agent's changes before they ship.

We're looking for a few teams to try Rerunner on one real repository and tell us where it's wrong.