Rerunner

Docs

The draft specification for the Rerunner check: what triggers a run, how a suite is written, how verdicts are decided, and how it fits into CI.

Draft spec. Field names and formats may change before the first stable release.

Overview

Rerunner is a CI check for changes to an agent setup: the instruction files, rules, skills, MCP tools, settings and model that decide how a coding agent behaves. On a pull request that changes any of them, it runs three layers, cheapest first:

  1. Context diff

    Resolves what the agent loads on main and on the pull request, shows the difference, and runs deterministic checks. Any failure here blocks the merge.

  2. Regression run

    Runs each task in the suite on both sides, several times each, in a fresh environment, and compares pass counts. Statistical differences are reported, not enforced.

  3. Root cause Roadmap

    Restores each changed piece of the setup one at a time and re-runs a regressed task to name the likely cause.

The result is one pull request comment and a check status. The product overview explains the reasoning; this page is the detail.

Suite format

A suite is a short TOML file in your repository. Each task has the prompt the agent receives and a check that decides whether the run succeeded. Write tasks from work your agent really does: small, recent, and with a check you already trust.

# rerunner.toml (draft format)

model = "sonnet"        # pin the model so both sides are comparable
runs = 3                # runs per task, on main and on the pull request
max_budget_usd = 0.50   # cost guard per run

[[task]]
id = "account-export"
prompt = """
Add GET /account/export. It returns the signed-in user's data as JSON.
Follow the API conventions in this repository.
"""
check = "npm test -- account && npm run lint:auth"   # exit code 0 means the run passed
expect_loaded = ["CLAUDE.md", ".claude/rules/api.md"]
timeout_s = 900
max_turns = 20

[[task]]
id = "invoice-pdf"
prompt = "Add a PDF download for invoices at GET /invoices/:id/pdf, with tests."
check = "npm test -- invoices"
must_not_load = [".claude/rules/legacy-*.md"]   # retired rules must never load again
An example suite in the draft format, not a captured file.
FieldWhereWhat it does
modelsuiteThe model both sides run with. Pinning it keeps the comparison fair.
runssuiteRuns per task, per side. Defaults to 3.
max_budget_usdsuiteA cost guard for each run.
[[task]]suiteOne block per task.
idtaskThe task's name in reports.
prompttaskWhat the agent is asked to do.
checktaskA shell command that decides success. Exit code 0 means the run passed.
expect_loadedtaskInstruction files that must load. One that loads on every main run and on no pull request run blocks the merge.
must_not_loadtaskFiles that must never load. One that loads on the pull request blocks the merge.
timeout_staskTime limit for each run, in seconds.
max_turnstaskLimit on the agent's turns in each run.

Mined suites Roadmap

Rerunner will propose tasks from your merged pull requests: the prompt from what the pull request did, the check from the tests it added, kept only if the check fails before the merge and passes after it. You review every proposed task.

Verdict rules

Each task gets one of four verdicts from its pass counts on main and on the pull request.

PASS

No difference: the pull request passes as often as main. A task that fails on both sides is reported, since the task or its check probably needs attention.

REGRESSED

The pull request passes less often than main, and a one-sided Fisher exact test on the pass counts gives p < 0.10.

IMPROVED

The pull request passes more often than main, by the same standard.

INCONCLUSIVE

The counts differ but not clearly enough to say the change caused it. A task where the agent errored on every run, on both sides, is also inconclusive.

Worked examples

The test asks: if the change made no difference, how likely is a split at least this lopsided? Take all the runs of a task together, count the passes, and ask how often chance alone would put that few of them on the pull request's side.

account-export six runs per side
main 6/6
PR 1/6

7 of 12 runs passed. The chance that at most 1 of those 7 lands on the PR side:

C(6,1)·C(6,6) / C(12,7) = 6 / 792 ≈ 0.008

p ≈ 0.008, below 0.10REGRESSED

Illustration — not captured output.
date-parsing six runs per side
main 4/6
PR 2/6

6 of 12 runs passed. The chance that 2 or fewer of those 6 land on the PR side:

(1 + 36 + 225) / C(12,6) = 262 / 924 ≈ 0.28

p ≈ 0.28, not below 0.10INCONCLUSIVE

Illustration — not captured output.

Why p < 0.10, and what three runs can tell you

Statistical results never block a merge, so the threshold only decides what gets flagged for a human. With the default three runs per side, only a complete flip clears it: 3 of 3 against 0 of 3 gives p = 0.05, while 3 of 3 against 1 of 3 gives p = 0.20 and is reported as inconclusive. For finer distinctions, raise the run count on the tasks that matter most; cost grows with it.

What blocks a merge

  • A deterministic context check fails: broken import or missing file, hidden Unicode, likely secret, invalid settings or .mcp.json, unparseable rule frontmatter, or context over an agent's size limit.
  • A file in expect_loaded loaded in every main run and in no pull request run.
  • A file in must_not_load loaded in a pull request run.

CI integration

Rerunner runs as a step in your pull request pipeline, on your own runners, with your code on GitHub or GitLab. There is no hosted service.

Trigger

Pull requests that change the agent setup or the suite: CLAUDE.md and AGENTS.md files and their imports, .claude/, skills, .mcp.json, settings, the pinned model, or the suite file. Other pull requests don't start a run.

Report

One pull request comment (merge request comment on GitLab), updated on every push: context diff, check results, pass counts and verdict per task, cost per side. Plus a check status you can make required.

Forks

Never runs automatically on pull requests from forks. A maintainer starts those runs by hand, after review, in a sandbox.

Credentials

A dedicated model API key with a spending limit, and a code host token with read access to contents and write access to pull request comments and checks.

Environment

A fresh, disposable runner or container per run, no production secrets, restricted network where your CI supports it. See sandbox guidance.

Cost

Every run calls your model provider under your account, so cost grows with tasks × runs × two sides. Each run's cost is reported; max_budget_usd caps a single run.

Agents and roadmap

Rerunner is vendor-neutral: it doesn't make a coding agent, and one suite serves whichever agents your team uses.

Claude Code

First supported agent. It reports which instruction files it loaded and why, lists its skills and MCP servers at startup, and reports cost per run, which is what makes the deterministic checks and expect_loaded possible. Details on Built on Claude.

Codex, Cursor, Gemini CLI, OpenCode Roadmap

For agents without a loading event, Rerunner will rely on each tool's documented loading rules and its headless output. Size limits are checked per agent: Codex stops adding AGENTS.md content at 32 KiB by default; Claude Code skips a CLAUDE.md over 4 MiB and warns past its recommended length.

Codex docs: AGENTS.md · Claude Code docs: memory

The full sequence is on the roadmap.

Glossary

Coding agent

An AI program that carries out programming tasks on its own: reads the codebase, edits files, runs commands and tests. Claude Code, Codex, Cursor and Gemini CLI are examples.

Agent setup

Everything besides the task that decides how an agent behaves: instruction files, rules, skills, MCP tools, hooks, settings and the pinned model, kept in the repository and changed through pull requests.

Instruction file CLAUDE.md, AGENTS.md

A plain text file that tells a coding agent how the team works. CLAUDE.md is read by Claude Code; AGENTS.md is a shared format read by Codex and others.

Rules

Smaller instruction files for one part of the codebase. Some load only when the agent touches matching files, so a moved file or a mistyped pattern can stop a rule loading unnoticed.

Skills

Packaged instructions for one kind of job. The agent always sees each skill's short description and reads the full skill only when it decides it needs it.

MCP tool

A tool the agent can use beyond the codebase, such as a database or a ticket tracker, connected through the Model Context Protocol. The agent decides when to call it largely from its description.

Context

Everything the agent reads before and while it works. A context diff lists what the agent sees differently after a change.

Regression

A task the agent used to complete reliably that it now completes less often.

CI

Continuous integration: the automatic checks that run on every pull request before it merges. Rerunner is one of those checks, for the agent instead of the code.

Try it on one repository.

We're looking for a few design partners to run Rerunner on a real repository and tell us where it's wrong.