Questions, answered honestly
If yours isn't here, write to hello@rerunner.dev.
Frequently asked
Select a question to open its answer.
Isn't this what claude plugin eval does?
No, and the two fit together. claude plugin eval tests a plugin or a skill on its own, comparing runs with and without it in an empty working directory. By design it doesn't load the project's CLAUDE.md, .claude/ directory or .mcp.json. Plugin evals docs
Rerunner is designed to test that project setup: the same tasks on the base branch and on the pull request, inside your repository, on every pull request that changes it. If you build a plugin, plugin eval is for the plugin, and Rerunner is meant for the repositories that use it. More on how we fit alongside Anthropic's tools.
Why not just evaluate the model?
Model benchmarks tell you how a model does on someone else's tasks with someone else's context. Your agent's behavior also depends on your CLAUDE.md, rules, skills, tools and settings, and those change far more often than the model does.
Research has found that context files can lower success or raise cost, and that the right context can help a great deal. Which one you get depends on the repository. See the evidence. The only way to know what a change does to your agent is to run your agent on your tasks.
Won't LLM randomness make CI flaky?
That's the main thing the design guards against. Statistical differences never block a merge. Each task runs several times per side, and a difference is only called REGRESSED or IMPROVED when a one-sided Fisher exact test gives p < 0.10. Otherwise it's INCONCLUSIVE, and the report says so.
Only deterministic failures block: a broken import, an expected rule that loaded in every base run and in no pull request run, a forbidden file that loads. Those happen on every run, so they don't flake. How verdicts work
What does a run cost?
It depends on your tasks. Cost is roughly tasks × runs per task × two sides, and each run costs about what that task costs your agent today. Runs will only happen on pull requests that change the agent's setup, prompt caching helps with repeated runs, and every run's cost will be reported.
Use a dedicated API key with a spending limit, plus the suite's per-run cost guard. We'll publish real figures once pilots give us some, not before.
Which agents will it support?
Claude Code first. Codex, Cursor, Gemini CLI and OpenCode are on the roadmap. Claude Code reports which instruction files loaded and why, which makes deterministic checks possible. For the others we'll rely on each tool's documented loading rules and its headless output.
Can it test a model upgrade?
That's one of the cases it's designed for. The pinned model is a setting in your repository, so changing it is a pull request like any other: the suite would run on both models, task by task, with cost per successful task on each side. See the model upgrade scenario.
Do you see our code?
No. Rerunner is designed to run in your CI, on your runners, and to send nothing to us. There is no hosted service. The agent itself sends prompts and repository content to its model provider under your account, as it does when you use it directly. Data handling
Is the source code available?
Not at this stage. Rerunner is closed source for now. It will be delivered as a check that runs in your own CI, and design partners will see what it runs and what it reports, in their own logs.
What stage are you at?
In development, pre-revenue, and looking for design partners. We have no customers yet. Nothing on this site is a customer quote or a benchmark result, and the sample reports are illustrations of what Rerunner is designed to produce.
How is this different from eval tools like promptfoo or Braintrust?
They're general-purpose evaluation tools, built mostly for prompts and LLM applications, and they're good at it. promptfoo even has a guide to evaluating coding agents. With enough glue, you could assemble something like Rerunner from them.
Rerunner is narrower, on purpose. The system under test is a coding agent in your repository, and the comparison is always base against pull request, on the same tasks. It's designed to know which instruction files, rules, skills and MCP servers the agent loaded, and why, and its merge gate will block only on deterministic failures, with an explicit “inconclusive” instead of a score.
How is this different from Stet?
Stet is the closest tool we know of, and it's worth a look. It turns your merged pull requests into tasks and scores them with your tests, plus measures such as review quality, footprint and cost. It compares setups across Claude Code, Codex and Cursor, including your current AGENTS.md against a candidate change, and recommends what to do, such as promote, hold or roll back. stet.sh, stet.sh/why
Rerunner is designed for a different job: a merge gate on every pull request that changes the agent's setup, base against pull request, with several runs per task on each side. It will block only on deterministic failures, and say “inconclusive” when a difference could be chance. Stet already covers more agents; Rerunner is in development and starts with Claude Code.
Will it work with GitHub and GitLab?
It's designed to run as a check in your own CI, with your code on GitHub or GitLab, and to post its report as a comment on the pull request or merge request.
Do we have to write the tasks ourselves?
In the first version, yes: a short TOML file with a prompt and a check for each task. Start with three to five tasks your agent does often. Later, Rerunner will propose tasks mined from your merged pull requests, for you to review. Suite format
How is this different from a linter, or /doctor prompt-audit?
Static checks read files. They catch broken paths, stale references and contradictions, and Rerunner's context diff is designed to run deterministic checks like these too. But reading can't tell you how the agent will behave, so Rerunner will also run the agent on real tasks, on both sides of the change.
Does it replace code review?
No. Your reviewers still decide. Rerunner is meant to give them evidence about what a change to the agent's setup does, the way tests do for code.
Are you affiliated with Anthropic?
No. Rerunner is an independent company. We're not affiliated with, partnered with or endorsed by Anthropic. We build on Claude Code's public, documented interfaces. Where Claude fits
Test-drive your agent's changes before they ship.
We're looking for a few teams to try Rerunner on one real repository and tell us where it's wrong.