Rerunner Early access

For the people who own your agents' setup

Someone in your company maintains the files every coding agent reads, even if nobody gave them the job. Here is what Rerunner is designed to give each of them, and the moment they'll want it. Unfamiliar terms are explained in the glossary.

Platform and DevEx teams

Where you are

You maintain the shared CLAUDE.md, rules, skills and MCP configuration that every engineer's agent depends on, often across many repositories. A change you make lands everywhere at once.

“We're about to change the shared rules, or move everyone to a new model. What breaks?”

What you get

  • A check in each repository's CI that runs that repository's own tasks on the change, base against pull request.
  • Deterministic blocks when a rule stops loading where a task expects it.
  • Pass rates and cost per successful task for model upgrades, task by task.
  • One suite format, whichever agents your engineers use, as more agents are supported.

AI enablement leads

Where you are

You roll out coding agents across teams, write internal skills and decide which MCP servers people should use. Leadership wants to know whether it's working and what it costs.

“Does this new skill actually help, or does it just make every task more expensive?”

What you get

  • A before and after for every change to skills, tools and instructions, on real tasks from real repositories.
  • Cost, tokens and turns per task, next to pass rates, so “better” and “cheaper” are both measured.
  • Honest results you can share: counts, not scores, and “inconclusive” when the evidence is thin.

Staff engineers and tech leads

Where you are

You own your repository's CLAUDE.md and AGENTS.md, and you review the pull requests that change them. Reading them tells you what they say, not what the agent will do.

“This pull request trims CLAUDE.md by a third. Is anything important gone?”

What you get

  • A context diff in the pull request: which files the agent will load differently, and when.
  • Broken imports, missing files and invalid settings caught before anyone merges.
  • A suite of your own tasks in a short TOML file you can read and review like code.

AppSec and security teams

Where you are

Rules files, hooks and MCP configuration decide what an agent reads and runs. They change in pull requests and are part of your software supply chain, whether or not anyone treats them that way.

“A pull request adds an MCP server and edits a rules file. What exactly will the agent see and run?”

What you get

  • Checks for hidden Unicode and likely secrets in agent context, on every change.
  • A clear record of which servers, hooks and files each run loaded.
  • A design that runs in your own CI, never runs fork pull requests automatically, and expects a sandbox. See Security.

Not a fit yet

We'd rather tell you now.

  • You mainly use an agent other than Claude Code. Rerunner is being built for Claude Code first. Codex, Cursor, Gemini CLI and OpenCode are on the roadmap, for later.
  • You want a hosted service or a dashboard. Rerunner is designed to run in your own CI and report on the pull request. There is no hosted service.
  • You don't edit your agents' setup. If CLAUDE.md, rules and skills rarely change, there's little to test yet.
  • You're evaluating a chatbot or an LLM app. General-purpose eval tools fit that better. Rerunner is for coding agents working in a repository.

Test-drive your agent's changes before they ship.

We're looking for a few teams to try Rerunner on one real repository and tell us where it's wrong.