Rerunner Early access

The files that steer your agent change every week.

Companies now let AI coding agents write code. Each agent follows a written handbook the team keeps for it: which rules to follow, which tools to use, which model runs. Teams edit that handbook all the time, and one small edit can quietly change how the agent works. It might stop using the secure login service, for example. Nobody notices, because normal tests check the code, not the agent, and the edit looks harmless in review. The cost shows up later, as a security hole, a broken feature or a bigger bill.

The rest of this page is the detail. A coding agent's behavior depends on its model and on everything it loads: instructions, rules, skills, tools and settings. All of it changes in ordinary pull requests, and most of it is reviewed by reading. Below: what changes, what the research says, and why the effect has to be measured in your own repository. Unfamiliar terms are in the glossary.

What changes

Six kinds of edits that can change what an agent does without changing a line of your product's code.

Instruction files

CLAUDE.md (the instruction file Claude Code reads), AGENTS.md (a shared format that Codex and other agents read) and the files they import. Claude Code joins every CLAUDE.md from the filesystem root down to the working directory and loads the ones in subdirectories on demand. Adding a personal CLAUDE.local.md stops Claude from reading AGENTS.md. Claude Code docs: memory

Path-scoped rules

Rules are smaller instruction files for one part of the codebase. Rules in .claude/rules/ with a paths: glob (a file-name pattern) load only when the agent reads or edits a matching file. A glob that stops matching means the rule never loads. And if the YAML at the top of a rule doesn't parse, Claude Code loads the rule as if it had no paths at all, so a typo turns a scoped rule into one that is always on. Claude Code docs: memory

Skills

A skill is a packaged set of instructions for one kind of job. A skill's name and description sit in context in every session; its body loads only when the skill is used. Editing one description changes when the agent reaches for that skill, and how much context each task carries. Claude Code docs: skills

MCP tools and config

MCP (Model Context Protocol) is a standard way to connect an agent to outside tools, such as a database or a ticket tracker. Tool descriptions decide whether and how an agent calls a tool. .mcp.json decides which servers start at all. A server that fails to connect shows up in the session's startup event, not as a failing test. Claude Code docs: headless

Settings and the model

The pinned model, permissions and hooks (commands that run automatically at set moments in a session). A model upgrade changes everything at once: some tasks get better, others get worse, and instructions written for the old model may now work against you.

The agent itself

Agents update often, and their loading rules change with them. Claude Code's setting for reading AGENTS.md, for example, requires version 2.1.277 or later. Claude Code docs: memory

What the evidence says

Context can help a lot, do nothing, or hurt. Its cost moves on its own.

  • Context files are not a free win. Across the benchmark in a 2026 study, providing them did not generally improve task success, and raised inference cost by over 20% on average. This held for files written by developers and for files generated by a model, and agents took more steps with them. The authors recommend that any gains be “rigorously validated”.

    Source: Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?, arXiv 2602.11988

  • But the right context can matter a great deal. In Vercel's evaluations on Next.js 16 APIs that were absent from model training data, an 8 KB AGENTS.md documentation index took the pass rate from 53% to 100%, while skills with default settings left it at 53%.

    Source: Vercel blog, January 2026

  • Longer input is less reliable input. Chroma tested 18 models and found that performance degrades as input length grows, even on simple tasks, and in non-uniform ways.

    Source: Chroma, Context Rot, July 2025

  • Fewer tokens is not a smaller bill. In a 2026 study run with Claude Code, the largest compression setup cut tool-output tokens by 38.4% and raised the billed cost by 6.8%; cache writes and reads dominated input-side cost. The authors recommend measuring cost per successful task.

    Source: Token Reduction Is Not Cost Reduction, arXiv 2607.12161

Put together: you can't tell from a diff whether a change to your agent's context will help, hurt or just cost more. The answer depends on your repository, your tasks and your agent. You find it by running them.

Failures you don't see

Documented and reported cases where context fails without an error.

  • A broken import fails silently. In our own testing with Claude Code 2.1.293, a CLAUDE.md import pointing at a file that doesn't exist produced no warning in headless mode. The file simply wasn't loaded.

    Observed by us on 8 October 2026. Import behavior is described in Claude Code docs: memory.

  • Codex stops adding AGENTS.md files once their combined size reaches 32 KiB by default, and an open issue reports that this happens with no warning.

    Sources: Codex docs: AGENTS.md, openai/codex issue 13386

  • A personal note can switch off shared instructions. A reported issue reproduced that adding CLAUDE.local.md turns off the AGENTS.md fallback, dropping a project's shared instructions for that developer.

    Source: anthropics/claude-code issue 96117 (September 2026)

  • Subagents may not see your rules. An issue reported that Task subagents did not load the project's CLAUDE.md or .claude/rules/, so project configuration was ignored without notice.

    Source: anthropics/claude-code issue 29423 (February 2026)

  • Rules files can hide instructions from reviewers. The “Rules File Backdoor” technique used zero-width and bidirectional Unicode characters to hide instructions in Cursor and Copilot rules files.

    Source: The Hacker News, March 2025

None of these makes a unit test fail. Some are deterministic, and a static check can catch them before any agent runs. Others only show up in what the agent does.

More than one agent

One repository, several readers of the same files.

Most organisations don't run a single coding agent. In GitLab's 2026 AI Accountability Report, 91% of organisations had at least two AI coding tools in active use. IT Brief

Each tool reads the repository by its own rules. By default Claude Code reads AGENTS.md only when there is no CLAUDE.md, and doesn't read AGENTS.override.md at all. Codex reads AGENTS.override.md first in each directory and caps the combined size. The same AGENTS.md can steer two agents differently, and an edit that helps one can hurt the other. Claude Code docs, Codex docs

Meanwhile, these files rarely get the review that code gets. As Mitch Ashley of Futurum put it, they “sit outside the review path that already covers pipeline configuration”. DevOps.com, September 2026

What follows

Treat your agent's setup like code. Diff it, check what can be checked deterministically, and run the agent on real tasks before and after the change. Report what you measured, with counts, and say so when the difference could be noise.

That is what Rerunner is for.

How it works

Read the essay: why agent changes need tests

Test-drive your agent's changes before they ship.

We're looking for a few teams to try Rerunner on one real repository and tell us where it's wrong.