The files that steer your agent change every week.
Companies now let AI coding agents write code. Each agent follows a written handbook the team keeps for it: which rules to follow, which tools to use, which model runs. Teams edit that handbook all the time, and one small edit can quietly change how the agent works. It might stop using the secure login service, for example. Nobody notices, because normal tests check the code, not the agent, and the edit looks harmless in review. The cost shows up later, as a security hole, a broken feature or a bigger bill.
The rest of this page is the detail. A coding agent's behavior depends on its model and on everything it loads: instructions, rules, skills, tools and settings. All of it changes in ordinary pull requests, and most of it is reviewed by reading. Below: what changes, what the research says, and why the effect has to be measured in your own repository. Unfamiliar terms are in the glossary.
What changes
Six kinds of edits that can change what an agent does without changing a line of your product's code.
- Instruction files
CLAUDE.md (the instruction file Claude Code reads), AGENTS.md (a shared format that Codex and other agents read) and the files they import. Claude Code joins every CLAUDE.md from the filesystem root down to the working directory and loads the ones in subdirectories on demand. Adding a personal
CLAUDE.local.mdstops Claude from reading AGENTS.md. Claude Code docs: memory- Path-scoped rules
Rules are smaller instruction files for one part of the codebase. Rules in
.claude/rules/with apaths:glob (a file-name pattern) load only when the agent reads or edits a matching file. A glob that stops matching means the rule never loads. And if the YAML at the top of a rule doesn't parse, Claude Code loads the rule as if it had no paths at all, so a typo turns a scoped rule into one that is always on. Claude Code docs: memory- Skills
A skill is a packaged set of instructions for one kind of job. A skill's name and description sit in context in every session; its body loads only when the skill is used. Editing one description changes when the agent reaches for that skill, and how much context each task carries. Claude Code docs: skills
- MCP tools and config
MCP (Model Context Protocol) is a standard way to connect an agent to outside tools, such as a database or a ticket tracker. Tool descriptions decide whether and how an agent calls a tool.
.mcp.jsondecides which servers start at all. A server that fails to connect shows up in the session's startup event, not as a failing test. Claude Code docs: headless- Settings and the model
The pinned model, permissions and hooks (commands that run automatically at set moments in a session). A model upgrade changes everything at once: some tasks get better, others get worse, and instructions written for the old model may now work against you.
- The agent itself
Agents update often, and their loading rules change with them. Claude Code's setting for reading AGENTS.md, for example, requires version 2.1.277 or later. Claude Code docs: memory
What the evidence says
Context can help a lot, do nothing, or hurt. Its cost moves on its own.
-
Context files are not a free win. Across the benchmark in a 2026 study, providing them did not generally improve task success, and raised inference cost by over 20% on average. This held for files written by developers and for files generated by a model, and agents took more steps with them. The authors recommend that any gains be “rigorously validated”.
-
But the right context can matter a great deal. In Vercel's evaluations on Next.js 16 APIs that were absent from model training data, an 8 KB AGENTS.md documentation index took the pass rate from 53% to 100%, while skills with default settings left it at 53%.
Source: Vercel blog, January 2026
-
Longer input is less reliable input. Chroma tested 18 models and found that performance degrades as input length grows, even on simple tasks, and in non-uniform ways.
Source: Chroma, Context Rot, July 2025
-
Fewer tokens is not a smaller bill. In a 2026 study run with Claude Code, the largest compression setup cut tool-output tokens by 38.4% and raised the billed cost by 6.8%; cache writes and reads dominated input-side cost. The authors recommend measuring cost per successful task.
Source: Token Reduction Is Not Cost Reduction, arXiv 2607.12161
Put together: you can't tell from a diff whether a change to your agent's context will help, hurt or just cost more. The answer depends on your repository, your tasks and your agent. You find it by running them.
Failures you don't see
Documented and reported cases where context fails without an error.
-
A broken import fails silently. In our own testing with Claude Code 2.1.293, a CLAUDE.md import pointing at a file that doesn't exist produced no warning in headless mode. The file simply wasn't loaded.
Observed by us on 8 October 2026. Import behavior is described in Claude Code docs: memory.
-
Codex stops adding AGENTS.md files once their combined size reaches 32 KiB by default, and an open issue reports that this happens with no warning.
-
A personal note can switch off shared instructions. A reported issue reproduced that adding
CLAUDE.local.mdturns off the AGENTS.md fallback, dropping a project's shared instructions for that developer.Source: anthropics/claude-code issue 96117 (September 2026)
-
Subagents may not see your rules. An issue reported that Task subagents did not load the project's CLAUDE.md or
.claude/rules/, so project configuration was ignored without notice.Source: anthropics/claude-code issue 29423 (February 2026)
-
Rules files can hide instructions from reviewers. The “Rules File Backdoor” technique used zero-width and bidirectional Unicode characters to hide instructions in Cursor and Copilot rules files.
Source: The Hacker News, March 2025
None of these makes a unit test fail. Some are deterministic, and a static check can catch them before any agent runs. Others only show up in what the agent does.
More than one agent
One repository, several readers of the same files.
Most organisations don't run a single coding agent. In GitLab's 2026 AI Accountability Report, 91% of organisations had at least two AI coding tools in active use. IT Brief
Each tool reads the repository by its own rules. By default Claude Code reads AGENTS.md only when there is no CLAUDE.md, and doesn't read AGENTS.override.md at all. Codex reads AGENTS.override.md first in each directory and caps the combined size. The same AGENTS.md can steer two agents differently, and an edit that helps one can hurt the other. Claude Code docs, Codex docs
Meanwhile, these files rarely get the review that code gets. As Mitch Ashley of Futurum put it, they “sit outside the review path that already covers pipeline configuration”. DevOps.com, September 2026
What follows
Treat your agent's setup like code. Diff it, check what can be checked deterministically, and run the agent on real tasks before and after the change. Report what you measured, with counts, and say so when the difference could be noise.
That is what Rerunner is for.
Test-drive your agent's changes before they ship.
We're looking for a few teams to try Rerunner on one real repository and tell us where it's wrong.