All posts

4 min read

Is a Codebase With 500 .md Files Good or Bad for AI?

I saw this post from Cory House on X:

Cory House
@housecor

Just joined a team with nearly 500 .md files for AI.

Nearly half a million words.

Just reading all the AI related instructions, skills, hooks, etc would take me about 30 hours!

I keep seeing teams with elaborate setups like this.

I'm increasingly skeptical it's worthwhile.

Oct 7, 2026

My first thought was, who keeps all of this up to date? My guess is that a lot of these files were generated from the codebase in the first place. But the code keeps changing, and the files usually don't.

I had my own experience with this on a recent project, but I also wanted to check what the industry says. So does more markdown really make your AI better?

An AI agent reading from the codebase and a small CLAUDE.md, while 500 markdown files that get out of date are blocked

What Did the Research Find?

I found a study from ETH Zurich and LogicStar.ai that tested exactly this. They ran coding agents like Claude Code, Codex and Qwen Code on 438 real GitHub issues, with and without a context file like AGENTS.md or CLAUDE.md. For 138 of them, the repo already had a context file written by its own developers.

The result surprised me. The context files did not generally help the agents solve more tasks, but they made the runs cost over 20% more on average.

Why Would More Context Make It Worse?

The same study found that agents really do what the file says. When the file mentioned the uv tool, agents used it 1.6 times per task on average. When it didn't, they almost never used it.

That sounds great until the file is wrong. If the AI follows your instructions this well, it will follow the old ones too.

I found this out in a recent project. A rule I added early on was still sitting in a markdown file, and the AI kept following it after it stopped being true. Now I keep temporary rules like that in the code instead, where they change with everything else.

There is also a limit to how much an AI can keep in mind. The Claude Code docs say it directly:

Bloated CLAUDE.md files cause Claude to ignore your actual instructions!

Another study gave models up to 500 simple instructions at once, and even the best models followed only about 68% of them.

So What Should Go in the Markdown?

This is where the research matches what I already believed: the source of truth should always be the codebase. The AI can read, search and run your code, so you don't need to copy it into markdown.

The markdown should only hold what the code can't tell the AI:

It should skip what the AI can find on its own, like project overviews and copies of code. Anthropic's list of what to leave out of CLAUDE.md even includes "Anything Claude can figure out by reading code".

To be fair, the AI doesn't load all 500 files every time, since skills and nested AGENTS.md files only load when they're needed. But someone still has to keep all of them correct, and if it takes 30 hours just to read them, I doubt anyone is doing that.

What I Would Do

If you already have a lot of AI files, ask this question from the Claude Code docs for every line:

Would removing this cause Claude to make mistakes?

If the answer is no, cut it. And if a rule must happen every single time, make it a hook, a lint rule or a test instead, so it actually gets checked.


If you want to read more, the study is here, Anthropic's tips for CLAUDE.md are here, the 500 instructions study is here, and Michael Nygard's post on decision records is here.


See all posts