Stop Blaming the LLM: Why Your Architecture Is Making Your Coding Agents Write Slop
Your AI coding agent isn't hallucinating because the model is broken. It's writing slop because your codebase relies on tribal knowledge that dies outside the context window.

It usually happens around 1:00 AM. You sit at a workstation in Gbagada or a quiet desk in Akure, running Cursor or Claude Code against a legacy frontend module. You type a deceptively simple instruction: "Add a reorder button to the order details card and wire it to the checkout flow."
Thirty seconds later, the diff lands. The code compiles. On the surface, it looks clean.
Then you inspect the pull request.
Instead of consuming the date-formatting utility your team wrote three months ago, the agent invented its own string parser. Instead of using the shared UI toast system, it pulled in an unvetted npm package. It ignored your bespoke state-management conventions and wired up an ad-hoc local store that breaks hydration.
The gut reaction—heard in Slack channels from San Francisco to Lagos—is to blame the model. Model A is hallucinating again. Model B is too expensive. The context window is too small. We need better system prompts.
Software engineer Kayra Berk Tuncer recently pulled back the curtain on this exact frustration. His conclusion cuts straight through the developer cope: the model is fine; your architecture is broken.
The interesting thing about this story is not merely that AI coding agents write messy pull requests. It is actually that software architecture is undergoing its first fundamental inversion in two decades: we are no longer structuring codebases purely for human cognitive load. We are now forced to architect them for context-window economics.
The Architecture Problem Nobody Wants to Admit
For decades, software architecture was designed to solve one specific bottleneck: the messy, forgetful human brain. We built patterns like Model-View-Controller, separated concerns, and wrote documentation because human engineers cannot hold 200,000 lines of code in working memory.
When you—the human who wrote or inherited the system—sit down to build a feature, you rely on an invisible asset: tribal context. You know instinctively that shared UI buttons live in src/components/ui, that auth tokens are intercepted at the network boundary, and that utils/date.ts holds the only acceptable formatting helper.
An LLM has zero tribal context. Every time an agent spins up, it is an amnesiac contractor arriving on day one with a 128k-token attention span.
If your codebase is a sprawling, loosely coupled monolith where business logic leaks across folder boundaries, the agent faces an impossible trade-off:
- Feed it the entire repository: The context window floods with irrelevant CSS tokens, configuration files, and adjacent modules. Reasoning degrades, latency spikes, token costs explode, and the model drowns in noise.
- Constrain its context to the target folder: The model cannot see your shared utilities or internal UI libraries, so it hallucinates fresh implementations or imports external dependencies to fill the vacuum.
Prompt engineering cannot fix a structural leak. If an agent needs to read sixteen unrelated files across four directories just to render a button, your system design is actively hostile to autonomous tooling.
The Mechanics of "Agent-Native" Monorepos
Tuncer’s proposed remediation points toward a disciplined structural pattern: monorepos functioning as deterministic context management systems.
By leveraging tools like Turborepo with pnpm workspaces, an engineering team can isolate micro frontends while enforcing mechanical boundaries:
repo/
├── apps/
│ ├── shell/ # Thin orchestration host
│ ├── module-orders/ # Isolated domain module
│ └── module-catalog/ # Isolated domain module
├── packages/
│ ├── ui/ # Shared design system components
│ ├── utils/ # Pure, deterministic helper functions
│ ├── analytics/ # Single source of truth for tracking
│ └── config/ # ESLint, Tailwind, TypeScript presets
├── turbo.json
└── package.json
When an agent is tasked with refactoring module-orders, you don't expose the entire codebase to its context scraper. The workspace boundaries explicitly define what exists.
Because dependencies are managed via explicit workspace protocols (workspace:*), the agent only consumes what it imports. It doesn't crawl unrelated apps; it references deterministic contracts. The context stays lean. The reasoning capacity remains sharp. The agent stops hallucinating duplicate utilities because the shared utility package is an explicitly defined, bounded dependency.
The Short Answer
Your coding agent isn't producing garbage because the LLM is stupid; it's producing garbage because your codebase relies on implicit assumptions instead of explicit structural boundaries. If your repository cannot be parsed in modular, self-contained slices without pulling in hundreds of kilobytes of irrelevant context, an autonomous agent will inevitably invent abstractions, duplicate logic, and bloat your dependencies.
What Is Really Happening
We are witnessing the death of implicit developer architecture.
For the past ten years, messy startups got away with "vibes-based" directory structures. A senior engineer kept the system coherent through PR reviews and institutional memory. But as teams integrate Claude Code, Cursor, and background coding bots to handle tickets, that institutional memory is useless.
LLMs consume context linearly and pay an attention penalty for every unnecessary token they ingest. When you give an agent a loose architecture, you are asking an amnesiac to guess your unspoken rules. The result is token burn, broken dependencies, and PRs that take longer to review than to write from scratch.
The Assumption I'd Challenge
The widespread assumption across engineering leadership is that waiting for larger context windows and smarter frontier models will solve agent slop.
I challenge that directly. Giving an LLM a 2-million-token context window does not fix architectural rot; it subsidizes lazy engineering at astronomical API costs.
More context does not equal better comprehension. Attention drift is a documented reality in transformer architectures: as context fills up, retrieval accuracy in the "middle" degrades. If you feed an agent a bloated codebase, the model's reasoning degrades regardless of the model's marketing claims. Architectural discipline beats token brute-forcing every single time.
The Strategic Options
Engineering teams facing degraded agent output have three options:
The Brute-Force Route (Status Quo): Keep the legacy directory structure. Write massive, 2,000-word system prompt instructions (
CLAUDE.mdor.cursorrules) begging the agent not to install random packages or write duplicate utils.- Trade-off: High token bills, brittle rules that break on complex tasks, and permanent cognitive fatigue during code review.
The Micro-Repository Split: Carve every domain into independent Git repositories to guarantee small context windows.
- Trade-off: Operational hell. Cross-repository dependency synchronization kills velocity. When an agent needs a utility from repo A inside repo B, you end up manually copying code or managing complex npm version bumps.
The Agent-Native Monorepo: Restructure the application into isolated workspace modules (Turborepo/pnpm) with strict dependency boundaries. Shared code lives in discrete packages; domain logic lives in self-contained apps.
- Trade-off: Requires upfront refactoring labor, but drops agent hallucinations and token consumption immediately.
My Recommendation
Choose the Agent-Native Monorepo architecture.
Treat context as a scarce operational resource. Stop viewing your repository as a digital junk drawer where files are grouped by arbitrary folder names (/components, /helpers, /services). Instead, treat every folder as a bounded system with explicit inputs and outputs.
If a shared function exists, it belongs in an internal package with clear TypeScript interfaces. If an agent cannot determine how your app formats currency by looking at an imported @repo/utils declaration, your architecture—not the model—has failed.
What I Would Do Next
If you are leading an engineering team that wants 10x output from coding agents without drowning in technical debt, take these steps by Monday:
- Audit Agent PR Failures: Pull your last 20 agent-generated PRs. Categorize the errors: Did it duplicate a function? Did it violate state management? Did it pull an unnecessary dependency? Identify the missing context that caused each mistake.
- Lock Down Dependency Installation: Strip your agents of the permission to run arbitrary
npm installcommands without human confirmation. Force them to solve problems using the packages already declared in the lockfile. - Isolate Shared Utilities into a Internal Package: If your helpers are scattered inside
src/utils/, pull them out into a dedicated workspace package (packages/utils) with strict export boundaries. Make it impossible for the agent to miss what already exists. - Draft Machine-Readable Domain Contracts: Instead of writing novels in your prompt files, provide concise, typed API contracts that define how a module interacts with the rest of the application.
What Would Change My Mind
I would reconsider this position if autonomous agents develop zero-shot global repo graph parsing that costs pennies and experiences zero attention degradation across millions of lines of unstructured code.
If frontier labs introduce architectures that can ingest messy, unstructured monorepos and infer unwritten human intentions with 99.9% deterministic accuracy at near-instant speeds, structural boundaries will matter less.
Until that day arrives, context economics are real. The builders winning with AI in production aren't the ones with the cleverest prompts; they are the ones whose codebases are clean enough for a machine to understand.
Related from Tech
Let's build your next big product.
Accepting project-based freelance, remote engineering roles, and hybrid positions.