Skip to content
← All writing
September 6, 2026 · 7 min read

The code works, nobody understands it: the real cost of AI-assisted development

  • ai
  • code quality
  • architecture
  • code review
  • engineering

I have spent a fair amount of time recently inside codebases written by other developers. Most of it was written with AI. You can tell, and not because of anything as obvious as a stray comment.

You can tell because the code works and nothing in it explains itself.

Every function does what it says. Tests pass, more or less. The feature is live and the ticket is closed. And then you try to change something, and you find there is nothing underneath — no shape, no decision, no reason why this file exists here and not there. It is a wall of working code with no architecture behind it.

That is a different problem from bad code, and I think we are still underrating it.

What it actually looks like

The pattern repeats enough that I can now spot it in about ten minutes.

The same logic appears three times, slightly differently each time. Not copy-paste duplication that a linter would catch — three separate implementations of the same idea, each generated in a different conversation, each locally reasonable. Nobody ever saw them side by side.

Naming drifts. getUser, fetchUserData, loadUserInfo, all in one module, all doing roughly the same thing. Each name was fine when it was generated. Together they mean nothing.

Defensive checks are everywhere. Null guards on values that cannot be null. Try/catch around code that does not throw. It looks careful. It is actually the opposite — it is what you write when you do not know what the data can do, so you guard against everything.

Comments restate the code instead of explaining it. // loop through the users above a loop through the users. The line that would have helped — why this loop skips inactive accounts — is not there, because the person shipping it did not know either.

And there are no boundaries. Business logic in the controller, HTTP concerns in the service, a database call inside a component. Not because someone chose a pragmatic trade-off, but because each piece was generated in isolation and dropped wherever the surrounding file happened to be open.

Individually, none of these are serious. Together they mean the codebase has no through-line.

Why the code is hard in a new way

Reading unfamiliar code has always been work. But normally there is something to recover.

When a person writes a system, they leave a trail. The structure reflects how they were thinking. Even bad decisions are decisions — you can find them, understand the constraint behind them, and decide whether it still holds. Legacy code is hard, but it is coherent. It was one mind, applied over time, and you can reconstruct that mind by reading.

Code assembled from a hundred disconnected generations has no such trail. There is no consistent point of view to recover, because there was never one there. You cannot ask "why is it like this?" and get an answer, because the honest answer is that nobody decided. It came out that way.

So the usual method fails. You cannot read your way to understanding. You end up reverse-engineering behaviour instead — running it, changing things, seeing what breaks. That is dramatically slower, and it is why touching this kind of code feels so much worse than the line count suggests.

This is a pressure problem, not a tooling problem

I want to be careful here, because the easy version of this argument is wrong.

I use AI every day. It is genuinely good, and I am not interested in the position that says real engineers type everything by hand. That is nostalgia, not a standard.

The actual issue is that AI removed a specific piece of friction, and that friction was doing a job nobody noticed.

It used to be that you could not produce working code for a system you did not understand. Understanding was not a virtue you chose — it was a prerequisite. To make the thing work at all, you had to know what the pieces did. The learning happened whether you wanted it or not, as a side effect of shipping.

That link is now broken. You can ship working code for a system you cannot explain. The output looks the same from the outside. And when there is deadline pressure — which there always is — the path that skips understanding is faster, and nothing in the process punishes you for taking it.

Not this sprint, anyway. The cost lands on whoever opens the file in six months. Often that is the same developer, who by then has no more context than a stranger would.

So I do not think this is a discipline failure by individual developers. It is what happens when the incentive is speed, the tool removes the last reason to slow down, and no part of the workflow asks whether anyone understood what was merged.

Where I think the line is

The distinction I care about is not how much AI you use. It is whether you can defend what you shipped.

Concretely, before something of mine merges:

Can I explain every line without the assistant? Not recite it — explain why it is there and what happens if it is removed. If I cannot, I do not understand my own change yet.

Did I decide the structure, or did I accept it? Where a thing lives, what it depends on, where the boundaries sit — these are architecture, and they compound. This is the part I keep for myself. I will happily take generated code inside a shape I chose. I will not take the shape.

Would I have written it this way? Generated code is usually correct and rarely idiomatic for your codebase. It does not know your conventions or the three patterns you already have for this. If it does not match, change it before merging, not later. Later does not come.

Do I know why it works? "The tests pass" is evidence, not understanding. Passing tests on code you cannot explain means you have verified behaviour you did not specify.

None of that slows you down much. Reviewing a generated function properly takes a couple of minutes. Reverse-engineering it a year later takes days.

The part that actually worries me

Ship velocity recovers. A messy codebase can be refactored — that is normal work, and I have done plenty of it.

What does not recover as easily is the skill. Debugging, architecture, knowing why one structure survives change and another does not — none of that is taught. It is accumulated, by struggling with systems until the patterns become visible. Every time that struggle is skipped, the reps are skipped too.

A developer who has shipped for two years without that is not lazy and not untalented. They have been rewarded, consistently, for closing tickets. They are simply missing the reps, and there is no natural moment where anyone tells them.

If you are early in your career right now, that is the thing worth protecting. Use the tools — you would be foolish not to. But make sure you can still answer why. The generated code is not the valuable part. Being the person who can tell whether it is any good is.