AI-generated code: why “it works” does not mean “it is maintainable”
Coding agents quickly produce code that passes the tests. But code that works is not necessarily healthy code. What research on design smell detection teaches us, and how to keep technical debt under control.
ADSERVIO INSIGHTS · GENAI

KEY POINTS
- Code that passes the tests can still have design flaws: tests check what the code does, not how it is designed. Code smells slip under that radar.
- The debt is paid later: these smells are associated with more changes, more defects and longer modification times.
- No model sees everything: across 17 models tested in a PhD thesis carried out at Adservio, none scores above 0.30 on Feature Envy.
- The key is how code is represented: what the model is given to read matters more than its size or sophistication.
SECTION 1
Code that passes the tests is not necessarily healthy code
Coding agents have changed the pace of software development. In a few minutes they write a function, a service, sometimes a whole module, that compiles and passes the tests. For a team under pressure, in banking, energy, healthcare or retail, it is tempting to consider the job done.
That is where the risk lies. A test checks a behaviour: for this input, that output. It says nothing about design: a method too long to understand, a class that centralises everything, logic duplicated in three places. The program works, but it becomes harder to understand, change and evolve.
These flaws have had a name since Martin Fowler's work: code smells, symptoms in the code that point to a deeper design problem. They matter all the more now that models generate working code very quickly, and that this code tends to accumulate poor design practices that are paid for, in the medium term, as technical debt.
This article draws on the research of Djamel Mesbah, carried out at Adservio as part of a CIFRE PhD with Université Paris-Saclay and ISEP, on the automatic detection of these smells with artificial intelligence. It has led to several scientific publications, listed at the end of the article.
@cite:la-programmation-intuitive-peut-elle-produire-des-logiciels
SECTION 2
Why design smells are costly
Code smells are not a matter of style. Many empirical studies associate them with lower maintainability and greater evolution effort.
### More changes, more defects
Systems that concentrate smells such as God Class, Feature Envy or Long Method show more changes and defects, longer modification times and more complex change sets during maintenance. The components concerned accumulate technical debt: it becomes harder to understand the design intent, locate a change and propagate it safely.
### A cognitive load that adds up
An isolated smell remains manageable for a developer. It is their accumulation that causes trouble: several co-existing smells clearly increase the workload and reduce comprehension accuracy. At the scale of a codebase continuously fed by agents, this accumulation is precisely what needs watching.
### Four smells to know
@diagram:code-ia-defauts
SECTION 3
Can detection be handed over to AI?
If AI produces these smells, can it also spot them? Research gives a nuanced answer, and that is what makes it useful for teams.
### No model family dominates
Seventeen models from three broad families were compared on the same set of Java samples annotated by professional developers (the MLCQ dataset), under an identical protocol: classical learning on metrics, sequence models that read code as text, and neural networks on graphs built from the syntax tree. Graph-based models win on class-level smells (Blob, Data Class), sequence models on overly long methods. No family wins everywhere.
@diagram:code-ia-scores
### Large language models: useful, but not enough
Queried directly through prompts, models such as GPT-4 and LLaMA reach limited performance on this task. Adding a few examples to the prompt clearly improves their answers, and GPT-4 remains the most reliable. But both struggle with Feature Envy and Blob. A study is under way on more recent language models, able to reason step by step (chain of thought) and equipped with a tool harness, to measure their progress.
### One smell resists everything
Feature Envy remains an open challenge for every approach tested. The reason is telling: it is a relational smell. To see it, you need to know which classes call each other and with which types, information that today's representations of an isolated piece of code do not carry.
SECTION 4
The lesson: it is not the model, it is the representation
The main takeaway fits in one sentence: what makes a smell detectable is not the sophistication of the model, but how the code it reads is represented.
@diagram:code-ia-representations
More computing power does not change this. Distributing training reduces cost and computing time, but does not lift the detection ceiling: the limit lies in how code is represented.
For an IT department as for a development team, the lesson goes beyond research: a bigger model does not replace thinking about what it is given to see. That holds for smell detection, and it holds for coding agents themselves.
SECTION 5
Measuring how clean generated code is
Reference benchmarks such as SWE-bench measure whether generated code works and solves the problem at hand. Another question remains largely open: is that code clean? Which design smells do models introduce into what they produce?
This is the direction explored by ongoing work at Adservio. The question becomes central as coding agents spread: producing code that passes the tests does not guarantee maintainable code.
SECTION 6
How we integrate AI agents into our projects
At Adservio, these lessons shape how we integrate AI agents into our clients' software lifecycle, whatever their sector. Four principles structure our projects.
@diagram:code-ia-chaine
### Measure design, not only behaviour
Tests remain essential, but they are no longer enough. Design smell detection enters the integration pipeline alongside the tests, so that technical debt becomes visible when it is created, not six months later.
### Combine approaches rather than bet on one model
Since no model family dominates, robust tooling combines several approaches depending on the smell sought, rather than entrusting all detection to a single model, however powerful.
### Keep humans on relational smells
Smells that require understanding relations between classes, such as Feature Envy, still escape tools. Design review by an experienced developer remains the safety net on these points.
### Frame agents on proven patterns
A coding agent gives its best results on recurring tasks whose patterns have been validated and capitalised. We start with those cases, measure the gains, then widen progressively. The gain comes from framing the agents, not from setting them free.
@cite:claude-code-nous-a-economise-97-du-travail-puis-c-est-devenu
SECTION 7
In short
Coding agents speed up software production, but speed says nothing about design quality. Measuring smells as soon as they appear and framing the agents is what turns a productivity gain into a lasting asset.
SECTION 8
Sources
Mesbah D., El Madhoun N., Al Agha K., Zouaoui A., « A Survey on Code Smells Detection using Machine Learning Techniques », Information and Software Technology, Elsevier, accepted, 2026. https://www.sciencedirect.com/science/article/pii/S0950584926002314
Mesbah D., El Madhoun N., Al Agha K., Chalouati H., « Leveraging Prompt-Based Large Language Models for Code Smell Detection: A Comparative Study on the MLCQ Dataset », EIDWT, Springer, 2025. https://link.springer.com/chapter/10.1007/978-3-031-86149-9_42
Mesbah D., El Madhoun N., Al Agha K., Chalouati H., « Exploring NLP Techniques for Code Smell Detection: A Comparative Study », AINA, Springer, 2025. https://link.springer.com/chapter/10.1007/978-3-031-87769-8_9
Mesbah D., El Madhoun N., Al Agha K., Chalouati H., « Beyond the Code: Unraveling the Applicability of Graph Neural Networks in Smell Detection », NBiS, Springer, 2024. https://link.springer.com/chapter/10.1007/978-3-031-72325-4_15
FAQ
Frequently asked questions
What is a code smell?
A symptom in the code that signals a deeper design problem, as defined by Martin Fowler. The program works, but it becomes harder to understand, change and maintain. Examples: a method that is too long, a class that centralises too many responsibilities.
Is AI-generated code of lower quality?
Code produced by LLMs often works, but it tends to accumulate poor design practices. Measuring exactly which ones is the subject of ongoing research at Adservio: passing the tests does not guarantee maintainable code.
Can AI be used to detect design smells?
Partly. Graph-based models detect class-level smells well, sequence models catch overly long methods, and GPT-4 does better with a few examples in the prompt. But some relational smells, such as Feature Envy, resist every approach tested.
How can the technical debt introduced by coding agents be limited?
By adding design smell detection to the integration pipeline, combining several detection tools, keeping human review on relational smells and framing agents on proven patterns.
Why should executives and business teams care?
Because code that is hard to evolve means a new feature that arrives later or costs more. Maintainability is a matter of time and budget, not only of engineering: the technical debt introduced today is paid with every change requested tomorrow.
ABOUT ADSERVIO
Adservio is an AI-native digital transformation partner: AI-augmented IT departments, software engineering, DevOps, MLOps, cybersecurity and AI governance.
Let's talk about your project: hello@adservio.fr · adservio.fr/contact