AI-generated CSS rarely fails in obvious ways. The button renders, the card has a shadow, the layout collapses correctly on a phone. What goes wrong is slower and quieter. Each prompt adds another slightly different grey, another font size close to the existing ones, another !important to beat a rule the model could not see. After a dozen iterations, the stylesheet still works, but nobody can say why it works, and every change becomes risky.
The way to catch this is to measure it. Five numbers reveal most of the damage: unique colours, unique font sizes, !important declarations, ID and overly specific selectors, and rules that override each other instead of being edited. Record them before an AI session, record them after, and treat any jump as a code-review question. The rest of this article explains why AI-written CSS drifts in these particular ways, how each pattern appears in the metrics, and how to structure prompts and reviews so less drift gets in.
Why AI-written CSS degrades differently
Human-written CSS decays too, but people usually make one of two mistakes: they don't know the system, or they are in a hurry. AI coding tools make a third kind of mistake, which comes from how they work.
Every prompt is a local fix. When you ask an assistant to "make the header sticky" or "fix the spacing on mobile", it optimises for the visible result of that single request. Unless the tool is explicitly given the relevant files, it may not see the design tokens defined in another file, the utility class that already does the job, or the rule three components away that will conflict. The shortest path to "it works now" is usually to add a new rule rather than edit an existing one.
Adding feels safer than editing. Changing an existing rule can break something elsewhere, and a model that cannot see the rest of the project has no way to check. So AI tools lean towards appending: a new selector, a slightly more specific override, a new class. Each addition is harmless alone. Together they become layers of overrides that nobody designed.
"Make it work" pressure produces specificity escalation. If a style does not apply, the fastest ways to force it are a more specific selector or !important. Both fix the immediate symptom and make the next fix harder, and the next prompt often escalates again.
Training data lags the platform. Models learned from years of CSS written before container queries, :has(), native nesting, cascade layers and color-mix() were widely supported. Current models know these features, but they often fall back on older patterns: media queries for component-level layout, extra wrapper elements, hardcoded values where custom properties would do.
Nobody reviews the details. In vibe-coding workflows especially, the review is visual: does it look right? CSS is uniquely easy to accept on appearance alone. A stylesheet that passes the visual check but contains 40 unique colours will not show a problem until someone tries to change the brand colour.
This is not a fringe concern. In the 2025 Stack Overflow Developer Survey, 84% of respondents said they were using or planning to use AI tools. Yet 46% said they distrusted the accuracy of AI output, against 33% who trusted it. The most commonly cited frustration, named by 66%, was solutions that are "almost right, but not quite". A December 2025 CodeRabbit analysis of pull requests in open-source projects reported that AI-co-authored changes contained roughly 1.7 times as many major issues as human-written ones. That study covered code in general, not CSS specifically, but the pattern is the same: plausible output that needs verification.
The patterns to look for
Colour drift
This is the most common and most measurable pattern. You ask for "a slightly darker hover state", and the model writes #2b2f36 instead of using --color-surface-strong. Another prompt yields #2a2e35. A third, generated in a different session, uses rgb(43, 47, 54). On screen they are indistinguishable. In the stylesheet they are three unrelated values, and a later brand change will miss two of them.
StyleStats reports this as total unique colors, normalised to lowercase hex. A design system with a deliberate palette usually has a stable count. A rising count across AI sessions is the clearest single sign that tokens are being bypassed.
Font-size sprawl
The typographic version of colour drift. The type scale defines 14, 16, 20 and 24 pixels, and AI-generated components introduce 15px, 17px, 0.95rem and 1.05em because each looked right in isolation. Total unique font sizes tracks this. It rises less dramatically than colours but is just as telling, because type scales are usually small and deliberate.
Specificity escalation and !important
When a generated rule does not apply because an existing rule wins, the assistant often responds with a longer selector (.page .content .card .card-title), an ID selector, or !important. Specificity, the cascade rule that decides which competing selector wins, means each of these makes the next override harder.
StyleStats counts important keywords and ID selectors directly. Few healthy stylesheets need more than a handful of !important declarations, usually in utility classes or for accessibility overrides. A sudden jump after an AI session almost always marks a place where the model fought the cascade instead of understanding it.
Stacked overrides and low cohesion
Appending rather than editing produces a recognisable pattern: the same selector, or nearly the same selector, defined several times, each carrying one or two declarations that override an earlier version. StyleStats' average of cohesion metric, the average number of declarations per selector, tends to fall as small one-off rules accumulate. Lowest cohesion selector points to the worst offender, which is often a good place to start a clean-up.
Legacy layout habits
Floats for layout, heavy vendor prefixes, positioning hacks and media queries tied to viewport width for components that should respond to their container are all signs of patterns absorbed from older code. StyleStats' float properties metric should be close to zero in a modern codebase. For other patterns you need a linter or a manual review.
Universal and fragile selectors
AI-generated resets and "fix everything" rules sometimes include broad universal selectors (* { … }) or unqualified attribute selectors that match more than intended. They are rarely harmful alone but make behaviour harder to reason about. Universal selectors and unqualified attribute selectors are both counted.
Dead styles from abandoned iterations
Vibe coding is iterative: try a layout, reject it, try another. The CSS from rejected attempts often stays behind, because nobody asks the model to remove it and the model has no reason to. That is the unused-CSS problem at high speed, and it appears as growth in selector count and stylesheet size with no matching growth in features.
What the metrics catch, and what they don't
StyleStats is a smoke test, not a verdict. Its documentation says plainly that most of its numbers are most useful as trends over time rather than absolute thresholds. The value is in the comparison: what changed between the last merge and this one.
AI failure patternMetric to watchWhat a warning looks likeColours outside the design tokensTotal unique colorsCount rises after a session that added no new visual designOff-scale typographyTotal unique font sizesNew values between existing steps of the scaleCascade fightsImportant keywords, ID selectorsAny new !important outside utilities or accessibility overridesOverride stackingAverage of cohesion, lowest cohesionAverage falls, many one- or two-declaration rulesOutdated layoutFloat propertiesAny new float used for layoutOver-broad rulesUniversal selectors, unqualified attribute selectorsNew * or bare [attr] selectors
Some patterns need other tools. Magic numbers such as top: 37px or width: 347px and escalating z-index values do not appear in these metrics. Stylelint catches many of them with rules that forbid ID selectors, !important, duplicate selectors and duplicate properties, or that limit selector specificity. Plugins can require values from a list of tokens. Visual regression testing catches layout breakage that no static metric can see. The metrics tell you how much the stylesheet changed in risky ways. Linting and reviews tell you where.
A special case, AI-generated Tailwind
Many AI coding tools default to Tailwind CSS, often combined with component libraries, which changes where drift appears.
In a utility-first codebase, colour drift rarely shows up as new hex values in a hand-written stylesheet. It shows up as arbitrary values in class names: bg-[#2b2f36], text-[15px], mt-[13px], z-[999]. Each one compiles into a new rule and a new value that sits outside the theme. Since Tailwind v4 defines the design system in CSS with the @theme directive, the theme's colours and spacing are available as CSS variables. Arbitrary values are an escape hatch that bypasses them.
Two practical consequences follow. First, measure the compiled output, not the source. Point StyleStats at the built CSS file, because that is where arbitrary values become countable colours and font sizes. Second, search the source for square brackets. A quick search for -[ in class attributes shows how often the model left the design system, which is often more revealing than the compiled totals.
A review workflow that holds up
None of this requires giving up AI assistance. It requires treating generated CSS like any other contribution: review it, and back the review with numbers.
1. Record a baseline
Before a significant AI-assisted session, or on the main branch as a regular job, run StyleStats against the production stylesheet and save the JSON output. The tool supports JSON, HTML, Markdown and CSV formats. JSON is best for comparison, and Markdown is convenient for pull-request comments.
2. Compare after the session
Run the same command on the branch and compare the key numbers. An illustrative example of what a problematic session can look like:
MetricMain branchAfter AI sessionUnique colors1831Unique font sizes712Important keywords311ID selectors02
Nothing in that diff shows as a visual bug. All of it is a conversation worth having before merging: which of the 13 new colours are intentional, and why the cascade needed eight more !important declarations.
3. Automate it in CI
The StyleStats documentation describes the typical pattern: build the project, run StyleStats against the output stylesheet in JSON format, save the report as a build artifact, and compare it with the previous run. Start by reporting, not blocking. Once the team trusts the numbers, a rising count of unique colours or !important can be made a required check.
4. Constrain the prompts
Most drift can be prevented upstream. Assistants follow explicit constraints much better than implicit conventions. Useful instructions, often kept in a project rules file that the tool reads automatically:
Give the model the token file, and require colours, spacing and type sizes to come from it.
Forbid
!importantand ID selectors, and ask it to explain any conflict it cannot resolve with the existing cascade.Ask it to edit the existing rule that controls a property rather than add a new one, and to name the rule it changed.
In Tailwind projects, forbid arbitrary values unless explicitly approved.
Ask for modern features where they fit: container queries for component layout,
:has()instead of JavaScript class toggles, cascade layers to separate resets, components and utilities.At the end of an iteration, ask it to remove styles from abandoned attempts.
5. Review the diff, not just the screenshot
Visual approval is necessary but not enough. Read the CSS diff with three questions in mind. Does each new value come from the system? Does each new selector need to exist, or could an existing rule have been edited? Did any change weaken the cascade?
What AI does well in CSS
It would be unfair to present this as a list of failings. AI assistants are good at many CSS tasks that humans find tedious or error-prone: converting a design into a first layout, writing complex grid templates, producing accessible focus styles, translating between syntaxes, explaining why a rule doesn't apply, and refactoring a messy block when asked explicitly. Some of the best uses of AI in stylesheets are clean-up tasks. Given the StyleStats report and the token file, a model can often consolidate duplicate colours or collapse stacked overrides faster than a person.
Drift is not caused by the tool's lack of ability. It happens when a stateless generator works on a system it cannot fully see, and the output is judged only by appearance. Metrics fix the second problem, and giving the model the right context reduces the first.
The more of a codebase AI writes, the more the review shifts from reading every line to watching the right signals. For stylesheets, those signals were defined long before vibe coding existed: how many colours, how many sizes, how many overrides, and whether they grew.
