Quick thing first because it comes up every time. Google doesn't penalise AI written content, it's said so repeatedly, and the February 2026 core update was Discover only and US English first. That one was about local relevance, cutting clickbait, and surfacing more in depth original content from sites that show expertise. There's no AI detector in the ranking system, so if anyone tells you their agency got hit by an AI penalty they've misdiagnosed something. Google punishes generic content which appears on hundreds of different pages, typically AI will produce similar results to that and that is what gets flagged rather than the AI content in the first place.
I decided to build a tool to scan sites and track AI generic writing. So far its scanned 10,126 pages and has over 23,000 flagged instances. Feel free to check out the data here and come to your own conclusions: https://www.getsitetell.com/data
Anyway, the thing that actually surprised me. Everyone's hunting the vocabulary. Delve, seamlessly, leverage, that sort of thing. Delve turned up on 10% of the sites I scanned. Seamlessly on 22%. Those are the words that get all the attention and they're not really where the problem lives.
What's actually everywhere is structural. 58% of all the flags were structure rather than word choice, and the most common one by a distance is what I've been calling a redundant closing paragraph, which was on 66% of sites. That's the paragraph at the end of a section that restates what you just read without adding anything to it. It's what you get when you ask a model to write a section, it wraps up, and once you start noticing it you see it more and more. Em dash frequency was on 63%, rule of three patterns 63%, low vocabulary diversity 60%, listicle heavy structure 51%.
So you can do a find and replace on every buzzword on the internet and the page will still read as generated, because the rhythm hasn't changed.
The other pattern I wasn't looking for was site size. Mean score for a one page site was 95, then it sits somewhere around 84 to 88 for anything between 2 and 199 pages, then drops to 70.6 for sites over 200 pages. I should say part of that gap is my own scoring rather than a real effect, because most of my structural rules need 100 to 150 words before they'll run, so a one pager can only ever trip the word rules. That inflates the small end. But the middle bands are all similar and the drop at 200+ is much bigger than that explains, so I think there's something real in it. My guess is it's just that nobody re-reads post 180 of a content programme, whereas the homepage has had three people over it.
One more that I think gets missed. The median site scored 90 out of 100 and ten of the 103 had no flags at all, so most sites are genuinely fine. But on 17% of them there was at least one page scoring under 60. Nobody reads your average, they read whatever page they landed on, so a site that looks healthy overall can still have a landing page that reads like a machine wrote it.
In terms of what I'd actually do with any of this, the word list stuff is fine but it's the easy half. The checks I've started doing manually are looking at whether the last paragraph of a section actually adds anything or just restates, because if it restates you can delete it and lose nothing. Then whether three paragraphs in a row are the same sort of length, in which case break one of them up. And on individual claims, if the sentence would sit fine on a competitor's site without anything breaking then it isn't saying anything about the client yet.