How to Tell AI Slop From Human Writing: Ten Structural Tells From 13,500 Blog Posts

Structure alone tells AI-generated blog posts from human ones at 98.0 macro-F1 on 13,500 posts. The ten tells and what human writing does differently.

Jochen MadlerSep 4
Sitefire research chart, 'AI writing has a shape. You can see it without the words.': structural rarity of 13,500 blog posts by source, human posts concentrated in the rarest configurations at 0.84 against five AI models between 0.33 and 0.55

Take 13,500 blog posts. Throw away every word. Keep only how each post is built: what it promises in the title, where it states its thesis, whether it announces its own sections, how it closes. A classifier reading nothing but that structure still tells the AI-generated posts from the human ones at 98.0 macro-F1, a balanced accuracy score across both classes, on companies it never saw in training.

At Sitefire we ran this comparison on 2,250 blog posts written by people at 268 company websites between 2008 and 2022, before ChatGPT existed, against 11,250 AI-generated versions of the same briefs by five AI models: GPT-5.4, Claude Sonnet 4.6, Gemini 3 Flash, DeepSeek V3.2, and Kimi K2.5. The design replicates StoryScope (Russell et al., 2026), a study that found the same structural signature in AI-written fiction. Commercial blog posts have no plot and no characters, and the signature transfers anyway.

This post is the practical version of our study, down to what to check in your own drafts on Monday.

Can you tell AI writing from human writing without reading the words?

Yes. Structure alone separates AI-generated blog posts from human ones at 98.0 macro-F1, while word count alone scores 45.5, which is chance.

Word-level detectors are the usual tool, and on unedited AI-generated text they are close to perfect: on the same posts the best of them score 100.0. They also read only the surface. Change the words and the detector has nothing left to read, which is the weakness Krishna et al. (2023), Sadasivan et al. (2023), and Weber-Wulff et al. (2023) documented with rewording attacks.

Structure is a different layer. Our study measured each post on 214 features covering what a post says, in what order, with what evidence, and in what voice. The 187 features that describe structure, with every style feature removed, reach 98.0. The 27 style features on their own reach 88.1. A classifier given only the post's length cannot separate the classes at all.

What the classifier readsDetection score (macro-F1)
Structure only (187 features)98.0
Structure, after every AI post is reworded98.1
Style features only (27 features)88.1
Style features, after rewording87.1
Word count only45.5
Best word-level detector, unedited text100.0

The classifier was tested on posts from companies it never saw in training. Company websites resemble themselves, so a split by company is the strict version of the test, and 98.0 is the number under that split (95% confidence interval 96.7 to 99.2).

What gives an AI-generated post away?

Ten structural tells give an AI-generated post away, each a choice about how the post is built rather than which words it uses. The four largest: a payoff promised in the title, a summary stage, a restated thesis at the close, and no path for the reader to participate.

A tell is a feature value that one class picks far more often than the other. The table lists the ten with the largest human-AI gap, the share of AI posts that show the AI-leaning value, and the share of human posts that do. Read them in the order a reader meets them.

Where in the postThe tellAI postsHuman posts
TitleThe payoff is promised in the title89%26%
OpeningThe thesis comes before the first section93%51%
OpeningThe problem is placed in the title or first sentence53%24%
MiddleA legacy-versus-modern contrast frames the argument76%26%
MiddleThe voice is an editorial explainer, not a person71%40%
MiddleThe stakes escalate as the post goes on88%52%
CloseA summary or synthesis stage88%27%
CloseThe closing move restates the thesis77%12%
CloseNo path for the reader to participate97%38%
LengthOver 800 words83%56%

Put the AI-leaning values together and you get the tidy, self-announcing blog post. It promises the payoff in the title, then states its thesis and announces its structure before the first section. It speaks as an editorial explainer about a category rather than as a person about a case, and it closes with a stage that summarizes or restates the thesis.

You have read this post. The title promises the payoff: "How to Cut Onboarding Time in Half." The opening announces the flow: "In this post, we'll cover why onboarding stalls, three fixes that work, and how to measure the difference." The close restates the thesis: "In short, structured onboarding saves time." Any one of these moves is defensible. Across 13,500 posts, the set is the shape of AI writing.

Two of the tells have nothing to do with signposting. 93% of AI posts state their thesis before the first section, against 51% of human posts. And 97% of AI posts offer the reader no way in: no place to reply, join, apply, or try. Human posts leave that door open 62% of the time. The AI post is written at you.

Does rewording remove the tells?

No. Reworded by the same AI model that wrote it, every AI-generated post in the test set still reads as AI-generated: detection moves from 98.0 to 98.1. Only 5 of 1,450 reworded posts pass as human.

We had each of the five AI models reword its own posts span by span, targeting the seven categories of AI-writing artifacts that professional editors identified in earlier work, while keeping every claim and link. Afterward, on average, 73% of a post's 13-word sequences no longer appeared verbatim.

Style features felt it. Their score dropped from 88.1 to 87.1, as a wording attack should. Structure did not move. The reworded post still promised the payoff in its title, still announced its sections, still closed on a restatement, because none of that lives in the words.

The test covers only rewording by the same AI model that wrote the post. Dedicated humanizer tools, which are built to defeat detectors and can restructure a post as well as reword it, were not part of our study.

If you want to know whether a post reads as AI-written, look at what it does, not which words it uses. Swapping vocabulary changes the score a word-level detector gives you and leaves the shape your readers see untouched.

Do all five AI models write the same way?

Largely yes. The five AI models share one structural shape, and each adds a small accent of its own. 79.3% of posts are attributed to the correct one of six sources, against a 16.7% chance rate.

When the classifier had to pick which of the six sources wrote a post, it identified the human posts almost perfectly, and nearly every error was one AI model mistaken for another.

Rarity makes the same point from the other side. For every post we measured how far its structure sits from its 25 nearest neighbors, as a rarity percentile ranked against all other posts. Human posts average a rarity percentile of 0.84. AI-generated posts average 0.44, and the five AI models cluster: DeepSeek V3.2 at 0.55, Claude Sonnet 4.6 at 0.48, Gemini 3 Flash at 0.46, Kimi K2.5 at 0.36, GPT-5.4 at 0.33. Every one sits well below the human mean.

Detection after rewording runs 97.6 to 98.1 across the five AI models, so none of them escapes the shape. Which one sits where is a detail. The shape is a property of how these AI models were trained to write blog posts, not of any one vendor.

What does human writing do differently?

Human blog posts occupy the rare regions of structural space: a mean rarity percentile of 0.84 against 0.44 for AI-generated posts. Of the rarest 1% of all posts, 149 are human and 4 are AI.

The human-leaning tells read as absences: no announced flow, no restated close, no legacy-versus-modern setup. Stakes stay where the writer put them instead of escalating on schedule, posts run under 800 words more often, and the problem arrives when the writer gets there, not in the title.

Rarity is a proxy for originality, and the gap is large: Cohen's d of 1.83, where the fiction study found 0.83. Within a group of six posts written from the same brief, the human one is the structurally rarest in 85.5% of cases. On rarity alone, with no classifier at all, the classes separate at an AUC of 0.90.

None of this says human posts are better written, only that they are less alike. The human post is the one that does not tell you it is about to make a point, does not sum itself up at the end, and leaves you somewhere to go next.

What to check in your own drafts on Monday

1. Run the ten-tell check on your last five posts. Count how many of the ten AI-leaning values each post shows. Seven or more, and the shape is the problem, not the words. A human post in our study shows three or four on average. An AI-generated post shows eight.

2. Restructure, do not reword. Rewording moved the detection score by a tenth of a point. Cutting the announced flow, replacing the restated close with the actual next step for the reader, and giving the reader a path in changes the post's shape. Those are the tells with the largest gaps: 97% versus 38% for the missing path, 77% versus 12% for the restated close.

3. Test yourself. Spot the Slop, Sitefire's five-round game, puts two posts side by side. One was written by a person, one by an AI model, and the reveal shows the tell that gave it away. Most people who take it learn the shape faster than any checklist teaches it.

At Sitefire we ship articles into customers' CMSs every week, and our study set the quality bar for them: a post has to pass on structure, not on vocabulary, because the words are the easy part.

The Bottom Line

AI-generated blog posts can be identified without reading a single word, at 98.0 macro-F1, from structure alone. The structure survives rewording, is shared across five AI models, and is checkable by a reader in the time it takes to skim a post: payoff in the title, thesis before the first section, announced flow, summary close, no way in.

Human writing sits in the regions of structural space the AI models rarely reach, and it gets there by leaving things out. The signature fits in one sentence: AI writes the tidy, self-announcing post it was taught to write; humans just write the thing.

Frequently Asked Questions

Does this apply to my industry, or only to B2B blogs?

The 2,250 human posts come from 268 company websites across software, e-commerce, fintech, health, developer tools, and services, and detection held above 96 in every industry group with enough posts to measure. The fiction study found the same signature in a completely different genre, which is the strongest evidence the shape is general.

Do Google and AI search engines use signals like this?

Google's policy on scaled content abuse targets mass-produced pages by outcome rather than method, and YouTube already uses structural signals such as upload pacing and templated narrative patterns against mass-produced video. Structural features are a candidate signal of exactly that kind. Whether Google or any AI model uses them today is not something our study can show.

Can I make an AI-generated post pass by restructuring it?

Restructuring changes the shape, which is what the classifier reads, so a post rebuilt without the announced flow, the restated close, and the closed door would look more human on these ten tells. Our study did not test that attack. The useful reading is the other way round: restructuring is what makes a post read as written by a person.

How was the study validated?

Two people scored a sample of posts by hand against the same feature definitions the scoring AI model used. Human-human agreement reached kappa 0.928 and human-model agreement 0.946, both above the fiction study's 0.739 and 0.839. The classifier was tested only on companies it never saw in training, and eight further robustness checks are reported in the study.

Key Takeaways

  • Structure alone identifies AI-generated blog posts at 98.0 macro-F1 on 13,500 posts, tested on companies the classifier never saw.
  • Ten structural tells carry the signal. The largest are a payoff promised in the title (89% of AI posts vs 26% of human), a summary stage (88% vs 27%), a restated close (77% vs 12%), and no path for the reader to participate (97% vs 38%).
  • Rewording every AI post with the same AI model that wrote it moves detection from 98.0 to 98.1, while style features lose ground and structure does not.
  • The five AI models share one shape and can still be told apart 79.3% of the time against a 16.7% chance rate.
  • Human posts sit in the rare regions of structural space (mean rarity percentile 0.84 vs 0.44), and of the rarest 1% of posts, 149 are human and 4 are AI.
  • Restructuring, not rewording, is what changes how a post reads.

Methodology note

Our study is "SlopShape: Identifying AI-Generated Commercial Web Content" (Madler, 2026), a domain replication of StoryScope. Human posts: 2,250 blog posts from 268 company websites, captured in Wayback Machine snapshots dated 2008 to 2022. AI-generated posts: a content brief inferred from each human post, with the publisher anonymized, given to GPT-5.4, Claude Sonnet 4.6, Gemini 3 Flash, DeepSeek V3.2, and Kimi K2.5, yielding 11,250 AI-generated posts. The AI model never saw the human original. Features: 214 features, 187 structural and 27 style, scored by an AI model and validated in a human annotation session (kappa 0.928 human-human, 0.946 human-model). The ten structural tells are the study's ten core features. Classifier: gradient-boosted trees, tested on held-out companies never seen in training. Not tested: dedicated humanizer tools, edited or human-AI collaborative text, and text from AI models other than the five studied.

Sitefire builds content tools. The study's methods, code, prompts, and aggregate results are public in the verification repository. Per-post data is available to researchers on request.

Sources


Share this article

Keep reading