Free tools Windows power users keep installed
One-click scans. No signup required.
iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
SlopShape looks for patterns in how a commercial blog post organizes information—not just words or phrases that sound machine-written. In its September 28, 2026 version 3 preprint, the structure-only classifier scored 97.0 macro-F1 on a controlled test of company domains and 96.1 macro-F1 after the tested AI systems reworded their own posts. Those are results for a specific benchmark, not proof that a detector can identify the author of any web page.
What SlopShape detects
SlopShape is a research classifier for distinguishing human-written from AI-generated B2B commercial blog posts. It examines document-level characteristics such as the order in which information appears, how evidence is introduced, how the argument is framed, and what voice the post adopts. It is not a raw-HTML or DOM fingerprint, and its target is narrower than web content generally.
The paper describes a recurring combination in AI-generated posts as a “tidy, self-announcing” pattern: a post states its thesis early, previews its organization, contrasts a legacy approach with a modern one, ends with a summary, and uses a confident institutional voice. The authors treat that as a combined signal. None of those traits on its own establishes who wrote a particular page.
How its architecture turns a post into classifier features
The pipeline has two main stages: an LLM converts a post into a structured feature representation, then a classifier uses that representation to predict authorship class. The LLM is part of the measurement system; the method is not simply a conventional classifier reading page markup.
#1 Best Overall
- Represent the post. The pipeline creates a JSON-like template for each document, organizing observations under a commercial-content schema.
- Find candidate features. Researchers compare human posts with AI-generated mirrors to identify potentially distinguishing characteristics.
- Screen the candidates. Features are checked for whether they can be answered, whether supporting evidence can be located in the text, and whether the scoring is reliable.
- Score the documents. An LLM applies the retained instrument to each post and encodes its answers as a feature representation.
- Train and test classifiers. The resulting representations are used to predict human versus AI authorship. A separate task predicts which of five AI systems generated a post.
The version 3 instrument contains 203 features across 11 dimensions: purpose, audience, structure and flow, explanation, evidence, voices, actionability, commercial integration, timeliness, page format, and writing style. Of these, 176 are structural and 27 are style features. The headline structure-only results exclude the writing-style features.
What the benchmark tested—and what the scores mean
The corpus contains 2,250 human B2B blog posts archived from 2008 through 2022, drawn from 268 company domains, and 11,250 AI mirrors. The mirrors were produced in August 2026 from reverse-engineered briefs using GPT-5.4, Claude Sonnet 4.6, Gemini 3 Flash, DeepSeek V3.2, and Kimi K2.5. The final data were split by company domain, so the companies in the test set did not appear in training.
The table reports results from Jochen Madler’s 2026 version 3 preprint. Macro-F1 is the unweighted average of the F1 scores for the classes in a task; it is not the percentage of all pages correctly classified.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errors| Reported result | What it measures |
|---|---|
| 97.0 macro-F1 | Structure-only human-versus-AI classification on held-out company domains. |
| 96.1 macro-F1 | The same structure-only task after each AI test post was reworded by the AI system that generated it. |
| 68.6% correct; 16.7% chance | Six-way source attribution: human writing or one of five AI systems. This is a different task from binary detection. |
| 8.8 macro-F1 points | The structure-only classifier’s advantage over the style-only classifier in this study. |
| 0.939 Cohen’s kappa | Agreement between the two annotators in the gold-annotation session. |
| 0.951 mean Cohen’s kappa | Agreement between human annotators and the LLM scorer on the feature items included in that session. |
The paper also reports that the structure-only and style-only classifiers tended to make errors on different posts. That pattern supports the view that structure contributed information beyond surface wording in this benchmark; it does not show that the same separation will hold for other genres or populations.
How much does rewording change the result?
In the reported self-rewording test, each AI system reworded posts that it had generated, and the structure-only score changed from 97.0 to 96.1 macro-F1. This is evidence that the measured structural signal persisted under that particular kind of paraphrase.
The test did not cover humanizer products, extensive human editing, collaborative human-AI writing, or deliberate attempts to alter a post’s structure. It therefore cannot establish resistance to every way a person might revise or disguise AI-assisted content.
Rank #3
Where the evidence stops
The human and AI writing came from different periods and processes
The human posts predate ChatGPT, while the AI mirrors were generated in August 2026. The authors check for publication-year effects but say that a same-period comparison of human and AI writing remains open. The collection is also software-heavy and US-heavy; 75.5% of its snapshots are from 2020–2022.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchThere is also an information difference: original human authors wrote with full business context, while the AI systems received reverse-engineered briefs. Some detected distinctions could reflect context left out of those briefs rather than authorship alone. The authors also scored human pages and AI outputs in different formats; their format-sensitivity checks removed features affected by that difference.
The LLM scorer is part of the method
The LLM helps screen and extract templates, discover candidate features, and score documents. The paper reports repeatability checks and human validation, but the gold-annotation validation was a limited session involving two company-affiliated annotators. Those checks provide evidence about the instrument’s performance under the study’s protocol, not an independent demonstration of error rates across the web.
Rank #4
Real writing can be mixed or formulaic
The measured labels are source classes in a constructed experiment. A real marketing post may combine human drafting, AI assistance, editing, and a conventional SEO template. The study does not establish how often those cases would be misclassified, nor does it establish false-positive rates for current human-written marketing content. The authors also caution that their feature discovery was optimized to separate six sources in this corpus; they do not claim commercial writing is inherently easier to distinguish than fiction.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.What readers can do with SlopShape’s related checker
Sitefire’s Slop Checker accepts one blog-post URL and returns a “Human score” based on ten structural signals from the study. The product page says the study measured English posts and that scores for other languages are indicative. Treat the result as a screening signal, not an authorship verdict: the benchmark does not establish the origin of an individual post, especially where writing is mixed or the content differs from the tested corpus.
Code, materials, and commercial context
The public release repository says it includes code, prompts, the feature instrument, aggregate artifacts, and verification materials. It lists the code under the PolyForm Noncommercial License 1.0.0; other release materials are all rights reserved and made public for audit or verification. Some document-level data are available to researchers under a noncommercial research agreement. The repository also discloses the author’s commercial interest in Sitefire and says commercial use of the code requires a separate agreement.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

