Can a Blog Post Sound Human and Still Have an AI-Shaped Sales Pitch?
Quick summary: can a sales pitch sound human but keep an AI-shaped structure?
- Structure matters: the study examined how commercial posts build an argument, rather than just their wording.
- Rewording did not remove the pattern: the article reports that structural classification remained strong after substantial rewriting in one controlled test.
- The findings have limits: this comparison of older human-written B2B posts and AI-generated versions does not establish a universal AI detector.
- Clear writing is still useful: an early thesis, organised sections or a summary does not prove that AI wrote a post.
- Use detection cautiously: test errors, bias and new writing contexts, and prioritise evidence and provenance over guesses about individual authors.
The pitch can change its wording and keep its shape. That is the surprising result behind a new study of AI-written company blogs: an AI-shaped sales pitch may survive after its wording changes. Researchers report that a classifier could still separate AI-generated posts from human originals after the AI rewrote most of its phrasing. The classifier was looking at how a post made its case, not simply which words it used.
That matters to developers building content systems, evaluation pipelines, and moderation tools. It also raises a more interesting question than “Can we spot AI prose?” Can a post sound human sentence by sentence while retaining an AI-shaped way of selling?
The short answer is: this study offers evidence that an AI-shaped sales pitch can be measured in one controlled setting. It does not establish a universal AI detector, or show that every tidy marketing post was written by a model.
What the study actually tested
SlopShape: Identifying AI-Generated Commercial Web Content is a September 2026 arXiv preprint by Jochen Madler of Sitefire. It extends StoryScope, a University of Maryland and Google DeepMind study of AI-generated fiction. SlopShape is a separate study, not a paper by those institutions.
The SlopShape corpus contains 2,250 human-written B2B company blog posts from archived web pages published before ChatGPT. For each human post, the researchers inferred a brief and asked five models to write a matching post: GPT-5.4, Claude Sonnet 4.6, Gemini 3 Flash, DeepSeek V3.2, and Kimi K2.5. That produced 11,250 AI mirrors. The comparisons are paired: the AI and human posts for a given prompt share a company domain and topic.
“AI-shaped sales pitch” is shorthand for a pattern of structural choices measured in this commercial writing corpus. It is not a psychological claim about what a model intends, and it is not a judgment that every post with a clear thesis or summary is bad writing.

How an AI-shaped sales pitch becomes a feature vector
The pipeline turns prose into structured answers before training the classifier. An LLM-based instrument scores each post against questions about purpose, structure, evidence, voice, commercial integration, and other dimensions. Examples include where the post first promises its payoff, whether it announces its structure, how it discloses sources, and what its conclusion does.
The released instrument has 214 questions: 187 designated structural and 27 designated style features. Answers are encoded as categorical, binary, ordinal, or multi-select values. The paper then trains an XGBoost classifier on those feature vectors. In plain terms, a post becomes a row of measurements, and the classifier learns combinations of measurements that separated the generated mirrors from the human originals in this dataset.
This is different from feeding the raw post to a detector that looks for token patterns. The structural model does not need the phrase “in this article” to identify an announced outline. It can ask whether the post promises a payoff in its title, states its thesis early, announces its sections, and ends by restating the thesis.

Why the sales-pitch framing fits
The strongest commercial-post features were not a list of suspicious adjectives. They described how a post delivered its argument. AI versions more often promised the payoff in the title, stated the thesis before the first section, announced the flow, used an editorial-explainer voice, and closed with a summary or restatement. That combination, rather than any single tell, is what this article calls an AI-shaped sales pitch. Several human-leaning features were the inverse: the thesis arrived later, the stakes escalated, and the post did not announce every stage in advance.
Imagine two posts about reducing software onboarding time. One opens with “This guide explains the three steps that cut onboarding time,” lists the steps, and ends by repeating the promise. A second opens with a concrete failure, shows the support ticket that exposed it, compares two fixes, and lets the recommendation emerge from the evidence. The first is not automatically AI-written, and the second is not automatically human. They illustrate the kinds of structural decisions the paper measures.
The paper also measures structural rarity: posts are considered rarer when their feature vectors sit farther from nearby posts in the measured space. In the corpus-wide analysis, 149 human posts and 4 AI posts fell above the 99th-percentile rarity threshold. That is an interesting distributional result, not a score for literary originality or quality, and it is specific to the study’s reference set.
The authors report 98.0 macro-F1 for structural features on a test split grouped by company domain. Macro-F1 averages the F1 score for each class, so it gives the human class equal weight to the AI class. This matters because each human post has five AI mirrors. A classifier that guessed “AI” for every row would already get 83.3% raw accuracy on this paired dataset, while failing to identify any human post.
There is a metric wrinkle in the social post that introduced the result. It describes 98% accuracy and says the model got 19 of 1,740 posts wrong. Nineteen errors out of 1,740 imply about 98.9% raw accuracy. The paper reports 98.0 macro-F1 for the structural classifier. Those are different metrics, so quote the paper’s metric precisely rather than treating “98%” and “19 wrong” as interchangeable.
Rewording changed the copy, not the score
For the rewording test, the authors asked each model to edit its own AI-generated posts, following an editing protocol aimed at known AI-writing patterns. Across the 1,450 AI posts in the test subset, about 73% of the original 13-word sequences no longer appeared verbatim. Yet the structural classifier scored 98.1 macro-F1 on the rewritten set, compared with 98.0 on the original set.
That is a strong result for this particular rewrite test. It does not mean that every paraphrase, human edit, or collaborative workflow will leave the signal intact. The authors explicitly say they tested this one rewording setup, not a broad range of humanizer tools or substantial human-AI editing. They also report that the rewritten feature answers did change; the changes simply did not move enough examples toward the human pattern for the classifier to lose performance.
A small, runnable evaluation pattern
Developers can borrow the evaluation discipline without copying the paper’s classifier. The example below assumes you have an annotated CSV with one row per post, a company-domain group, a human/AI label, and categorical structural measurements. It holds out entire company domains, preventing posts from the same company from appearing on both sides of the split.
Install the dependencies with python -m pip install pandas scikit-learn, save the snippet as evaluate_structure.py, and run it with python evaluate_structure.py after saving your annotations as annotated_posts.csv.
import pandas as pd
from sklearn.compose import ColumnTransformer
from sklearn.impute import SimpleImputer
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import classification_report
from sklearn.model_selection import GroupShuffleSplit
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import OneHotEncoder
# CSV columns: company_domain, label, and the features below.
# label must contain exactly "human" or "ai".
feature_columns = [
"payoff_location",
"thesis_position",
"outline_announced",
"source_disclosure",
"conclusion_behavior",
"primary_voice",
]
data = pd.read_csv("annotated_posts.csv")
X = data[feature_columns]
y = data["label"].str.lower()
groups = data["company_domain"]
split = GroupShuffleSplit(
n_splits=1,
test_size=0.20,
random_state=42,
)
train_index, test_index = next(split.split(X, y, groups=groups))
categorical = make_pipeline(
SimpleImputer(strategy="most_frequent"),
OneHotEncoder(handle_unknown="ignore"),
)
features = ColumnTransformer(
[("structure", categorical, feature_columns)]
)
classifier = make_pipeline(
features,
LogisticRegression(max_iter=1000, class_weight="balanced"),
)
classifier.fit(X.iloc[train_index], y.iloc[train_index])
predicted = classifier.predict(X.iloc[test_index])
print(
classification_report(
y.iloc[test_index],
predicted,
labels=["human", "ai"],
target_names=["human", "AI"],
zero_division=0,
)
)
This is a teaching baseline, not a reproduction of SlopShape, and it will not recreate the paper’s score. The example uses logistic regression, while the paper uses XGBoost and its own LLM-scored feature instrument. Most importantly, model output is only as useful as the labels, feature definitions, sampling, and held-out test design behind it.
Testing an AI-shaped sales pitch detector
A held-out company split is a useful safeguard against learning a brand’s house style. SlopShape split its classification corpus by company domain, putting no domain in more than one split. That is stricter than randomly dividing individual posts. But it does not answer every deployment question.
- Match the target use. These experiments test first-pass AI mirrors of old B2B posts. They do not establish performance on personal essays, documentation, product pages, or drafts edited collaboratively with humans.
- Separate development from evaluation. Discover features and tune thresholds using training and validation data. Keep a company- and time-separated test set untouched until the end.
- Report class-specific errors. Include confusion matrices, precision, recall, macro-F1, and calibration at the intended operating threshold. A high overall score can hide poor performance on the less common or higher-cost class.
- Test distribution shifts. Try newer models, other industries, different languages, new content formats, human-edited AI drafts, and human writing produced with grammar or editing tools.
- Validate the measuring instrument. If an LLM converts text into feature values, test repeatability and compare its answers with independent human annotations. SlopShape reports both repeatability checks and a human gold-annotation session, which strengthens the measurement case but does not eliminate judgment calls in the schema.
These safeguards matter beyond this one paper. The RAID benchmark evaluates detectors across more than six million generations, 11 models, eight domains, 11 attacks, and four decoding strategies. Its authors find that detectors can fail under adversarial edits, changed sampling strategies, and unseen models. That benchmark studies general AI-text detection, not SlopShape’s structural instrument, but it shows why in-distribution performance is not the same as deployment robustness.
Two design limits deserve special attention. The human posts predate ChatGPT, while the AI mirrors were generated in 2026. The paper tests whether its features predict publication year among human posts, but a same-period human-versus-AI corpus would address the time mismatch more directly. Also, human writers had the full business context; the models received reverse-engineered briefs. Some measured differences could come from that missing context, not only from authorship.
The instrument itself uses an LLM to assign feature values. The authors report repeatability checks and agreement against a small human annotation session, but the feature schema still encodes judgment about what counts as structure. The paper also discloses that its author operates Sitefire, a commercial GEO product. It releases code and aggregate artifacts, while detailed per-post answers and generated mirrors are available to researchers under a non-commercial agreement.
Fairness also needs its own evaluation. A 2023 study of earlier AI detectors found that they disproportionately labelled essays by non-native English writers as AI-generated. Those detectors relied on different signals from SlopShape’s structural features, so the finding does not prove this instrument has the same bias. It does establish a reason to measure false positives across writer groups before using any detector to make consequential judgments.
What AI-shaped sales pitch detection does not prove
The study’s structural patterns are descriptive correlations in a particular generated-versus-human comparison. They are not rules for good writing. A clear title, early thesis, signposted sections, and a concise conclusion can help readers. Removing those features just to seem less machine-like would be a poor engineering or editorial objective.
A better use is diagnostic. If a content system produces the same argument shape every time, inspect the workflow: Are briefs too thin? Does the prompt force a summary-first template? Does every draft lack original evidence, named sources, a real counterexample, or a decision that came from product experience? These are questions about quality and provenance, not detector evasion.
For a related discussion of what style cues can and cannot establish, see Should We Really Be Afraid of an Em Dash?. For a different approach to content provenance, see AI Watermarks Are Not a Silver Bullet. A statistical detector estimates resemblance to a training distribution. A provenance signal records information about origin. Neither, by itself, tells you whether a post is accurate, useful, or honestly disclosed.
What an AI-shaped sales pitch means for developers
SlopShape makes a compelling case that commercial AI writing can have measurable structural regularities, and that those regularities can survive substantial rewording in a controlled test. Its most practical contribution for developers is the method: represent choices explicitly, hold out whole sources, report class-aware metrics, and validate the annotator as carefully as the classifier.
The unresolved question is whether the same “AI-shaped pitch” survives when writers and models work together in ordinary publishing workflows. Until that is tested, use these results to study repeated patterns in a corpus, not to declare who wrote an individual blog post.
Sources and further reading
- Madler, “SlopShape: Identifying AI-Generated Commercial Web Content” (2026 preprint), with code and verification artifacts.
- Russell et al., “StoryScope: Investigating Idiosyncrasies in AI Fiction” (2026), the University of Maryland and Google DeepMind study SlopShape extends.
- Dugan et al., “RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors” (ACL 2024).
- Liang et al., “GPT Detectors Are Biased Against Non-Native English Writers” (2023).
Research note: SlopShape is an arXiv preprint. Its author states that he operates Sitefire, a commercial generative-engine-optimization product. The paper releases code and aggregate artifacts; access to post-level feature answers and generated mirrors is available to researchers under a non-commercial agreement.
The post Can a Blog Post Sound Human and Still Have an AI-Shaped Sales Pitch? appeared first on Alpesh Kumar.