The Experiment We Killed
We appended AI-written sections to 15 live pages. Nine read as machine-made and were rolled back the next day. The rule it taught us.
Not every experiment works, and the useful Case Notes include the ones that did not. This is one. We tried to deepen a set of live pages with AI-generated sections, measured the quality honestly, and killed it the next day.
The point is not that AI content fails. It is how it failed, why, and the standing rule that failure bought us. That rule now governs every page we touch.
Here is the verdict, counted honestly the day after.
Key takeaways
- We appended AI-written sections to 15 live pages to target near-ranking queries.
- Nine of the 15 read as obviously machine-written. We rolled them back within a day; 4 needed light rework, 2 were kept.
- The fix was not ‘never use AI.’ It was a process: one writer per page, fed the page’s own voice first, with a voice-fit check before anything ships.
The setup: a reasonable-sounding idea
Several pages ranked just below page one for queries they did not fully answer. The plan was to append two or three new sections to each, targeting those near-ranking queries, and lift the pages the rest of the way.
On paper it is sound. More depth, aimed at proven demand, on pages already close. The flaw was not the idea. It was how the writing was produced.
The method that failed, and why
A script appended new sections to 15 live pages. The agent writing them received only the existing section titles and a list of target queries. It was given no prior body text from the page and no voice profile.
So it wrote competent, generic prose with no memory of how the page actually sounded. Read in isolation each section was fine. Read inside the existing article, the seam was obvious. The voice changed mid-page.
The verdict: rolled back the next day
We reviewed every page the following day and graded it honestly. Nine of the 15 read as plainly AI-augmented and were rolled back. Four needed light rework. Two were good enough to keep.
| Outcome | Pages | What it means |
|---|---|---|
| Rolled back | 9 | Read as machine-written, restored to the prior version |
| Reworked | 4 | Salvageable with a human editing pass |
| Kept | 2 | Read as on-voice, left in place |
The rollback itself was clean because every edit had a saved prior version to restore. Eight pages were restored byte-for-byte; one from a local snapshot. Honest measurement is only useful if you can act on it, and being able to revert is part of being able to experiment.
The rule it bought
The failure was specific, so the fix could be specific. The problem was not AI. It was AI written blind to the page’s own voice. The standing rule now is:
| Rule | Why |
|---|---|
| One writer per page, not one batch | Each page gets its own pass, not a generic template applied across many |
| Feed the page’s own voice first | The writer reads the existing article before adding to it, so the seam disappears |
| Voice-fit check before shipping | A self-review gate catches off-voice writing before it reaches the live site |
| Pilot, then batch | Prove the approach on a few pages before scaling it |
That rule is why a later answer-first rewrite across the site ran the opposite way: one page at a time, drafts approved first, with saved versions ready as rollback. The same failure cannot happen the same way twice.
What this teaches
- Measure quality honestly, even when you authored it. We graded our own work as if a stranger wrote it. Nine rollbacks is not a comfortable number to report. Reporting it is what made the lesson real.
- AI fails on voice, not on facts. The sections were factually fine. They failed because they did not sound like the page. Generic-but-correct is still a tell.
- Context is the whole job. The agent wrote blind to the page’s own text. Feed a writer the existing voice first and the seam closes. The input, not the model, was the flaw.
- Always keep a clean rollback. Every edit had a saved prior version, so killing the experiment cost minutes. You can only run bold experiments if you can undo them.
- Pilot before you batch. Fifteen pages was already too many to fail on at once. Prove an approach on a handful, grade it, then scale. The pilot is the safety valve.
- A failed test that yields a rule is not wasted. This experiment produced nothing publishable and one durable process rule. The rule has paid for the failure many times since.
What this does not claim
Honesty keeps the note useful. This is not evidence that AI content cannot rank or that AI has no place in production: two of the fifteen pages were kept, and a later, more carefully run pass succeeded. It is also a quality judgment, made by us on our own pages, not a ranking penalty measured in Search. What the experiment supports is narrow and practical: AI sections written without a page’s own voice and context read as machine-made often enough to fail, and the fix is process, not abstinence.
Frequently Asked Questions
Is AI content bad for SEO?
Not inherently. Google rewards helpful, high-quality content regardless of how it was produced, and treats thin, unhelpful content the same way whoever made it. The risk with AI is quality and voice: text written without the page’s own context often reads as generic, which fails readers first and search second.
Does Google penalize AI-generated content?
Google does not penalize content for being AI-assisted. Its guidance targets content created primarily to manipulate rankings rather than help people. AI used to scale low-value pages is the risk. AI used carefully, edited for voice and accuracy, is treated like any other content.
Why did the AI sections read as machine-written?
Because the writer was given no prior body text and no voice profile, only section titles and target queries. Each section was competent in isolation but did not match how the page actually sounded, so the seam was obvious when read in context. The flaw was the missing input, not the model.
How can you use AI for content without it sounding generic?
Give it context. Feed the writer the page’s existing voice and text before it adds anything, work one page at a time rather than batching a template, and run a voice-fit review before publishing. The difference between on-voice and generic is almost entirely in the input and the editing.
What is thin content and how is it different from this?
Thin content is low-value text that does not satisfy the searcher, often mass-produced. The pilot here was not thin in facts; it failed on voice and fit. Both share a cause worth avoiding: producing pages faster than you can make them genuinely good.
Should you roll back content changes that do not work?
Yes, and you should set up to do it before you experiment. Every edit in this pilot had a saved prior version, so rolling back nine pages took minutes. Keeping clean, restorable versions is what lets you test boldly without risking the live site.
The takeaway in one line: the question is never whether to use AI, it is whether your process forces it to sound like you, because the audience can tell when it does not.
Want an honest read on whether your content sounds human and ranks for it? That is part of what a free diagnosis checks.
Continue Reading:
More From This Series
- Build the Cluster, Not the Post
- Clusters Compound. Scattered Posts Plateau.
- The CTR Fix: Same Ranking, More Clicks
- From Zero: A Cold-Start Cluster
More from TDM Insights
- Use AI to Write Without Sounding Like AI
- How to Produce Content That Ranks and Gets Cited
- Writing for E-E-A-T: Prove Experience on the Page
- Content Refresh: Win Back Rankings You Lost
- How to Write a Content Brief (That Gets It Right)
Explore TDM Insights Topics