Agentic Wikiwiki / ai-blog-draft-workflow
← Wiki index
taskbeginner

How do you write a blog post with AI without it sounding like AI?

A production process for one long blog post: gather the evidence before the first prompt, keep the model inside it, and publish against a written acceptance record. What fixes AI-sounding text is your own material, not a warmer prompt.

Job outcome

You can run one long blog post through an AI-assisted process where every fact traces to a source a human opened, and at least three elements could only have come from you.

Last verified 2026-07-27

This is a production process for one long blog post — roughly 1,200 to 3,000 words of how-to, analysis or case study — in which you assemble the evidence before the first prompt, hold the model inside that evidence, and publish against a written record of what was checked. It is a process rather than a prompt because the "sounds like AI" problem sits deeper than tone. When researchers measured what changes as people write with a model, diversity fell on two levels at once: word choice and ideas. In a controlled experiment in which 38 writers hired on Upwork for their writing or copyediting experience produced 300 essays, texts written with a feedback-tuned model were more similar to one another than texts written alone (homogenization score 0.1660 against 0.1536) and carried lower lexical and lower content diversity (Padmakumar and He, ICLR 2024 — measured in 2023 on that generation of models). Asking for a warmer voice repairs the first level and leaves the second untouched. What moves the second level is material the model cannot have: your numbers, your interviews, your dated observations.

When should you use this — and when should you not?

Use it when the post makes claims a reader could check, and when you hold raw material worth protecting — performance data from your own campaigns, an interview you conducted, a project that failed in a specific way. A model inside the process is the majority pattern: in the 2026 B2B survey by the Content Marketing Institute and MarketingProfs (n = 1,015, fielded 24 June to 14 August 2025, sponsored by CMS vendor Storyblok — worth remembering when reading how the questions are framed), 95% of respondents said their organization uses AI-powered applications, and the most-cited use was content creation tools for generating or optimizing marketing copy (89%); among the marketers using AI for content creation, 58% said content quality had improved and 12% said it had declined (CMI and MarketingProfs, 8 October 2025). In Orbit Media's twelfth blogger survey (n = 808, fielded August 2025, a self-selected sample from the author's own network, as the authors state), about one in ten respondents used AI to write complete articles (Orbit Media Studios).

One framing can go before you start: Google does not penalize text for having been produced with AI. Its spam policy names scale and intent, not method. Scaled content abuse is "when many pages are generated for the primary purpose of manipulating search rankings and not helping users", aimed at "large amounts of unoriginal content that provides little to no value to users, no matter how it's created" (Google Search Central spam policies, page last updated 2026-05-15). The announcement of 5 March 2024 adds that this holds "no matter whether content is produced through automation, human efforts, or some combination of human and automated processes" (Google Search Central blog), and the earlier guidance is blunter: "Appropriate use of AI or automation is not against our guidelines" (Sullivan and Nelson, 8 February 2023).

Skip the process in these situations:

Situation What to do instead
You hold no raw material of your own on the topic Go get some — one interview, one small experiment, one set of your own numbers. Run the process afterwards; no prompt substitutes for the missing input.
A short piece resting on a single source: an announcement, a personal essay, an incident note Write it directly. The evidence pack and the acceptance log cost more than they return at that length.
Health, money or legal topics without review by a qualified expert Route the draft through that expert first; Google's helpful-content guidance holds such topics to stronger trust signals (Creating helpful, reliable, people-first content).
The goal is "N posts a week to cover a keyword cluster" Stop. That is the definition of scaled content abuse quoted above, and it applies whoever or whatever writes the pages.
A contract or an editorial policy restricts AI assistance Settle it before drafting — this is a question of consent, not craft. See ai-content-pre-publish-compliance.
Your work will be judged by someone who runs an AI detector Negotiate a different gate, using the evidence in the failure section below. Rewriting to please a detector degrades the text.

What do you need before you start?

This page assumes the acceptance habits taught earlier in the ai-for-marketers path; if you have not met them, ai-draft-pre-publish-check is the shortest entry point.

Input What it is Where it comes from
Proprietary raw material At least three items a model could not have: your own metrics, a quote from an interview you ran, a dated observation, a breakdown of something that went wrong Your analytics, your calendar, your inbox, a 30-minute call with a colleague or customer
Reader and promise One paragraph: who this is for, what they can do afterwards, what it does not promise You write it. Not the model — this is the sentence everything else is measured against
Evidence pack A table with one row per claim: claim, source URL, exact quote or number, date accessed, type (own / primary / secondary) Built by hand before the first prompt
Voice sample Two or three of your own texts you like, plus a list of words and turns of phrase you never use Your archive; if you maintain one, brand-voice-doc-for-ai already holds it
Structural constraint Target length, required sections, publication format The brief — see seo-content-brief-with-ai
Disclosure decision Your answer to "would a reader wonder how this was made?" and what you will do about it Decided before drafting, recorded in the acceptance log

On that last row: Google states no disclosure requirement. Its FAQ says disclosures "are useful for content where someone might think 'How was this created?'" and that giving AI an author byline is "probably not the best way" to follow its recommendation (Sullivan and Nelson, 8 February 2023); the "Who, How, and Why" self-assessment asks whether the use of automation is self-evident to visitors (Creating helpful, reliable, people-first content). There is no rule to comply with, only a question to answer — and answering it before drafting is cheaper than answering it in a comment thread.

How do you do it, step by step?

1. Build the evidence pack with the model switched off. Every claim you intend to make gets a row. Anything you cannot back with a link you opened or an artifact you own does not enter the pack, which means it cannot enter the post. Use this row shape:

| # | claim (one sentence) | source URL | exact quote or number | accessed | type: own / primary / secondary |

At least three rows must be typed own. That count is not decoration: it is the only part of the post that no competitor and no model can reproduce.

2. Write the promise yourself, then have the model attack it. You write one paragraph; the model looks for claims the pack cannot support.

You are a skeptical editor. Below is a one-paragraph promise for an article and the
evidence pack meant to back it. Do NOT rewrite the promise.

PROMISE: [promise paragraph]
EVIDENCE PACK: [paste table]

Return exactly this structure:
1. CLAIMS IN PROMISE - every distinct claim the promise makes, as bullets.
2. UNBACKED - for each claim with no matching evidence row: claim | why unbacked.
3. OVERREACH - claims asserting cause and effect the evidence cannot support.
4. NARROWER PROMISE - one alternative paragraph the pack fully supports.
Do not add facts. Use no information outside the evidence pack.

3. Build the outline out of the pack, not out of a template. Each section has to earn its place from specific rows. A section with no rows behind it is a research gap, not an invitation to write generally. Ask for an explicit shape per section as well — since the measured loss of diversity reaches structure and content, not only vocabulary, varying shape on purpose is part of the fix.

Build an outline for a [target length]-word article for [reader description].

Constraints:
- Every section MUST be justified by at least one evidence row, cited by row number.
- If a necessary section has no supporting row, list it under NEEDS RESEARCH instead.
  Never fill it with general knowledge.
- Order sections by the reader's decision sequence, not by a template.
- Label each section with one shape: [narrative | table | steps | example | counterargument].
  No more than two consecutive sections may share a shape.

PROMISE: [promise paragraph]
EVIDENCE PACK: [paste table]

Output format:
## OUTLINE
- H2 title | shape | evidence rows: [#, #] | what the reader can do after this section
## NEEDS RESEARCH
- topic | why it matters | what source would settle it

4. Draft one section at a time under a fact lock. The model gets only the rows for that section and is forbidden to supply facts from anywhere else.

Write ONE section of the article: "[section H2]".

Hard rules:
- Use ONLY facts from these evidence rows: [rows].
- If a sentence needs a fact that is not in those rows, write [GAP: what is missing]
  instead of inventing or approximating it. Never write an unsourced number, name,
  date or study.
- Attribute every number in plain language: who measured it, when, on what sample.
- Target [N] words. Shape: [narrative | table | steps | example | counterargument].
- Write for [reader description]. Assume no technical background.
- Do not open with a definition or a scene-setting framing.

EVIDENCE ROWS: [paste only this section's rows]
VOICE SAMPLE (imitate rhythm and vocabulary, not content): [200-300 words of your own writing]

Output: the section text, then a final line "SOURCES USED: [row numbers]".

5. Trace every factual sentence back to a row. Ask for the mapping, then resolve it by hand.

Below is a draft and the evidence pack meant to back it. For EVERY sentence stating
a fact - a number, a date, a named study, a platform policy, a quote, an attribution -
output one row:

| sentence (verbatim) | claimed source row | status |

status is exactly one of: MAPPED (the row supports the sentence as written),
DRIFT (a row exists but the sentence overstates, rounds or shifts its meaning),
UNMAPPED (no row supports it). List DRIFT and UNMAPPED first. Fix nothing.

DRAFT: [paste]
EVIDENCE PACK: [paste]

Then open every URL yourself and confirm the quote is on the page: the model does not verify its own links, and the mapping is a hypothesis until a human checks it. spot-fabricated-stats covers what a plausible-but-invented citation looks like; ai-draft-pre-publish-check covers the wider pre-publication sweep.

6. Replace the generic, then count the markers. First pass: find every sentence that would still be true if you swapped the industry, the company and the year, and replace it with something from an own row — or delete it. Second pass is mechanical and needs no model: count occurrences of a watchlist of style words and record a decision for each hit. The watchlist with the strongest published backing comes from an "excess vocabulary" analysis of 15,103,888 English-language PubMed abstracts from 2010 to 2024, which found 379 excess style words in 2024 texts, led by delves (28x the expected frequency), underscores (13.8x) and showcasing (10.7x) (Kobak et al., Science Advances, 2 July 2025). The authors' published list of all 900 excess words they identified from 2013 to 2024 also annotates intricate, meticulously, crucial, pivotal, realm, notably, comprehensive, garnered and aligns as style words (excess-word list, Kobak et al.). That corpus is scientific abstracts, not marketing blogs — carrying the list over to blog copy is a plausible extrapolation, not a measured result, and no equivalent public measurement on web content turned up in our search. Deleting these words does not improve a text by itself; the count is there to show you where you handed over your voice along with the typing. Editing method: edit-ai-draft. Repeatable scoring: score-ai-output.

7. Add what the model cannot know, then close the log. Write the parts no model has access to: what went wrong, what you still do not know, where you changed your mind. Put a human name on the byline, record the disclosure decision, and store the acceptance log beside the article rather than in a folder nobody finds in six months.

What does a good result look like?

The article is the visible output; the artifacts make it defensible. Filled in, they look like this — the rows below come from this page's own evidence pack:

# claim source URL exact quote or number accessed type
4 The best AI detector tested scored below 80% eprints.whiterose.ac.uk (Weber-Wulff et al. 2023) "even the highest accuracy values are below 80%"; best tool 76% 2026-07-27 primary
7 AI-assisted writing lowers corpus diversity arxiv.org/abs/2309.05196 homogenization 0.1660 vs 0.1536; 38 writers, 300 essays 2026-07-27 primary
9 No public equivalent of that marker measurement exists for marketing blogs our own research log, 2026-07-27 search on 2026-07-27 found no corpus study of web content 2026-07-27 own

The traceability pass then looks like this. Only the first row would clear:

sentence claimed row status
"The best tool tested reached 76% accuracy." 4 MAPPED
"Detectors are wrong most of the time." 4 DRIFT — the row says below 80%, not below 50%
"Most editors now run detectors." UNMAPPED — cut it or go find a source

A post that ships in this shape can be handed to another person, six months later, and re-checked without your presence. That is the actual deliverable.

How do you know the output is good?

The property shared by every criterion above is that someone who did not write the post can answer it from the artifacts, without rereading the article and without asking your opinion. "Does it sound human?" fails that test; "is there a row for this sentence?" passes it. Most of the checks are counting exercises, one is a text search, one is a record — none needs judgment about voice.

Two omissions are deliberate. There is no detector score, for the reasons in the next section. And there is no threshold on the marker count — no "fewer than N style words per 1,000". The published measurement gives frequency ratios across a corpus, not an editorial cutoff for a blog post (Kobak et al.), and inventing a number would plant a fabricated statistic inside a workflow built to catch fabricated statistics. The record itself is the criterion: the count exists, and each hit has a decision written beside it.

Disclosure works the same way. Google states no requirement, only a question (Sullivan and Nelson, 8 February 2023), so failing a post for not disclosing would enforce a rule Google has not issued — whatever your own publisher or client requires is a separate matter. What is required is that a decision was made deliberately and can be produced on request.

One failure signature is worth naming: if the traceability pass returns rows you cannot resolve without going back to research, the outline ran ahead of the evidence. The repair is step 3, not a heavier edit.

Where does this usually break?

Failure What it looks like What stops it
Invented source A confident sentence with a real-sounding study, number or URL behind it The fact lock in step 4 plus manual URL opening in step 5; spot-fabricated-stats
Detector used as the acceptance gate Rewriting exactly the passages a tool dislikes, which are usually the plainest ones Replace the gate with the checks above
Paraphrase laundering Running the draft through a rewriting tool until a detector goes quiet Name the trade honestly — see below
Voice prompt instead of substance "Write conversationally, use contractions", and the article still says what twenty others say Steps 1 and 6a: specificity comes from own rows
The workflow scaled into a page factory The same pipeline aimed at a keyword cluster, N posts a week Volume changes the class of the site, not the quality of one post
The trail lost at publication The draft goes into the CMS; the evidence pack stays in someone's folder Store the acceptance log with the article — step 7
Writing worse to look human Adding ornament so the prose does not read as machine-plain See the warning below
Folklore markers Social-media checklists of "AI tells" applied as rules Use only markers whose measurement you can name

Detector accuracy was measured in 2023, on that generation of models, and even then the instruments failed. Across seven detectors run on 91 TOEFL essays by non-native English writers, the average false positive rate was 61.22%, against 5.19% on 88 essays by US eighth-graders; 18 of the 91 essays were flagged as AI by all seven tools at once (Liang et al., Patterns, 2023). A separate study of 14 tools across 54 test cases found the best performer at 76% and stated plainly that "even the highest accuracy values are below 80%", concluding the tools are "neither accurate nor reliable" and biased toward calling text human-written (Weber-Wulff et al., 2023). OpenAI withdrew its own classifier with the note "As of July 20, 2023, the AI classifier is no longer available due to its low rate of accuracy" (Search Engine Journal, 25 July 2023; same quote reported by Search Engine Land). No independent re-run of comparable rigor on 2025–2026 models surfaced in our search, so read these figures as measurements of that generation, not as a verdict on today's tools.

The paraphrase result is the one to sit with. In the same 14-tool study, the case labeled "machine-generated with subsequent machine paraphrase" scored 26% overall accuracy — the lowest of the six document types tested (Weber-Wulff et al.). Read it backwards: the single action most reliable at getting text past a detector adds no fact, no example and no experience to that text. Optimizing for the signal walks away from the thing the signal was meant to stand for.

Surface-level voice fixes fail for a related reason. The measured loss of diversity reaches content, not only wording: in an experiment with 293 writers and 600 evaluators, access to AI-generated ideas raised individual creativity ratings while making the resulting stories more similar to one another (Doshi and Hauser, Science Advances, 2024). An article can sound like you and still say what everyone else said.

What else do people ask?

Do I have to tell readers that AI was involved?

Google does not require it as of 2026-07-27: it frames disclosure as a question rather than a rule, and calls giving AI an author byline "probably not the best way" to make AI's role clear to readers (Sullivan and Nelson, 8 February 2023). We checked Google, not every publisher. Our position — a position, not a platform rule — is that a named human byline plus a written answer to "would a reader wonder how this was made?" costs nothing and survives scrutiny. Client contracts and internal policies are a separate question: ai-content-pre-publish-compliance.

Is an em dash a sign of AI writing?

We found no primary measurement of it. The published marker work covers vocabulary, not punctuation (Kobak et al., 2025), and the same holds for the other circulating "tells" such as sentence triplets and the "not X, but Y" construction. Treat any marker whose measurement you cannot name as folklore, and do not let it drive edits.

What if I have no material of my own?

Then this process yields a competent summary of what already exists, and no amount of prompting changes that. The next step is to acquire material: one customer interview, one small experiment you can report on, one dated observation of your own work. customer-research-synthesis covers turning interview material into evidence without inventing a consensus the interviews do not support.

Sources

Verification

4 log entries
dateactionresult
2026-07-27researchapplied
2026-07-27draftapplied
2026-07-27correctionapplied
2026-07-27fact-checkpass-3-0

Backlinks