How do you check AI content before you publish it?
An ordered pass for checking AI-written marketing content before it goes live: pull out the claims, trace them outside the AI tool, verify numbers and quotes, clear regulated claims, test whether the piece adds anything, record a decision.
You can check AI-written content the way an editor would: every claim traced to a source you opened, every risky claim backed or removed, and a decision on record.
Last verified 2026-07-27
A pre-publish check is a fixed, ordered pass over AI-written content, run before anyone outside your team sees it. It is not proofreading: proofreading asks whether the sentences read well, this pass asks which sentences say something about the world, whether each of those survives contact with a source you opened yourself, and whether the piece earns its place at all. It ends in one of three decisions you write down — ship, fix and re-check, or scrap and re-brief. The order below is cheapest kill first, so you find out whether the draft is salvageable before you spend an hour making it sound like you.
When should you use this — and when should you not?
Run it on anything an assistant helped write that will be published under your name or your company's: blog posts, landing pages, emails, case studies, social posts with a claim in them.
Skip it or replace it when:
- The draft makes no checkable statements — a short caption, an anecdote, pure opinion. Go straight to the substance test.
- Nothing outside the team will see it. The pass protects a published claim, and there isn't one.
- The piece is regulated or high-stakes — health, finance, legal, employment, or any promise about results. The pass is necessary but not sufficient here: see the compliance review.
- You gave the model no sources and the topic moves fast. Checking a draft whose every fact was invented costs more than regenerating it with a source pack.
- You have already blown your scrap threshold. Patching sentences leaves the argument standing on premises that do not exist.
- It is a translation or localisation — a different failure set, and this pass under-detects there.
What do you need before you start?
Open all of it before the first check. Without these the pass quietly degrades into proofreading.
| Have this open | Why it matters |
|---|---|
| The brief the draft came from | The substance test measures against intent, not against "good writing" |
| The source pack you gave the model | Separates "it repeated my source" from "it invented one" |
| The full unedited output | Once you start smoothing sentences you lose the record of what was asserted |
| A browser, outside the AI tool | Checking inside the tool that produced the claim is not verification |
| The files that back your claims | A claim you cannot back is a claim you delete |
| A scrap threshold, in writing | The decision gets harder once you are attached to the draft |
| A named editorial owner | Someone has to own the published claim |
Budget real time for the check, and distrust your sense of how long it took. In a randomised trial of 16 experienced open-source developers across 246 tasks, participants took 19% longer when allowed to use early-2025 AI tools, having predicted a 24% speed-up and still believing afterwards that they had been sped up by 20% (METR, published 10 July 2025). METR now marks that productivity result out of date and reports weak evidence of a speed-up with later tools (METR, 24 February 2026). It is software rather than marketing either way, so read it as evidence about self-assessed speed, not about your own throughput.
How do you do it, step by step?
What does the draft actually claim?
Paste the unedited draft into a fresh chat and ask for mechanical extraction, not judgement. You are using the model as a highlighter.
You are an extraction tool, not an editor. Do not fix, rate, or comment on the text.
From the draft below, extract EVERY statement a reader could check against the outside
world: numbers, percentages, dates, money amounts, named people, named organisations,
product names, quotations, cited sources or links, superlatives ("the most", "the first"),
and any statement of cause and effect.
Output a Markdown table with exactly these columns:
| # | Verbatim sentence | Claim type (number / date / entity / quote / link / causal / superlative / regulated) | Source named in the draft (URL or "none") |
Rules:
- One row per claim. If a sentence contains two claims, give two rows.
- Column 2 must be copied verbatim. Do not paraphrase.
- If the draft names no source, write "none". Do not guess a source.
- Do not add claims that are not in the draft.
DRAFT:
[PASTE FULL UNEDITED DRAFT]
Write down the row count — it is the denominator for everything that follows.
Do not ask "is anything here inaccurate?" instead. On reasoning tasks, models revising their own output with no external feedback "struggle to self-correct… and at times, their performance even degrades after self-correction" (Huang et al., ICLR 2024), which is exactly why every step below hands the model something external to compare against. Anthropic makes the same point from the vendor side: its recommended mitigations "significantly reduce hallucinations, [but] they don't eliminate them entirely. Always validate critical information" (Anthropic documentation).
Which claims survive outside the AI tool?
In a browser, by you. For each row, open the named source or search for the claim, and find the sentence that supports it. Two established moves do most of the work: SIFT — stop, investigate the source, find better coverage, trace claims and quotes to their original context — and lateral reading, leaving the page you are judging to see what other sources say about it (University of Chicago Library, adapted from Mike Caulfield).
Record one of exactly four verdicts, plus the URL you opened:
| Verdict | What it means |
|---|---|
TRACED |
You opened a source and read a sentence that supports the claim as written |
DRIFTED |
A real source exists, but says something narrower, older or different |
NO SOURCE |
You found nothing that supports it |
FABRICATED |
The cited source, study, quote or URL does not exist |
The trap worth naming: a link that opens is not a verified claim. A May 2026 preprint that tested 14 models on 130 research queries reports that even the strongest frontier models kept link validity above 94% and topical relevance above 80%, while their factual accuracy against the cited sources ran at only 39–77% — and "link validity" there means only that the URL returned accessible content (arXiv 2605.06635 — preprint, not peer-reviewed). Deep source archaeology on one suspicious statistic belongs to catching fabricated statistics; here you need only a verdict per row.
Do the numbers, dates and quotes match exactly?
Re-check the TRACED rows of type number, date and quote character by character. A claim can point at the right source and still carry the wrong digit. In the BBC's assessment of 100 news questions published in February 2025, 19% of answers citing BBC content introduced factual errors — wrong statements, numbers and dates — and 13% of quotes attributed to BBC articles were altered or absent from the article cited (as reported by The Register, 12 February 2025). The larger EBU follow-up, covering over 3,000 responses generated in late May and early June 2025 and rated by 22 public service media organisations, found 45% of responses had at least one significant issue and 31% had significant sourcing issues (EBU and BBC, report published October 2025 — note that these are news organisations with a commercial position opposite AI companies).
Use the model as a diff engine and adjudicate yourself:
Below are (A) sentences from my draft and (B) passages I copied verbatim from the
original sources. For each pair report ONLY differences, as a table:
| # | Draft says | Source says | Difference type |
Difference types: NUMBER MISMATCH / DATE MISMATCH / UNITS OR BASE CHANGED /
SCOPE BROADENED / HEDGE REMOVED / QUOTE ALTERED / NO DIFFERENCE.
If the source passage does not mention the claim at all, write NOT PRESENT IN SOURCE.
Do not judge which version is correct. Do not rewrite.
A — DRAFT SENTENCES:
[PASTE]
B — SOURCE PASSAGES (verbatim, copied by me):
[PASTE]
Every HEDGE REMOVED and SCOPE BROADENED is a real defect: that is how a cautious finding turns into a marketing overclaim. Attach an as-of date to every surviving number while you are here — "45% of responses had a significant issue" means nothing without the study and the period it covers.
Which claims need backing before they can run?
Walk the table for the regulated, causal and superlative rows, plus every testimonial, review and customer quote. This is a gate, not a judgement call: each item has a named backing file, or it comes out. An AI-written customer quote is not a stylistic choice — the US Federal Trade Commission's rule on consumer reviews and testimonials (16 CFR part 465, effective 21 October 2024) covers consumer reviews and testimonials that materially misrepresent that the reviewer exists or had the experience described, and the Commission stated that "AI-generated reviews are covered by the final rule" (Federal Register, 22 August 2024).
Read the draft below and list, as a table, every sentence that is any of:
(a) a claim about a result, outcome, saving, gain or performance;
(b) a health, medical, financial, legal or safety claim;
(c) a testimonial, review, quotation or first-person customer experience;
(d) a comparison to a named competitor;
(e) a superlative ("best", "fastest", "only", "#1").
| # | Sentence | Category (a-e) | What evidence would a regulator expect here? |
Do NOT assess whether the claim is legal or acceptable. Do NOT rewrite. Output the table only.
DRAFT:
[PASTE]
Resolve every row to BACKED (file: ___), SOFTENED TO NON-CLAIM, or REMOVED. Regulated-industry review goes deeper in the compliance page.
Does the piece add anything a generic answer would not?
Open a fresh chat with no memory, no context and no source pack, ask it the one question your piece answers, then diff that answer against your draft.
PROMPT 1
Answer this question for a [AUDIENCE] in about [WORD COUNT] words.
Do not use any sources, files or browsing. Plain generic answer only.
QUESTION: [THE ONE QUESTION YOUR PIECE ANSWERS]
PROMPT 2
Compare TEXT A (my draft) with TEXT B (the generic answer). Output two lists:
LIST 1 — substantive elements present in A but absent from B: specific data with a date,
a named source, a first-hand example, a concrete procedure, a stated trade-off,
a counter-argument, a stopping rule.
LIST 2 — elements present in both.
Do not evaluate quality. Do not rewrite. Quote the sentence for each item.
TEXT A: [DRAFT]
TEXT B: [GENERIC ANSWER]
If list 1 is empty, the piece has no reason to exist. Scrap it. Google's guidance points the same way without promising anything about rankings: it tells creators to "focus on accuracy, quality, and relevance", and warns that using generative AI to produce many pages without adding value for users may fall under its scaled content abuse policy (Google Search Central, last updated 10 December 2025).
And what about voice and polish?
Last, and bounded: read it aloud, cut the hedges the model piles up, check that the promise in the intro is the one delivered. Two non-obvious rules belong here — do not use an AI-text detector as your publish gate, and do not read fluency as accuracy (both are in the table below). Rewriting is a separate job: see editing an AI draft to a publishable standard and building a brand voice document.
What do you record before you hit publish?
Three lines in the document, not in your head:
Claims extracted: N
Verdicts: TRACED __ / DRIFTED __ / NO SOURCE __ / FABRICATED __
Decision: SHIP | FIX AND RE-CHECK | SCRAP AND RE-BRIEF
Editorial owner: [name]
The point of that block is the scrap threshold you set earlier. A workable default is to scrap rather than patch once FABRICATED plus NO SOURCE passes roughly a third of the claims — but treat that as a convention you adopt so the decision is not made under sunk-cost pressure, not as a measured finding. We found no study establishing where patching becomes more expensive than regenerating.
What does a good result look like?
A finished pass on a 900-word post for an invented product, Northgate Analytics:
| # | Claim type | Verdict | What happened |
|---|---|---|---|
| 1 | number | FABRICATED |
"73% of teams report…" attributed to an analyst firm with no such report; sentence deleted |
| 2 | number | DRIFTED |
Source says "among enterprise respondents", draft said "among all marketers"; scope restored |
| 3 | quote | TRACED |
Opened the source, matched word for word, hedge intact |
| 4 | link | TRACED |
Page opens and contains the supporting sentence, quoted in the worksheet |
| 5 | regulated | removed | AI-written customer testimonial; no real customer, no permission |
Claims extracted: 12
Verdicts: TRACED 9 / DRIFTED 1 / NO SOURCE 0 / FABRICATED 2
Decision: FIX AND RE-CHECK
Editorial owner: Dana R.
The artefacts are the point: a second person can pick up the worksheet, open four URLs and reach the same verdicts without asking you a single question.
How do you know the output is good?
- Every checkable statement in the draft has a written verdict — traced, drifted, no source, or fabricated — with no blanks left.
- Nothing marked fabricated or no source survives in the published version; it was cut, or rewritten so it no longer asserts anything.
- For every source cited in the published piece, you opened the page yourself and can point to the sentence that supports the claim.
- Every number, date and quotation matches its source word for word, keeps the source's hedges, and carries the date the data refers to.
- Every results claim, testimonial and customer quote is either backed by a named file you can produce on request, or it is gone.
- The piece contains at least one thing a reader could not get by asking any assistant the same question, and the decision and the named owner are written down.
The criterion about opening the source yourself is the one that fails silently, because a citation that resolves feels like a citation that checks out. The measurement above separates the two properties: among the strongest frontier models, link validity above 94% alongside factual accuracy of 39–77% means most broken citations look perfectly healthy from the outside.
The criterion about numbers, dates and quotes is where careful sources turn into overclaims. "May reduce" becoming "reduces" is a two-word edit that changes what you are asserting, and no spell-check will flag it.
The last criterion separates checked from worth publishing. Everything before it removes defects; only the generic-answer diff asks whether the piece should exist. Turning that judgement into a repeatable score is its own method — see scoring AI output consistently.
Test the whole set by handing the worksheet and the final text to a colleague who did not write the brief. If they cannot answer every criterion yes or no from the artefacts alone, the pass was not finished.
Where does this usually break?
| Failure | What it looks like | Why it happens |
|---|---|---|
| Asking the model to check itself | "Is anything here inaccurate?" returns a confident "looks good" | Self-correction on reasoning tasks without external feedback is unreliable and can degrade output (Huang et al., ICLR 2024) |
| Treating a working link as verification | The citation opens, so the sentence is accepted | Link validity and factual support are separate properties: among the strongest frontier models, above 94% validity against 39–77% accuracy (arXiv 2605.06635, preprint) |
| Using an AI-text detector as the gate | Green light, ship it | OpenAI withdrew its own classifier on 20 July 2023; on its own challenge set it caught 26% of AI-written text while flagging 9% of human text as AI (reported figures) |
| Rubber-stamp review | Someone "reviewed" it; nothing was checked | Clicking approve leaves no artefact a second person can re-check |
| The speed illusion | The check is skipped because AI already saved the time | Self-reported and measured time diverge, and the gap survives direct experience (METR, 2025) |
| Reading fluency as accuracy | Confident prose trusted, hedged prose edited out | Training and evaluation "reward guessing over acknowledging uncertainty", so output is most assertive where it is least grounded (Kalai et al., September 2025) |
| Surface editing only | Voice fixed, substance untouched | Work that looks finished and pushes the thinking downstream cost a surveyed group of US desk workers about two hours per incident to sort out (BetterUp and Stanford Social Media Lab, September 2025, n = 1,150; BetterUp sells workforce products and commissioned the research) |
| Trusting long research chains | "It used 200 sources, so it must be solid" | In an ablation on two frontier models, factual accuracy fell about 42 points on average as tool calls grew from 2 to 150 (79%→17% and 80%→58%), while surface citation metrics stayed above 92% (arXiv 2605.06635, preprint) |
One honest limitation: we found no study showing that a structured pre-publish check catches more defects than an unstructured read. The evidence above establishes that these failures are real; the ordering of the steps is a design argument, not a measured result.
What else do people ask?
Do I have to tell readers the content was AI-assisted?
It depends on where you publish and what the piece is, and the answer is not ours to give. Under EU AI Act Article 50(4), which applies from 2 August 2026, deployers of a system that generates or manipulates text "published with the purpose of informing the public on matters of public interest" must disclose that it was artificially generated — and the obligation does not apply where the content "has undergone a process of human review or editorial control and where a natural or legal person holds editorial responsibility for the publication" (Article 50; European Commission guidelines). Whether ordinary product marketing counts as informing the public on matters of public interest is genuinely unsettled; much of it plausibly does not. Separately, Google suggests you "consider adding information on how your content was created in a way that makes sense for your audience" (Google Search Central). Client contracts and platform rules may add more — see the compliance page.
Can I skip tracing if the tool already shows citations?
No, and their presence tells you less than it appears to. Some assistants cite only when browsing is switched on and return nothing when answering from memory; some display sources found after the fact rather than sources used to write the sentence. Citation display is a product feature that changes between releases. The claim either traces to a sentence you read, or it does not.
Where does this fit in a repeatable process?
This pass is the acceptance gate. Briefing the model so the draft arrives with fewer defects is upstream — see the blog-draft workflow — and the full marketing track is laid out in AI for marketers.
Want the next practical guide?
Sources
- Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents — arXiv 2605.06635 (preprint)accessed 2026-07-27
- EBU and BBC — News Integrity in AI Assistants, international PSM study (full report, October 2025)accessed 2026-07-27
- The Register — report on the BBC's February 2025 assessment of AI assistants and newsaccessed 2026-07-27
- Large Language Models Cannot Self-Correct Reasoning Yet — Huang et al., ICLR 2024accessed 2026-07-27
- Anthropic documentation — Reduce hallucinationsaccessed 2026-07-27
- BetterUp Labs and Stanford Social Media Lab — workslop researchaccessed 2026-07-27
- Google Search Central — Using generative AI contentaccessed 2026-07-27
- METR — Measuring the impact of early-2025 AI on experienced open-source developer productivityaccessed 2026-07-27
- METR — We are Changing our Developer Productivity Experiment Design (24 February 2026)accessed 2026-07-27
- Why Language Models Hallucinate — Kalai et al., arXiv 2509.04664accessed 2026-07-27
- Federal Register — Trade Regulation Rule on the Use of Consumer Reviews and Testimonials, 16 CFR part 465accessed 2026-07-27
- EU AI Act — Article 50, transparency obligations (consolidated text)accessed 2026-07-27
- European Commission — Guidelines on Article 50 transparency obligationsaccessed 2026-07-27
- Synthedia — OpenAI shuts down its AI-written text classifieraccessed 2026-07-27
- University of Chicago Library — the SIFT method, adapted from Mike Caulfield (CC BY 4.0)accessed 2026-07-27
Verification
4 log entries
| date | action | result |
|---|---|---|
| 2026-07-27 | research | applied |
| 2026-07-27 | draft | applied |
| 2026-07-27 | correction | applied |
| 2026-07-27 | fact-check | pass-3-0 |
Backlinks
- How do you pick which AI-generated ads are worth testing?
- How do you write a blog post with AI without it sounding like AI?
- What can actually get you in trouble when you publish AI content?
- AI for marketers
- How do you get AI to actually sound like your brand?
- How do you turn customer feedback into copy without AI inventing the pattern?
- How much of an AI draft do you actually rewrite?
- How do you know if AI output is any good?
- How do you catch AI hallucinations before they reach a client?