Agentic Wikiwiki / edit-ai-draft
← Wiki index
taskbeginner

How much of an AI draft do you actually rewrite?

Edit AI content by consequence: verify claims, cut model-shaped structure, replace generic passages with your own material, then fix language on a timer. Includes a stop rule — and the finding that the widely quoted "25–40% revision benchmark" has no source behind it.

Job outcome

You can decide what to change in AI content, in what order, and stop on a rule someone else can check — instead of editing until you lose interest.

Last verified 2026-07-27

Editing AI content is triage, not rewriting. What decides your afternoon is which defects you fix first, which you leave alone, and what tells you to stop. Below: four passes ordered by what each defect costs if it ships, and a stop rule someone else can check. Producing a better first draft is a different job (ai-blog-draft-workflow), as is the final gate before publishing (ai-draft-pre-publish-check).

Is there a number that tells you how much to rewrite?

No. As of 2026-07-27 the figure in circulation traces back to no measurement. The guide that says "human revision should land in the 25–40% range of the original draft" presents it as a settled 2026 benchmark and cites nothing for it (Atom Writer, April 2026); a competing guide budgets "45 minutes to an hour" for the same 1,500-word draft where that one budgets 25–35 minutes, also unsourced (briefer, March 2026). It is unusable anyway: "percent revised" needs a reference text to diff against, and a blog post has none.

What research measures is retention, not effort. In keystroke-logged sessions, 38 writers hired on Upwork for prior writing or copyediting experience accepted about 70% of the suggestions they requested, and the model had written about a third of the characters in the finished essays — 32% in the feedback-tuned condition (Padmakumar & He, ICLR 2024, 2023-era models). That is a measure of how much machine text is in the finished piece, not of how long the edit took.

So the stopping criterion is a set of checks and a clock, not a percentage.

When should you use this — and when should you not?

Editing is the right frame only when the draft is already near the job. In a preregistered experiment with 758 knowledge workers at a global consulting firm, those with AI completed tasks inside the model's competence "25.1% more quickly"; on a managerial task chosen to sit outside it they were "19% less likely to produce correct solutions" (Dell'Acqua et al., Organization Science, 2026 — quoted from the published abstract, full text paywalled; the experiment ran on GPT-4 and first circulated as a 2023 working paper).

Do not edit — do this instead Why editing fails here
The draft answers a different question, or the wrong reader You would spend the budget relocating a text. Re-brief.
The value depends on information the model never had — your data, your customer's words, a product you used Editing cannot manufacture first-hand material; Google's quality questions ask for "original information, reporting, research, or analysis" (Google Search Central).
Short, high-stakes copy: ad headline, pricing page, subject line Per-word overhead exceeds writing it (ad-copy-variants-with-ai).
Regulated claims — health, finance, employment, safety Needs a compliance pass with a named reviewer (ai-content-pre-publish-compliance).
Your only goal is to make it "not look AI-written" Spend with no measurable return; see the last question below.
You have already rewritten more than half of it You are writing. Start the clock honestly.

What do you need before you start?

The brief in one line — reader, the one job, the single claim — is what Pass 0 judges against; if you cannot write it, the problem is upstream and editing does not reach it. The source list is what Pass 1 checks against. Two or three pieces you would be happy to be judged by give Pass 4 a comparison target instead of a vibe; a written voice document does this better (brand-voice-doc-for-ai). One piece of material only you have — an analytics number, a customer sentence, a screenshot — is what Pass 3 inserts. A time budget, written down (say, 25 minutes for a 900-word post) is what makes the stop rule evaluable.

How do you do it, step by step?

Defects are not equally expensive, and the order is the method. Spend from the top; the bottom row is the one you may leave unfinished.

Tier Defect Cost if you publish it Pass
Outcome-bearing Fabricated or misattributed fact, number, quote, source Wrong in public; the fix costs more than the page earned 1
Outcome-bearing Answers the wrong question or the wrong reader The page does no job at all 0
Outcome-bearing Nothing in it only you could have supplied Interchangeable with every competitor's page 3
Reader-cost Padding: restated conclusions, symmetric bullets, filler triads The reader leaves before the point 2
Reader-cost A hedge concealing that nobody checked Either an unsourced claim or an unusable one 1, then 4
Cosmetic Marker vocabulary, em dashes, mechanical bolding, title case Some readers notice; no measured effect either way 4, time-boxed

One rule runs across all passes. Leave alone sentences already clear, accurate and on-voice; structure that serves the task, even though a model chose it; hedges that are true, since an uncertain claim should hedge and you only remove one you can replace with a source; and any single em dash or triad, because the diagnostic is the pattern, not the instance.

Pass 0 — is this draft worth keeping? (2 minutes)

Read the first screen and the H2s only. Answer three questions in writing: does this answer the brief's question, is this the structure you would have chosen, is there one paragraph you would keep verbatim. Two "no" answers kill the draft — re-brief rather than repair, and log the kill.

Decide the outline once, explicitly, because it goes invisible later. Writers co-writing with an opinionated assistant wrote 63% of their own sentences and still shifted towards the assistant's framing; those who spent four to six minutes on the post still shifted 0.20 on the measured scale, against 0.38 for the fastest (Jakesch et al., CHI 2023). Working slowly does not undo the anchor.

Pass 1 — can every claim be traced to a source you opened? (40% of the budget)

Verification is where the work sits: fact-checking and hallucination review was the largest single source of extra work for 2,003 marketing leaders surveyed for Optimizely by Savanta, named by 48% (fieldwork May–June 2026; Optimizely — vendor-commissioned, questionnaire not public as of 2026-07-27). In an audit by 22 public-service broadcasters, 31% of the 2,709 evaluated answers had significant sourcing problems (EBU / BBC, October 2025 — news questions put to consumer assistants in mid-2025, so a reason to open sources yourself, not an error rate for your drafts).

Extract every checkable assertion, then open the source yourself — not the one the model names, the one you can reach. Each claim ends as verified with a URL, softened into your own reasoning without false precision, or cut (spot-fabricated-stats covers the fabricated-source case).

The point is not that machines lie. A peer-reviewed CHI paper describes an earlier study as finding "a 74% average reduction in time" from post-editing machine translation (Green, Heer & Manning, CHI 2013). In the original, 74% is the gain in throughput; the time saving is 43%, and the paper prints the arithmetic (Plitt & Masselot, PBML 93, 2010). Peer review at both ends, no AI anywhere, a factor-of-1.7 error — because nobody opened the source.

Extract every checkable assertion from the text below. Do not verify them,
do not add sources, do not comment on quality.

TEXT: [paste draft]

Output a markdown table, one row per assertion, with columns:
| # | Assertion (verbatim) | Type (number / date / name / quote / capability / causal claim) | Source named in text? (yes: [what] / no) |

Rules:
- Copy assertions verbatim. Do not paraphrase.
- Include claims stated as background or common knowledge.
- If the text implies causation ("X drives Y"), list it as a causal claim.

You add a fifth column by hand: verified <url>, softened, or cut. The pass ends when no row is blank.

Pass 2 — what structure exists only because a model produces structure? (25%)

Wikipedia's cleanup editors maintain a field guide to the recurring shapes, built from real flagged articles and explicitly descriptive rather than prescriptive (Wikipedia:Signs of AI writing): sections that restate the opening, "not just X, but Y" constructions implying a contrast nobody earned, vertical lists in which every bullet opens with a bolded inline header, and sentences saying something "plays a crucial role" without saying what it does. Treat the vocabulary as a production artefact rather than a style worth preserving — a post-ChatGPT jump in "excess vocabulary" across 15 million-plus PubMed abstracts isolated 379 excess style words for 2024, two thirds of them verbs (Kobak et al., Science Advances, 2025; biomedical abstracts, so the rates do not transfer), and an independent group found the same overrepresentation (Juzek & Ward, COLING 2025).

You are a structural editor. Do not rewrite. Do not improve style.
Find only the following patterns and report where they occur.

TEXT: [paste draft]

Report as a table:
| Pattern | Verbatim excerpt | Paragraph | Cut, compress, or keep? |

Patterns:
1. A closing paragraph or section that restates earlier content
2. Three-item lists where the third item adds nothing
3. "Not just X but Y" / "It's not only... it's..." constructions
4. Bullets with identical shape and near-identical length
5. Sentences asserting importance without stating a consequence
6. Meta-sentences addressed to the reader about the article itself
7. Terms bolded on every mention

Report nothing outside these seven patterns. If a pattern does not occur, write "none".

The pass ends when the word count has gone down. If it went up, you were rewriting.

Pass 3 — what in this draft could only have come from you? (25%)

Run the swap test on every remaining paragraph: replace your brand name with your closest competitor's. If nothing becomes false, the paragraph carries no information about you — cut it or make it specific. Then insert your own material. In the 16th annual B2B benchmark from CMI and MarketingProfs (1,015 respondents, fielded June–August 2025, sponsored by Storyblok), among those using AI for content creation 87% said productivity had improved while only 39% said content performance had (CMI & MarketingProfs). Faster production of interchangeable text is what that gap looks like from the inside.

For each paragraph below, answer one question only:
"If the brand name were replaced with a direct competitor's name, would any
sentence in this paragraph become false?"

TEXT: [paste draft, paragraphs numbered]

Output: | Paragraph | Anything become false? (yes/no) | Which sentence | If no, what specific information is missing |

Do not rewrite anything. Do not suggest wording.

Pass 4 — which sentences survive, and when does the timer stop you? (10%, hard cap)

Hedges that hide the absence of evidence — "can help", "may contribute to" — are allowed only when you can name what makes you uncertain; otherwise state the claim plainly with the Pass 1 source, or delete the sentence with it. Do not just strike the qualifier: "AI may improve conversion" becoming "AI improves conversion" manufactures a new unsourced claim. Then search out the marker vocabulary that survived Pass 2. When the timer ends, the pass ends; consistency across drafts belongs to a scoring routine, not a longer polish (score-ai-output).

What does a good result look like?

The artefacts, not the prose. A finished edit leaves a claim log with no blank rows and one line in your edit log.

# Assertion (verbatim) Source in text Outcome
1 "Email marketing returns $36 for every $1 spent" no cut — no primary source reachable
2 "Our onboarding sequence recovered 11% of trial drop-offs in Q2" no verified — own dashboard export, 2026-06-30
3 "Short subject lines drive higher open rates" "studies show" softened — stated as our observation
2026-07-27 | q3-trial-onboarding-post | budget 25m | actual 31m | not killed | word count rose in pass 2, re-cut

Ten of those lines tell you whether AI is saving you time. Nothing else will.

How do you know the output is good?

Each check points at an artefact rather than at your judgement, so a colleague can run them without having watched you work. Two require you to keep something: the claim log from Pass 1, and the word count of the draft as it arrived.

The one people skip is the stop rule. "Still editing" is not a state this task can end in: either the checks pass and you publish while you can still see sentences you would phrase differently, or the budget expires and you write down a decision — kill, re-brief, or escalate. A failing check with a large fix is a Pass 0 verdict arriving late, not a reason to keep going.

Two things deliberately stay out: how the text scores on an AI detector, and how much changed. Google's guidance on generative-AI content tells you to "focus on accuracy, quality, and relevance" (Google Search Central), and its spam policy treats mass-produced pages of little value as abuse "no matter how it's created" (Google Search Central — spam policies) — stated policy, not a measured ranking effect, and the only claim about search this page makes.

Where does this usually break?

You start with tone. The budget disappears into sentences Pass 2 would have deleted. The symptom is a word count that rose during editing.

You restyle the model's sentences instead of replacing them. This is what leaves an "extensively edited" draft still reading like everyone else's. In the keystroke-logged essay study the rise in sameness — statistically significant with the feedback-tuned model, not the base one — came from the model's own contributions, while "user-written text is unaffected" (Padmakumar & He); an ideation study found homogenization "stems from the LLM providing different users with similar ideas, rather than by increasing individual user-level fixation" (Anderson, Shah & Kreminski, C&C 2024 — 33 participants, convergent rather than conclusive). Paraphrase preserves the problem; substitution removes it. That is Pass 3, not Pass 4.

You keep editing a draft that failed Pass 0 — sunk-cost repair at Pass 4 prices.

You delete hedges rather than resolving them, converting a vague sentence into a false one.

You edit sentences while shipping a governance problem. In the Optimizely/Savanta survey, 25% of marketing leaders said they frequently or always publish AI content they know is not fully on-brand under deadline pressure, and 30% that they pass AI work off as their own, human-created original work (Optimizely). Self-reports about behaviour people prefer not to admit, so read them as a floor — and neither is fixable at the sentence level.

What else do people ask?

Does editing remove the sameness that makes AI content read like everyone else's?

No study behind this page measures it. The diversity research compares AI-assisted writing with unassisted writing; as of 2026-07-27 none of it compares a text before and after a human editing pass, in either direction — so every "just edit it properly" claim, including the softer one on this page, rests on inference.

The wider literature also disagrees with itself. Studies finding that AI narrows collective variety include the essay study above and a 293-writer experiment where AI-assisted stories were rated more creative individually while converging on each other (Doshi & Hauser, Science Advances, 2024). A larger, more ecologically valid experiment — 800-plus participants across 40-plus countries, each cohort's ideas seeding the next — found the opposite sign, reporting that high AI exposure "did not affect the creativity of individual ideas but did increase the average amount and rate of change of collective idea diversity" (Ashkinaze et al., ACM Collective Intelligence, 2025). What they agree on is what this page relies on: the interchangeable text is the model's own contribution.

Should I edit to make it not read as AI-written?

Not as a goal in itself. Across six experiments with 4,600 participants, people could not detect AI-generated self-presentations, judging instead on flawed cues — first-person pronouns and contractions read as "human" (Jakesch, Hancock & Naaman, PNAS, 2023). That covers short texts from 2022-era models, so it does not license "readers can never tell"; it does mean surface AI-ness is not what your reader is grading.

How long should this take?

Whatever you wrote down, until your own log says otherwise. The nearest hard numbers come from machine translation, where 12 translators on nearly 145,000 source words improved throughput by 74%, saving 43% of translation time (Plitt & Masselot, 2010) — a setting with a reference text and a narrow space of correct answers, which marketing drafting has neither of. The rest of the track: ai-for-marketers.

Sources

Verification

4 log entries
dateactionresult
2026-07-27researchapplied
2026-07-27draftapplied
2026-07-27correctionapplied
2026-07-27fact-checkpass-3-0

Backlinks