How do you build a content brief with AI — and check it's not made up?
A no-tool method for building a content brief with a general-purpose chatbot: you capture the results page yourself, keep every metric out of the model's hands, and validate the finished brief against the live SERP before a writer touches it.
You can assemble a content brief with a general-purpose chatbot and show, line by line, where each claim came from and how you would know if it were wrong.
Last verified 2026-07-27
A content brief is the instruction set a writer gets before drafting: the query the page targets, the intent behind that query, the outline, the original material the page must carry, and the constraints. A general-purpose chatbot can assemble one in minutes, and it fails in one specific way — asked for search volume it has no database for, or for "what's currently ranking", it will produce a confident answer built from memory rather than observation. The method below treats the model as an analyst working strictly from evidence you captured, keeps every metric out of its hands, and ends with a validation pass run against the live results page. It needs no paid SEO tool and no code.
When should you use this — and when should you not?
Use it when you are briefing one page, you can open the live results page for the country and language you actually target, and the decision the brief supports is editorial: what to cover, in what order, with what first-hand material.
Speed is not the scarce resource here. In the 2026 B2B survey from Content Marketing Institute and MarketingProfs (n = 1,015, fielded 24 June – 14 August 2025; sponsored research), 95% said their organisation uses AI-powered applications, and among those using AI for content creation 87% reported better productivity but only 39% reported better content performance (CMI/MarketingProfs). The check, not the generation, is what most workflows are missing.
Do not use this workflow when:
- You cannot look at the live results page yourself. The method rests on a human-captured snapshot; without one, nothing in the brief is grounded.
- You need volume or difficulty figures precise enough to justify budget. Free official sources give rounded averages or a relative index, not decision-grade absolutes. That decision needs a data product or your own Search Console history, not a prompt.
- You are producing briefs in bulk to fill a site. Google's spam policies define scaled content abuse as generating many pages "for the primary purpose of manipulating search rankings and not helping users", listing first among examples "using generative AI tools or other similar tools to generate many pages without adding value for users" (Google Search Central). A faster brief factory is a faster route there.
- The topic is health, finance, safety or civics. For informational and YMYL pages, Google's rater guidelines make "accuracy and consistency with well established expert consensus" a quality criterion (General Guidelines, 11 September 2025). A brief reverse-engineered from a results page is not a qualified reviewer.
- The query is a brand or navigational one — a "Website query" in Google's own taxonomy. There is no content gap to fill.
What do you need before you start?
| Input | Where it comes from | Why it is required |
|---|---|---|
| One target query, frozen in writing | You | Every later check is "against this query" |
| A results-page snapshot: top 10 titles, URLs and page formats; the People Also Ask questions; whether an AI Overview appeared; capture date, country, language | You, by hand, in a logged-out/incognito window | The only ground truth in the workflow |
| Search Console performance data, if the site exists | Search Console → Performance report | Real impressions, clicks and average position for queries you already appear for |
| Volume order-of-magnitude (optional) | Google Keyword Planner, free with an Ads account | Sanity check only — the figures are rounded and averaged |
| Seasonality (optional) | Google Trends | Direction of travel, never a volume |
| One sentence on the reader and the job the page does for them | You | Stops the model optimising for "comprehensiveness" |
Deliberate non-input: any number the chatbot itself produces.
How do you do it, step by step?
Step 1 — Freeze the query and capture the results page yourself. Open a private window, set the country and language you target, run the query, and record the top ten organic results (title, URL, and what kind of page each is), the People Also Ask questions, and whether an AI Overview appeared. Put the capture date and location at the top. This is step one and not step three because the model cannot do it: frontier models ship with a documented, per-model knowledge cutoff (OpenAI developer docs), and with no retrieval tool switched on, anything it "sees" on your results page is reconstruction.
QUERY: [target query]
CAPTURED: [YYYY-MM-DD] | COUNTRY: [xx] | LANGUAGE: [xx] | MODE: incognito
AI OVERVIEW PRESENT: [yes/no]
PEOPLE ALSO ASK: [q1] / [q2] / [q3] / [q4]
1. [title] | [url] | format: [guide|listicle|tool page|comparison|forum|video|docs]
2. ...
10. ...
Step 2 — Have the model classify intent in Google's own vocabulary. Google's rater guidelines classify queries as Know (with Know Simple as a subtype), Do, Website and Visit-in-person, and have an explicit category for queries with multiple user intents (General Guidelines §12.7). Borrowing that taxonomy costs nothing and, unlike invented vocabularies, is stable and citable.
You are analysing a search results page I captured manually. Use ONLY the data
between <snapshot> tags. You have no other knowledge of this results page.
<snapshot>
[paste your snapshot]
</snapshot>
Classify the dominant user intent using exactly one of these four categories
from Google's Search Quality Rater Guidelines: Know (Know Simple if a short
fact answers it), Do, Website, Visit-in-person.
Answer in this format and nothing else:
DOMINANT INTENT: [one category]
EVIDENCE: [snapshot result numbers that support it]
COMPETING INTERPRETATION: [second-most-likely intent, or "none"]
DOMINANT FORMAT: [page format shared by the largest number of top-5 results]
FORMAT COUNT: [n of 5]
UNKNOWN: [anything you cannot determine from the snapshot]
Do not state search volume, keyword difficulty, CPC or traffic estimates.
If asked for a number you cannot derive from the snapshot, write UNKNOWN.
If EVIDENCE comes back empty, or FORMAT COUNT is 1 of 5, the results page is mixed. Record that in the brief instead of forcing one intent.
Step 3 — Generate a grounded, metric-free outline. Every heading must be traceable to something you captured.
Using ONLY the snapshot and the intent verdict above, draft a content outline.
Rules:
1. Tag every H2/H3 with its origin:
[SERP n,n] the topic appears in those snapshot results
[PAA] it comes from a People Also Ask question
[GAP] it appears in none of them and you are proposing it
2. No word counts, search volume, keyword difficulty or traffic numbers.
If you think one is needed, write: NEEDS EXTERNAL DATA: [what].
3. Do not invent statistics, studies or sources inside the outline.
4. Maximum 8 H2s.
Output format:
## [H2 heading] [origin tag]
- [H3] [origin tag]
INTENT SERVED: [one sentence]
End with:
GAP HEADINGS: [count]
NEEDS EXTERNAL DATA: [bulleted list, or "none"]
The tags make one thing visible at a glance: how much of the brief is consensus you copied and how much is genuinely yours.
Step 4 — Add what the results page does not have. Decide the original material the page will carry — your own test, your own data, a customer quote, a worked example, a template — and write it into the brief as a required asset with a named owner. Google's raters are told that for most pages, main-content quality "can be determined by the amount of effort, originality, and talent or skill that went into the creation of the content", and that main content which is "copied, paraphrased, embedded, auto or AI generated, or reposted from other sources with little to no effort, little to no originality, and little to no added value" earns the Lowest rating (General Guidelines §3.2 and §4.6.6). An outline reverse-engineered from the results page is, by construction, the consensus. Use the model as a challenger here, not an author:
<outline>[paste]</outline>
<snapshot>[paste]</snapshot>
Act as a critic. Answer only these three questions:
1. WHAT IS MISSING: which questions would a [audience] still have after reading
every result in the snapshot? Up to 5.
2. WHAT IS OVER-COVERED: which outline sections are already served by 3+
snapshot results?
3. ORIGINAL ASSET CANDIDATES: propose 3 kinds of first-hand material that would
make this page non-substitutable. For each, state what the author must do.
Do not rewrite the outline. Do not add statistics.
Step 5 — Fill every number from a named, dated source. For each NEEDS EXTERNAL DATA item, fetch the figure yourself and record source plus retrieval date beside it. Three free surfaces, read as they are actually defined:
- Search Console — an impression means a user "has seen (or potentially seen) a link to your site", and the reported position is "the topmost position occupied by a link to your property or page in search results, averaged across all queries in which your property appeared" (Search Console Help). Impressions are not searches.
- Keyword Planner — Google states plainly that "your search volume statistics are rounded", and that the figure is averaged over a 12-month period for exact matches (Google Ads Help). Order of magnitude, not precision.
- Google Trends — each point is "divided by the total searches of the geography and time range it represents" and scaled 0–100, and Google warns that "different regions that show the same search interest for a term don't always have the same total search volumes" (Trends Help). Direction, never volume.
Then run the model as a guard over your own brief:
Below is my brief. Output a table of EVERY number, percentage, date or named
statistic it contains.
<brief>[paste]</brief>
Columns: VALUE | WHERE IT APPEARS | SOURCE CITED IN BRIEF | RETRIEVAL DATE CITED
Mark any row missing a source or a retrieval date as UNSOURCED.
Last line: the count of UNSOURCED rows.
Do not supply the missing sources. Do not guess. Do not add numbers.
Step 6 — Re-open the results page and check the brief against it. For each [SERP n] tag, confirm the result is still there and still that format, and confirm the AI Overview state still matches. If more than two of the top five have changed, redo step 1. Add one line to the brief: SERP re-verified [YYYY-MM-DD]; [n] of top 10 unchanged.
What does a good result look like?
A finished brief is short and every line answers "where did this come from". The shape below is what hand-off looks like:
TARGET QUERY: how to run a customer interview
SNAPSHOT: captured 2026-07-20 | US | en | incognito
SERP RE-VERIFIED: 2026-07-26 — 9 of top 10 unchanged
INTENT: Do (evidence: results 1,2,4,7) | competing: Know
FORMAT: step-by-step guide (4 of 5)
OUTLINE
## What is a customer interview? [SERP 1,3]
## How long should one run? [PAA]
## What do you ask when the user won't talk? [GAP] — absent from all top 10
INTENT SERVED: gives the reader a runnable script.
ORIGINAL ASSET: recording of our own 20-minute interview + annotated transcript.
OWNER: [name]. Appears in 0 of top 10.
NUMBERS
- 12-month avg volume ~2,400 (rounded) | Keyword Planner | retrieved 2026-07-20
UNSOURCED: 0
GAP HEADINGS: 1
Note what is absent: no keyword difficulty score, no CPC, no "target word count", no claim about the results page that is not in the snapshot.
How do you know the output is good?
- Every number, percentage and named statistic in the brief carries a source name and a retrieval date; a metric audit of the finished brief returns zero unsourced rows.
- Every H2 and H3 in the outline carries an origin tag ([SERP n], [PAA] or [GAP]); the count of untagged headings is zero.
- Every [SERP n] tag was re-checked against the live results page within 14 days of hand-off, and the brief records that date plus how many of the top ten were unchanged.
- The brief names one dominant intent from Google's four rater categories, or explicitly declares the query multi-intent, and the declared page format matches at least three of the top five captured results.
- The brief names at least one first-hand asset that appears in none of the captured top-ten results, with a named owner.
Treat the criteria above as a review someone else could run without asking you a single question. That is the test of a good brief — not whether it reads well.
Run the metric audit last, on the finished document, and read each UNSOURCED row as one figure to fetch or delete. A first pass that returns zero usually means the prompt was run on the outline instead of the whole brief.
The heading audit is pure counting: headings without an origin tag are headings nobody can trace, and they are usually the ones the model invented. [GAP] headings are not failures — they are the valuable part — but each needs one sentence saying why the gap is real.
The freshness check is the one people skip, and the only criterion that touches the live web. Re-open the results page yourself; do not ask the model whether it changed. Put the date and the unchanged count in the brief so the writer inherits evidence rather than your reassurance.
Intent and original asset are best judged by a second reader: hand over the brief and the snapshot, and ask a colleague to name the dominant intent independently. A different answer means the results page is mixed and the brief should say so. For the asset the question is blunt — could any of the top ten results have produced this material? For a repeatable rubric, see scoring AI output consistently.
Where does this usually break?
Fabricated metrics. Asked for search volume, difficulty or CPC, a general chatbot will supply a number. It has no keyword database to draw one from — Ahrefs, which sells that database and so has an interest in the claim, states that "any volume, KD, or SERP figures it generates unprompted are fabricated" (Ahrefs); SpyFu, a competing vendor, says the same (SpyFu). The mechanism is documented independently of both: models "guess when uncertain, producing plausible yet incorrect statements instead of admitting uncertainty", because the way they are graded rewards guessing over abstention (Kalai et al., arXiv 2509.04664). Note what none of these sources supplies: a measured error rate for chatbot-stated volumes. This is not a "usually wrong" percentage — it is a capability the tool does not have, so treat every such figure as unusable however right it looks. Deeper detection: catching fabricated statistics and sources.
The imagined results page. Ask "what's ranking for X?" with web search off and you get a plausible page assembled from training data. Detection is mechanical: ask for exact URLs and open them; dead links, wrong publishers or pages that plainly do not rank mean fabrication. The variant that catches experienced people is silent tool state — retrieval is off, or the search returned nothing and the model answered anyway. Require the URLs it actually retrieved; no URLs, no retrieval.
Word count as a target. Brief templates routinely prescribe a length derived from the pages currently ranking. Google's stated quality criteria are effort, originality, talent or skill, and accuracy (General Guidelines §3.2); length is not among them.
Consensus cloning. A brief that is 100% [SERP]-tagged with no original asset describes a page whose content is already on the web — the Lowest-rating territory quoted in step 4, and, repeated at scale, the scaled-content-abuse policy. The counter is countable: GAP HEADINGS: 0 plus no owned asset means start over.
A snapshot nobody can reproduce. Location, language, device and personalisation change what appears; Search Console's documentation says the position you see in your own search "might be different than the average because of many variables, such as your search history, location, and so on" (Search Console Help). Without country, language and incognito state recorded, the writer cannot re-run your capture.
Stale snapshots. Composition moves faster than briefs get used. Semrush — a vendor selling into this market — tracked AI Overview trigger rate across 10M+ keywords at 6.49% of queries in January 2025, peaking at 24.61% in July and falling to 15.69% in November (Semrush, refreshed 15 December 2025). Whatever the exact figures, a brief validated months ago may describe a different page.
Trusting any single volume estimate. Even paid, industry-standard figures disagree with each other and with Google's own data: across 72,635 keywords, Ahrefs found Keyword Planner overestimated versus Search Console impressions in 91.45% of cases, with 54.28% off by more than half (Ahrefs, 3 November 2021). The yardstick there is Search Console impressions, which that study calls the "single source of truth" for keyword data — a proxy, not a count of searches. It also predates AI Overviews: read it for the direction, not for a current accuracy rate.
What else do people ask?
Can't I just switch web search on and let the model read the results page?
You can, and it helps — but it changes what you verify rather than removing the verification. A retrieval-enabled model returns a summary of pages it fetched, in a session whose location and personalisation you did not control, and it will not announce that retrieval silently failed. Ask for the URLs, open them, and record the capture conditions yourself.
Do I need a paid SEO tool for this?
Not for the brief. Paid tools genuinely have live results data and entity extraction at scale, and this method does not replicate that. What it gives you instead is a brief where every line is traceable and a check a colleague can re-run. If a budget decision hangs on precise volume, that is a tool purchase, not a prompt.
Should the brief tell the writer how to get cited in AI answers?
Not as a promise. Google's own documentation on AI features states that "there are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary" (page last updated 10 December 2025) (Google Search Central), and its helpful-content guidance is about originality, first-hand expertise and clarity rather than a format (Google Search Central). Treat any brief instruction that promises citations in AI answers as untested: write the brief for the reader, then measure what happens on your own pages.
Related workflows: turning the brief into a reliable draft workflow, synthesising customer research without inventing consensus, and checking marketing content before it publishes. Full track: ai-for-marketers.
Want the next practical guide?
Sources
- Google — General Guidelines (Search Quality Rater Guidelines), September 11, 2025accessed 2026-07-27
- Google Search Central — Spam policies for Google web searchaccessed 2026-07-27
- Google Search Central — Creating helpful, reliable, people-first contentaccessed 2026-07-27
- Google Search Central — AI features and your website (last updated 10 December 2025)accessed 2026-07-27
- Google Ads Help — About Keyword Planner dataaccessed 2026-07-27
- Google Trends Help — FAQ about Google Trends dataaccessed 2026-07-27
- Google Search Console Help — Performance report definitionsaccessed 2026-07-27
- Why Language Models Hallucinate — Kalai, Nachum, Vempala & Zhang, arXiv 2509.04664accessed 2026-07-27
- OpenAI developer docs — Models (per-model knowledge cutoff)accessed 2026-07-27
- Ahrefs — AI Keyword Researchaccessed 2026-07-27
- SpyFu — Don't Risk Your Work on Made-Up Dataaccessed 2026-07-27
- Ahrefs — AI Overviews Reduce Clicks by 34.5% (17 April 2025)accessed 2026-07-27
- Ahrefs — Update: AI Overviews Reduce Clicks by 58% (4 February 2026)accessed 2026-07-27
- Pew Research Center — Google users are less likely to click on links when an AI summary appearsaccessed 2026-07-27
- Semrush — AI Overviews Study (refreshed 15 December 2025)accessed 2026-07-27
- Ahrefs — GSC vs. GKP: Comparing Search Volumes for 72k Keywordsaccessed 2026-07-27
- Content Marketing Institute & MarketingProfs — B2B Content and Marketing Trends: Insights for 2026accessed 2026-07-27
Verification
4 log entries
| date | action | result |
|---|---|---|
| 2026-07-27 | research | applied |
| 2026-07-27 | draft | applied |
| 2026-07-27 | correction | applied |
| 2026-07-27 | fact-check | pass-3-0 |