How do you pick which AI-generated ads are worth testing?
A method for cutting AI-generated ad copy down to a set worth running: group candidates by the promise they make, tie every claim to a fact you can substantiate, and size the test your traffic can actually support.
You can cut a pile of AI-generated ad copy down to a short list worth running, and state before launch what your traffic could and could not prove.
Last verified 2026-07-27
Generating ad copy is the cheap half of the job: any assistant returns fifty headlines in a minute, and most of them are one promise in fifty costumes. The expensive half is selection — grouping candidates by the promise they actually make, tying each to a fact you can substantiate, and working out what your click volume could prove before you launch anything. The method works with any assistant that takes a written brief and returns a table, and it belongs to the ai-for-marketers track.
When should you use this — and when should you not?
Use it when you are rewriting copy in a live ad group, you know that group's current click-through and conversion rate, and you can name the objection you think the copy fails to answer. Four conditions should stop you first, and the first one stops most small accounts.
Your traffic cannot power a test. Look up the clicks each version would get:
| Clicks per version | Smallest lift you could detect (absolute) | The same lift, stated relative |
|---|---|---|
| 200 | 4.8 pp | +161% |
| 500 | 3.1 pp | +102% |
| 1,000 | 2.2 pp | +72% |
| 5,000 | 1.0 pp | +32% |
| 10,000 | 0.7 pp | +23% |
| 50,000 | 0.3 pp | +10% |
Baseline conversion rate 3%, 95% confidence, 80% power, two versions of equal size. Computed for this page from the sample-size rule n = 16σ²/Δ² published in Kohavi et al., Controlled experiments on the web (2009) — read your row rather than redoing the algebra.
At 200 clicks per version, only a change that more than doubles your conversion rate is detectable. Ad copy does not do that, so at that volume the dashboard's winner is noise. One rule to carry out of the arithmetic: the effect is squared, so halving the effect you want to detect quadruples the traffic you need — and a much bigger contrast needs far less traffic, which is the escape route in step 5. For scale, across thousands of experiments a year at Bing, "most fail, and those that succeed improve key metrics by 0.1% to 1.0%, once diluted to overall impact" (Kohavi et al., Seven Rules of Thumb, KDD 2014) — web experimentation rather than paid search, so an analogy, but one pointing the same way as the table.
The platform will not hold your copy still. Google's "Optimize" rotation setting "prioritizes ads that are expected to perform better than other ads within an ad group", and Google Ads "will automatically use the 'Optimize' ad rotation setting" whenever you run Smart Bidding; even the "Do not optimize" alternative warns that "the percentage of impressions for each ad may not be even due to the quality of the ad" (Use ad rotation). Meta's dynamic creative "helps you find the best creative combination per impression" (Dynamic Creative). Either way, "ad A beat ad B" is confounded by the platform's own allocation. Use the experiment surface in step 5, or make no causal claim.
The message is claims-heavy or regulated. Health, finance and guaranteed-outcome copy sits inside policies that reject specific phrasings; clear the claim before generating fifteen versions of it. ai-content-pre-publish-compliance comes first.
You have no baseline. Without the ad group's current click-through and conversion rate there is nothing to compare against and no way to size anything.
What do you need before you start?
Put these on one page first, because each becomes a [variable] in the prompts below. The load-bearing one is the substantiation list — facts you can prove: prices, terms, ratings, credentials, counts. It turns "does this claim check out?" from a judgement call into a lookup. Your brand-voice-doc-for-ai supplies the other half: words legal or brand will not allow.
The input people skip is the hypothesis set — three to five competing beliefs about why the buyer does not click or convert, each falsifiable by a headline: price, speed, risk, specificity, authority. If you have interview or review data, customer-research-synthesis turns it into hypotheses without inventing a consensus. Without this list you get variety of wording instead of variety of meaning, and nothing transfers to the next campaign.
Load the format limits into the brief too, so the assistant writes to the slot the copy will occupy.
| Platform | Asset | Limit |
|---|---|---|
| Google responsive search ad | Headlines | up to 15, 30 characters each |
| Google responsive search ad | Descriptions | up to 4, 90 characters each |
| Google responsive search ad | Display path fields | 15 characters each |
| Meta asset feed | Titles / descriptions | up to 5 each, 255 characters max |
| Meta asset feed | Body text | up to 5, 1,024 characters max |
| Meta asset feed | Total assets | 30 maximum; images and videos 10 each |
As of 2026-07-27 — Google figures from About responsive search ads, Meta from the Asset Feed Spec options. Two caveats: Google's help pages publish no last-updated date, so that stamp is this page's access date, and the Meta numbers are the Marketing API specification — Ads Manager may expose fewer slots.
How do you do it, step by step?
1. Write the hypotheses before you open the assistant. Three to five, numbered, in your own words. This step decides whether anything downstream is measurable.
2. Generate per hypothesis, not per request.
You are helping me write search ad copy. Do not invent facts.
PRODUCT: [product/service]
AUDIENCE: [audience and the job they are trying to do]
OFFER: [what is on the other side of the click]
FACTS I CAN SUBSTANTIATE (the only permitted basis for any claim): [proof list]
FORBIDDEN WORDS AND CLAIMS: [constraints]
Competing hypotheses about the buyer's main objection:
[H1 ...] [H2 ...] [H3 ...]
Write [N] headlines per hypothesis, each MAX [30] CHARACTERS INCLUDING SPACES.
Rules:
- Every headline must stand alone; it may be shown without the others.
- Headlines under the same hypothesis make the SAME promise in different words.
- Headlines under DIFFERENT hypotheses make DIFFERENT promises.
- No superlatives and no guarantees unless they appear in FACTS above.
Output a markdown table with columns:
| hypothesis | headline | character_count | which fact it relies on |
3. Find out how many candidates you really have. Paste that table into a fresh chat.
Here are candidate ad headlines with the hypothesis each is meant to express:
[paste table]
For each PAIR, judge whether they make the SAME promise to the buyer or a
DIFFERENT one. Ignore wording; judge the promise.
Output:
1. Clusters — headlines making the same promise, grouped.
2. For each cluster, one sentence naming the promise.
3. A flag on any cluster containing headlines I labelled with DIFFERENT
hypotheses, since that means my labelling is wrong.
If fifteen headlines collapse into two clusters, you have two candidates rather than fifteen, and the comparison you planned cannot answer your question. Google's guidance pushes the same way: "Create unique headlines and descriptions" and "Avoid repeating the same content across different assets" (About Ad Strength).
4. Run a policy pass before anything goes live. Ask for an ALLOW / RISK / BLOCK verdict per headline against the rules that bite ad copy most often: nothing that "entice[s] the user with an improbable result (even if this result is possible) as the likely outcome", and guarantees only with "a clear and easily accessible refund (money-back) policy" (Misrepresentation); no claim absent from your substantiation list; no "[n]on-standard, gimmicky, or unnecessary repetition of names, words, or phrases" (Editorial); no implied "affiliation with or endorsement by another organization, brand, or private citizen" (Misleading representation). On Meta, ads must comply with the standards on Fraud, Scams and Deceptive Practices, which targets "exaggerated claims", and on Misinformation. The verdict is a first filter, never a clearance: regulated categories go to a human.
5. Decide which regime you are in, and write it down in one sentence: "I am in regime X, running until [date or volume], and I will call a winner only if [criterion]."
Regime A — enough traffic (detectable lift around 10–20% relative or better). Run a real experiment, which on Google means ad variations rather than two ads in one ad group: ad variations randomise at the user level, since "[y]our ad variation and original ads will be split by assigning cookies to users… users may discover only one version of your ad, regardless of how many times they search" (Set up an ad variation). Rotation documents no such split. The report shows a difference against the original with an interval, and a blue asterisk marks a metric "at least 95% likely" to reflect your change "rather than… random chance" — but check the setting first: "You can pick your own confidence intervals (80% is the default confidence interval)" (Monitor your ad variations). An 80% interval is a far weaker bar than the 95% asterisk, and both live on one screen.
Regime B — moderate traffic. Do not test individual headlines; test hypotheses, as two complete asset sets expressing different promises. Bigger contrasts need far less traffic.
Regime C — low traffic, the common case. Do not claim measurement. Ship a diverse, policy-clean portfolio, let the platform allocate, judge qualitatively — and say out loud that this is not a test.
6. Load it without breaking the plan. On Google, leave rotation on "Optimize" — "'Do not optimize' isn't recommended for most advertisers" (Use ad rotation) — and avoid pinning unless wording is legally required: "pinning isn't recommended for most advertisers and can affect ad strength", and "a minimum of one headline and one description will be selected to show in different combinations and orders" (About responsive search ads). You are not shipping fifteen ads; you are shipping one system that assembles ads. That system now reaches past the single ad: under "Enhanced Flexibility", headline text "may show at the beginning of the description", and "[u]nused responsive search ad headlines and description assets from another ad within the same ad group can serve as link-based assets" (About responsive search ads) — one more reason an ad-level comparison is not clean. On Meta, stay inside the feed limits and expect combination-level delivery — the Marketing API documents assembly of combinations from the assets you supply, and nothing beyond that, so assume nothing beyond it.
7. Read the result at the stopping point, not before.
Results from an ad copy test.
CONTROL: [copy] — [impressions], [clicks], [conversions]
NEW VERSION: [copy] — [impressions], [clicks], [conversions]
Pre-registered stopping rule: [what I wrote in step 5]
Smallest lift this test could detect: [X% relative]
Tasks:
1. State the observed relative difference in click-through and conversion rate.
2. State whether it exceeds the smallest detectable lift. If it does not, say
explicitly that this test could not have detected it.
3. Name every explanation other than the copy: unequal exposure, seasonality,
auction changes, small numbers.
4. Do NOT declare a winner. Output evidence and caveats only.
8. Bank the learning at the hypothesis level. Record which promise won, not which sentence: one line per test with hypothesis, regime, outcome, date, and the lift the test could have detected.
What does a good result look like?
An illustrative portfolio for a small invoicing tool whose substantiation list is a fourteen-day trial, no card required, ten-minute setup and 2,400 customers. The figures are invented, to show the shape of the artefact.
| Hypothesis | Headline | Characters | Fact relied on |
|---|---|---|---|
| H1 risk | No card. 14-day free trial | 26 | trial terms |
| H1 risk | Invoicing that starts free | 26 | trial terms |
| H2 speed | Set up invoicing in 10 mins | 27 | setup time |
| H2 speed | Send your first invoice today | 29 | setup time |
| H3 authority | Used by 2,400 businesses | 24 | customer count |
| H3 authority | 2,400 businesses invoice here | 29 | customer count |
Three distinct promises, two candidates each, every row attributed, every count verified by hand. Beside it sits one dated sentence written before launch — "Regime C: 340 clicks a month, smallest detectable lift about +120%, so no winner will be declared; running four weeks, then judged on policy-clean coverage and fit" — and, after the run, a one-word verdict: supported, not supported, or underpowered.
How do you know the output is good?
- Every headline and description is inside the platform's character limit, verified by counting the characters yourself rather than by asking the assistant
- Every candidate is labelled with exactly one hypothesis, and clustering by promise does not collapse two hypotheses into one
- Every claim in every candidate points to a fact on your substantiation list, and no candidate cites a fact you did not supply
- Before launch you wrote down the smallest lift your click volume can detect, the stopping point, and whether a test is possible at all
- The final verdict is one of supported, not supported, or underpowered — and underpowered is chosen whenever the observed difference is smaller than the lift you could detect
- No conclusion rests on Ad Strength, on an asset-level ratio metric alone, or on a reading taken before the stopping point
Three checks are mechanical. Count characters yourself instead of trusting the count in the table — a self-reported count is not evidence of anything, and the limit is hard: Google's headline fields "support up to 30 characters" (About responsive search ads). Confirm attribution is complete: every row names a fact, and every named fact is on the list you supplied. Confirm the clustering returns as many distinct promises as you claimed hypotheses; fewer means the pile is cosmetic, and no amount of traffic will rescue it. Keeping these checks consistent from run to run is what score-ai-output is for.
The fourth check is arithmetic and must exist in writing, dated, before launch: the smallest lift this test could detect, and the stopping point. Written afterwards, they are a rationalisation.
What does not count as evidence matters as much, and here the platform contradicts the way its own reports get read. Ad Strength is not a result: "The Ad strength rating of an ad doesn't directly influence your ad's serving eligibility" and "[i]t isn't used to calculate Ad Rank, Quality Score, or auction wins" (About Ad Strength) — it is feedback on asset diversity. Asset-level ratios are not verdicts: "Ratio metrics at the asset level, such as CTR, CPC, CPA, ROAS, should be used as directional indicators only. These metrics don't accurately reflect the overall performance of a single asset in isolation" (the ad-level asset report). That sentence sits in Google's own help centre, on the page documenting the report it is warning you about.
The honest vocabulary is three words wide, not two: "underpowered — no conclusion" is a legitimate and, at most account sizes, expected outcome. A method that always produces a winner is producing them out of noise.
Where does this usually break?
Peeking. Checking daily and stopping the moment it looks significant manufactures results. P-values and confidence intervals "are wholly unreliable if users endogenously choose samples sizes by continuously monitoring their tests", and "[e]ven with 10,000 samples (a sample size which is quite common in online A/B testing), Type I error can easily increase fivefold" (Johari, Pekelis and Walsh, Always Valid Inference). A nominal 5% false-positive rate behaving like roughly 25% means the significance marker you are reading is not the assurance it looks like. Note the tie when reading that paper: two authors were employed by, and one advised, the experimentation vendor whose product implements their proposed fix. The critique of peeking is standard and independent of that; the remedy available to you is simply to pre-commit the stopping point.
Optimising Ad Strength, and repeating the statistic that encourages it. Google reports that advertisers improving Ad Strength from "Poor" to "Excellent" see "15% more conversions on average", footnoted "Source: Google Internal Data. Date range: August 15, 2025 to August 20, 2025" (About Ad Strength). Read the footnote before the number: the publisher is the platform, the data is internal and unauditable, the window is five days, and no comparison group, sample size or method is stated anywhere on the page — advertisers who improve their Ad Strength are advertisers who chose to, which is not a randomised comparison. Google's own pages do not word it identically; the guide to responsive search ads renders it as "15% more clicks & conversions". Diversifying assets may still be sensible; this number does not show that it caused anything.
Quoting the write-up instead of the report. A widely shared summary of the State of PPC 2026 survey states that "AI saves PPC professionals an average of 5.2 hours per week" (State of PPC 2026: 7 Key Takeaways); the source report says "one to five hours per week" and publishes no such average at all. This is the discipline the page asks of AI output, turned on your own sources — the reflex ai-draft-pre-publish-check exists to build.
| Failure | What it looks like | What to do |
|---|---|---|
| Cosmetic variation | Step 3 collapses everything into one or two clusters | Regenerate per hypothesis; do not proceed with a pile |
| Testing without power | A winner declared on a few hundred clicks | Read your row in the table; accept regime C |
| Hallucinated specifics | A price, rating, award or count you never supplied | Reject the row; see spot-fabricated-stats |
| Repetition across assets | The same phrase reused across headlines in one ad group | Editorial policy prohibits gimmicky repetition; rewrite |
| Broken historical comparison | Asset performance history stops before mid-2025 | "Full performance statistics is only available for dates on or after June 5, 2025" (asset report) |
Believing someone else's winner will win for you. The KDD 2014 rule is titled "Your Mileage WILL Vary", and of the famous red-versus-green call-to-action result its authors write: "Since we do not see a lot of sites with red call-to-action buttons, we believe this is not a general result that replicates well." The same paper quotes Twyman's law: "Any figure that looks interesting or different is usually wrong!" (Seven Rules of Thumb). A case study claiming a 300% lift from a headline swap is a reason to check its methodology, not a template.
What else do people ask?
Does generating fifty headlines actually help?
Only if they differ in what they promise, and only up to what the format holds — fifteen headlines and four descriptions on a Google responsive search ad. Past that the surplus is unused; below that, the binding constraint is variety of meaning, not volume. Fifty near-identical headlines give the platform nothing meaningful to choose between and give you nothing to learn.
Most PPC teams already use AI for ad copy — doesn't that settle it?
Adoption is real and is not evidence of quality. The State of PPC 2026 survey reports adoption of large language models "ranging from 10% for budget management to 59% for ad copy", up "from 42% to 59%"; that question went to individual contributors and freelancers, n = 656, fielded 17 November to 24 December 2025 (State of PPC 2026). The 2026 charts do not state which response bucket the figure sums; matching them against the 2024 edition indicates "often or always", so read it that way rather than as "59% of marketers". The sample is self-selected, recruited by the survey's nine commercial partners and via social media, and skews European and Google-centric.
What the same report says next is more useful: quality and accuracy "emerge as the single biggest challenges PPC teams face when using AI and automation", and practitioners "are not struggling to adopt AI, they are struggling to operationalize it safely". Note also what is absent — the 2026 edition dropped the satisfaction question its 2024 predecessor carried, where writing ads scored 69% satisfied, second best of sixteen tasks, meaning roughly three in ten users of the second-best-rated application were not satisfied (State of PPC 2024). The flagship practitioner survey names verification as its top problem, then stops measuring it.
Want the next practical guide?
Sources
- Google Ads Help — About responsive search adsaccessed 2026-07-27
- Google Ads Help — About Ad Strength for responsive search adsaccessed 2026-07-27
- Google Ads Help — Your guide to responsive search adsaccessed 2026-07-27
- Google Ads Help — About the ad-level asset report for responsive search adsaccessed 2026-07-27
- Google Ads Help — Use ad rotationaccessed 2026-07-27
- Google Ads Help — About ad variationsaccessed 2026-07-27
- Google Ads Help — Set up an ad variationaccessed 2026-07-27
- Google Ads Help — Monitor your ad variationsaccessed 2026-07-27
- Google Ads Policy — Misrepresentation: unreliable claimsaccessed 2026-07-27
- Google Ads Policy — Editorialaccessed 2026-07-27
- Google Ads Policy — Misleading representationaccessed 2026-07-27
- Meta Marketing API — Asset Feed Spec optionsaccessed 2026-07-27
- Meta Marketing API — Dynamic Creativeaccessed 2026-07-27
- Meta Advertising Standards — Fraud, Scams and Deceptive Practicesaccessed 2026-07-27
- Meta Advertising Standards — Misinformationaccessed 2026-07-27
- Kohavi, Longbotham, Sommerfield, Henne — Controlled experiments on the web: survey and practical guide (Data Mining and Knowledge Discovery 18(1), 2009)accessed 2026-07-27
- Kohavi, Deng, Longbotham, Xu — Seven Rules of Thumb for Web Site Experimenters (KDD 2014)accessed 2026-07-27
- Johari, Pekelis, Walsh — Always Valid Inference: Continuous Monitoring of A/B Tests (arXiv 1512.04922)accessed 2026-07-27
- The State of PPC 2026 — Global Report (PPCsurvey.com, an initiative of TrueClicks)accessed 2026-07-27
- The State of PPC 2024 — Global Report (PPCsurvey.com)accessed 2026-07-27
- PPC Chief — State of PPC 2026: 7 Key Takeaways from 1,306 PPC Professionals (secondary summary, cited as an example of a mis-quoted statistic)accessed 2026-07-27
Verification
4 log entries
| date | action | result |
|---|---|---|
| 2026-07-27 | research | applied |
| 2026-07-27 | draft | applied |
| 2026-07-27 | correction | applied |
| 2026-07-27 | fact-check | pass-3-0 |