Agentic Wikiwiki / score-ai-output
← Wiki index
taskpractitioner

How do you know if AI output is any good?

Replace "this feels off" with an acceptance procedure: pass/fail checks drawn from your own failures, a frozen sample, measured agreement between two people, and an AI judge validated against human labels before you quote it.

Job outcome

You can run an acceptance procedure two reviewers apply the same way, and report an AI quality number next to the error rate of whatever produced it.

Last verified 2026-07-27

Scoring AI content means replacing "this feels off" with a written acceptance procedure: a short list of pass/fail checks drawn from your own past failures, a frozen sample of real output, and a measured answer to whether two people — and later an AI judge — apply those checks the same way. Call the list a checklist; the evaluation literature calls it a rubric, a written set of checks you grade against. The result is not a prettier score. It is a number you can defend, quoted next to the known error rate of whatever produced it.

When should you use this — and when should you not?

Use it when the same kind of AI content ships repeatedly, more than one person decides whether it ships, and quality arguments keep getting settled by whoever is more senior. The procedure pays for itself on a stream, not on a single piece.

Postpone or skip it here:

What do you need before you start?

What Where it comes from Why it is required
30–50 real outputs of one kind Your own archive: sent emails, published posts, delivered briefs Synthetic examples do not reproduce your failures
At least a third clear failures Deliberate selection — do not sample only successes With no failures in the sample, a missed failure cannot be measured
10–20 notes on what went wrong Free text, written while reading, no categories These become the checks; nothing else does
A second reviewer A colleague who understands the work "Consistent" is meaningless with one person
Brand or editorial rules, if any Existing docs — only the parts checkable from text See the brand voice document
A spreadsheet and any AI chat tool Anything you already pay for All arithmetic here is add-and-divide

Traffic numbers, conversion data and anything outside the artifact itself are deliberately not inputs.

How do you do it, step by step?

Pass one (steps 1–3) produces a checklist whose reliability you have measured — useful on its own, and where most teams should stop for a month. Pass two (steps 4–6) hands it to an AI judge and measures the judge. Automating an unmeasured checklist is only a faster way to be inconsistent.

Step 1 — What do you look at before writing any checks?

Open 20–30 recent AI outputs, read them in sequence, and write one line each: what is wrong with this piece, in your own words, no categories. Stop when about 20 consecutive items produce nothing new — the stopping heuristic from error-analysis practice, whose stated rule of thumb is to review at least 100 items in total (Husain & Shankar). Grouping notes you already wrote is fine work for an AI; inventing the criteria is the part you cannot outsource.

Step 2 — How do you turn failures into checks a stranger could apply?

Each frequent failure becomes a question answerable Yes or No from the text alone, naming an observable behaviour and requiring a quoted passage as proof. Keep 5–7; an artifact passes only if it passes all of them. Weak: Is the tone on-brand? Workable: Does every paragraph describing the product use first-person plural ("we")?

Rewrite each failure below as a binary acceptance check.

Rules:
- The check must be answerable Yes/No by reading the text alone.
- The check must name an observable behaviour in the text, not a judgement
  ("uses a specific number", "names the audience"), never "is good", "is engaging".
- The check must be answerable without knowing our business results.
- If a failure cannot be turned into such a check, output it under "NOT CHECKABLE".

Output format:
| check_id | check question (Yes/No) | what evidence must be quoted to answer it |

Failures:
[PASTE YOUR GROUPED NOTES FROM STEP 1]

Why binary rather than 1–10 is answered below; the short version is that pass/fail labels are easier for two people to reproduce, and reproducibility is the object here.

Step 3 — How do you check that two people apply the checklist the same way?

Pick 30–50 artifacts including at least a third obvious failures, number them, shuffle them, and freeze the set. Two people then label it independently, without seeing each other's verdicts, quoting evidence for each one. Freezing matters because your own standards move while you work: evaluation criteria are documented to shift as people label, which is why criteria-generating tools need human validation on a held sample (Shankar et al., UIST 2024). A frozen set is the only way to tell "the model got worse" from "I got stricter".

Then, per check, build the 2×2. Illustrative example, not a measurement:

Reviewer B: pass Reviewer B: fail row total
Reviewer A: pass 22 5 27
Reviewer A: fail 3 10 13
column total 25 15 40

The two agreed on 22 + 10 = 32 of 40 items, so raw agreement is 80%. Now subtract the agreement luck alone would produce: A says pass 27/40 of the time and B 25/40, so both saying pass by coincidence is 0.675 × 0.625 = 0.42; both saying fail is 0.325 × 0.375 = 0.12; together 0.54. The chance-corrected figure — Cohen's kappa — is the share of the remaining room you actually captured: (0.80 − 0.54) ÷ (1 − 0.54) = 0.56 (Cohen, 1960).

Eighty per cent felt like a strong result and turned out to be 0.56. Where reviewers disagreed, read both quoted passages and rewrite the wording of the check, not the people. Relabel and repeat until the figure stops improving.

Step 4 — How do you hand the checklist to an AI judge?

One check, one call. Ask for reasoning first, then a quote, then a one-word verdict — reasoning before scoring is recommended by both major vendors (OpenAI; Anthropic, whose grading examples also carry the note that it is generally best practice to grade with a model other than the one that generated the output).

You are grading one piece of [ARTIFACT TYPE] against exactly one criterion.

CRITERION
[PASTE ONE CHECK QUESTION]

WHAT COUNTS AS EVIDENCE
[PASTE THE "required evidence" CELL FOR THIS CHECK]

INSTRUCTIONS
1. Quote the single passage from the text most relevant to this criterion.
   If no such passage exists, write: NO RELEVANT PASSAGE.
2. In no more than two sentences, explain how that passage does or does not
   satisfy the criterion.
3. Output the verdict on the last line, alone, as exactly one word: PASS or FAIL.
4. Judge only this criterion. Ignore spelling, length, style, and everything else.

TEXT
[PASTE ARTIFACT]

Step 5 — How do you measure the judge's two different mistakes?

Run the judge over the same frozen sample and compare against the agreed human labels. Record two rates separately: misses (human said fail, judge said pass) and false alarms (human said pass, judge said fail). One accuracy figure hides the asymmetry — a miss costs you a published mistake, a false alarm costs ten minutes. Practice guidance frames the target the same way, as high true-positive and true-negative rates on a held-out labelled set, and puts judge development at 100+ labelled examples (Husain & Shankar).

Before quoting any judge number, run the three stress tests detailed in the break section below: reverse the order of paired comparisons, inflate a passing text without adding information, and repeat an identical run. Each produces a number and a date.

Step 6 — When do you redo this?

The published error-analysis cadence is a review cycle every 2–4 weeks over 100+ fresh items, with 10–20 items a week in between focused on outliers (Husain & Shankar). Recalibration is mandatory whenever you switch models, change the generating prompt, or change the artifact format.

What does a good result look like?

One screen that survives the sentence "well, I don't like it". Illustrative values, not measurements:

check_id judge PASS rate miss rate false alarm rate reviewer kappa last recalibrated
C1 no unsourced statistics 78% (n=40) 6% 11% 0.71 2026-07-20
C2 audience named in first 100 words 92% (n=40) 3% 4% 0.84 2026-07-20
C3 one concrete customer example 61% (n=40) 18% 9% 0.52 2026-07-20

C3 is the useful row: a kappa of 0.52 and an 18% miss rate mean the check is worded too loosely for people and unreliable for the judge, so its 61% is not yet a fact about quality. Reported, that reads:

Judge PASS rate: [X]% on [N] items, run [DATE]
Judge miss rate: [M]% · false alarm rate: [F]% (measured on [K] human-labelled items, [DATE])
Human agreement on checklist: kappa = [VALUE] between 2 reviewers, [DATE]
Checklist version: [vN] · last recalibrated: [DATE]

How do you know the output is good?

Each criterion above is audited by looking, not by judging: walk the checklist hunting for one adjective of taste, open the calibration folder and confirm the item count and the date, open the agreement tab and confirm a chance-corrected figure sits beside every percentage.

Two cautions about thresholds. The familiar bands — 0.41–0.60 "moderate", 0.61–0.80 "substantial" — are labels Landis & Koch attached to ranges in 1977, reproduced since in reference tables (AHRQ methods table). They are a convention, not a measurement: that table supplies adjectives, not evidence that 0.61 is where a checklist becomes usable. Use kappa as a tracking number — did rewording the check raise it? — not as a passing grade.

Second, calibrate expectations to the fuzziness of what you measure. In one widely used summarisation dataset, expert human agreement ran 0.798 on consistency but 0.588 on fluency and 0.398 on relevance, as reported in a 2025 study of judge self-consistency (Rating Roulette). A check sitting near 0.4 is almost always a wording problem, not a people problem.

Where does this usually break?

Does the order you show two versions in change the answer?

Yes, measurably. In the 2023 study that introduced MT-bench judging, swapping the order of two answers left the verdict unchanged in only 65.0% of cases for GPT-4, 46.2% for GPT-3.5 and 23.8% for Claude v1 under the study's default judging prompt, with the bias running toward whichever answer came first; rewording that prompt moved Claude v1 to 56.2%, so the size of the effect is partly a property of how you ask (Zheng et al., NeurIPS 2023). A 2024 paper showed the effect exploited outright: by changing presentation order, a weaker model "won" 66 of 80 prompts under a ChatGPT judge (Wang et al., ACL 2024). Fix: run both orders, average them, and read by hand every pair whose verdict flipped.

Does padding a text make it score higher?

Length moves verdicts — but check which way yours moves rather than assuming. Rewriting list-style answers more verbosely, adding no information, flipped the verdict toward the padded version in 91.3% of cases for Claude v1 and GPT-3.5, and 8.7% for GPT-4 (Zheng et al.). On a 2024 benchmark, the baseline model's win rate swung from 22.9% to 64.3% purely from instructions about verbosity, tightening to 41.9–51.6% once length was controlled (Dubois et al., COLM 2024); OpenAI's guidance lists the same tendency as something to control for. The direction is not universal: in an informal 2025 experiment on code revisions, judges preferred the shorter answer (Couch). Fix: inflate five passing artifacts and watch which way your judge moves.

Does the judge agree with itself?

Often less than you would guess. Running the same prompt three times, a 2025 study measured self-agreement of 0.33–0.79 across three models on a binary summary-consistency task, and 0.27–0.56 on a three-way ranking task — where even the most self-consistent model returned identical judgements on all three runs in only 61.3% of cases (Rating Roulette). On the binary task, majority voting over the three runs beat a single run on balanced accuracy, while switching sampling off (temperature 0) scored slightly below the sampled runs — "just make it deterministic" is not the fix. Fix: run the judge three times and take the majority when repeat agreement is below your bar.

The remaining failure modes are cheaper to read about than to discover:

What goes wrong Why it happens What to do
You report raw agreement Percentages flatter: in a 2024 evaluation, judges above 90% raw agreement still assigned scores more than 10 points away from the human-assigned ones, and the best judge trailed the human alignment score by 8 points on the chance-corrected measure (Thakur et al.) Always print kappa beside the percentage
The calibration set has no failures A miss rate cannot be computed against labels that are all "pass" Keep at least 30% human-labelled failures
The AI wrote the labels too Generated evaluators inherit the flaws of the models they grade and need human validation (Shankar et al.) Human labels are the only ground truth in steps 3 and 5
The checklist was calibrated once Models, prompts and formats all move Recalibrate on the cadence in step 6

What else do people ask?

Is a 1–10 scale really worse than pass/fail?

For this job — a small team labelling artifacts for acceptance, and an AI judge repeating that labelling — yes: pass/fail labels are easier to reproduce, the difference between a 3 and a 4 is defined nowhere, OpenAI's guidance recommends pairwise comparison or pass/fail for reliability, and practice guidance treats numeric labels as advanced and usually unnecessary (OpenAI; Husain & Shankar).

Do not widen that into "scales are bad". In classical psychometrics the result is the opposite: 2-, 3- and 4-point scales showed relatively poor reliability, validity and discriminating power, with indices rising up to around seven categories (Preston & Colman, 2000). That research averages many items across many respondents; you are asking two editors for one verdict on one artifact. Vendors differ too — Anthropic documents a 1–5 scale with explicit anchors as a legitimate grading pattern (Anthropic). Workable rule: binary by default, a scale only once you have measured that two people reproduce its steps.

Can the model that wrote the text also judge it?

Prefer not. Beyond Anthropic's note that grading is generally best done with a different model, models demonstrably identify their own output, and that self-recognition tracks with preferring it (Panickssery et al.). The original judging study also estimated that judges rate their own answers more favourably than humans do, but its authors flagged the data as thin and the differences small (Zheng et al.) — a reason for caution, not a number to quote. If one model is all you have, validate the judge on a second model and compare the error rates.

What if I am the only reviewer?

Label the frozen sample, wait a week, label it again without looking at your first pass, and run the same arithmetic against yourself. Acceptance checks do not replace source and statistics checking or editing to a publishable standard; the machinery behind an AI grading another AI's work is described on llm-as-judge.

Sources

Verification

4 log entries
dateactionresult
2026-07-27researchapplied
2026-07-27draftapplied
2026-07-27correctionapplied
2026-07-27fact-checkpass-3-0

Backlinks