---
name: ai-output-verification
description: >
  Audits AI-produced research work before anyone relies on it, using a
  verification procedure built around the specific ways AI output fails:
  fabricated specifics, plausible citations, silent gap-filling, restatement
  presented as synthesis, quietly dropped contradictions and uniform confidence.
  Use when someone says "can we trust this AI analysis", "check the AI's output",
  "verify this before it goes to the client", "did the model make this up",
  "the AI summarised our transcripts", "audit this AI-generated report", or
  "how do I know these themes are really in the data".
category: 13 Research Quality, Ethics and Governance
ref: "13.03"
tier: 1
inherits: [K2, K3, K4, K5]
---

# AI Output Verification

## 1. One-line description
Verifies an AI-produced research output against the source material it was given, targeting the failures characteristic of AI rather than those characteristic of people, and returns a verification report with a trust judgement that is permitted to be "this output cannot be relied on".

## 2. What this skill is used for

**The research problem it solves.** AI research output fails differently from human research output, and a reviewer applying human-error instincts will look in the wrong places. A tired analyst produces work that is visibly rough: gaps, inconsistencies, hedging in the wrong spots, a section that trails off. An AI system produces work that is uniformly fluent, internally consistent, correctly formatted, appropriately toned, and wrong in ways that leave no surface trace. **The output looks most finished exactly where it is most invented**, because a fabricated specific has no messy provenance to complicate it.

The characteristic failure profile is stable enough to audit against directly:

- **Fluent, confident wrong answers.** Errors arrive in the same register as correct results. There is no tell in the prose.
- **Fabricated specifics.** Numbers, participant identifiers, dates, sample sizes and named details that were never in the input, generated because the format required something in that position.
- **Plausible citations.** Real-sounding sources that do not exist, handled in full by **13.02**.
- **Restatement presented as synthesis.** A finding paraphrased at a higher level of abstraction and labelled an insight, with no explanatory content added (K2 §2.2).
- **Silent gap-filling.** Where the source material is thin or absent, the output completes the pattern rather than reporting the gap. This is the single most damaging failure, because the filled gap is invisible and the reader has no reason to look for it.
- **Over-resolution of ambiguity.** Where a passage supports two readings, the more coherent one is chosen and the ambiguity disappears, which is exactly wrong when the ambiguity is the finding (K5 §2.6).
- **Quiet dropping of contradictory material.** Evidence that does not fit the emerging structure is not argued away; it simply does not appear.
- **Uniform confidence.** The same declarative register applied to a finding on n=1,200 and one on three interviews, so the reader cannot tell them apart (K3 §1).

Verification is not scepticism about AI. It is the recognition that these failure modes are systematic, which means they can be tested for systematically, which is a far better position than the one research has with human error.

**Where it sits.** Cross-cutting, and it runs immediately after any AI-performed step and before that step's output is used for anything. Its findings feed **13.01 Research Quality Review** as an input rather than replacing it.

**Typical use cases.**
- Auditing AI-produced coding, theme development or classification before the coded data is analysed.
- Verifying an AI-drafted analysis, summary or report section before it reaches a client or a stakeholder.
- Checking an AI synthesis of transcripts, open ends or documents against the sources it was given.
- Establishing whether an AI-generated set of insights or recommendations is evidenced or merely coherent.
- Auditing an inherited output where nobody is certain how much of it was AI-produced.
- Producing the human-verification record that K4 §7 and K5 §7 require to be disclosed.

**Who uses it.** Any researcher using AI in delivery work, at every level; quality leads setting a standard for a team; client-side reviewers receiving AI-assisted deliverables; research leads who must sign their name to work a system produced (K5 §2.8).

## 3. When to use it

- Any AI-produced output is about to be used for anything beyond the researcher's own orientation.
- An output contains specifics (numbers, quotes, identifiers, citations, dates, sample sizes) that a reader would take as measured.
- An AI system performed coding, classification, extraction or summarisation at a scale nobody has read in full.
- An output reads unusually well and no one can immediately say where any particular claim came from.
- The AI was given a large or heterogeneous input set and it is not clear how much of it was actually used.
- A deliverable carries a named researcher's accountability, per K5 §2.8.
- An output is being reused downstream, where an unverified claim will propagate into documents that inherit its authority.
- Any client, regulatory or organisational requirement obliges disclosure of AI involvement and human verification.

## 4. When NOT to use it

- **The question is the methodological soundness of the underlying study.** Whether the design could answer the question, whether the sample supports the claims, whether the analysis was right: **13.01 Research Quality Review**. This skill checks whether an output faithfully represents its inputs. An output can be perfectly faithful to a study that should never have been run that way.
- **The question is only about sources and citations.** **13.02 Source and Citation Verification** covers that in full, and this skill hands the citation component to it rather than duplicating it.
- **The question is whether the work is biased.** Selective subgroup reporting, confirmation bias, a loaded research question, the commissioner's interest: **13.04 Bias Detection**. Some overlap exists (dropped contradictions appear in both) and the taxonomies are different.
- **The question is whether the AI should have been used at all**, or what may be sent to it, or what must be disclosed: **13.06 AI Research Governance** for the operational rules, **13.05 Research Ethics and Consent Design** where participant data or consent is involved. Verifying an output whose production breached a consent does not make it usable.
- **The source material is not available.** Verification is comparison against inputs. Without the transcripts, the dataset, the documents or the tables the output was built from, you can check internal consistency and plausibility and nothing else. Say that explicitly, label the exercise a plausibility review, and do not let it be recorded as verification (K4 §6.4). This situation is common and it is the point at which most verification quietly stops being verification.
- **The output is exploratory scaffolding the researcher will rebuild anyway.** A first-pass structure, a set of candidate themes to react to, a list of angles: verifying these wastes effort on material that will not survive. Verify at the point the output starts being relied on, and be clear with yourself about when that point arrived, because it usually arrives earlier than intended.
- **Verification is being performed to produce a sign-off record rather than to find errors.** A pass rate produced by checking the things most likely to be right is worse than no verification, because it converts an unchecked output into a certified one. If the checking is not capable of returning "cannot be relied on", it is not verification.
- **Nobody will act on a negative result.** Where an output will ship regardless of what verification finds, say so before starting, per K4 §9. The honest response is to record what was not verified rather than to spend the effort and have it ignored.

## 5. Required inputs

**Required.** Without these, verification cannot be performed as verification.
- **The AI output**, complete, in the form it will be used.
- **The complete input set the AI was given**: transcripts, dataset, tables, documents, prior outputs. Not a description of them. The verification boundary is the input set, and anything specific in the output that is not derivable from it is a candidate fabrication.
- **The instruction or prompt, and the processing sequence.** What the AI was asked to do, in what order, and whether the output is the result of one step or a chain of them. A chain matters: an error introduced at step two is inherited by every step after it, and it will look like consistent corroboration.
- **A statement of which parts are AI-produced and which are human-written**, where the output is mixed. If nobody can say, treat the whole output as AI-produced.

**Optional, and what each one adds.**
- **The model or system version and settings.** Makes a finding reproducible, and matters when an output is re-run and does not reproduce (13.06).
- **Intermediate outputs from each step in a chain.** Localise where an error entered, which turns a rejected output into a fixable one.
- **A human-produced comparison on a subset.** The strongest available check on classification and coding work: a human codes a sample independently, and the agreement rate is the measured accuracy that K4 §7 requires to be disclosed.
- **The analysis plan or code frame the output was supposed to apply.** Converts "does this look right" into "did it do what it was told", which is a far more answerable question.
- **Prior verification records on the same pipeline.** Establish where this configuration has failed before, which is the best available guide to where to look first.

## 6. Questions to ask before starting

1. **Exactly which steps did the AI perform, and on what inputs?** Determines the verification boundary and the expected failure profile. *Default if unanswered:* treat everything as AI-produced from the full input set, which is the conservative position and usually the true one.
2. **Is this a single step or a chain, and are intermediates available?** Determines whether an error can be localised or only detected. *Default:* assume a chain, and verify the earliest available output first, since errors there are inherited everywhere.
3. **What is the output for, and what happens if a claim in it is wrong?** Sets coverage and the escalation threshold. *Default:* verify to client-delivery standard.
4. **Has any human already checked any of this, and how?** Determines what can be relied on and prevents double-verifying the checked while leaving the unchecked. *Default:* assume nothing has been checked. A read-through is not a check.
5. **Is the input set complete, and did the AI actually receive all of it?** Truncation, file failures and context limits produce outputs built on a subset while reading as though built on everything. *Default:* verify coverage explicitly at step 2, since this is invisible from the output.
6. **What did the researcher expect the output to say?** Determines where over-resolution and agreeable gap-filling are most likely, because an output produced under a stated expectation will tend to meet it. *Default:* ask, and record the answer before verifying.

## 7. Step-by-step methodology

**Step 1. Establish the verification boundary.** Write down what the AI was given, what it was asked, what it produced, and where human hands touched it. Then state the rule that governs everything after: **anything specific in the output that cannot be located in the input set, or derived from it by a stated operation, is a candidate fabrication until shown otherwise.** Not "possibly wrong". A candidate fabrication, carrying the burden of proof. This inversion is the whole discipline, and it is what separates verification from reading critically. *Correct result:* a one-paragraph boundary statement and a list of the output's specific claims.

**Step 2. Check input coverage before checking output content.** Establish whether the AI actually received and used the whole input set. Take a sample of source items from across the set (early, middle and late, since truncation is positional) and check whether each is represented in the output. Then check the arithmetic: if 40 transcripts went in and the output's prevalence counts sum to a maximum of 26 participants, something was not read. Truncation, failed file loads and silently dropped inputs produce outputs that describe a subset with the confidence of a census, and this is invisible from the output alone because the output never mentions what it did not see. *Correct result:* a coverage statement naming any input not represented, and a verdict on whether the output describes the whole input set.

**Step 3. Predict where this output will fail, then look there first.** The failure profile varies by task, and verification effort should not be spread evenly. **Extraction and coding** fails on boundary cases, on rare categories, and by over-applying the dominant code. **Summarisation** fails by dropping the minority position and the qualifying clause. **Synthesis** fails by restatement and by over-resolving ambiguity. **Drafting** fails on specifics: numbers, names, dates and citations generated to fill positions. **Quantitative description** fails on bases, filters and direction. Write the two or three predicted failure types for this output before opening it properly. Verification without a hypothesis becomes reading, and reading fluent text finds nothing.

**Step 4. Verify numbers against source, exhaustively where they matter.** Every number in the output falls into one of three classes: **recomputed** (you calculate it from the source data and compare), **located** (you find it in a supplied table or document and compare, including its base, filter and unit), or **unsupported** (neither possible). Coverage is not uniform, and the rule is fixed:

**Exhaustive, never sampled:** every number in the executive summary; every headline claim; every figure in a chart title, callout or recommendation; every base size; every prevalence count in qualitative work.

**Sampled:** everything else, stratified across sections and claim types rather than taken in document order, because errors cluster by section.

**The escalation rule:** any fabrication found in a sampled class converts that class to exhaustive coverage. One invented number is not one error. It is evidence about the process that produced every other number in the output.

For each number check four things beyond the digits: the base it is calculated on, the filter applied, the direction (a reversed scale or an inverted comparison produces a number that is right and means the opposite), and the unit. *Correct result:* every number carries a status of recomputed, located, or unsupported, and every unsupported number is either removed or marked.

**Step 5. Verify every quote. No sampling, ever.** Locate each quote in the source, character by character, and check three things: that the words match (permitted edits per K4 §2.3 only), that the participant identifier is right, and that the participant actually belongs to the segment the quote is attributed to. Quotes are the highest-consequence class because a fabricated quote is indistinguishable from a real one to everyone downstream and there is no later opportunity to catch it. Two specific AI failures recur: **merged quotes**, where two participants' words are combined into one fluent utterance, and **smoothed quotes**, where hesitation, self-contradiction and dialect have been tidied into something more quotable, which changes the evidence. Where a quote cannot be located in source, it is removed, not softened. Hand the selection and attribution standard to **07.04**.

**Step 6. Verify every citation.** Coverage is total, per **13.02**, and the reason is the base rate: in AI-assisted work the plausible-but-unreal citation is an expected output rather than a rare error. Run 13.02 and bring its log in as a component of this report rather than restating its procedure.

**Step 7. Trace a stratified sample of claims back through the evidence chain.** Take claims from each K2 level (finding, interpretation, insight, implication, recommendation), weighted toward the top of the chain because that is where the failures concentrate, and walk each one backwards to source per K2 §2. You are looking for four specific things:

- **Level inflation.** A finding restated at a higher altitude and labelled an insight, adding no explanatory content. The test: does it say *why*, or does it say the same thing in a bigger voice (K2 §2.2)?
- **Silent interpretation.** An inference written in the grammar of an observation, with no signal word between the finding and the judgement (K2 §3.2).
- **The orphan recommendation.** An action with no traceable finding beneath it. Common, plausible, and usually a general truth about the domain rather than anything this study established.
- **The unsupported bridge.** A connective sentence asserting a relationship that no input establishes, holding two evidenced sections together.

*Correct result:* every traced claim ends at a specific piece of source material or is recorded as untraceable, and untraceable is a defect rather than a gap to be filled later.

**Step 8. Check significance and causal language, then check confidence calibration.** Two passes, both mechanical, both high-yield.

**The language pass.** Search the output for every comparative and causal construction: "significantly", "notably", "clearly", "markedly", "drives", "leads to", "because", "impact of", "results in", "improving X will". For each, establish whether a test was run (K4 §3.1) and whether the design licenses causation (K4 §3.2). AI output reintroduces causal language downstream of an analysis that was careful about it, because causal phrasing is more fluent than associative phrasing and fluency is what the generation optimises for. Also check the direction of every comparison against source, since inverted comparisons survive proofreading easily.

**The calibration pass.** Read the output for its confidence register alone, ignoring content. Then ask: **does the confidence vary?** A genuinely calibrated output states some things declaratively and others with their alternative explanation attached, and a reader can tell them apart (K3 §4). AI output tends toward one register throughout, usually confident, which means the reader cannot distinguish the claim on 1,200 responses from the one on three interviews. Map each significant claim's stated confidence against the confidence its evidence supports under K3 §3. **Uniform confidence is itself the finding**, and it is reported as one rather than corrected claim by claim, because the underlying problem is that no calibration was performed.

**Step 9. Audit for what is absent. This is the step that finds silent gap-filling, and it is the one most often skipped**, because checking what is present feels like verification and produces a satisfying tick list.

Work from the inputs, never from the output's own structure. Build the inventory of what went in: every question fielded, every participant, every document, every theme in the code frame, every subgroup. Then check each against the output. What never appears? Then run three specific searches:

- **The missing contradiction.** Find the strongest evidence in the source that cuts against the output's main conclusion. It exists in almost every real dataset. Is it in the output? If not, it was dropped rather than argued with, and that is a K4 §4.1 breach whether or not it was deliberate.
- **The resolved ambiguity.** Find passages in the source that support two readings. Does the output acknowledge the ambiguity or has it silently chosen? AI resolves toward the more coherent reading, and coherence is not evidence (K5 §2.6).
- **The absent gap.** Does the output say anywhere that something could not be established? An output covering every objective with equal confidence and no gaps is describing a research project that does not exist. **The absence of a "what we could not establish" section is a finding about the output, not a stylistic omission.**

*Correct result:* a list of inputs not represented, contradictions not reported, ambiguities resolved without acknowledgement, and gaps not disclosed.

**Step 10. Probe the output adversarially, and read the response for what it tells you.** Challenge specific claims and observe what happens. Three probes, in this order:

- **"Which participant said this, and where in the transcript?"** A grounded claim answers with an identifier and a location that check out. An ungrounded one produces a new plausible identifier, which is itself a fabrication and a strong signal about the original claim.
- **"Show me the evidence that cuts against this conclusion."** Genuine analysis produces specific contrary material from the source. Gap-filled analysis produces a generic methodological caveat, because it has no contrary material to produce.
- **A false assertion, stated confidently.** Contradict a claim you have verified as correct. An output that immediately abandons a correct claim under mild pressure is not tracking evidence, and everything else it asserts should be treated accordingly.

**Read the responses as leads, not verdicts.** Folding is not proof the claim was wrong and defending is not proof it was right; both send you to the source. This is the fastest way to find the invented claims in a long output, and it is worth doing before the exhaustive passes so that it can direct them.

**Step 11. Assemble the verification report and issue a trust judgement.** Four forms, and the fourth must remain genuinely available: **usable as produced**; **usable after the listed corrections**; **usable only for the narrower subset specified**, where part of the output is verified and part is not recoverable; and **cannot be relied on**, where the density or the position of the failures means the output cannot be repaired by correction and the work needs redoing. The threshold for the fourth is not a count. It is whether the verified parts can be separated from the unverified ones. A fabricated citation is repaired by deletion. **A fabricated prevalence count in a synthesis is not repairable, because it means the synthesis was not performed on the data**, and every other count in it is now unknown.

Then record the disclosure that K4 §7 and K5 §7 require: what the AI did, what a human verified, at what coverage, what was found, and who performed the verification. That record is the deliverable that makes AI use defensible, and it belongs in the output rather than in a file (13.06).

## 8. Analytical framework

Verification runs on two axes at once. The first is the failure taxonomy, which says what to look for:

| Failure | Where it hides | The check that finds it |
|---|---|---|
| Fabricated specific | Numbers, IDs, dates, sample sizes, names | Locate in input, or it is invented (steps 4, 5) |
| Plausible citation | Reference lists and in-text sources | Locate the item (13.02, step 6) |
| Restatement as synthesis | Insight and implication sections | The "does it say why" test (step 7) |
| Silent gap-filling | Thin or absent input areas | The absence audit against the input inventory (step 9) |
| Over-resolved ambiguity | Qualitative interpretation | Return to the passages that support two readings (step 9) |
| Dropped contradiction | Anywhere the story is clean | Find the strongest contrary evidence in source and look for it (step 9) |
| Causal drift | Fluent prose downstream of careful analysis | The language pass (step 8) |
| Uniform confidence | The whole document | The calibration pass, read for register alone (step 8) |

The second is coverage, which says how much:

    Exhaustive: headline claims, executive summary, recommendations,
                all quotes, all citations, all bases, all prevalence counts
    Sampled:    everything else, stratified by section and claim type
    Escalation: any fabrication found → that class becomes exhaustive
                any fabricated quote or prevalence count → the output's
                trust judgement starts at "cannot be relied on"

**Applying it.** Run the absence audit before the presence checks if time is tight. Checking what is present will find errors; only checking what is absent will find the gap-filling, and gap-filling is the failure with no surface signature and the largest downstream cost. The most common verification failure in practice is a thorough, well-documented, correct check of everything the output contains, which says nothing about what the output should have contained.

## 9. Output format

**A. Verification scope.** What the AI produced, from what inputs, under what instruction, in how many steps; what human work is mixed in; what was available to verify against; the coverage rule and why; and who verified, when.

**B. Trust judgement.** One of the four forms, in the first fifty words, with the findings that drive it.

**C. Verification log.**

| Ref | Claim or element | Type | Location in output | Source checked | Result | Severity | Action |
|---|---|---|---|---|---|---|---|

Types: number, quote, citation, prevalence, base, comparative claim, causal claim, insight, recommendation, bridge. Results: verified, verified with correction, unsupported, not located, fabricated, misattributed, uncheckable.

**D. Absence audit.** Inputs not represented; contradictions in source absent from the output; ambiguities resolved without acknowledgement; objectives with no gap statement. This section is usually the most informative and it is written even when empty, with the searches that were run.

**E. Calibration assessment.** Whether confidence varies, and the claims whose stated confidence exceeds what their evidence supports under K3 §3.

**F. Corrections required**, each with a specific action, and consequential edits where removing a claim breaks something downstream.

**G. Disclosure record**, per K4 §7 and K5 §7: what the AI did, what was verified, at what coverage, by whom, and what was found. Written to be reproduced in the deliverable.

**When the evidence is thin.** Where the input set is incomplete, the report says which parts could not be verified and the trust judgement is qualified to the verified subset. Where the source material was not available at all, the output is labelled a plausibility review and is not recorded as verification (K4 §6.4). **A verification that finds serious problems is a successful verification.** The report does not soften a fabrication into a "possible inconsistency", and it does not pad a clean result to look diligent: where an output verifies clean, the report says so briefly and lists exactly what was checked and at what coverage, so the assurance is inspectable rather than asserted.

## 10. Quality checks

Run on the verification. K4 §8 runs anyway.

1. Was input coverage established before output content was checked?
2. Was every quote verified, with no sampling, including participant ID and segment membership?
3. Were all citations verified per 13.02, at total coverage?
4. Were all executive summary and headline numbers recomputed or located, rather than sampled?
5. Was the sample for the remainder stratified across sections and claim types rather than taken in document order?
6. Was the escalation rule applied wherever a fabrication was found?
7. Was the absence audit run against the input inventory rather than against the output's own structure?
8. Was the strongest contradicting evidence in the source actively searched for, and its presence or absence in the output recorded?
9. Was confidence read as a register across the whole document, not only claim by claim?
10. Was every causal and comparative construction traced to a design and a test?
11. Were adversarial probes run, and were the responses treated as leads to check rather than as answers?
12. Does the trust judgement follow from the findings, and was "cannot be relied on" genuinely available?
13. Is the disclosure record complete enough for someone else to know what was and was not checked?
14. Where verification could not be completed, is the exercise labelled accordingly rather than recorded as verification?

## 11. Common failure modes

| Failure | How to recognise it | How to prevent it |
|---|---|---|
| **Verifying only what is present** | A thorough log, every claim checked, and a silently dropped contradiction still in the output | Step 9. Run the absence audit first when time is short |
| **Reading instead of verifying** | The verifier reports the output "reads well and seems consistent" | Internal consistency is what AI output has by construction. Every check ends at a source, not at a judgement |
| **Uniform sampling** | 10% of everything checked, including quotes and headline numbers | The coverage rule in step 4. Some classes are never sampled |
| **Stopping at the first pass rate** | "38 of 40 checks passed" reported as a positive result | Two failures in a sample means the population failure rate is unknown. Escalate the class to exhaustive |
| **Accepting a resolving reference as verification** | A citation is real, and the paper does not say what is claimed | Read the passage (13.02) |
| **Verifying a chain from its output only** | An error introduced at step two is confirmed by every later step and looks corroborated | Verify the earliest available intermediate first |
| **Assuming complete input** | The output describes 40 interviews; 26 were actually read | Step 2, always, before anything else |
| **Treating a probe response as a verdict** | The output defended the claim confidently, so the claim was accepted | Probes generate leads. Every lead ends at source |
| **Verification as sign-off theatre** | The record exists, the coverage is unstated, and no finding could have stopped delivery | State the coverage rule and the escalation threshold before starting |
| **Fixing rather than judging** | Each error is corrected as it is found, and no one assesses what the error density means | Log first, judge second, correct third. Density is the judgement |
| **The repaired fabrication** | A fabricated prevalence count is corrected and the synthesis is retained | A fabricated count means the analysis was not performed on the data. Correction does not restore it |
| **Verifier drift** | Later sections checked faster and more leniently than earlier ones | Fixed coverage rule, stratified sampling, and a record of who checked what and when |

## 12. AI guardrails

Universal prohibitions are inherited from K4. This skill is largely an application of K4 §2 and §8, and the following are the rules specific to performing the verification itself.

1. **Never verify an output using the system that produced it as the authority on whether it is right.** Its account of its own reasoning is generated text, not a record. Use it to locate claims in source, never to confirm them.
2. **Never accept a specific that cannot be located in the input set.** Not for numbers, participant identifiers, dates, sample sizes, prevalence counts, or named details. Plausibility is not provenance.
3. **Never sample quotes, citations, bases or headline numbers.** These classes are exhaustive without exception, and a verification that sampled them has not been performed.
4. **Never correct a fabrication silently.** Every one is logged, because the count and position of fabrications is the evidence behind the trust judgement, and correcting as you go destroys it.
5. **Never treat the absence of a detected error as evidence of correctness in an unchecked area.** Report coverage, and state what was not checked.
6. **Never let a challenged claim be repaired with new supporting detail.** An output that produces a fresh participant identifier, a fresh figure or a fresh source when challenged has fabricated twice, and the second fabrication is the more informative one.
7. **Never soften a finding's classification.** A fabricated quote is a fabricated quote, not an attribution inconsistency. The vocabulary of the log is what the trust judgement is built from.
8. **Never issue a trust judgement the findings do not support**, and never withhold "cannot be relied on" because the output is needed. Deadline pressure is the condition under which this judgement most needs to remain available (K4 §9).
9. **Never record a plausibility review as a verification.** Where the source material was unavailable, the label changes and so does what the record may be used to claim.

## 13. Best-practice principles

1. **AI output fails differently from human output, so look in different places.** Human work shows its weakness on the surface. AI work is most polished exactly where it is most invented, because a fabricated specific has no awkward provenance.
2. **Fluency is not evidence and internal consistency is not corroboration.** Both are properties of how the text was generated, and both are present in a completely fabricated output.
3. **Check what is missing before you check what is there.** This is the single highest-yield habit in the skill. Present-claim checking finds errors; absence auditing finds the gap-filling that has no signature.
4. **Anything specific that is not in the input is invented until located.** Make this the default and verification becomes a mechanical search rather than an exercise of judgement.
5. **One fabrication changes the base rate for everything else.** It is not an isolated error, it is information about the process, and the coverage rule should respond to it immediately.
6. **Uniform confidence is a finding.** Where everything is asserted in one register, no calibration was performed, and correcting individual claims does not fix that.
7. **The cleanest story deserves the most suspicion.** Real evidence is untidy. An output where every stream converges and nothing contradicts has usually had the contradictions removed rather than resolved.
8. **Verify the earliest step in a chain first.** Errors introduced early are inherited late and arrive looking like independent agreement.
9. **Some errors are correctable and some invalidate the work.** A bad citation is deleted. A fabricated prevalence count means the analysis was not performed on the data, and no amount of correction restores it.
10. **Probes find leads faster than reading does, and prove nothing on their own.** Use them to direct the exhaustive checks, and always end at the source.
11. **Record the coverage, not just the result.** "Verified" without a coverage rule is an assertion. "Verified: all quotes, citations and headline numbers exhaustively; remaining claims at one in five, stratified" is a fact somebody else can rely on.
12. **The verification record is the thing that makes AI use defensible.** Not the quality of the output, which nobody downstream can assess, but the existence of a specific, dated, named account of what a human checked (K5 §7).

## 14. Worked example

*Fictional scenario, used for illustration only. The organisation, participants, figures and quotes below are invented for the purpose of demonstrating method.*

**INPUT.** An international NGO ran 40 depth interviews with beneficiaries of a rural water programme across three districts, to inform a funder report. An AI system was given the 40 transcripts and asked to produce a thematic synthesis with prevalence, illustrative quotes and recommendations. The synthesis is 14 pages, reads extremely well, and reports six themes with prevalence counts, 22 quotes and five recommendations. The research lead must sign it.

**PROCESS.**

*Steps 1 and 2, and the first finding.* Boundary established: 40 transcripts in, one processing step, no human editing. Input coverage is checked first. Participant identifiers appearing anywhere in the output run from P01 to P27, with none above P27. The transcripts are ordered by district, and the third district's interviews are P28 to P40. The output therefore describes two districts and presents itself as describing three, with no statement anywhere that any material was unread. Checking the file handover confirms the last thirteen transcripts were not successfully loaded. **This single check invalidates every prevalence count in the output**, and it was invisible from the output itself, which discusses district three in general terms drawn from the first two.

*Step 3.* Predicted failures for a synthesis task on qualitative input: restatement as insight, over-resolved ambiguity, dropped contradictions, and prevalence counts as the highest-risk fabricated specific.

*Steps 4 and 5.* Prevalence counts: all six are wrong on the corrected base, and two are wrong even against the 27 transcripts actually read (one theme reported at "18 of 40" appears in 11 of the 27). Quotes: 20 of 22 locate correctly. One cannot be found in any transcript, in any wording. One is a merger of two participants' sentences from different interviews, joined fluently, attributed to a single participant. Both are removed and logged as fabrications, not as attribution errors.

*Step 7, and the judgement call.* A recommendation reads: "Establish community maintenance committees to sustain the water points." It is sensible, standard in the sector, and traces back to nothing. Two participants mention maintenance; neither mentions committees, and neither describes a governance problem. This is an orphan recommendation drawn from domain knowledge rather than from the interviews. The judgement call is that the recommendation may well be good advice, which is exactly why it is dangerous: it will be read by the funder as a finding from beneficiaries. Disposition is removal from the findings, with a note that it may be raised as the NGO's own recommendation, clearly separated from what participants said.

*Step 9, the absence audit, and the largest finding.* Working from the transcripts rather than the output: nine of the 27 read participants describe the water points as functioning well and their main difficulty as the distance to them, which is a different problem with different implications and no funding relationship to the maintenance narrative. The output does not contain this. It is not argued against; it is absent. The output's central theme, "sustainability of infrastructure", was constructed from the 18 who fit it, and the nine who did not simply do not appear. This is a K4 §4.1 breach and it is the reason the output cannot be repaired by correcting counts.

Two ambiguities are also found resolved: several participants describe the programme in terms that could be gratitude or could be deference to an interviewer associated with the funder, and the output reads them uniformly as satisfaction, with no acknowledgement of the alternative reading (K5 §2.2 and §2.6).

*Step 10.* Probing the missing quote: asked which participant said it and where, the output supplies "P19, discussing the dry season", and P19's transcript contains no such passage. The second fabrication confirms the first. Probing the central theme: asked for evidence cutting against it, the output produces a generic caveat about qualitative sample sizes rather than the nine participants who contradict it, which it had never registered.

*Step 11.* Trust judgement: **cannot be relied on**. Not because of the error count, but because of the position of the errors. The prevalence counts are wrong at their base, the central theme was built by omission, and two quotes were fabricated. The verified components (20 located quotes, the thematic vocabulary, the structure) are not separable from the unverified ones, because the theme structure is what the omission produced. The synthesis is redone on the complete transcript set, with the code frame reviewed by a human before it is applied and the distance-to-water-point counter-theme carried explicitly.

**OUTPUT.** A verification report opening with the trust judgement; a scope statement recording that 13 of 40 inputs were never read; a log of 51 checked elements with two fabricated quotes, six incorrect prevalence counts, one orphan recommendation and three unsupported bridges; an absence audit naming the nine contradicting participants and two unacknowledged ambiguities; a calibration assessment noting a single confident register throughout, including on the themes resting on the smallest counts; and a disclosure record stating what the AI did, what was checked, at what coverage, by whom and what was found, for inclusion in the funder report per K4 §7. `RESEARCHER SIGN-OFF REQUIRED` per K5 §3.1 on the reworked synthesis, and a K5 §2.2 review point on the gratitude-or-deference reading, which needs someone with district context rather than another pass by the system.

## 15. Advanced usage

**Verifying classification and coding at scale.** Where an AI has coded thousands of items, item-level verification is impossible and the right instrument is a measured agreement rate: a human independently codes a random sample, agreement is computed per code rather than overall, and the per-code rates are disclosed under K4 §7. Overall agreement hides the failure that matters, because rare and boundary categories are where accuracy collapses while the aggregate stays high. Stratify the sample to over-represent rare codes, and report the rate the client should rely on, which is the worst per-code rate among codes that carry a finding.

**Verifying a chain of AI steps.** Verify backwards from the output to locate the failure and forwards from the earliest step to bound the damage. An error introduced in an early summarisation step is inherited by everything after it and presents as consistency, which is why a chain's output can be internally flawless and entirely wrong. Where intermediates were not retained, that is itself a governance finding (13.06).

**Building verification into the workflow.** Verification is cheapest at the step boundary, where the input set is small and the source is at hand, and most expensive at the end, where it costs the rewrite. Teams that verify per step and keep a running log get a defensible record almost free; teams that verify at the end usually verify a sample and call it coverage.

**When the standard approach does not fit.** Where an output is too large to verify meaningfully and too integrated to verify in parts, do not thin the coverage until it fits the time available. State that the output cannot be verified at the coverage its use requires, and either narrow what it is used for or rebuild it in verifiable steps. **"This could not be verified" is a legitimate and useful conclusion**, and it is considerably more useful than a verification whose coverage nobody wrote down.

## 16. Skill chain

**Recommended previous skills:**
- Any skill whose output was AI-produced. This one runs against the output of the library, not after a particular predecessor.
- **13.02 Source and Citation Verification.** Runs as the citation component of step 6 and hands back a log at total coverage.
- **13.06 AI Research Governance.** Hands over the record of what the AI was given, what version performed the step, and what the disclosure obligations are, which this skill needs at step 1.

**Recommended next skills:**
- **13.01 Research Quality Review.** Takes the verification report as an input and assesses whether the underlying research is sound, which verification does not address.
- **13.04 Bias Detection.** Takes the dropped contradictions and selective reporting findings and assesses whether they form a directional pattern rather than isolated omissions.
- **07.01 Thematic Analysis**, **07.02 Open-Ended Response Coding** or the relevant analysis skill, where the verdict is that the work must be redone rather than corrected.
- **12.03 Research Report Compilation**, which applies corrections and consequential edits properly rather than patching the output.

**Runs well alongside:**
- **07.04 Quote and Evidence Extraction**, which supplies the attribution and verbatim standard step 5 checks against.
- **05.02 Statistical Testing**, which supplies the standard for the significance language pass in step 8.
- **K2**, whose evidence chain step 7 audits directly, **K3**, whose confidence levels the calibration pass is run against, and **K5 §7**, which the disclosure record satisfies.

---
A Yazi Supplied Skill and resource.
