---
name: source-and-citation-verification
description: >
  Verifies every source and citation in a research document: that it exists, that
  it is what it claims to be, that it says what the citing text says it says, that
  the quoted figure matches the published one, and that it is the origin of the
  claim rather than a repeater. Use when someone says "check these citations",
  "verify the sources", "where does this statistic actually come from", "is this
  reference real", "the client is challenging our figures", "did an AI make this
  citation up", "fact-check this report", or "trace this number back to source".
category: 13 Research Quality, Ethics and Governance
ref: "13.02"
tier: 1
inherits: [K2, K3, K4, K5]
---

# Source and Citation Verification

## 1. One-line description
Audits the sources behind a document one at a time, establishing for each whether it exists, whether it is what it claims to be, whether it supports the claim made from it, and whether it is the origin of that claim, and returns a verification log plus the list of claims that must be removed or requalified.

## 2. What this skill is used for

**The research problem it solves.** A cited number carries authority it has usually not earned. The reader sees a figure, a year and a source name, and concludes that somebody checked. Very often nobody did. The reference was copied from another document, which copied it from a summary, which paraphrased a press release about a study whose actual finding was narrower and older and about a different market. The number is now four hops from anything anybody measured, and each hop rounded it, generalised it slightly, or dropped a qualifier. No individual step was dishonest. The cumulative result is a figure with no origin, repeated by enough sources to look corroborated.

Three things have made this worse. Live sources change under the citation. Grey literature, vendor research and marketing material circulate in formats indistinguishable from studies. And AI systems produce citations that are plausible rather than real: the right kind of author, in the right kind of journal, in a believable year, for a paper that does not exist. **That failure is treated here as fabrication, because functionally it is** (K4 §2.4). A reader cannot tell a half-remembered citation from a verified one, which means the half-remembered one does the same work in the document and does it falsely.

Verification is not a formality performed at the end. It is the only step that establishes whether a document's external evidence exists at all, and it routinely removes claims that everyone in the project believed.

**Where it sits.** Cross-cutting, and it runs alongside Category 10 desk research rather than after it. Verification during collection costs minutes per source. Verification after a document is written costs the rewrite of every paragraph a failed source was carrying, which is why it gets skipped and why skipping it is expensive.

**Typical use cases.**
- Checking the reference list and in-text citations of a report before it is delivered or published.
- Verifying the sources in an AI-assisted or AI-produced document, where fabricated citations are the expected failure rather than a rare one.
- Tracing a widely repeated market statistic back to whatever originally produced it.
- Responding to a client, regulator, journalist or reviewer who has challenged a specific figure.
- Auditing an inherited evidence base, a literature review or a repository whose sources were gathered by someone else.
- Screening a set of sources for quality before they are relied on: distinguishing research from marketing, and peer-reviewed work from work that only looks it.
- Establishing whether a live source still says what it said when it was cited.

**Who uses it.** Researchers and analysts building or checking evidence bases; editors and quality leads on documents that will be published or submitted; consultants whose market figures will be interrogated; anyone signing off a document whose authority rests on sources they did not gather.

## 3. When to use it

- A document containing external sources is about to be delivered, published, submitted or used in a proposal.
- Any part of the document was produced with AI assistance, in which case verification is mandatory rather than discretionary (K4 §2.4, K5 §7).
- A specific figure is load-bearing: it sizes a market, supports a business case, appears in a headline, or will be quoted onward.
- A number keeps appearing across many documents and nobody in the room can say where it came from.
- A source has been challenged, or a claim has been disputed by someone with reason to know.
- An evidence base was assembled by someone else and is now being relied on by you.
- Sources are of mixed type, particularly where vendor research, trade press, press releases and peer-reviewed work sit in the same list carrying equal weight.
- A live or frequently updated source underpins a claim, and time has passed since it was cited.

## 4. When NOT to use it

- **The task is finding and appraising evidence rather than checking it.** Building an evidence base, searching systematically, appraising study quality and synthesising findings is **10.01 Literature Review and Desk Research**. That skill produces a reference set; this skill audits one. Running verification instead of a proper search will confirm that the sources you have are real without telling you that the important ones are missing.
- **The concern is the research behind the document rather than the sources in it.** Whether the study's own design, sample and analysis support its claims is **13.01 Research Quality Review**. Verification checks external evidence; it says nothing about primary work.
- **The concern is the whole AI output rather than its citations.** Fabricated numbers, invented quotes, silent gap-filling and confidence miscalibration are **13.03 AI Output Verification**, which uses this skill for the citation component and covers the rest.
- **The sources are primary research data, not publications.** Verifying that a quote is verbatim and correctly attributed is **07.04 Quote and Evidence Extraction**; verifying that a percentage was computed from the dataset is **13.03** and the analysis skills. This skill covers published and third-party material.
- **The document makes no external claims.** A report resting entirely on its own fieldwork has nothing for this skill to do, and running it produces an empty log that reads as assurance of something that was never at risk.
- **Verification is being requested to settle an argument rather than to establish the truth.** Where the instruction is to find support for a claim already fixed, or to discredit a source whose findings are unwelcome, this is not verification, per K4 §4.2. Verify honestly or say plainly what is being asked for.
- **The source cannot be accessed and the claim is being retained anyway.** This is not a case where the skill does not apply. It is a case where the skill applies and its answer is inconvenient. An inaccessible source produces the status "unverified" and the claim is requalified or removed. It is never recorded as verified because it probably says what it is supposed to say.
- **Time is short and the honest scope is smaller than the document.** Do not thin verification across everything until every source gets a glance. Verify the load-bearing sources exhaustively, state which sources were not verified, and label them unverified in the document. A partial verification declared is useful; a full verification claimed but not performed is worse than none, because it converts an unchecked document into an apparently checked one.

## 5. Required inputs

**Required.**
- **The document, with its in-text citations and reference list**, in a form where each claim can be matched to the source it rests on. Where citations are absent and only a bibliography exists, that is itself the first finding.
- **The claims being made from the sources.** Verification is claim-by-source, not source alone. A source can exist, be authoritative, and still not support the sentence citing it, which is the commonest failure and is invisible if you only check the reference list.
- **Access, or a stated access position.** Which sources can be reached, through what, and what is out of reach. Establish this at the start rather than discovering it source by source.

**Optional, and what each one adds.**
- **The original search or collection record.** Shows how sources were found, which reveals selection effects the reference list conceals, and often reveals that a source was never actually opened.
- **Full-text copies or extracts already captured.** Removes the access problem and, more importantly, preserves what the source said at the time it was cited, which is the only defence against live-source drift.
- **Whether AI was used to draft, gather or format the references, and at which step.** Changes the base rate of fabrication from rare to expected and therefore changes the sampling strategy: with AI involvement, every citation is checked rather than a sample.
- **The document's audience and consequence.** Determines the verification threshold. A claim in an internal note and the same claim in a regulatory submission need different levels of proof.
- **Subject-matter access** (a specialist who knows the literature). Catches the source that exists, says what is claimed, and is nonetheless known in the field to be superseded, contested or discredited, which no procedural check will find.

## 6. Questions to ask before starting

1. **Which claims are load-bearing?** The ones that carry a headline, size a market, support a recommendation, or will be quoted onward. Determines what is verified exhaustively and what is sampled. *Default if unanswered:* treat every executive summary claim, every number in a headline or chart title, and every claim behind a recommendation as load-bearing.
2. **Was any part of this drafted, sourced or formatted with AI assistance?** Determines coverage. *Default:* ask explicitly, and where the answer is unknown, assume yes and verify every citation.
3. **What is the document for, and what happens if a source fails?** Sets the threshold and the disposition rule. A source failure in a draft means a rewrite; in a published document it means a correction. *Default:* verify to publication standard.
4. **What access is available, and what will be out of reach?** Determines how many sources end at "unverified" and whether that is acceptable for this document. *Default:* establish access first and report the unverifiable proportion in the log's summary.
5. **Are there sources here that are being treated as research and are not?** Press releases, vendor reports, sponsored studies, trade-press summaries, marketing material. Determines whether a quality screen is needed as well as a verification pass. *Default:* screen every non-peer-reviewed source for type.
6. **Is any source live or frequently updated?** Determines whether a version, edition or capture is needed rather than a bare link. *Default:* capture an extract with an access date for every live source at the moment of verification.

## 7. Step-by-step methodology

**Step 1. Build the claim-source register, and find the mismatches immediately.** Go through the document and list every claim that rests on external evidence, with the source it cites and the location of both. Then build two exception lists that the reference list alone will never show you. **Claims with no source:** statements of fact about the world, particularly numbers, presented with no citation at all. **Sources with no claim:** references in the list that no sentence in the document actually uses, which are either decoration, residue from an earlier draft, or (with AI-assisted documents) an invented list assembled to look like a bibliography. Both are findings before verification begins. *Correct result:* a register with one row per claim-source pair, plus the two exception lists.

**Step 2. Rank by load, and set the coverage rule.** Mark each row as load-bearing or contextual. **Load-bearing:** the claim fails if the source fails, and the claim is doing work (a headline, a market size, a business case input, a recommendation's premise, anything in an executive summary). **Contextual:** the claim would survive as background if the source were removed. Load-bearing rows are verified exhaustively. Contextual rows are sampled, unless AI involvement is confirmed or suspected, in which case coverage is total, because a fabricated citation is as likely in a contextual position as anywhere and it fails the whole document's credibility when found. Record the coverage rule in the log so that a reader knows what was and was not checked.

**Step 3. Run the five verification questions against each source. In this order, because a failure at one question makes the next ones moot.**

**(a) Does it exist?** Locate the actual item, not a mention of it. A search result showing the title, a reference to it in another document, or a matching author-and-topic pair is not the item. If the item cannot be located, it is not yet verified, and the possibilities are that the citation is malformed, the item is inaccessible, or the item does not exist. These are different findings and the log distinguishes them.

**(b) Is it what the citation says it is?** Check every element against the item itself: authors, title, year, publisher or journal, volume, edition, version, and type. Type matters more than the rest, because it is the element most often wrong and least often checked: a working paper cited as a peer-reviewed article, a conference abstract cited as a study, a commentary cited as research, a press release cited as a report, an executive summary cited as the report it summarises.

**(c) Does it say what the citing text says it says?** Open it and read the relevant passage, not the abstract. This is where verification actually earns its cost, and it fails more often than existence does. The recurring patterns: the source says the thing but hedged, and the citing text dropped the hedge; the source says it about a subgroup, a market or a period, and the citing text says it generally; the source describes it as a possibility and the citing text reports it as an established finding; the source says the opposite and was cited from its title.

**(d) Is the quoted figure the published figure?** Check the number, and then check the four things around it that change what it means: the **base** (of whom), the **unit** (percentage, percentage points, index, currency, real or nominal), the **period** (when, and over what interval), and the **definition** (what was counted as the thing). A figure that matches on the digits and differs on any of these is a misquotation, not a rounding.

**(e) Is this source the origin of the claim, or a repeater?** Look at what the source itself cites for the statement. If it cites someone else, this is not the origin, and step 4 applies. Repeaters are legitimate to read and illegitimate to cite as though they were origins.

*Correct result:* every checked row carries an answer to all five questions, or a recorded stopping point with the reason.

**Step 4. Collapse every citation chain to its origin.** Where a source is a repeater, follow its own reference, and repeat. Record the chain: A cites B cites C cites D. Stop when you reach a source that reports its own measurement, or a dead end.

Three things this reveals, and each is a distinct finding. **The one-origin problem:** five sources appear to corroborate a figure, and all five trace to the same single study, so the apparent corroboration is an artefact of repetition and the claim rests on one piece of evidence, not five. This is the single most valuable output of the whole skill, because a figure's frequency is routinely mistaken for its strength. **The dead end:** the chain terminates at a source that states the figure with no attribution at all, which means the number has no established origin and cannot be cited as fact regardless of how widely it appears. **The origin that does not say it:** the chain terminates at a real study whose actual finding is narrower, older, or about something adjacent.

**Then check the figure for drift along the chain.** Numbers change as they are repeated, and the mechanisms are consistent: rounding at each hop; a unit slipping (percentage to percentage points, an index read as a percentage); a base widening (a figure about enterprise buyers becoming a figure about buyers); a scope creeping (one market becoming a region); a date detaching (a 2019 measurement cited without its year and read as current); and a qualifier dropping (an estimate, a projection or a modelled figure becoming a measurement). Record the value at each hop. Where the value at the document differs from the value at origin, the document's figure is wrong even where every intermediate step looked reasonable. *Correct result:* for each traced claim, a chain with a named origin or a named dead end, the figure at each hop, and a drift note where the values differ.

**Step 5. Identify the plausible-but-unreal citation, and treat it as fabrication.** This is the signature failure of AI-assisted work and it has a recognisable signature. A real author who works in the field, paired with a real journal that publishes in the field, a believable year, and a title that reads exactly like a paper that would exist. Frequently a DOI that is malformed or resolves to something else, page numbers that fall outside the volume, or a volume and year that do not correspond. Sometimes an amalgam: two real papers merged into one reference. Sometimes a real paper with an invented finding attached.

The handling is not negotiable. **A citation that cannot be located is not cited.** It does not become "personal communication", it does not get softened to "research suggests", and it is not retained on the basis that the underlying claim is probably true. Where the claim matters, find a real source for it. Where no real source can be found, the claim is unsourced and is either removed or restated as the author's own assertion (K4 §2.4). And where one fabricated citation is found in a document, **the base rate for the rest of that document has changed**: coverage goes to total, including sources already sampled and passed.

**Step 6. Resolve secondary citation properly.** A source cited for something it in turn cites from elsewhere is a secondary citation, and there are exactly two honest treatments. **Go to the origin, read it, and cite it**, which is always preferred. Or, where the origin is genuinely unobtainable, **cite the secondary explicitly**: the origin, then the fact that it is cited in the secondary source you actually read, so the reader knows which document you have seen. What is not permitted is citing the origin as though you had read it, which is the near-universal default and is the mechanism by which errors propagate for decades: the origin's actual wording, base and caveats never get checked by anyone in the chain.

**Step 7. Handle live sources, versions and dates.** Anything that can change under the citation (a web page, an online statistical table, a register, a regularly revised dataset, a standard, a guideline) needs three things recorded: the **version or edition** where one exists, the **access date**, and a **captured extract** of the specific passage or figure relied on. Then check whether it has changed since it was cited. Revisions are common and quiet: a statistical series gets rebased or revised, a guideline is superseded, a page is updated with no visible history. Where the current version no longer supports the claim, the finding is that the claim was true of a version and needs its version stated, or is no longer true. Where a source has been withdrawn or retracted, the claim comes out, and any downstream claim resting on it comes out with it.

**Step 8. Screen source type and quality, separately from verification.** A source can pass all five questions and still be unsuitable. Three screens.

**Is it research, or is it something wearing research's clothes?** Press releases about studies, vendor and consultancy reports produced to support a commercial position, sponsored surveys with no published method, infographics, and trade-press summaries of any of these. None of these is disqualified automatically; a vendor's data on its own market can be the best available. All of them require the type to be named in the citation and the interest to be disclosed. **The operative test is whether a method is available**: a "survey of 2,000 professionals" with no sampling description, no fieldwork dates, no instrument and no base breakdown is an assertion with a number attached, and it is cited as such or not at all.

**Is the outlet what it appears to be?** Low-quality and predatory publishing produces material that carries the surface features of peer review without the substance. Recognisable signals, none conclusive alone: acceptance times inconsistent with any real review process; fees disclosed only after acceptance; a scope so broad it covers unrelated fields; a title closely imitating an established publication; an editorial board that cannot be corroborated or that lists people who are unaware of it; metrics claimed from bodies that do not issue them; and no discoverable, describable peer-review process. Where these appear, the item is treated as unreviewed material and is either dropped or cited as a preprint-equivalent with that stated.

**Is it grey literature, and does that matter here?** Government and agency reports, NGO research, working papers, theses and internal studies are often the best or only evidence, and much of it is more rigorous than the published alternative. The requirement is not to exclude it but to cite it as what it is, with its issuing body, its date and its status, so a reader can weight it themselves.

**Step 9. Assign a status to every source and a disposition to every claim.** Statuses: **verified** (all five questions answered, item seen); **verified with correction** (the item supports the claim once the citation's details are fixed); **partially verified** (existence and identity confirmed, content not checkable, usually a paywall or a language barrier); **unverified** (could not be located or accessed, with the reason); **misquoted** (the source exists and the figure or wording in the document does not match it); **misattributed** (the claim belongs to a different source, usually the one the cited source was quoting); **not the origin** (chain traced, origin named); **withdrawn or superseded**; and **not found, treated as fabricated**.

Then the disposition, which is the part that changes the document: correct the citation; requalify the claim to what the source actually supports; attribute the claim to its origin; label it unverified in the document itself; or remove it, together with anything downstream that rests on it. **The rule that governs all of them: an unverifiable source is labelled unverified, never cited as verified** (K4 §7). *Correct result:* a log where every row has a status and every failed row has a disposition with an owner, and a short list of the claims the document must change before it goes anywhere.

## 8. Analytical framework

The verification ladder, applied per claim-source pair. Each rung is only meaningful if the one below it holds:

    Exists → Is what it claims to be → Says what is claimed
        → Figure matches on value, base, unit, period and definition
            → Is the origin, not a repeater

And, running underneath it, the chain collapse:

    Document → Repeater → Repeater → Origin (or dead end)
                                       ↑
                          The only place the claim is actually evidenced.
                          Everything to the left is transmission.

**Applying it.** Work up the ladder and stop at the first failure, recording where you stopped: there is no value in checking whether a nonexistent paper says what is claimed. The two frames intersect at the top rung, and that intersection is where the most consequential findings sit, because a source can pass every rung below and still be a repeater whose origin says something different.

**The corroboration test.** Before treating multiple sources as agreement, collapse each to its origin and count the origins, not the sources. Three independent measurements agreeing is corroboration. Three documents repeating one measurement is one measurement, and describing it as widely reported is true and irrelevant. This distinction is worth more than any other single output of the skill.

## 9. Output format

**A. Verification summary.** Sources in the document; sources verified; the coverage rule applied and why; counts by status; the number of claims requiring change; and the headline judgement on whether the document's external evidence base can be relied on.

**B. The verification log.** One row per claim-source pair.

| Ref | Claim (as written) | Location | Source as cited | Item located? | Type as cited / actual | Supports claim? | Figure in document / at source | Origin or repeater | Chain to origin | Status | Disposition | Checked by / date |
|---|---|---|---|---|---|---|---|---|---|---|---|---|

**C. Citation chains.** For every traced figure, the chain written out with the value at each hop, the origin named, and the drift mechanism where the value changed.

**D. Claims requiring action.** The list the document owner works from, ordered by severity.

| Claim | Location | Problem | Required action | Consequential edits |
|---|---|---|---|---|

Required actions are specific: "remove", "requalify to: [text the source supports]", "reattribute to [origin]", "label unverified", "correct citation details". "Review" is not an action.

**E. Unverified register.** Every source that could not be verified, with the reason (paywall, withdrawn, language, print-only, not located) and what would be needed to verify it. This is the document's honest disclosure and it appears in the document itself, not only in the log.

**F. Source quality notes.** Where type was misrepresented, where a commercial interest exists, where an outlet's review status could not be established.

**When the evidence is thin.** The log is not padded and statuses are not upgraded to make a document look better sourced. A source that could not be checked is `unverified`, with the reason, every time. Where verification could not be completed within the available time, the log states which sources were checked and which were not, and the document carries that disclosure. **A partial verification honestly scoped is a real output. A full verification claimed but not performed converts an unchecked document into an apparently checked one, which is materially worse than doing nothing** (K4 §1).

## 10. Quality checks

Run on the verification itself. K4 §8 runs anyway.

1. Was every source located as an item and seen, rather than confirmed from a search result, an abstract or a mention elsewhere?
2. Was the relevant passage read, rather than the abstract or the title?
3. For every figure, were the base, unit, period and definition checked as well as the digits?
4. Was every load-bearing claim traced to an origin, and are the origins counted rather than the sources?
5. Are claims with no source and sources with no claim both listed?
6. Where AI involvement is confirmed or suspected, was coverage total rather than sampled?
7. If any citation was found to be fabricated, was coverage extended to the whole document including sources already passed?
8. Is every unverifiable source recorded as unverified with its reason, rather than passed because the claim is probably true?
9. Does every failed row carry a specific disposition and an owner, rather than a note to review?
10. Were consequential edits identified, so that removing a source also removes what rested on it?
11. Was every live source captured with a version or access date and an extract of the passage relied on?
12. Was source type checked against what the citation claims, particularly for press releases, working papers, summaries and vendor material?
13. Does the log record who checked each row and when, so the verification is itself auditable?
14. Is the coverage rule stated in the output, so a reader knows what was not checked?

## 11. Common failure modes

| Failure | How to recognise it | How to prevent it |
|---|---|---|
| **Verifying existence and stopping** | Every reference resolves, and three of them do not say what the document says they say | The ladder in Section 8. Existence is the first rung, not the check |
| **Abstract verification** | The passage cited is a general claim in the abstract; the paper's actual finding is narrower | Read the relevant passage in the body, always |
| **Digit matching** | The number is right and it was about a different base, period or population | Check base, unit, period and definition alongside the value |
| **Corroboration by repetition** | "Widely reported", five sources, one origin | Step 4. Collapse every chain and count origins |
| **The silent secondary** | The origin is cited; the author read only the document quoting it | Cite what you read. "Cited in" where the origin is unobtainable |
| **Live-source drift** | The link works, the page has changed, the figure has been revised | Version, access date and captured extract at the time of verification |
| **Type laundering** | A press release, working paper or executive summary cited as a study | Check type against the item. Name the type in the citation |
| **Method-free statistics** | "A survey of 2,000 professionals found..." with no sampling, dates, base or instrument | Require an available method. Otherwise it is an assertion with a number |
| **AI: the plausible citation** | Right author, right journal, believable year, paper does not exist. Often a malformed DOI or impossible page range | Locate the item. Not found is fabrication, and it triggers total coverage of the document (K4 §2.4) |
| **AI: the real paper, invented finding** | The citation resolves perfectly and the paper says nothing like the claim | Read the passage. Never accept a resolving reference as verification of content |
| **AI: the amalgamated reference** | Two real papers merged; the author list belongs to one and the title to another | Check every element against the located item, not against plausibility |
| **AI: softening on challenge** | A citation that cannot be found becomes "some research suggests" rather than being removed | The claim is unsourced. Remove it or attribute it to the author |
| **Verification theatre** | A tick against every source, completed faster than the sources could have been opened | Record who checked each row, when, and what passage was read |
| **The retained inaccessible source** | A paywalled source marked verified because it is reputable and probably right | Reputation is not verification. Status is `partially verified` or `unverified` |

## 12. AI guardrails

Universal prohibitions are inherited from K4 and not repeated. **K4 §2.4 governs this skill in full and is its central rule.**

1. **Never record a source as verified without having located and opened the item.** A search result, a citation in another document, an abstract, a title match, or a recollection that the paper exists are not verification, individually or together.
2. **Never reconstruct a citation from memory to complete a reference.** Missing volume numbers, page ranges, DOIs, publishers and years are looked up in the item or left marked incomplete. A reference completed from plausibility is a fabricated reference even where the paper is real.
3. **Never retain a claim whose citation could not be located by converting it into vaguer language.** "Research suggests", "studies show" and "it is widely reported" applied to an unverifiable source are the same fabrication with the evidence removed.
4. **Never treat wide repetition as corroboration.** Count origins, not sources, and state the count.
5. **Never upgrade a status because the claim is probably true, the source is reputable, or the deadline is close.** Probability is not verification and reputation is not access.
6. **Never quote a figure without checking the base, unit, period and definition attached to it at source**, and never carry a figure between documents without re-verifying it at source (K2 §7).
7. **Never cite a source for a claim it attributes to someone else** without either going to the origin or making the secondary citation explicit.
8. **Never assess a source's quality from its title, its formatting or the confidence of its language.** Type, method availability and review status are established from the item and its issuing body.
9. **Where one fabricated citation is found, never continue with a sampled coverage rule.** Verify everything, including what has already passed, and say in the output why coverage changed.
10. **Never present a partial verification as a complete one.** The coverage rule and the unverified register appear in the output and in the document.

## 13. Best-practice principles

1. **Verify while you collect, not at the end.** A source checked at the moment of reading costs a minute. The same source checked after the document is written costs the paragraph, the argument the paragraph was carrying, and the credibility of everything near it.
2. **The question is never "does this source exist".** It is "does this source support this sentence". Existence is the cheapest and least informative rung on the ladder.
3. **A figure's frequency is not its strength.** The most repeated statistics in most fields are the least verified, precisely because everyone assumes somebody else checked.
4. **Always go one hop further than feels necessary.** The hop you skip is the one where the base changed.
5. **Capture the passage, not the link.** A link records where you looked. An extract with a date records what it said, and it is the only thing that survives the source being revised.
6. **Type is the most consequential citation element and the least checked.** A working paper, a press release and a peer-reviewed article carry entirely different weight and are indistinguishable once formatted into a reference list.
7. **Ask what method is available before treating anything as a study.** Where no method can be obtained, you are citing an assertion, and the citation should say so.
8. **An unverifiable source labelled unverified is a respectable component of a document.** An unverifiable source presented as verified is a defect, and the difference costs one sentence.
9. **Removing a source means removing what rested on it.** Verification's consequential edits are where documents quietly break: the claim goes, the paragraph that built on it stays, and the argument now has a hole nobody can see.
10. **Treat a fabricated citation as information about the document, not about that reference.** It tells you the process that produced this document does not verify, which means nothing in it has been checked.
11. **Record who checked, and when.** Verification that cannot itself be audited provides the same assurance as no verification, and it is what a challenge will ask for first.
12. **Say what could not be checked.** The unverified register is the most credible part of a verification report, because it is the part that had nothing to gain.

## 14. Worked example

*Fictional scenario, used for illustration only. The organisation, sources and figures below are invented for the purpose of demonstrating method.*

**INPUT.** A B2B consultancy has drafted a market-entry report for a client considering launching a logistics software product. The report's opening claim: "The mid-market logistics software segment is growing at 24% a year and will reach 4.1 billion by 2028." Four sources are cited for it. The report contains 31 sources in total, and the analyst confirms that AI assistance was used to draft two background sections and to format the reference list.

**PROCESS.**

*Steps 1 and 2.* The register produces 38 claim-source pairs. Two exception lists appear immediately: six claims stating figures with no citation at all, and four references in the list that no sentence in the document uses. Because AI assistance is confirmed, the coverage rule is set to total rather than sampled. The growth and market-size claim is load-bearing: it opens the report and the business case is built on it.

*Step 3.* Of the four sources cited for the headline claim, three are located and one is not. The one that is not has the signature of the plausible fabrication: a real analyst house, a report title that reads exactly as their report titles read, a believable year, and no locatable item under that title or any near variant. Of the three located, one is a trade-press article, one is a vendor's own market report, and one is a summary page on an industry association's site. None is a study. All three are repeaters.

*Step 4, and the finding that decides the review.* Collapsing the chains: the trade-press article cites the vendor report. The association summary cites the trade-press article. The vendor report attributes the figures to "internal analysis" with no method, no base and no definition of the segment. So four apparently independent sources collapse to one origin, that origin is a commercially interested party, and it publishes no method. The apparent corroboration is entirely an artefact of repetition.

*The figure drift.* The vendor report states 24% growth for "the logistics software market" globally over 2021 to 2023, and a 4.1 billion figure that is explicitly described as a modelled projection under a stated adoption assumption. By the time it reaches the draft, three things have happened: the segment has narrowed to "mid-market" (a restriction the source never made), the period has detached so a historical 2021 to 2023 rate reads as forward-looking, and "modelled projection" has become a flat forecast. Every digit matches. The claim is wrong on base, period and status.

*Step 5, and the judgement call.* The analyst proposes retaining the claim with softer wording, on the grounds that the direction is surely right and the number is used everywhere in the sector. This is refused: the source that would support it does not exist, and the three that do exist support a different claim about a different segment over a past period. Softening the language while keeping the figure is the same fabrication with the evidence hidden (K4 §2.4). The narrower claim the evidence does support is available and is offered: one commercially interested source, unmethoded, reports 24% growth in the broader global market over 2021 to 2023, and there is currently no verifiable evidence on the mid-market segment specifically. That is a weaker sentence and a true one, and it is also a finding about the market: nobody has measured this segment, which is itself relevant to the client's decision.

*Steps 6 to 9, the rest of the document.* Two further fabricated citations are found in the AI-drafted background sections, one of them a real author paired with a real journal and a title that does not exist. Four sources are secondary citations presented as origins; two origins are obtained and cited directly, two are unobtainable and are rewritten as "cited in". One regulatory guideline has been superseded since it was cited, and the claim resting on it is now wrong. Three sources are paywalled and are recorded as partially verified: identity confirmed, content not checked. Nine of 31 sources require citation corrections, mostly on type: two working papers cited as journal articles and one executive summary cited as the full report.

**OUTPUT.** A verification log covering all 38 claim-source pairs; three citations recorded as not found and treated as fabricated; a chain diagram showing four sources collapsing to one commercially interested origin with the figure drifting on base, period and status at three hops; a claims-requiring-action list with eleven entries, of which two are removals (with their consequential edits named, since the business case's growth assumption came from the removed figure), five are requalifications and four are citation corrections; an unverified register naming the three paywalled sources and what would be needed to check them. Overall judgement: the external evidence base cannot be relied on as it stands, and the headline claim must be removed rather than softened. `RESEARCHER DECISION REQUIRED` per K5 §3.1: whether the client is told the segment is unmeasured, or the report is delayed to commission a sizing exercise, is a commercial judgement for the engagement owner, but it must be made knowing the figure has no origin.

## 15. Advanced usage

**Verifying at scale.** Where a document or repository carries hundreds of references, verify in two passes. A structural pass first, checking the resolvable elements of every reference (identity, type, existence, version), which is fast and catches fabrication and type errors. Then a content pass on the load-bearing subset, checking what each source actually says, which is slow and cannot be delegated to matching. Report the two coverages separately, because a document where every reference resolves and no content was checked is a document where nothing has been verified.

**Building the verification into the evidence base.** Where a team maintains a repository or a rolling literature base, carry the status field with each source permanently rather than re-verifying per document, and add a review date for live sources. The register then becomes reusable, and the cost falls to verifying new material only. The discipline that makes this work is that a status without a checker and a date is not a status.

**Contested and politicised figures.** Where a figure is disputed rather than merely repeated, the chain collapse usually reveals two origins with different definitions rather than a disagreement about a measurement. Report both origins with their definitions side by side rather than adjudicating: the disagreement is nearly always definitional and naming it is more useful than picking a side.

**When the origin is unobtainable.** Some chains end at a study that is out of print, held privately, or was never published. The honest treatment is to state the chain, name the terminal source, state that it could not be obtained, and label everything downstream as resting on an unexamined origin. This is a legitimate and informative output, and it is materially different from the claim having been checked.

## 16. Skill chain

**Recommended previous skills:**
- **10.01 Literature Review and Desk Research.** Hands over the reference set and the search record, which this skill audits. The search record reveals selection effects a reference list conceals.
- **10.03 Evidence Synthesis Across Sources** and **10.05 Competitive and Market Intelligence.** Hand over multi-source claims whose apparent corroboration this skill tests by collapsing chains to origins.
- **12.03 Research Report Compilation.** Hands over the compiled document with its source inventory, so verification runs against the version that will actually be delivered.

**Recommended next skills:**
- **13.01 Research Quality Review.** Takes the verification log as an input and assesses the methodological soundness of the primary work, which verification does not touch.
- **12.03 Research Report Compilation.** Takes the claims-requiring-action list and applies the removals, requalifications and consequential edits properly, rather than patching the document.
- **13.06 AI Research Governance.** Takes the fabrication findings and the coverage rule into the project's disclosure and audit trail.

**Runs well alongside:**
- **13.03 AI Output Verification**, which uses this skill as its citation component and covers the numbers, quotes, gaps and calibration that citations do not.
- **K2 §4.3**, which defines the reference format this skill checks against, and **K4 §7**, which requires an unverifiable source to be marked unverified rather than cited as verified.

---
A Yazi Supplied Skill and resource.
