---
name: research-report-qa
description: >
  The last gate before a research report leaves: a systematic seven-pass check
  of numbers, claims, evidence, language, consistency, completeness and
  presentation, producing a defect log with severity and required fixes and a
  clear pass or fail judgement. Use for "check this report before it goes out",
  "QA the deck", "final check on the report", "have we claimed significance
  anywhere we shouldn't", "do the numbers match across the document", "verify
  the quotes", "is the summary stronger than the body".
category: 12 Report Design and Compilation
ref: 12.06
tier: 1
inherits: [K2, K3, K4, K5]
---

# Research Report QA

## 1. One-line description
Runs the last structured check before a research report is delivered: a systematic pass over numbers, claims, evidence, language, internal consistency, completeness and presentation, producing a defect log with severity and required fixes and a pass or fail judgement on the document.

## 2. What this skill is used for

**The research problem it solves.** Most research reports are checked, and most of them are checked in the way that catches the fewest defects: someone reads the document for sense and typographic errors, in one pass, on the last afternoon, having already read it three times. That process is close to useless against the failures that actually matter, and its failures are systematic rather than random. A number is verified against the previous draft rather than against source, which propagates the error rather than finding it and makes the third draft more confidently wrong than the first. A base is missing on the one chart where it changes the reading. A quote reads well because it was tidied during a rewrite, and nobody opens the transcript. A difference is described as "significantly higher" when no test was run, which is a claim of reliability the report cannot support. A finding that says "appears to be associated with" in the analysis says "drives" in the executive summary, having gained certainty at each restatement. The same figure appears in three places and disagrees with itself in one of them. A recommendation everybody likes has no finding beneath it. And the limitation that qualifies the headline sits on page 60. Every one of these is invisible to a read-through and visible to a structured pass, and the difference between the two is the difference between a report that survives interrogation and one that does not. This skill is the structured pass, and it produces a judgement rather than a set of comments.

**Where it sits.** Reporting, at the end. After compilation, design and executive artefacts, before delivery. It is the last gate, and it is a gate rather than a review: its output includes a decision about whether the document may leave.

**Typical use cases.**
- A compiled report about to go to a client, where the author's own checks do not constitute independent review.
- A deck built by several people, where consistency across sections has never been checked by one person.
- A report whose claims will be interrogated line by line by a client, a regulator or a rival supplier.
- A document that will be published, quoted publicly, or submitted, where a defect becomes a matter of record.
- A tracking wave where comparability claims must be verified before a trend chart is published.
- A report inherited late, where the checking has to establish what is defensible rather than confirm that it is.

**Who uses it.** Research directors signing off deliverables; senior researchers checking a colleague's report; quality leads running a formal gate; consultants checking their own work with a structure that compensates for familiarity; anyone whose name will be on the document.

## 3. When to use it

- A report, deck or executive artefact is finished and about to be delivered.
- The document was compiled from several people's work and nobody has checked it as one thing.
- Claims will be challenged, quoted, extracted or acted on by people who will not see the evidence.
- Numbers appear in more than one place in the document, which is every report.
- The report makes comparative, trend or significance claims, which are the most frequently overstated and the easiest to check.
- Recommendations are included, and each has to trace to findings before the report carries them.
- An accessibility or presentation standard has been claimed and has not been tested.
- Time is short and the instinct is to skim, which is precisely when a defined pass structure earns its cost.

## 4. When NOT to use it

- **The question is whether the research itself is sound.** This skill checks the document against its own evidence: does the report say what its sources support, consistently, completely and defensibly. It does not assess whether the study was well designed, whether the sample was appropriate, whether the instrument measured what it claimed, or whether the analysis was the right analysis. That is **13.01 Research Quality Review**. The boundary runs both ways and matters: **this skill checks the document, 13.01 checks the research.** A report can pass here and fail there, which is the case of a perfectly assembled document reporting an unsound study, and it can fail here and pass there, which is a good study badly written up. Where QA reveals that a defect is methodological rather than documentary, it escalates to 13.01 rather than resolving it. In the other direction, where 13.01 finds the study sound, the document still needs this gate.
- **The report is still being written.** QA on a moving document wastes the pass and creates false confidence, because the version checked is not the version delivered. Freeze the draft first (step 1). The tell is a defect log with entries already superseded by the time it is delivered.
- **The document has not been compiled properly.** Where claims arrive without sources and quotes without identifiers, this is not a QA problem, it is a compilation defect, and the report goes back to **12.03 Research Report Compilation** or **12.04 Research Evidence Integration** rather than being checked in that state. QA can report that a document is unverifiable; it cannot make it verifiable.
- **The task is restructuring or rewriting.** If the report is wrongly structured, the answer is **12.01 Research Report Architecture**, not a QA pass that lists structural complaints as defects. QA identifies that the structure is a problem and hands it on. A QA pass that turns into a rewrite loses the original's decisions and produces a second version in circulation with no record of the change.
- **Only source and citation checking is needed.** Verifying that cited sources exist and say what they are claimed to say is **13.02 Source and Citation Verification**, which this skill invokes as a pass rather than replaces.
- **The concern is that the analysis or the report is biased.** Systematic checking for selection, framing and confirmation effects is **13.04 Bias Detection**. This skill's completeness pass catches omitted contradicting evidence, which is one symptom, not the diagnosis.
- **The output required is reassurance.** Where the request is for confirmation that a report is fine before a deadline, and there is no intention to act on findings, do not run the pass. A QA log that is not acted on is worse than none: it documents known defects in a delivered report, which is a materially worse position than not having checked.
- **The author is the only available checker and the stakes are high.** Self-QA using this structure is far better than a read-through and is normal practice under time pressure. It is not equivalent to independent review, because familiarity blinds a reader to their own draft, and where a document is high-stakes, state plainly that only self-QA was performed.

## 5. Required inputs

**Required.** Without these the skill cannot run. If absent, ask. If no answer is available and work must proceed, state the assumption at the point where it bites, per K5 §5.

- **The frozen final draft**, in the form it will be delivered, including the executive summary, appendix, charts and any separate executive artefact.
- **The source material behind every claim**: analysis tables with question references and bases, transcripts, the dataset where available, client-supplied files, secondary sources. **Checking against the previous draft is not checking**, and the absence of source material is a stopping condition rather than a constraint to work around.
- **The evidence map**, per K2 §5. Without it, the claim pass becomes a reconstruction exercise and takes four times as long. Where none exists, that is itself a defect and the report goes back to 12.03 or 12.04.
- **The agreed objectives**, which are the completeness test.
- **The questionnaire or discussion guide as fielded**, with routing, without which base descriptions cannot be verified.

**Optional, and what each one adds.**

- **The conventions register** from compilation: turns the consistency pass from a judgement into a check against a stated standard, which is faster and produces defensible defects rather than opinions.
- **The raw dataset**: lets a disputed figure be recomputed at source rather than compared between derived files, which is the only way to settle a conflict rather than adjudicate it.
- **The previous wave's report and questionnaire**: the only way to verify a comparability claim rather than accept it.
- **The statistical test outputs**: let every significance claim be verified against an actual test and threshold rather than against the analyst's memory of one.
- **A named second checker for the numbers pass**: catches the transcription errors that the person who made them cannot see, which is a well-established limitation of self-checking rather than a comment on anyone's care.
- **The accessibility specification** from 12.02, where a standard has been claimed: gives the presentation pass a standard to test against rather than an impression to form.

## 6. Questions to ask before starting

1. **Is this the final draft, and is it frozen?** Determines whether the pass is worth running at all. *Default if unanswered:* ask explicitly and record the version and date checked in the log, so a later version cannot inherit this pass's judgement.
2. **What does this document claim, and who will interrogate it?** Determines the depth of each pass and where to concentrate. A report going to a regulator, a public audience or a rival supplier is checked to a different depth from an internal readout. *Default:* check to the standard of a hostile reader, since it is the only assumption that never leaves you exposed.
3. **Was a conventions register agreed, and does one exist?** Determines whether consistency is checkable against a standard or must be inferred. *Default:* infer the conventions from the document, record what they appear to be, and treat departures as defects.
4. **Has any figure been transcribed by hand anywhere?** Determines where transcription errors will be, which is where numbers were retyped rather than exported. *Default:* assume every number in a chart title, a callout and the executive summary was retyped, and recompute all of them.
5. **Is this independent QA or self-QA?** Determines what can be claimed about the check, and it is disclosed in the log either way. *Default:* state which it was; do not present self-QA as independent review.
6. **What is the release decision, and who makes it?** Determines who receives the pass or fail judgement and who has authority to release with open defects. *Default:* the named researcher signing the deliverable, per K5 §2.8.

## 7. Step-by-step methodology

**What QA is here.** Not proofreading, and not a read-through with comments. A defined sequence of passes, each looking for one class of defect across the whole document, producing a defect log and a judgement. **The pass structure is the method**: reading a report once looking for everything reliably finds fluency problems and misses evidence problems, because attention cannot hold seven checking frames at once. Each pass below covers the whole document, including the executive summary, chart titles, callouts, headings, footnotes, appendix and any separate executive artefact. Material outside the body is where defects concentrate, because it is written last and read least.

**Step 1. Set up and freeze.** Fix the version and record it. Assemble the source material and confirm the evidence map exists; if it does not, stop and return the report, since checking a document with no traceable sources produces a list of things you could not verify rather than a QA result. Build the checklist from the document itself: list every number, every quote, every citation, every recommendation, every comparative or trend claim, and every significance claim, with its location. That inventory is what makes the passes complete rather than impressionistic, and building it takes a fraction of the time it saves. *Correct result:* a frozen version, an item inventory with locations, and confirmation that sources exist for everything in it.

**Step 2. The numbers pass.** Verify every number against its original source, not against the previous draft and not against another location in the same document. **Checking a figure against the last version is how an error becomes established**: each draft confirms the one before it, and by version four the wrong number has been checked three times. Open the analysis file, find the figure, confirm the value, the base description, the base size, the filter, the weighting state and the question reference. Recompute anything transcribed by hand, and treat every number in a chart title, a callout, a headline and the executive summary as transcribed, because it usually was. Then check every base and denominator specifically: does the base description match the questionnaire routing, is the base the base of that specific claim rather than of the study, and is the base shown wherever a proportion appears, per K4 §7. Check every percentage that should sum, every difference that is stated as a difference, and every index or derived figure for its own calculation. Where two locations in the document disagree, that is a consistency defect for step 5, and the number itself is still verified against source here. *Correct result:* every number confirmed against source with its base, and every discrepancy logged with its source location.

**Step 3. The evidence pass.** Take every claim in the document and confirm it has the evidence it asserts. **Quotes**: check every one word for word against transcript, confirm the participant identifier, and confirm the participant belongs to the segment the quote is attributed to, which is a frequent and damaging error. Confirm that edits fall within those permitted by K4 §2.3 and that the convention is stated once in the document. A quote that cannot be located in a transcript is a critical defect, not a query. **Citations**: verify each source exists, is correctly referenced, and says what the report says it says, per **13.02**; a citation to a real source for a claim it does not make is the harder failure and the one a spot check misses. **Recommendations**: confirm each one traces to named findings, per K2 §5, by walking the map backwards from recommendation to finding to evidence. A recommendation whose row cannot be populated backwards is an orphan and is a critical defect, whoever agrees with it. **Claims generally**: confirm each has a map row, that the row's sources resolve to real files, and that the claim's level on the K2 chain matches how it is written, so an interpretation is not presented as an observation, per K2 §3.3. **Provenance labels**: confirm client-supplied, previous-wave, third-party and unverified material is labelled at every appearance rather than only the first, per K2 §6. *Correct result:* every quote, citation, recommendation and major claim verified or logged, with no claim resting on a source that cannot be opened.

**Step 4. The language pass.** Three explicit sub-passes over the finished text, each run separately and each covering headings, chart titles, callouts, bullet fragments and the executive summary, because those are where the language defects concentrate and where a body-text reading does not go.

**The significance-language check, per K4 §3.1.** Search the document for "significant", "significantly", "notably", "markedly", "clearly", "substantially higher", "a real difference" and equivalents. For each, confirm a test was run, and that the test, the result and the threshold are reported. Where no test was run, the language changes to observed comparison with both bases shown, or the claim is removed. Also check the reverse defect: a tested and significant difference reported without its test, which understates what the study established.

**The causal-language check, per K4 §3.2.** Search for "drives", "leads to", "causes", "results in", "impact of", "improving X will", "because", "due to" and equivalents, including the causal readings smuggled into headlines and chart titles, which is where they most often survive. For each, ask whether the design licenses a causal claim: a randomised experiment, a valid quasi-experimental design with its assumptions stated, or a longitudinal design with temporal order established. Where it does not, rewrite to association. Where the report uses "driver" as a conventional term for a modelling output, confirm the document states once that these are statistical associations whose causal direction this design did not establish.

**The confidence-calibration check, per K3.** For every significant conclusion, confirm the language matches the level the analysis assigned: high confidence stated declaratively with no hedge, moderate confidence carrying its alternative explanation or specific limitation in the same passage, hypotheses labelled as hypotheses with their validation named. Check for the two symmetrical defects: a strong finding hedged into ambiguity, and a moderate finding written declaratively. Then check for uniform hedging across the whole document, where every conclusion carries the same qualifier and the qualifiers therefore carry no information, per K3 §7. Also run the overgeneralisation check, per K4 §3.3: findings extended beyond the sample, the period or the market, and a company's own customers described as consumers. *Correct result:* three sub-passes logged separately, with each flagged instance either evidenced or rewritten.

**Step 5. The consistency pass.** Internal consistency is checked mechanically rather than by reading. **The same figure everywhere**: take each number from the step 1 inventory and check every location it appears, including the summary, the chart, the chart title, the body text, the appendix and the executive artefact. A figure in three places that matches in two is the classic defect and it is invisible to a linear read, because the three locations are never read together. **Summary versus body**: compare every executive summary claim to the body claim it summarises, and compare both to the analysis. **Claims routinely strengthen on their way into the summary**, losing a qualifier at each step, and this is the single most productive check in the whole pass. Look specifically for: a hedge removed, a base dropped, a segment claim generalised to the whole sample, an association become a cause, and a moderate reading become a statement. **Caveat travel**: confirm every material limitation sits with the finding it qualifies, in body text, rather than only in a method section or appendix, per K4 §4.3. A caveat that appears only in the appendix has been dropped, whatever the appendix says. **Terminology and conventions**: one term per concept throughout, consistent base descriptions, consistent rounding, consistent segment names, consistent confidence vocabulary. *Correct result:* every repeated figure reconciled, every summary claim checked against its body and its analysis, and every material caveat located with its finding.

**Step 6. The completeness pass.** Check for what is missing, which no amount of reading the document will surface. **Objectives**: every agreed objective is answered in the report or listed explicitly as unanswered. **Contradicting evidence**: evidence that cuts against the conclusions is present in the document, per K4 §4.1. This is checked against the analysis outputs, not against the report, since a report cannot show you what it omitted. **Required disclosures**: run K4 §7 as a list against the document. Small bases flagged, weighting scheme and effective base stated, non-probability samples declared with no margin of error attached, untested differences described as untested, unusual fieldwork periods noted, findings resting on a single item identified, exclusions stated with counts, unverified sources marked, and AI involvement disclosed with what a human verified, per K5 §7. **What we could not establish**: present, with content, per K3 §5.2. **Review points**: consolidated where the deliverable warrants it, and every recommendation carrying a sign-off marker, per K5 §3.1 and §3.3. *Correct result:* an objective-by-objective coverage confirmation, a disclosure checklist worked line by line, and a statement of what the report does not answer.

**Step 7. The presentation and accessibility pass.** Check the document as an object. Every chart carries its base, question reference and source, legibly and in a consistent position. Every page headline is a claim that is true of its page and safe quoted alone. No meaning is carried by colour alone; charts survive greyscale. Measured contrast meets the applicable thresholds; no caveat or base is set below the type floor. Heading levels are real and in order, tables have marked header rows, meaningful images have alternative text. Page numbering, cross-references, table of contents and appendix references resolve. Where an accessibility standard has been claimed, it has been tested rather than asserted, and where it has not been tested, the claim is removed. Design standards belong to **12.02**; this pass checks compliance with them rather than setting them. *Correct result:* a presentation defect list, and either a tested accessibility statement or no accessibility claim.

**Step 8. Log, classify and judge.** Record every defect with location, description, severity and the required fix, in the format at Section 9. Severity is assigned by consequence rather than by effort. **Critical**: a wrong number; an unverifiable or fabricated quote; a recommendation with no traceable finding; a significance or causal claim the design does not support; a missing base where the base changes the reading; a claim in the summary stronger than the analysis supports; a material limitation absent from the document; a required disclosure omitted. **Major**: the same figure inconsistent between locations; a caveat present only in the appendix; a provenance label missing after first mention; an objective neither answered nor listed as unanswered; a headline claiming more than its page; an untested accessibility claim. **Minor**: terminology inconsistency; formatting and rounding inconsistency; a broken cross-reference; a chart convention departure with no effect on reading. Then give the judgement. **Fail while any critical defect is open**, without exception and regardless of deadline: a critical defect means the document asserts something its evidence does not support, and delivering it knowingly is a worse position than delivering late. **Fail on a pattern of major defects** even where each is individually tolerable, because a document with fifteen inconsistencies has a systemic problem that spot fixes will not address. **Pass with conditions** where only minor defects remain and the fixes are specified. Record whether the pass was independent or self-QA, and name who holds the release decision, per K5 §2.8. *Correct result:* a defect log, a severity distribution, a stated judgement, and a named person who receives it.

## 8. Analytical framework

The seven passes, each over the whole document:

    Numbers → Evidence → Language → Consistency → Completeness
        → Presentation → Judgement

**Applying it.** Run them in order and run them separately. The order is not arbitrary: numbers must be right before consistency between locations means anything, and evidence must be verified before language calibration can be assessed, because you cannot judge whether "shows" is too strong until you know what the source says. Running them together produces the read-through this skill exists to replace.

**The two governing checking rules:**

    Check against source, never against the previous draft.
    Check the summary against the analysis, never against the chapter.

Both address the same mechanism. A claim degrades in small steps, and each step is checked against the step before it, which is why a defect gets more confident with every version rather than less.

**Severity by consequence:**

| Severity | Definition | Effect on judgement |
|---|---|---|
| Critical | The document asserts something its evidence does not support, or omits something a reader needs to act correctly | Fail while open, regardless of deadline |
| Major | The document is internally inconsistent, incomplete or misleadingly emphasised, without asserting a falsehood | Fail on a pattern; fix before release |
| Minor | The document is inconsistent in ways that do not change what a reader concludes | Pass with conditions |

The classification is by consequence to the reader, not by how hard the fix is. A wrong number that takes ten seconds to correct is critical; a terminology inconsistency running through forty pages is minor.

## 9. Output format

**A. Header**
Document and version checked, date, who checked it, whether the pass was independent or self-QA, and which passes were run.

**B. Defect log**

| ID | Pass | Location | Defect | Severity | Required fix | Source checked against | Status |
|---|---|---|---|---|---|---|---|

Every entry names a specific location and a specific fix. "Tighten the language in chapter 3" is not a defect entry. "p.14 headline states 'drives'; design is cross-sectional; rewrite to 'is associated with', per K4 §3.2" is.

**C. Verification record**

| Item type | Total in document | Verified against source | Defects found |
|---|---|---|---|

Numbers, quotes, citations, recommendations, significance claims, causal claims, comparative and trend claims.

**D. Language pass results**
Instances found and their disposition, listed separately for the significance check, the causal check, the confidence-calibration check and the overgeneralisation check.

**E. Disclosure checklist**
K4 §7 worked line by line, with present, not applicable, or defect against each.

**F. Judgement**
Pass, pass with conditions, or fail. The reason. The named person who receives it and holds the release decision. Where the judgement is fail and release is nonetheless required, that is a decision for a named human and is recorded as such, per K5 §3.1, with the open critical defects listed in the record.

**When the evidence is thin.** A QA pass that could not verify things says so rather than passing them. Where source material was unavailable, the item is logged as unverified, not as verified, and the verification record shows the shortfall. **An unverified claim is not a passed claim.** Where a substantial proportion of the document could not be checked, the correct judgement is that the report is unverifiable in its current state, which is a compilation defect that returns to 12.03 or 12.04, and the log says so rather than reporting a low defect count achieved by not looking.

## 10. Quality checks

Checks on the QA pass itself, run before the judgement is issued.

1. Was every number checked against its original source rather than against a previous draft or another location in the document?
2. Was every quote checked word for word against a transcript, including its participant identifier and the segment it is attributed to?
3. Were the executive summary, chart titles, headings, callouts and appendix included in every pass, rather than only the body?
4. Was the separate executive artefact, where one exists, checked to the same standard as the report?
5. Were the significance, causal and confidence checks run as separate explicit passes rather than folded into a general read?
6. Was every repeated figure checked in every location it appears?
7. Was every summary claim compared to the analysis it came from, and not only to the chapter that paraphrased it?
8. Was completeness checked against the objectives and the analysis outputs rather than against the report?
9. Was every recommendation traced backwards to named findings?
10. Was the K4 §7 disclosure list worked line by line rather than assessed by impression?
11. Does every defect entry name a specific location and a specific required fix?
12. Is severity assigned by consequence to the reader rather than by how hard the fix is?
13. Are items that could not be verified logged as unverified rather than absorbed into the pass?
14. Does the log state whether this was independent QA or self-QA?
15. Is the judgement stated explicitly, and does it hold where a critical defect remains open?

## 11. Common failure modes

| Failure | How to recognise it | How to prevent it |
|---|---|---|
| **Checking against the previous draft** | Numbers confirmed quickly, with no source file open | Step 2. Each draft confirms the last, so by version four the error has been checked three times |
| **The single-pass read-through** | One person reading the document once for everything, finding typographic errors | Step 8's pass structure. Attention cannot hold seven checking frames at once |
| **Body-only checking** | The executive summary, chart titles and appendix were skimmed | Every pass covers the whole document. Defects concentrate in material written last and read least |
| **Spot-checking quotes** | Three of eleven quotes verified, chosen by convenience | Every quote, every time. A tidied quote reads better than an untidied one, so fluency selects against the ones that need checking |
| **The plausible citation accepted** | The source exists, the reference is correct, and nobody read the source | Step 3. A real source cited for a claim it does not make is the harder failure |
| **Summary drift missed** | The summary was checked against the chapter, which was checked against the analysis | Step 5. Compare summary to analysis directly. Each restatement drops a qualifier |
| **The caveat that only lives in the appendix** | The limitation exists in the document, but not with the finding | Step 5. K4 §4.3. Appendix presence is not travel |
| **Completeness assessed from the report** | Nobody checked whether contradicting evidence in the analysis reached the document | Step 6. A report cannot show you what it omitted; check against the analysis outputs |
| **Severity by effort** | A wrong number rated minor because it is a quick fix; a forty-page terminology issue rated critical | Section 8 table. Severity is consequence to the reader |
| **The soft judgement** | A log with open critical defects and a conclusion of "broadly fine, some tidying needed" | Step 8. Fail while any critical defect is open. The judgement is the deliverable |
| **QA on a moving document** | Defects superseded before the log is delivered; two versions in circulation | Step 1. Freeze first, record the version |
| **Unverified counted as verified** | A high pass rate achieved because the source files were unavailable | Section 9. Unverified is its own category and appears in the verification record |
| **The AI checker that confirms** | Every claim reported as consistent and traceable, with no source file ever opened | Guardrail 1 in Section 12. Verification is an act on source material, not an assessment of plausibility |
| **QA becoming a rewrite** | The checker fixes the prose, and the log becomes a changelog | Section 4. Log the defect and the required fix; the author fixes it. A rewritten report needs re-checking |

## 12. AI guardrails

Skill-specific only. Universal prohibitions are inherited from K4 and are not repeated here.

1. **Never report a claim as verified unless the source was opened and the value read.** Plausibility is not verification, internal coherence is not verification, and a claim that looks consistent with the rest of the document has not been checked. This is the central risk of automated QA, because a fluent document reads as a correct one.
2. **Never verify a figure against another location in the same document or against a previous draft.** Both are the error propagating, and both produce a confident pass on a wrong number.
3. **Never mark an item verified when the source was unavailable.** Unverified is a distinct status and appears in the verification record, per Section 9.
4. **Never soften a defect's severity because of a deadline, a client relationship or the effort required to fix it.** Severity is a statement about consequence to the reader.
5. **Never issue a pass while a critical defect is open**, and never convert a fail into "pass with conditions" by reclassifying the defect that caused it.
6. **Never rewrite the report during QA.** Log the defect and the required fix. A checker who rewrites has produced an unchecked document, and the author's decisions are lost with no record.
7. **Never accept a significance or causal claim because it is written confidently.** Confidence in the prose is the strongest predictor that the claim was never tested, and the two language passes exist precisely because fluency conceals these defects.
8. **Never treat the absence of contradicting evidence in a report as evidence that none exists.** Completeness is checked against the analysis outputs, never against the document.
9. **Never present self-QA as independent review.** State which was performed, since the two catch different defects and familiarity is a real and measurable limitation.
10. **Never generate a defect that has not been located in the document.** An invented or generic defect wastes the author's time and reduces the credibility of the log, which is the mechanism by which real defects stop being acted on.

## 13. Best-practice principles

1. **The pass structure is the method.** A single read looking for everything finds fluency problems and misses evidence problems, and the difference is not effort, it is attention.
2. **Check against source, never against the previous draft.** This is the principle that catches the largest number of consequential defects, and it is broken by almost every informal checking process.
3. **Check the summary against the analysis, never against the chapter.** Claims lose a qualifier at each restatement, and the summary is the most-read and least-checked part of any deliverable.
4. **Everything outside the body is where defects live.** Chart titles, callouts, headings, the executive summary and the appendix are written last, read least and quoted most.
5. **Fluency is a warning sign, not a reassurance.** The best-written sentence in a report is disproportionately likely to be the one that overclaims, because it was rewritten for effect.
6. **Every quote, every time.** Spot-checking quotes selects against the ones that were tidied, because tidied quotes read better.
7. **A recommendation with no traceable finding is deleted, not queried.** The fact that everyone agrees with it is what made it possible for it to arrive unsupported.
8. **Severity is consequence to the reader, not effort to fix.** A wrong number corrected in ten seconds is critical; a terminology inconsistency across forty pages is minor.
9. **The judgement is the deliverable.** A log of comments without a decision leaves the release call to whoever is most tired, and QA that does not gate is documentation of known defects.
10. **A defect log that is not acted on is worse than no QA.** It converts an unknown defect into a known one that was delivered anyway.
11. **Unverified is not passed.** A low defect count achieved by not looking is the most dangerous output this skill can produce.
12. **Self-QA with structure beats a read-through by someone else without one**, and neither is independent review with structure. Say which you did.
13. **The gate has to hold on the day it is inconvenient.** A QA process that passes under deadline pressure is not a gate, and every checker will face this at some point.

## 14. Worked example

*Fictional scenario, used for illustration only. The organisation, defects, figures and quotes below are invented for the purpose of demonstrating method.*

**INPUT.** A design consultancy has completed a 40-page usability and adoption report for a government digital service, drawn from 900 survey responses, 24 moderated usability sessions, and service analytics supplied by the department. It goes to the department's programme board in three days, and will be shared with a scrutiny committee. Two researchers wrote it. A third runs QA with the evidence map, the analysis workbooks, the transcripts and the questionnaire.

**PROCESS.**

*Step 1.* Version frozen and recorded. Item inventory built: 96 numbers, 11 quotes, 6 secondary citations, 5 recommendations, 14 comparative claims, 4 significance claims, and 9 charts.

*Step 2, numbers.* Ninety-six numbers checked against workbooks. Three defects. A completion-rate figure appears as 62% in the report and is 58% in the source; the 62% traces to an earlier draft of the workbook that was superseded, and had been checked twice against the previous version of the report, which is exactly how it survived. A chart on page 22 shows a subgroup percentage with the study base of 900 rather than the subgroup base of 71. And an appendix table sums to 103%, caused by a multi-response question presented as single-response.

*Step 3, evidence.* Eleven quotes checked against transcripts. Nine verify. One has been smoothed during a rewrite ("I couldn't work out where to go next" against the transcript's "I couldn't, um, work out, like, where you go next"), which is within the permitted edits of K4 §2.3 but the report does not state the convention, so the fix is to add the convention statement rather than to alter the quote. One cannot be found in any transcript: it appears in a researcher's session notes as a paraphrase and was promoted to a quotation during writing. Critical defect; the quote is removed and replaced with the closest verified verbatim. Five recommendations traced backwards. Four trace cleanly. The fifth, "introduce a save-and-return function", traces to a stakeholder workshop comment and to no finding in the study, although 24 usability sessions make it entirely plausible. Critical defect: orphan recommendation, per K2 §2.2.

*Step 4, language.* The significance pass finds four instances of "significantly". Two have tests behind them and the tests are not reported, which understates what was established; the fix is to add the test and threshold. Two have no test: "significantly more likely to abandon" on an untested comparison of 71 against 402, rewritten to an observed difference with both bases shown, per K4 §3.1. The causal pass finds seven causal constructions, five in chart titles and headings, which is where they concentrate. Four are supported by the usability sessions where the participant demonstrated the sequence and are retained with the design named. Three are cross-sectional survey associations written as "drives" and are rewritten. The confidence pass finds one defect and it is the most consequential in the report: the analysis records the link between form length and abandonment as moderate confidence with an unresolved alternative explanation (that longer forms are used for more complex cases), the body chapter states it with the alternative explanation, and the executive summary states it flatly with the alternative explanation absent. The qualifier was lost in exactly one restatement.

*Step 5, consistency.* The 58% figure appears in five places and had been corrected in three of them during an earlier round, which produced a document disagreeing with itself. Two material caveats live only in the appendix: the analytics extract's date range and the non-probability nature of the survey panel. Both move to body text next to their findings, per K4 §4.3. Terminology: "user", "citizen" and "applicant" are used interchangeably for the same population across three chapters, logged as minor.

*Step 6, completeness.* Four of five objectives are answered. The fifth, on accessibility of the service to assisted-digital users, was not covered by the fieldwork, and the report neither answers it nor says so. Major defect: it goes into "what we could not establish". Contradicting evidence checked against the analysis outputs: a segment where satisfaction rose is present in the workbook and absent from the report, and it materially qualifies the headline. Critical defect, per K4 §4.1. Disclosure list: AI-assisted coding of the open ends is disclosed, but the disclosure does not state that a human verified a sample, per K5 §7; added.

*Step 7, presentation.* Three charts distinguish two series by colour alone and fail in greyscale. Two base lines are set below the type floor. The document claims conformance to an accessibility standard that has not been tested; the claim is removed rather than asserted, per guardrail 7 of 12.02.

*Step 8, judgement.* Four critical defects, six major, nine minor. **Fail**, with the fixes specified and re-check required on the numbers and evidence passes after correction. The deadline is three days and does not alter the judgement. The log goes to the named research director, who holds the release decision.

**OUTPUT.** A defect log of 19 entries with location, severity, required fix and source checked against; a verification record showing 96 numbers, 11 quotes, 6 citations, 5 recommendations and 18 comparative or significance claims checked, with two items unverifiable; separate results for the four language sub-passes; a worked disclosure checklist; and a fail judgement with four open critical defects named, delivered to the person holding release authority.

## 15. Advanced usage

**QA under severe time pressure.** Where a full pass is impossible, do not sample randomly. Run the highest-yield subset in this order: every number in the executive summary and every chart title, every recommendation traced backwards, every quote, and the significance and causal language passes. That subset takes a fraction of the full time and catches the majority of critical defects, because critical defects concentrate in exactly those places. Then state plainly in the log which passes were not run, so nobody reads a partial check as a full one.

**Self-QA.** Familiarity is the enemy and the structure is the compensation. Three things measurably help: leave a gap between writing and checking, check the document in a different form from the one you wrote it in, and run the passes in reverse order of your confidence, starting with the section you are surest about, because that is where you will look least carefully. State in the log that it was self-QA.

**Building a standing gate.** Where QA runs on every deliverable, keep the defect log as a series rather than discarding it. Defect patterns are informative about process rather than about individuals: a recurring base-description defect means the conventions register is not being applied; recurring summary drift means summaries are being written from chapters rather than from analysis; recurring orphan recommendations mean recommendations are entering after the evidence map was built. Each of those is fixed upstream, and the log is the only way to see it.

**Checking a report you cannot verify.** Where sources are unavailable and the report must still be assessed, the honest output is a statement of what could and could not be checked, with the unverifiable proportion stated. Do not produce a defect log implying a verification that did not happen. This situation is a compilation defect and returns to **12.03** or **12.04**.

**When QA finds a methodological problem.** A defect that turns out to be about the study rather than the document (a sample that cannot support the claims, an instrument that did not measure what is being reported, an analysis that was the wrong analysis) is escalated to **13.01 Research Quality Review** rather than logged as a document defect, and the report does not leave in the meantime. Marked `RESEARCHER DECISION REQUIRED` per K5 §3.1, since the consequences reach beyond the document.

## 16. Skill chain

**Recommended previous skills:**
- **12.03 Research Report Compilation.** Hands over the compiled report, its evidence map, its conventions register and its discrepancy log, which are the inputs the passes run against.
- **12.04 Research Evidence Integration.** Hands over the source inventory, verification statuses and comparability record, which make the evidence pass a check rather than an investigation.
- **12.02 Research Report Design.** Hands over the design and accessibility specification, which gives the presentation pass a standard to test against.
- **12.05 Executive Research Reporting.** Hands over the executive artefact, which is checked to the same standard as the report and not to a lighter one.

**Recommended next skills:**
- **13.01 Research Quality Review.** Takes any defect that turns out to be methodological rather than documentary, and assesses the study rather than the document.
- **13.02 Source and Citation Verification.** Takes the citation pass where sources need appraisal rather than confirmation.
- **13.03 AI Output Verification.** Runs alongside where substantial parts of the document were AI-assisted, checking for the defects specific to that.
- **13.04 Bias Detection.** Takes the report where the completeness pass suggests a pattern of omission rather than an isolated gap.

**Runs well alongside:**
- **K4 §7 and §8**, which supply the disclosure list and the self-check this pass operationalises.
- **K3 §8**, which supplies the calibration self-check the confidence pass runs.
- **K5 §2.8**, which governs the sign-off and release decision this skill hands to a named human.

---
A Yazi Supplied Skill and resource.
