---
name: quote-and-evidence-extraction
description: >
  Selects, verifies and attributes the verbatim evidence that goes into a
  research report, so quotes match the claims they sit under and none of them
  are invented. Use for "pull quotes for this theme", "find a good quote for
  this slide", "which verbatims support this finding", "select evidence for the
  report", "check these quotes", "anonymise these quotes", "we need a quote for
  the executive summary".
category: 07 Qualitative Analysis
ref: 07.04
tier: 1
inherits: [K2, K3, K4, K5]
---

# Quote and Evidence Extraction

## 1. One-line description
Selects the verbatim evidence for a report on representativeness rather than eloquence, verifies every quote word for word against source, matches each one to the specific claim it is asked to carry, and refuses to manufacture verbatim support where none exists.

## 2. What this skill is used for

**The research problem it solves.** This is where fabricated quotes enter research reports. Not through malice, but through a chain of small, reasonable-looking steps: a theme needs evidence, the memory of a striking phrase surfaces, a near-match is tidied to fit the slide, two participants' words are combined because neither said the whole thing, an identifier is dropped because it cluttered the layout, and a composite sentence nobody said arrives in an executive summary carrying more persuasive weight than any statistic in the deck. Every downstream reader treats it as fact, because a quote is indistinguishable from a real one once it is in quotation marks. The second failure is quieter and more common: quotes selected for how well they read. The articulate participant gets quoted four times, the inarticulate one who said something more important gets quoted never, and the report's evidence base narrows to the handful of people who spoke in complete sentences. This skill is the control on both.

**Where it sits.** The bridge between analysis and reporting. It takes coded, themed or case-level material and hands verified, attributed, claim-matched evidence to writing, compilation and presentation.

**Typical use cases.**
- Pulling supporting verbatims for each theme in a qualitative report.
- Selecting the one quote that will open a section or a presentation.
- Populating an evidence appendix that has to withstand a client's own analyst checking it.
- Anonymising quotes for a report that will be shared beyond the commissioning team.
- Auditing an existing deck's quotes back to source when a finding has been challenged.
- Choosing a verbatim that represents a dissenting minority without inflating it.
- Assembling verbatim evidence for a regulatory, public or published output where every word will be scrutinised.

**Who uses it.** Qualitative researchers assembling deliverables; report writers and designers who receive quotes and have no way to check them; insight leads signing off client-facing work; anyone who has been asked for "a good quote for this slide" and knows that is the wrong question.

## 3. When to use it

- A report, deck, summary or paper needs verbatim evidence, and someone will act on what it says.
- You hold coded or themed qualitative material and need to select from it rather than remember from it.
- The evidence must be attributable: every quote traceable to a participant who exists in the corpus.
- The output will be read by people who did not run the fieldwork and cannot check the quotes themselves.
- The sample is small or specialist, and anonymisation needs doing properly rather than by removing names.
- A minority or dissenting view needs to appear in the report without being given the weight of a majority.
- Quotes already in a document need auditing against source, either routinely or because a claim has been challenged.
- The deliverable is external-facing, published, or attributed to a named researcher, where a fabricated or misattributed quote is a professional and legal exposure.

## 4. When NOT to use it

- **The themes have not been built yet.** Selecting quotes before the analysis is done means the analysis will be built backwards from the quotes, which is the mechanism by which a report ends up saying whatever the most quotable participants said. Run **07.01 Thematic Analysis** or **07.03 Interview and Transcript Analysis** first. The boundary in the other direction: 07.01 and 07.03 hand their theme records and annotated passages here, and neither of them selects final report evidence, because selection is a distinct discipline with its own biases.
- **The deliverable is a distribution, not a narrative.** Where the output is coded frequencies from short survey text, verbatims are illustrations of code definitions rather than evidence for claims. That belongs in **07.02 Open-Ended Response Coding**, which produces illustrative examples under a different standard.
- **You cannot access the source text.** Selecting a quote from a summary, a previous deck, or a colleague's notes cannot be verified, and an unverifiable quote is functionally a fabricated one per K2 §4.2. If the source is not available, say so and do not use the quote.
- **The claim is not supported by the material.** Where a stakeholder needs a quote for a conclusion the data does not carry, the honest answer is that no quote exists because no evidence exists, per K4 §9. Supplying the closest available verbatim to prop up an unsupported claim is worse than supplying nothing, because it launders the claim.
- **The sample is too small or too specialist for the participant to be anonymisable.** Where a job title plus a market plus a sector identifies one person, a quote is a named quote whatever identifier is printed next to it. Stop and take it to **13.05 Research Ethics and Consent Design**. This is not solved by removing the company name.
- **Consent does not cover verbatim reproduction.** Participants sometimes consent to research participation and not to being quoted, or consent to internal use and not to publication. Check before selecting, not before publishing.
- **The purpose is decoration.** Where quotes are being requested to fill white space, break up text or add colour to a quantitative deck, they are not evidence and should not be presented as though they were. Either they support a claim or they come out. See **07.06** on the specific failure of using qualitative material as decoration for quantitative findings.
- **The material is a translation and the wording carries the claim.** A translated quote is the translator's sentence. It can evidence what a participant meant; it cannot evidence how they said it. Where the finding turns on their exact language, the claim cannot be made from translated text, per K5 §2.2.

## 5. Required inputs

**Required.** Without these the skill cannot run. If absent, ask. If no answer is available and the work must proceed, state the assumption at the point where it bites, per K5 §5.

- **The source text**, in full and searchable: transcripts, coded open ends, or the passages with their surrounding context. Not an extract list. Verification is impossible against an extract list, and context is what determines whether a quote means what it appears to mean.
- **The specific claims the evidence must support**, written out. Not theme names, which are labels, but the sentences that will appear in the report. Selecting evidence for a theme name rather than a claim is how a quote ends up under a statement it does not support.
- **Participant identifiers and the characteristics that will be printed alongside them**, so attribution can be checked and the re-identification screen can be run.
- **The coded or themed material**, so the candidate pool is drawn from the analysis rather than from recall.

**Optional, and what each one adds.**

- **Audio or timecodes**: settles disputes about wording, resolves passages the transcript garbled, and is the only way to verify a quote whose transcript is uncertain.
- **Prevalence data for each theme**: makes it possible to state how typical a quote is, which is the difference between illustration and misrepresentation.
- **The consent wording participants signed**: determines what may be reproduced, for which audience, and whether attribution characteristics are permitted.
- **The report layout or slide template**: reveals the length constraint before selection rather than after, so quotes are chosen at usable length instead of being trimmed past the permitted edits later.
- **The audience for the deliverable**: internal working document, client deck, public report and regulatory submission carry different verification and anonymisation standards.
- **The list of quotes already used elsewhere in the project**: prevents the same three participants carrying the whole report across multiple documents.

## 6. Questions to ask before starting

1. **What exactly is each quote being asked to evidence?** *Default if unanswered:* extract the claim sentence from the draft report and confirm it before selecting. Never select against a theme name alone.
2. **Who will read this, and can it leave the organisation?** Determines the verification standard, the anonymisation standard and whether attribution characteristics are permitted. *Default:* assume the document will be forwarded, and apply the external standard.
3. **How many quotes does the deliverable need, and how long can each be?** Determines whether selection is possible at all, because a claim that needs 40 words of context cannot be evidenced in a 15-word pull-out. *Default:* select at natural length, and flag any claim that cannot be evidenced within the constraint rather than trimming to fit.
4. **What identifiers may be shown?** Participant number, segment, role, market, tenure: each addition raises re-identification risk multiplicatively in a small sample. *Default:* the minimum set that makes the quote interpretable, and run the screen at step 10.
5. **Is a dissenting or minority view being represented?** Determines the framing language, which is what stops one quote reading like a widespread position. *Default:* every minority quote carries its base in the same visual unit as the quote.
6. **Are quotes reused from earlier documents in this project?** *Default:* check, and prefer fresh evidence, so the report does not narrow to a handful of participants across every deliverable.
7. **Is audio available for verification?** *Default:* verify against the transcript, and mark any quote whose transcript is uncertain rather than presenting it as settled.

## 7. Step-by-step methodology

**The non-negotiable.** Every quote that leaves this process has been read in its source, in context, and matched character by character. There is no version of this step that is skipped for time, and there is no deliverable so small that it does not apply. A quote is the most persuasive object in a research report and the least checkable, which is exactly why the check happens here and not downstream, where it will not happen at all.

**1. Start from the claim, not from the quote.** Write out every sentence in the report that will carry verbatim evidence, as it will appear. For each, note what the quote has to do: demonstrate that the experience exists, show what it felt like, show the reasoning behind a behaviour, show the boundary of the theme, or show the counter-case. *Correct result:* a claim list, each entry with an evidential job. Selecting a quote and then writing a claim around it inverts the logic of the report, and it is how a memorable phrase becomes a finding.

**2. Build the candidate pool from the coded material, not from memory.** For each claim, pull every passage coded to the relevant theme or code, across all participants, without filtering. Include the flat ones. *Correct result:* a pool typically five to ten times larger than the number of quotes needed. **If the pool is small, that is a finding about the evidence, and it is recorded now rather than discovered at step 11.** Working from the passages you remember guarantees selection on memorability, because memorability is what made you remember them.

**3. Screen the pool for typicality before screening for quality.** For each candidate, ask whether it represents how participants generally expressed this theme, or whether it is an unusually strong, unusually articulate or unusually extreme version of it. Read the pool as a set: what does the ordinary expression of this theme sound like? A quote that is more vivid than 90 percent of the pool is not a better example of the theme, it is a worse one, because it misrepresents the intensity of the material. *Correct result:* each candidate marked typical, strong, or outlier, with the modal expression of the theme identified. **The standing bias in this task runs toward the articulate respondent**, who is over-represented in every qualitative report ever written, and who is systematically different from the rest of the sample: more educated, more confident, more used to being asked their opinion, and more likely to give the answer that sounds like an answer. Correcting for it requires deliberately selecting some evidence that is less well expressed.

**4. Apply the coverage constraint.** Count how many distinct participants appear across the whole evidence set, not per theme. Set a budget before selecting: no participant carries more than a small share of the report's quotes, and the total number of distinct voices is stated. **A twenty-participant study whose report quotes four people twelve times has reported four interviews.** Where one participant keeps winning selection, that is the articulate-respondent bias operating, and the correct response is to find the same point in someone else's words even if it reads less well. *Correct result:* a distinct-participant count and a per-participant quote count, both reportable, with any concentration justified explicitly.

**5. Run the quote-to-claim match, one pair at a time.** For each selected quote and its claim, test three things. Does the quote actually say what the claim says, or does it say something adjacent that a reader will assume matches? Does it require the surrounding context to mean what it appears to mean, and is that context available to the reader? Would the quote support a different claim equally well, and if so, is the claim doing work the evidence does not license? *Correct result:* every pair passes all three, or the quote is replaced or the claim is weakened. The commonest failure is the adjacent match: a claim about trust evidenced by a quote about frustration, which reads perfectly and evidences nothing. The reader supplies the connection and attributes it to the participant.

**6. Select the boundary and the counter-case deliberately.** For each substantial theme, include at least one quote that shows where the theme stops, or that cuts against it. This is not balance for its own sake. A theme evidenced only by its clearest instances has not been shown to have edges, and a reader cannot tell how far it extends. *Correct result:* every substantial theme carries a boundary or counter quote, or an explicit statement that none was found in the corpus, which is itself a finding per K4 §4.1.

**7. Handle minority and dissenting views with their weight attached.** A dissenting quote is often the most valuable evidence in a report and the easiest to misrepresent, because a single strong verbatim reads like a position rather than like one person's position. Three controls. First, **the base travels with the quote in the same visual unit**, not in a footnote and not on the previous slide: "3 of 24 participants". Second, the framing sentence states the prevalence before the quote is read, not after. Third, the quote is not selected for being the most forceful expression of the minority view, because that compounds the weighting problem. *Correct result:* a reader who saw only the slide would correctly estimate how many people held this view.

**8. Verify every quote against source, word for word.** Open the source, locate the passage, and compare character by character. Check the speaker label. Check that the passage is not spliced from two places in the transcript. Check that an ellipsis is not doing violence to the sense. Where audio exists, verify the quotes you most want to use, because the most quotable line is the one most likely to have been misheard by transcription or misremembered by the analyst. **Verify the quote you like best first.** *Correct result:* every quote carries a verification mark and the location it was verified against. **This step is required and is not a quality-assurance nicety.** An unverified quote in a client-facing document is the single highest-consequence failure in qualitative reporting, per K5 §6.

**9. Apply only the permitted edits, and write the convention statement.** Editing is limited to the four permitted operations in **K4 §2.3**: filler removal where meaning is unchanged, marked elision, square-bracketed clarification of a referent, and flagged correction of an obvious transcription error. Nothing else. Not grammar. Not tense. Not word order. Not the removal of a hedge that weakens the point. **The output must carry a stated convention**, once, visibly, saying what has been done to the quotes: for example, that quotes are verbatim, that filler has been removed, that omissions are marked, and that square brackets indicate clarification added by the researcher. *Correct result:* a convention statement in the deliverable and an edit log showing what was changed in each quote. Where a quote needs more than the permitted edits to make the point, the quote does not make the point, and step 11 applies.

**10. Run the re-identification screen on every attribution.** Build the attribution string exactly as it will be printed, then ask: how many people in the population this sample was drawn from match this combination of characteristics? **Job title plus market is frequently one person. Job title plus sector plus company size is almost always one person.** In specialist B2B, expert and employee research, a participant number provides no protection at all, because the identifying information is in the descriptors. Assess the risk not just to the reader you intend but to the colleague who receives the deck onward, who may know exactly who the sole procurement director for that region is. Where risk is real, reduce the descriptor set, band the characteristics, or move the quote to a paraphrase attributed to the group. *Correct result:* every attribution passes the screen, or has been reduced, or has been escalated per K5 §2.4. Do not treat removing the participant's name as anonymisation.

**11. Take the honest position where no good quote exists.** Sometimes a real, well-evidenced theme has no quotable verbatim behind it: everyone expressed it in fragments, in half-sentences, or in answer to a probe. This is common and it is not a failure of the analysis. **The correct action is to report the theme with its prevalence and describe it in the analyst's voice, clearly marked as the analyst's summary rather than a participant's words.** Per K4 §9, offer the closest genuine quote with an honest note that it is partial, or a labelled analyst summary, or a prevalence statement with no verbatim at all. What is never permitted: composing a sentence that captures what participants meant, merging two participants' phrasing, or lifting a probe-shaped answer and presenting it as spontaneous. *Correct result:* a theme reported without verbatim support, with the reason stated, which a research director reads as rigour rather than as a gap.

**12. Assemble the evidence set with its register.** Produce the quotes in report order, each with claim, participant ID, characteristics as printed, verification status, edit log, typicality rating and prevalence context. *Correct result:* a register from which anyone can re-verify any quote in the report without asking you where it came from.

**13. Audit the assembled set as a whole.** Read every quote in the deliverable in sequence, ignoring the surrounding text, and ask four questions. Do these quotes sound like different people, or do they sound like one voice? Would a reader seeing only the quotes reach the same conclusions as the report? Is any participant over-represented? Is any theme carried entirely by its most extreme instance? *Correct result:* a set that reads as a sample rather than as a chorus. **If the quotes all sound alike, something has been rewritten**, and the edit log is the place to find it.

## 8. Analytical framework

The chain the output is built on:

    Claim → Candidate pool → Typicality screen → Coverage check → Claim match
        → Verification → Permitted edits → Attribution screen → Evidence register

**Applying it.** The order is load-bearing. Selection happens before verification because verification is expensive and should be spent on the final set; but nothing enters the deliverable that has not been through verification, so the two are never traded against each other. The attribution screen sits after editing because bracketed clarifications frequently add identifying detail that was not in the original.

**The four evidential jobs a quote can do.** Naming the job prevents the adjacent match at step 5.

| Job | What the quote must show | Common misuse |
|---|---|---|
| **Existence** | That this experience or view occurs at all | Treated as evidence of prevalence |
| **Texture** | What it is like, in the participant's own terms | Selected for vividness rather than typicality |
| **Mechanism** | The reasoning or sequence behind a behaviour | A post-hoc explanation reported as the cause |
| **Boundary** | Where the theme stops, or the counter-case | Omitted entirely, leaving a theme with no edges |

**The eloquence trap.** Three quotes can each be genuine, verified and correctly attributed, and together misrepresent the study, because all three came from the three most fluent participants. Fluency correlates with education, confidence and familiarity with being asked one's opinion, and none of those is what the study is measuring. The correction is procedural, not editorial: the coverage budget at step 4 and the typicality screen at step 3, applied before anything is chosen for how it reads.

**Illustration versus memorability.** A quote that illustrates a theme is one a reader could use to recognise the theme in new material. A memorable quote is one a reader will repeat. These overlap sometimes and are not the same property, and the memorable one travels further, which means it does more damage when it is atypical. Where a quote is selected for memorability, that is a legitimate communication decision, and it is made explicitly, with a typical quote placed alongside it.

## 9. Output format

**A. Evidence note (front matter)**
Source corpus and where it can be found; the verification standard applied and by whom; the quote convention statement, per step 9; the distinct-participant count and total quote count; the anonymisation approach and the re-identification screen result; statement that AI performed selection and what a human verified, per K4 §7.

**B. Evidence register**

| Ref | Claim it evidences | Quote (as printed) | Participant ID | Characteristics as printed | Source location | Verified (Y/N, against what) | Edits applied | Typicality | Theme prevalence | Re-identification risk |
|---|---|---|---|---|---|---|---|---|---|---|

**C. Coverage summary**

| Participant | Quotes used | Themes covered | Share of total quotes |
|---|---|---|---|

With a stated concentration threshold and a justification for any participant above it.

**D. Claims with no verbatim support**
Each theme or claim for which no adequate quote exists, the reason, and how it is reported instead (analyst summary, prevalence statement, or partial quote with a note).

**E. Quotes considered and rejected**
Retained rather than discarded, with the reason: atypical, unverifiable, re-identifying, adjacent to the claim, or duplicative of a participant already used. This is what makes the selection rule inspectable, per K4 §4.2.

**F. Review points**
Quotes where attribution risk required a judgement; quotes whose transcript is uncertain; any claim where the strongest available evidence is weaker than the draft text implies.

**When the evidence is thin.** No slot is filled to look complete. A theme with no quotable material appears in section D, not in section B with the nearest available approximation. A quote that could not be verified is not published with a caveat; it is not published. Where the coverage summary shows three participants carrying the report, that is stated in the front matter rather than smoothed, because it changes how the whole report should be read. Where the strongest available quote is weaker than the claim, the claim is softened rather than the quote strengthened.

## 10. Quality checks

Run before anything is presented. Sits on top of K4 §8.

1. Has every quote been located in the source and compared word for word, including the ones that seemed obviously fine?
2. Is any quote spliced from two locations in a transcript without a marked elision?
3. Is any quote a merge of two participants' words? (This is the check that matters most and takes the least time.)
4. Does every quote carry a participant identifier that exists in the corpus?
5. Does every attributed characteristic match that participant's actual record, rather than the segment the claim is about?
6. Does the deliverable carry the quote convention statement, visibly, once?
7. Are all edits within the four permitted operations, and is each logged?
8. Does every quote say what its claim says, rather than something adjacent that reads as though it matches?
9. Does every minority-view quote carry its base in the same visual unit as the quote?
10. How many distinct participants are quoted, and does any one of them exceed the concentration threshold?
11. Would a reader who saw only the quotes reach the report's conclusions?
12. Do the quotes sound like different people?
13. Has the re-identification screen been run on the attribution string as it will actually be printed?
14. Does consent cover verbatim reproduction for this audience?
15. Is any theme evidenced solely by its most extreme instance?
16. For every theme with no quote, is the absence reported rather than filled?

## 11. Common failure modes

| Failure | How to recognise it | How to prevent it |
|---|---|---|
| **The composite quote** (the most damaging failure in the library) | A quote captures the theme unusually completely, in balanced clauses, with nothing extraneous in it | Never merge speakers. Verify word for word. Suspect any quote that is too good |
| **The remembered quote** | The analyst recalls the phrasing and cannot immediately point to it in the source | Build the pool from coded material, not recall. Verification is not optional |
| **Quote polishing** | Quotes read fluently, in a consistent register, with no repetition or repair | Permitted edits only, per K4 §2.3. If they all sound alike, something was rewritten |
| **The articulate respondent monopoly** (the standing bias) | Four participants carry the report; the quotes are notably well-phrased | Coverage budget at step 4, typicality screen at step 3, both before selection |
| **The adjacent match** | The quote is genuine, the claim is reasonable, and the quote does not actually evidence it | Step 5, tested one pair at a time, out loud |
| **The unweighted minority** | One vivid quote reads as a widespread view; the base is on another page | Base in the same visual unit. Prevalence stated before the quote |
| **Attribution drift** | A quote is labelled with the segment the claim is about rather than the speaker's own | Check the speaker's record, not the claim's subject |
| **Silent re-identification** | Anonymised by participant number, identified by role plus market | Run the screen on the printed string. Count how many people match |
| **Ellipsis abuse** | A marked omission removes a qualifier and reverses the sense | Read the full passage. An ellipsis that changes the meaning is a rewrite |
| **Decorative quoting** | Quotes appear where the layout needed something, evidencing nothing specific | Every quote has a named claim and an evidential job, or it comes out |
| **The manufactured quote for a real theme** | A theme with no verbatim support acquires one late in the process | Step 11. Report the theme with prevalence and an analyst summary, clearly labelled |
| **Verification by plausibility** (the signature AI failure) | The quote is confirmed because it is consistent with the transcript, not because it was found in it | Locate the string. Consistency is not verification |
| **Reuse concentration across deliverables** | The same three quotes appear in the topline, the deck and the summary | Track quotes used across the project and prefer fresh evidence |

## 12. AI guardrails

Skill-specific only. Universal prohibitions are inherited from K4 and are not repeated here. Quote handling follows **K4 §2.3** in full and is the operative standard for this skill.

1. **Never produce a quote that has not been located in the source text in this session.** Not from memory, not from a previous document, not from a summary. If it cannot be found, it does not exist.
2. **Never merge, splice or compose.** Two participants' words never become one quote. Two locations in one transcript never become one quote without a marked elision that preserves the sense.
3. **Never present a paraphrase as verbatim**, and never soften this by calling it "a representative quote". A representative quote is a real one that is typical, not a written one that is representative.
4. **Never edit beyond the four permitted operations**, and never publish edited quotes without the convention statement in the deliverable.
5. **Never attribute a quote to a segment without checking the speaker belongs to it.**
6. **Never print an attribution string without running the re-identification screen on it**, particularly in specialist, B2B, expert and employee samples.
7. **Never let a quote stand under a claim it does not directly evidence.** Adjacent is not evidence.
8. **Never present a minority view without its base in the same visual unit as the quote.**
9. **Never fill an empty evidence slot.** Where no adequate quote exists, report the theme with prevalence and a labelled analyst summary, per K4 §1 and §9.
10. **Never confirm a quote by plausibility.** Consistency with the transcript is not verification against it.
11. **Never select on eloquence without correcting for it.** Where the strongest-reading candidates cluster in a few participants, deliberately select less fluent evidence from others.

## 13. Best-practice principles

1. **The quote you love most is the one to verify first.** Perfection is a symptom. Real speech is redundant, hedged and slightly off the point.
2. **Select for the claim, never for the slide.** A quote chosen because it fits the space will end up trimmed past the permitted edits by whoever lays out the page.
3. **A quote proves the experience exists. It never proves how common it is.** Prevalence comes from the coding, and the two are placed together so no reader has to guess.
4. **Fluency is a characteristic of the participant, not of the finding.** Reports that quote only fluent people have quietly reported a subsample.
5. **Read the passage around the quote before taking the quote.** Meaning lives in the ten seconds either side more often than in the sentence itself.
6. **The counter-quote strengthens the report.** Showing where a theme stops is what tells a reader it has edges, and a theme with no edges reads as an assertion.
7. **Anonymity is a property of the descriptor set, not of the identifier.** P07 protects nobody if the descriptors narrow to one person.
8. **Quotes travel further than any other part of a report.** They get lifted into emails, board packs and press releases, where the base, the caveat and the context do not follow. Choose as though the quote will be read alone, because it will be.
9. **A theme with no good quote is still a theme.** Saying so is a mark of a competent analysis, not a hole in it.
10. **Track who you have quoted across the whole project.** Concentration accumulates document by document, and nobody notices until an external reader does.
11. **When a stakeholder asks for a better quote, they are usually asking for a different claim.** Find out which, and address that.
12. **The evidence set is finished when someone else could re-verify every line of it without asking you a question.**

## 14. Worked example

*Fictional scenario, used for illustration only. All participants, quotes and findings below are invented for the purpose of demonstrating method.*

**INPUT.** A UX research team at a workplace software provider has run 18 usability and diary sessions with administrators of a scheduling tool, across three customer organisations. Thematic analysis produced five themes. The team needs verbatim evidence for a report going to the product leadership group and, in edited form, to the three participating customers.

**PROCESS.**

*Steps 1 to 2.* Claim list extracted from the draft: eleven sentences requiring evidence. For the lead claim, "administrators do not trust the conflict warnings and route around them", the candidate pool is pulled from the two relevant codes across all participants: 43 passages from 14 participants.

*Step 3, and the first judgement call.* The pool's modal expression of this theme is flat and procedural: participants describe checking the roster manually afterwards, in short factual sentences. One passage, from P11, is vivid and quotable and describes the warnings as something they had "stopped even seeing". It is markedly stronger than 90 percent of the pool. The tempting move is to lead with it. Assessment: it is an outlier in intensity, and using it alone would tell leadership that administrators are actively dismissive when the evidence shows something duller and more actionable, which is that they treat the warnings as one input among several and verify manually as routine. Decision: **P11 is used, and is placed second, after a typical passage from P04 that describes the manual check without dramatising it**, with the intensity difference noted in the evidence register. Recorded as a communication decision made explicitly, not as a selection on quality.

*Step 4.* Coverage check across the whole draft report: 11 claims, 19 quotes, drawn from 7 of 18 participants, with P11 appearing four times. Over the threshold. Two of P11's quotes are replaced with less fluent material from P02 and P16 saying substantially the same thing. Final set: 19 quotes from 12 participants, maximum 3 from any one.

*Step 5.* Match test catches an adjacent match. A claim about administrators lacking confidence in the system was evidenced by a quote about the system being slow. Genuine quote, reasonable claim, no evidential link: the reader supplies the connection. The quote is moved to the performance claim and a different passage, from P09, is found for the confidence claim.

*Step 8, and the second judgement call.* Verification against source finds that the P11 quote reads, in the transcript, "I'd stopped even seeing them, or I don't know, maybe I see them and don't read them". The draft deck carries only the first clause. The clause that was cut is not filler and does not qualify a hedge: it substantively changes the finding from "administrators ignore the warnings" to "administrators are not sure whether they process them". Cutting it would exceed the permitted edits. Decision: the fuller quote is used, at greater length, and the claim is softened to match. **The claim moved to fit the evidence, not the other way round.**

*Steps 9 to 11.* Convention statement written into the report: quotes are verbatim, filler removed, omissions marked, square brackets indicate researcher clarification. Re-identification screen run on attributions: the draft printed "Administrator, healthcare client, 400+ sites", which identifies one of the three participating organisations and, combined with the role, plausibly one person to anyone at that organisation, who will receive the report. Reduced to "Administrator, large multi-site client". One theme, about handover between shift administrators, has no adequate quote: everyone described it in fragments across several turns and no single passage carries it. Reported as a theme with prevalence (9 of 18) and a labelled analyst summary, with a note that no single verbatim captures it.

*Step 13.* Whole-set audit. The quotes now vary in register and length, and two are noticeably inarticulate, which is correct. A reader given only the quote set would reach the report's conclusions with one exception: they would over-weight system speed, which is fixed by removing one duplicative quote.

**OUTPUT.** Nineteen verified quotes from twelve participants, each matched to a written claim; an evidence register with source locations, edits and typicality ratings; a coverage summary; one theme reported with prevalence and no verbatim, with the reason stated; a rejected-quotes list with selection reasons; and one softened claim, recorded as having been changed to match the evidence.

## 15. Advanced usage

**Auditing an existing deck.** Where quotes are already in a document, run steps 8 to 10 as an audit rather than a selection: locate every quote in source, check attribution against the participant record, and check the printed attribution string for re-identification. Report unverifiable quotes as unverifiable rather than removing them silently, because their presence is a finding about the process that produced the deck. This pairs with **13.03 AI Output Verification**.

**Executive summary quotes.** The summary is the most-read and least-checked part of any report, per K2 §5, and a quote there travels furthest with the least context. Apply the strictest standard: typical rather than vivid, prevalence adjacent, and attribution reduced to the minimum that keeps it interpretable. A quote that needs the body of the report to be read correctly does not belong in the summary.

**Public and regulatory outputs.** Verification moves from sample to census, audio verification becomes the standard rather than the exception where it exists, and the anonymisation assessment is documented as a decision with a named owner. Consent wording is checked against the specific publication route, not against research participation in general.

**Small and specialist samples.** Where every participant is identifiable in principle, the choice is between reducing descriptors until the quote is uninterpretable and moving to paraphrase attributed to the group. Make it explicitly and early, because it changes what the report can show. In employee research the risk is not only external: the participant's own manager is often the reader.

**Cross-language reporting.** Present translated quotes with the original alongside where the audience can read it, and mark every translated quote as translated. Never make a claim about a participant's specific word choice from a translation. Where a phrase is untranslatable and carries the finding, describe it rather than substitute an English approximation and quote that.

**Building an evidence appendix that survives challenge.** Order by claim, not by theme, so a reader disputing a specific sentence in the report can find its evidence in one step. Include the rejected-quotes list. A challenger who can see what was not used, and why, is significantly less likely to suspect selection.

## 16. Skill chain

**Recommended previous skills:**
- **07.01 Thematic Analysis.** Hands over theme records with prevalence, counter-evidence and coded extracts carrying participant IDs.
- **07.03 Interview and Transcript Analysis.** Hands over positioned passages classified as volunteered, elicited, prompted or confirmed, so evidence strength is known before selection.
- **07.02 Open-Ended Response Coding**, where the evidence pool is coded short-form text.

**Recommended next skills:**
- **11.03 Research Report Writing** and **11.02 Executive Summary Development.** Receive claim-matched, verified evidence with its convention statement.
- **12.03 Research Report Compilation.** Receives the evidence register, so quotes keep their source references through assembly, per K2 §6.
- **11.05 Research Presentation Development**, where quotes will appear on slides and travel without their context.

**Runs well alongside:**
- **13.03 AI Output Verification**, which audits a finished deliverable against this skill's standard.
- **13.05 Research Ethics and Consent Design**, wherever the re-identification screen returns a real risk.
- **K5**, at the two review points this skill mandates: verification of every client-facing quote, and any attribution carrying re-identification risk.

---
A Yazi Supplied Skill and resource.
