---
name: thematic-analysis
description: >
  Rigorous AI-assisted thematic analysis of rich qualitative material: depth
  interview transcripts, focus groups, long-form open ends, diaries and
  ethnographic notes. Use for "what are the themes", "code these transcripts",
  "build a code frame", "analyse our qual", "what came out of the interviews",
  "synthesise these focus groups", "find the patterns in this qualitative data".
category: 07 Qualitative Analysis
ref: 07.01
tier: 0
inherits: [K2, K3, K4, K5]
---

# Thematic Analysis

## 1. One-line description
Builds defensible themes from rich qualitative material through systematic coding, tests those themes against the whole dataset rather than the convenient parts of it, and reports prevalence, contradiction and strategic importance as three separate things.

## 2. What this skill is used for

**The research problem it solves.** Rich qualitative material is easy to summarise and hard to analyse. The default failure, by humans under deadline and by AI systems by construction, is to read a set of transcripts, notice what recurs, write it up in confident prose, and hang three quotes off it. That process cannot distinguish a pattern from an impression. It cannot say how many people held a view, whether anyone held the opposite, whether it was volunteered or extracted by the moderator, or whether the theme survives the transcripts not used to build it. It also quietly deletes the most valuable material in qualitative work: the participant who contradicts themselves, the minority who see what the majority cannot, and the ambivalence that never resolves. This skill replaces impression with method.

**Where it sits.** Analysis. It takes prepared qualitative material and hands structured, evidenced themes to insight development.

**Typical use cases.**
- Twenty to forty depth interview transcripts from a single study.
- Six to twelve focus group transcripts, analysed across groups.
- A few hundred substantial open ends, each a paragraph or more.
- Longitudinal diary entries analysed across participants and over time.
- Ethnographic or accompanied-shop field notes.
- A qualitative phase within a mixed-method study, prepared for integration.

**Who uses it.** Qualitative researchers and research directors who want an auditable analysis rather than a plausible one; generalists analysing qual outside their main discipline; insight teams inheriting fieldwork somebody else ran; academics defending a coding process to an examiner.

## 3. When to use it

- You hold rich, discursive material where meaning is carried in how people explain themselves, not just in what they mention.
- The dataset spans multiple participants and you need to say something about the set, not about one person.
- The research question is about experience, reasoning, motivation, decision-making, barriers, meaning or language, rather than about how many.
- You need themes that will survive a stakeholder asking "how do you know that" and "who said it".
- Prevalence matters to the client, but so does whether a small number of people said something consequential.
- The material contains disagreement, and the disagreement is part of the answer.
- Findings will be integrated with quantitative data later and need to be stated with matching discipline.
- The analysis will be reused across a repository, so the code frame and definitions need to outlive the project.

## 4. When NOT to use it

- **The material is thin.** Short answers, single clauses, one-line open ends and rating justifications do not contain enough context to code for meaning. Coding them thematically produces a topic list dressed as an analysis. Use **07.02 Open-Ended Response Coding**, which is built for high-volume short text and reports prevalence honestly.
- **The volume is too large for reading.** Above roughly 5,000 items, no genuine familiarisation pass is possible and any claim to have engaged with the whole corpus is false. Use **06.06 Large-Scale Text Analytics** to structure the corpus, then use this skill on a purposively sampled subset if depth is still needed. Note the sampling rule in the output.
- **The unit of interest is one person.** A single transcript analysed for a single participant's account, journey or logic is **07.03 Interview and Transcript Analysis**. Running thematic analysis on one transcript produces themes with a base of one, which is not a theme, it is a summary.
- **The question is really a counting question.** If the deliverable is "how many customers mention delivery", the honest answer comes from a quantitative instrument or from coded frequency, not from thematic analysis. Thematic analysis will produce a number and the number will be wrong, because qualitative samples are not designed to estimate proportions.
- **The sample was recruited to represent, and will be reported as representative.** Non-probability qualitative samples support depth, not incidence. If the deliverable requires population estimates, this analysis cannot supply them and should not be presented next to them without an explicit statement, per K4 §3.3.
- **The material is a translation you cannot check, or comes from a cultural context you cannot read.** Idiom, deference, indirectness and humour are where thematic analysis is most easily wrong. Per K5 §2.2, a human with the relevant context must review before themes are fixed. Where no such reviewer is available, code descriptively, do not interpret, and say why.
- **Consent does not cover the analysis or the reporting.** Where participants agreed to one use and the analysis is for another, or where a small subgroup would be identifiable in the reporting, stop. This is **13.05 Research Ethics and Consent Design**, not an analysis problem.
- **You are being asked to confirm a conclusion that already exists.** Where a stakeholder has supplied the themes and wants evidence attached, thematic analysis is being used as a search for confirming quotes. Say so, per K4 §4.2, and offer either a genuine analysis or an explicit and labelled evidence-assessment of the stated hypothesis.
- **The transcripts are unreliable.** Automated transcription with high error rates, missing speaker labels, or recordings where large sections are inaudible cannot be coded for meaning. Fix the transcripts, or state which sections are excluded and how much of the corpus that removes.

## 5. Required inputs

**Required.** Without these the skill cannot run. If absent, ask. If no answer is available and the work must proceed, state the assumption at the point where it bites, per K5 §5.

- **The full text of the material**, complete rather than excerpted. Excerpts pre-select the evidence and make theme testing impossible. If only excerpts exist, say what the analysis can and cannot establish from them.
- **Participant identifiers** attached to every passage, stable across the whole corpus. Without IDs the analysis cannot report prevalence or attribute quotes, and per K2 §4.2 an unattributable quote is functionally a fabricated one. If IDs are missing, stop and ask for them.
- **The research objectives**, or the questions the study exists to answer. Thematic analysis without objectives produces a description of the transcripts rather than an answer.
- **The discussion guide or question set** used in fieldwork. Required in order to distinguish prompted from unprompted material, which changes what prevalence means.

**Optional, and what each one adds.**

- **Participant characteristics** (segment, market, tenure, behaviour, recruitment criteria): makes it possible to say whether a theme is general or belongs to a subgroup, and to check that a quote attributed to a segment came from a member of it.
- **Moderator or fieldwork debrief notes**: surface what the room felt like, which questions landed badly, and where a participant's words and manner diverged. This is context AI cannot recover from a transcript.
- **A prior code frame** from an earlier wave or a related study: enables deductive coding, comparability across waves, and repository reuse. Also introduces a risk of forcing old codes onto new material, so it is applied as a starting frame, not a fixed one.
- **Recordings or timecodes**: permit verification of a disputed quote against source.
- **Quantitative results from the same study**: allow convergence and divergence to be checked during theme testing rather than discovered at integration.
- **Client or stakeholder hypotheses**: useful as things to test, and dangerous as things to confirm. Record them before analysis so that confirmation is visible if it happens.

## 6. Questions to ask before starting

1. **What decision does this research inform, and who makes it?** Determines which themes are important as distinct from which are frequent, and therefore how the output is ordered. *Default if unanswered:* order by evidential strength, and flag that strategic ordering requires the researcher.
2. **Inductive, deductive or hybrid?** Determines whether the code frame is built from the material, imposed from the objectives, or both. *Default:* hybrid, with a deductive spine from the objectives and inductive freedom below it, and both layers marked in the frame.
3. **Is the unit of analysis the participant, the passage, or the group?** Determines how prevalence is counted and what a "theme raised by 9" means. *Default:* the participant, counted once per theme regardless of how often they returned to it.
4. **Were all participants asked about everything?** Routing, time pressure and guide changes mean some topics were put to some people only. Determines the denominator. *Default:* report prevalence against those who were asked, and state the denominator explicitly.
5. **Is a prior code frame in play, and is it binding?** Determines whether comparability to a previous wave outranks fidelity to this material. *Default:* apply the prior frame as a starting point, log every new code, and report where the old frame did not fit.
6. **What is the expected use of quotes?** Determines how much verification effort goes where. Anything client-facing gets every quote checked against source, per K5 §6. *Default:* full verification for all quotes leaving the analysis.
7. **Are there subgroups that must be reportable separately?** Determines whether the analysis holds subgroup structure through coding rather than reconstructing it afterwards. *Default:* code subgroup-blind, then test themes by subgroup, so that subgroup expectations do not shape the coding.

## 7. Step-by-step methodology

**Which tradition this follows, and why.** The default is **reflexive thematic analysis** in the Braun and Clarke sense, with a codebook discipline borrowed from framework and template analysis. Reflexive TA is the default because it insists themes are *constructed* by an analyst making decisions, not "found" in data waiting to emerge. That insistence is what AI-assisted analysis needs, because a system that believes themes emerge will present its own pattern-completion as a property of the data. The codebook discipline is added because AI applies codes across a whole corpus, and undefined codes produce unstable output at scale. The hybrid is deliberate and it is contested: Braun and Clarke argue codebooks and coding-reliability measures sit awkwardly with reflexive TA's premise that coding is interpretative. The position here is that when the coder is a machine, an auditable frame is the price of letting it code at all, and the interpretative work relocates to theme construction and the human review points, where it belongs.

Use a different tradition when: objectives are fixed, cross-case comparison is the point, and the client needs a matrix (use **framework analysis**, charting at step 8); the goal is theory-building and the design permits theoretical sampling and constant comparison (use **grounded theory**, noting that commercial timelines rarely permit it, so "grounded-theory-influenced coding" usually means open and axial coding without theoretical sampling, which should be stated rather than implied); the interest is one person's lived experience of a phenomenon (use IPA, and see **07.03**).

**1. Fix the analytical frame before reading.** Write down, before the first transcript, the research objectives, the coding approach, the unit of analysis, the prevalence denominator, and any stakeholder hypotheses in play. *Correct result:* a short pre-analysis note a colleague could use to hold you to your own method later. Recording hypotheses in advance is what makes confirmation bias visible if it occurs.

**2. Familiarise with the whole corpus.** Read everything through once without coding. Note first impressions, recurring language, moments of hesitation or contradiction, and anything surprising, in a memo held separately from the analysis. *Correct result:* a memo of impressions explicitly labelled as impressions. These are hypotheses to test, not findings. Where the corpus is too large to read in full, the analysis is out of scope here (see Section 4).

**3. Generate initial codes across the whole corpus.** Work through the material segment by segment and attach short, descriptive codes at the level of meaning, not topic. Code generously: at this stage over-coding is cheaper than under-coding. Code for what is said, how it is said, what is assumed, and what is conspicuously absent where you would expect it. Retain the exact text span and the participant ID for every code instance. *Correct result:* typically 60 to 150 initial codes for a 20 to 30 transcript study, each pointing to real text spans. A frame of 12 codes at this stage means the material has been summarised, not coded.

**4. Build the code frame, and apply hygiene to it.** Collapse duplicates, split codes doing two jobs, and organise into a shallow hierarchy of parent and child codes. Every code gets: a name, a one-sentence definition, an inclusion rule, an exclusion rule, and at least one boundary example showing a passage that nearly qualifies and does not. Decide explicitly which parts of the frame are mutually exclusive (usually structural or classificatory codes, such as decision stage) and which permit multiple assignment (usually meaning codes, where a passage legitimately carries two ideas). *Correct result:* a frame another analyst could apply to a fresh transcript and reach broadly the same assignments. If two codes cannot be told apart from their definitions alone, they are one code.

**5. Stop for human review of the frame.** Per K5 §6, the code frame is reviewed by a researcher **before** it is applied at scale. Present the frame, the boundary examples, the codes you were unsure about, and any place where the material resisted the frame. *Correct result:* an approved frame, with a record of what changed in review. Applying an unreviewed frame to 30 transcripts multiplies a definitional error by 30.

**6. Apply the frame to the full dataset.** Code every passage against the approved frame. Where a passage does not fit, do not force it: open a candidate code and log it. New codes appearing late in the corpus are informative, so track them (see step 11). *Correct result:* a complete coded dataset, every instance carrying participant ID, the verbatim span, and whether the material was volunteered or came in response to a direct question.

**7. Check coding consistency.** Where AI has done the coding, re-code a random sample of at least 15 percent of the corpus in a fresh pass, blind to the first, and compare. Report disagreement rates by code, not just overall, because disagreement concentrates in codes with weak definitions. Send every disagreement to a human, and fix the definition rather than the instance. **Call this stability, not reliability.** Two passes by the same system measure whether the frame is applied consistently, not whether it is applied correctly, and an agreement coefficient over them dresses stability up as validity. Where genuine independent coders exist, report agreement conventionally and name the statistic and threshold. *Correct result:* per-code disagreement rates, a list of definitions tightened, and the residual disagreements a human resolved.

**8. Construct candidate themes.** A theme is not a topic bucket. A theme has a **central organising concept**: a single idea explaining why these extracts belong together and what they collectively say. "Price" is a topic. "Price is read as a signal of whether the organisation respects the customer's time" is a central organising concept. Cluster codes into candidates, write the organising concept for each in one sentence, and discard clusters for which you cannot write one, because they are usually domain summaries. Where framework analysis is in use, chart cases against themes in a matrix here. *Correct result:* five to eight candidate themes, each with a written organising concept and the codes it draws on.

**9. Test each candidate theme against the entire dataset.** This is the step most often skipped and the one that separates analysis from impression. For each candidate, return to the full corpus and ask four questions. Does it hold in the transcripts that were not used to build it? Is there material that contradicts it? Is it internally coherent, or is it two themes wearing one name? Is it distinct from its neighbours, or are the boundaries arbitrary? Split, merge, redraw or discard accordingly. *Correct result:* a revised theme set where at least one candidate has changed shape. If nothing changed, the test was not run properly.

**10. Interrogate contradiction, ambivalence and divergence deliberately.** Run three separate sweeps. First, **within-participant contradiction**: find participants who said incompatible things and record what each side of the contradiction was attached to, because the conditions under which a person switches position are frequently the finding. Second, **between-participant divergence**: identify participants who diverge from the majority reading, and characterise them by who they are and what they have experienced, rather than as noise. Third, **ambivalence**: find themes where participants hold two things at once and neither resolves. **Do not resolve these into a clean statement.** Write the theme so that the tension survives into the report: "participants valued X and were constrained by X, and did not experience this as a contradiction" is a stronger finding than either half alone. *Correct result:* every theme carries an explicit note of the evidence that cuts against it, and at least one theme is stated as a tension rather than a position. A theme set with no contradictions in it has almost certainly been smoothed.

**11. Assess saturation honestly, or decline to claim it.** Track new codes against the order in which material was analysed. **Code saturation** (no new codes appearing) arrives earlier than **meaning saturation** (no new dimensions within existing codes), and only the second matters for interpretation. Where new codes were still appearing in the final transcripts, saturation was not reached, and the statement says so and names what remained open. Where the sample was fixed before analysis, as it usually is commercially, saturation can only be assessed after the fact and never designed for. Do not claim saturation from a sample size alone: no number saturates by itself. *Correct result:* either evidence-based saturation with the tracking shown, or an explicit statement that saturation is not claimed and which themes are provisional as a result.

**12. Report prevalence in participants, and separate it from importance.** Count participants, not mentions, unless mentions are genuinely the unit of interest and that is stated. Use counts, never percentages, on qualitative bases, per K4 §7. Show the denominator, and where only some were asked, show that denominator rather than the sample size. Mark each theme as predominantly volunteered or prompted. **Then assess importance separately**, using the criteria in Section 8, and present the two side by side rather than collapsing them. *Correct result:* a prevalence table and an importance assessment that do not automatically agree, with reasoning visible where they disagree.

**13. Name and define each theme.** The name is a **label, not a finding**. It must be short, memorable, distinctive, and must not embed the conclusion, because names travel into decks and get quoted as findings, per K2 §7. Pair every name with a one-sentence definition, a boundary statement of what the theme excludes, and the tension it holds if it holds one. *Correct result:* a name a stakeholder can repeat, sitting above a definition that a stakeholder cannot mistake for the evidence.

**14. Assemble evidence, then interpret, in that order.** For each theme, pull the supporting extracts with participant IDs, including at least one that shows the theme's boundary or its counter-case. Verify every quote against source. Only once the evidence is assembled, write the interpretation, marked as interpretation per K2 §3.2, with a confidence level per K3. Connect each theme explicitly to the research objective it answers, and name the objectives no theme addressed. *Correct result:* a theme record where a reader can see the evidence before the argument, and where the point at which judgement enters is visible.

## 8. Analytical framework

The chain the output is built on:

    Extract → Code → Code frame → Candidate theme → Tested theme → Theme record
                                                                       ↓
                            Prevalence  ·  Importance  ·  Divergence  ·  Interpretation

**Applying it.** Everything to the left of "Theme record" is procedure that can be audited. The four elements below it are reported separately and never merged, because merging them is how a frequent theme becomes an important one by default and how an interpretation becomes an observation.

**The frequency-importance grid.** Every theme is placed on two axes, and the interesting themes are off the diagonal.

| | Low importance | High importance |
|---|---|---|
| **High prevalence** | Report briefly. Often context, hygiene factors, or an artefact of what the guide asked | Lead with it |
| **Low prevalence** | Note in the appendix or drop | **The most commonly missed material in qualitative work.** Report it, with its base, and say why it matters |

**Judging importance.** A theme is important to the degree that it: sits close to the decision the research exists to inform; carries high consequence if true; explains other themes rather than sitting alongside them; is held by participants who occupy a structurally significant position (the people who left, the gatekeepers, the highest-value users, the ones the organisation never hears from); is tractable, in that someone could act on it; and is not already known. A theme raised by five of thirty participants who all recently defected can be worth more than a theme raised by all thirty, because the thirty are describing the category and the five are describing the failure. Prevalence is evidence about how widely something is held. It is not evidence about how much it matters, and the two are reported as separate columns for exactly that reason.

**Suppressed prevalence.** Low counts are sometimes an artefact of method rather than a fact about the world: the guide did not ask, the topic is sensitive, the theme only arises unprompted, or the moderator moved on. Where you can see this in the material, say so, because it changes how a low count should be read.

## 9. Output format

**A. Analysis note (front matter)**
Objectives; tradition and coding approach used, with the reason; unit of analysis; prevalence denominator and how it was derived; corpus description (participants, material type, what was excluded and why); statement that AI performed coding and what a human verified, per K4 §7; consolidated review points, per K5 §3.3.

**B. Code frame**

| Code | Parent | Definition | Include | Exclude | Boundary example | Origin (deductive / inductive) | Participants | Instances |
|---|---|---|---|---|---|---|---|---|

**C. Theme records** (one per theme, in importance order, with prevalence shown alongside)

    Theme name (label only)
    Definition: one sentence
    Central organising concept: what holds these extracts together
    Boundary: what this theme is not
    Prevalence: n of N participants [denominator described], volunteered / prompted
    Codes drawn on: list
    Evidence: 3 to 6 verbatim extracts, each [PID, characteristics]
    Counter-evidence: extracts that cut against it, with IDs
    Tension held: the ambivalence, if the theme holds one, stated unresolved
    Divergent voices: who differs, who they are, and what they see
    Objective addressed: which research objective this answers
    Interpretation: marked as interpretation, with confidence per K3
    Importance: assessment and reasoning, separate from prevalence

**D. Prevalence table**

| Theme | Participants | Denominator | Volunteered / prompted | Subgroup concentration | Notes on suppressed prevalence |
|---|---|---|---|---|---|

**E. Divergence and contradiction register**
Every within-participant contradiction and between-participant divergence found in step 10, with IDs, whether it was resolved, and how.

**F. Saturation statement**
Either the tracking evidence, or an explicit statement that saturation is not claimed and which themes are provisional as a result.

**G. What could not be established**
Objectives no theme addressed; questions the material cannot answer; themes that failed testing and why.

**When the evidence is thin.** The format does not get filled to look complete, per K4 §1. A theme carried by three participants is reported with a base of three and a confidence level to match, not promoted to fill a slot. A theme with no counter-evidence records "none found in this corpus", which is a finding, not an empty field. Where fewer themes survived testing than the deck template expects, the output has fewer themes and says why. Where a research objective produced nothing, it appears in section G rather than being answered thinly.

## 10. Quality checks

Run before anything is presented. Sits on top of K4 §8.

1. Does every quote carry a participant ID, and does every ID exist in the corpus?
2. Has every client-facing quote been verified against the source text, word for word?
3. Is any quote a merge of two participants, or attributed to a segment the speaker does not belong to?
4. Is prevalence counted in participants, with the denominator stated, and free of percentages on small bases?
5. Does the denominator reflect who was actually asked, rather than the full sample by default?
6. Is each theme marked as predominantly volunteered or prompted?
7. Was each theme tested against transcripts that were not used to construct it?
8. Does every theme record contain counter-evidence, or an explicit statement that none was found?
9. Is at least one theme stated as an unresolved tension, and if none is, has the corpus genuinely no ambivalence in it?
10. Are divergent participants characterised rather than dismissed as outliers?
11. Is prevalence reported separately from importance, with the importance reasoning visible?
12. Is every theme name a label rather than a finding, and would it be safe if quoted alone on a slide?
13. Was the code frame reviewed by a human before it was applied at scale, and is that recorded?
14. Does the saturation statement claim only what the tracking supports?
15. Is every research objective either answered by a theme or listed in "what could not be established"?
16. Could a colleague reconstruct how the top three themes were built, from the frame and the records alone?

## 11. Common failure modes

| Failure | How to recognise it | How to prevent it |
|---|---|---|
| **Coherence over-resolution** (the signature AI failure) | Ambivalent material has become a clean, confident statement. The participant who said "I love it and I have stopped using it" appears only on one side | Run step 10 as a separate sweep. Write the tension into the theme definition. Treat any theme with no internal friction as suspect until checked |
| **Word frequency mistaken for analysis** | Themes track the most-repeated nouns; the theme list resembles the discussion guide | Code for meaning, not topic. Require a central organising concept for every theme. Check whether high-prevalence themes are guide artefacts |
| **Theme built from the first five transcripts** | Evidence for a theme clusters in the earliest IDs | Step 9, mandatory. Record which transcripts each theme was tested against |
| **Topic summary presented as a theme** | The theme name is a domain noun and the definition restates it | If the organising concept cannot be written in one sentence, it is not a theme |
| **Prevalence inflation** | "Most participants felt..." with no count, or counts of mentions presented as counts of people | Count participants. Show n and denominator on every prevalence claim |
| **The vanishing minority** | A clean four-theme story with no divergence register | Step 10, sweep two. Characterise divergent participants rather than dropping them |
| **Quote polishing** | Quotes read fluently, in a consistent register, with no hesitation or repair | Verify against source. Permitted edits only, per K4 §2.3. If quotes across participants sound alike, something has been rewritten |
| **Composite or drifted attribution** | A quote captures the theme unusually well; an ID appears against a segment it does not belong to | Never merge speakers. Check the speaker's characteristics against the segment claimed |
| **Theme name as finding** | The name contains a verb and a conclusion, and stakeholders quote it as evidence | Names are labels. Conclusions live in the interpretation field |
| **Deductive frame smothering the material** | Almost nothing was coded outside the prior frame; "other" is large | Log every forced fit. Report where the prior frame did not accommodate this wave |
| **Saturation claimed from sample size** | "We reached saturation with 24 interviews" with no tracking | Track new codes by order of analysis, or decline the claim |
| **Interpretation written first** | The interpretation is elegant and the evidence supporting it is sparse or repetitive | Assemble evidence before writing meaning. Step 14 order is not cosmetic |
| **Confirmation of a supplied hypothesis** | Themes map suspiciously well onto what the stakeholder said before fieldwork | Record hypotheses at step 1 so the overlap is visible and can be examined |

## 12. AI guardrails

Skill-specific only. Universal prohibitions are inherited from K4 and are not repeated here. Quote handling follows **K4 §2.3**, including the list of permitted edits.

1. **Never resolve ambivalence toward the more coherent reading.** Where a passage supports two readings, or a participant holds two incompatible positions, both are recorded and the theme is written to hold the tension. Per K5 §2.6 this is a review point, not a decision to make alone.
2. **Never report a theme without the count of participants behind it and the denominator that count is out of.** "Many", "most" and "several" are not prevalence statements.
3. **Never let prevalence stand in for importance, or importance stand in for prevalence.** They are separate fields in the output and separate judgements in the reasoning.
4. **Never present a code frame as validated because a second AI pass agreed with the first.** Agreement between passes is stability. Report it as stability and say so explicitly.
5. **Never apply a code frame at scale before a human has reviewed it.** Per K5 §6. Record the review and what changed.
6. **Never drop a passage because it does not fit the frame.** Log it, open a candidate code, and report the residue. The material that resists the frame is where the frame is wrong.
7. **Never characterise a divergent participant as an outlier without describing who they are.** Divergence is attributed and explained, or it is not reported as divergence.
8. **Never claim saturation without the tracking evidence**, and never infer it from sample size.
9. **Never write a theme name that contains its own conclusion.** Names are labels and they travel unaccompanied.
10. **Never interpret cultural or linguistic material outside your competence.** Code descriptively and flag for a human reader with the relevant context, per K5 §2.2.

## 13. Best-practice principles

1. **Themes are constructed, not discovered.** The language of themes "emerging" hides the analyst's decisions and makes them unauditable. Say what you decided and why.
2. **The interesting material is where people disagree with themselves.** Consistency across a transcript is often social performance. The moment a participant qualifies, hesitates or reverses is usually the moment they stopped performing.
3. **Read the absence.** What nobody said, in a discussion where you would expect it, is evidence. It is also the most fragile kind of evidence, so state it as an observation about the corpus, not about the world.
4. **Prompted and unprompted are different data.** A theme every participant discussed because the guide asked about it is weaker evidence of salience than a theme a third raised unbidden.
5. **A small number of the right people beats a large number of the wrong ones.** Weight the base by who is in it, not only by how big it is. Five people who churned are a different kind of evidence from thirty who stayed.
6. **The counter-case makes the theme stronger, not weaker.** A theme that has been tested against contradicting material and survived is more defensible than one that has never met any. Report the test.
7. **Code frames rot without boundary examples.** Definitions drift under application; a passage that nearly qualifies and does not is worth more than three paragraphs of definition.
8. **Resist the four-theme deck.** Studies do not reliably produce a tidy number of themes. Where six survive testing, report six. Where three do, report three and say what the fourth candidate failed on.
9. **Verify the quote you love most first.** The extract that captures a theme perfectly is the one most likely to have been smoothed, trimmed past the permitted edits, or remembered rather than read.
10. **Keep the participant visible.** A theme abstracted so far from its speakers that no individual is recognisable in it has usually lost the thing that made it true.
11. **Write meaning last, and write it separately.** If the interpretation is drafted while coding, coding becomes evidence-gathering for a conclusion already reached.
12. **The analysis is finished when someone else could rebuild it.** Not when it reads well.

## 14. Worked example

*Fictional scenario, used for illustration only. All participants, quotes and findings below are invented for the purpose of demonstrating method.*

**INPUT.** A national charity running a volunteer-staffed community advice service commissions research into why trained volunteers stop volunteering within twelve months. Twenty-two depth interviews with volunteers who left, plus the discussion guide and a joiner cohort file. Objective: understand what drives early attrition and what would retain volunteers.

**PROCESS.**

*Steps 1 to 3.* Hypotheses recorded before analysis: the client believes attrition is about time pressure and unpaid travel costs. Familiarisation memo notes that travel costs are mentioned often and briefly, while training is discussed at length and with unusual warmth. Open coding across all 22 transcripts produces 94 initial codes.

*Steps 4 to 6.* Codes collapse into a frame of 31 codes under 6 parents. Two codes, "training quality" and "training expectations", are nearly indistinguishable, so a boundary example is written: a volunteer praising the trainer is *training quality*; a volunteer describing what they expected the role to be *because of* the training is *training expectations*. Human review splits a third code and rejects one as a topic label. Frame approved, then applied across the corpus.

*Step 7.* A blind re-code of four transcripts (18 percent) disagrees on 11 percent of assignments, concentrated in one code with a weak exclusion rule. The rule is tightened. Reported as stability, not reliability.

*Steps 8 to 9.* Five candidate themes. One, provisionally "Cost of volunteering", covers travel and time; tested against the corpus, it fragments, because time pressure is described by leavers as a *consequence* of feeling ineffective rather than a cause of leaving. It is redrawn.

*Step 10, and the judgement call.* Fourteen of 22 volunteers describe the training in strongly positive terms. Twelve of those same fourteen describe arriving at their first shift unable to do the thing the training prepared them for, because the real cases were messier than the training cases. The tempting move is to resolve this into "training was inadequate". The transcripts do not support that: participants insist the training was good, and several say so after describing its failure. The tension is the finding, so it is written as one: volunteers experienced excellent training and profound unpreparedness at the same time, and did not treat these as contradictory, because they blamed themselves rather than the training. Flagged for researcher review as ambiguous qualitative evidence per K5 §2.6.

Separately, four volunteers describe a safeguarding incident they did not know how to escalate. Four of 22 is low prevalence. Assessed on importance: high consequence, close to the decision, held by participants in a structurally significant position (all four left within eight weeks), and plausibly suppressed, since the guide never asked about safeguarding and all four raised it unprompted. It is reported prominently with its base of four shown, and marked as a hypothesis for direct testing per K3 §4.3.

*Steps 11 to 14.* New codes were still appearing at transcript 20, so saturation is not claimed. Themes named as labels ("The competent stranger", "Nobody to ask"), each with a definition beneath. Quotes pulled with IDs and verified. Interpretation written last.

**OUTPUT.** Four themes with prevalence, counter-evidence and tensions intact; a divergence register; a four-participant safeguarding finding reported despite low prevalence, with reasoning shown; a statement that travel cost, the client's lead hypothesis, was mentioned by 9 of 22 but never as a reason for leaving; and a saturation statement declining the claim.

## 15. Advanced usage

**Multi-wave and tracking qualitative.** Freeze the code frame between waves and log additions separately, so that change in the material is distinguishable from change in the frame. Report new codes as a finding about the wave, not as frame maintenance.

**Cross-market work.** Code within market first, then compare frames before merging. A code that exists in one market and not another is evidence, and merging early destroys it. A native-language reviewer is required per K5 §2.2, and the reporting states which themes were built on translated material.

**Focus groups.** The unit of analysis needs a decision that individual interviews do not force: the participant, the group, or the interaction. Record who spoke first, who deferred, and where consensus formed rather than existed. Group consensus is a social product; a theme "held by 8 of 10 in the group" may be one confident person and nine who did not object.

**Negative case analysis.** Actively hunt for the participant who most disconfirms each theme and try to break the theme with them. A theme that survives a deliberate attempt to falsify it earns a higher confidence level under K3 §3.5.

**Combining with quantitative material.** Where a survey ran alongside, test themes against it during step 9 rather than at integration. Divergence found early can be investigated; divergence found late gets averaged away. Hand to **07.06**.

**When the standard approach does not fit.** Very short fieldwork with two or three interviews cannot support thematic analysis; run **07.03** per transcript and report as cases. Highly structured, objective-led work with a client who needs a comparison matrix should switch to framework analysis at step 8. Material where the interest is the language itself, rather than what the language reports, needs discourse or narrative analysis, which this skill does not cover.

## 16. Skill chain

**Recommended previous skills:**
- **07.03 Interview and Transcript Analysis.** Hands over cleaned, speaker-labelled transcripts with per-participant accounts already understood, so that cross-participant theme building starts from comprehension rather than from raw text.
- **04.02 Data Cleaning** where open ends or diary data arrive as part of a dataset.
- **02.02 Discussion Guide Design**, retrospectively, to establish what was prompted and what was volunteered.

**Recommended next skills:**
- **07.04 Quote and Evidence Extraction.** Takes the theme records and assembles verified, correctly attributed evidence for specific claims in the report.
- **07.06 Qualitative and Quantitative Integration.** Takes themes with prevalence and tension intact and tests convergence and divergence against the quantitative stream.
- **08.01 Finding to Insight Development.** Takes tested themes and works the step from what the data shows to why it is happening.

**Runs well alongside:**
- **13.03 AI Output Verification**, run against the coded output and the quote set before the analysis is trusted.
- **13.04 Bias Detection**, particularly where stakeholder hypotheses were recorded at step 1.
- **K5**, at the three review points this skill mandates: the code frame before scaled application, ambiguous passages, and cultural or linguistic context.

---
A Yazi Supplied Skill and resource.
