---
name: open-ended-response-coding
description: >
  Systematic coding of high-volume short open-ended survey text into a defined
  code frame with honest prevalence. Use for "code these open ends", "build a
  code frame for this question", "what did people say in Q12", "categorise these
  verbatims", "apply last wave's code frame", "why did they give that score",
  "summarise these survey comments", "quantify the open ends".
category: 07 Qualitative Analysis
ref: 07.02
tier: 1
inherits: [K2, K3, K4, K5]
---

# Open-Ended Response Coding

## 1. One-line description
Turns a large volume of short open-ended survey responses into a defined, auditable code frame and reports prevalence against the correct denominator, including the responses that said nothing.

## 2. What this skill is used for

**The research problem it solves.** An open-ended question produces the most honest material in a survey and the most abused. It arrives as a column of a few thousand fragments, most of them under fifteen words, and it gets handled in one of three bad ways: it is summarised into a paragraph of impressions with no counts behind it, it is machine-clustered into groups nobody can define, or it is coded into a frame that was never written down, so no two analysts and no two waves produce the same answer. Each failure destroys the same thing, which is the ability to say how many people said what, out of how many, and to be checked on it. This skill supplies the discipline: read before building, define every code so a second coder would agree, keep the frame stable where comparability matters, count against a denominator you have stated, and treat the blank response as data rather than as absence.

**Where it sits.** Analysis. It takes a prepared survey dataset and hands coded frequencies to descriptive analysis, cross-tabulation and reporting, and hands coded qualitative structure to insight development.

**Typical use cases.**
- Reason-for-score follow-ups behind a rating or recommendation question, a few hundred to a few thousand responses.
- Unaided awareness, unaided brand or unaided attribute mentions requiring a mention list.
- "Anything else you would like to tell us" free text at the end of a questionnaire.
- Wave-on-wave tracking of a standing open end where the frame must stay comparable.
- Complaint, cancellation-reason or exit-survey text at volume.
- Public consultation responses where every submission must be accounted for.

**Who uses it.** Quantitative researchers who have inherited a qualitative column; insight and CX teams running continuous feedback; public sector analysts handling consultation returns; students and generalists who need coded frequencies they can defend.

## 3. When to use it

- You hold short open-ended text, typically 100 to 5,000 items, where most responses are a phrase, a sentence or two, and few run past a short paragraph.
- The deliverable is a code frame with prevalence: what was said, by how many, out of how many.
- The responses sit inside a survey, so the base, the routing and the question wording are all knowable and all matter.
- The same question will be asked again, and this wave has to be comparable to the last.
- A quantitative result needs explanation, and the open end is the only place the explanation exists.
- The volume is beyond what a person will read carefully, but not so large that reading a sample is meaningless.
- The client wants a number from qualitative text and needs to be told precisely what that number can and cannot support.

## 4. When NOT to use it

- **The material is rich rather than short.** Paragraph-length reflective answers, diary entries and long-form written responses carry meaning in how the respondent explains themselves, and coding them to a frame flattens that. Use **07.01 Thematic Analysis**, which builds themes with central organising concepts, tests them across the dataset, and preserves contradiction. The boundary in the other direction: 07.01 explicitly hands short, thin material here, because thematic analysis on one-line answers produces a topic list dressed as an analysis.
- **The material is conversational.** A transcript, a chat log or an AI-moderated interview has sequence, prompting and repair in it. Coding it as though each turn were an independent survey answer destroys the context that makes it interpretable. Use **07.03 Interview and Transcript Analysis**.
- **The volume is beyond reading at all.** Above roughly 5,000 items, a genuine sample read still works but frame construction by reading breaks down, and any claim to have engaged with the corpus becomes false. Use **06.06 Large-Scale Text Analytics** to structure the corpus first, then code a sampled subset with this skill and state the sampling rule.
- **The question did not work.** Where a large share of responses show that respondents misread the question, answered a different question, or could not answer it, the correct output is a note on the instrument, not a code frame. Coding a broken question produces a precise distribution of confusion. Take the wording to **02.04 Question Bias Detection** and report the failure.
- **The answered base is too small to report as proportions.** Below 30 responses, report verbatim or as counts only, per K4 §7. Below roughly 100 answered, a code frame is still useful for structure but the percentages should be read directionally and the base disclosed at every mention.
- **The client wants attitudes, and the open end measures salience.** An unprompted question measures what came to mind, not what people believe. If the deliverable is "what proportion of customers think X", an open end cannot supply it and a prompted, closed measure is required. Coding will produce a number and the number will answer a different question.
- **The task is really quote selection.** Where the deliverable is verbatim evidence for a report rather than a distribution, the coding is a means and **07.04 Quote and Evidence Extraction** is the skill. Do not present a code frame as a quote set or a quote set as prevalence.
- **The responses are identifying and the reporting is public.** Free text is where respondents put names, addresses, account numbers and details that identify them. Where the output will be published or shared beyond the research team, disclosure control comes first. This is **13.05 Research Ethics and Consent Design**, and it is not solved by removing the obvious cases.

## 5. Required inputs

**Required.** Without these the skill cannot run. If absent, ask. If no answer is available and the work must proceed, state the assumption at the point where it bites, per K5 §5.

- **The full set of responses, unfiltered**, including blanks, single characters, punctuation-only entries and anything that looks like junk. A file that has already had the empties stripped has had the denominator changed before you saw it, and you cannot recover it.
- **The exact question wording, including any preamble and any preceding question.** An open end is inseparable from what was asked immediately before it. The same words after a satisfaction rating and after a complaint question produce different data.
- **The routing.** Who was asked this question and who was not. This determines the denominator and it is not the same as the sample size.
- **A respondent identifier on every response**, so a code can be crossed against segments and a quote can be traced back, per K2 §4.2.

**Optional, and what each one adds.**

- **The prior wave's code frame, with definitions**: makes wave-on-wave comparison possible. Without the definitions the codes are labels, and matching labels across waves without matching definitions produces a trend line built on drift.
- **The closed question the open end explains** (the score, the rating, the choice): permits coding to be crossed against the answer it justifies, which is usually where the value is. A reason given by detractors and by promoters is not the same reason.
- **Respondent characteristics and weighting variables**: allow the coded frequencies to be reported by segment and weighted consistently with the rest of the study.
- **The client's reporting categories or existing taxonomy**: tells you what the frame will eventually have to map onto, so you can build nets that survive contact with the deck rather than restructuring afterwards.
- **Previous wave's coded data, not just the frame**: allows a check that this wave's application matches the last wave's application, which is a different question from whether the frame matches.

## 6. Questions to ask before starting

1. **Is the reporting base all respondents, or only those who answered?** This is the single most consequential decision in open-end work and the commonest reporting error. Both are legitimate and they produce different numbers. *Default if unanswered:* report against those who answered, show the answered base and the total base side by side, and state the convention once at the top of the output.
2. **Can a response carry more than one code?** Multi-coding is usually correct and it means the percentages sum above 100. *Default:* multi-code permitted, capped at three codes per response, with the sum-above-100 convention stated in the output.
3. **Is a prior frame in play, and is it binding?** On a tracker the answer is almost always yes, and comparability outranks elegance. *Default:* treat the prior frame as binding, add new codes beneath it, log every addition, and never silently restructure.
4. **What is a code allowed to be about?** Some frames code the topic, some code the sentiment-bearing statement, some code the requested action. Mixing them produces a frame whose codes are not comparable to each other. *Default:* code the substantive content of what was said, hold sentiment separately (see **07.05**), and say which convention you used.
5. **What happens to the non-answers?** Blanks, "n/a", "no", "none", "nothing", "no comment" and single punctuation marks are not one thing. *Default:* classify them into explicit non-answer codes, report them, and never delete them from the base silently.
6. **Will these codes be crossed against anything?** If the frame will be cross-tabbed by segment or by score, code granularity has to survive the smallest cell. *Default:* build at the granularity the material supports, and net upward for reporting rather than coding coarsely from the start.
7. **What is the tolerance for "other"?** *Default:* a working target below 10 percent of coded responses, and a hard trigger at 15 percent, at which point the frame is rebuilt rather than the bucket accepted.

## 7. Step-by-step methodology

**The governing sequence.** Read, build, define, review, pilot, apply, audit, count. Every failure in open-end coding comes from skipping one of the first four or the seventh. The most common shortcut, going straight from the raw column to an applied frame, produces codes that cannot be defended because nobody, including the person who made them, can say what they mean.

**1. Fix the reporting frame before reading anything.** Write down the total base, the routed base, the expected answered base, the multi-code policy, the denominator convention, and whether a prior frame binds. *Correct result:* a four-line pre-analysis note. This exists so that the denominator is chosen before you know which denominator flatters the story.

**2. Read a sample, and build nothing while you read.** Draw 100 to 200 responses at random from the answered set, or stratified across the closed question the open end explains if there is one. Read them all. Note the register, the typical length, the vocabulary respondents use for the thing being asked about, and anything that suggests the question was misread. *Correct result:* a short read-note that names the three or four things people are clearly talking about, in their own words, plus any instrument problem. Building codes while reading the first fifty responses produces a frame shaped by the first fifty responses, which is the same error as building a theme from the first five transcripts.

**3. Diagnose the question before you code the answers.** Check what precedes the open end, whether it is prompted or unprompted, whether it asks for one thing or invites a list, and whether the phrasing steers. "What could we do better?" and "How was your experience?" produce different distributions from the same population, and neither is a measure of satisfaction. *Correct result:* an explicit statement of what this question measures, which is nearly always salience under a specific prompt, recorded in the output front matter. Where the wording is leading, say so here rather than discovering it in the debrief.

**4. Separate the non-answers, and classify them.** Sort every response that carries no substantive content into named categories, at minimum: blank or missing, explicit refusal ("prefer not to say", "none of your business"), explicit nothing ("nothing", "none", "no issues", "n/a"), non-responsive ("test", "asdf", punctuation only), and unintelligible. These are different findings. An explicit "nothing, it was fine" after a satisfaction question is a substantive answer about the absence of problems and belongs in the coded data. A blank is a decision not to answer, which speaks to question burden and placement. A refusal is about the question's sensitivity. *Correct result:* every response in the file has landed somewhere, and the counts for each non-answer category are reportable. Deleting them silently is a K4 §4.4 breach: it changes the denominator without a log.

**5. Build the frame bottom up, or load the prior frame.**

*Bottom up, for a new question.* Working from the read sample plus a further open pass over a larger sample, write codes at the level of the smallest distinct thing people are actually saying. Resist the urge to start with the client's reporting categories: those are where you net to, not where you code from. Aim for granularity that would survive a cross-tab, typically 25 to 60 codes for a substantial general question, fewer for a narrow one. *Correct result:* a flat list of candidate codes, each traceable to real responses you can point at.

*Prior frame, for a tracker.* Load the existing frame with its definitions intact and apply it first. Do not rebuild it because you would have built it differently. **A tracker frame that is imperfect and stable is worth more than a better frame that breaks the trend**, because the trend is the deliverable and a restructured frame silently converts frame change into apparent market change. Where the old frame genuinely does not fit this wave's material, add new codes beneath the existing structure, log every addition with the wave it appeared in, and report the additions as a finding about the wave. Where an old code has become useless, retire it but keep it visible in the trend table as retired rather than deleting the history. Restructuring is permitted only as a deliberate, announced break, with the old and new frames run in parallel for at least one wave so the discontinuity is measured rather than absorbed.

**6. Write definitions that a second coder could apply.** Every code gets a name, a one-sentence definition, an inclusion rule, an exclusion rule, and at least one boundary example: a real response that nearly qualifies and does not, with the reason. Boundary examples do more work than definitions, because drift happens at edges, not at centres. *Correct result:* two codes that cannot be told apart from their definitions alone have been merged, and any code whose exclusion rule is empty has been examined, because a code that excludes nothing is usually a net wearing a code's clothes.

**7. Build the structure: codes, nets, super-codes.** Organise the flat list into a shallow hierarchy. A **net** is a reporting aggregation of codes that belong to one topic (all price-related codes into a price net). A **super-code** is a cross-cutting aggregation that pulls codes from several nets because they share a property (everything that constitutes a service failure, wherever it sits). Two rules govern both. First, **a net is computed on unique respondents, not by summing its codes**: a respondent who gave two price codes counts once in the price net, and a net built by addition will overstate itself, sometimes badly. Second, nets are reporting objects, never coding objects. Nobody codes to a net. *Correct result:* a frame with codes at the bottom, nets defined as explicit membership lists, and the de-duplication rule written down.

**8. Stop for human review of the frame before it touches the full dataset.** Per K5 §6, present the frame, the definitions, the boundary examples, the codes you were unsure about, the proposed nets, and the non-answer categories. *Correct result:* an approved frame and a record of what changed. A definitional error applied to 3,000 responses is 3,000 errors, and it is cheaper to argue about a definition than to re-code a wave.

**9. Pilot on a fresh sample and measure fit before mass application.** Apply the approved frame to 150 to 200 responses that were not used to build it. Measure three things: the proportion falling to "other", the proportion needing a code you do not have, and the average codes per response against your multi-code policy. *Correct result:* an other rate below 10 percent and a list of the responses that resisted. Above 15 percent other, **the frame is wrong and the correct action is to return to step 5, not to accept the bucket**. A large other bucket is not a residual, it is the frame telling you it was built on the wrong sample or at the wrong level.

**10. Apply the frame to the full set.** Code every response, retaining the response text, the respondent ID, the codes assigned, and the assignment order where multi-coded. Where a response does not fit, do not force it: send it to a candidate list rather than to "other", and review the candidate list at the end. *Correct result:* a complete coded dataset in which every response, including every non-answer, carries at least one code and the residue is visible.

**11. Audit a sample of the coded output.** Draw 5 to 10 percent of coded responses at random, weighted toward the codes with the weakest exclusion rules, and re-code them blind in a fresh pass. Report disagreement by code rather than overall, because disagreement concentrates in badly defined codes and an overall figure hides it. Where AI performed the coding, call this **stability, not reliability**: two passes by the same system measure consistency of application, not correctness, and dressing it up with an agreement coefficient converts a consistency check into a false validity claim. Send every disagreement to a human, and fix the definition rather than the individual instance. *Correct result:* per-code disagreement rates, a short list of tightened definitions, and a re-run of the affected code across the dataset.

**12. Compute prevalence against a stated denominator, and show both.** For every code report: count, percentage of those who answered, percentage of all who were asked, and the two base sizes. Do not choose one. **The two numbers can differ by a factor of two on a question with a 45 percent answer rate**, and a report that shows only the answered base while the reader assumes the total base is the commonest reporting error in open-end work, in both directions. Net figures are computed on unique respondents per step 7. Where the study is weighted, coded frequencies are weighted on the same scheme as the rest of the survey, and the effective base is shown per K4 §7. *Correct result:* a frequency table where a reader cannot mistake which base a number is on.

**13. Write the prevalence statement that says what the number means.** Every open-end frequency carries an interpretive constraint and it goes in the output, not in the analyst's head. **An unprompted open end measures what was top of mind at the moment of asking, under that specific prompt. It does not measure what people think, and absence is not disagreement.** Nine percent mentioning waiting times does not mean 91 percent are content with waiting times: it means 91 percent did not raise it unprompted, which is compatible with indifference, with satisfaction, with not having thought of it, and with having mentioned something they cared about more. Where the client needs the incidence of a belief, name the closed question that would measure it. *Correct result:* a standing statement in the front matter plus a specific note wherever a low count is likely to be misread as a low incidence.

## 8. Analytical framework

The chain the output is built on:

    Response → Non-answer screen → Code(s) → Net → Prevalence (on a stated base) → Salience claim

**Applying it.** Each arrow is a place where a number can go wrong, and each is separately auditable. The non-answer screen fixes what is in the denominator. Coding fixes what is counted. Netting fixes double-counting. The base fixes the percentage. The salience claim fixes what the percentage is allowed to mean. A report that shows a percentage without the last two steps has published a number whose meaning has never been decided.

**The four denominators.** Confusion here is endemic, so name them explicitly in every output.

| Base | What it is | When it is the right one |
|---|---|---|
| Total sample | Everyone in the survey | Almost never for an open end, unless everyone was asked |
| Asked | Those routed to the question | The honest default for "how common is this in the population studied" |
| Answered | Those who gave any response, including "nothing" | The default for "of the people who told us something, what did they say" |
| Substantively answered | Those who gave codeable content | Useful for frame diagnostics, dangerous in reporting, because it is the flattering base |

**Multi-coding and what it does.** Where responses can carry more than one code, column percentages sum above 100 and the sum has no meaning. Report codes as independent proportions, never as a share of mentions unless mentions are genuinely the unit and that is stated. A "share of mentions" table on multi-coded data quietly changes the denominator to the number of codes, which is not a number of people and cannot be crossed against anything.

**The other-rate diagnostic.** Treat the other bucket as an instrument reading on the frame, not as a category.

| Other rate | Reading |
|---|---|
| Under 5% | Frame fits. Check that "other" is not being used to avoid a difficult judgement |
| 5% to 10% | Normal residue. Inspect it once for an emerging code |
| 10% to 15% | The frame is missing something specific. Find it and add a code |
| Above 15% | The frame is wrong. Rebuild from a fresh sample |

## 9. Output format

**A. Coding note (front matter)**
Question wording and what preceded it; routing and who was asked; total, asked, answered and substantively answered bases; multi-code policy and the sum-above-100 convention; denominator convention used throughout; whether a prior frame was applied and whether it was binding; statement that AI performed the coding and what a human reviewed and audited, per K4 §7; the salience statement from step 13; consolidated review points per K5 §3.3.

**B. Code frame**

| Code | Net | Definition | Include | Exclude | Boundary example | Origin (new / prior wave / added this wave) |
|---|---|---|---|---|---|---|

**C. Frequency table**

| Code | Net | n | % of answered | % of asked | Multi-coded with (top co-occurrence) | Prior wave % | Change |
|---|---|---|---|---|---|---|---|

Nets shown as unique-respondent figures, flagged as such. Change column populated only where the frame and the base are comparable, and left empty with a reason where they are not.

**D. Non-answer report**

| Category | n | % of asked | Note |
|---|---|---|---|

**E. Frame diagnostics**
Other rate; codes with the highest audit disagreement; codes added this wave and why; codes retired; any code with fewer than five responses, flagged for merging next wave.

**F. Illustrative verbatims**
Two to four per major code, with respondent IDs, verified against source, edited only per K4 §2.3 with the convention stated. These illustrate code definitions; they are not evidence selection, which is **07.04**.

**G. What this question cannot tell you**
The salience constraint applied to the specific findings; codes whose low count is likely an artefact of the question rather than of the world; anything the client asked that this open end does not answer.

**When the evidence is thin.** A frame is not padded to a target number of codes. Where only six things were said, the frame has six codes and says so. Where the answered base is under 30, the frequency table is replaced by counts and verbatims, per K4 §7, and no percentage is published. Where a code has three responses behind it, it is reported with n=3 rather than promoted into a net to make it look substantial. Where the other rate could not be brought below 15 percent, the output says the frame did not fit and describes what the residue contains, rather than presenting a frame that quietly hides a sixth of the data.

## 10. Quality checks

Run before anything is presented. Sits on top of K4 §8.

1. Does every response in the original file appear in the coded output, including blanks and junk?
2. Is the denominator stated on every percentage, and are the answered and asked bases both visible?
3. Are net figures computed on unique respondents rather than by summing member codes?
4. Where responses are multi-coded, is the sum-above-100 convention stated, and is there no "share of mentions" table masquerading as a share of people?
5. Is the other rate reported, and is it below the threshold at which the frame should have been rebuilt?
6. Does every code have an inclusion rule, an exclusion rule and a boundary example?
7. Could two codes be told apart from their definitions alone, without seeing the examples?
8. Was the frame reviewed by a human before mass application, and is that recorded?
9. Was a sample of coded output audited after application, with disagreement reported per code?
10. On a tracker, is every frame change logged, and is any trend line drawn across a frame break flagged as such?
11. Are non-answer categories reported separately rather than merged or dropped?
12. Does the output state that this measures salience under a prompt, and is any low count that could be misread as low incidence annotated?
13. Are weighted figures weighted on the same scheme as the rest of the study, with effective base shown?
14. Has any respondent-identifying detail in the verbatims been handled before the output leaves the analysis?

## 11. Common failure modes

| Failure | How to recognise it | How to prevent it |
|---|---|---|
| **The moving denominator** | A percentage in the deck does not reconcile with the frequency table, or two slides use different bases for the same question | Fix the convention at step 1, show both bases in every table, state the convention in front matter |
| **The other bucket as a filing cabinet** | Other is 20 percent and nobody has looked inside it | Step 9 threshold. Above 15 percent, rebuild. Inspect the residue every time |
| **Silent deletion of blanks** | The answered base equals the asked base, or the file arrived pre-cleaned | Demand the unfiltered column. Classify non-answers as codes, per step 4 |
| **Frame drift across waves** (the tracker killer) | The trend moves at exactly the wave the frame was tidied | Freeze the frame, add rather than restructure, run parallel frames across a deliberate break |
| **Net inflation** | A net is larger than the sum of its parts should allow, or exceeds 100 percent implausibly | Compute nets on unique respondents. Check every net against a de-duplicated count |
| **Coding the topic and the sentiment in one frame** | Codes like "price" sit next to codes like "unhappy" and cannot be crossed | Decide the convention at question 4. Hold sentiment separately, per **07.05** |
| **Absence read as disagreement** (the signature reporting error) | "Only 9 percent mentioned X, so X is not an issue" | The salience statement in the front matter, plus a specific note on the finding |
| **Codes that are labels, not definitions** | The frame is a list of nouns with no rules; two analysts disagree and neither can point at anything | Step 6. Inclusion, exclusion, boundary example, every code, no exceptions |
| **AI clustering presented as a code frame** | Groups have machine-generated names, no definitions, and cannot be reproduced | Clusters are a hypothesis about structure. Convert to defined codes or do not report them as a frame |
| **Stability reported as reliability** (the signature AI failure) | An agreement coefficient between two AI passes is offered as validation | Call it stability. Adjudicate disagreements with a human. Fix definitions, not instances |
| **Over-fine coding that cannot be reported** | Forty codes each with n under 10, and the deck nets them into four anyway | Code at cross-tab granularity, net upward, and check the smallest cell before finalising |
| **The frame built from the first hundred** | Codes appearing late in the file have no home; other rises through the dataset | Build from a random or stratified sample, pilot on fresh material at step 9 |

## 12. AI guardrails

Skill-specific only. Universal prohibitions are inherited from K4 and are not repeated here. Quote handling follows **K4 §2.3**.

1. **Never report an open-end percentage without naming its denominator in the same table.** Answered and asked are both shown, always.
2. **Never delete, merge or exclude a non-answer from the base without logging it.** Blank, refusal and "nothing" are three different findings and none of them is missing data to be discarded.
3. **Never accept an other bucket above the step 9 threshold.** A large other is a defect in the frame, not a category in the output.
4. **Never restructure a tracker code frame without declaring a break** and running both frames in parallel for at least one wave. Comparability outranks tidiness.
5. **Never compute a net by summing its member codes** where multi-coding is permitted. Nets are unique-respondent counts.
6. **Never present agreement between two AI coding passes as reliability or validation.** It is stability, and it is reported as stability.
7. **Never apply a code frame at scale before a human has reviewed it**, per K5 §6, and never present a frame as reviewed when only its output was skimmed.
8. **Never infer a belief, an attitude or an incidence from an unprompted mention rate.** Salience is not opinion, and silence is not disagreement.
9. **Never generate a code from a plausible category rather than from responses you can point at.** Every code is traceable to real text, per K2 §4.2.
10. **Never publish free-text verbatims without checking them for material that identifies the respondent**, particularly in small subgroups and public-facing outputs.

## 13. Best-practice principles

1. **Read before you build.** The frame you would write after a hundred responses and the frame you would write after five hundred are different frames, and only one of them survives the dataset.
2. **A code is a rule, not a word.** If you cannot write the sentence that excludes something, you have named a topic rather than defined a code.
3. **The boundary example is worth more than the definition.** Drift happens at edges. One well-chosen near-miss stops more disagreement than a paragraph of prose.
4. **Stability beats elegance on a tracker.** The wave-on-wave comparison is usually the client's whole reason for asking, and a better frame that breaks it has destroyed more value than it added.
5. **Count people, not mentions**, unless mentions are genuinely the unit of interest and you have said so. Mentions are a base nobody can cross-tabulate.
6. **The blanks are a finding.** A question with a 30 percent answer rate has told you something about the question, and often about the respondent's willingness to engage with the survey by that point.
7. **Open ends explain scores better than they measure the world.** The strongest use of a coded open end is crossed against the closed question it justifies. Read the codes by score band before reading them in total.
8. **Granularity is cheap at coding time and impossible afterwards.** You can always net upward. You cannot split a coarse code without re-coding.
9. **A frequency table from an open end is a qualitative object wearing quantitative clothes.** Treat it with quantitative discipline on the base and qualitative humility on the meaning.
10. **Look at what nobody said, and say so carefully.** A topic the client expected and nobody raised is worth reporting as an observation about this question, never as an observation about the population.
11. **Codes with n under five are candidates for merging, not for headlines.** Report them, hold them for the next wave, and do not build a narrative on them.
12. **The frame is finished when someone else could apply it to a new wave and get your answer.** Not when it looks complete.

## 14. Worked example

*Fictional scenario, used for illustration only. All figures, responses and findings below are invented for the purpose of demonstrating method.*

**INPUT.** A city transport authority runs a quarterly rider survey. After a five-point satisfaction rating, respondents routed as either satisfied or dissatisfied (n=2,410 of 3,180 total) are asked: "In your own words, why did you give that rating?" This is wave 6. A frame exists from wave 1, with 34 codes under 7 nets. The client's question: has anything changed, and why did the satisfaction score fall two points?

**PROCESS.**

*Steps 1 to 3.* Bases fixed: total 3,180, asked 2,410, answered 1,844 (76 percent of asked), substantively answered 1,702. Multi-coding permitted, capped at three. Denominator convention set to answered, with asked shown alongside. The prior frame is binding: this is a tracker and the trend is the deliverable. A random sample of 180 responses is read before anything is opened. The read-note flags that riders are using a word for a new ticketing app that does not exist anywhere in the wave 1 frame, because the app launched between waves.

*Step 4.* Non-answers classified: 566 blank, 88 "n/a" or "no", 41 "nothing to add", 32 non-responsive strings, 7 refusals. Note recorded that "nothing to add" from a satisfied rider is a substantive statement about the absence of problems and is coded as such rather than counted as a non-answer, while a blank is not.

*Step 5, and the judgement call.* Applying the wave 1 frame, two codes overlap badly: "ticket machine problems" and "buying a ticket". In wave 1 these were distinguishable because ticket machines were the only channel. With an app in the field the distinction has collapsed and roughly a third of relevant responses could sit in either. The tempting move is to merge them into a clean "ticketing" code. That would break five waves of trend on two separately reported lines. The decision taken: **preserve both codes unchanged, add two new codes beneath the existing ticketing net for app-specific content, and report the ambiguity explicitly** as a note on the two legacy codes rather than resolving it by restructuring. The frame is now imperfect and comparable, which is the correct trade on a tracker. Flagged for researcher review as a methodological trade-off under constraint, per K5 §2.7, with a recommendation to declare a deliberate frame break at wave 8 and run parallel frames for one wave.

*Steps 6 to 9.* Definitions written for the two new codes, each with a boundary example (a rider complaining the app crashed is *app reliability*; a rider complaining they could not work out which fare to buy in the app is *fare comprehension*, which is a legacy code and stays there). Human review approves the frame and splits one of the new codes. Pilot on 200 fresh responses returns an other rate of 7 percent.

*Steps 10 to 12.* Full application. Audit of 8 percent of coded responses, blind, disagrees on 9 percent of assignments, concentrated in the two legacy ticketing codes, which is exactly where the frame is known to be weak. Reported as stability, with the concentration named. Prevalence computed on both bases.

*Step 13, and the second judgement call.* Reliability of service appears in 11 percent of answered responses, down from 16 percent at wave 5. The draft narrative offered is "reliability concerns have halved as a share of complaints". This is wrong twice: it is not halved, and a fall in unprompted mentions is not a fall in the underlying problem. It is equally consistent with reliability being displaced from top of mind by the new app, which appears in 19 percent of responses at its first wave. The finding is written as a change in salience, with the displacement stated as the most likely reading at moderate confidence per K3 §4.2, and the recommendation is to read it against the closed reliability rating rather than instead of it.

**OUTPUT.** A frame of 36 codes under 7 nets, two added this wave and logged; a frequency table on both bases with wave 5 comparisons and one line marked non-comparable; a non-answer report showing the 24 percent who were asked and did not answer; an other rate of 7 percent; a salience statement explaining why the fall in reliability mentions is not evidence of improved reliability; and a flagged recommendation for a declared frame break at wave 8.

## 15. Advanced usage

**Coding against the score.** Where the open end explains a rating, build the frequency table by score band before building it in total. The same code carries opposite meaning across bands: "staff" from a top-box rater and "staff" from a bottom-box rater are the same word and different findings. Where this happens often, it is a signal that the frame is coding topics where it should be coding statements, and the fix is at step 5, not in the reporting.

**Emergent code detection on trackers.** Track the other rate wave on wave rather than only within a wave. A rising other rate across three waves is the earliest available signal that something is happening in the category that the frame cannot see, and it usually precedes any movement in the closed metrics.

**Very large volumes.** Above roughly 5,000 items, use **06.06** to segment the corpus first, then draw a stratified sample for frame construction, apply at scale, and audit at a higher rate than usual because errors no longer surface by inspection. State the sampling rule in the output.

**Multi-language open ends.** Code in the original language where possible, with a native reviewer per K5 §2.2. Where translation is unavoidable, code the translation, mark those codes as translation-dependent, and expect the frame to fit worse in the translated set. A code that exists in one language's responses and not another's is evidence, not an error, and merging frames across languages too early destroys it.

**Feeding the qualitative chain.** Where coded responses turn out to be richer than expected, hand the substantive subset to **07.01** for thematic treatment rather than forcing depth out of a frequency table. Where the deliverable needs verbatim evidence, hand the coded set to **07.04** with the codes as the search structure.

**Public consultation.** Where every submission must be accounted for individually, the audit rate rises to 100 percent for any code that carries a policy consequence, and the non-answer report becomes a formal part of the output rather than a diagnostic. Volume-based prevalence is reported with an explicit statement that consultation respondents are self-selected and the counts are not population estimates, per K4 §3.3.

## 16. Skill chain

**Recommended previous skills:**
- **04.02 Data Cleaning.** Hands over a dataset with the open-end column intact and unfiltered, and a cleaning log that shows nothing was removed from it.
- **02.01 Survey Questionnaire Design**, retrospectively, to establish what the question asked, what preceded it, and whether it was prompted.
- **04.05 Weighting and Base Management**, where the study is weighted and coded frequencies must match the weighting scheme used elsewhere.

**Recommended next skills:**
- **05.03 Cross-Tabulation.** Takes coded frequencies and crosses them against segments and against the closed question the open end explains.
- **07.05 Sentiment and Emotion Analysis.** Sits downstream of this skill: sentiment is scored on responses that have already been coded for content, because sentiment without a cause is not a finding.
- **07.04 Quote and Evidence Extraction.** Takes the coded set and selects verified, representative verbatims for the report.
- **08.01 Finding to Insight Development.** Takes coded prevalence with its bases and salience constraints intact.

**Runs well alongside:**
- **07.01 Thematic Analysis**, where a subset of the responses proves rich enough to warrant depth.
- **13.03 AI Output Verification**, run against the coded dataset and the frequency table.
- **K5**, at the two review points this skill mandates: the code frame before mass application, and the post-application audit.

---
A Yazi Supplied Skill and resource.
