---
name: interview-question-development
description: >
  Writes the individual questions and probes that go inside a discussion guide or
  interview. Use when someone says "how should I word this question", "these
  questions feel leading", "write me some interview questions", "how do I ask why
  without asking why", "how do I ask about the last time they did this", "how do I
  ask about something embarrassing", "what probes should I use", or when an
  interview produced fluent answers that turned out to be useless.
category: 02 Instrument Design
ref: 02.03
tier: 1
inherits: [K2, K3, K4, K5]
---

# Interview Question Development

## 1. One-line description
Writes and repairs the individual questions and probes inside a qualitative instrument, so that each one asks the participant to perform a mental act humans can reliably perform, and returns an account rather than a rationalisation.

## 2. What this skill is used for

**The research problem it solves.** A guide can have the right sections, the right order and the right time budget and still produce nothing usable, because the failure happens at the sentence. Most weak interview questions ask people to do things people cannot do: predict their own future behaviour, explain their own motivation, average their own conduct into a general tendency, rank options they have never compared, or recall the frequency of something unremarkable. The participant complies. They always comply, because they are cooperative, articulate and unwilling to look blank, and what comes back is fluent, confident and constructed on the spot. The damage is that this failure is invisible in the transcript. A rationalisation reads exactly like an account. A plausible reconstruction reads exactly like a memory. Nothing in the recording marks the moment where the participant stopped reporting and started composing, so the analyst codes invention as evidence and nobody ever finds out. The second failure is smaller and easier to see: leading wording, imported vocabulary, evaluative probes, questions that can be answered "yes" by someone with nothing to say.

**Where it sits in the research lifecycle.** Inside **02.02 Discussion Guide Design**, not instead of it. The guide fixes which sections exist, what each must achieve, how many minutes it gets and what order everything runs in. This skill writes the sentences inside those sections. It is also used to repair questions mid-fieldwork, and to write questions for automated interviewing, where wording cannot be rescued live by a moderator.

**Typical use cases.**
- Turning a section purpose ("the moderator should have heard one specific occasion when...") into the question that opens it and the probes that deepen it.
- Repairing a stakeholder question list full of "why do you prefer", "would you use" and "how important is".
- Designing a critical incident or episode reconstruction sequence for a decision that happened months ago.
- Writing questions on a sensitive, stigmatised or socially regulated behaviour so that the true answer is easy to give.
- Laddering from a stated attribute to the consequence and value beneath it without putting the value in the participant's mouth.
- Writing questions for an automated or asynchronous interview, where every probe must exist in advance.

**Who uses it.** Qualitative researchers and moderators; UX and product researchers who interview without a qualitative specialist nearby; research managers reviewing a guide before fieldwork; anyone who has read a transcript full of confident answers and could not work out why it produced nothing.

## 3. When to use it

- A section purpose exists and the opening question has to be written.
- A draft guide reads plausibly but the questions all begin "why", "would you" or "how important".
- The first interviews produced short, general or normative answers and the reason is not obvious.
- The study needs a decision reconstructed that the participant made some time ago, and accuracy of the account matters more than fluency.
- The subject is one people manage the presentation of: money, health, work performance, parenting, consumption, compliance, effort.
- The instrument will be run by more than one moderator, or by an automated interviewer, so wording carries comparability rather than being adjusted in the room.
- A stakeholder has supplied questions and someone has to decide which survive contact with a real participant.
- The analysis needs a specific unit (episodes, sequences, vocabulary) and the current questions do not collect it.

## 4. When NOT to use it

- **The guide's structure has not been decided.** Writing good questions into the wrong sections, in the wrong order, for a session that cannot hold them, is wasted craft. Sequence contamination is irreversible and wording is not, so structure is settled first. Go to **02.02 Discussion Guide Design**. This skill does not decide which sections exist, how long they run, where stimulus is exposed, how a group's dynamics are managed, or what goes in the moderator layer. All of that belongs to 02.02 and this skill runs inside its section brief.
- **The objectives are not settled.** A perfectly worded question against an unclear objective produces excellent evidence about nothing. **01.01 Research Brief Interrogation** or **01.02 Business Problem to Research Question** first.
- **You are auditing an instrument rather than writing one.** For a hostile review of an existing guide or questionnaire, including one an AI produced, **02.04 Question Bias Detection**. Authoring and auditing should not be the same pass by the same mind, because an author reads what they meant and an auditor reads what is there.
- **The instrument is a survey.** Closed questions live or die on their response frames, and their wording rules are different in specific ways: no probe exists to rescue an ambiguous stem, and the option list carries as much bias as the question. **02.01 Survey Questionnaire Design** and **02.07 Scale and Measurement Selection**.
- **The objective is prevalence.** No amount of question craft makes twelve interviews countable. A beautifully worded question asked of twelve people still produces range and mechanism, not magnitude.
- **The behaviour can be observed.** Where the objective is how a task is performed, watch it. A question about a task returns the tidied, narratable version, and better wording improves the tidiness rather than the accuracy. Reserve questions for what observation cannot reach.
- **The material requires clinical competence the interviewer does not have.** Trauma, bereavement, active mental health difficulty, abuse. Wording is not the control that makes that session safe. **13.05 Research Ethics and Consent Design**, and stop until the duty-of-care structure exists.
- **The question exists to elicit a quote for a conclusion already reached.** Craft applied to this produces a more persuasive fabrication. Say so and stop.

## 5. Required inputs

**Required. Without these the skill cannot run.**
- **The section purpose and the unit of evidence it must produce.** From **02.02**: what the moderator should have heard by the end, and whether that is an episode, a sequence, a set of positions, a vocabulary or an evaluation. Without it, questions are written toward a topic, and topic-directed questions are the ones that come back general. Stop and ask.
- **What the participant has already been told.** The recruitment framing, the consent wording, and what the invitation named. A question cannot ask unaided about something the invitation has already put in the participant's head, and this is discovered too late more often than any other constraint.
- **Where this question sits in the sequence**, including whether it is above or below the stimulus exposure boundary, and what has been said before it.

**Optional, and what each one adds.**
- **Participant vocabulary from prior research, verbatims or service records.** Lets the question be written in the words the participant uses rather than the words the client uses, which is the single cheapest improvement available to any question.
- **The analysis approach (07.01, 07.03).** Tells you what the coded unit will be. A question that returns a general attitude cannot be coded as an episode no matter how the analyst tries.
- **Moderator experience level.** Determines whether probes are written as classes with one example or as fully worked alternatives.
- **Mode.** Telephone removes visual anchors and shortens tolerable question length; asynchronous text makes answers composed rather than spontaneous; automated moderation requires every probe to be pre-specified with its trigger condition.
- **Known sensitivity of the subject for this population.** Determines framing, load direction and whether an indirect route is required.
- **Language and market list.** Some question forms do not survive translation, and some (sentence completion, tense-dependent recall anchoring) are grammar-dependent.

## 6. Questions to ask before starting

1. **What mental act does answering this question require, and can a person perform it?** The governing question of this skill. Recall of a specific recent event, reconstruction of a sequence, description of a state, comparison of things actually compared: all reliable. Prediction, introspection into causes, self-averaging, and ranking of things never compared: all unreliable, and all answered fluently anyway. Default if unresolved: assume the harder act and design the question to require the easier one.
2. **Is the objective the behaviour, the account of the behaviour, or the meaning of it?** These need different questions, and conflating them is why a single question tries to do three jobs. Default: split into separate questions and let the sequencing rules in 02.02 order them.
3. **What has the participant already been told this is about?** Constrains everything unaided. Default: assume the invitation named the sponsor's category, and treat spontaneous awareness of it as unavailable.
4. **How recent and how salient is the thing being asked about?** A salient rare event (a hospital admission, a house move, a complaint) can be reconstructed a year later. A routine low-salience act (opening an app, buying milk) cannot be reconstructed reliably from last week. Default: shorten the window until the participant could plausibly name the occasion, or ask about the most recent occasion instead of a typical one.
5. **What is the socially preferred answer here, and how easy is it to give the other one?** Determines framing and load. Default: assume there is a preferred answer for anything involving money, effort, health, care, work performance or consumption, and design the permission in.
6. **Who is asking, and does that change what can be said?** A researcher visibly working for the organisation being discussed will not hear the same thing as an independent one. Default: assume the participant knows or assumes a sponsor, and avoid questions whose honest answer is a criticism of the person in the room.

## 7. Step-by-step methodology

**Step 1. Name the evidence unit before writing a word.** For each question, state what a good answer physically consists of: a reconstructed episode with a sequence of actions, a list of alternatives considered and rejected, a vocabulary set, a described state, a position on a contested issue, a reaction to a shown thing. This is not the same as the topic. "Switching" is a topic. "One occasion on which the participant seriously considered switching and did not, with what stopped them, in their words" is an evidence unit. Correct result: a one-line statement of what the analyst will hold in their hand after this question, which the wording must then deliver.

**Step 2. Classify the mental act the draft question requires, and swap any unreliable act for a reliable one.** This is the core move and it explains most of the rules that follow.

| Act the question demands | Reliability | Do this instead |
|---|---|---|
| Recall a specific recent salient event | Good | Keep it |
| Reconstruct a sequence with anchors | Good, with support | Provide the anchor (last time, the day it happened, what you were doing just before) |
| Describe a current state or an artefact in front of them | Good | Keep it |
| Compare two things they have genuinely compared | Good | Keep it |
| Report a general tendency ("usually", "typically") | Weak. Produces an on-the-spot summary weighted to the most recent and most vivid instance | Ask about the last occasion, then the one before, and let prevalence be an analysis question |
| Count a routine low-salience behaviour over a long window | Weak. Produces rounding to salient numbers and anchoring to the response options | Shorten the window, or ask "when was the last time" and reconstruct backwards |
| Explain their own motivation | Weak. Returns the most culturally available account, delivered confidently | Route through sequence, consequence, contrast, or ladder. See Steps 4 and 5 |
| Predict their own future behaviour | Very weak, and systematically generous | Ask what they did last time, what they nearly did, what stopped them, what would have had to be different |
| Rank or rate things never compared | Very weak. Manufactures a preference at the moment of asking | Ask only about comparisons the participant has actually faced, or move it to a designed trade-off exercise in category 06 |
| Introspect an unconscious process | Not possible | This is an analyst's inference from evidence, not a participant's answer. Collect the evidence and label the inference per **K2 §3** |

Correct result at this step: every question in the draft classified, and every question in the bottom five rows either rewritten or removed. A guide that survives this untouched has not been read carefully.

**Step 3. Prefer the concrete episode to the general tendency, and know the exception.** "Tell me about the last time you did X" beats "what do you usually do when X" for three reasons: the participant is retrieving rather than composing; the answer contains detail the researcher did not think to ask for, which is where most unexpected findings come from; and an episode is an analysable unit, whereas a generalisation is already the participant's own analysis, performed badly, with the working thrown away. The exception is real: where the objective genuinely is the participant's own summary judgement (their self-concept, their stated position on an issue, their sense of a relationship over time), the general question is the right one, because the summary is the thing being studied. The test is whether you would be willing to code the answer as fact about the world or only as fact about how they describe themselves. Where an episode is needed and none exists (the participant has never done the thing), do not accept a hypothetical substitute: record the absence, which is itself a finding.

**Step 4. Route around "why".** Direct "why" questions do not return motivation. They return the most available, most socially acceptable, most narratively tidy account, produced under time pressure by a person who has usually never been asked before, and they return it in a form indistinguishable from insight. Four substitutes, chosen by what is actually wanted.

- **Sequence.** "What happened just before that?" and "and then what?" Reconstruct the run-up and the aftermath, and the operative cause is usually visible in the ordering without anyone having to name it.
- **Consequence.** "What did that mean for you?" and "what would have happened if you had not?" Reaches the stake without asking for a cause.
- **Contrast.** "How was that different from the time before?" and "what would have had to be different for you to have done the other thing?" A counterfactual anchored to a real event is answerable; an unanchored hypothetical is not.
- **Deferred why.** If a direct why is used at all, use it last, on an episode already reconstructed, and treat the answer as a claim to be weighed against the sequence, not as the finding. Where the two disagree, the disagreement is the interesting thing, and per **K4 §4.1** it is reported rather than resolved.

**Step 5. Use critical incident technique when the objective is what actually happens at the moment of truth.** The protocol, in order. Define the incident class precisely and out loud ("a time when you needed help with the account and could not get it"), so that participant and researcher are describing the same class of event. Ask them to select a specific instance and anchor it in time and place ("when was it, where were you, what were you doing"). Elicit the sequence in the order it happened, without interrupting to interpret. Then walk it again for detail: what exactly they did at each step, what they saw, who else was involved, what they tried that did not work. Then elicit the outcome and the consequence. Then, and only then, ask what about the situation made the difference. Two or three incidents per participant is realistic in a depth interview; one is a story, and four is a summary. The technique's power is that it produces behaviour-level detail no attitude question reaches, and its cost is time: budget 8 to 12 minutes per incident and price it against the section allowance from 02.02.

**Step 6. Use laddering when the objective is the structure beneath a stated preference, and stop it before it invents one.** Laddering climbs from attribute to consequence to value: "you said it had to be the small pack, what does the small size do for you?", then "and what does not wasting any mean to you?" Three disciplines make it evidence rather than a party trick. Start from something the participant actually said, never from an attribute you supplied. Climb one rung at a time, using their own words in the next question. And stop when the answers become abstract virtues that could apply to anything ("I like value", "I want the best for my family"), because everything above that rung is the participant supplying the culturally expected top of the ladder, not reporting a structure. Two or three rungs is usually the honest limit. The output of a ladder is one participant's articulated chain, which is evidence of how they account for a preference, not proof of an unconscious hierarchy, and the analysis must say which it is treating it as.

**Step 7. Open with a grand tour question, then narrow.** A grand tour question asks the participant to walk the researcher through a domain they know and the researcher does not: "talk me through what happens on a normal Tuesday from when you get in", "walk me through what you did the last time an order went wrong, from the first thing you noticed". It works because it hands the participant the expertise, sets the expected answer length at "several minutes" rather than "a sentence", and surfaces vocabulary, sequence and unanticipated actors that no direct question would have asked about. Follow with mini tours into the parts that matter ("you mentioned checking with the depot, talk me through what that involves"). Two failure modes to avoid: a grand tour so broad the participant does not know where to start (anchor it in a time and a place), and a grand tour used where a specific incident is needed, which returns the routine rather than the exception.

**Step 8. Write the stem.** One idea. Open, so it cannot be closed with a word. Short enough to be said in one breath, because a long question teaches the participant that long is the register and they answer the last clause. In the participant's vocabulary, never the client's. Past or present tense, not conditional. No evaluative adjective, no candidate answer, no embedded assumption about what they did, felt or knew. No noun the participant has not used, unless introducing it is the deliberate purpose of the question and it sits below the exposure boundary. Then read it aloud: anything you stumble over will be misheard.

**Step 9. Make the undesirable answer easy to give.** For anything with a socially preferred answer, permission is designed into the question and cannot be retrofitted by a confidentiality assurance. Four devices, used in combination. **Normalise before asking:** "some people manage to keep on top of this and plenty of people do not". **Load the question toward the undesirable side:** "how many times have you had to put it off?" presupposes it happened, which is easier to correct than to confess. **Ask about the behaviour, not the identity:** people will describe missing three payments and will not accept being someone who misses payments. **Go third person where the direct route is closed:** "what do people you know say about this?" produces the range of sayable positions, from which the participant's own often emerges, but it is evidence about the norm rather than about the individual and must be labelled that way at analysis. Never combine an assurance of confidentiality with a question that makes the honest answer sound bad; the assurance signals that the topic is shameful.

**Step 10. Ask about the past without inviting reconstruction.** Three rules. Anchor to an event, not a period: "the last time" beats "in the last six months", and a landmark ("since you moved", "since the baby") beats a calendar date. Ask what happened before you ask what they thought, because an opinion elicited first gets retrofitted into the account. And never ask a participant to explain their past self's reasoning as though they had access to it ("why did you decide to..."): ask what they knew at the time, what options were in front of them, and what they did, and let the analyst reconstruct the decision. Where the event is old or the participant is visibly composing, say so in the fieldnotes; the analyst needs to know which accounts were retrieved and which were built.

**Step 11. Write the probe ladder for each question.** For each primary question, write the classes of probe that will deepen it and one worked example of each, per the four-level architecture in **02.02 Step 7**. Then test every written probe: no imported noun, no candidate answer, no evaluation, not answerable "yes". Add the two probes most guides omit: the evidential probe ("can you think of a time that happened?"), which converts a generalisation back into an episode, and the negative-case probe ("was there ever a time it went the other way?"), which is the cheapest available defence against a participant telling a consistent story for the researcher's benefit.

**Step 12. Test the question before fieldwork, on a person, out loud.** Three tests, in order. The **anticipation test**: write down the three answers you expect. If you can predict them, the question is closed, leading, or asking for something the participant will produce from the culture rather than from their life. The **empty answer test**: could someone with nothing to say still answer this fluently? If yes, it will not discriminate. The **cognitive test**: ask two or three people from the target population, then ask them what they thought the question meant and how they worked out their answer. This is the only method that reliably catches ambiguous terms, unanswerable recall windows and imported vocabulary, and it takes under an hour. Log the changes it produces, per the revision discipline in **02.02 Step 13**.

## 8. Analytical framework

Every question is built and checked on one chain:

    Section purpose → Evidence unit → Mental act required → Question form → Wording → Probe ladder → Codable material

**Forward** is the design test. This section needs the moderator to have heard this specific thing, which is this unit of evidence, which requires the participant to perform this mental act, which is reliably performed only in this question form, worded this way, deepened by these probes, producing material the planned analysis can code. If any arrow cannot be drawn, the question is not finished.

**Backward** is the elimination test, and the link that fails most often is the third. A question with a clear purpose and clean wording still fails if the act it demands is prediction, introspection or self-averaging, because the answer will arrive fluently and be worthless. The second most common break is at the last link: the question returns generalised opinion where the analysis needed episodes, and no coding effort can convert one into the other.

Nested inside the chain is a second distinction the analyst inherits: **what the participant did, what the participant says they did, and what the participant says it means**. These are three different kinds of evidence with three different reliabilities, and the question determines which one is collected. Per **K2 §3**, an instrument that blurs them hands the analyst material that cannot be labelled correctly downstream.

## 9. Output format

**1. Question set, by section**, in guide order. Each entry carries: section and purpose it serves, the question verbatim as it will be asked, the evidence unit it is written to produce, the mental act it requires, probe classes with one worked example each, and any prompt list marked "only if not mentioned".

**2. Question specification table**, for review.

| Q ref | Section | Purpose served | Evidence unit | Mental act | Question form | Risks noted |
|---|---|---|---|---|---|---|

**3. Rewrite log**, wherever a supplied or draft question was changed.

| Original question | Defect | Rewritten as | What is now collectable that was not |
|---|---|---|---|

**4. Removed questions**, with the reason and what would justify reinstating them.

**5. Sensitive-question notes.** For each, the framing device used, the residual risk, and the direction of the remaining bias.

**6. Cognitive testing notes**, where testing has happened: what was misread, by whom, and what changed.

**7. Open items and review points**, per **K5 §3**.

**When inputs are thin**, the format is not filled in anyway. A question whose evidence unit cannot be stated is marked as an open item, not smoothed into a plausible sentence. Vocabulary that has not been supplied is marked as a placeholder rather than invented, per **K4 §2.1**. No example answer is written alongside any question, ever: see Section 12.

## 10. Quality checks

Run before the question set is shared. These sit on top of **K4 §8**.

1. Every question states the evidence unit it is written to produce.
2. No question requires prediction of the participant's own future behaviour.
3. No question asks the participant to explain their own motivation directly, except as a deliberate deferred check on an already-reconstructed episode.
4. No question asks for a ranking or comparison the participant has not actually made.
5. No question asks for a count of a routine low-salience behaviour over a long window.
6. Every general-tendency question has been challenged, and the ones that survive are those where the participant's own summary is genuinely the object of study.
7. Every recall question is anchored to an event or a landmark, not only to a calendar period.
8. No stem contains two ideas, an evaluative adjective, a candidate answer, or an assumption about what the participant did.
9. No stem or probe introduces a noun the participant has not used, above the exposure boundary.
10. Every question passes the empty answer test: someone with nothing to say could not answer it fluently.
11. Every question passes the anticipation test, or the reason it is deliberately closed is recorded.
12. Every sensitive question carries a normalising or permissive frame, and the residual bias direction is stated.
13. Every primary question has probe classes with worked examples, including at least one evidential and one negative-case probe in each section.
14. Question wording uses participant vocabulary, evidenced from source, or is marked as unverified.
15. Nothing in the question set anticipates or illustrates what a participant will say.

## 11. Common failure modes

| Failure | How to recognise it | How to prevent it |
|---|---|---|
| **The fluent rationalisation** | Confident, tidy, well-structured answers to "why", all resembling each other | Route around why (Step 4). Treat any direct-why answer as a claim to weigh against the sequence, not as the finding |
| **The generous forecast** | "Would you use it?", "how likely are you to...?", answered enthusiastically by everyone | Never ask for prediction. Ask what happened last time and what would have had to be different |
| **The self-averaging question** | "Usually", "typically", "on average", "how often do you normally" | Ask about the last occasion. Prevalence is an analysis output, not a participant's job |
| **Manufactured preference** | A ranking of five things the participant has never held in mind at once | Ask only about comparisons faced in life. Real trade-offs need a designed exercise in category 06 |
| **Imported vocabulary** | The participant adopts the client's word back and the transcript looks like agreement | Write in participant vocabulary from evidence; watch for the moment a term first enters the transcript and check who said it |
| **The buried assumption** | "How did the delay affect you?" asked of someone who never mentioned a delay | Every stem checked for what it presupposes the participant did, felt or noticed |
| **The two-part question** | Two ideas in one sentence; the participant answers the second | One idea per stem, always. Long questions get answered from the end |
| **The confessional trap** | A sensitive question asked directly, after a confidentiality assurance, with the socially bad answer clearly marked | Normalise, load toward the undesirable answer, ask about behaviour not identity |
| **Ladder to nowhere** | The chain terminates in "I want what is best for my family" and is written up as a value | Stop the ladder when the answers stop being specific to this person's life. Two or three rungs is the honest limit |
| **The unanchored past** | "In the last year, how often did you..." for something unmemorable | Anchor to an event or landmark, shorten the window, or ask about the most recent occasion |
| **AI: symmetric question sets** | Every section given the same four questions in the same shapes, evenly balanced, none removed | Questions come from the evidence unit, which differs by section. A set with nothing rewritten or removed has not been developed |
| **AI: the plausible participant voice** | Generated example answers, illustrative quotes, or "participants are likely to say" alongside the questions | **K4 §2.2, §2.3**. A question set is written before anyone has spoken. Anticipated answers migrate into analysis documents and become indistinguishable from evidence |
| **AI: invented vocabulary** | Questions written in confident category language, product names or terms that were never supplied | **K4 §2.1, §6.2**. Mark vocabulary as a placeholder and name what would confirm it |
| **AI: the polished why** | A generated guide full of well-worded "why do you feel that way" questions, which read as good practice and are the core defect | The prohibition is on the mental act, not on the phrasing. A beautifully worded introspection request is still an introspection request |
| **Wording used to fix a structural problem** | Ever more careful phrasing applied to a question that is in the wrong place in the sequence | Order is irreversible and wording is not. Send it back to **02.02** |

## 12. AI guardrails

Skill-specific only. **K4** applies in full and is not repeated here.

1. **Never generate an example answer, an illustrative quote, or a description of what participants will probably say, alongside a question set.** This is the highest-risk output in this skill. Anticipated responses written during design travel into analysis documents and become indistinguishable from evidence (**K4 §2.2, §2.3**).
2. **Never write a question that requires an act from the bottom five rows of the Step 2 table**, however well phrased, and however clearly the stakeholder asked for it. Explain what the answer would actually measure and offer the substitute (**K4 §9**).
3. **Never invent the participant's vocabulary.** Category terms, product names, colloquialisms and the words a population uses for its own experience come from supplied material, prior transcripts or documented sources. Where absent, mark as a placeholder needing verification (**K4 §2.1, §6.2**).
4. **Never present a laddering output or a projective response as an established motivation.** It is an articulated chain from one person, at one moment, in one interview. The interpretive step is the analyst's and must be visible (**K2 §3**).
5. **Never quietly rewrite a stakeholder's question.** Every change goes in the rewrite log with the defect named and what the rewrite now makes collectable. Silent improvement is unfalsifiable and removes the stakeholder's chance to say the original mattered.
6. **Never claim a question is neutral or unbiased.** Report the checks run and what they found. Residual risks, including a framing that could not be avoided and a sensitive question whose remaining bias has a known direction, are disclosed with that direction stated.
7. **Never generate a question for a sensitive topic without stating the residual bias and the duty-of-care implication.** A well-designed sensitive question reduces misreporting; it does not eliminate it, and it may still distress.
8. **Never treat wording as a remedy for a sequencing, sampling or method error.** Where the real defect is upstream, say which skill owns it rather than producing a better sentence in the wrong place.

## 13. Best-practice principles

1. **Ask people what they did, not what they think.** Behaviour is remembered, reasons are constructed. The reasons are worth having, but they are evidence about how a person accounts for themselves, not about what happened.
2. **The most reliable question in qualitative research is "tell me about the last time".** It converts an unanswerable question about tendency into an answerable one about an event, and it yields detail the researcher did not know to ask for.
3. **A fluent answer is not a good answer.** Fluency signals availability, not accuracy, and the most fluent answers are usually the most culturally rehearsed ones.
4. **If you can predict the answer, do not ask the question.** Predictability means you are collecting the culture, not this person's life.
5. **Every noun in your question is a gift to the participant.** They will use it back, and you will read your own word in the transcript and mistake it for their concept.
6. **Permission beats assurance.** "Plenty of people find this hard" collects more truth than any promise of confidentiality, which mainly signals that the topic is shameful.
7. **A hypothetical anchored to a real event is answerable; a free-floating one is not.** "What would have had to be different that day?" works. "Would you buy it?" does not.
8. **Ask about the exception as deliberately as the rule.** The negative-case probe is the cheapest defence against a participant building a consistent story for you.
9. **Short questions get long answers, and long questions get short ones.** A question that takes fifteen seconds to ask teaches the participant that the researcher likes talking.
10. **Silence is a question.** The three seconds after an answer is where the qualification, the exception and the real version usually arrive.
11. **The interview is where the questions get tested; the guide is only the hypothesis.** Cognitive testing before fieldwork and honest inspection after the first two sessions catch more defects than any amount of desk redrafting.
12. **Distinguish what they did, what they say they did, and what they say it means, at the point of writing the question.** If the instrument blurs these three, the analyst cannot separate them later, and the report will present accounts as behaviour.

## 14. Worked example

**INPUT**

A fictional software firm, Ledgerline, sells expense management software to mid-sized companies. Its product team can see from usage data that a large share of employees photograph a receipt and then never complete the submission. They cannot see why. Method already selected: 45-minute remote depth interviews with 14 employees at customer companies who have abandoned at least one submission in the past month. The guide (**02.02**) has three sections: how expenses fit into the working week, the abandoned submission, and reaction to a redesigned flow behind an exposure boundary at minute 32. The stakeholder-supplied question list arrives with fourteen questions.

**PROCESS**

*Step 1, evidence units.* Section 2's purpose is "the moderator should have heard one specific abandoned submission, reconstructed from the moment the receipt was photographed to the moment the participant stopped, with what they were doing around it". The evidence unit is therefore an episode with a sequence and an interruption point, not an opinion about the software. This single line eliminated four of the fourteen supplied questions immediately, because none of them could produce an episode.

*Step 2, mental acts.* The supplied list classified: "Why do you find expense submission frustrating?" requires introspection and presupposes frustration. "How often do you submit expenses?" requires counting a routine low-salience act. "How important is it to you that the process is quick?" requires rating an attribute never compared against anything. "Would you use a feature that reminded you?" requires prediction. All four are the reliable-looking, unusable kind: every participant would answer all of them without hesitation.

*Step 3 and 7, the rewrite.* The section now opens with a grand tour anchored in time: "Talk me through what happened the last time you took a photo of a receipt and did not finish putting it in. Start from where you were when you took the photo." This asks for retrieval, sets the answer length, and surfaces the surrounding context (in a taxi, at a client's reception desk, at the end of a day) which turned out to matter more than anything about the interface.

*Step 4, the judgement call.* The product lead objected to the removal of "why do you find it frustrating", on the grounds that the team needs to know the reason and the question asks for it directly. The objection was answered rather than overruled, and the reasoning recorded. A direct why returns the most available explanation, and for software the most available explanation in every population is "the app is slow" or "it is fiddly", because that is the culturally supplied account of all software friction. It would have been returned confidently by most of the fourteen and would have sent the team to optimise load times. The route taken instead was sequence and contrast: "what were you doing just before you opened it?", "what happened next?", "was there a time you did finish one, and what was different about that day?" Recorded as a rewrite-log entry rather than a silent deletion, so the product lead could see what was traded.

*Step 6, laddering, and where it was stopped.* One participant said the submission had to be done "properly". The ladder ran two rungs: what does doing it properly get you (not having it queried by finance), and what would being queried mean (a conversation with a manager about a personal purchase on a shared card). It was stopped there. A third rung would have arrived at "I want to be seen as trustworthy", which is true of almost everyone and specific to no one, and would have been a value the analyst supplied rather than a structure the participant reported. Logged as a two-rung chain from one participant, labelled as that participant's articulated account per **K2 §3**.

*Step 9, sensitivity.* One objective touched a socially managed area: whether people ever gave up on a claim and absorbed the cost themselves, which implies both disorganisation and a small personal loss. Asked directly, this returns denial. Asked as "how many expenses would you say you have ended up just writing off in the last few months?", it presupposes the behaviour and makes the honest answer easy, with a normalising preamble ("almost everyone we speak to has a couple they never got round to"). The residual bias was recorded as likely under-reporting, direction known, size not.

*Step 12, testing.* Three cognitive tests found that "submission" was the internal word: participants said "putting it through" or "claiming it". Every stem was rewritten in that vocabulary. One participant also read "the last time you did not finish" as an invitation to describe a technical error, so the stem was anchored more firmly to their own action rather than the system's.

**OUTPUT**

A question set of nine primary questions across three sections, with probe classes and worked examples including an evidential probe and a negative-case probe in each section; a rewrite log with six entries naming the defect in each supplied question and what the replacement makes collectable; a removed-questions list with four entries; sensitive-question notes recording the residual under-reporting on written-off claims; cognitive testing notes covering the vocabulary change; and one **K5** review point, on whether interviews conducted by the product team itself, whose members participants may recognise, will suppress criticism enough to warrant an independent moderator.

## 15. Advanced usage

**Questions for automated and asynchronous interviewing.** When no moderator is present, the probe cannot be selected from what was just said, so the entire probe architecture must be pre-specified with trigger conditions: if the answer is under a certain length, probe for elaboration; if it names an artefact, probe for specification; if it contains a generalisation, probe evidentially for an episode. This changes the question design itself. Prefer stems that are self-anchoring ("the last time" rather than "when you do this"), avoid anything requiring clarification, and expect answers to be composed rather than spontaneous, which raises the register and reduces the exceptions people mention. Hand over to **03.02 AI-Moderated Interview Design** for the specification format.

**Expert and elite interviews.** The usual balance inverts. The participant knows more than the researcher, controls the time, and is practised at delivering a prepared account. Grand tour questions still work and are often the only thing that gets past the prepared account, but the critical incident protocol has to be applied gently and specifically, and the most productive questions are frequently the ones that ask for the exception ("when did that approach not work?") or the mechanism ("who actually signs that off?"). Prediction questions are more tempting here and no more reliable: an expert's forecast is an opinion with better vocabulary.

**Cross-language question development.** Translate the mental act, not the sentence. Recall anchoring is tense-dependent and some languages handle "the last time" differently; normalising preambles carry different social weight; third-person framing is more effective in some cultures and reads as evasive in others; and directness itself is a variable, so a question that is neutral in one market is rude in another and gets a managed answer. Develop the question in each language with a local researcher against the evidence unit, then back-check that the unit is still what comes back.

**Repairing a live study.** When the first interviews return generalisations, the cause is almost always one of three things: the question asked for a tendency, the question asked for a reason, or the moderator moved on after the first answer. Diagnose by reading two transcripts and marking every point at which a general answer went unprobed. The repair is usually one evidential probe added in each section, and it can be made between sessions, logged per the revision discipline in **02.02 Step 13**.

## 16. Skill chain

**Recommended previous skills**
- **02.02 Discussion Guide Design.** Hands over the section brief: purposes, exit conditions, sequence, time allowances and the exposure boundary, which are the frame this skill writes inside.
- **01.07 Analysis Plan Development.** Hands over the unit the analysis needs, which determines the evidence unit each question must produce.

**Recommended next skills**
- **02.04 Question Bias Detection.** Audits the drafted questions independently. Writing and auditing should not be the same pass.
- **03.02 AI-Moderated Interview Design.** Takes the questions and probe classes and converts them into a specification an automated interviewer can execute.

**Runs well alongside**
- **07.03 Interview and Transcript Analysis**, downstream, which inherits the evidence units as its coding targets and can tell whether the questions delivered them.
- **13.05 Research Ethics and Consent Design**, wherever the questions reach sensitive, stigmatised or distressing material.
- **02.01 Survey Questionnaire Design**, where qualitative questioning is generating the vocabulary and option lists a later survey will need.

---
A Yazi Supplied Skill and resource.
