---
name: discussion-guide-design
description: >
  Designs discussion guides and topic guides for moderated qualitative research:
  depth interviews, paired depths, focus groups and workshops. Use when someone
  says "write a discussion guide", "draft a topic guide", "we need a focus group
  guide", "build the interview guide", "design the moderator guide", or asks how
  many questions fit in an hour, what order to put topics in, where to show the
  stimulus, how to write probes that do not lead, or how to stop one person
  dominating a group.
category: 02 Instrument Design
ref: 02.02
tier: 0
inherits: [K2, K3, K4, K5]
---

# Discussion Guide Design

## 1. One-line description
Turns research objectives into a flexible instrument a moderator can actually run: a small number of sections, each with a stated purpose, a defended time budget, an opening question and a probe architecture, ordered so that nothing later contaminates anything earlier.

## 2. What this skill is used for

**The research problem it solves.** The commonest failure in qualitative instrument design is category error: the guide is written as a questionnaire in prose. It contains thirty or forty numbered questions, each phrased as a complete sentence to be read aloud, covering every topic anyone mentioned. In the room it produces an interrogation. The moderator, conscious of the list, asks and moves on, asks and moves on, and never asks the second question, which is where qualitative research actually happens. The transcript is wide and shallow: forty topics touched, none reconstructed, no episode described in enough detail to be analysed. The opposite failure is rarer but equally costly: a guide so loose that four moderators run four different studies and nothing can be compared. Between them sits a third failure, sequencing, which is irreversible. A stimulus shown in minute ten cannot be unshown, and everything after it is a reaction to a document rather than an account of a life.

**Where it sits in the research lifecycle.** After the method has been chosen and the sample defined, before recruitment screening is finalised and before fieldwork. It is the last point at which the shape of the eventual analysis can still be changed cheaply.

**Typical use cases.**
- Building a depth interview guide from an agreed objectives list.
- Designing a focus group guide, including the exercises and the dynamics management.
- Converting a stakeholder topic list into a guide that fits the session length.
- Designing wave 1 of a longitudinal or repeat-contact study, where the spine must hold across waves.
- Adapting a guide across markets and languages without losing comparability.
- Rebuilding a guide after the first two sessions showed it was asking the wrong things.

**Who uses it.** Qualitative researchers and moderators writing their own guides; research managers commissioning fieldwork and reviewing what comes back; UX and product researchers running interviews without a qualitative specialist nearby; client-side insight managers who need to challenge a guide before it is fielded.

## 3. When to use it

- The method is settled as moderated qualitative and a guide now has to exist.
- A session length has been sold or booked and the topic list is visibly too long for it.
- Stimulus (concept, prototype, pack, ad, document, price) will be shown, and where it sits in the session has not been decided.
- The subject matter is sensitive, stigmatised or socially regulated, and placement and framing will determine whether anyone tells the truth.
- Several moderators will run the same study and their sessions must be comparable.
- The study is multi-market or multi-language and one guide has to survive translation.
- The study involves repeat contact with the same participants and the guide must work more than once.
- The first sessions have run and the guide is not producing what the objectives need.

## 4. When NOT to use it

- **The objectives are not settled.** A guide cannot resolve an unclear research question. It will encode the confusion into a conversation and produce a transcript that is interesting and unusable. Go to **01.01 Research Brief Interrogation** or **01.02 Business Problem to Research Question** first.
- **The objective is prevalence, magnitude or share.** Moderated qualitative produces range, vocabulary, mechanism and account. It does not produce "how many". A guide written to answer a counting question yields numbers with a base of twelve that will be quoted in a board pack as though they were measurement. Use **02.01 Survey Questionnaire Design**. Where both are needed, design the qualitative to generate the vocabulary and hypotheses and the quantitative to size them.
- **The method has not been chosen, or was chosen by habit.** "We always do three groups" is not a design. Whether the instrument should be a group, a depth, a dyad, an ethnographic visit or no primary research at all belongs to **01.04 Research Method Selection**. Writing a guide for the wrong instrument produces a well-crafted wrong study.
- **A group is proposed for material that requires privacy.** Debt, health, shame, illegality, family conflict, workplace grievance, anything the participant manages the presentation of in front of peers. Group composition cannot fix this, and moderator skill cannot fix it. The correct response is to change the instrument, not to soften the guide.
- **The behaviour can be watched.** Where the objective is how a task is actually performed, asking someone to narrate it produces a rationalised, tidied account. Observation, usability testing or behavioural data measures it directly. Reserve the guide for what observation cannot reach: what they considered and rejected, what happened before, what it meant.
- **You are writing or auditing individual question wording.** For question craft, **02.03 Interview Question Development**. For a bias audit of an existing guide, **02.04 Question Bias Detection**. Running a design skill over a finished guide tends to produce an unrequested rewrite.
- **The session will be AI-moderated.** The probe architecture must then be specified in advance rather than exercised in the room, which is a different design problem with its own failure modes. Design the objectives and section purposes here, then hand over to **03.02 AI-Moderated Interview Design**.
- **The session is a facilitated workshop with a delivery outcome rather than a research objective.** Alignment sessions, ideation and co-creation with stakeholders are facilitation. They can be valuable. They do not produce research evidence, and a guide that treats them as though they do launders stakeholder consensus into a findings deck.
- **The topic requires clinical competence the moderator does not have.** Trauma, active mental health difficulty, bereavement, abuse. A guide cannot make an unqualified moderator safe to run that session, for the participant or for the moderator. See **13.05 Research Ethics and Consent Design**, and stop until the ethical and duty-of-care structure exists.
- **The purpose is to collect quotes supporting a conclusion already reached.** No design fixes this. Say so.

## 5. Required inputs

**Required. Without these the skill cannot run.**
- **Research objectives, and for each one the decision it informs.** If decisions are not supplied, ask once. If no answer comes, proceed and mark the objective `[decision not supplied]`, which flags every section hanging off it as a candidate for cutting when time runs short.
- **Instrument and session length.** Depth, paired depth, group or workshop; and the number of minutes. These are the two facts that determine how much can be covered. Without them, stop and ask: a guide written for an unspecified length is a topic list, not an instrument.
- **Participant definition.** Who is in the room, what they have in common, what varies, and what they have been told the session is about. Recruitment framing constrains the opening: a guide cannot ask an unaided question about a subject the invitation already named.

**Optional, and what each one adds.**
- **Stimulus, in the form participants will see it.** Determines the exposure point, the fidelity warning the moderator must give, the rotation plan, and how much time the reaction genuinely needs. Without it, the exposure point can be marked but the section cannot be built.
- **Moderator identity and experience level.** Determines how much of the probe architecture must be written out. An experienced moderator on familiar subject matter needs probe classes. A junior moderator, or a subject the moderator does not know, needs worked probe examples.
- **Number of moderators and sessions.** More than one moderator raises the comparability requirement and changes what must be fixed versus free.
- **Analysis approach (07.03).** Tells you what unit the analysis needs: episodes, decision journeys, attitudinal positions, language. A guide that does not collect the unit the analysis needs produces transcripts that resist coding.
- **Previous wave, or the same participants' prior sessions.** Enables the carry-forward brief and the fixed spine, and identifies where panel conditioning has to be managed.
- **Market and language list.** Determines which elements are fixed for comparability and which are localised.
- **Known constraints on the room.** Observers behind glass or on a video link, client attendance, recording restrictions, an interpreter, a venue with a hard stop. Each changes what the guide must tell the moderator.

## 6. Questions to ask before starting

1. **What is the decision this research feeds, and which objectives actually bear on it?** Objectives that do not bear on a decision are the first thing cut when the time budget bites, and it is better to cut them now, in writing, than in minute fifty of a session. Default if unanswered: build the guide, rank sections by stated importance, and mark the ranking as an assumption.
2. **How long is the session, and is that number negotiable?** Length is the binding constraint on depth. A 45-minute interview and a 90-minute interview are different instruments, not the same instrument at different speeds. Default: assume 60 minutes for a depth and 90 for a group, and state the assumption.
3. **Will stimulus be shown, and what must be measured before it is?** This is the irreversible decision in the guide. Everything spontaneous has to happen first. Default: assume any stimulus is shown no earlier than two thirds of the way through, and flag it.
4. **How sensitive is this material for these participants, and what will they manage the impression of?** Determines placement, framing, whether a group is viable at all, and what the moderator must be told to do if someone becomes distressed. Default: treat money, health, family, employment status, care responsibilities, immigration, addiction and legal trouble as sensitive, and place accordingly.
5. **Who is watching or listening, and do participants know?** Observers change what is said. So does a client who intervenes. Default: assume observers are present, require that participants are told, and write an observer protocol into the moderator layer.
6. **Will these participants be seen again?** Repeat contact changes the opening, the consent, and what the moderator may reveal. Default: assume single contact, and flag re-contact consent as an open item if longitudinal work is plausible.
7. **Who is moderating, and how many of them?** Determines how prescriptive the guide must be and how much has to be fixed for comparability. Default: assume a single competent moderator unfamiliar with the category, and write probe classes with two worked examples each.

## 7. Step-by-step methodology

**Step 1. Convert each objective into a section purpose.** For every objective, write one sentence beginning "By the end of this section the moderator should have heard...". This is the governing move of the whole skill. "Explore attitudes to switching" is not a purpose; it is a topic, and a topic gives the moderator nothing to steer by. "By the end of this section the moderator should have heard at least one specific occasion on which the participant considered switching and did not, reconstructed in enough detail to name what stopped them" is a purpose. It tells the moderator what a successful section sounds like, which is the only way they can know whether to probe again or move on. An objective that cannot be turned into a hearable purpose is not ready to be researched, and should be raised rather than designed around.

**Step 2. Build the section brief.** This is the spine artefact and it is built before any question is written. One row per section, with these columns: section name, objective ref, purpose (the sentence from Step 1), minutes allocated, entry condition (what must already have happened), exit condition (what the moderator must have heard before moving on), and cut priority (the order in which sections are shortened if the session runs late). Two rules govern it, and they are the discipline of the skill. **No section without an objective:** a section with an empty objective column is deleted or its objective is agreed and added. **No objective without a section:** an objective with no section is unresearched, and unresearched objectives are how a qualitative study fails silently, because the transcripts are full and nobody notices what is missing until analysis. Build the brief before drafting questions. Researchers who draft first almost always rationalise the questions they have already written.

**Step 3. Budget the time honestly, then cut to fit.** Qualitative time arithmetic is unforgiving and routinely ignored. Work in this order, using planning conventions and labelling them as such rather than as measured norms. From the session length, subtract the fixed overheads: introduction, consent and settling (roughly 4 to 6 minutes in a depth, 8 to 12 in a group, more if there is an exercise to explain), and the close (3 to 5 minutes, which must include the "anything we have not covered" question because it is frequently where the best material arrives). What remains is the working time. Then price the content. A primary question explored to genuine depth, with three or four probes and the pauses that let a participant think, takes 4 to 8 minutes, not 30 seconds. A reconstructed episode ("walk me through the last time") takes 8 to 12 minutes and cannot be rushed without becoming a summary. So a 60-minute depth has roughly 48 minutes of working time and supports four to six primary lines of questioning, which means eight to twelve questions the moderator will actually ask, not forty. A 45-minute depth supports two to four. A 90-minute session supports six to eight, and no more, because attention degrades in the last twenty minutes and the material collected there is thinner. In a group the arithmetic is harsher because airtime divides. A 90-minute group of six has roughly 70 minutes of working time, which is under 12 minutes of speaking time per participant even if distribution were even, and it never is. Any go-round costs its per-person time multiplied by the number of participants plus moderator linking: a "one minute each" round with six people is 8 to 10 minutes. Any exercise costs explanation, doing and debrief, and the debrief is the data, so an exercise with a 5-minute task is a 15-minute section. Price every exercise this way before deciding to keep it. When the total exceeds the session, cut whole sections in cut-priority order and record what was cut and which objective is now unaddressed. Do not absorb the overage by shortening every section, which converts one good study into six shallow ones.

**Step 4. Set the funnel.** Order sections broad to narrow, and inside each section order the questioning broad to narrow. Four sequencing rules, in priority order. **Unaided before aided**, without exception: once a brand, feature, competitor, reason or attribute has been named by the moderator, spontaneous salience of it cannot be recovered in that session. **Spontaneous before prompted**: the prompt list exists to check coverage after the participant has finished, not to structure their answer. **General before specific**: the overall account comes before the diagnostic, because a specific failure raised first re-frames everything after it. **Behaviour before attitude** on the same object: attitudes reported after an episode has been described tend to be rationalised into consistency with the episode, which is at least visible; attitudes reported first contaminate the episode. Then read the sequence once more asking two questions of every section: what does this teach the participant that changes the next section, and what does it tell them about what the researcher wants to hear?

**Step 5. Design the opening.** The opening does four jobs and the guide must name all four: set the contract (who you are, who the research is for at the level consent covers, how long, recording, observers, that there are no right answers, that they can decline any question or stop), establish that talking at length is the expected behaviour, calibrate the moderator to this participant's vocabulary and pace, and collect context the later sections will need. The warm-up question should be easy, concrete, about them, and genuinely relevant to the subject without naming the specific thing under study. "Tell me about a normal weekday morning in your house" earns its place before a study of breakfast. "How do you feel about breakfast cereal" does not: it is the study, asked cold, and it will get an answer the participant assembles on the spot. Two failures to avoid: a warm-up so long it eats the working time, and a warm-up so generic it teaches the participant that short answers are acceptable.

**Step 6. Write the primary questions only.** Write the question the moderator will actually say to open each line of enquiry. Open, short, one idea, in the participant's vocabulary, and anchored to something real. Prefer the episodic over the general: "tell me about the last time" outperforms "what do you usually do", because usual behaviour is a summary the participant constructs and the last time is an event they can reconstruct. Prefer the concrete over the abstract, the descriptive over the evaluative, and the past over the future. Do not write questions that ask people to predict their own behaviour, and do not write questions that ask them to explain their own unconscious motivation. Both are covered in Step 7 and Section 11. Aim for the number Step 3 allows, and if you have more, you have not finished cutting.

**Step 7. Build the probe architecture.** A guide has four levels and confusing them is what produces both the forty-question guide and the guide that cannot be moderated.

| Level | What it is | Written in the guide as |
|---|---|---|
| **Topic** | An area of enquiry owned by an objective. Never spoken aloud. | A section heading with its purpose |
| **Question** | What the moderator says to open the line. Broad, open, spontaneous. | Verbatim, because the wording matters for comparability |
| **Probe** | A follow-up whose content is determined by what the participant just said. | A class, with one or two worked examples. Never a script |
| **Prompt** | An area raised only if it has not come up spontaneously, always after the unaided exploration is exhausted. | A list, explicitly marked "only if not mentioned" |

Probes are the interview. Write them as classes so the moderator selects rather than reads: elaboration ("tell me more about that part"), specification ("what did that actually look like"), clarification ("when you say it was a hassle, what was the hassle"), sequence ("what happened just before that"), contrast ("how was that different from the time before"), consequence ("and then what happened", "what did that mean for you"), evidential ("can you think of a time that happened"), echo (repeating the participant's own word as a question), and silence, which should appear in the guide as a permitted and expected move rather than an accident. Then test every written probe against four rules. It must not introduce a noun the participant has not used. It must not offer a candidate answer ("was it because it was expensive?"). It must not carry an evaluation, because repeated interest in one direction teaches the participant which answers are rewarded. And it must not be answerable "yes" by someone with nothing to say. Balance matters as much as wording: a guide that probes positives twice and negatives once is leading through allocation of attention, even if every individual probe is neutral. Where the client has a hypothesis, it belongs in the prompt list at the end of the section, never in a probe.

**Step 8. Place the sensitive material.** Sensitive sections go late, after rapport exists and after everything that would be damaged by a break-off or a withdrawal has been collected, and before the participant is too tired to do the work. In a 60-minute depth this is usually the third quarter, not the last five minutes, because sensitive material needs recovery time and the guide must budget for it: a closing section that returns to safer ground is part of the duty of care, not padding. Frame with normalisation rather than reassurance ("some people find this straightforward, others find it a real struggle") because permission collected in the question outperforms confidentiality promised in the preamble. Give the moderator an explicit escalation instruction: what to say if the participant becomes distressed, what to do if they disclose harm, and that stopping is always available and never a failure. In a group, sensitive material is a signal to re-examine the instrument, not to move the section.

**Step 9. Place the stimulus and mark the exposure boundary.** Draw a hard line in the guide at the point of first exposure and label it. Everything above the line is pre-exposure data and can never be collected again. Everything below is a reaction. Check the section brief above the line for anything that must be pre-exposure and has drifted below it: spontaneous vocabulary, current behaviour, unaided awareness, the existing mental model, the problem as the participant frames it. Then set four things. Fidelity: the moderator must state what the participant is looking at and what it is not ("this is a rough idea, not a finished product"), because absent that, participants critique the artwork. Rotation: where more than one stimulus is shown, rotate order across sessions and record the order used in each, otherwise a concept effect and an order effect are indistinguishable at analysis. Individual response before discussion, in a group: a written or private first reaction before anyone speaks, because the first spoken reaction anchors the room. Removal: whether the stimulus stays visible, because a visible stimulus keeps pulling the conversation back to itself.

**Step 10. Decide on projective and enabling techniques, and make them earn the time.** Distinguish the two. **Enabling techniques** help a participant articulate something they hold but cannot easily put into words: card sorts, laddering, timeline and journey mapping, ranking with a forced trade-off, drawing a process. **Projective techniques** approach material indirectly because a direct question would be answered normatively or not at all: sentence completion, third-person framing ("what would people you know say about..."), personification, collage, obituary and letter-writing tasks. A technique earns its place when three things are true: a direct question has been tried or can be predicted to produce flat, normative or socially managed answers; the technique reaches material the direct route cannot; and there is an interpretation rule agreed before fieldwork saying what the output means. Where any of these is missing, the technique is decoration, and decoration in a 60-minute session costs a whole line of questioning. Two disciplines govern their use. First, the output is not the finding: the collage, the completed sentence and the sorted cards are prompts for talk, and the talk is the data. Budget the debrief as the larger half of the exercise. Second, a projective output never establishes an unconscious motivation. It generates material the analyst interprets, and per **K2 §3** that interpretive step must be visible in the output rather than presented as something the participant said.

**Step 11. If it is a group, design the dynamics rather than relying on the moderator.** A group does not produce an aggregate of individual views. It produces a socially negotiated account, and that is its distinctive value: the range of positions available, the vocabulary people use with each other, what is sayable and what gets laughed at, and how a view moves under challenge. It is a poor instrument for the private, the counter-normative, the individual decision journey, and anything from which prevalence will be inferred. Design three things into the guide. **Against dominance:** private written response before open discussion, timed go-rounds, direct nomination by name, splitting into pairs for five minutes, and physical artefacts that occupy hands and equalise turn-taking. These are structural remedies that work; asking the moderator to "manage" a dominant participant in the moment is not a design. **Against premature convergence:** private pre-commitment on paper before the first spoken view, an explicit invitation to disagreement written into the guide as a question ("who sees this completely differently"), a designed slot for the minority position, and a moderator instruction to give no signal of the preferred answer. **For composition:** homogeneity on the dimension that governs candour, and heterogeneity where contrast is the point. Never mix power or status on a dimension that matters to the topic (managers with their own staff, parents with their own teenagers), because the lower-power participant will manage every answer. For paired depths, write the guide for the dyad and include at least one moment where each participant answers individually, and record in the limitations that neither will say what they would not say to the other.

**Step 12. Write the moderator layer.** A guide contains two documents interleaved: what is said to the participant, and what only the moderator sees. The second is the part most guides omit and the part that determines whether fieldwork survives contact with reality. It must contain: the purpose and exit condition for every section; the time allocation and the cut-priority order for when the session runs late; what has and has not yet been revealed, so nothing is said early; stimulus handling instructions; the escalation instructions from Step 8; what to do in the recurring awkward situations (the participant has never used the category, asks what the client wants, asks the moderator's own opinion, names a competitor, or turns out not to match the recruitment criteria); the observer protocol; and any recruitment claim that must be verified in the room. Typographically separate the two so that a moderator glancing down mid-sentence never reads an instruction aloud.

**Step 13. Pilot the guide, then revise it in field and record the revision.** Run one or two sessions and inspect four things: whether each section reached its exit condition, whether the timings held, which questions produced short or defensive answers, and where the participant's own vocabulary differed from the guide's. Then revise. A guide is a hypothesis about the conversation, and fieldwork is the test. Treating a guide as fixed once fieldwork has shown it is wrong is not rigour, it is the opposite, and it wastes the remaining sessions. Revision has one condition: it is logged, with what changed, when, why, and which sessions used which version, so the analyst knows whether a difference between session three and session eight is a participant difference or an instrument difference. Where a trended or multi-market study depends on comparability, revision to the fixed spine is a researcher decision under **K5 §2.7**, not an editing decision.

## 8. Analytical framework

The guide is built on one chain, applied to every section and readable in both directions:

    Objective → Section purpose → Opening question → Probe ladder → Exit condition → Analysis unit

**Forward** is the design test. This objective requires hearing this specific thing, which is opened by this question, deepened by these probe classes, and complete when this exit condition is met, producing this unit of material for analysis (an episode, a decision sequence, a vocabulary set, a range of positions). If any arrow cannot be drawn, the section is not finished.

**Backward** is the elimination test. This section produces this material, which supports this analysis, which answers this objective. Sections most often fail on the last two links: they produce interesting talk that no planned analysis uses, or they produce generalised opinion where the analysis needed an episode.

Nested inside the chain is the four-level vocabulary from Step 7: **Topic → Question → Probe → Prompt**. Confusing these levels is the single most reliable diagnostic of a badly designed guide. A guide with forty items has usually written every probe and every prompt as a question. A guide that cannot be moderated has usually written the topics as questions and left the probes to chance.

The section brief holding the chain is a live document. It goes to the moderator, it is updated when the guide is revised in field, and at analysis it becomes the map from every theme back to the objective that justified collecting it. Per **K2**, traceability starts here: a finding can only be traced to a section that existed for a reason.

## 9. Output format

**1. Design summary.** Instrument, session length, participant definition, number of sessions and moderators, objectives, mode, and the constraints applied. States the recruitment framing, so a reader can see what participants already knew.

**2. Section brief.**

| Section | Obj ref | Purpose ("by the end, the moderator should have heard...") | Minutes | Entry condition | Exit condition | Cut priority |
|---|---|---|---|---|---|---|

Every objective appears at least once. An objective with no section is shown with the section columns empty and marked `UNADDRESSED`, which is a defect to be resolved, not a formatting state to be left alone.

**3. The guide itself**, in session order, with the moderator layer visually distinct from the spoken layer. Each section carries: heading, purpose and exit condition (moderator only), time allocation and elapsed-time marker, the opening question verbatim, probe classes with worked examples, a prompt list marked "only if not mentioned", exercise or stimulus instructions where applicable, and any handling instruction specific to that section.

**4. The exposure boundary**, marked as a single labelled line in the guide, with a note of everything that must be complete above it.

**5. Timing plan.** Cumulative elapsed time at each section boundary, so a moderator can tell at a glance whether they are behind, plus the cut-priority order.

**6. Moderator briefing notes.** Handling instructions, escalation and duty-of-care instructions, observer protocol, stimulus rotation schedule and how to record which order was used, and anything to be verified in the room.

**7. Cut log.**

| Topic proposed | Source | Why cut | Objective affected | Reinstate if |
|---|---|---|---|---|

**8. Open items and review points.** Per **K5 §3**, at the point of the decision, with a consolidated list at the front.

**When inputs are thin**, the format is not filled in anyway. An objective with no stated decision is written `[decision not supplied]`. A stimulus that has not been supplied produces a marked exposure point and an empty section, not an invented concept. A purpose that cannot be written because the objective is unclear is recorded as an open item, not smoothed into a plausible sentence. See **K4 §1**: the presence of a row is not a reason to fill it.

## 10. Quality checks

Run before the guide is shared. These sit on top of **K4 §8**, which runs anyway.

1. Every section states, in one sentence, what the moderator should have heard by the end of it.
2. Every section traces to an objective, and every objective traces to at least one section.
3. The number of primary questions is consistent with the session length under the Step 3 arithmetic, and the arithmetic is shown rather than asserted.
4. Section timings sum to the session length including introduction and close, and a cut-priority order exists.
5. No unaided or spontaneous question sits below the exposure boundary, and nothing above the boundary names a brand, feature or reason the participant has not raised.
6. Every probe passes the four tests in Step 7: no imported noun, no candidate answer, no evaluation, not answerable "yes".
7. Probing attention is balanced across positive and negative territory, not just individually neutral.
8. No question asks a participant to predict their own future behaviour, or to explain their own unconscious motivation directly.
9. Prompt lists are marked "only if not mentioned" and sit at the end of their section.
10. Sensitive sections sit late but not last, carry a normalising frame, and are followed by recovery time.
11. Every exercise and technique has a stated purpose, an interpretation rule agreed in advance, and a debrief longer than the task.
12. In group guides, at least one structural remedy for dominance and one for premature convergence are designed in, not left to the moderator.
13. The moderator layer contains handling instructions for the predictable awkward situations, and escalation instructions where the topic warrants them.
14. Spoken text and moderator instructions are visually distinct and cannot be confused mid-session.
15. Where multiple markets, waves or moderators are involved, what is fixed for comparability and what is free to vary is stated explicitly.

## 11. Common failure modes

| Failure | How to recognise it | How to prevent it |
|---|---|---|
| **The questionnaire in prose** | Thirty to forty numbered questions, each a full sentence to be read aloud, no probes written | Build the section brief first, apply the Step 3 arithmetic, and count primary questions against what the session length allows |
| **Over-specification** | The guide reads as a script; a moderator following it cannot pursue an unexpected answer without losing their place | Write topics and probe classes, not sentences to be read. Fix the opening question and the exit condition; leave the route between them to the moderator |
| **Under-specification** | Six bullet points with no purposes, no probes and no timings; four moderators run four different studies | Purposes and exit conditions are mandatory even in a loose guide. It is the purpose that creates comparability, not the wording |
| **Leading probes** | "Was that because it was too expensive?", "So it was frustrating?", "Don't you find that annoying?" | The four probe tests in Step 7. Rewrite as echo, elaboration or specification, using only the participant's own words |
| **Prediction questions** | "Would you buy this?", "How likely are you to use it?", "Would you switch if we did X?" | Stated future behaviour in a moderated setting measures politeness and imagination. Ask what they did last time, what they considered, and what would have had to be different |
| **Direct motivation questions** | "Why do you prefer that brand?" asked once, answer taken at face value | People give the most available account, not the operative reason. Route through behaviour, sequence and consequence, or use laddering, then interpret and label the interpretation |
| **Stimulus too early** | Concept shown in the first third "so we have plenty of time to discuss it" | Draw the exposure boundary in Step 9 and check everything above it. The pre-exposure data cannot be recovered; the reaction time can be found by cutting elsewhere |
| **Decorative projectives** | An exercise with no stated interpretation rule, or one included because it will look good to observers | Step 10's three conditions. Price the exercise at task plus debrief and see whether it still beats the line of questioning it displaces |
| **The group used for individual data** | A guide asking each participant to reconstruct their own private journey, in a room of eight | Groups produce negotiated accounts and range. If the objective needs the individual case, the instrument is wrong. Back to **01.04** |
| **Dominance left to the moderator** | No written remedy; the transcript shows one participant with a third of the words | Structural remedies in Step 11: private response first, go-rounds, nomination, pairs |
| **Sensitive material placed by convenience** | Income or health in the warm-up because "it's quick" | Step 8. Sensitive late, normalised, with recovery time budgeted |
| **The frozen guide** | Session five is failing the same way session one failed, and nobody has changed anything | Pilot, inspect against exit conditions, revise, and log the revision with the sessions affected |
| **AI: fluent comprehensiveness** | A generated guide that is well written, well ordered, covers every objective evenly, and is three times too long for the session | Require the section brief, the timing arithmetic and the cut log as outputs, not just the guide. A guide with nothing cut has not been designed |
| **AI: symmetric probe sets** | Every question given the same four probes regardless of what would actually be interesting | Probes are selected by what the participant says. Write classes with examples, and vary them by what the section is for |
| **AI: invented stimulus and content** | Concept copy, product features, price points, brand names or claim wording that was never supplied | **K4 §2.1**. Mark the exposure point and leave the content as a labelled placeholder naming what is required |
| **AI: fabricated norms** | "Sessions of this type typically achieve...", "the standard focus group is..." with a number attached | **K4 §2.4, §2.5**. Timing figures are planning conventions, labelled as such, or they are not stated |
| **AI: the moderator layer omitted** | A guide with questions and no purposes, timings, handling instructions or escalation guidance | The moderator layer is half the deliverable. Section 9 lists what it must contain |

## 12. AI guardrails

Skill-specific only. **K4** applies in full and is not repeated here.

1. **Never produce a guide without the section brief and the timing arithmetic behind it.** A list of questions with no purposes, no minutes and nothing cut is a topic list presented as an instrument, and it transfers the design work to the moderator at the worst possible moment.
2. **Never invent stimulus content.** Concept descriptions, feature lists, price points, claim wording, competitor sets, brand names and category terminology come from supplied material. Where they are needed and absent, mark the exposure point and output a labelled placeholder naming what is required (**K4 §2.1, §6.2**).
3. **Never state a session norm, group size convention, saturation point or timing as a measured value.** Planning conventions are labelled as planning conventions (**K4 §2.4, §2.5**). "Saturation at twelve interviews" is a claim about a specific study, not a general fact, and is not asserted.
4. **Never write a probe that supplies its own answer, and never write a prompt list above the exposure boundary.** Both are irreversible contamination introduced by the instrument itself, and neither is visible in the transcript as a defect: the leading probe's effect looks exactly like a finding.
5. **Never generate example participant answers, illustrative quotes or "what participants will say" alongside a guide.** A guide is written before anyone has spoken. Anticipated responses in a design document migrate into analysis documents and become indistinguishable from evidence (**K4 §2.2, §2.3**).
6. **Never silently drop a stakeholder topic.** Every cut is recorded in the cut log with its reason and the objective it affects, so the conversation about scope is documented rather than defensive.
7. **Never claim a guide is neutral or unbiased.** Report the checks performed and what they found. Residual risks that could not be designed out (a recruitment framing that names the subject, a client-supplied prompt list, a hypothesis the client insisted on prompting) are disclosed with their likely direction.
8. **Never design a group guide for material that requires privacy, even when asked.** Explain what the group will produce instead (a normative account with the deviant cases suppressed), name the alternative instrument, and let the researcher decide (**K4 §9**, **K5 §2.4**).
9. **Never present a guide adaptation across markets or languages as equivalent without saying what was held fixed.** Comparability is a claim about design, and it has to be evidenced by the fixed-versus-free specification, not asserted.

## 13. Best-practice principles

1. **A guide is a memory aid for the moderator, not a script for the respondent.** The test: could a competent moderator run this while looking at the participant rather than the page? If it has to be read, it is the wrong document.
2. **If you cannot say what a successful section sounds like, you cannot moderate it.** The purpose sentence is not documentation. It is the thing the moderator steers by when a participant goes somewhere unexpected and a decision has to be made in two seconds.
3. **Time is the binding constraint and depth is what it buys.** Every question added subtracts depth from the questions already there. A guide is a budget, and the arithmetic is not optional.
4. **Ask about the last time, not about usually.** "Usually" invites a summary the participant constructs on the spot; "the last time" invites a reconstruction they can actually perform. Episodes are also the analysable unit, which generalisations are not.
5. **People are good reporters of what happened and poor reporters of why.** The introspective route to motivation returns the most culturally available explanation. Route through sequence and consequence, or ladder from attribute to consequence to value, and label the interpretation as yours.
6. **Never ask anyone to predict their own behaviour.** Ask what they did, what they nearly did, what stopped them, and what would have had to be different. Those are answerable. "Would you use this?" is not, and the answer is systematically generous.
7. **The probe is where qualitative research happens.** The question opens the door. Everything of value is on the other side of the second, third and fourth follow-up, and most weak transcripts are weak because the moderator moved on.
8. **Silence is an instrument, and it belongs in the guide.** Write it in explicitly, because moderators fill silence when it is not sanctioned, and the sentence after the pause is often the one that matters.
9. **Order is irreversible in a way wording is not.** A badly worded question can be rescued by a probe. A brand named too early cannot be unnamed, a concept shown too early cannot be unshown, and no amount of skill recovers what the sequence destroyed.
10. **Design the room, not just the questions.** Composition, seating, private-first responses and go-rounds do more for group quality than moderation technique, and unlike technique they can be specified in advance and repeated.
11. **A guide is a hypothesis about the conversation.** Field it, inspect it against its own exit conditions, revise it, and log the revision. A guide that never changes has either been designed exceptionally well or is not being examined.
12. **What the guide does not say, the moderator will have to invent live.** Every situation you can foresee and do not write down becomes an improvised decision made under time pressure with a participant waiting, and those decisions are where comparability quietly dies.

## 14. Worked example

**INPUT**

A fictional national charity, the Kestrel Trust, runs a hardship grant for households in short-term financial crisis. Internal records show that a substantial share of people who begin an application do not submit it, and that referral partners believe many eligible people never apply at all. Method already selected (**01.04**): 60-minute depth interviews, video and telephone offered, with eighteen people who either abandoned an application or were referred and never started one. Objectives supplied: (a) understand how people in this situation seek help, and where the Trust sits in that; (b) understand what stops an eligible person from applying; (c) get reaction to the redesigned application form. Recruitment framing agreed as "a conversation about managing when money is tight", not naming the Trust.

**PROCESS**

*Step 1, purposes.* (a) becomes "by the end, the moderator should have heard the sequence of things this person actually tried when money last ran short, in order, including the options they rejected". (b) becomes "the moderator should have heard at least one specific moment at which applying was possible and did not happen, with the participant's own account of what was in the way". (c) becomes "the moderator should have heard where in the form the participant stops, hesitates or asks a question, and what they think is being asked of them".

*Step 3, time.* 60 minutes. Overheads: 5 for introduction and consent (higher than usual because of the sensitivity and the recording), 5 for the close including recovery and the "anything we have not covered" question. Working time 50. Priced: help-seeking sequence 12 minutes as a reconstructed episode, alternatives and rejected options 10, the non-application moment 8, an indirect section on how asking for help is regarded 5, form exposure 8, warm-up and context 7. Total 50. Six lines of questioning, not the fourteen topics the internal stakeholder list contained. Eight topics cut and logged, including "views on the Trust's brand", which fails against every objective.

*Step 4 and 9, the judgement call.* The programme team asked for the redesigned form to be shown in the first fifteen minutes, on the reasoning that form feedback is the deliverable most likely to be acted on and should not be squeezed by overrunning. This was refused and the refusal recorded with its reasoning. Objective (b) requires reconstructing a decision made months earlier in a household under strain, and once a participant has been positioned as a reviewer of a form, they answer as a reviewer: barriers get reframed as form problems, and the barriers that have nothing to do with the form (not knowing the Trust exists, believing they would not qualify, not wanting to be the kind of person who applies) become invisible. The exposure boundary was drawn at minute 45. The team's underlying concern was addressed differently, by proposing a separate short usability session with a different sample, where the form can be the whole subject rather than an eight-minute reaction at the end of a demanding interview. Recorded as a **K5** review point, because whether that session is funded is a business decision this study cannot make.

*Step 7, probes.* The stakeholder draft contained "Was it because the form was too long?" and "Did you feel embarrassed about applying?". The first supplies a candidate answer; the second imports a noun the participant has not used and invites either a defensive denial or an agreeable yes. Both were replaced with echo and specification probes built from the participant's own words, plus a sequence probe ("what were you doing in the days just before you would have needed to send it") which reaches practical obstruction without naming it.

*Step 10, technique.* Objective (b) includes something direct questioning predictably fails on: the self-image cost of applying. "Did stigma stop you" produces denial in almost every case. A sentence completion was retained ("People who ask for help with money are..."), with an interpretation rule agreed in advance: the completion is a prompt for the talk that follows, never a finding on its own, and the analyst's reading of it is labelled as interpretation per **K2 §3**. Budgeted at 5 minutes, of which 1 is the task and 4 the debrief. A collage exercise proposed by the same stakeholder was cut: no interpretation rule, 15 minutes at task-plus-debrief pricing, and it would have displaced the help-seeking reconstruction, which is the study's core.

*Step 8, sensitive placement.* The financial-crisis reconstruction sits at minutes 12 to 24, not in the warm-up and not at the end. The close returns deliberately to safer ground and includes signposting agreed with the Trust in advance. The moderator layer carries explicit instructions on distress and on disclosure of a current crisis.

*Step 13, pilot.* Two pilot interviews showed the help-seeking section reaching its exit condition comfortably, but the non-application section producing generalised answers, because participants had already narrated the obstacles and treated the question as a repeat. The two were merged into one longer reconstruction with the non-application moment probed inside it, releasing 4 minutes. Revision logged, noting that interviews 1 and 2 used version 1.0.

**OUTPUT**

A six-section, 60-minute depth guide with an exposure boundary at minute 45, a section brief with purposes and exit conditions, a timing plan with cumulative markers and a cut priority order, a moderator layer covering distress escalation, observer protocol and the "what does the Trust want me to say" question, a cut log with eight entries, a revision log showing version 1.0 and 1.1 and which interviews used each, and two **K5** review points: whether a separate form usability session is commissioned, and whether the sentence completion is appropriate for participants recruited through a crisis referral route, which is an ethical judgement about a vulnerable sample rather than a design one.

## 15. Advanced usage

**Multi-market and multi-language guides.** Comparability lives at the level of objective, section purpose and sequence, not at the level of the sentence. Fix five things across markets: the objectives, the section order, the unaided-before-aided boundary, the stimulus and its exposure point, and the probe classes. Allow five things to vary: exact question wording, the warm-up, examples, the time distribution across sections, and the choice of enabling technique, because some do not travel (sentence completion is grammar-dependent, metaphor and humour rarely survive translation, and card sorts assume a literacy level). Translate the intent rather than the words: brief a local moderator on the purpose, have them write the question in language a local participant would use, then back-check that it still serves the purpose. Where an interpreter is unavoidable, roughly halve the effective content, because simultaneous translation costs time and consecutive translation costs the probe. Log every adaptation, so that at analysis a market difference can be distinguished from an instrument difference, and require a local reading before findings are fixed per **K5 §2.2**.

**Longitudinal and repeat-contact guides.** Build a fixed spine that runs unchanged in every wave and a wave-specific module that carries the new content. Anything to be compared across waves is asked in the same words, in the same position, after the same prior context: a question that moves from minute 40 to minute 15 has changed even if its wording has not. Avoid opening with "what has changed since we last spoke", which invites a summary judgement rather than a reconstruction; re-anchor on specifics the participant gave last time, then ask what happened since. Each guide needs a carry-forward brief: what this participant said before that the moderator must know, and, separately, what must not be reminded to them before the unaided section. Expect panel conditioning. Repeat participants become more articulate, more expert and increasingly fluent in the study's own vocabulary, which looks like insight and is an artefact. Mitigate by keeping unaided sections first and clean, refreshing part of the sample where the design allows, and reporting conditioning as a limitation rather than designing around it silently.

**Mode.** Choose on methodological merits, and design the guide to the mode rather than porting one guide everywhere. Face to face gives the full nonverbal channel, physical stimulus and the best conditions for rapport and long sessions, at the cost of geography, expense and a visibly artificial setting. Video reaches dispersed samples and handles screen-based stimulus well, but flattens group dynamics because turn-taking serialises, loses peripheral behaviour, fatigues faster (so cut roughly a fifth from what the same length would carry face to face), and skews toward the digitally confident. Telephone removes the visual channel entirely, which rules out stimulus and nonverbal reading, but the reduced social presence can raise disclosure on sensitive topics and reaches people video excludes; guides must be shorter, more structured, and explicit about how the moderator handles silence they cannot see the reason for. Asynchronous text spreads the session across days, suits diary-shaped and reflective material, and reaches people who cannot give an uninterrupted hour, but the answers are composed rather than spontaneous, the register shifts toward the presentable, real-time probing is replaced by a next-day exchange, and attrition across days is a design constraint rather than a nuisance. Where sessions will be AI-moderated, the probe architecture must be specified in advance rather than exercised live, which is a different instrument design problem: hand over to **03.02 AI-Moderated Interview Design**.

**When the standard approach does not fit.** Very short sessions (20 to 30 minutes, common in intercept and clinical settings) support one line of questioning and should be designed as one, not as a compressed hour. Expert and elite interviews invert the usual balance: the participant sets the pace, the guide becomes a checklist of what must be covered before the time is taken away, and the probe classes matter more than the questions. Sessions with observers who intervene need a written protocol agreed before fieldwork, because an unbriefed observer question at minute 50 can undo the entire sequencing design.

## 16. Skill chain

**Recommended previous skills**
- **01.04 Research Method Selection.** Confirms that moderated qualitative is the right instrument, and which shape of it, and hands over what the design can and cannot answer.
- **01.05 Research Plan Development.** Hands over session count, length, participant definition, fieldwork window and the objectives in their agreed final form, which are the inputs the time budget is built on.

**Recommended next skills**
- **02.03 Interview Question Development.** Takes the primary questions and works the wording, the episodic framing and the probe examples in detail.
- **02.04 Question Bias Detection.** Audits the drafted questions and probes independently. Design and review should not be the same pass.
- **03.02 AI-Moderated Interview Design.** Takes the section purposes and probe architecture and converts them into the explicit specification an automated moderator requires.

**Runs well alongside**
- **13.05 Research Ethics and Consent Design**, wherever the topic is sensitive, the sample vulnerable, or re-contact planned.
- **02.06 Screener and Quota Design**, because group composition decisions made here have to be delivered by the screener.
- **07.03 Interview and Transcript Analysis** downstream, which inherits the section brief as the map from theme back to objective.

---
A Yazi Supplied Skill and resource.
