---
name: survey-questionnaire-design
description: >
  Designs a professional survey questionnaire from research objectives. Use when
  someone says "write me a survey", "design a questionnaire", "draft the survey
  questions", "turn these objectives into a survey", "we need a customer survey",
  "build the instrument", or asks how long a survey should be, what order the
  questions should go in, or whether a question is biased.
category: 02 Instrument Design
ref: 02.01
tier: 0
inherits: [K2, K3, K4, K5]
---

# Survey Questionnaire Design

## 1. One-line description
Turns research objectives into a fielded questionnaire in which every question earns its place by enabling a specific decision, and every objective is measured by something.

## 2. What this skill is used for

**The research problem it solves.** Most weak surveys are not weak because a question was worded badly. They are weak because nobody asked what the answers were for. Questions accumulate: one from the brief, three from a stakeholder who "would find it interesting", four carried over from last year because they were there last year. The instrument gets long, respondents satisfice, the data gets worse at exactly the point the survey gets expensive, and at analysis the researcher discovers that the one objective the client actually cared about was never measured. Bias in wording is the second problem. Absence of purpose is the first, and it is the one that cannot be fixed after fielding.

**Where it sits in the research lifecycle.** After the method is chosen and the sample and analysis are planned. Before screener detail, logic review, bias audit and scale selection, each of which reviews or deepens part of what this skill produces.

**Typical use cases.**
- Building a questionnaire from a signed-off objectives list or proposal.
- Converting a stakeholder wish list into a defensible instrument, with the cuts justified.
- Designing wave 1 of a tracker, where wording decisions will be locked for years.
- Rebuilding a survey that runs too long, breaks off too often, or produces unusable data.
- Adapting an instrument written for one mode (interviewer, desktop) to another (mobile, voice).

**Who uses it.** Research executives and managers drafting instruments; research directors reviewing them; insight managers client-side who commission surveys and need to challenge what comes back; UX and product researchers running their own quantitative work without a survey specialist nearby.

## 3. When to use it

- Objectives exist and are agreed, and an instrument now has to exist.
- A stakeholder has supplied a list of questions and someone has to decide which survive.
- A survey is being repeated and this is the moment to fix it, before the trend locks the wording.
- The study needs subgroup comparisons and the classification questions have to be right, not improvised at the end.
- The audience, mode or device profile has changed and the existing instrument no longer suits it.
- A survey is running long and the length has to come out of it without losing an objective.
- A questionnaire is being written for a population whose language, literacy or context differs from the researcher's own.

## 4. When NOT to use it

- **The objectives are not settled.** A questionnaire cannot resolve an unclear research question, it can only encode the confusion in a fixed form and multiply it by the sample size. Go to **01.01 Research Brief Interrogation** or **01.02 Business Problem to Research Question** first. If a stated objective has no decision attached to it, say so and stop rather than designing around it.
- **The question is "why", or the vocabulary is not yet known.** Surveys measure the distribution of things you already know how to name. If you cannot yet write the answer options because you do not know what people would say, a survey will produce a tidy distribution of your assumptions. Use **02.02 Discussion Guide Design** and qualitative work first, then return.
- **The behaviour can be observed directly.** Where transaction, usage, log or till data exists, asking people to recall it produces a worse measurement at higher cost. Ask only for what the behavioural data cannot supply: reasons, alternatives considered, satisfaction, things that happened elsewhere.
- **The objective is a trade-off, a price point or a preference share.** Rating scales cannot produce these. Stated importance ratings in particular are among the least reliable measures in commercial research, because almost everything is rated important. This needs a designed exercise (conjoint, MaxDiff, a pricing method) specified in category 06, not a battery of importance questions.
- **A tracker is already in field.** Improving the wording of a tracked question breaks the series. The gain from a better question is usually smaller than the loss of the trend. Where a change is genuinely necessary, it is a parallel-run decision, not a design decision, and it belongs to a researcher, not to this skill.
- **The sample cannot support what the objectives require.** If the study needs to compare four segments and the plan yields 40 in the smallest, the instrument is not the problem. Fix **01.06 Sampling Strategy** first. No amount of question quality rescues an inadequate base.
- **You are reviewing rather than designing.** For an existing questionnaire, **02.04 Question Bias Detection** and **02.05 Survey Logic and Flow Review** are the right tools. Running a design skill over a finished instrument tends to produce a rewrite nobody asked for.
- **The topic is one where self-report is known to be unreliable and no design adjustment is available.** Sensitive, socially regulated or heavily normative behaviour reported to a stranger will be misreported. Where the design cannot mitigate this (indirect questioning, self-completion mode, validated instruments), report that the measure will be biased in a known direction, or do not field it.

## 5. Required inputs

**Required. Without these the skill cannot run.**
- **Research objectives**, and for each one, the decision it informs. If objectives are present but decisions are not, ask once. If no answer is available, proceed and record the decision column as `[not supplied]`, which will flag which questions cannot be justified.
- **Target population and sample definition**, including the subgroups that must be reportable. Subgroup requirements drive the classification section and the minimum base logic. Without this, stop and ask.
- **Mode and expected device profile.** Self-completion, interviewer-administered, mobile-dominant, voice. Wording, option-list length and scale choice all change with mode. Without this, assume self-completion on a small screen and state the assumption, because it is the most constraining common case.

**Optional, and what each one adds.**
- **Analysis plan (01.07).** Lets each question be written to the exact form the analysis needs: the banner points, the nets, the derived variables. Without it you will discover at analysis that a question was asked in a form that cannot be tabulated the way the story requires.
- **Previous wave or comparable instrument.** Makes trend comparability possible and shows which wording is already locked.
- **Quota and screener plan (02.06).** Removes duplication between screener and classification, which is a common source of unnecessary length.
- **Known incidence of the target group.** Determines how much screening the questionnaire must do and whether the screener needs to be a separate, shorter instrument.
- **Regulatory, legal or ethical constraints.** Special-category data, minors, financial advice, health claims. Changes what may be asked and what consent wording must accompany it.
- **Length or cost ceiling.** Converts the elimination step from a judgement into a budget, which is easier to hold against a stakeholder.

## 6. Questions to ask before starting

1. **For each objective, what decision changes depending on the answer?** This is the governing question of the whole skill. If the answer is "nothing, it would just be good to know", the objective does not need measuring. Default if unanswered: proceed, mark the objective as decision-unspecified, and flag every question hanging off it as a candidate for removal.
2. **Which subgroups must be reportable separately?** Determines classification questions and the base sizes the sampling plan must deliver. Default: report on total only, and say so.
3. **What is the maximum acceptable length, and what is it based on?** Incentive, audience patience, and cost all bear on this. Default: assume a 12-minute self-completion ceiling for a general-population survey and 15 to 20 minutes for an incentivised specialist audience, and state that the estimate is a planning heuristic, not a measurement.
4. **Is any of this being compared to something?** A previous wave, a competitor benchmark, a norm database. Comparability constrains wording far more than quality does. Default: assume standalone, and flag any question where a common convention exists.
5. **What has already been decided that the survey cannot change?** Research that arrives after the decision is made needs a different, smaller instrument, or none. Default: assume the decision is open.
6. **How sensitive is the subject matter for this audience?** Determines placement, mode, the need for reassurance wording, and whether a non-response option is required. Default: treat income, health, finances, employment status and any protected characteristic as sensitive.
7. **Who else will want questions in this survey, and when will they see it?** Stakeholder additions arriving after the draft is written are the single largest source of length inflation. Default: assume at least one round of additions and hold length budget in reserve for it.

## 7. Step-by-step methodology

**Step 1. Convert objectives into decisions.** Before writing anything, list the objectives and next to each one write the decision it serves, the person who will make it, and what they would do differently under each plausible answer. An objective that survives this is a measurement target. An objective that does not is a topic of interest, and topics of interest are where questionnaires go to die. Where an objective cannot be tied to a decision, do not delete it unilaterally: mark it and raise it. A correct result at this step is a short table in which every row has a named decision or an explicit `[not supplied]`.

**Step 2. Build the objective-to-question mapping grid.** This is the starting artefact and the spine of everything that follows. One row per intended question, with these columns: objective ref, decision served, construct being measured, draft question, question type, response frame, base (who sees it), and the analysis output it feeds (a specific table, chart, cut or model term). Two rules govern the grid, and they are the discipline of the whole skill.

- **No question without a parent objective.** A row with an empty objective column is deleted or promoted to a new agreed objective. There is no third option.
- **No objective without a question.** An objective with no rows is unmeasured, and unmeasured objectives are how a study fails silently. Every objective must be traceable forward into at least one question, and back again.

Build the grid before drafting wording. Researchers who draft first and map afterwards almost always rationalise the questions they have already written.

**Step 3. Apply the elimination test to every row.** Ask of each: what decision will the answer to this question enable? Then apply four cuts. **Redundancy:** is this measured by another question, or derivable from one? **Availability:** does the organisation already hold this, in CRM, transaction or operational data? **Actionability:** if the answer came back at any value, would anyone do anything? **Answerability:** can a respondent actually know this, remember it accurately, and be willing to say it? Questions asking people to explain their own motivations, predict their future behaviour, or recall the frequency of routine low-salience events fail the fourth cut more often than researchers expect. Log every removal with its reason. The log is what turns a defensive conversation with a stakeholder into a documented one.

**Step 4. Set the architecture.** Order the surviving questions into five blocks, in this sequence.

- **Screener.** Eligibility and quota only. Nothing that is not a pass or fail. Keep it short and neutral so it does not teach respondents what to say to qualify. Detailed design belongs to **02.06**.
- **Warm-up.** Easy, relevant, non-threatening, on-topic. Its job is to establish that the survey is about what the invitation said and that answering is easy. It is not the place for the most important question, and it is not the place for classification.
- **Core.** The substance, ordered by the sequencing rules in Step 5.
- **Sensitive.** Income, health, personal circumstances, anything a respondent might not want to answer. Placed late, so that a break-off at this point costs the core data nothing.
- **Classification.** Demographics and firmographics not already captured in the screener. Late, because they are dull, and because they are needed for weighting and cuts rather than for the respondent's engagement.

**Step 5. Sequence the core.** Four rules, in priority order. **Unaided before aided**, always: once a brand, feature or attribute list has been shown, spontaneous awareness of it cannot be measured again in that interview, and this contamination is irreversible. **General before specific**: overall satisfaction, overall opinion and overall likelihood come before the diagnostics that explain them, because asking about a specific failure first primes the overall rating downward (and a specific delight primes it upward). **Behaviour before attitude** where both are asked about the same object, since attitudes reported after describing behaviour tend to be rationalised into consistency with it. **Related questions grouped, unrelated topics separated by a transition**, so respondents are not forced to switch context repeatedly. Then check the sequence for priming in both directions: what does each question teach the respondent that changes how they answer the next one, and what does it imply about what the researcher wants to hear?

**Step 6. Write each question.** One idea per question. Neutral framing. Plain words the respondent uses, not the words the client uses. A defined, realistic recall period. A base that is unambiguous. Then run the defect catalogue in Step 8 over the draft before anyone else sees it.

**Step 7. Build the response frame.** Answer options are where most quantitative surveys actually break, because a wording flaw is visible and a coding flaw is not. Options must be **exhaustive** (every respondent can find themselves), **mutually exclusive** (no respondent could legitimately choose two), **balanced** (equal numbers of positive and negative points, with symmetric intensity in the labels), and **ordered consistently** across the instrument. Numeric ranges must not overlap at the boundary, and must not leave a gap. Then decide, question by question and never by default, whether the frame needs a non-substantive option:

- **"Don't know"** belongs where not knowing is a real and meaningful state (factual questions, questions about a brand the respondent may not have encountered, questions about an event they may not have witnessed). It does not belong on an attitude question where everyone has some view, because it becomes an escape hatch for effortful thinking and inflates non-response. Where it is included, it is offered as a distinct, visually separated option, not as a scale point, and it is excluded from the base at analysis.
- **"None of these"** belongs on any multi-select where genuinely selecting nothing is possible. Without it, respondents who fit no option will pick the nearest one, and a false positive is worse than a missing value. It is exclusive: selecting it clears everything else.
- **"Other, please specify"** belongs where the list may be incomplete. A high "other" count is a finding about the list, not about the respondents. Use it in wave 1 and in any list you did not derive from qualitative work or existing data, and plan to back-code it. Do not use it as insurance against having failed to research the list.
- **"Prefer not to say"** belongs on every sensitive question. Its absence does not produce honesty, it produces break-off or a false answer.

**Step 8. Run the defect sweep.** Read every question against the catalogue below, in one pass per defect rather than one pass per question, because the eye is much better at spotting a pattern than a list.

| Defect | How it shows up | Why it damages the data | Fix |
|---|---|---|---|
| **Leading** | The question suggests its own answer. "How much did you enjoy the new layout?" | Shifts the distribution toward the suggested response; the size of the shift is unknown and unrecoverable | Neutral stem, balanced options. "How would you describe the new layout?" with a full positive-to-negative scale |
| **Loaded** | Carries an embedded assumption or an emotive frame. "Do you agree that wasteful spending should be cut?" | Measures agreement with the framing, not with the proposition | Strip the evaluative adjective, present the proposition neutrally, and offer both sides in the stem where relevant |
| **Double-barrelled** | Two ideas, one answer. "How satisfied are you with the speed and accuracy of the service?" | A respondent who is happy with one and not the other cannot answer, and the analyst cannot tell which drove the score | Split into two questions, or, if length forbids, keep the one that serves the decision and drop the other |
| **Overlapping options** | "18 to 25, 25 to 35, 35 to 50" | Respondents at the boundary choose arbitrarily, and the categories are no longer comparable to anything | Close the ranges: 18 to 24, 25 to 34, 35 to 49 |
| **Incomplete scale** | Options that do not cover the possible range, or a missing midpoint where a genuine neutral exists | Forces respondents into a position they do not hold, inflating whichever end is nearer | Cover the range; decide on the midpoint deliberately (see 02.07) rather than by habit |
| **Unbalanced scale** | "Excellent, very good, good, fair, poor" | Three positive points against one negative point; the mean is pushed upward before anyone answers | Symmetric points and symmetric label intensity either side of the centre |
| **Order effects** | Aided list before unaided recall; a diagnostic before the overall; a long list never rotated | Earlier questions contaminate later ones; unrotated lists accumulate primacy (visual) or recency (aural) bias in fixed positions | Apply the Step 5 sequence; randomise or rotate lists and batteries, pinning "Other" and "None of these" |
| **Priming** | An introduction that explains the client's position, or a block of questions on one attribute before a choice question | Respondents answer the salient frame you just gave them | Neutral introductions; place choice and overall measures before attribute exploration |
| **Unnecessary questions** | Nothing in the analysis plan uses it; nobody can name the decision | Consumes the respondent's finite attention, which is a shared budget: a wasted question degrades the answers to the ones that matter | The Step 3 elimination test, applied without exception |
| **Excessive length** | Beyond the agreed budget; long grids; repeated batteries | Break-off, straightlining, speeding and satisficing, concentrated in the back half of the instrument and in the least engaged respondents, so the loss is not random | Cut by objective, not by trimming words. See Step 9 |
| **Poor recall period** | "How many times in the past year did you...?" for a routine event | Produces recall error and rounding to salient numbers; long periods for frequent low-salience behaviour are effectively unanswerable | Shorten the window to one the respondent can actually reconstruct, or ask about the most recent occasion instead of a count |
| **Ambiguous language** | "Regularly", "recently", "typical", "household", "your main provider" | Different respondents answer different questions, and the variance looks like real variance | Define the term inside the question, or replace it with a specific quantity or date |
| **Jargon** | Internal product names, sector acronyms, category language | Non-response, guessing, and a systematic bias toward respondents who happen to know the term | Use the respondent's vocabulary; where a term must be used, define it in the question, not in a preamble |
| **Acquiescence bias** | Agree/disagree batteries, yes/no framings, long same-direction grids | A general tendency to agree, stronger in some populations than others, inflates agreement and correlates the whole battery | Prefer item-specific response frames over agree/disagree; where a battery is unavoidable, reverse some items and check reversal at analysis |
| **Social desirability** | Reported voting, exercise, charitable giving, healthy eating, safety behaviour, "would you recommend" | Over-reporting of approved behaviour and under-reporting of disapproved behaviour, in a known direction of unknown size | Self-completion mode where possible, normalising preamble ("some people do, some people don't"), indirect or third-person framing, response options that make the less desirable answer easy to select |
| **Inappropriate scale** | A 10-point scale on a construct with three meaningful levels; a satisfaction scale on a question about frequency; different scales for questions that will be compared | Precision that does not exist, or answers that cannot be combined | Match the scale to the construct and keep it consistent within the instrument. Detail belongs to **02.07** |
| **Missing answer options** | No "don't know" on a factual question; no "none of these" on a multi-select; no "prefer not to say" on income | Forces false answers, which are indistinguishable from true ones in the data | Apply Step 7, per question, deliberately |

**Step 9. Budget the length.** Estimate completion time from the question inventory using a stated convention (for example, a simple closed question at roughly 10 to 15 seconds, a grid row at roughly 5 to 8 seconds, an open text at 30 to 60 seconds), state that this is a planning heuristic rather than a measured value, and calibrate it against observed data from comparable studies where any is available. If the estimate exceeds the budget, cut whole objectives or whole blocks, and say which. Do not shorten by deleting scale points, removing "don't know", compressing three questions into one double-barrelled question, or converting single questions into grids. Each of these buys a minute and costs data quality, and the cost is invisible until analysis.

**Step 10. Adapt to mode.** Self-completion needs questions that are entirely self-contained, because no one can clarify them, and it carries a primacy tendency in visible lists. Interviewer-administered can carry more complex routing but needs option lists short enough to hold in working memory (roughly four to five for aural presentation), carries a recency tendency, and raises social desirability because a human is listening. Small screens punish wide grids, long option lists and horizontal scales; a battery that reads as a matrix on a desktop should be designed to work as a sequence of single items, and a non-substantive option at the bottom of a long scrolling list is effectively hidden. Voice removes the visual list altogether: options must be few, labels must be short and repeatable, and any scale must have its anchors read out each time. These are properties of the mode, not of any particular tool, and they hold wherever that mode is used.

**Step 11. Specify routing at design level.** Mark, for every question, who sees it and why, and make sure every respondent has a complete path through the instrument. Check three things here and leave the rest to review: no dead end (a respondent who reaches a question with no valid answer available), no orphan (a question no one can reach), and no base that becomes too small to report once the routing is applied. Detailed logic review, including interaction between routing and quotas, belongs to **02.05**.

**Step 12. Pilot, then soft-launch.** Cognitive test the draft with a small number of people from the target population, asking them to say aloud what each question means to them and how they arrived at their answer. This finds ambiguity, jargon and unanswerable recall requests, which reading the draft never does. Then soft-launch to a small share of the sample and inspect: completion time against the estimate, break-off by question, straightlining in batteries, "other" verbatims (which reveal missing options), "don't know" rates (which reveal unanswerable questions), item non-response on sensitive questions, and the distribution on each key measure (a variable with 96% in one category will not support the analysis it was written for). Fix before the rest of the sample is spent. A pilot that produces no changes usually means nobody looked hard enough.

## 8. Analytical framework

The instrument is built on one chain, applied to every question and readable in both directions:

    Objective → Decision → Construct → Question → Response frame → Analysis output

**Forward** is the design test: this objective informs this decision, which requires knowing this construct, which is measured by this question, in this response frame, producing this specific table or model input. If any arrow cannot be drawn, the question is not finished.

**Backward** is the elimination test: this question produces this output, which feeds this decision, which serves this objective. If the chain breaks anywhere on the way back, the question comes out. In practice the chain most often breaks at the "analysis output" link, where a question turns out to produce a number nobody would ever put on a page.

The grid holding this chain is a live document, not a planning exercise. It is updated whenever a question changes, it accompanies the questionnaire to the client, and at analysis it becomes the map from every table back to the objective that justified it. It is also the K2 evidence chain starting one step earlier than usual: traceability begins at instrument design, because a finding can only be traced to a question that was written for a purpose.

## 9. Output format

**1. Design summary.** Objectives, mode, target population, estimated length with the convention used, and the constraints applied.

**2. Objective-to-question mapping grid.**

| Obj ref | Objective | Decision served | Construct | Q no. | Question type | Response frame | Base | Analysis output |
|---|---|---|---|---|---|---|---|---|

Every objective appears at least once. Every question appears exactly once. Objectives with no question are shown with the question columns empty and marked `UNMEASURED`, which is a defect to be resolved, not a formatting state to be left alone.

**3. The questionnaire.** In field order, with block headings (screener, warm-up, core, sensitive, classification). Each question carries: number, base instruction, question text exactly as the respondent will see or hear it, response options in field order, single or multi-select, randomisation or rotation instruction with pinned items named, scale definition, and routing instruction.

**4. Design rationale notes.** Short notes on the decisions a reviewer would otherwise query: why this scale, why this recall period, why this question is placed here, why this list is rotated, why a midpoint is present or absent.

**5. Removed-questions log.**

| Proposed question | Source | Reason removed | Decision it would have served | Reinstate if |
|---|---|---|---|---|

**6. Length estimate.** By block, with the per-question-type convention stated and labelled as a planning estimate.

**7. Open items and review points.** Per **K5 §3**, at the point of the decision, with a consolidated list at the front.

**When the inputs are thin**, the format does not get filled in anyway. An objective with no stated decision is written as `[decision not supplied]`, not as a plausible-sounding decision. A construct with no agreed measure is named as an open item rather than measured by an invented question. A length estimate with no calibration data says so. See **K4 §1**: the presence of a column is not a reason to populate it.

## 10. Quality checks

Run before the draft is shared. These sit on top of **K4 §8**, which runs anyway.

1. Every question traces to an objective, and every objective traces to at least one question.
2. Every question has a named decision, or an explicit note that the decision was not supplied.
3. Every question has a named analysis output, and that output is one a researcher would actually present.
4. No question contains two ideas, an evaluative adjective, an embedded assumption, or an undefined quantifier.
5. Every option list is exhaustive and mutually exclusive, with no overlapping or gapped numeric ranges.
6. Every scale is balanced in both point count and label intensity, and consistent with other scales it will be compared to.
7. Non-substantive options ("don't know", "none of these", "other", "prefer not to say") have been decided per question, with a reason, and not applied or omitted by default.
8. Unaided questions precede every aided list on the same subject, with no exception anywhere in the instrument.
9. Overall evaluations precede diagnostics on the same object.
10. Sensitive questions sit after the core, and each carries a non-response option.
11. Every list of more than about five items and every attribute battery has a rotation or randomisation instruction, with anchored items pinned.
12. Every question states its base, every respondent has a complete path, and no routed base falls below the reportable minimum agreed in the sampling plan.
13. Mode has been applied: no wide grid or long scrolling list where small-screen completion is expected; no list longer than about five options where questions are read aloud.
14. The estimated length is within budget, and any overage has been resolved by cutting objectives rather than by compressing questions.
15. Classification questions match the cuts the analysis plan requires, in the categories the analysis will use, including any external benchmark or weighting target the study must match.

## 11. Common failure modes

| Failure | How to recognise it | How to prevent it |
|---|---|---|
| **The stakeholder wish list** | Questions arrive by email and are added without a mapping grid row | Route every addition through Step 2 and Step 3, and hand back the removed-questions log rather than a refusal |
| **Objective drift** | The instrument is complete and one objective has no question against it | The grid is checked in both directions before drafting stops, not after |
| **Design by precedent** | "It was in last year's survey" is the only justification offered | Precedent is a reason for wording consistency, never for inclusion. Re-run the elimination test on inherited questions |
| **Length denial** | The estimate is over budget and the response is to shorten question text | Cut objectives or blocks, and record what was cut and what is no longer answerable |
| **The interesting question** | Everyone likes it and nobody can say what would be done with the answer | Ask for the table it produces. If nobody can describe the chart, remove it |
| **Scale drift within an instrument** | Three different scale lengths and two different label sets across comparable questions | Fix the scale conventions before drafting, not during. See 02.07 |
| **Analysis-blind wording** | At analysis, a question cannot be netted, banner-cut or combined the way the story needs | Write the analysis output column before the question text |
| **AI: fluent but purposeless questions** | A generated draft is well written, well ordered and 40% longer than it needs to be, with no removals | Require the mapping grid and the removed-questions log as outputs, not just the questionnaire. A design with no removals has not been designed |
| **AI: symmetry over substance** | Every construct gets the same five-point agree/disagree battery because it looks consistent | Choose the response frame from the construct, and prefer item-specific frames over agree/disagree (acquiescence) |
| **AI: invented context** | Generated preambles or option lists containing brand names, product features, prices or categories that were never supplied | K4 §2.1. Every substantive option list comes from the brief, from data, or from qualitative work. Otherwise it is marked as a placeholder needing client input |
| **AI: plausible norms** | "Surveys of this type typically achieve..." with a number attached | K4 §2.4. State timing and benchmark figures as planning conventions or not at all |
| **Silent mode assumption** | The draft contains a nine-item grid and nobody established the device profile | Mode is a required input. Where unstated, assume small-screen self-completion and say so in the design summary |
| **Pilot as formality** | The pilot runs, nothing changes, fieldwork proceeds | Specify in advance what the pilot will be inspected for (Step 12) and what result would trigger a change |

## 12. AI guardrails

Skill-specific only. **K4** applies in full and is not repeated here.

1. **Never generate a question without a mapping grid row.** If the objective, decision or analysis output cannot be stated, the question is not produced. Producing it and hoping the researcher will check inverts the responsibility.
2. **Never invent the substantive content of an option list.** Brand lists, feature lists, competitor sets, price points, channel names, category descriptors and internal terminology come from supplied material. Where they are needed and not supplied, output a labelled placeholder and name what is required (K4 §2.1, §6.2).
3. **Never state a completion time, break-off rate, response rate or industry norm as a measured value.** Timing conventions used to budget length are labelled as planning heuristics (K4 §2.4, §2.5).
4. **Never quietly drop a stakeholder question.** Every removal is logged with its reason in the removed-questions log. Silent deletion is unfalsifiable and destroys trust in the design.
5. **Never present a wording change to a tracked question as an improvement without flagging the break in comparability.** The trend loss is the material consequence and it is the researcher's decision (K5 §2.7).
6. **Never assume the direction, polarity or point count of a scale carried over from another instrument.** Confirm it or mark it unverified (K4 §6.2).
7. **Never fill an unmeasured objective with a question that approximates it.** An adjacent measure presented as if it measured the objective is the instrument-design version of manufacturing evidence. Mark it `UNMEASURED` and say what would be required.
8. **Never claim a questionnaire is unbiased.** Report the defect sweep as performed and what it found. Residual bias that could not be designed out is disclosed, with its likely direction, in the design rationale notes.

## 13. Best-practice principles

1. **Respondent attention is a fixed budget, and every question spends it.** The cost of a pointless question is not the 12 seconds it takes. It is the quality of every answer after it. This is why the elimination test is a data-quality tool and not an efficiency measure.
2. **Design backwards from the table.** An experienced designer writes the chart title first, then the question that produces it. Questions written forward from a topic produce data that cannot be presented.
3. **The respondent's vocabulary beats the client's, always.** If the category calls it a "solution" and people call it "the machine", the question says machine. Client terminology in a questionnaire is a systematic filter for respondents who work in the category.
4. **Answer options carry as much bias as question stems, and get a fraction of the scrutiny.** Most reviewers read the questions and skim the lists. Read the lists twice.
5. **Prefer item-specific response frames to agree/disagree.** "How easy or difficult was it to...?" outperforms "The process was easy: agree/disagree" on acquiescence, on interpretability and on comparability, at no cost in length.
6. **Make the undesirable answer easy to give.** Social desirability is reduced far more by a normalising preamble and a permissive option list than by any amount of confidentiality assurance. "Many people find it hard to keep up with this. Which best describes you?" collects more truth than a promise of anonymity.
7. **A recall period is a design parameter, not a phrase.** Match it to the salience and frequency of the behaviour: a long window for rare salient events, a short window for frequent unremarkable ones, and last-occasion questioning where a count would be guessed.
8. **Neutrality is symmetry, not the absence of adjectives.** Offering both sides in the stem ("Some people think X, others think Y. Which is closer to your view?") is more neutral than a stripped-down stem with an unbalanced option list.
9. **Randomise everything that can be randomised, and pin everything that cannot.** Rotation converts a fixed bias into random noise, which the analysis can live with. "Other" and "None of these" never rotate.
10. **The first question sets the contract.** If the invitation promised a survey about deliveries and the first question is about household income, the respondent learns that the survey is not what it said, and the answers change.
11. **Write the classification section against the analysis plan and any weighting target, not from memory.** Age bands that do not match the population statistics the study will be weighted to are a self-inflicted wound discovered too late to fix.
12. **A questionnaire is a research instrument and a piece of writing.** Read it aloud end to end before it goes anywhere. Anything you stumble over, a respondent will stumble over, and they have less patience than you.

## 14. Worked example

**INPUT**

A fictional public library service, the Riverbend Regional Library Service, is deciding whether to extend evening opening at three of its eleven branches. Objectives supplied: (a) understand current usage patterns; (b) assess demand for evening opening; (c) understand barriers to visiting; (d) "gauge overall satisfaction with the service". Mode: self-completion, expected to be mostly on phones, sample drawn from the library's registered borrower list. Budget: 8 minutes.

**PROCESS**

*Step 1, decisions.* (a) usage patterns determines which branches are candidates. (b) demand determines whether to extend at all and where. (c) barriers determines whether evening opening is even the right intervention, since if the main barrier is parking, later hours will not help. (d) has no decision attached. It is on the list because the service reports satisfaction annually. Marked `[decision not supplied]` and raised.

*Step 2, grid.* Fourteen candidate questions mapped. Objective (b) initially has one question: "Would you use the library if it was open later in the evening?"

*Step 3, elimination.* Three questions cut. Two on borrowing volumes, which the service already holds in its lending records (availability cut). One on "how important is the library to the community", which fails the actionability cut: no answer to it changes the opening-hours decision. All three logged, the last with a reinstate condition of a separate advocacy objective being agreed.

*The judgement call.* Objective (d) was raised with the client. The answer was that the annual satisfaction number is genuinely used, in the service's published performance report. That is a decision of a kind, so one overall satisfaction question was retained, using the previous year's exact wording to preserve the trend, and placed at the top of the core block, before any question about barriers or opening hours. Placing it after the barriers block would have primed it downward and broken the comparison with a year in which no such priming existed. Recorded in the design rationale notes as the "general before specific" rule applied to a trended measure.

*Step 6 and 7, wording.* The demand question fails three ways. It is leading (a respondent is being asked to endorse a service improvement, which almost everyone will), it is hypothetical (stated future behaviour, which over-reports), and it has no recall anchor. Rewritten as two questions with a real behavioural anchor: first, "In the last month, was there a time you wanted to visit a Riverbend library and could not?" (yes / no / don't know), then, for those who said yes, a multi-select of reasons including "It was closed when I could have gone", with "Other, please specify", "None of these" and no "don't know" (having just said it happened, the respondent can describe it). Only then a direct question on which evening time slots they would use, worded as a choice among specific slots rather than as a yes/no to a better service.

*Step 9, length.* Eleven questions, estimated 6 to 7 minutes on the stated convention, inside budget with room for one stakeholder addition.

*Step 12, pilot.* Cognitive testing with six borrowers found that "branch" was read as "the one I normally use" by some and "any Riverbend library" by others. Wording changed to name the specific branch in the question stem.

**OUTPUT**

An eleven-question instrument, the mapping grid showing all four objectives measured (with (d) annotated as trend-continuity rather than decision-driven), a removed-questions log with three entries, design rationale notes covering the satisfaction placement and the hypothetical-demand rewrite, and one **K5** review point: whether the barriers question list, derived from staff opinion rather than from research, is complete enough to field without a qualitative check.

## 15. Advanced usage

**Multi-market instruments.** Design in the language of analysis, then translate and back-translate, and expect the scale to be the hardest part: response-style differences across cultures (extreme responding, midpoint avoidance, acquiescence) can be larger than the substantive differences the study is trying to measure. Where cross-market comparison is the objective, favour item-specific frames and behavioural anchors over attitude scales, and plan the standardisation approach in the analysis plan rather than discovering the problem in the cross-tabs.

**Tracker wave 1.** Wording decisions made now will be unchangeable for years. Spend disproportionate effort on the questions that will be trended, use established conventions where they exist so the data can be benchmarked, and design deliberate slack: a rotating module of non-trended questions gives future waves somewhere to put new content without touching the core.

**Embedded experiments.** A questionnaire can carry a split-sample test of its own design: two wordings, two option orders, two scale lengths, randomly assigned. This turns a wording argument into a measurement, and it is the only way to establish the size of a suspected framing effect. It costs base size, so it needs planning in the sampling stage, not at the end.

**Modular design under stakeholder pressure.** Where multiple stakeholders each want content, build a core all respondents see and rotate optional modules across random subsamples. Everyone gets their question, the instrument stays short for each respondent, and the cost is paid in base size on the modules, transparently, rather than in data quality everywhere, invisibly.

**When the standard approach does not fit.** Very low-incidence populations may need the screener split into a separate short instrument. Highly sensitive topics may need indirect techniques (list experiments, randomised response) that trade individual-level data for reduced bias. Both are specialist choices and both need the analysis approach agreed before fielding.

## 16. Skill chain

**Recommended previous skills**
- **01.04 Research Method Selection.** Confirms a survey is the right instrument and hands over the rationale for what it can and cannot answer.
- **01.06 Sampling Strategy.** Hands over the population definition, the subgroups that must be reportable and the base sizes available, which set the classification section and the routing limits.
- **01.07 Analysis Plan Development.** Hands over the tables, cuts and derived variables the questionnaire must produce, which is the analysis-output column of the grid.

**Recommended next skills**
- **02.04 Question Bias Detection.** Takes the draft and audits wording independently. Design and review should not be the same pass.
- **02.05 Survey Logic and Flow Review.** Takes the routing specified at design level and reviews it in detail, including interaction with quotas.
- **02.06 Screener and Quota Design.** Takes the screener block and builds it out against the quota plan.
- **02.07 Scale and Measurement Selection.** Takes the scale decisions made at design level and settles point counts, labels, midpoints and comparability.

**Runs well alongside**
- **13.05 Research Ethics and Consent Design**, wherever sensitive data, minors or vulnerable participants are involved.
- **03.04 Fieldwork Monitoring and Response Quality**, which inspects in field the quality signals the soft launch first looks for.
- **02.02 Discussion Guide Design**, where qualitative work is needed before option lists can be written.

---
A Yazi Supplied Skill and resource.
