---
name: question-bias-detection
description: >
  Audits an existing survey, discussion guide or screener for bias, as a hostile
  reviewer rather than an author, and returns a severity-rated defect list with a
  specific rewrite for each. Use when someone says "is this question biased",
  "review this questionnaire", "check this survey before we field it", "this feels
  leading", "audit the instrument", "the client wrote these questions", or when a
  draft instrument has been produced by AI and needs checking before it goes out.
category: 02 Instrument Design
ref: 02.04
tier: 1
inherits: [K2, K3, K4, K5]
---

# Question Bias Detection

## 1. One-line description
Audits a drafted instrument as a hostile reviewer, question by question and whole-instrument, and returns every defect found with its mechanism, its likely direction of effect, its severity, a specific rewrite, and an honest statement of what the rewrite does not fix.

## 2. What this skill is used for

**The research problem it solves.** Authors cannot audit their own instruments, and not because they lack skill. An author reads what they meant. They know which word was chosen carefully, which list was researched, which scale was inherited, and that knowledge fills in every gap the respondent will meet cold. A reviewer who reads the draft the way a respondent will meet it, with no context and no charity, finds things the author cannot see by looking harder. Two further problems compound it. First, bias is asymmetric in visibility: a leading stem is obvious and a missing "not applicable" is invisible, yet the second usually does more damage, because it converts a respondent who has no view into one who appears to have one. Second, and least discussed, the most consequential bias in most instruments is not in the wording at all. It is in the objectives, which frequently presuppose the answer, and no question-level review will ever find it, because every question is faithfully serving an objective that should not have been agreed. A separate audit pass, run by someone who did not write the thing and is not trying to be helpful, is the only reliable control.

**Where it sits in the research lifecycle.** After an instrument exists and before it is fielded, translated, programmed or costed. It runs against instruments written by a colleague, by a client, by a stakeholder, by a previous wave, or by an AI system, which has its own characteristic defects. It is also the first thing to run when a fielded study produced an implausibly clean result.

**Typical use cases.**
- Reviewing a questionnaire before programming, when changes are still cheap.
- Auditing a client-supplied or stakeholder-supplied instrument and returning the case for each change in a form that survives a defensive conversation.
- Checking an AI-generated draft, which is typically fluent, symmetric, plausible and quietly biased in specific repeating ways.
- Reviewing a discussion guide's questions and probes for leading, imported vocabulary and unbalanced probing attention.
- Diagnosing a fielded instrument where a result looks too good, a distribution is implausibly skewed, or a tracked measure moved for no visible reason.
- Reviewing a translated or adapted instrument for terms that are neutral in one language and loaded in another.

**Who uses it.** Research directors and managers reviewing work before it ships; client-side insight managers challenging an agency draft; researchers auditing their own instrument after a deliberate gap; quality and governance functions; anyone handed a questionnaire and asked whether it is any good.

## 3. When to use it

- A questionnaire, guide or screener is drafted and about to be programmed, translated or fielded.
- A stakeholder or client supplied questions and someone has to say, with evidence, which ones cannot be used.
- An instrument was produced with AI assistance and has not been independently reviewed.
- A tracked instrument is being changed and the change needs to be assessed before the trend is affected.
- A study returned a result nobody believes, or one everybody believes a little too readily.
- An instrument is being reused from another market, another audience or another study, where wording that was neutral may no longer be.
- The research is high-stakes: a regulatory submission, a public claim, a pricing decision, an investment case, a published finding.
- The author is the person with an interest in the answer, which is the condition under which framing bias is most likely and least visible.

## 4. When NOT to use it

- **The instrument does not exist yet.** An audit needs an object. For authoring a survey, **02.01 Survey Questionnaire Design**; for a discussion guide, **02.02 Discussion Guide Design**; for question wording inside a guide, **02.03 Interview Question Development**. This skill does not author instruments. Where the audit finds an instrument beyond repair, it says so and hands back to the authoring skill rather than quietly rewriting the whole thing, which converts a review into an unrequested redesign and removes the author's ownership.
- **The problem is routing rather than wording.** Skip patterns, orphan questions, dead ends, base definitions, piping and quota interaction are a different audit with a different method, and a bias review will miss all of them. **02.05 Survey Logic and Flow Review**.
- **The problem is which scale to use rather than whether this one is balanced.** This skill flags an unbalanced or inappropriate scale as a defect. Deciding what should replace it, including point count, labelling, midpoint and comparability, belongs to **02.07 Scale and Measurement Selection**.
- **The problem is who is being asked rather than what.** A perfectly neutral instrument administered to the wrong people produces a biased result that no wording review can detect. **02.06 Screener and Quota Design** and **01.06 Sampling Strategy**.
- **The instrument is already in field and the wave is committed.** Auditing at that point produces findings nobody can act on and a report that undermines the study without improving it. The correct output then is not a defect list but a limitations statement, written for the report, naming the direction of each residual bias. Say this rather than delivering an audit that reads as an accusation.
- **The objectives themselves are unsound and the client will not revisit them.** An instrument faithfully serving a leading objective cannot be fixed at the question level. Escalate to **01.01 Research Brief Interrogation** or **01.02 Business Problem to Research Question**. Auditing the wording in this situation produces a clean instrument that will still return the answer someone wanted.
- **The purpose of the review is to justify a decision already taken.** An audit commissioned to bless an instrument, or to kill it, is not an audit. Say so.
- **Bias in analysis, interpretation or reporting.** This skill stops at the instrument. Selection bias in analysis, confirmation bias in interpretation and framing bias in a report belong to **13.04 Bias Detection**.

## 5. Required inputs

**Required. Without these the skill cannot run.**
- **The instrument in the form respondents will meet it**, including question order, response options in field order, randomisation and rotation instructions, introductions and any preamble text. An audit of question text stripped of its options and its order will miss most of what matters. If only partial material is available, state exactly what was reviewed and what was not, and do not present the audit as complete.
- **The research objectives**, because a question can only be judged biased relative to what it is supposed to measure, and because the objectives themselves are part of the audit. If they are not supplied, ask once, then run the audit and mark the objective-level pass as `NOT PERFORMED`, which is a material limitation and is stated at the top of the output.
- **The target population and mode.** A term that is plain for one audience is jargon for another, and a scale that works on a desktop grid is a defect on a phone. Without these, assume a general-population, small-screen, self-completion context and state the assumption.

**Optional, and what each one adds.**
- **The analysis plan.** Lets the audit judge whether a defect is recoverable at analysis or fatal, which is the distinction that makes a defect list actionable rather than merely correct.
- **Previous wave wording.** Distinguishes a defect that must be fixed from a defect that is locked into a trend, where fixing it costs more than leaving it and disclosing it.
- **Sponsor visibility and invitation text.** The single most under-audited source of framing. What the invitation said, and whether the respondent knows who is asking, changes the answers to everything.
- **The source of each option list.** A list derived from qualitative work is a different object from one written from the client's product taxonomy, and only the second is likely to be missing the options respondents would have chosen.
- **Language and translation status.** Enables the loaded-term and equivalence pass. Without it, that pass is marked as not performed rather than assumed clean.
- **Known incidence or expected distributions.** Lets the audit predict which questions will produce unusable distributions before fielding rather than after.

## 6. Questions to ask before starting

1. **Who wrote this, and what answer would suit them?** Not an accusation, a diagnostic. Framing bias clusters around the questions whose answers matter to the author, and knowing where to look doubles the yield of the audit. Default if unanswered: assume the sponsor has a preferred answer on the primary measure and audit that section hardest.
2. **What has already been fixed and cannot change?** Tracked wording, regulatory phrasing, a client's mandated brand battery, a translated instrument already signed off. Defects in immovable material still get logged, but they are reported as disclosures for the limitations section rather than as change requests. Default: assume everything is changeable and mark anything the author identifies as locked.
3. **When does this field, and what is the cost of each class of change?** A defect list delivered two days before programming, sorted by severity, is useful. The same list unsorted is not. Default: assume programming has not started and flag any defect that would require re-programming.
4. **What is the mode and the device profile?** Determines whether a long option list, a wide grid, a horizontal scale or an aural list is itself a defect. Default: small-screen self-completion, stated.
5. **Does the respondent know who is asking?** Sponsor awareness inflates ratings of the sponsor, suppresses criticism, and changes what "would you recommend" means. Default: assume the respondent can infer the sponsor from the invitation or the content, and audit accordingly.
6. **What will each question be used for?** A defect in a question that feeds the headline is a different problem from the same defect in a question nobody will tabulate. Default: treat every question as reportable, and flag any that appear to serve no analysis.

## 7. Step-by-step methodology

**Step 0. Adopt the reviewer's position, and say what you were given.** Read the instrument as a respondent will meet it: cold, in order, with no access to the author's intent, no charity about what a word probably means, and no assumption that anything was considered. Where you find yourself thinking "they obviously mean X", that is a defect, not a resolution. Then record the audit scope: what was supplied, what was not, and which passes could therefore not be run. An audit that quietly omits the objective-level pass because no objectives were supplied is worse than one that says so.

**Step 1. Audit the objectives, the introduction and the framing, before any question.** This is the pass most reviews skip and the one that finds the largest defects. Four checks. **Do any objectives presuppose their answer?** "Understand why customers value the new service" cannot return the finding that they do not. "Identify the barriers to adoption" presupposes that adoption is desirable and the obstacle is on the customer's side. Rewrite the objective symmetrically and see whether the instrument still makes sense; frequently half of it does not. **Does the invitation or the survey introduction prime?** An introduction that names the sponsor, explains why the topic matters, thanks the respondent for being a valued customer, or describes what the organisation is trying to do, has framed everything that follows. **Is the universe of possible answers constrained by the objectives?** If every objective concerns the organisation's own service, the respondent has no route to say that the category is irrelevant to them, and the report will not contain that finding. **Is there a missing objective whose absence is the bias?** Where a study asks how to improve something, and never asks whether it should exist, the omission is a design choice with a direction. Log objective-level findings separately from question-level ones, because they cannot be fixed by a rewrite and they usually need a conversation with whoever agreed the brief.

**Step 2. First pass, whole instrument, in order, as a respondent.** Read start to finish once without stopping to annotate individual questions, and record only whole-instrument effects, because these are invisible when reading question by question and are the ones that survive every other review. Look for: **order effects** (an aided list before an unaided question on the same subject, a diagnostic before an overall rating, a specific complaint before a general satisfaction measure); **priming** (a block establishing a frame that the key question then sits inside, an educational preamble, a definition that carries an evaluation); **carryover and consistency pressure** (a respondent who has just committed to a position will answer the next question consistently with it, whether or not that is what they think); **anchoring** (a number, a price, a scale range or an example that sets the scale for later answers); **battery fatigue** (three grids in a row, after which the data is straightlining regardless of wording); and **the cumulative signal** (what does the whole instrument, taken together, tell the respondent the researcher wants to hear). Record the position of each effect and what it contaminates, because these defects are located between questions rather than in them.

**Step 3. Second pass, stem by stem.** Now go question by question against the stem catalogue, running one defect class across the whole instrument at a time rather than all classes on each question, because the eye is far better at pattern-matching than at checklisting.

| Stem defect | Detection test | Mechanism |
|---|---|---|
| **Leading** | Does the stem suggest which answer is expected or approved? | Shifts the distribution toward the suggested response by an unknown and unrecoverable amount |
| **Loaded** | Strip every adjective and adverb. Does the meaning change? | Measures agreement with the framing rather than the proposition |
| **Double-barrelled** | Can a respondent hold different views on the two parts? Look for "and", "or", "as well as" | The answer is uninterpretable and the analyst cannot tell which half drove it |
| **Assumed premise** | Does the stem presuppose the respondent did, felt, noticed or knew something? | Respondents who did not are forced into a false answer or a break-off |
| **Undefined quantifier** | "Regularly", "recently", "typically", "often", "main", "household" | Different respondents answer different questions; the variance looks substantive |
| **Jargon and internal language** | Would a respondent outside the category use this word unprompted? | Non-response, guessing, and a bias toward respondents who happen to know the term |
| **Absolute framing** | "Always", "never", "all", "completely" | Pushes respondents to disagree with a true proposition stated too strongly |
| **Negation and double negation** | Any "not", especially in an agree/disagree item | Misreading rates are high and the error is not random across literacy levels |
| **Sponsor tell** | Does the wording reveal who is asking or what they hope? | Inflates ratings of the sponsor and suppresses criticism |
| **Hypothetical and predictive** | Does it ask what the respondent would do? | Stated intention over-reports approved and novel behaviour, systematically |
| **Recall beyond capacity** | Does it ask for a count of a routine low-salience act over a long window? | Rounding to salient numbers, anchoring to the options offered, and a false appearance of precision |

**Step 4. Third pass, response frames.** Answer options carry as much bias as stems and receive a fraction of the scrutiny, so this pass is where a careful audit earns its fee. For every closed question, check: **balance** (equal numbers of positive and negative points, and symmetric label intensity, so that "excellent, very good, good, fair, poor" is caught as three-against-one); **exhaustiveness** (can every respondent find themselves, and what happens to the ones who cannot); **mutual exclusivity** (could a respondent legitimately choose two); **numeric ranges** (overlaps at the boundary, gaps, and open ends that swallow the interesting tail); **list balance** (a list of eight positive attributes and two negative ones is a leading question with no leading words, and a list containing every one of the sponsor's features and two of the competitor's is a market-share estimate of the list-writer's mind); **option order** (fixed lists accumulate primacy bias on screen and recency bias when read aloud, so any list over about five items without a rotation instruction is a defect); and **missing non-substantive options**. That last one is the most frequently missed defect in commercial research, so test each explicitly: is "don't know" needed here because not knowing is a genuine state; is "not applicable" needed because some respondents will not be in scope for this question; is "none of these" needed because selecting nothing is possible on this multi-select; is "prefer not to say" needed because this question is sensitive. Each absence has a direction: a missing "don't know" inflates whichever substantive answer is most available, a missing "not applicable" inflates the midpoint or the neutral option, a missing "none of these" inflates the nearest listed option, and a missing "prefer not to say" produces break-off or a false answer. State the direction; do not just flag the absence.

**Step 5. Fourth pass, response style and population.** These are properties of the interaction between the instrument and the people answering it, and none of them is visible in a single question. **Acquiescence:** count the agree/disagree items and the yes/no items, and count how many run in the same direction. A long same-direction agree battery will produce a correlated block that looks like a construct and is partly a response style, and the effect is stronger in some populations than others, so it also biases subgroup comparisons. **Social desirability:** identify every question with an approved answer (exercise, saving, voting, recycling, safety, healthy eating, charitable giving, reading terms and conditions, professional compliance) and check whether the wording makes the unapproved answer easy to give. **Extreme and midpoint responding:** flag where cross-cultural or cross-market comparison will be made on scale means, because response-style differences across cultures can exceed the substantive differences being measured. **Culturally and linguistically loaded terms:** words that carry evaluation in one language or community and not another, terms with a clinical meaning and a colloquial one, category words that imply a judgement (calling a payment method "convenient", a product "premium", a behaviour "failure to"), and terms whose translated form is not equivalent in register. **Literacy and cognitive load:** sentence length, subordinate clauses, and any question requiring the respondent to hold more than about two facts in mind. Where the instrument will be translated, mark every term whose equivalence needs checking rather than assuming translation will handle it.

**Step 6. Rate each defect for severity, using a stated rubric.** A defect list without severity is a wish list, and reviewers who mark everything as important get everything ignored.

| Severity | Definition | Action |
|---|---|---|
| **Critical** | The question cannot be interpreted, or will produce a confidently wrong answer that looks correct. Includes: assumed premises that exclude legitimate respondents, double-barrelled questions feeding a headline measure, missing options that force false answers, and irreversible order contamination | Do not field without change |
| **Major** | A known directional bias of material size on a question that will be reported | Change before fielding; if immovable, disclose the direction in every output using the question |
| **Minor** | A defect that adds noise or reduces precision without a clear direction | Change if the cost is low; otherwise note |
| **Note** | An observation, a house-style inconsistency, or a risk that depends on how the data is used | Record only |

Two rules keep the rubric honest. Severity is set by consequence, not by how obviously wrong the question looks: a badly worded question nobody will report is minor, and a subtly unbalanced scale on the primary measure is critical. And severity rises when the defect is unrecoverable, which is the next step.

**Step 7. Classify each defect as recoverable or unrecoverable at analysis.** This is the distinction that makes an audit useful to a research director. A missing "don't know" is unrecoverable: the respondents who did not know are now inside the substantive distribution and cannot be identified. An unbalanced scale is partially recoverable: the level is wrong but the comparison between subgroups may hold, because the bias applies to everyone. A double-barrelled question is unrecoverable, because no analysis separates the two halves. Order contamination is unrecoverable within the wave, though sometimes detectable if the order was randomised. A leading stem is unrecoverable in level and may be recoverable in trend if the same wording ran last wave. Say which for every defect, because a fielded instrument's limitations section is built from exactly this column.

**Step 8. Write a specific rewrite for every defect above Note.** A flag is not a finding. Each rewrite is the actual replacement text, with its response options where the frame changes, and it must be materially different rather than a synonym swap. Then state, in one line, what the rewrite does not fix. A leading question rewritten neutrally still sits after the block that primed it. A missing "prefer not to say" added to an income question does not remove the mode effect. Naming the residual is what stops the audit from over-promising, and per **K3 §4.2** it is the caveat that must travel with the claim.

**Step 9. Re-audit the rewrites.** Rewrites introduce defects at a rate high enough that skipping this step is a known way to make an instrument worse. The commonest are: fixing a leading stem by making it long and cognitively harder; fixing a double-barrelled question by splitting it and blowing the length budget; adding "don't know" to an attitude question where everyone has a view, which manufactures non-response; adding a midpoint to force balance where no genuine neutral exists; and lengthening an option list past the point a respondent will read it. Run Steps 3 to 5 over the rewritten text.

**Step 10. Assemble the register, then say whether it can be fielded.** The defect register is the deliverable, but a register alone leaves the decision unmade. Close with an explicit statement: field as is, field with the critical defects fixed, or do not field. Where objectives-level bias was found, say plainly that no question-level change addresses it, and route it upward. Per **K5 §2.7** the decision to field with known defects is the researcher's, not the auditor's, and the auditor's job is to make sure it is made knowingly.

## 8. Analytical framework

Every defect is worked through one chain, and a defect that cannot complete the chain is not yet a finding:

    Defect → Mechanism → Direction of effect → Severity → Recoverable at analysis? → Rewrite → Residual

**Mechanism** is what forces the answer to move: not "this is leading" but "the stem supplies an evaluative adjective, so respondents with no view adopt it". Without a mechanism, the audit is an opinion, and an author can dismiss it as taste.

**Direction** is which way the data will be wrong. Most defects have one and it can be stated, and stating it is what allows a researcher to decide whether the defect matters for this study. Where the direction genuinely cannot be predicted, say so, because "adds unpredictable noise" is a different and usually lesser problem than "inflates the headline by an unknown amount in a known direction".

**Recoverability** separates the defects that must be fixed now from those that can be managed at analysis, and it is the column that turns an audit into a plan.

**Residual** is what remains after the rewrite. It is the honesty check on the whole document, and it is also the raw material for the study's limitations section, which per **K2** is where the audit's value survives into the eventual report.

## 9. Output format

**1. Audit scope and limitations.** What was supplied and reviewed, what was not, which passes were run, which were not and why, and the assumptions made about mode, population and sponsor visibility.

**2. Verdict.** One paragraph: field as is, field with critical defects fixed, or do not field. With the count of defects by severity.

**3. Objective and framing findings**, separately from question findings, because they cannot be fixed by rewriting.

| Ref | Objective or framing element | Problem | Consequence for the study | What is needed |
|---|---|---|---|---|

**4. Defect register**, the core deliverable, sorted by severity then by question number.

| Ref | Q no. | Defect type | Quoted text as written | Mechanism | Direction of effect | Severity | Recoverable? | Rewrite | Residual |
|---|---|---|---|---|---|---|---|---|---|

**5. Whole-instrument findings**, which have no single question number: order, priming, carryover, battery placement, cumulative signal, and the sponsor question.

**6. Response-style and population findings**, including acquiescence load, social desirability exposure, translation and equivalence risks.

**7. Rewrite re-audit note**, confirming Step 9 was run and recording anything it caught.

**8. Limitations text**, drafted ready for the eventual report, covering every defect the researcher decides not to fix.

**9. Open items and review points**, per **K5 §3**.

**When the inputs are thin**, the format does not get filled in anyway. A pass that could not be run is marked `NOT PERFORMED` with the reason, not omitted silently. A direction of effect that genuinely cannot be predicted is written as unknown rather than guessed, per **K4 §1**. No defect is invented to populate a category, and a clean pass is reported as clean, with what was checked.

## 10. Quality checks

Run before the audit is delivered. These sit on top of **K4 §8**.

1. The objectives and the introduction were audited, or the omission is stated at the top of the output.
2. The instrument was read once end to end in field order before any question was annotated.
3. Every defect names its mechanism, not just its label.
4. Every defect states a direction of effect, or explicitly states that the direction is unknown.
5. Every defect above Note has a specific rewrite in full text, not a description of a rewrite.
6. Every rewrite states its residual: what it does not fix.
7. Every rewrite has itself been audited against the same catalogue.
8. Severity has been assigned by consequence, not by conspicuousness, and the rubric is stated in the output.
9. Every question's response frame was audited, not only its stem.
10. Non-substantive options were checked question by question, with the direction of any absence stated.
11. Order, priming and carryover findings are recorded as whole-instrument defects, with the positions they contaminate.
12. Acquiescence load was counted, not estimated by impression.
13. Immovable defects (tracked wording, regulatory phrasing) are recorded as disclosures rather than change requests.
14. The output ends with an explicit field or do-not-field verdict, and the decision is left with the researcher.
15. Nothing in the audit claims the rewritten instrument is unbiased.

## 11. Common failure modes

| Failure | How to recognise it | How to prevent it |
|---|---|---|
| **The author audits their own draft** | The review finds three typos and no defects | A different reader, or the same reader after a deliberate gap and reading in field order as a respondent |
| **Stem-only review** | Every finding concerns question wording; not one concerns an option list | Step 4 is a separate pass, run across the whole instrument |
| **Flag without rewrite** | A list of adjectives ("leading", "loaded") with no replacement text | Every defect above Note carries the actual replacement wording |
| **Undifferentiated severity** | Twenty-eight defects, all described as important | Apply the rubric by consequence. If everything is critical, nothing is |
| **The missed introduction** | Every question audited; the survey preamble, the invitation and the sponsor disclosure never read | Step 1. The framing is part of the instrument and is often the largest single effect in it |
| **Objectives taken as given** | A flawless audit of an instrument faithfully serving a leading objective | Step 1's first check. Rewrite each objective symmetrically and see what survives |
| **Rewrite introduces a worse defect** | The fixed question is longer, harder, and now has a manufactured midpoint | Step 9, without exception |
| **Audit as accusation** | The register reads as a case against the author and gets rejected wholesale | Report mechanism and direction, not judgement. The finding is about the data, not the person |
| **Ignoring what cannot change** | Tracked wording flagged as a defect to fix, and the trend broken to fix it | Locked material is disclosed, not changed. The trend is usually worth more than the fix |
| **The clean bill of health** | An audit that finds nothing and says the instrument is unbiased | No instrument is unbiased. Report what was checked, what was found, and what remains |
| **AI: fluency mistaken for neutrality** | A generated instrument reads beautifully and every scale is unbalanced in the same direction | Audit AI drafts against the same catalogue with extra attention to option lists and scale symmetry, which is where generated instruments fail most consistently |
| **AI: invented option lists** | Brand lists, feature lists, competitor sets and category descriptors that no source supplied | **K4 §2.1**. Ask where each list came from; an unsourced list is a defect of unknown direction |
| **AI: symmetric batteries everywhere** | Every construct measured by the same five-point agree/disagree grid | Flag the acquiescence load and route the redesign to **02.07** |
| **AI: the plausible defect** | An audit that reports defects that are not there, because the format expected findings | **K4 §1**. Quote the text as written for every defect. A defect that cannot be quoted does not exist |
| **AI: agreeing with the author** | An audit of a supplied instrument that finds it broadly sound and suggests minor polish | The reviewer's position is adversarial by design. An audit with no critical or major findings on a real draft is an unusual result and should be interrogated |

## 12. AI guardrails

Skill-specific only. **K4** applies in full and is not repeated here.

1. **Never report a defect without quoting the text as written.** A defect that cannot be quoted from the supplied instrument is a hallucinated defect, and it destroys the author's trust in every real finding in the register (**K4 §1, §6.4**).
2. **Never audit an instrument you were not fully given.** If only the question text was supplied without options, order or introduction, say so and mark the corresponding passes `NOT PERFORMED`. Do not infer the missing parts and audit the inference.
3. **Never state a bias magnitude as a measured value.** "This will reduce agreement by around 10 points" is a fabricated effect size unless it comes from a cited split-sample test. Direction can be stated; size is stated only where evidence supports it (**K4 §2.1, §2.4**).
4. **Never present a rewrite as unbiased.** Every rewrite carries its residual. An audit that hands back a clean instrument has overstated its own product.
5. **Never quietly rewrite beyond the defect.** A rewrite addresses the defect found. Redesigning the question, the scale or the section is a different job and belongs to the authoring skill, with the author's agreement.
6. **Never soften a critical finding because the author is the client.** State the mechanism and the consequence plainly, offer the rewrite, and leave the decision with the researcher (**K4 §9**, **K5 §2.7**).
7. **Never manufacture findings to fill the register.** A section that audits clean is reported as clean, with what was checked. The pressure to produce defects because a review was commissioned is the same structural pressure K4 §1 exists to resist.
8. **Never treat the absence of a flagged term as evidence of neutrality.** Bias by option list, by order, by omission and by objective produces no flaggable phrase, and these are the defects that survive automated review.

## 13. Best-practice principles

1. **Read it as a respondent, in order, once, before you annotate anything.** The defects that live between questions are invisible to a question-by-question review, and they are frequently the largest.
2. **Bias you can see is rarely the bias that matters.** A leading adjective is caught by everyone. A missing "not applicable", an unbalanced option list and a priming preamble are caught by almost no one and do more damage.
3. **Read the option lists twice and the stems once.** The scrutiny distribution in most reviews is exactly the wrong way round.
4. **Every defect has a direction. Find it.** "This is biased" invites argument. "This will inflate the top-two-box by an unknown amount, in this direction, and the effect is not recoverable at analysis" ends it.
5. **Severity is about consequence, not conspicuousness.** The question nobody will report can be badly worded without costing anything. The primary measure cannot.
6. **A rewrite is the deliverable; a flag is a complaint.** Reviewers who supply the replacement text get their changes made, and reviewers who supply adjectives get a defensive meeting.
7. **Audit your own rewrites.** The correction is written quickly, under the satisfaction of having found something, and that is exactly when a new defect gets introduced.
8. **The objectives are part of the instrument.** A question can be perfectly neutral and still be unable to return the finding that would matter, because the objective never permitted it.
9. **What the respondent was told before question one is part of the instrument too.** The invitation, the sponsor, the incentive and the introduction frame everything, and they are usually written by someone who has never read the questionnaire.
10. **Distinguish the fixable from the disclosable.** Some defects are locked into a trend or a regulation. The right output for those is a limitations sentence, not a change request, and writing that sentence now is a gift to the person who will write the report.
11. **Never certify an instrument as unbiased.** Report what was checked, what was found, what was fixed and what remains. Certification is a claim nobody can support.
12. **An audit that finds nothing on a real draft has probably not been run properly.** Instruments written by competent people still contain major defects, because authors cannot see their own assumptions, which is the entire reason this pass exists.

## 14. Worked example

**INPUT**

A fictional retail bank, Northwater Bank, has drafted a 14-question customer survey ahead of a decision on whether to change its overdraft fee structure. The stated objectives are: (a) understand why customers value Northwater's overdraft facility; (b) identify barriers to customers using the new mobile alerts feature; (c) measure satisfaction with the current fee structure. Mode: self-completion, emailed to the customer base, mobile-dominant. The survey introduction reads: "As a valued Northwater customer, we would like your views on how we can keep improving the services you rely on."

**PROCESS**

*Step 1, objectives and framing.* Two objective-level findings, both critical, and neither fixable by rewriting a question. Objective (a) presupposes that customers value the facility, so no wording of any question beneath it can return the finding that some customers resent it, avoid it, or do not know they have it. Rewritten symmetrically as "understand how customers regard the overdraft facility, including those who avoid or resent it", which immediately revealed that the instrument contained no route for a customer to say the facility had cost them money unexpectedly, which is the finding the fee decision actually turns on. Objective (b) presupposes that non-use is a barrier problem on the customer's side rather than a value problem with the feature. The introduction was flagged separately: "valued customer" and "keep improving" prime positively, and "the services you rely on" presupposes reliance. Rewritten neutrally, with the sponsor named plainly and no evaluative framing.

*Step 2, whole instrument.* Three whole-instrument findings. The instrument opens with four questions on the mobile app's features, then asks overall satisfaction with Northwater at Q9: the overall measure now sits downstream of a block that has reminded the respondent of every feature they have, which primes it upward. An aided list of banking features at Q3 precedes an unaided question at Q6 about what the respondent uses most, which is irreversible contamination. And Q11 asks about fee fairness immediately after Q10 has stated the current fee, which anchors it.

*Step 3, stems.* Q7, "How helpful did you find the new mobile alerts in helping you avoid unexpected charges?", is critical on three counts: leading ("helpful"), double-barrelled (helpfulness of the alerts and the avoidance of charges are separate), and carrying an assumed premise (that the respondent received alerts and had charges to avoid). Q12, "Do you agree that a simple, transparent fee structure is better for customers?", is loaded to the point of being unanswerable in the negative and measures agreement with the framing rather than any position on Northwater's fees.

*Step 4, response frames.* The satisfaction scale at Q9 runs "extremely satisfied, very satisfied, satisfied, not very satisfied", which is three positive points against one negative, guaranteeing a high mean before anyone answers, and it feeds the headline. Rated critical because it is the primary measure and the resulting level is uninterpretable. Q5, a multi-select of reasons for using the overdraft, contains seven reasons, six of which are practical and none of which permits "I did not realise I had gone into it", which staff feedback and complaints data both suggested is common. Rated critical: the missing option forces respondents into a false answer, and the direction is a systematic understatement of the exact behaviour the fee decision concerns. Q13, on household income, has no "prefer not to say".

*Step 5, response style and population.* Six of the fourteen questions are agree/disagree items, all worded in the same direction, and all concerning Northwater's own performance. The acquiescence load was recorded as a major finding with a stated direction, and the redesign routed to **02.07** rather than resolved in the audit.

*The judgement call.* Q9's unbalanced satisfaction scale has run unchanged for six waves and the tracked series is used in the bank's published customer commitments. Fixing it breaks the trend. The audit did not decide this. It recorded the defect as critical for level and noted that the bias applies equally in every wave, so wave-on-wave movement remains readable even though the absolute number is not, and it presented two options with their costs: change and re-baseline with a parallel run, or retain and disclose that the absolute satisfaction figure is not comparable to any externally scaled measure. Marked **RESEARCHER DECISION REQUIRED** per **K5 §2.7**, because it trades measurement quality against a public commitment, which is a business judgement the audit does not own.

*Steps 8 and 9, rewrites and re-audit.* Q7 split into two questions, one on receipt and one on usefulness, with the premise removed and a routing note handed to **02.05**. Q12 rewritten as a two-sided stem offering both positions. Q5 given the missing option plus "other, please specify" and "none of these". The re-audit caught that the first rewrite of Q7 had introduced a "don't know" on an attitude question where every recipient would have a view, manufacturing non-response; it was removed and replaced with a "not applicable" on the receipt question instead, where not receiving alerts is a genuine state.

**OUTPUT**

An audit with two critical objective-level findings routed to **01.01**; a defect register of 21 entries (5 critical, 9 major, 5 minor, 2 notes), each with quoted text, mechanism, direction, recoverability and full rewrite text; three whole-instrument findings on order, priming and anchoring with the positions they contaminate; a response-style finding on acquiescence load routed to **02.07**; drafted limitations text covering the tracked scale under either decision; a verdict of do not field without the five critical fixes; and one **K5** decision point on the tracked satisfaction scale.

## 15. Advanced usage

**Auditing an AI-generated instrument.** Generated drafts fail in a characteristic and repeatable pattern, so the audit can be targeted. They are fluent, so stem-level defects are rarer than in human drafts, and the reviewer's attention should move immediately to the places generation is weakest: option lists (frequently invented, frequently unbalanced, frequently containing every positive attribute and one token negative), scale symmetry (label intensity is often asymmetric even when the point count is even), non-substantive options (usually omitted, because the generator produces the substantive frame and stops), acquiescence load (agree/disagree batteries proliferate because they are easy to generate symmetrically), and length (generated instruments are consistently longer than the objectives justify, with no removals). Ask for the provenance of every substantive list; per **K4 §2.1** an unsourced list is a defect of unknown direction, not a neutral default.

**Split-sample testing to settle an argument.** Where the author disputes that a defect matters, and the study can afford it, the disagreement can be measured rather than debated: field both wordings to random halves and compare. This converts a taste argument into an effect size, and it is the only way to establish the magnitude of a suspected framing effect rather than its direction. It costs base size, so it belongs in the sampling plan, and it is worth doing on a tracked question before wording is locked for years.

**Auditing across languages.** Run the audit on the source instrument first, then again on each translated version with a native researcher, because equivalence is not a translation property. Look specifically for: scale labels whose intensity differs after translation (the second point of a five-point scale is rarely equivalent across languages), terms that are clinical in one language and colloquial in another, politeness registers that make disagreement harder, and negation structures that do not survive. Where cross-market comparison is the objective, flag response-style differences as a design-level risk for the analysis plan rather than a wording defect.

**Auditing after fielding.** When a study is complete and a result looks wrong, the same passes are run diagnostically rather than correctively, and the data itself becomes evidence. High "other" verbatim counts indicate a missing option; high item non-response indicates an unanswerable or unacceptable question; implausible skew on a rating indicates an unbalanced scale or a leading stem; straightlining concentrated in one battery indicates fatigue or acquiescence; and a distribution that differs sharply between rotated and non-rotated positions indicates an order effect large enough to see. The output is a limitations statement and a wording specification for the next wave, not a defect list nobody can act on.

## 16. Skill chain

**Recommended previous skills**
- **02.01 Survey Questionnaire Design** or **02.02 Discussion Guide Design.** Hand over the instrument, the objective-to-question mapping and the design rationale, which tell the auditor what each question was meant to do.
- **01.07 Analysis Plan Development.** Hands over what each question feeds, which determines severity and recoverability.

**Recommended next skills**
- **02.05 Survey Logic and Flow Review.** Takes the corrected instrument and audits its routing, which this skill does not cover.
- **02.07 Scale and Measurement Selection.** Takes the scale defects found here and settles what should replace them.
- **02.01** or **02.02**, where the audit concludes the instrument needs re-authoring rather than repair.

**Runs well alongside**
- **13.04 Bias Detection**, the cross-cutting skill that applies the same discipline to analysis, interpretation and reporting.
- **01.01 Research Brief Interrogation**, wherever the audit finds the bias is in the objectives rather than the instrument.
- **03.04 Fieldwork Monitoring and Response Quality**, which detects in field the consequences of defects that were not caught here.

---
A Yazi Supplied Skill and resource.
