---
name: research-method-selection
description: >
  Chooses the research method that actually fits the question and explains why,
  including the honest verdicts a supplier avoids: that desk research may be
  enough, that the question is unanswerable as posed, and that no research
  should be commissioned. Use for "which method should we use", "qual or
  quant", "is a survey the right approach", "how should we research this",
  "can we answer this with what we already have", "what is the least-bad
  design on this budget".
category: 01 Research Strategy and Design
ref: 01.04
tier: 0
inherits: [K2, K3, K4, K5]
---

# Research Method Selection

## 1. One-line description
Selects the research method whose evidence can carry the claim the decision requires, states what that method cannot deliver, and records why the alternatives were rejected.

## 2. What this skill is used for

**The research problem it solves.** Method is chosen badly more often than it is chosen well, and almost always for reasons that have nothing to do with the question: habit, the shape of the budget, what a stakeholder finds persuasive, or whichever method the person answering has most practice in. The consequence is not usually a visibly failed study. It is a competently executed study that answers a question nobody needed answered, whose findings are then stretched to cover the question that mattered. Method error is unrecoverable downstream, because no amount of analytical care rescues a design that could never have produced the evidence required. This skill forces the choice to start from the shape of the question and the claim the decision needs, and makes the reasoning inspectable.

**Where it sits in the research lifecycle.** After the business problem has been converted into research questions (01.02) and hypotheses articulated where they exist (01.03). Before the plan, the sample and the instrument. It is the hinge of the design phase.

**Typical use cases.**
- A brief names a method ("run a survey") and the method does not fit the question.
- It is unclear whether any research is needed.
- Several objectives sit in one brief and need different methods, or need sequencing.
- The ideal design is infeasible and the least-bad alternative must be chosen and defended.
- A proposed design has to be reviewed before commissioning.
- A tracker is being re-designed and continuity constrains the options.

**Who uses it.** Research managers and directors specifying studies; client-side insight leads deciding what to commission; consultants writing proposals; product and UX researchers choosing between study types; postgraduate researchers selecting an approach.

## 3. When to use it

- You have research questions but no agreed method, and the choice is still open.
- A method has been named in a brief and you need to test whether it fits, or replace it.
- The brief contains a mix of question types and you suspect one method is being asked to do all of them.
- The budget, deadline or audience access has changed and the design has to be re-cut.
- You need to justify a method choice to a stakeholder who prefers a different one.
- You are reviewing someone else's proposed design and need to know what it will not be able to say.
- You suspect the answer already exists in internal data or published sources and want that tested before commissioning.
- You need to decide whether to run one study or a sequence.

## 4. When NOT to use it

- **Before the research question exists.** If the objectives are still a topic ("understand our customers"), this skill will produce a plausible method for an unspecified question, which is worse than no method at all because it looks like progress. Use **01.02 Business Problem to Research Question** first, and return here.
- **To set sample size, quotas or the sample frame.** This skill states what sample structure a method implies and what would make it infeasible. The actual sampling design belongs to **01.06 Sampling Strategy**. A method recommendation that specifies exact cell sizes is overreaching.
- **To design the instrument.** Choosing depth interviews does not write the guide. Hand to **Category 02 Instrument Design**.
- **To choose a supplier, panel, platform or fieldwork partner.** This skill has no view on suppliers and must not develop one. Channel is discussed here only as a methodological variable.
- **When the decision has already been taken and the research is being commissioned to support it.** No method is correct for this. The honest output is to name what is happening and offer the two legitimate alternatives: a genuine test of the decision, which the commissioner may not want, or no study.
- **Mid-series on an established tracker, without the continuity constraint in front of you.** Recommending a better method for a tracking study is actively misleading if the recommendation breaks comparability with the existing series. The trade-off between a better instrument and an unbroken trend is a real decision, and it belongs with **05.05 Trend and Tracker Analysis** in the room.
- **When the question is whether the research is ethical or lawful to run.** Feasibility is not permission. Route to **13.05 Research Ethics and Consent Design** before any design is finalised, and always where participants may be vulnerable, where the topic is sensitive, or where a subgroup could be identifiable in the output.
- **When the population cannot be reached at all.** If there is no honest route to the people whose answers would settle the question, the correct output is that the study cannot be run as specified, not a weaker design presented as though it were an approximation of the right one.
- **For an academic methodology chapter requiring a paradigm and philosophy justification.** The logic here is decision-first and commercial. Use **15.08 Research Design and Methodology Chapter** or **15.16 Doctoral Methodology Justification and Rigour**.

## 5. Required inputs

**Required. Without these the skill cannot run.**

- **The decision the research serves, and who owns it.** Not the topic. The action that will or will not be taken. If absent, stop and ask. If nobody can name it, the correct output is that no method can be selected yet, and the work returns to 01.02.
- **The decision deadline.** A design that lands after the decision is made has zero value regardless of quality. Ask. If genuinely unknown, state the assumption and the elapsed time each option needs.
- **The research questions or objectives**, in the form 01.02 produces them.
- **The population of interest, and whether there is any known route to it.** If access is unknown, that is itself an output of this skill, not an input to be assumed.
- **The hard constraints.** Deadline, budget order of magnitude, internal analysis capability, regulatory or client restrictions. Constraints do not choose the method, but they determine which recommendations are real.

**Optional, and what each one adds.**

- **Hypotheses (01.03).** Convert a vague objective into a testable one, and often reduce the design from exploratory to confirmatory, which is cheaper and sharper.
- **An inventory of internal behavioural, transactional or operational data.** Makes it possible to conclude that part of the question is already answered, and to design the primary research around only the residue.
- **Prior studies on the same audience.** Establish what does not need re-measuring, and set the expected effect size that determines whether a quantitative design is worth powering.
- **The stakeholder audience for the findings.** Determines the defensibility threshold. A board investment case and an internal design sprint need different evidence for the same finding.
- **Tracker history and prior instrument.** Makes the continuity constraint visible before a design change breaks a series.
- **Known incidence of the target audience in the general population.** Turns feasibility from a guess into an estimate, and often rules out an approach immediately.

## 6. Questions to ask before starting

1. **What decision does this change, and when is it made?** Everything downstream is set by this. *Default if unanswered:* do not proceed. This is the one question that stops the work.
2. **What claim do you need to be able to make at the end, and to whom?** "Six of twelve customers described the same barrier" and "38% of the customer base, plus or minus 3 points" are different claims requiring different designs. Asking this early prevents a qualitative study being commissioned and then quoted as a proportion. *Default:* assume the finding will be quoted to a sceptical audience and design to the higher standard, and say you have assumed it.
3. **What is already known, measured or published?** *Default:* assume nothing is known, and flag that the desk research check has not been run, because it frequently removes an objective entirely.
4. **Which subgroups must be reported separately?** Subgroup reporting, not total sample, drives quantitative feasibility, and it is the most common reason a study is under-powered for the analysis everyone intended. *Default:* assume total sample only, and state that any subgroup reading would require re-specification.
5. **Is the audience reachable, honestly, and by what channels?** Low-incidence, high-value or offline audiences change the answer completely. *Default:* flag access as unverified and treat the design as conditional on a feasibility check.
6. **What is fixed and what is flexible: deadline, budget, or scope?** At least one usually gives. Knowing which one determines whether the answer is a narrower study, a slower one, or a different method. *Default:* assume deadline is fixed and scope is flexible, which is the most common real configuration, and say so.
7. **Does this need to be comparable to anything, past or future?** *Default:* assume it is a one-off, and note that if it is not, method change costs the comparison.

## 7. Step-by-step methodology

### Step 1. Recover the decision, and refuse to proceed without it

Write the decision as a sentence containing an actor, an action and a date: "The category team decides in March whether to fund a second phase of the self-service programme." If that sentence cannot be written, no method is selectable, and producing one anyway is the characteristic failure of this task. What a correct result looks like: one sentence, owned by a named role, with a date, plus a note on what happens by default if no research exists. That default matters. If the organisation will do the same thing either way, the value of any design is zero and the honest recommendation is no study.

### Step 2. Decompose the brief into question shapes

Briefs arrive as topics. Methods answer shapes. Split every objective until each fragment has exactly one shape. Most weak designs come from a single method being asked to cover three shapes at once.

| Shape | Sounds like | What it demands of the evidence |
|---|---|---|
| **How many / how much** | Incidence, size, frequency, proportion, magnitude | Measurement on a sample that supports estimation for the reporting groups |
| **Why** | Reasons, meaning, motivation, barriers, sense-making | Depth, follow-up, contradiction surfaced rather than resolved |
| **How** | Practice, sequence, workaround, what actually happens | Observation of behaviour, or capture close to the moment it occurs |
| **What would happen if** | Reaction, adoption, effect of a change | Manipulation and a comparison group, or the closest approximation available |
| **Which of these / what is it worth** | Ranking, priority, trade-off, willingness to give something up | Forced choice under constraint, not rating in isolation |
| **Is it changing** | Trend, movement, direction, before and after | Repeated measurement on a stable instrument |
| **Is it already known** | Established fact, context, market structure, precedent | Existing evidence, appraised for quality and currency |

What a correct result looks like: a table of question fragments, each carrying one shape, with the fragments that were merged in the brief now visibly separate. Frequently one brief yields four fragments and only two of them need primary research.

### Step 3. Convert each shape into the claim type and the evidence that would carry it

For each fragment, write the sentence you expect to be able to say at the end, then ask what would have to be in your hands for that sentence to be defensible. This is the step that catches unanswerable questions. "What features would make non-adopters switch?" expects a sentence about future behaviour under a product that does not exist, obtained by asking people who have never encountered it. No design delivers that reliably. The question is not expensive, it is unanswerable as posed, and the correct move is to re-specify it into something that is answerable: what stops them now, what have they done when they hit that barrier, and which of a defined set of changes they trade off against each other.

Three re-specifications recur often enough to check for by default. Stated future behaviour becomes current behaviour plus revealed trade-off. Attributed causes ("why did you leave?") become sequence reconstruction plus behavioural data. Importance ratings, which almost always come back uniformly high, become forced trade-off.

### Step 4. Run the four honest verdicts before selecting any method

Every fragment goes through these four tests before any method is named. A supplier's incentive is to skip this step. Skipping it is the most expensive thing in the design phase.

1. **Is it already answered?** Internal transaction, service, usage or CRM data answers "how many" and "how often" questions far better than asking people to recall. Published statistics answer market structure. If a fragment is already answered, remove it from the design and say where the answer is. See **10.01 Literature Review and Desk Research**.
2. **Is it unanswerable as posed?** If Step 3 could not name evidence that would carry the claim, say so plainly and offer the re-specification. Do not substitute a method that produces a number about a related thing.
3. **Is the quantitative version worth doing?** A quantitative design that cannot support the subgroup comparisons it exists to make is worse than a qualitative one, because it produces precise figures that invite over-reading. Where the reportable base for the cut that matters would fall below the thresholds in K3 §3.2, or where the expected difference is smaller than the design could detect, the honest recommendation is a smaller, well-run qualitative study that makes no numerical claim.
4. **Should any research be commissioned?** No, where the decision is already made, where the cost of the research exceeds the value of the decision, where the result would not change the action under any outcome, or where the timing means findings arrive after the decision. State it, and state what the money would buy instead: a better analysis of existing data, a smaller test, or a decision taken on judgement with a monitoring plan behind it.

What a correct result looks like: at least one fragment removed, re-specified or explicitly declined, or a positive statement that all four tests were run and all fragments survived.

### Step 5. Select the method family, then specify the design

Only now consult the repertoire in 7.9. Select on the shape and the claim type, not on the topic. Then specify: who, how many or how few, over what period, with what stimulus, and analysed how. A method name is not a design. "Depth interviews" is not a recommendation until it says with whom, how many, how recruited, how long, and what the analysis will produce.

Two specification rules carry most of the weight. **Sample follows the smallest unit you must report on**, not the total. **The analysis is specified before the fieldwork**, because a method whose analysis the team cannot run is not available, whatever the budget says. See **01.07 Analysis Plan Development**.

### Step 6. Choose the channel on methodological grounds and state what it costs

Channel is a design variable, not an administrative one. Each channel changes who is reachable, what can be asked, and how people answer. Choose it explicitly and disclose the cost.

- **Online self-completion.** Efficient, supports complex logic and visual stimulus, no interviewer effect. Excludes or under-represents people with poor connectivity or low digital confidence, and where samples are drawn from opt-in sources they carry professional-respondent and inattention risk. Never described as representative without naming the frame.
- **Telephone.** Reaches offline and older populations that online misses, allows clarification and light probing. Contact and cooperation rates are low and falling in most markets, calls skew the achieved sample, question formats are constrained to what can be held in the head, and social desirability rises with a live interviewer.
- **Face to face.** The only channel for extended stimulus, physical products, in-context observation, and populations with literacy or connectivity barriers. Expensive, slow, geographically clustered in ways that affect the sample, and subject to interviewer effects that need controlling.
- **Mobile and messaging-based.** Low friction, reaches mobile-first populations others miss, captures responses close to the moment, and supports short repeated contacts well. Constrains question length and format, makes grids and complex stimulus impractical, tends to produce shorter open-ended answers, and rewards designs built for small screens rather than designs ported onto them.
- **In-app or intercept.** Excellent for questions about a specific journey, captured in context. Samples only the people already present and engaged, which makes it structurally unable to answer questions about non-users or lapsers.

### Step 7. Test feasibility, then degrade the design deliberately

Check three things: incidence and reachability of the audience, elapsed time against the decision date, and whether the analysis is executable in-house. Where the ideal design fails one of them, do not silently substitute a weaker method. Degrade in a stated order, and price each degradation in conclusions rather than in money.

The usual order, least damaging first: reduce scope (fewer objectives, fully answered), then reduce precision (fewer subgroups reported, wider intervals), then reduce depth (shorter instrument, fewer sessions), then change method family. Changing method family is last because it changes what can be claimed, not just how confidently.

Every degradation is written as a cost sentence: "Dropping the regional cut removes the ability to say whether the barrier differs outside the metropolitan markets, which is one of the two questions the programme team asked." This judgement, that a compromise is acceptable, is a K5 §2.7 decision and does not belong to the AI. Mark it `RESEARCHER DECISION REQUIRED`, name what changes under each option, and stop until it is answered.

### Step 8. Write what the design cannot answer, and record the rejected options

Two lists close the recommendation. The first is the questions this design will not answer, expressed as the sentences that will not be sayable. It travels forward into the plan, the report limitations, and any executive summary. The second is the methods considered and rejected, each with the reason, so that when a stakeholder asks "why not a focus group", the answer exists in writing and does not have to be reconstructed defensively six weeks later.

### 7.9 Method repertoire (reference for Step 5)

Effort is expressed in three bands (Light, Moderate, Heavy) across three surfaces: design effort, fieldwork elapsed time, analysis effort. Absolute costs and timings are not stated because they vary by market, audience and scope. Never invent them (K4 §2.1); where a planning range is needed, obtain it and label it as supplied.

**Quantitative survey**
- *Answers:* how many, how much, how often, how groups differ, how stated attitudes distribute.
- *Cannot answer:* why, in any depth. What people would do in a situation they have not experienced. Any mechanism the respondent is not consciously aware of. Causation.
- *Effort:* design Heavy (everything must be correct before launch), fieldwork Light to Moderate, analysis Moderate.
- *Failure modes:* precision about the wrong construct; answers from people with no basis for an answer; satisficing and straight-lining; attitudes manufactured by the scale that did not exist before the question; subgroup bases too small for the cuts everyone intended.
- *Implies:* a defined frame, size driven by the smallest reportable subgroup, a pre-specified analysis plan, and testing discipline per K4 §3.1.

**Qualitative depth interview**
- *Answers:* why, how people make sense of something, sequence and context of a decision, the vocabulary people use, contradictions.
- *Cannot answer:* how many. Anything requiring population estimation. Comparisons between subgroups expressed numerically.
- *Effort:* design Light to Moderate, fieldwork Moderate, analysis Heavy and routinely underestimated.
- *Failure modes:* rationalised accounts accepted as causes; the moderator's frame imported into the answers; recruiting the articulate rather than the relevant; counting participants and presenting it as prevalence.
- *Implies:* purposive recruitment against a defined variation, a guide rather than a script, saturation judged not assumed, and thematic analysis with prevalence expressed as counts of participants (07.01, 07.03).

**Focus group**
- *Answers:* how a view forms and shifts in company, shared language and social norms, reactions to stimulus where debate is informative, range of positions surfaced quickly.
- *Cannot answer:* individual decision journeys, anything sensitive or stigmatised, prevalence, or the strength of a privately held view.
- *Effort:* design Moderate, fieldwork Light per unit of coverage, analysis Moderate.
- *Failure modes:* dominance and conformity mistaken for consensus; a group used because it is faster than eight interviews when the question is individual; over-recruitment of the socially confident; groups run on a topic where people will not be honest in front of peers.
- *Implies:* segment-homogeneous groups, a moderator managing floor time, and analysis at group level as well as participant level.

**AI-moderated interview**
- *Answers:* why and how, at a scale and consistency human moderation cannot reach; open-ended depth on larger samples; the same probing logic applied to every participant without fatigue or drift.
- *Cannot answer:* what requires reading a room, cultural nuance, or a moderator's judgement to depart from the design entirely. Weak where the interesting material is what a participant will not volunteer, and where trust has to be built before disclosure.
- *Effort:* design Heavy (probing logic must be specified in advance), fieldwork Light, analysis Moderate to Heavy given volume.
- *Failure modes:* shallow probes that accept the first answer; leading follow-ups generated on the fly; participants writing less than they would say; the appearance of qualitative depth across a sample large enough that people start quoting it as prevalence; and the temptation to treat AI-coded output as verified when it is not (K4 §7).
- *Implies:* explicit probe rules and stopping rules (03.02), a human review of a sample of transcripts, disclosure of AI involvement to participants and in the output, and analysis that keeps participant identifiers end to end (K2 §4.2).

**Diary study**
- *Answers:* how something unfolds over time, frequency and context of behaviour close to the moment, the difference between recalled and actual routine.
- *Cannot answer:* population incidence. Anything about people who will not sustain a multi-day commitment, which is a real and non-random group.
- *Effort:* design Moderate, fieldwork Heavy in elapsed time and participant management, analysis Heavy.
- *Failure modes:* attrition that quietly reshapes the sample; the act of recording changing the behaviour recorded; entries that thin out after day three; over-burdening participants until compliance becomes performance.
- *Implies:* short repeated tasks, an incentive structure matched to burden (03.05), an attrition plan, and analysis both within participant over time and across participants.

**Ethnography and immersive observation**
- *Answers:* how, in the fullest sense. The gap between what people say and what they do. Context, environment, workarounds, social meaning, and the things too ordinary for participants to mention.
- *Cannot answer:* how many, anything requiring breadth, and anything on a short timetable.
- *Effort:* design Moderate, fieldwork Heavy, analysis Heavy.
- *Failure modes:* observer effect; over-reading a small number of settings; time and cost consumed before the decision date; findings too rich to be operationalised by a team that needs a number.
- *Implies:* very small n, extended access, field notes as primary evidence with their own traceability, and consent arrangements that are more complex than for interview work (13.05).

**Structured observation**
- *Answers:* what happens, how often, in what sequence, in a defined setting, without relying on recall.
- *Cannot answer:* why it happened. Intent. Anything occurring outside the observed setting or period.
- *Effort:* design Moderate (the coding scheme is the study), fieldwork Moderate, analysis Light to Moderate.
- *Failure modes:* a coding scheme that misses the behaviour that mattered; inter-observer inconsistency; reactivity; observing an unrepresentative period such as a promotional week.
- *Implies:* a pre-specified coding frame, observer training and reliability checks, and a sampling plan across times and locations rather than across people.

**Social listening and public content analysis**
- *Answers:* what is being said unprompted, in whose words, with what volume and by what visible communities. Useful for language, emergent issues, and topic monitoring.
- *Cannot answer:* what a population thinks. Silent majorities are silent. Also weak on sarcasm, negation, code-switching and multilingual text.
- *Effort:* design Light to Moderate, collection Light, analysis Moderate to Heavy if done properly.
- *Failure modes:* platform demographics mistaken for population demographics; bot and promotional content in the base; volume spikes read as opinion shifts; sentiment scores reported without the reason behind them (07.05).
- *Implies:* an explicit query definition and its exclusions, a stated platform and period, and no population claims of any kind.

**Desk research and secondary analysis**
- *Answers:* what is already established, market structure and size, precedent, regulatory context, and often a good part of "how many".
- *Cannot answer:* anything specific to your customers, your proposition or your unreleased product. Anything requiring a question nobody has yet asked.
- *Effort:* design Light, collection Light to Moderate, appraisal Moderate.
- *Failure modes:* sources cited without verification, which is the fabrication risk most likely to occur here (K4 §2.4); figures whose method and definition differ from yours; stale data presented as current; a single secondary figure repeated across many sources and mistaken for corroboration.
- *Implies:* documented search, source-quality appraisal, dating of every figure, and full referencing per K2 §4.3.

**Experiment (randomised, field or online)**
- *Answers:* what would happen if. The only family that licenses causal language (K4 §3.2).
- *Cannot answer:* why the effect occurred, unless a mechanism was measured alongside. Anything about conditions outside the ones tested.
- *Effort:* design Heavy, fieldwork Moderate to Heavy, analysis Moderate.
- *Failure modes:* under-powering, so a real effect is missed and reported as no difference; contamination between arms; peeking and stopping when the result looks good; testing a variant so small that a null result says nothing; generalising a lab or online result to a live context.
- *Implies:* randomisation, a control condition, a pre-registered primary outcome and analysis, and a power calculation done before fieldwork rather than a sample size chosen for convenience (06.04).

**Concept testing**
- *Answers:* how a defined idea is received relative to alternatives or to a benchmark; which elements land and which confuse; comparative preference between concepts.
- *Cannot answer:* what the concept will do in market. Stated purchase intention is a weak predictor and must never be reported as forecast demand.
- *Effort:* design Moderate (stimulus preparation dominates), fieldwork Light to Moderate, analysis Light to Moderate.
- *Failure modes:* the execution being tested rather than the idea; monadic versus comparative design chosen without thought; order effects left unrotated; absolute scores read without a benchmark; concepts written at different levels of polish so the best-written one wins.
- *Implies:* controlled and equalised stimulus (03.03), rotation, a comparison point, and diagnostics alongside any preference measure.

**Usability testing**
- *Answers:* whether people can complete a task with a specific artefact, where they fail, and why they fail at the point of failure.
- *Cannot answer:* whether they want it, whether they would find it, or how common the problem is in the population.
- *Effort:* design Light to Moderate, fieldwork Light, analysis Light.
- *Failure modes:* small-n problem counts converted into percentages; testing with people who are not the intended users; task wording that gives away the answer; treating a preference expressed at the end as evidence of value.
- *Implies:* representative tasks, think-aloud or observed performance, small n reported as counts and severity, and no percentages at these bases (K4 §7).

**Conjoint analysis**
- *Answers:* how people trade off attributes against each other, the relative value of features and price levels, and what a defined configuration is worth against another.
- *Cannot answer:* anything outside the attributes and levels you specified. Attributes omitted from the design do not exist in the results, which is the method's central and least understood limitation.
- *Effort:* design Heavy, fieldwork Moderate, analysis Heavy and requiring specialist capability.
- *Failure modes:* an attribute list built without prior qualitative work; too many attributes for the respondent to process; simulator outputs presented as demand forecasts; hypothetical bias treated as absent; a design run because the word "pricing" appeared in the brief.
- *Implies:* prior qualitative or internal work to define attributes, a realistic and balanced experimental design, adequate sample for the estimation approach, and analytical capability in-house or contracted (06.01).

**MaxDiff**
- *Answers:* the relative priority of a list of items, cleanly, without the scale-use bias and uniform-high ratings that importance scales produce.
- *Cannot answer:* absolute importance, whether any item matters at all, or trade-offs against price. It ranks the list you supplied and nothing else.
- *Effort:* design Moderate, fieldwork Light to Moderate, analysis Moderate.
- *Failure modes:* an item list that omits the real driver; items written at inconsistent levels of abstraction; results read as though the top item were sufficient rather than merely first; too many items for the number of tasks shown.
- *Implies:* a carefully constructed and balanced item list, a design with adequate item coverage per respondent, and reporting as relative scores with the list stated alongside.

**Longitudinal and tracking research**
- *Answers:* is it changing, in what direction, by how much, and does the change coincide with something.
- *Cannot answer:* why it changed, without a design element built to capture that. A tracker that shows a fall and cannot explain it is the most common disappointment in commercial research.
- *Effort:* design Heavy at the outset, fieldwork Moderate and recurring, analysis Light per wave and Heavy at the point of any change.
- *Failure modes:* wave-on-wave noise read as movement; instrument or sample-source changes that break comparability and are then not disclosed; the tracker outliving the decision it was built for; panel conditioning in a true longitudinal panel; questions retained past their usefulness because the trend would be lost.
- *Implies:* a stable instrument, a stable sample specification, a defined change threshold agreed before fieldwork, and a change log covering every alteration to the series (05.05).

**Mixed-method designs**
- *Answers:* combinations that no single method reaches: prevalence with mechanism, breadth with depth, or a measure whose construct was built from the audience's own language.
- *Cannot answer:* more than the sum of its parts if the strands are not connected. Two studies reported side by side are not a mixed-method design.
- *Effort:* design Heavy, fieldwork Moderate to Heavy, analysis Heavy including integration.
- *Failure modes:* mixed method chosen as a hedge because the question was never sharpened; sequential strands where the first is fielded before it can inform the second; the qualitative strand reduced to decorative quotes for the quantitative deck; divergence between strands averaged away rather than reported (K2 §4.4).
- *Implies:* an explicit sequence and purpose for each strand (exploratory first, explanatory second, or parallel with convergence tested), a stated integration point, and analysis that reports divergence as a finding (07.06).

## 8. Analytical framework

    Question shape → Claim type → Evidence required → Method family
        → Design specification → Residual (what cannot be answered)

**Question shape** is the grammar of the question, from the Step 2 table. It is recovered from the objective, not from the topic.

**Claim type** is the sentence you need to be able to say: descriptive, comparative, explanatory, behavioural, causal, trade-off, or temporal. This is where defensibility enters, because the claim type determines the standard of evidence, not the other way round.

**Evidence required** is the concrete material that would carry that sentence: a proportion on an adequate base, twelve accounts of the same sequence, an observed completion rate, a difference between randomised arms. If this cell cannot be filled, the question is unanswerable as posed and the framework stops here rather than continuing to a method.

**Method family** is selected from 7.9 by matching the evidence required, never by matching the topic. Topics do not have methods. Questions do.

**Design specification** turns the family into a study: who, how many, over what period, with what stimulus, analysed how, in what channel, at what cost in effort.

**Residual** is the part of the original brief that this design will not answer. Every design has one. A framework run that produces an empty residual has almost certainly been run carelessly, and the residual list travels forward into the plan and the report's limitations.

Apply it fragment by fragment, not to the brief as a whole. Where two fragments resolve to the same family, combine them into one study. Where they do not, the honest output is either two studies, a sequence, or a decision about which question is dropped.

## 9. Output format

**Method Recommendation Note**, in this order.

**1. The decision.** One sentence: actor, action, date. Plus the default action if no research is run.

**2. Question shape analysis.**

| Fragment | Shape | Claim needed | Evidence that would carry it | Verdict |
|---|---|---|---|---|
| ... | ... | ... | ... | Primary research / Already answered / Re-specify / Decline |

**3. Recommended design.** Method family, and then the specification: population and recruitment basis, scale (with the reporting unit that drives it), period, channel with its stated cost, stimulus if any, and the analysis it implies. Cross-reference 01.06 for the sampling design and 01.07 for the analysis plan.

**4. Why this and not the obvious alternative.**

| Option considered | Why rejected |
|---|---|
| ... | ... |

At least two rows, and one of them must be the method the brief named or the stakeholder expects.

**5. What this design cannot answer.** A list of sentences that will not be sayable at the end. Written in plain language, not as methodological caveats.

**6. Constraints and compromises.** Each degradation, in the Step 7 order, with its cost expressed in conclusions rather than in money. Each carries a `RESEARCHER DECISION REQUIRED` marker per K5 §2.7 where the trade-off is genuinely open.

**7. Effort profile.** The three bands (design, fieldwork elapsed, analysis) with the drivers named. Absolute costs and timings only where they were supplied, labelled as supplied.

**8. Confidence in the recommendation**, per K3. The recommendation itself is a claim and carries a level. Moderate is a legitimate and common answer, and where it is moderate, name the assumption that would change it (usually incidence, reachability, or the true effect size).

**When the evidence is thin.** Do not fill the format. The following are complete, acceptable outputs of this skill, and each of them is more useful than a design produced to fill a slot:

- *No method can be selected yet.* The decision is unnamed or the objectives are still a topic. Say what is missing and route to 01.02.
- *No primary research is required.* The question is answered by existing internal or published evidence. Say where, and what would still be worth checking.
- *The question is unanswerable as posed.* State why no evidence could carry the claim, and offer the re-specification.
- *No research should be commissioned.* The decision is fixed, the timing does not work, or the value of the answer is below its cost. State which, and what the alternative use of the effort is.
- *Feasibility is unverified.* Where incidence or reachability is unknown, the design is presented as conditional and the check is named as the next step. Do not present an unverified design as though feasibility had been established.

## 10. Quality checks

Run before the recommendation is presented. These sit on top of K4 §8, which runs anyway.

1. Is the decision written as actor, action and date, and would a different research result change that action?
2. Has every objective been decomposed to a single question shape, with nothing left as a compound?
3. Were all four honest verdicts in Step 4 run explicitly, and is the result of each recorded even where the answer was "proceed"?
4. Does each recommended method match the claim type, rather than the topic or the stakeholder's expectation?
5. Is the reporting unit that drives scale stated, and is it the smallest subgroup rather than the total?
6. Is the analysis executable by the team that will have to run it, and has that been checked rather than assumed?
7. Is the channel choice justified methodologically, with its weakness stated in the same paragraph as its selection?
8. Does the note contain a "cannot answer" list, and is it written in language a non-researcher would understand?
9. Are at least two rejected options recorded with reasons, including the one the brief named?
10. Is every compromise priced in conclusions rather than in money, and routed to a human where the trade-off is open?
11. Are all cost, timing, incidence and feasibility figures either supplied or absent, with none invented?
12. Is any method named here one that no supplier, channel or platform reference could be inferred from?
13. Does the recommendation carry its own confidence level, and does it name the assumption that would change it?
14. If the recommendation is "no research", has it been stated as plainly as a positive recommendation would have been?

## 11. Common failure modes

| Failure | How to recognise it | How to prevent it |
|---|---|---|
| **Method-first design** | The method appears in the brief and in the recommendation, unchanged, with the justification written afterwards | Step 2 before Step 5, always. The shape table is built before the repertoire is opened |
| **The survey default** | A quantitative survey is proposed for a "why" question, usually with an open-ended box added to cover it | Test the claim type in Step 3. An open end on a survey is not a substitute for depth |
| **Mixed method as a hedge** | Both qual and quant proposed, with no stated sequence, integration point or reason | Require the purpose of each strand in writing. If neither strand can be dropped without losing an objective, it is a design. If it can, it was a hedge |
| **Feasibility assumed** | A design specified in full detail for an audience whose incidence and reachability were never checked | Make feasibility a conditional in the output, not an assumption inside it |
| **Powering by budget** | The sample size matches what the budget bought, and the subgroup analysis everyone wants is quietly impossible | Derive scale from the reporting unit in Step 5, then compare with budget, in that order |
| **Silent degradation** | The final design is weaker than the ideal and nowhere says so | Step 7 requires each degradation to be written as a cost sentence and routed per K5 §2.7 |
| **The unasked "is it already known"** | A study commissioned to measure something the organisation already records | Verdict 1 in Step 4, run before any method is named |
| **Trade-off methods by keyword** | Conjoint proposed because the brief mentions price, MaxDiff because it mentions priorities, with no attribute or item work behind either | Both methods require prior work to define what goes into them. No attribute list, no design |
| **AI pattern-matching the topic to a method** | The recommendation would be identical for any brief in the same sector | Check whether the recommendation would change if the decision changed. If not, the decision was never used |
| **AI inventing feasibility, cost or timing** | Confident specifics on incidence, fieldwork duration or price that nobody supplied | K4 §2.1. Use the effort bands, and mark supplied figures as supplied |
| **AI over-promising the design** | The "cannot answer" list is short, generic, or absent | Every design has a substantial residual. Write it before the recommendation is considered finished |
| **AI resolving a constraint trade-off on its own** | A compromise chosen and presented as settled | K5 §2.7. The AI names the options and the cost of each. The researcher chooses |

## 12. AI guardrails

Skill-specific only. K4 applies in full and is not repeated.

1. **Never name or imply a platform, panel, supplier, tool or vendor**, and never let a channel comparison drift toward one. Where a channel is recommended, its methodological weakness is stated in the same passage.
2. **Never state incidence, feasibility, cost, timing or achievable sample as fact unless supplied.** Use the effort bands. A supplied figure is labelled as supplied and attributed.
3. **Never recommend a method whose analysis has not been confirmed as executable** by the team that will run it. An unrunnable design is not a recommendation, it is a deferral of the problem.
4. **Never let a design imply a claim it cannot support.** In particular, do not recommend a cross-sectional design for a question whose answer will be read causally without stating, in the recommendation itself, that causal language will not be licensed (K4 §3.2).
5. **Never present a mixed-method design as inherently superior.** It is superior only where each strand answers a fragment the other cannot, and where the integration point is specified.
6. **Never omit the "no research needed", "already answered" or "unanswerable as posed" verdicts because a study is expected.** The expectation of a commission is not evidence about the question.
7. **Never resolve a constraint trade-off unilaterally.** Specify the options and what each costs the conclusions, then stop for a human decision per K5 §2.7.
8. **Never describe any sample as representative** without naming the frame it is representative of and the coverage gap it carries (K3 §6).
9. **Never carry a method recommendation forward once the decision it served has changed.** Re-run from Step 1. A design is only correct relative to a decision.
10. **Never let stakeholder preference enter as a methodological argument.** Record it as a constraint, visibly, so that a reader can see it was a preference and not a finding about the question.

## 13. Best-practice principles

1. **Design backwards from the sentence you need to be able to say.** Write the headline finding as though it already existed, with its base and its caveat, then ask what would have had to happen to produce it honestly. Designs built this way rarely include anything that is not needed.
2. **The method follows the claim, never the topic.** There is no such thing as a brand method or a pricing method. There are questions about brands that need measurement and questions about brands that need depth.
3. **Scale is set by the smallest thing you must report on.** Nobody has ever been disappointed by a total-sample base. They are routinely disappointed by a subgroup of 40.
4. **Every method's weakness is the mirror of its strength.** Depth interviews find what a survey cannot precisely because they do not sample enough people to count. Reading a limitation as a defect rather than a trade-off leads to designs that try to be everything.
5. **Prefer measured behaviour to reported behaviour whenever both are available**, and treat stated future behaviour as the weakest evidence in the repertoire. It is not worthless, but it is directional and it is never a forecast.
6. **Ask what would have to be true for the cheap option to be sufficient.** Often the answer is "nothing much", and the expensive design was habit. Occasionally the answer is precise and important, and it becomes the justification for the expensive design.
7. **Sequence beats combination.** Two strands in the right order, where the first shapes the second, deliver more than the same two run in parallel and reported together.
8. **Cost the analysis, not just the fieldwork.** The analytical burden of qualitative work, of open-ended volume, and of trade-off methods is where projects overrun. A design the team cannot analyse on time is not cheaper, it is late.
9. **Do not buy precision you will not use.** A tighter interval on a number that will be reported as "about a third" is money spent on nothing.
10. **The decision deadline is a design input, not a scheduling detail.** Elapsed time is a methodological constraint of the same standing as budget, and it eliminates more designs than money does.
11. **Where the finding will be contested, design for the objection.** Anticipate the specific challenge the sceptic will make, and put the element that answers it into the design rather than into the defence afterwards.
12. **A design nobody can explain in three sentences will not be trusted, and often should not be.** Complexity that cannot be justified in plain language is usually complexity added for reassurance.

## 14. Worked example

**Sector: public utility (fictional).**

    INPUT

A brief from Northvale Water, a fictional regional utility: "We need a survey of
1,000 customers to find out why they are not using our online billing portal and
what features would make them switch to it."

    PROCESS

**Step 1, recover the decision.** "The digital programme lead decides in November
whether next year's budget funds portal features or funds a migration campaign
onto the existing portal." Two decisions were hiding inside one brief, and they
need different evidence. Default if no research runs: the budget goes to
features, because that is what the roadmap already assumes.

**Step 2, decompose.** Four fragments. (a) How many customers are not using the
portal, and who are they: *how many*. (b) Why are they not using it: *why*.
(c) What stops them at the point of trying: *how*. (d) What features would make
them switch: *what would happen if*, asked as stated intention.

**Step 3, claim types.** For (d), the sentence required is about future
behaviour toward features that do not exist, from people who have not
encountered them. No evidence carries that claim. It is unanswerable as posed.

**Step 4, honest verdicts.** Fragment (a) is already answered: the utility's own
account system records registration and login activity against every account,
with tenure, tariff and region attached. A survey would sell the client back its
own data, less accurately, because it would rely on recall. Removed, with an
internal analysis specified instead. Fragment (d) is re-specified into what each
customer has already done when they hit the barrier, and which of a defined set
of changes they trade off against each other. The trade-off version is deferred,
because a MaxDiff item list cannot be built before (b) and (c) have been run.

**Step 5, select and specify.** Fragment (c) resolves to usability testing on the
live registration journey with recent non-adopters, reported as failure points
and severity counts. Fragment (b) resolves to depth interviews with a purposive
spread across tenure, tariff type and region, drawn from the non-adopter
population the account data has now defined.

**The judgement call.** The budget supports one strand. Usability testing shows
where people fail but not whether failure is why they stopped, and a substantial
share of non-adopters have never attempted registration at all, so a
usability-only design would study the wrong population. Depth interviews are
selected, with task observation appended to the second half of each session
where participants are willing. This buys most of fragment (c) at the cost of a
smaller observed base and a less controlled task.

**Step 7, constraint.** Face-to-face sessions are not feasible across the rural
service area within the timetable, and remote sessions require the connectivity
that the least digitally engaged customers are least likely to have. This is the
compromise that bites, because it excludes the segment most likely to explain
non-adoption.

> **Researcher decision required.** Remote-only fieldwork excludes low-
> connectivity customers, who are plausibly over-represented among non-adopters.
> The options are: accept the exclusion and state it as a coverage limitation on
> every finding; add a small telephone-only strand, which loses the observation
> element for those participants; or extend the timetable past the November
> decision. This is a trade-off between coverage and the decision date and it
> belongs to the researcher (K5 §2.7). What turns on it: whether the study can
> speak about the offline segment at all.

    OUTPUT

Depth interviews with portal non-adopters, purposively spread, with appended
task observation where possible, delivered ahead of the November decision.
Fragment (a) answered from internal account data. A MaxDiff on feature
trade-offs specified as a conditional second phase, dependent on the interviews
producing a defensible item list.

**Cannot answer:** how common each barrier is across the customer base; whether
barriers differ by region; what proportion would adopt if a given change were
made. Confidence in the recommendation: moderate, resting on the unverified
assumption that non-adopters can be recruited from the account data at
reasonable incidence.

## 15. Advanced usage

**Stage the design against the decision, not the budget.** Where the decision has stages, so should the research. A cheap first strand that removes options is often worth more than a comprehensive study that arrives once the options have narrowed anyway. Ask what the first strand would have to show to make the second unnecessary, and design it to be capable of showing that.

**Reason about the value of the information.** Before specifying anything, ask what the organisation would do under each possible result. Where two or more results lead to the same action, that part of the design is not earning its place. This test removes more scope than any budget conversation.

**Match rigour to reversibility.** A decision that can be reversed cheaply in three months justifies a fast, weaker design and a monitoring plan. An irreversible commitment justifies the design that will survive challenge. Applying the same standard to both wastes money on one and risks the other.

**Run a pre-mortem on the finding.** Assume the study has been delivered and dismissed. Work out which objection did it: base too small, wrong audience, stated intention, a channel that missed the relevant people, or a finding the stakeholder already believed. Then put the element that answers that objection into the design.

**Design for the null result.** Ask what the output looks like if nothing differs and nothing is found. If that output is unusable, the design has an unstated dependency on a positive result, which is a design fault and frequently the origin of post-hoc storytelling.

**Trackers and continuity.** On an established series, method improvement and comparability are in direct conflict. The usual resolution is a parallel-run bridge wave, which is expensive and often the only honest way to change an instrument without discarding the history. Where a bridge is unaffordable, changing the method means accepting that the series restarts, and that decision belongs to a human.

**Deliberately choosing the weaker method.** Sometimes the right recommendation is the faster, weaker design, because it lands before the decision. This is legitimate, and it is only legitimate when the weakness is stated in the same breath and carried into the report's limitations rather than forgotten between the design note and the deck.

## 16. Skill chain

**Recommended previous skills**
- **01.02 Business Problem to Research Question.** Hands over the decision and the objectives in answerable form. Without it, this skill has nothing to select against.
- **01.03 Hypothesis Development.** Hands over testable propositions, which frequently convert an exploratory design into a cheaper confirmatory one.

**Recommended next skills**
- **01.05 Research Plan Development.** Takes the selected method and turns it into a plan with phases, dependencies, timings and deliverables.
- **01.06 Sampling Strategy.** Takes the sample structure this skill implies and turns it into a frame, a size and a quota design.
- **Category 02 Instrument Design.** Takes the method and builds the questionnaire, guide or protocol it requires.

**Runs well alongside**
- **10.01 Literature Review and Desk Research**, for Step 4 verdict 1.
- **01.07 Analysis Plan Development**, as the test of whether the chosen method's analysis is executable.
- **13.05 Research Ethics and Consent Design**, wherever the design touches sensitive topics or vulnerable participants.
- **05.05 Trend and Tracker Analysis**, wherever continuity constrains the method choice.

---
A Yazi Supplied Skill and resource.
