---
name: analysis-plan-development
description: >
  Specifies the analysis before the data exists. Use for "write the analysis plan",
  "what are we going to test and how", "which subgroups are we cutting and why",
  "how do we avoid p-hacking", "what counts as a meaningful difference", "how do we
  handle don't knows and missing data", "should we correct for multiple
  comparisons". Pre-specifies primary and secondary comparisons, tests, thresholds,
  derived variables and nets, missing data and outlier rules, the
  multiple-comparison approach, and keeps exploratory analysis in a separate,
  labelled register from confirmatory analysis.
category: 01 Research Strategy and Design
ref: "01.07"
tier: 1
inherits: [K2, K3, K4, K5]
---

# Analysis Plan Development

## 1. One-line description
Writes down, before any data exists, exactly what will be analysed and how: the primary and secondary comparisons, the tests and thresholds, what counts as a meaningful difference, the derived variables, the missing data and outlier rules, the multiple-comparison approach, and the boundary between confirmatory and exploratory analysis.

## 2. What this skill is used for

**The research problem it solves.** Analytical flexibility is invisible and enormous. Between arrival of a dataset and a finished chart there are dozens of defensible choices: which subgroups to cut, whether "don't know" stays in the base, how a net is composed, where a scale is dichotomised, which cases to exclude, which of forty possible comparisons to report. Made after seeing the data, every one of them is made in the presence of knowledge about which way it moves the answer, and none of that appears in the output. The result is not usually deliberate manipulation. It is an honest researcher following a story that the data appeared to offer, arriving at a finding that would not survive replication, and reporting it with the confidence appropriate to a prediction. A study with twenty untested subgroup comparisons and a 95% threshold will produce a significant result by chance as a matter of routine, and the one that appears is the one that gets the slide. This skill removes the flexibility by spending it in advance, in writing, dated.

The corollary matters as much as the rule: **analysis discovered after seeing the data is not invalid**. Exploration is where a great deal of the value in any dataset lives. The failure is reporting an exploratory result as though it had been predicted. This skill makes the distinction structural rather than a matter of memory.

**Where it sits in the research lifecycle.** After the question set (01.02), the hypotheses where they exist (01.03), the method (01.04) and the sample design (01.06), and before fieldwork. Writing it earlier than that is guesswork; writing it later than that defeats its purpose.

**Typical use cases.**
- A quantitative study is about to field and nobody has written down what will be tested.
- A study will be contested, and pre-specification is the defence.
- Multiple subgroups are planned and the multiplicity problem has not been faced.
- An experiment or test needs a primary outcome and analysis before it launches.
- A tracker needs a fixed analysis convention so waves remain comparable.
- A team has been asked to explain the difference between a confirmatory and an exploratory finding.
- The stakeholder wants a "significant difference" and needs the meaningfulness conversation before rather than after.

**Who uses it.** Quantitative research managers and directors, evaluation and policy analysts, product and growth researchers running tests, client-side insight leads who will have to defend a finding, and postgraduate researchers writing an analysis chapter. It is also the skill a qualitative lead should read for its confirmatory/exploratory discipline, which applies with modification to structured qualitative work.

## 3. When to use it

- A quantitative or mixed-method study is designed and fieldwork has not started.
- Hypotheses exist and someone must specify how each will be tested.
- A study has subgroup objectives, so multiplicity and base adequacy are in play.
- An experiment, pilot or A/B test is being set up and needs a primary outcome.
- A recurring study needs a stable analysis convention so its trend means something.
- A dataset has arrived and the team is about to start cutting it. Writing the plan now is late but far better than not writing it.
- A previous study produced a finding nobody can replicate and you want to know why.
- The design uses a scale, an index or a derived measure whose construction is not yet defined.

## 4. When NOT to use it

- **Before the sample design exists.** The plan will specify comparisons that the achieved bases cannot support, which is the commonest way an analysis plan becomes decorative. Use **01.06 Sampling Strategy** first, and if the plan's comparisons require bases the sample cannot deliver, the sample design is re-opened rather than the plan quietly softened.
- **After the data has been seen, presented as though it were written before.** Writing an analysis plan late is legitimate and useful, and dating it honestly is compulsory. A plan reconstructed after sight of the data and dated to the design phase is a misrepresentation of the evidence, not a documentation exercise (K2 §7).
- **For genuinely exploratory work.** Discovery analysis, early-stage data exploration, and open-ended qualitative work do not have a confirmatory plan to write. What they need instead is an exploratory protocol: the questions being explored, the analyses run, the count of them, and the labelling convention that keeps their output out of the confirmatory register. Imposing a confirmatory plan here does not add rigour; it manufactures the appearance of prediction.
- **To run the analysis.** This skill specifies; it does not execute. Descriptive output belongs to **05.01 Descriptive Analysis**, testing to **05.02 Statistical Testing**, trend interpretation to **05.05 Trend and Tracker Analysis**, and qualitative coding to **07.01 Thematic Analysis**.
- **To decide what a finding means.** Interpretation, insight and implication are downstream and belong to **08.01 Finding to Insight Development**. An analysis plan states what will be produced, not what it will show, and a plan that predicts its own conclusions has confused the two.
- **To develop hypotheses.** The propositions, their prior basis and their disconfirming evidence belong to **01.03 Hypothesis Development**. This skill takes the register as an input and specifies the tests.
- **When the study is purely descriptive with a single audience and no comparisons.** A frequency run with no subgroup cuts does not need a multiplicity strategy. It still needs the derived variable, missing data and base conventions, so use those steps and skip the rest rather than producing apparatus nobody will read.
- **When the analysis being requested cannot be performed by anyone available.** A plan specifying a technique the team cannot execute is a deferral of the problem. Establish capability now (this is a Step 1 check) or specify an analysis that can actually be run.

## 5. Required inputs

**Required.**
- **The objectives and research questions**, in the form 01.02 produces them. Every analysis in the plan must name the objective it serves, and objectives with no analysis are the plan's most valuable finding.
- **The sample design and the achievable bases per reporting cell** (01.06). Comparisons specified on bases that will not exist are a plan for disappointment.
- **The instrument, or the measures as they will be asked.** Response options, scale points, the presence or absence of a "don't know", and routing all determine what the analysis can do, and each is a decision that must be made before the plan can be written.
- **The decision the study serves**, because the meaningfulness threshold is defined by it and by nothing else.

**Optional, and what each one adds.**
- **The hypothesis register (01.03).** Converts a list of planned comparisons into a set of confirmatory tests with pre-specified disconfirming patterns, which is a much stronger position.
- **Prior wave or comparable study data.** Supplies expected proportions and variances, which make the base adequacy check real rather than nominal, and reveals which comparisons are likely to be inconclusive before the study is run.
- **The smallest difference that would change the decision**, agreed with the decision owner. The single most valuable input, and a business judgement (K5 §2.1).
- **The reporting template or deck structure, where one is fixed.** Reveals the cuts that will be demanded regardless of whether they are in the plan, which is better known now than at the debrief.
- **The data processing conventions in use** (how open ends will be coded, how weights will be applied, how derived variables have historically been built). Prevent an analysis plan that conflicts with what the processing will actually deliver.
- **Any regulatory, audit or publication requirement.** Determines the level of pre-specification required and whether the plan must be lodged externally.

## 6. Questions to ask before starting

1. **What is the smallest difference that would change what you do?** Why it matters: it is the meaningfulness threshold, and without it the study can only report significance, which is a property of the sample size. Default if unanswered: do not invent one. Specify the analysis as conditional, name this as the missing input, and report effect sizes so the conversation can happen with numbers in front of it.
2. **Which comparisons is this study actually for?** Why it matters: designating one to three primaries is what makes the rest secondary and keeps the multiplicity burden manageable. Default: derive the primaries from the decision and mark them provisional pending confirmation.
3. **Which subgroups will be analysed, and why each one?** Why it matters: subgroups analysed because the variable exists are the main engine of false positives. Default: restrict to subgroups named in the objectives, and record any others as exploratory in advance.
4. **Does "don't know" belong in the base?** Why it matters: it is a single convention that can move a headline several points, and deciding it after seeing which way it moves is the cleanest example of a post-hoc choice. Default: exclude from the percentage base but report it separately and show both where it exceeds a material share.
5. **What will happen if a planned comparison's base falls short?** Why it matters: the answer decided now is better than the substitution improvised later. Default: report the comparison descriptively with the base shown and no test, and state that in the plan.
6. **How many tests will be run in total?** Why it matters: it determines whether a multiplicity correction is needed and which. Default: count them, and if the count is uncomfortable, reduce the tests rather than adjusting the threshold after the fact.
7. **Who runs this analysis, and can they?** Why it matters: an unexecutable plan is a deferral. Default: specify only techniques the named analyst can run, and flag any that require capability not currently available.

## 7. Step-by-step methodology

**Step 1. Convert each objective into an analysis question with a named output.**
For every objective, write what will be produced: the measure, the base, the form (a proportion, a mean, a distribution, a difference, a ranking, a model coefficient), and the artefact it appears in. If an objective cannot be converted, it was never properly formed, and that discovery belongs here rather than at analysis; send it back to 01.02. A correct result is a table in which every objective has at least one named output and every planned output names an objective. Orphan outputs (the demographic profile nobody asked for, the standard battery that arrived with the template) are removed now, and each one removed is a test not run.

**Step 2. Designate the primary comparisons.**
One to three. A primary comparison is one the study exists to make, on which the decision turns, and which will be reported whatever it shows. Each is specified completely: the measure, the two or more groups or waves being compared, the direction of interest, the test that will be used and why it suits the measure and the design, the threshold, and the base required in each arm. The discipline of designating primaries is what makes everything else secondary, and the number matters: a study with nine primaries has none. Where a hypothesis register exists (01.03), the primary comparisons are the tests of the primary hypotheses and must match them exactly, including the disconfirming patterns.

**Step 3. Specify secondary comparisons, distinguished by consequence.**
Secondary comparisons contextualise, support and explain. They do not by themselves change the decision, and that is the definition, not their statistical status. Specify them with the same completeness as the primaries, then state their multiplicity treatment, which is usually different: secondaries carry a higher risk of chance findings and a correspondingly weaker interpretive standing. Anything that cannot be specified at this stage is not a secondary comparison; it is exploratory, and it goes to Step 11.

**Step 4. Define derived variables, nets, indices and scale handling.**
This is where a large share of unnoticed flexibility lives, and every decision is made now. For **nets**: exactly which codes combine, and why those. For **top-box and top-two-box conventions**: which points, and an explicit statement of the cost, since dichotomising discards distributional information and can make a shift in the middle of a scale invisible. For **scales**: whether treated as ordinal or interval, with the justification, since means on ordinal scales are common practice and are a defensible convention rather than a neutral one, and the choice must be consistent across the study. For **indices and composites**: the items, the weighting, the treatment of missing items, the reliability check that will be run, and, critically, the rule for what happens if that check fails, decided now so that a failing index is not quietly retained because the deadline is near. For **recodes and banding**: exact cut points, and the reason for each, because age bands and value bands chosen after seeing the distribution are among the most effective ways to manufacture a difference. A correct result lets a second analyst construct every derived variable identically from the raw data without asking a question.

**Step 5. Pre-specify the subgroup analysis.**
List every subgroup that will be analysed, and against each write the reason a decision turns on it. Availability is not a reason. Then attach the minimum base, per K3 §3.2 and the sample design, and the pre-decided action if the base falls short: report descriptively with the base shown, combine with an adjacent category on a rule stated now, or drop the cut. Then count: the number of subgroups multiplied by the number of measures analysed within them is the size of the multiplicity problem, and seeing that number written down is usually what reduces the list. Any subgroup analysis not in this list is exploratory when it happens, whatever it finds.

**Step 6. Decide missing data handling in advance.**
Four decisions. **Item non-response**: what happens when an item is unanswered, and whether the base for that item differs from the study base (it usually does, and unstated shifting bases are a frequent source of apparent contradiction between charts). **"Don't know" and "prefer not to say"**: in or out of the percentage base, decided now and applied consistently, with both versions reported where the share is material. **Case exclusion**: the completeness threshold below which a case is dropped, expressed as a rule. **Imputation**: whether any is used, of what kind, on which variables, and how imputed values will be flagged in the output. Then state how missingness will be reported, because the pattern of what people decline to answer is frequently informative in itself and is routinely discarded. Every one of these decisions is capable of moving a headline, which is why all four are made blind to outcome.

**Step 7. Decide outlier and data quality rules in advance, and apply them blind.**
Specify the rules and the thresholds: completion speed below which a case is reviewed, straight-lining detection on grid items and what it triggers, out-of-range and internally inconsistent values, duplicate detection, and open-end quality criteria where relevant. The essential discipline is that each rule is written as a rule and then applied to the whole dataset without reference to what the exclusion does to the results. Where an exclusion materially changes a headline, both figures are reported and the decision is documented per K4 §4.4. A quality rule invented after seeing the results is an exclusion criterion selected for its effect, whatever the intent behind it, and this is a K5 §2.7 decision point that returns to a named researcher.

**Step 8. Choose and state the multiplicity approach.**
Begin by counting the planned tests. Then define the families: comparisons are grouped into families within which the error rate is controlled, and the definition of the family is itself an analytical choice that must be written down. Then choose the approach with its cost stated. **No correction** is defensible for a small number of pre-specified primary comparisons, and its cost is that with twenty independent tests at the conventional threshold, roughly one false positive is the expectation rather than the exception. **Family-wise control** protects against any false positive in the family, at the cost of substantially reduced power, and it is appropriate where a single false claim would be expensive. **False discovery rate control** suits large batteries of exploratory comparisons where some false positives are tolerable and the aim is to keep their proportion bounded. And the option that is usually the best available and rarely considered: **run fewer tests**. Whichever is chosen, state it, state what it costs, and never select it after seeing which tests survive it.

**Step 9. Define what counts as a meaningful difference, separately from a significant one.**
This is the step the field skips, and it is the one that determines whether the study produces decisions or trivia. For each primary comparison, state the smallest difference that would change what the organisation does, agreed with the decision owner before fielding. Significance and meaningfulness then form four cases, and the plan states in advance how each will be reported.

| | Meaningful in size | Not meaningful in size |
|---|---|---|
| **Statistically significant** | Act. The claim the study exists to make | Report, with the size shown, and state explicitly that the difference is reliable and too small to act on. Large samples make trivial differences significant |
| **Not statistically significant** | Inconclusive, which is not the same as no difference. Report the estimate, the interval and the base required to have detected it | Evidence of no material difference at this base, reported as such |

The bottom-left cell is the one most often mis-reported as "no difference", and the top-right is the one most often over-claimed. Writing all four in advance is what prevents both.

**Step 10. Write the analysis sequence and the stopping rules.**
State the order in which analyses will be run, and the rule that results do not change the plan. Two specific stopping rules matter. Fieldwork does not stop early because the result looks good unless the design provided for interim analysis with an adjusted threshold; unplanned peeking with an option to stop inflates false positives substantially. And the analysis does not continue past the plan in search of a result: when the specified analyses are complete, what remains is exploratory and is labelled as such.

**Step 11. Open the exploratory register.**
A separate, dated section that will be filled after the data arrives. Its rules are stated now. Exploratory analysis is legitimate, valuable, and frequently where the interesting material is; it is not second-class work. What it must not do is arrive in a report wearing confirmatory clothing. Three requirements: every exploratory analysis is recorded, including those that found nothing, so the denominator is visible; every exploratory finding is reported with the language of a hypothesis per K3 §4.3, with what would validate it named; and exploratory findings generate the next study's hypotheses rather than this study's conclusions. Where an exploratory finding is strong enough to be acted on, the honest route is a confirmatory test in the next wave, and saying so is a stronger professional position than presenting it as established.

**Step 12. Check the plan against the sample, then lock and date it.**
Walk every specified comparison against the achievable bases from 01.06. Comparisons that cannot be supported are either re-specified now, resourced by changing the sample design, or explicitly demoted to descriptive reporting with no test. Then lock the plan with a version and a date, circulate it beyond the analysis team, and open the deviation log. Every departure from the plan is recorded with four fields: what changed, why, the date, and whether the change was made before or after the data was seen. That last field is the entire integrity mechanism, and it costs one word.

## 8. Analytical framework

    Objective
      → Analysis question      (what will be produced)
        → Measure and base     (from the instrument and the sample design)
          → Comparison         (primary or secondary, or none)
            → Test and threshold
              → Meaningfulness threshold   (from the decision, not the statistics)
                → Decision consequence     (what happens under each outcome)

with a hard vertical line through the whole document:

    CONFIRMATORY                    |  EXPLORATORY
    Specified and dated before data |  Discovered after data
    Reported as a test              |  Reported as a hypothesis (K3 §4.3)
    Multiplicity budget applied     |  Count of analyses disclosed
    Changes the decision            |  Generates the next study's hypotheses

Build the chain downward for each objective. Then apply two checks. **Downward feasibility**: can the base support the comparison, and can the analyst run the test? **Upward necessity**: does every specified analysis serve an objective, or does it exist because the data permits it? The second check is what keeps the test count low, and a low test count is worth more than any correction procedure.

The line between the two columns is not a matter of the analyst's memory. It is a matter of the date on the document.

## 9. Output format

An **Analysis Plan**, versioned and dated, in this order.

**1. Objectives to outputs.**

| Objective | Analysis question | Output (measure, base, form) | Where it appears |

**2. Primary comparisons.**

| # | Measure | Groups compared | Direction of interest | Test and why | Threshold | Base required per arm | Meaningful difference | Action if confirmed / not confirmed / inconclusive |

**3. Secondary comparisons.** Same columns, plus the multiplicity treatment applied to the family.

**4. Derived variables and conventions.**

| Derived variable | Construction (exact codes, items, cut points) | Justification | Reliability check and the rule if it fails |

Plus the study-wide conventions: scale treatment, top-box convention and its cost, rounding, and base labelling.

**5. Subgroup register.**

| Subgroup | Why a decision turns on it | Minimum base | Action if base falls short |

With the count of subgroup-by-measure comparisons stated.

**6. Missing data rules.** Item non-response, "don't know" treatment with the base convention, case exclusion threshold, imputation (or none), and how missingness will be reported.

**7. Data quality and outlier rules.** Each rule, its threshold, what it triggers, and the requirement that both figures are reported where an exclusion moves a headline materially.

**8. Multiplicity.** Number of planned tests, family definitions, the approach chosen, and its stated cost.

**9. Analysis sequence and stopping rules.**

**10. Exploratory register.** Empty at lock, with its rules stated: recording requirement, reporting language, and the prohibition on confirmatory presentation.

**11. Base and claim check.** Every comparison against the achievable base from 01.06, with the resolution where it fails.

**12. Version, lock date, and deviation log.**

| Deviation | Reason | Date | Before or after data was seen |

**When the evidence is thin.** Where the meaningfulness threshold has not been supplied, the plan says so and specifies that effect sizes and intervals will be reported so the judgement can be made with numbers present; it does not invent a threshold. Where a base will not support a planned test, the plan says the comparison will be reported descriptively without a test rather than specifying a test that will be run anyway. Where the plan is being written after fieldwork, the lock date says so and every analysis in it is exploratory. Format is not evidence (K4 §1).

## 10. Quality checks

Run before the plan is locked. These sit on top of K4 §8.

1. Does every objective have at least one named output, and every output an objective?
2. Are there three or fewer primary comparisons, each fully specified?
3. Does every primary comparison match the hypothesis register where one exists, including its disconfirming pattern?
4. Could a second analyst construct every derived variable, net and index identically from the raw data without asking a question?
5. Does every subgroup carry a reason a decision turns on it, rather than the existence of the variable?
6. Has the total number of planned tests been counted and written down?
7. Is the multiplicity approach stated with its cost, and was it chosen before any results existed?
8. Is the "don't know" base convention decided, stated, and applied consistently across the document?
9. Are the case exclusion and outlier rules written as rules that can be applied without reference to their effect on results?
10. Does every primary comparison have a meaningfulness threshold, or an explicit statement that one was not supplied and what will be reported instead?
11. Are all four significance-by-meaningfulness cells addressed, including the inconclusive one?
12. Has every specified comparison been checked against the achievable base, with a resolution where it fails?
13. Is the exploratory register present, empty, and governed by stated rules?
14. Is the plan dated, and does the date honestly reflect when it was written relative to the data?
15. Does the deviation log include the field recording whether a change preceded or followed sight of the data?

## 11. Common failure modes

| Failure | How to recognise it | How to prevent it |
|---|---|---|
| **The subgroup fishing expedition** | Every demographic crossed with every measure, and the significant ones reported | Step 5: subgroups need a prior reason and the comparison count is written down |
| **Post-hoc base convention** | "Don't knows" excluded in one chart and included in another, in the direction that helps | Step 6: one convention, decided in advance, applied everywhere |
| **The invented cut point** | Age bands or value bands that do not match any convention and split exactly where the difference is | Step 4: cut points and their reasons are fixed before data |
| **Significance as meaning** | A two-point difference on a large sample reported as a finding | Step 9's four cells, with the meaningfulness threshold set by the decision owner |
| **Null read as no difference** | A non-significant result reported as "no difference between groups" | The bottom-left cell: report the estimate, the interval, and the base that would have been needed |
| **The failing index retained** | A composite whose reliability check was run and quietly ignored | Step 4: the rule for a failed check is written before the check is run |
| **Multiplicity discovered late** | Correction discussed after seeing which comparisons survived | Step 8, before fielding, with the test count stated |
| **Exclusion by effect** | A quality rule introduced once its effect on the headline is known | Step 7: rules applied blind, both figures reported, decision documented (K4 §4.4) |
| **The plan that outran the sample** | Elegant comparisons on cells of 30 | Step 12's base check against 01.06 |
| **Retrospective plan** | An analysis plan dated to the design phase, written after the topline | The lock date is honest or the plan is worthless. Late plans are labelled exploratory |
| **AI: plausible thresholds** | A meaningful-difference threshold produced without the decision owner | It is a business judgement (K5 §2.1). Ask, or report effect sizes and say the threshold is unset |
| **AI: test selection by name** | A test chosen because it matches the words in the objective rather than the measure and design | Every test carries a stated reason referring to the measure's level and the design |
| **AI: exhaustive specification** | Forty comparisons specified because they are all possible | Upward necessity: every analysis names an objective, and fewer tests beats any correction |
| **AI: exploratory laundering** | An after-the-fact finding written in confirmatory language because it is compelling | The register and the date. Exploratory findings use hypothesis language per K3 §4.3 |

## 12. AI guardrails

Skill-specific. K4 applies in full and is not repeated here.

1. **Never date an analysis plan to a time before it was written.** If the plan is produced after the data exists, it says so, and every analysis it contains is exploratory. This is the skill's single most important rule, because everything else it protects depends on the date being true.
2. **Never present an analysis discovered in the data as confirmatory.** Say plainly that it is exploratory, state how many analyses were run, and use hypothesis language (K3 §4.3). Exploratory findings are valuable; mislabelled ones are not.
3. **Never invent a meaningfulness threshold.** The smallest difference that would change an organisation's action is a business fact the AI does not hold (K5 §2.1). Where it is missing, report effect sizes and intervals and name the gap.
4. **Never choose or change a multiplicity approach after seeing which comparisons survive it**, and never leave the approach unstated because no correction was applied. "No correction, and here is what that costs" is a legitimate stated position; silence is not.
5. **Never specify a test without naming why it suits the measure's level of measurement, the design and the base.** A test named without a reason invites a test run without a reason (K4 §6.3).
6. **Never let a data quality or exclusion rule be written, revised or applied with knowledge of its effect on the results.** Where an exclusion moves a headline materially, both figures are reported and the decision returns to a researcher (K5 §2.7).
7. **Never specify a comparison the sample cannot support.** Check against 01.06 and either re-specify, re-open the sample design, or demote to descriptive reporting with the base shown.
8. **Never allow the base convention to vary within a study.** A shifting base is invisible in the output and produces apparent contradictions no reader can diagnose.
9. **Never report a failure to reject as evidence of no difference**, and never omit the base that would have been required to detect the difference that matters.
10. **Never quietly drop a pre-specified analysis because its result was unhelpful.** Every specified analysis is reported, and the deviation log records anything that was not, with the reason and the date.

**Human review points** (see K5 §2 for the classes):
- **Researcher decision required** on the meaningfulness threshold for each primary comparison. Class 2.1, business relevance and materiality.
- **Researcher decision required** where a data quality rule's application materially changes a headline. Class 2.7, and it must be made and documented rather than absorbed.
- **Researcher decision required** where a planned comparison's base falls short and the choice is between combining categories, reporting descriptively and dropping the cut. Class 2.7.
- **Researcher sign-off required** on the locked plan before fieldwork, because the lock date is the study's integrity record and must be owned by a named person. Class 2.8.

## 13. Best-practice principles

1. **Spend the analytical flexibility in advance.** Every choice made before the data exists is a choice made without knowing which way it moves the answer. That is the whole of the method.
2. **Three primaries, at most.** A study with nine primary comparisons has none, and the discipline of choosing is what makes the multiplicity problem tractable.
3. **Fewer tests beats any correction.** Every correction procedure buys protection with power. Removing comparisons that no decision turns on buys the same protection for free.
4. **Significance is a property of your sample size; meaningfulness is a property of the business.** Only one of those is yours to decide, and a study that reports the first without the second has outsourced its judgement to n.
5. **The bottom-left cell is where studies mislead.** A meaningful-sized difference that missed significance is inconclusive, not absent, and reporting it as "no difference" is the commonest quantitative error in applied research.
6. **Decide the "don't know" convention first.** It is one line, it can move a headline several points, and deciding it late is the purest available example of a choice contaminated by knowing the answer.
7. **Write the rule for a failed reliability check before running it.** Otherwise the check is a formality that only ever confirms.
8. **Exploration is not the sin. Mislabelling is.** Say what was found by looking, say how much looking was done, and hand it to the next study as a hypothesis.
9. **Report the denominator of your exploration.** "One of forty subgroup comparisons was significant" is a completely different statement from "this subgroup differs", and only the first is honest about the evidence.
10. **Circulate the plan beyond the analysis team.** Its protective value comes from stakeholders having seen it before the result, which is exactly when it is cheapest for them to agree to it.
11. **Check the plan against the sample before locking.** An analysis plan that outruns its bases is a list of disappointments with a schedule attached.
12. **One field carries the integrity of the whole document: before or after the data was seen.** Everything else is administration.

## 14. Worked example

Fictional scenario, higher education. All figures are illustrative and belong to the scenario.

    INPUT

    A university plans a survey of first-year students to inform a decision, in
    June, about whether to extend a peer-mentoring scheme to all faculties.
    The scheme currently runs in three of nine faculties. Objectives: whether
    mentored students report stronger academic confidence and belonging than
    non-mentored students; whether any difference varies by entry route
    (direct, transfer, mature); and what mentored students say the scheme
    changed. Planned achieved sample from 01.06: 1,200 total, with roughly
    380 mentored, and entry-route cells of about 700, 300 and 200.

**Process.**

*Step 2, primaries.* Two, not three. The mean on the five-item belonging scale, mentored versus non-mentored; and the mean on the four-item academic confidence scale, same comparison. The direction of interest is stated as higher among mentored students, and both will be reported whatever they show. The test is specified with its reason: comparison of means on multi-item scales treated as interval, with unequal variances assumed, at the conventional threshold, on bases of roughly 380 and 820.

*Step 4, derived variables and the first judgement call.* Both scales are composites, and the plan fixes the items, the treatment of a missing item (mean of the answered items where at least four of five are present, otherwise the case is missing for that scale), the reliability check, and the rule if it fails: if internal consistency falls below the stated threshold, the composite is abandoned and the two most face-valid items are reported separately, with the failure disclosed. Writing that rule in advance was contested by a stakeholder who wanted the composite retained regardless, which is precisely why it was written down.

*Step 5, subgroups and the multiplicity problem.* The entry-route objective implies three subgroups crossed with two scales, giving six comparisons on top of the two primaries. But the mature-entry cell is around 200 total and would split to roughly 60 mentored and 140 not, which is below the base needed to detect anything but a large difference. The plan states this in advance: the mature-entry comparison will be reported with the estimate, the interval and the base, with no test, and explicitly labelled as unable to detect a difference of the size that would matter. Doing this in the plan rather than at analysis prevents the far more likely alternative, which is a non-significant result reported as evidence that the scheme does not work for mature entrants.

*Step 8, multiplicity.* Eight planned tests. The two primaries are treated as a family with no correction, on the grounds that they are pre-specified, few, and both will be reported whatever they show. The six subgroup comparisons form a second family with false discovery rate control, and the cost is stated: a genuine but modest difference in one route may not survive.

*Step 9, meaningfulness, and the second judgement call.* The decision owner was asked what difference on the belonging scale would justify extending the scheme across six further faculties. The first answer was "any significant improvement", which was declined, because at these bases a difference too small to justify the cost would reach significance. Pressed, the owner named a threshold based on the resourcing case. That number went into the plan, and it changed the interpretation of the study before it was run: a difference below it would be reported as reliable and insufficient, which is a finding the scheme's sponsors had not previously considered possible.

*Step 11, exploratory register.* The open-ended question about what the scheme changed is exploratory by nature, and its coding will generate hypotheses rather than test them. The register also anticipates the near-certain request for a faculty-by-faculty cut, which is not in the plan, cannot be supported at these bases, and would consist of nine comparisons chosen after the fact. The plan states in advance that such a cut, if produced, will be labelled exploratory with the number of comparisons disclosed.

    OUTPUT

    A locked, dated Analysis Plan: two primary comparisons fully specified with
    tests, thresholds, required bases and meaningfulness thresholds; six
    secondary comparisons with false discovery rate control and one of them
    pre-designated as untestable at its base; two composite scales with exact
    construction, missing-item rules and a stated abandonment rule; a "don't
    know" convention applied study-wide; case exclusion and straight-lining
    rules with blind application; a test count of eight; an empty exploratory
    register with its rules; a base check against the sample design; and a
    deviation log carrying the before-or-after-data field. Circulated to the
    scheme's sponsors three weeks before fieldwork opened.

## 15. Advanced usage

**Blind analysis.** Where the stakes are high and the pressure toward a result is real, run the specified analysis with the group labels masked, complete the interpretation of the patterns, and only then unmask. It is unusual in commercial work and it removes a large class of unconscious choices at almost no cost. It is particularly worth doing where the analyst has a view about which arm should win.

**Interim analysis done properly.** If there is any prospect of stopping early, the plan must specify it: when the look happens, the adjusted threshold, and the stopping rule in both directions (success and futility). Unplanned peeking with the option to stop is one of the most effective ways to produce a false positive, and it is common precisely because it feels prudent.

**Analysis plans for trackers.** A tracker's plan is a convention document as much as an analysis one: base definitions, derived variables, wave-on-wave testing rules, and, critically, the change threshold agreed before the data that would trigger an action. Any change to the analysis convention breaks the series in the same way an instrument change does, so the deviation log matters more here than anywhere else (05.05).

**Pre-specification as a client deliverable.** Circulating the analysis plan to the client before fieldwork changes the debrief conversation fundamentally. The question "why did you not cut it by X" is answered by a document they approved, and the request for a post-hoc cut becomes an explicit, jointly owned decision to run an exploratory analysis rather than a quiet extension of the confirmatory claims.

**Applying the discipline to qualitative work.** The confirmatory/exploratory distinction transfers with modification. A structured qualitative study can pre-specify its coding frame's top level, the prevalence conventions, the comparisons across participant groups it intends to make, and the rule for reporting a theme. What it cannot and should not pre-specify is the emergent lower-level codes, which are the method's purpose (07.01). Stating which parts were fixed in advance is what makes the analysis auditable.

**Recovering from a study with no plan.** Where a dataset arrives with no pre-specification, do not reconstruct one. Write the plan now, date it now, run the specified analyses, and label everything exploratory. Then state in the report how many comparisons were examined. This is a weaker position than pre-specification and a far stronger one than presenting selected findings as though they had been predicted.

## 16. Skill chain

**Recommended previous skills**
- **01.03 Hypothesis Development.** Hands over the register of hypotheses with their confirming and disconfirming patterns, which become the primary comparisons.
- **01.06 Sampling Strategy.** Hands over the achievable bases per reporting cell, without which every specified comparison is provisional.
- **01.02 Business Problem to Research Question.** Hands over the objectives that every analysis must serve.

**Recommended next skills**
- **Category 02 Instrument Design.** The plan constrains the instrument: response options, "don't know" provision, scale points and routing must support the analysis specified here, and it is much cheaper to discover that now.
- **05.01 Descriptive Analysis** and **05.02 Statistical Testing**, which execute the plan and inherit its conventions, thresholds and base definitions.

**Runs well alongside**
- **01.05 Research Plan Development**, which schedules the analysis window and names who runs it.
- **05.05 Trend and Tracker Analysis**, wherever the plan is a continuing convention rather than a one-off.
- **13.01 Research Quality Review** and **13.03 AI Output Verification**, which audit the finished analysis against the locked plan and its deviation log.

---
A Yazi Supplied Skill and resource.
