---
name: hypothesis-development
description: >
  Builds hypotheses that could turn out to be wrong, and knows when a study should
  have none. Use for "what are our hypotheses", "how do I write a testable
  hypothesis", "is this a hypothesis or just a hunch", "what would prove us wrong",
  "should this study have hypotheses at all", "the client already knows what they
  expect to find". Separates hypothesis from expectation, hunch and client belief,
  writes falsifiable statements with their prior basis, and pre-specifies the
  disconfirming evidence before fieldwork.
category: 01 Research Strategy and Design
ref: "01.03"
tier: 1
inherits: [K2, K3, K4, K5]
---

# Hypothesis Development

## 1. One-line description
Turns propositions into hypotheses that could be refuted, names the prior basis for each, specifies in advance what would confirm and what would disconfirm them, and refuses to impose hypotheses where the study is genuinely exploratory.

## 2. What this skill is used for

**The research problem it solves.** Commercial and applied research uses the word hypothesis loosely, and the looseness is expensive. Three different things travel under the name: a proposition with an evidential basis and a stated way of being wrong, an expectation with a basis and no test attached, and a stakeholder's belief with neither. The third kind is the one that does the damage, because once written into a design it shapes the instrument, the sample and the analysis toward its own confirmation, and it does so invisibly. The opposite failure is just as common and less discussed: hypotheses imposed on work that is genuinely exploratory, which narrows the study to what was already thinkable and guarantees it finds nothing new. This skill sorts propositions into their real categories, writes the surviving ones so that a result could refute them, and records the disconfirming evidence before any data exists, which is the only point at which that record is credible.

**Where it sits in the research lifecycle.** After the question set (01.02) and before method selection (01.04), because a testable hypothesis frequently converts an expensive exploratory design into a cheaper confirmatory one. The hypothesis register is then an input to the analysis plan (01.07), which specifies how each one will actually be tested.

**Typical use cases.**
- A question set is settled and the team needs to state what it expects and why.
- A stakeholder arrives with a strong belief and it needs to be either converted into something testable or set aside as a belief.
- A study is being designed where prior evidence exists and should be used rather than re-discovered.
- An exploratory study is being pushed toward premature hypotheses and someone needs to make the case against.
- A proposal or ethics application requires stated hypotheses.
- An experiment or a test-and-learn programme needs a pre-specified primary hypothesis before it can be powered.

**Who uses it.** Research managers and directors specifying studies, product and growth researchers designing tests, evaluation analysts, client-side insight leads translating stakeholder beliefs into research, and postgraduate researchers writing a proposal. For a beginner it is a discipline; for an experienced researcher it is mainly a way of resisting pressure.

## 3. When to use it

- The research questions are settled and the design has not been chosen.
- Prior evidence, theory or the organisation's own data supports a directional expectation worth testing.
- A stakeholder has told you what the study will find, and you need a principled way to handle that.
- You are designing an experiment, a test, or any comparison that will need a pre-specified primary outcome.
- The study will be contested, and the value of stating in advance what would change your mind is high.
- A previous study produced an interpretation that was never tested and is now treated as established.
- The team is arguing about what the study is for, and making the competing expectations explicit would settle it.
- You need to decide, deliberately, that this study should have no hypotheses.

## 4. When NOT to use it

- **When the work is exploratory by design.** Discovery interviews, problem-finding, ethnographic scoping, early concept work and any study whose purpose is to find out what the questions are. Imposing hypotheses here is not rigour, it is a methodological error: it converts an open enquiry into a checklist, biases the instrument toward what was already imagined, and makes the most valuable possible outcome (a finding nobody anticipated) harder to see and harder to report. The correct output in this case is a short statement that no hypotheses are appropriate, why, and what will be stated instead: an area of enquiry, the sensitising concepts in play, and a stopping rule.
- **When no prior basis exists.** A directional statement with nothing behind it is a guess, and at the design stage a guess is worse than silence because it will shape the instrument. Where there is no prior evidence, no mechanism and no operational data, say so and state no hypothesis. Absence is a legitimate and common output.
- **When the study is descriptive measurement.** "What proportion of households hold the product, by region" needs a sample and an analysis plan, not a hypothesis. Hypothesis-fitting descriptive work produces a null-hypothesis ritual that adds nothing and invites significance-hunting across dozens of subgroups.
- **When the design cannot test the hypothesis you would write.** A causal hypothesis on a cross-sectional design cannot be tested, only illustrated. Either change the design (**01.04 Research Method Selection**) or rewrite the hypothesis as associational and say what it does not establish (K4 §3.2).
- **When the proposition is a stakeholder's conclusion rather than a proposition about the world.** "Our pricing is the problem" is a position in an internal argument. It can sometimes be converted into a testable statement, and where it cannot, it is recorded as a belief held by a named person, which is information about the organisation, not about the market (K4 §4.2).
- **When the question set does not yet exist.** Hypotheses derived from an unformed question set inherit its ambiguity and lock it in. Use **01.02 Business Problem to Research Question** first.
- **When what is needed is the analysis specification rather than the propositions.** Tests, thresholds, subgroup rules, multiplicity handling and the confirmatory/exploratory boundary belong to **01.07 Analysis Plan Development**. This skill states what will be tested and what would refute it; 01.07 states how.
- **When a tracker is being designed whose purpose is monitoring.** A tracker's value is continuity, and its measures do not each need a hypothesis. What they need is a change threshold and an owner, which is a different discipline (**05.05 Trend and Tracker Analysis**).
- **For an academic conceptual framework requiring theory selection and construct operationalisation at doctoral standard.** Use **15.08 Research Design and Methodology Chapter**.

## 5. Required inputs

**Required.**
- **The research question and sub-questions**, in the form 01.02 produces them. Hypotheses are statements about the answers to specific questions; without the questions they attach to nothing.
- **The propositions currently in play**, in the words of whoever holds them, with the holder named. Not a tidied version: "we think younger customers are more price-sensitive" and "the youth segment churns on price" are different claims with different testability.
- **The prior evidence, theory or operational data available**, or an explicit statement that none exists. This determines whether anything survives Step 2, and it is the input most often assumed rather than gathered.

**Optional, and what each one adds.**
- **The intended design, if already chosen.** Lets you screen out hypotheses the design cannot test before they are written into a proposal.
- **Prior waves or comparable studies with their effect sizes.** Convert a directional expectation into a magnitude expectation, which is what makes powering and meaningfulness thresholds possible later.
- **The decision and its options (01.01, 01.02).** Lets each hypothesis be scored on what its confirmation or disconfirmation would actually change, which is how the set gets cut to a usable size.
- **The organisation's operational or behavioural data.** Often the strongest available prior basis, and frequently better than published evidence because it concerns this population.
- **A named sceptic.** Someone who disagrees with the expected direction. Their objection is the fastest route to a rival hypothesis, and a set containing a genuine rival is far stronger than a set of variations on one story.

## 6. Questions to ask before starting

1. **Should this study have hypotheses at all?** Why it matters: it is a real question with a real answer, and the default answer in commercial work is wrongly yes. Default if unanswered: examine whether the study's purpose is to find out what is happening (no hypotheses) or to test something specific (hypotheses), and state the judgement explicitly in the output.
2. **What is the prior basis for each proposition, and can it be named?** Why it matters: it is the sole test that separates a hypothesis from a hunch. Default: treat unsourced propositions as hunches and place them in the exploratory register, not the hypothesis register.
3. **Who holds this belief, and what happens to them if it is wrong?** Why it matters: it identifies the hypotheses that will be defended rather than tested. Default: record the holder for every proposition; the field costs nothing and is frequently the most useful column in the table.
4. **What would you accept as evidence that this is wrong?** Why it matters: if nobody can answer, the proposition is not falsifiable in practice, whatever its logical form. Default: mark the hypothesis as not operationally falsifiable and either rewrite it or drop it.
5. **What changes depending on the result?** Why it matters: a hypothesis whose confirmation and disconfirmation lead to the same action is decoration, and it will still consume instrument length. Default: apply per hypothesis and mark the failures.
6. **Can the intended design actually test this?** Why it matters: causal hypotheses on non-causal designs are the commonest unenforceable statement in applied research. Default: assume the design is cross-sectional unless told otherwise, and rewrite any causal statement accordingly with the loss recorded.
7. **What would a sceptic in this organisation say instead?** Why it matters: it produces the rival hypothesis, which is what turns a confirmation exercise into a test. Default: construct the strongest rival you can and mark it as constructed rather than held.

## 7. Step-by-step methodology

**Step 1. Sort every proposition into one of four registers.**
Take the propositions as stated, by whoever stated them, and classify each. A **hypothesis** has a prior basis that can be named and a stated way of being wrong. An **expectation** has a basis but no test attached, usually because it concerns something the study will not measure. A **hunch** has no basis: it may be right, and it is not evidence about anything. A **client or stakeholder belief** is a position held by a named person, which is a fact about the organisation and never a premise about the world. A correct result is a table where every proposition carries a register and a holder, and where the belief rows are visibly separated so that they cannot be promoted later by proximity. This separation is the single most useful act in the skill, and it is done before anything is rewritten, because rewriting first launders a belief into a hypothesis.

**Step 2. Establish and record the prior basis for each candidate.**
Three sources are admissible. **Prior evidence**: a named, dated study or dataset, with its population and its limitations noted, because a finding from a different population is a weaker prior than it looks. **Mechanism or theory**: a stated reason why the relationship would hold, articulated as a chain that someone could disagree with, not as a label. **Operational data**: the organisation's own records, which are frequently the strongest prior available and the least used. Not admissible: "it is well known", "everyone in the category believes", "it is intuitive", and prior AI output. A correct result is that some candidates lose their prior basis and drop out of the hypothesis register into the exploratory register, and the register is smaller than the input list. If nothing drops out, the step was not run honestly.

**Step 3. Apply the consequence test.**
For each surviving candidate, write what the organisation does if it is confirmed, and what it does if it is disconfirmed. Two different actions, or the hypothesis is not earning the instrument length, the sample or the analysis time it will consume. This is the same test 01.02 applies to questions, applied here to propositions, and it typically halves the set. Note the near-miss: "we would have more confidence" is not an action.

**Step 4. Write the falsifiable statement.**
Each surviving hypothesis is rewritten in a fixed form: **population + measure + relationship + direction (if warranted) + boundary conditions**. "Customers are price-sensitive" becomes "among household customers on the standard tariff, stated likelihood of switching rises as the monthly charge rises, and the rise is steeper for customers in the first year of tenure than for those beyond it". Three tests on the rewrite. First, can you write the sentence describing the world in which it is false, in the same measurement units? If not, it is not falsifiable and must be rewritten or dropped. Second, does every construct in it have a measurable referent, or does it contain a word like engagement, trust or loyalty with no operational definition attached? Unoperationalised constructs make a hypothesis unfalsifiable by making it always arguably true. Third, does it contain more than one claim? Compound hypotheses cannot be refuted cleanly, because half of them can survive. Split them.

**Step 5. Decide directional or non-directional, on the strength of the prior.**
A directional hypothesis states which way the relationship runs and is warranted only where the prior specifies a direction. A non-directional hypothesis states that a relationship exists without committing. The choice is not a matter of confidence or preference: it is a matter of what the prior evidence supports. Practice is genuinely contested on the associated testing choice. One position holds that a directional prior justifies a one-sided test, which uses the available sample more efficiently. The opposing position, which is the stronger one in most applied commercial and policy work, holds that one-sided testing forfeits the ability to report a large effect in the unexpected direction, and that the unexpected direction is frequently the commercially important result. State the position taken and the reason, and hand the testing decision to 01.07. A correct result records the direction, the prior that licenses it, and what will happen if the effect appears strongly in the other direction.

**Step 6. State the null in operational terms.**
The null is not "nothing happens". It is the specific state of the world the design is required to rule out, written in the study's own measures: "there is no difference in mean switching likelihood between first-year and later-tenure customers on the standard tariff". Three things must be said alongside it, because they are the ones misunderstood in practice. Failing to reject the null is not evidence that the null is true; it is a statement that this design did not detect a difference, and its meaning depends entirely on what the design could have detected. A design too small to detect the difference that would matter cannot produce an interpretable null, which makes base adequacy a hypothesis-development problem and not only a sampling one (see 01.06 and K3 §3.2). And the null is not the interesting statement: it is the thing the evidence must overcome, and reporting a study as though the null were the finding is a common and avoidable error.

**Step 7. Pre-specify the confirming evidence.**
For each hypothesis, write what a confirming result actually looks like, in measurable terms: the measure, the comparison, the groups, the direction, the minimum size of the difference that would count, and the base on which it must be observed. "Satisfaction is higher among the treated group" is not pre-specification. "The mean on the six-item service scale is at least four points higher among the treated group than the comparison group, on bases of at least 200 in each" is. Where the meaningful size cannot be set without the decision owner, that is a question for them and it is asked now, not after the data arrives, because after the data arrives every threshold is contaminated by knowledge of the result.

**Step 8. Pre-specify the disconfirming evidence, and write it before fielding.**
This is a separate step because it is the one that gets skipped, and skipping it is what makes a hypothesis unfalsifiable in practice while remaining falsifiable in form. For each hypothesis write the sentence: *we would abandon this hypothesis if...*, in the same measurable terms. Then check it against the design: could the study actually produce that result? If the study is structurally incapable of producing the disconfirming pattern, the hypothesis is untestable by this design and must be either rewritten or removed. Date and lock this section before any data collection. A disconfirmation rule written after the data has arrived is not a rule, it is a rationalisation, and the date on the document is what distinguishes them.

**Step 9. Write the three-outcome consequence table.**
Three rows per hypothesis, not two: confirmed, disconfirmed, and inconclusive. The inconclusive row is the one teams forget and the most likely outcome for any under-powered comparison. It states what the organisation does when the study neither confirms nor rules out, which is usually to proceed on the prior with a monitoring plan, and saying so in advance prevents an inconclusive result being reported as a confirmation by omission.

**Step 10. Check the set for coverage, contradiction and rivals.**
Read the hypotheses together rather than singly. Three checks. Are any of them mutually exclusive, so that the study can discriminate between them rather than merely accumulating support for one story? A set in which every hypothesis is consistent with the same narrative is a narrative, not a set. Is there at least one rival explanation for the phenomenon of interest, drawn from the sceptic or constructed deliberately? Designing to discriminate between competing explanations is stronger than designing to accumulate support for one. And is the set small enough for the design to test, given that each hypothesis consumes instrument length, sample and, at analysis, a share of the multiplicity budget (01.07)?

**Step 11. Screen the set for the hypothesis that exists to be confirmed.**
For each, ask three questions: whose belief is this, what does confirmation buy them, and what would disconfirmation cost them. Where the answers point to a hypothesis with a sponsor and a stake, three protections apply. The prior basis must be evidential rather than positional. The disconfirmation rule must be written and agreed by that sponsor before fieldwork, which is the moment at which it is cheapest for them to agree to it. And the study's reporting commitment must be stated in advance: the result will be reported whichever way it comes out. A hypothesis that cannot survive those three conditions is not a hypothesis, and continuing to call it one is the ethical failure this skill exists to prevent. Mark it `RESEARCHER DECISION REQUIRED` per K5 §2.4.

**Step 12. Lock, date and hand over.**
Publish the register with a date, before data collection. Everything added after that date is exploratory, is labelled so, and travels into the exploratory section of the analysis plan rather than the confirmatory one. This is not a formality: the date is the entire mechanism by which a reader can distinguish a prediction from an explanation.

## 8. Analytical framework

The sorting ladder, applied first:

    Client or stakeholder belief   (a fact about a person)
      → Hunch                      (no prior basis)
        → Expectation              (prior basis, no test attached)
          → Hypothesis             (prior basis, and a stated way of being wrong)

Movement up this ladder requires evidence, not rewording. A belief becomes a hypothesis only by acquiring a prior basis and a disconfirmation rule, and never by being phrased more formally.

The specification chain, applied to everything that reaches the top rung:

    Proposition
      → Prior basis          (evidence, mechanism, or operational data, named and dated)
        → Falsifiable statement (population, measure, relationship, direction, boundary)
          → Null, in operational terms
            → Confirming pattern    (measure, comparison, threshold, base)
              → Disconfirming pattern (written before fielding, dated)
                → Consequence        (confirmed / disconfirmed / inconclusive)

Read the chain downward when building and upward when checking. The upward read is where the failures show: a consequence that is the same in all three rows means the hypothesis fails Step 3; a disconfirming pattern the design cannot produce means it fails Step 8; a falsifiable statement whose prior basis is a person's name means it never left the belief register.

## 9. Output format

A **Hypothesis Register**, dated, in this order.

**1. Hypothesis decision.** A short statement of whether this study should carry hypotheses, and why. Where the answer is no, this section carries the reasoning and the alternative (area of enquiry, sensitising concepts, stopping rule), and the rest of the document is not produced.

**2. Proposition sort.**

| Proposition, as stated | Held by | Register (hypothesis / expectation / hunch / belief) | Reason for the classification |

**3. Hypothesis specification**, one block per hypothesis.

| Field | Content |
|---|---|
| H number and statement | Population, measure, relationship, direction, boundary conditions |
| Prior basis | Evidence (named, dated, with population), mechanism, or operational data |
| Directional? | Yes/no, and the prior that licenses it |
| Null, operationally | The state the design must rule out |
| Confirming pattern | Measure, comparison, threshold, required base |
| Disconfirming pattern | The result that would cause abandonment |
| Design can produce disconfirmation? | Yes/no |
| If confirmed / disconfirmed / inconclusive | The action in each case |
| Sponsor and stake | Who holds it, what confirmation buys them |

**4. Rival hypotheses.** Stated in the same form, marked as held or constructed.

**5. Exploratory register.** The hunches and unsupported propositions, retained rather than deleted, marked as generating exploratory analysis only and never confirmatory claims.

**6. Beliefs recorded, not tested.** Named holder, statement, and why it was not converted. This section protects the researcher and the client equally.

**7. Lock date and sign-off.** The date before which the register was fixed, and the named person who agreed it.

**When the evidence is thin.** The honest outputs of this skill include: *no hypotheses are appropriate for this study*, with the reasoning; *this proposition has no prior basis and is recorded as a hunch*; and *this hypothesis cannot be tested by the intended design*. Each is a complete result. Do not populate a hypothesis block because the template has one. A prior basis field that reads "industry consensus" is an empty field with words in it (K4 §1, §2.4).

## 10. Quality checks

Run before the register is locked. These sit on top of K4 §8.

1. Has the question "should this study have hypotheses at all" been asked and answered explicitly?
2. Does every proposition in the input appear in the sort table, including the ones that were rejected?
3. Does every hypothesis name a prior basis that could be checked by someone else, with a date and a population?
4. Can you write, for each hypothesis, the sentence describing the world in which it is false, in the study's own measures?
5. Does any hypothesis contain a construct with no operational definition?
6. Is any hypothesis compound, so that it could be half-refuted?
7. Is the direction, where stated, licensed by the prior rather than by preference, and is the handling of an opposite-direction result stated?
8. Is the null written operationally rather than as "no difference" in the abstract?
9. Does every confirming pattern name a threshold and a required base, not just a direction?
10. Was every disconfirming pattern written before any data existed, and is the register dated to prove it?
11. Could the intended design actually produce each disconfirming pattern?
12. Does the consequence table have three rows, including inconclusive?
13. Does the set contain at least one genuine rival explanation, or a statement of why none exists?
14. Has every hypothesis with a sponsor and a stake been through the Step 11 screen?
15. Is the exploratory register separate, and is it clear that nothing in it can produce a confirmatory claim?

## 11. Common failure modes

| Failure | How to recognise it | How to prevent it |
|---|---|---|
| **The belief in formal clothing** | A stakeholder's position, rewritten in hypothesis grammar, with no prior basis | Step 1 sorts before Step 4 rewrites. Rewording never changes the register |
| **The unfalsifiable construct** | "Customers will be more engaged" with no definition of engaged | Step 4 test two: every construct needs a measurable referent |
| **The compound hypothesis** | Two claims joined by "and", where one could hold and the other fail | Split at Step 4. Each hypothesis is refutable on its own |
| **Disconfirmation written afterwards** | The abandonment criterion appears in the analysis document, not the design one | Step 8, dated and locked before fielding. The date is the mechanism |
| **The untestable causal claim** | A causal hypothesis on a cross-sectional design | Step 8's design check, plus rewriting as associational per K4 §3.2 |
| **Hypotheses on exploratory work** | A discovery study with eight hypotheses and a structured guide | The Step in Section 4: state that none are appropriate, and say what replaces them |
| **The null misread** | A non-significant result reported as "no difference exists" | Step 6: state what the design could have detected, and the base needed |
| **Hypothesis inflation** | One per sub-question, symmetrical, none prioritised | Step 3's consequence test, applied individually |
| **No rival** | Every hypothesis consistent with the same story | Step 10: construct the sceptic's version deliberately |
| **AI: prior invention** | A confident prior basis citing evidence that was never supplied | K4 §2.4. A partially remembered source is not a source. Prior basis is supplied or absent |
| **AI: symmetric filling** | A hypothesis generated for every question because the section exists | Absence is a valid output. The exploratory register is where unsupported propositions go |
| **AI: direction from plausibility** | A directional hypothesis whose only basis is that the direction seems likely | Direction requires a prior that specifies it. Plausibility is not a prior |
| **AI: threshold invention** | A meaningful-difference threshold produced without the decision owner | The threshold is a business judgement (K5 §2.1). Ask, or mark it unset |

## 12. AI guardrails

Skill-specific. K4 applies in full and is not repeated here.

1. **Never state a prior basis that was not supplied.** No remembered studies, no "research shows", no constructed citation. If no prior exists, the proposition goes into the exploratory register and the hypothesis is not written (K4 §2.4).
2. **Never promote a stakeholder belief into a hypothesis by rewording it.** Register change requires new evidence, and the record of who held it survives the rewrite.
3. **Never write a hypothesis whose disconfirming pattern the design cannot produce.** An untestable hypothesis in a design document licenses a confirmation that no result could have prevented.
4. **Never generate a hypothesis because a template section is empty.** The statement "no hypotheses are warranted; no prior basis exists" is a correct and complete output.
5. **Never set a meaningfulness threshold on the AI's own judgement.** The smallest difference that would change an organisation's action is a business fact the AI does not have (K5 §2.1). Ask for it, or mark it unset and flag the consequence for the analysis plan.
6. **Never state a direction without naming what licenses it**, and never suppress the handling of a result in the opposite direction.
7. **Never date a register retrospectively or reconstruct one after data has been seen.** If the register was not locked before fielding, say so plainly; the study's hypotheses are then exploratory, and calling them confirmatory is a misrepresentation of the evidence (K2 §7).
8. **Never impose hypotheses on exploratory work to make it look more rigorous.** The narrowing is real, it is invisible in the output, and it damages the studies that most need openness.
9. **Never report a failure to reject as evidence of no effect** in any downstream document, and where a null is stated, name the effect size the design could have detected.

**Human review points** (see K5 §2 for the classes):
- **Researcher decision required** on any hypothesis with an identified sponsor and stake, per Step 11. Class 2.4, ethical appropriateness.
- **Researcher decision required** on whether the study should carry hypotheses at all where the work sits between exploratory and confirmatory. Class 2.7.
- **Researcher sign-off required** on the locked register before fieldwork, because the lock date is the study's integrity record and it must be owned by a named person. Class 2.8.

## 13. Best-practice principles

1. **A hypothesis you cannot lose is not a hypothesis.** The test is not whether it is plausible but whether a specific, obtainable result would end it.
2. **Write the disconfirmation before the confirmation.** Reversing the usual order is uncomfortable and it is the single highest-value habit in this skill, because it forces the design question of whether such a result is even obtainable.
3. **The prior basis is the whole distinction.** Between a hypothesis and a hunch there is nothing else. Insisting on a named, dated source is not pedantry; it is what stops the design encoding the loudest voice in the room.
4. **Record who holds each proposition.** The column costs nothing and predicts, with some accuracy, which findings will be contested and which will be quietly dropped.
5. **Prefer a small set with a genuine rival to a large set that all points the same way.** Discriminating between two explanations is worth more than accumulating support for one, and it costs less instrument length than people expect.
6. **The absence of hypotheses is a design position, not a gap.** State it deliberately, with the reasoning, and the study is stronger for it. Exploratory work stated as exploratory is rigorous; exploratory work dressed as confirmatory is not.
7. **A null result is only interpretable against what the design could have detected.** Agree the detectable difference at design time or accept that a null will mean nothing.
8. **Set the meaningfulness threshold with the person who will act on it.** Statistical significance is a property of the design; meaningfulness is a property of the business, and only one of those two is yours to decide.
9. **Hypotheses are a budget, like questions.** Each consumes instrument length, sample and multiplicity allowance. Adding one is a trade against another.
10. **Beware the hypothesis that explains everything.** A proposition consistent with every possible result is not powerful, it is empty, and it will be confirmed by whatever comes back.
11. **Keep the exploratory register rather than deleting it.** Hunches are legitimate sources of exploratory analysis and of the next study's hypotheses. What they must never do is arrive in a report as though they had been predicted.
12. **The date on the register is the evidence.** Everything about the confirmatory/exploratory distinction rests on being able to show what was written down before the data arrived.

## 14. Worked example

Fictional scenario, financial services. All figures are illustrative and belong to the scenario.

    INPUT

    A savings provider's product team is designing a study before deciding
    whether to invest in promoting an automatic round-up saving feature.
    Three propositions are in play:

    (a) Product lead: "Customers do not use round-up because they do not
        know it exists."
    (b) Marketing lead: "Younger customers will adopt it if we promote it."
    (c) A service manager: "The people who would benefit most are the ones
        who cannot afford the variability in outgoings."

**Process.**

*Step 1, sort.* (a) is stated as fact and held by the person who commissioned the study: a belief, until a prior is found. (b) has no basis offered: a hunch. (c) has a mechanism behind it (variable outgoings are hard to absorb on a tight budget) and can be checked against the organisation's own account data: a candidate hypothesis.

*Step 2, prior basis.* The team's own product analytics were pulled. Feature-page visits were recorded for a substantial minority of the customer base, and among those visitors activation was low. That is a prior, and it points against (a): awareness alone does not explain non-use, because a group that clearly encountered the feature did not activate. (a) is rewritten rather than dropped, into a weaker and testable form about the awareness gap in the non-visiting group. (b) acquires a partial basis from the same source (activation among visitors skewed to one age band) which supports a non-directional statement about age, not the directional one marketing wanted. (c) is supported by mechanism plus a visible association in the account data between balance volatility and non-activation.

*Step 4 and Step 5, specification and the judgement call.* Marketing pressed for a directional hypothesis on age, because the campaign plan assumed it. The prior showed an association in activation among visitors, but visitors are a self-selected group and the direction in the wider base is not established by it. The judgement call was resolved by writing the hypothesis non-directionally, recording the marketing expectation in the beliefs section with its holder, and stating explicitly that a strong result in the opposite direction would be reported and would change the campaign targeting. The cost of one-sided efficiency was judged smaller than the cost of being unable to report the surprise.

*Step 8, disconfirmation.* For (c): the hypothesis would be abandoned if activation rates showed no relationship with balance volatility across the tenure bands, on adequate bases, and if customers with high volatility gave reasons for non-use that did not mention the unpredictability of outgoings. Both are obtainable from the planned design. For the rewritten (a): abandoned if, among customers who have never visited the feature page, a description of the feature does not raise stated likelihood of activation relative to their current state.

*Step 11, the sponsor screen.* Hypothesis (a) belongs to the person who commissioned the study, and its confirmation would justify a promotional budget already drafted. The three protections were applied: the prior basis was drawn from account data rather than from the sponsor's view, the disconfirmation rule was agreed by the sponsor in writing before fieldwork, and the study's reporting commitment was stated. A researcher decision marker was placed on this hypothesis under K5 §2.4.

    OUTPUT

    A dated Hypothesis Register: three specified hypotheses (one of them a
    rewritten and weakened version of the sponsor's belief, one non-directional
    on age, one on affordability with a mechanism), each with a named prior
    basis, an operational null, confirming and disconfirming patterns with
    thresholds and required bases, and a three-row consequence table. A rival
    hypothesis (that non-use reflects a preference for manual control rather
    than either awareness or affordability) was constructed and added. The
    marketing directional expectation was recorded in the beliefs section with
    its holder named. Locked and signed off two weeks before fieldwork.

## 15. Advanced usage

**Strong inference.** Where two or more explanations for the same phenomenon are live, design to discriminate rather than to confirm. Write the hypotheses so that each predicts a pattern the others do not, then check that the design can observe the distinguishing pattern. This is more efficient than testing each in turn, and it is the difference between a study that ends an internal argument and one that supplies ammunition to both sides.

**Hypotheses in mixed-method designs.** In a sequential design, the qualitative strand generates the propositions and the quantitative strand tests them. Keep the registers separate and dated: propositions generated in strand one are exploratory with respect to strand one and confirmatory with respect to strand two, and only if they were written down between the two.

**Working with a strong prior.** Where the prior evidence is very strong, ask whether the study is worth running at all. If the confirming result is near-certain, the study buys little; the value sits in the disconfirming tail, and the design should be built to detect that rather than to reproduce the expected result at higher precision.

**Where the hypothesis is about a mechanism, not an outcome.** Outcome hypotheses ("the treated group scores higher") are easy to test and often uninformative. Mechanism hypotheses ("the effect operates through reduced perceived effort rather than through increased perceived value") require measuring the intermediate step, which is a design decision that must be made now rather than reconstructed at analysis.

**Pre-registration outside academia.** Locking and circulating a dated hypothesis register achieves most of what formal pre-registration achieves, at almost no cost, and it is increasingly expected in evaluation work and in any study likely to be challenged. Circulate it to the stakeholders, not only to the research team, because its protective value depends on them having seen it before the result.

**Retro-fitting a register to an inherited study.** Where a study is already fielded and someone asks what the hypotheses were, do not construct them. Record that the study was not hypothesis-driven and that all comparisons are therefore exploratory. That is the honest position and it is far more useful than a plausible reconstruction, which converts every subsequent finding into an apparent confirmation.

## 16. Skill chain

**Recommended previous skills**
- **01.02 Business Problem to Research Question.** Hands over the research question, the sub-questions and the information needs. Hypotheses are statements about the answers to those questions and cannot be written before them.
- **01.01 Research Brief Interrogation.** Hands over the stakeholder map and the record of any expected finding stated in the brief, which is the input to Steps 1 and 11.

**Recommended next skills**
- **01.04 Research Method Selection.** Receives the hypotheses and frequently converts an exploratory design into a cheaper confirmatory one.
- **01.07 Analysis Plan Development.** Receives the register and specifies how each hypothesis will be tested: the comparison, the test, the threshold, the multiplicity treatment, and the confirmatory/exploratory boundary.

**Runs well alongside**
- **01.06 Sampling Strategy**, because a hypothesis is only testable at a base that could detect the difference that would matter.
- **05.02 Statistical Testing**, which executes the tests the register specifies.
- **10.01 Literature Review and Desk Research**, for establishing and verifying the prior basis at Step 2.

---
A Yazi Supplied Skill and resource.
