---
name: pattern-identification
description: >
  Finds and tests the structure across a scattered set of findings before any
  insight is developed: names the pattern types worth looking for, screens out
  patterns produced by the instrument or by chance, and reports a small set of
  tested structural claims. Use for "what's the pattern here", "these findings
  feel related", "connect the dots", "is there a theme across all of this",
  "the findings are scattered", "does this hold everywhere or just in one place".
category: 08 Insight Development
ref: 08.02
tier: 1
inherits: [K2, K3, K4, K5]
---

# Pattern Identification

## 1. One-line description
Takes a set of established findings and determines, by structured search and deliberate falsification, whether real structure connects them, screening out the patterns produced by the instrument, by shared bases, and by chance, and reporting only the structural claims that survive.

## 2. What this skill is used for

**The research problem it solves.** A study finishes and produces thirty findings that sit in no particular order. Someone has to say what connects them. That step is done badly almost everywhere, because it is done by reading the findings list and noticing what feels related, and feeling related is a property of the reader rather than of the data. Three findings that share a word get grouped. Five findings that support a familiar story get promoted into a theme. A pattern across a small number of findings gets treated as structure when it is exactly what randomness produces at that scale. The result is a report organised around a shape that was imposed rather than discovered, and everything downstream inherits it: the insight explains a pattern that is not there, the implication follows from the insight, and the recommendation spends money. This skill supplies the search procedure, the screens, and the falsification test that separate a real structural claim from a satisfying arrangement of the same information.

**Where it sits.** Early synthesis. It runs after analysis has produced findings and before **08.01 Finding to Insight Development** attempts to explain any of them. Explaining a pattern that does not exist is the most expensive error available in synthesis, because it is invisible after the first step.

**Typical use cases.**
- A completed analysis with a long findings list and no organising structure.
- Mixed-method work where the two streams produced separate findings lists that nobody has connected.
- Multi-market, multi-wave or multi-site studies where the question is whether a result is general or local.
- Programme evaluations with many outcome measures and no obvious spine.
- Checking a proposed theme, drafted by a colleague or generated by AI, before it is built on.
- Deciding whether a set of subgroup differences is a pattern or a multiplicity artefact.

**Who uses it.** Research directors and senior analysts responsible for the shape of a report; evaluators working with many outcome measures; insight managers handed a findings deck and asked what it adds up to; anyone reviewing AI-drafted synthesis, where lexical grouping is systematically mistaken for structure.

## 3. When to use it

- Analysis is complete and the findings do not yet have an order that came from the data rather than the report template.
- Several findings look related and you want to know whether they are, before explaining why.
- A subgroup difference has appeared in several places and you need to know whether it is one effect or several coincidences.
- A finding holds in one market, wave or site, and the report is about to describe it as general.
- Something expected did not appear, and you suspect the absence is itself the structure.
- A colleague or an AI draft has proposed a theme and you are deciding whether to build on it.
- The study has many outcome measures and the number of possible comparisons is large enough that some will look patterned by chance.
- The findings arrived from separate workstreams and nobody has yet checked whether apparent agreement between them is real or an artefact of shared method.

## 4. When NOT to use it

- **No findings have been supplied.** Pattern identification operates on findings, each with a source and a base per K2 §4. Raw responses, transcripts or a dataset are not findings. Run the analysis first: **05.01** to **05.03** for quantitative, **07.01** for qualitative. Searching raw material for patterns without an analysis step in between is data dredging with a professional label, and it produces exactly the results K4 §1 warns about.
- **There are too few findings for structure to be distinguishable from coincidence.** Below roughly six to eight independent findings, almost any arrangement will look patterned, because the number of possible arrangements is small and the human eye completes them. Report the findings individually, say that the set is too small to support a structural claim, and stop. Naming a "pattern" across three findings is the single most common way a study acquires a spine it did not earn.
- **The findings are not independent.** Findings drawn from the same battery, the same routing branch, the same subgroup base or the same coder are not separate evidence. A pattern across them may be one finding described five times. Section 7 step 6 screens for this; where the screen removes most of the set, there is no pattern to identify and the honest output says so.
- **The question is why the pattern is true.** That is **08.01 Finding to Insight Development**. This skill establishes that structure exists and defines its scope and exceptions. It does not explain it, and it deliberately stops short of explanation, because a pattern accompanied by its explanation is much harder to reject and therefore much less likely to be tested.
- **The question is which findings matter.** Prioritisation, materiality and sizing are **08.05 Insight Prioritisation and Sizing**. A pattern is not important because it is a pattern.
- **A theme has already been chosen and evidence is being fitted to it.** Where a stakeholder, a brief or a previous deck has supplied the structure and the task is to arrange findings under it, this is confirmation search, prohibited by K4 §4.2. Say so, and offer instead an explicit test of the proposed structure against the whole findings set, including the findings it does not cover.
- **The apparent pattern is in the instrument rather than the world and cannot be resolved.** Where question order, wording, routing or sample composition could produce the observed structure and the method documentation is unavailable, the candidate cannot be cleared. Report it as unresolvable, name the documentation that would resolve it, and do not pass it to 08.01.
- **The pattern concerns a small or identifiable subgroup on a sensitive characteristic.** A structural claim about a group can be harmful in ways an individual finding is not, because it attributes a shape to the group. That is **13.05 Research Ethics and Consent Design** before it is an analytical question.

## 5. Required inputs

**Required.** Without these the skill cannot run. If the first is absent, stop.

- **A set of established findings**, each stated as a finding with its source reference, base description and base size per K2 §4. **If findings have not been supplied, do not proceed.** Say what is missing and ask for the analysis output.
- **The provenance of each finding**: which question or theme it came from, which stream, which base, which subgroup, which wave. Without provenance the instrument screen in step 6 cannot be run, and a pattern cannot be distinguished from an artefact of shared measurement. Where provenance is partial, the affected findings are marked and the pattern's confidence is capped.
- **The study design and instrument**, or at minimum the questionnaire, discussion guide or observation protocol. Question order and routing are the commonest source of false patterns and cannot be checked without them.

**Optional, and what each one adds.**

- **The full dataset or transcript corpus:** allows a candidate pattern to be tested against evidence that was not used to build it, which is the only test that can actually fail. Without it, falsification is limited to the findings already written up.
- **The design hypotheses, prior waves, or the theory the study was built on:** supplies pre-specified patterns, which are weighted differently from patterns found by searching. Without them, every pattern is exploratory and must be labelled as such.
- **Subgroup, site, market or wave structure:** makes it possible to distinguish a finding that holds everywhere from one that holds in one place, which is a different claim with different consequences.
- **Behavioural or operational data alongside the research data:** the strongest available independent measure, and the one most likely to break a pattern built on self-report.
- **Field notes, moderator debriefs or fieldwork logs:** capture events during fieldwork that produce apparent structure (a news event, a service outage, a competitor promotion) and that no dataset records.
- **The findings from the previous study on the same question:** a pattern that also held before is stronger; a pattern that appears only this time is often about this time.

## 6. Questions to ask before starting

1. **What structure, if any, was expected before the data was seen?** Determines which patterns are pre-specified and which are exploratory, and this distinction carries more weight than any other in the whole method. *Default if unanswered:* treat every pattern as exploratory, label it so, and cap confidence accordingly per step 3.
2. **How many findings are there, and how many are independent of one another?** Sets the base rate and determines whether any structural claim is supportable at all. *Default:* count them, report the search space explicitly, and stop if independence falls below the threshold in Section 4.
3. **What is the instrument's structure: question order, batteries, routing?** Determines whether apparent co-occurrence is shared measurement. *Default:* request the instrument; if unavailable, mark every co-occurrence pattern as instrument-unscreened.
4. **Does the study have a subgroup, site or wave structure to test against?** Determines whether universality can be assessed at all. *Default:* report every pattern as holding within the sample as a whole, with no claim about generality across subgroups.
5. **Which decision does this study inform?** Determines which of the seven pattern types is worth searching hardest for; a decision about sequencing needs sequence patterns, a decision about targeting needs segment-specific effects. *Default:* run all seven passes and let the evidence set the emphasis.
6. **Is any evidence available that was not used to produce these findings?** Determines whether falsification is possible or only notional. *Default:* state that testing was confined to the findings themselves, and cap confidence at moderate.
7. **Has a theme, story or structure already been proposed by anyone?** Determines whether the work is a search or a test, and these require different handling. *Default:* ask, because an unstated prior structure will otherwise shape the search invisibly.

## 7. Step-by-step methodology

**The position this method takes.** Human beings are exceptionally good at finding structure, including where there is none. That capability is not a bias to be corrected at the margin, it is the central mechanism of this task and the single largest risk in this skill. A findings list is exactly the kind of material it misfires on: a small number of items, described in words, each open to several groupings, with a strong professional incentive to produce a story. Everything below exists to slow that mechanism down: to make the search explicit rather than intuitive, to make the base rate visible, to screen out the structures that measurement itself creates, and to require that a candidate pattern be sent looking for the evidence that would kill it.

**1. Gate on inputs.** Confirm that findings have been supplied, each with a source and a base. Confirm the count. If findings are absent, stop and ask, per Section 5. If fewer than six to eight independent findings are available, say that the set is too small for a structural claim and report the findings individually. *Correct result:* a numbered findings register, each row a statement of fact about the data that could be checked against source.

**2. Add the properties that make search possible.** For each finding record: the measure or theme it came from, the stream, the base description and size, the population and period, the subgroup or site if it is a subgroup finding, the position in the instrument (which block, which item, what preceded it), the direction and rough magnitude, and whether a test was run. This register is the substrate for everything that follows, and the instrument-position column is the one people omit and then need. *Correct result:* a findings table with those columns filled, and the gaps explicitly marked as gaps rather than left blank.

**3. Pre-specify before searching.** Write down, before looking at the register for structure, what patterns you expect and why: from the design hypotheses, the theory the study was built on, the previous wave, the brief, or the client's stated model. Then, as patterns emerge later, label each one **pre-specified** or **exploratory**. Both are legitimate. They are not equally strong. A pre-specified pattern that appears was predicted and then observed. An exploratory pattern was selected from among all the patterns the data could have shown, and the number of those is large. Pre-specification is a discipline, not a formality: if it is written after the search, it is not pre-specification, and calling it one is the research equivalent of moving the target after the shot. *Correct result:* a dated list of expected patterns written before search begins, kept in the working document whether or not any of them appear.

**4. Compute the search space and state the base rate.** Count what could have looked patterned. With twelve findings there are sixty-six pairs. With nine sites and six outcome measures there are fifty-four site-level comparisons, and at a 95% threshold roughly two to three will look different by chance alone. With four subgroups tested across fifteen questions there are sixty subgroup contrasts. Write the number down and state the expected number of spurious results. This step takes two minutes and it changes what the rest of the work means: a pattern across three of sixty-six pairs is not evidence of structure, it is the arithmetic of sixty-six pairs. Where the search space is large and the pattern is exploratory, no pattern claim is made without holdout confirmation at step 10. *Correct result:* a written statement of the search space and the expected chance yield, appearing in the output, not only in the working file.

**5. Search the seven pattern types, one pass per type.** Do not free-associate across the register. Take each type in turn and ask the register its specific question. This finds structure that intuition skips, and it prevents the search collapsing onto whichever pattern arrived first.

- **Co-occurrence.** Which findings appear together, in the same people, sites or cases? The test is whether they cover the same units, not whether they sound similar.
- **Sequence.** Does one thing reliably precede another, in time, in a journey, or in a process? Sequence patterns are the most useful and the most often missed, because a findings list has no time axis unless you build one.
- **Threshold and non-linearity.** Is there a point at which the relationship changes: a level of usage, a length of tenure, a number of contacts, a price, a wait, beyond which the picture is different? Averages hide thresholds, and a threshold is frequently the only actionable structure in a study.
- **Segment-specific effects.** Does the finding hold for one group and not others, and is the group defined by something the organisation can identify?
- **Absence where something was expected.** What should have appeared and did not: the subgroup difference that failed to materialise, the decay that did not happen, the association everyone assumed. Absence is structure and it is routinely discarded because it does not look like a result.
- **Consistency across independent measures.** Does the same shape appear in measures that do not share a method, a base or an instrument position? This is the strongest form of pattern evidence available and is the only one that survives the screen at step 6 intact.
- **Universal versus local.** Does the finding hold everywhere tested, or in one place? These are different claims. A finding that holds in one market is a finding about that market, and the commonest overreach in multi-market work is reporting it without the qualifier.

*Correct result:* seven completed passes, each with either candidate patterns or an explicit "none found", so that later readers can see what was searched for rather than only what was found.

**6. Screen every candidate for shared cause in the instrument, before treating it as substantive.** This is the screen that separates a pattern about the world from a pattern about how the data was collected, and it must run before any candidate is developed. Check each candidate against six sources of shared cause. **Adjacency and priming:** were the items next to each other, so that answering one framed the next? **Shared battery or common stem:** do the items share a scale, a stem or a response format, so that a respondent's answering style produces correlation between them regardless of attitude? **Shared routing:** were the items asked only of the same filtered subgroup, so that the pattern is a property of the filter? **Shared base:** are the findings computed on the same people, so that they are one observation described several ways rather than several observations? **Shared method or source:** are all the findings self-report, or all from the same coder, the same moderator, the same site team? **Shared time window:** did they all come from a period containing an event? A candidate that fails this screen is not automatically false, but it cannot be advanced on the current evidence; it is either resolved with a different measure or reported as unresolved. Three questions that pattern together because they sat next to each other in a battery and primed one another do not tell you that the underlying attitudes cohere, and the giveaway is that the items correlate with each other about as strongly as any of them correlates with anything outside the battery. *Correct result:* a screen table, one row per candidate, six columns, with the verdict cleared, resolved, or unresolved.

**7. Specify each surviving candidate precisely, in three parts.** Write the **claim** (what the structure is, in a sentence that could be wrong), the **scope** (which findings it covers, which population, which period, which sites), and the **exceptions** (which findings inside that scope it does not cover). A pattern without a stated scope is unbounded and therefore untestable. A pattern without stated exceptions has almost certainly not been checked for any. *Correct result:* a three-part pattern statement, with the covered findings named by reference.

**8. Go and find the exceptions deliberately.** Return to the register and list every finding within the pattern's scope that it does not account for. Count them. The exception rate is a property of the pattern and is reported with it, in the same way a base is reported with a percentage. A pattern covering nine of eleven findings in scope is a strong pattern with two exceptions worth examining. A pattern covering four of eleven, presented without the other seven, is a selection. Exceptions are frequently more informative than the pattern: they mark the boundary of whatever mechanism is at work. *Correct result:* an explicit covered-and-not-covered list, with the exception rate stated.

**9. State what would break the pattern, then look for it.** For each candidate, write the observation that would refute it: a distribution that should not exist if the pattern holds, a site where it should not appear, a measure that should not move. Then go and look, in evidence that was not used to build the pattern. This is the step that makes the difference between a pattern and an arrangement, and it is the step that gets skipped, because the candidate at this stage is usually elegant and nobody wants to lose it. A candidate that cannot generate a breaking observation is not a structural claim, it is a description of the findings list. *Correct result:* a stated breaking condition per candidate, a record of where you looked, and the result, including the cases where the breaking evidence was found.

**10. Split the evidence and confirm out of sample.** Where the data allows, test the pattern in a portion of the evidence that was not used to find it: a holdout set of sites, a later wave, the second half of the interviews, a different measure of the same construct. A pattern found by searching and confirmed in a holdout is a substantially stronger claim than one found by searching alone, and the difference is not rhetorical: exploratory search over a large space is expected to produce apparent structure, and holdout confirmation is the cheapest available correction. Where a holdout is impossible, say so and cap the pattern at moderate confidence per K3 §3.1. *Correct result:* a stated holdout, or an explicit statement that none was available and what that costs the claim.

**11. Decide universal or local, and write the qualifier into the claim.** For each pattern, state where it was tested and where it held. Then write the claim with that qualifier attached permanently, not in a footnote: qualifiers travel with claims per K2 §7, and a pattern that held in two of nine sites will otherwise be quoted as a general result within a fortnight. Where a pattern holds everywhere tested, say how many places that is, because "consistent across all markets" means something different for two markets than for eleven. *Correct result:* every pattern statement carrying its scope inside the sentence.

**12. Rate each pattern and cap it.** Record six things: pre-specified or exploratory; number of independent findings covered; exception rate; instrument screen verdict; falsification attempted and outcome; holdout confirmed or not. Assign confidence per K3 §3, remembering that the level is set by the weakest factor that materially bites, not the average. An exploratory pattern over a large search space, unconfirmed in holdout, is low confidence and is reported as a hypothesis in the K3 §4.3 register regardless of how compelling it looks. *Correct result:* a confidence level per pattern with the specific factor that set the ceiling named.

**13. Prune to a small set of tested structural claims.** The output is not a list of everything noticed. It is the patterns that survived, usually two to five for a substantial study, each stated with scope, exceptions, confidence and breaking condition. Findings that belong to no pattern stay in the register as unpatterned findings; they are not evidence of a failed analysis, and some of them will be the most important things in the study. Do not force a residual pattern to absorb them. *Correct result:* a short pattern register plus an unpatterned-findings list, with the count of each.

**14. Mark the judgement points and hand over.** Per K5, mark where a human is needed: whether an exception is material (§2.1), whether a pattern reads differently in a market whose context you cannot access (§2.2), and whether an unresolved instrument candidate should block the finding entirely. Hand to **08.01** the pattern statement, the covered findings with their references, the exceptions, the confidence, and the breaking condition, so that explanation begins from a bounded structure rather than an impression. *Correct result:* a pattern record another researcher could audit without asking where any of it came from.

## 8. Analytical framework

**The sequence.**

```
Findings register → Pre-specification → Typed search (7 passes) → Instrument screen
   → Three-part specification → Exception mapping → Falsification → Holdout
      → Structural claim with scope, exceptions and confidence → 08.01
```

**The seven pattern types, with the question each one asks.**

| Type | The question | Fails when |
|---|---|---|
| Co-occurrence | Do these appear in the same units? | The findings merely sound similar, or cover different bases |
| Sequence | Does one reliably precede the other? | Order is assumed from the report's order rather than measured |
| Threshold and non-linearity | Is there a point where the relationship changes? | The average was read and the shape was never plotted |
| Segment-specific effect | Does this hold for one identifiable group only? | The segment was defined after the difference was noticed |
| Absence | What should have appeared and did not? | Nobody wrote down what was expected, so nothing can be absent |
| Consistency across independent measures | Does the same shape appear in measures that share no method? | The measures share a battery, base, coder or stem |
| Universal versus local | Does it hold everywhere tested, or in one place? | The qualifier is dropped between the analysis and the summary |

**The three-part pattern statement.** Every pattern is written in this form, and a pattern that cannot be written in it has not been specified.

```
CLAIM       The structure, in one falsifiable sentence
SCOPE       Findings covered (by reference), population, period, sites, waves
EXCEPTIONS  Findings within scope that this does not cover, and how many there are
```

**The base-rate statement.** Included in the output, not just the working file.

```
Findings in register: 14
Independent findings after the instrument screen: 10
Pairwise comparisons available: 45
Site-level comparisons available: 54, of which 2 to 3 expected to differ by chance at 95%
Patterns claimed: 3, of which pre-specified: 1
```

**The instrument screen.** One row per candidate.

| Candidate | Adjacency | Shared battery | Shared routing | Shared base | Shared method | Shared time window | Verdict |
|---|---|---|---|---|---|---|---|

**On narrative structure in noise.** The most important thing in this section is not a table. Three findings that form a story will feel more true than seven that do not, and the feeling is produced by the story rather than by the evidence. The professional countermeasures are all procedural, because the perception is not correctable by effort: state the search space before searching, label exploratory patterns as exploratory, require a breaking condition, and require the exceptions to be counted. A researcher who cannot say how many findings their pattern fails to cover has not tested it.

## 9. Output format

**A. Pattern record** (one per pattern)

```
CLAIM               One falsifiable sentence, with its scope inside it
STATUS              Pre-specified / exploratory
SCOPE               Population, period, sites, waves
COVERS              Findings by reference, with bases
EXCEPTIONS          Findings in scope not covered, by reference, with the rate
PATTERN TYPE        Which of the seven, or which combination
INSTRUMENT SCREEN   Cleared / resolved by [measure] / unresolved
INDEPENDENCE        How many genuinely independent findings support this
BREAKING CONDITION  What observation would refute it
FALSIFICATION       Where you looked, and what you found
HOLDOUT             Confirmed in [holdout] / no holdout available
UNIVERSAL OR LOCAL  Where tested, where held
CONFIDENCE          High / Moderate / Low-hypothesis, with the factor that set the ceiling
REVIEW POINTS       Per K5, at the point of the judgement
```

**B. Pattern register**

| # | Claim | Type | Status | Covers | Exceptions | Screen | Confidence |
|---|---|---|---|---|---|---|---|

**C. Base-rate statement**
The search space, the expected chance yield, and the number of patterns claimed. Required.

**D. Rejected candidates**
Every candidate that failed, with the reason: screened out as instrument-shared, refuted by the breaking evidence, failed in holdout, or too few independent findings. This section is usually longer than section B and is frequently the more useful of the two.

**E. Unpatterned findings**
Findings that belong to no pattern, listed in full, with a line stating that they were not forced into one. Required.

**When the evidence is thin.** Where no pattern survives, the output is sections C, D and E, and it says plainly that the findings do not support a structural claim. This is a legitimate and reasonably common result, particularly in studies with few independent measures. Do not produce a residual pattern to give the report a spine, do not lower the exception threshold to make a candidate fit, and do not relabel a single finding as a pattern by describing it in more general language. A study with eleven findings and no pattern, honestly reported, protects everything downstream. Per K4 §1, the presence of a "themes" section in a template is not evidence that themes exist.

## 10. Quality checks

Run before anything is presented. Sits on top of K4 §8.

1. Were findings supplied, with sources and bases, and is every pattern traceable to specific findings by reference?
2. Is the number of independent findings stated, after the instrument screen rather than before?
3. Was the search space counted and the expected chance yield stated in the output?
4. Was pre-specification written down before the search, and is every pattern labelled pre-specified or exploratory?
5. Were all seven pattern passes run, with "none found" recorded where nothing appeared?
6. Did every candidate pass the six-column instrument screen, and are unresolved candidates excluded rather than caveated?
7. Does every pattern statement carry its scope inside the sentence rather than in a footnote?
8. Are the exceptions counted and reported with the pattern, in the same way a base is reported with a percentage?
9. Does every pattern have a stated breaking condition, and is there a record of where you looked for it?
10. Was any pattern confirmed out of sample, and where none was, is that stated and the confidence capped?
11. Is any pattern claimed as general when it was tested in one place?
12. Are findings grouped by structure in the data rather than by similarity of the words used to describe them?
13. Is the unpatterned-findings list present and complete?
14. Has any explanation crept into a pattern statement, which would make it 08.01's work presented as this skill's output?
15. Are the K5 judgement points marked at the point of judgement rather than collected at the end?

## 11. Common failure modes

| Failure | How to recognise it | How to prevent it |
|---|---|---|
| **Narrative structure imposed on noise** (the signature failure) | The pattern is satisfying, covers three or four findings, and nobody has counted how many it fails to cover | Base-rate statement, exception count, breaking condition. All three, in writing |
| **Retrospective pre-specification** | The "expected pattern" was written after the search and matches what was found exactly | Date the pre-specification list and keep it whether or not anything on it appeared |
| **Lexical grouping mistaken for structure** (the signature AI failure) | Findings are grouped because their sentences share vocabulary or abstraction level, not because they cover the same units | Group by base, unit and measure. Ask of each grouping which respondents or cases it spans |
| **The battery masquerading as a pattern** | Five items from one block all correlate, and correlate with each other as strongly as with anything outside it | The instrument screen, run before development, not after |
| **Base double-counting** | Five findings from the same 140 respondents are described as five pieces of evidence | Count independent findings after the screen. State the number |
| **Multiplicity blindness** | Two or three subgroup differences from sixty comparisons are presented as a pattern | Count the comparisons and state the expected chance yield |
| **The elastic pattern** | The claim has been reworded twice, each time more abstractly, and now covers everything | A claim that covers every finding explains none. Re-specify at the level where exceptions exist |
| **Exception suppression** | The pattern is stated and the findings it fails to cover are simply absent from the page | Require the covered-and-not-covered list in the output format |
| **Local reported as universal** | "Customers do X" from one market, one wave or one site | Write the scope inside the claim sentence at step 11 |
| **The pattern that arrives with its explanation** | The candidate is stated as "X happens because Y" | Explanation is 08.01. A pattern with a mechanism attached is much harder to reject and much less likely to be tested |
| **Absence overlooked** | The report contains only things that appeared | Run the absence pass explicitly and record what was expected |
| **Confirmation of the stakeholder's model** | Every pattern maps onto the framework in the brief | Treat a supplied structure as a hypothesis to test, per K4 §4.2, and report the findings it does not cover |
| **Pattern count driven by the template** | Exactly three themes, one per section | The number of patterns is a property of the evidence. Section E absorbs the rest |

## 12. AI guardrails

Skill-specific only. Universal prohibitions are inherited from **K4**; evidence levels follow **K2 §2**; confidence language follows **K3 §4**.

1. **Never group findings by similarity of their wording.** Two findings expressed in the same register are not thereby related. Grouping is by shared units, bases, measures or timing, and the basis for each grouping is stated.
2. **Never claim a pattern without stating the search space.** The number of comparisons available and the expected chance yield appear in the output.
3. **Never present an exploratory pattern as if it had been predicted.** Pre-specification is dated and written before the search or it does not exist.
4. **Never advance a candidate that has failed the instrument screen.** Adjacency, shared batteries, shared routing, shared bases, shared method and shared time windows are checked before development, and an unresolved candidate is reported as unresolved rather than caveated into the register.
5. **Never count findings that share a base as independent evidence.** State the independent count after the screen, not the raw finding count.
6. **Never report a pattern without its exception count.** The findings within scope that it fails to cover are listed, not summarised.
7. **Never state a pattern without a breaking condition and a record of where you looked for it.** A candidate with no refuting observation is a description, not a claim.
8. **Never drop the scope qualifier.** Where a pattern held in some sites, waves or segments and not others, the qualifier is inside the claim sentence in every downstream document.
9. **Never explain the pattern here.** Mechanism, motive and cause are 08.01. This skill establishes that structure exists and where its edges are.
10. **Never manufacture a pattern to give a report a structure.** Where nothing survives, the output is the unpatterned findings list and a statement that no structural claim is supported.

## 13. Best-practice principles

1. **The exception rate is part of the pattern.** A structural claim without a count of what it fails to cover has not been tested, whatever else has been done to it.
2. **Count the search space before you search.** It takes two minutes and it permanently changes how a three-finding coincidence looks.
3. **Absence is structure.** The subgroup difference that did not appear, the decay that did not happen, the association everyone assumed and nobody found. These are among the most valuable results a study produces and they are almost never in the deck.
4. **A pattern in one place is a finding about that place.** The qualifier is not modesty, it is the content of the claim.
5. **Check the instrument before you check the world.** Question order and battery structure produce more apparent coherence than any real attitude does, and the check is cheap.
6. **The pattern that arrives first is the one your reading order produced.** Run the seven passes anyway, including the ones that seem unpromising, because the pass you skip is where the sequence pattern was.
7. **Sequence patterns pay for themselves.** They are the ones that tell an organisation when to intervene, and they are invisible in a findings list because a list has no time axis until someone builds one.
8. **Thresholds hide inside averages.** Where a relationship is claimed, look at its shape before describing its direction.
9. **Independent measures beat repeated measures every time.** Two findings from different methods on the same phenomenon are worth more than six findings from one battery.
10. **Elegance is a warning.** A pattern that accounts for everything neatly has usually been stretched, and the stretching happens in the wording rather than in the data.
11. **Keep the rejected candidates.** They are the demonstration that the surviving pattern was tested rather than selected, and they are the first thing a sceptical stakeholder asks for.
12. **Stop before the explanation.** The moment a candidate acquires a "because", it becomes much harder to abandon, and abandoning candidates is the whole point of this stage.

## 14. Worked example

*Fictional scenario, used to demonstrate method. All organisations, participants, sites and figures below are invented.*

**INPUT.** An international water and sanitation NGO evaluates a household filter programme across nine districts, twelve months after installation. Fourteen findings are supplied from a household survey (n=1,180), thirty-one interviews with lapsed users, and programme records. The evaluation lead has proposed a theme: "community disengagement over time".

Selected findings: **F1** filter use is 71% at six months and 38% at twelve (survey, n=1,180). **F2** discontinuation is concentrated in the two months following the date the first cartridge replacement falls due (programme records, n=1,180). **F3** households receiving a second home visit report use more often, 62% versus 41%, untested. **F4** in three of nine districts, use at twelve months exceeds 60%; those three districts each have a local retailer stocking cartridges. **F5** 22 of 31 lapsed users described not knowing where to obtain a replacement [P01 to P31]. **F6** satisfaction with water taste is high and does not differ between continuing and lapsed users. **F7** reported handwashing improved over the same period and did not decay. **F8 to F12** five "community engagement" items, adjacent in one survey block, each associated with continued use.

**PROCESS.**

*Steps 1 to 4.* Fourteen findings, of which F8 to F12 are provisionally one block. The search space is written down: forty-five pairs among ten independent findings, and fifty-four district-level comparisons across nine districts and six measures, of which two to three are expected to look different by chance at a 95% threshold. Pre-specification, from the programme's own theory of change, was recorded before search: use should decay gradually, and decay should be slower where home visits were more frequent.

*Step 5, typed search.* Sequence produces the strongest candidate: F2 places the break at a specific programme event rather than across time. Absence produces the second: F7 shows a behaviour from the same programme, in the same households, that did not decay at all, which is precisely what the "disengagement" theme predicts should happen and it did not. Universal versus local produces the third: F4 splits the districts. Consistency across independent measures shows the same shape in programme records (F2), district structure (F4) and interviews (F5), which share no method.

*Step 6, the screen, and the judgement call.* F8 to F12 are the tempting candidate, because five findings pattern together and the resulting story is large. They fail the screen on three columns: the items were adjacent, shared a common stem and a common scale, and their correlations with each other are about as strong as their correlations with the outcome, which is the signature of a shared response style rather than a coherent attitude. The judgement was whether to report them as one weak finding or exclude them. Resolved: reported as a single instrument-unresolved finding, excluded from the pattern register, and named in the rejected-candidates section with what would resolve it (the same construct measured by a non-adjacent behavioural item in the next round). The independent finding count drops from fourteen to ten, and the base-rate statement is recalculated.

*Steps 7 to 10.* Claim: *programme outcomes hold until the first consumable resupply and break at it, except where resupply is locally available.* Scope: nine districts, twelve months, filter households only. Exceptions counted: one of nine districts has a stocking retailer and use still fell below 50%, and F3 is only partly accounted for. Breaking condition stated: if discontinuation were spread evenly across the twelve months, or were the same in retailer and non-retailer districts, the claim fails. Falsification attempted against month-by-month programme records that were not used to build the pattern: discontinuation in retailer districts is flatter and later, and the claim survives. Holdout: the three districts added to the programme late were held out and show the same timing.

*Steps 11 to 13.* Universal within the districts tested, which is nine, and that number is written into the claim. Status: exploratory for the resupply timing, pre-specified for nothing (the programme's own predicted gradual decay was refuted, which is recorded). Confidence: **moderate**, capped by the untested status of F3 and by the single deviant district, with the sequence evidence and the independent-measure convergence supporting it.

**OUTPUT.** One pattern with scope, one exception district named, the refuted gradual-decay expectation recorded, F8 to F12 in the rejected-candidates section as instrument-unresolved, four unpatterned findings listed including F6, and the base-rate statement showing ten independent findings and fifty-four available district comparisons. Handed to 08.01 for explanation, with a K5 §2.1 review point on whether the deviant district is material or a local supply failure the country team already knows about.

## 15. Advanced usage

**When the pattern is the absence.** Where a study was designed to detect a difference and found none across several measures, the absence itself can be the primary structural claim: a distinction the organisation is spending money on does not exist for the people it serves. Treat it exactly as a pattern, with scope, exceptions and a breaking condition, and check the power of the design before claiming it, since an absence in an underpowered study is not evidence of absence.

**Multi-wave and repository work.** A pattern that has held across three waves is a different object from one seen once, and it should be re-tested rather than re-quoted, per K2 §7. Store each pattern with its date, its covered findings, its exception rate and its scope. Patterns decay silently: the mechanism that produced them changes and the sentence stays in the deck.

**Competing structures over the same findings.** Where two candidate patterns both survive, do not choose. Report both, state which findings each covers and which it does not, and name the analysis or the measure that would separate them. Two honestly reported structures are more useful than one selected structure, and the separation test is usually cheap.

**Where the findings arrive without provenance.** This is common when compiling from previous decks. Generate candidates, mark every one as instrument-unscreened, cap confidence at low, and state that the screen could not be run. Do not treat unscreened findings as independent, and say what documentation would allow the screen.

**Running this on someone else's structure.** Where a theme has been proposed by a stakeholder, an earlier report or an AI draft, invert the method: take the proposed structure as a single candidate, run the instrument screen on it, list every finding in scope it fails to cover, write its breaking condition, and go and look. Most proposed themes fail at the exception count, and producing that count is a fast and defensible way to redirect a report before it is built.

## 16. Skill chain

**Recommended previous skills:**
- **05.02 Cross-Tabulation and Segmentation Analysis** and **05.03 Statistical Significance Testing.** Hand over subgroup structure and tested differences, so that segment-specific and threshold patterns start from established variation rather than from eyeballed tables.
- **07.01 Thematic Analysis.** Hands over themes with prevalence, counter-evidence and preserved contradictions, which are findings for this skill and not yet patterns.
- **07.06 Qualitative and Quantitative Integration.** Hands over which streams bear on which questions, which is what makes the independent-measure pass possible.
- **02.01 Questionnaire Design** documentation, or the equivalent instrument, without which the screen at step 6 cannot run.

**Recommended next skills:**
- **08.01 Finding to Insight Development.** Takes the bounded pattern, with its scope and exceptions, and works out why it is true. This is the primary handover.
- **08.05 Insight Prioritisation and Sizing.** Takes patterns and insights and decides which are material.
- **12.01 Research Report Architecture.** Takes the pattern register as the candidate spine for the report, with the unpatterned findings retained rather than discarded.

**Runs well alongside:**
- **13.04 Bias Detection**, particularly where a structure was proposed before the analysis.
- **13.03 AI Output Verification.** Run against any AI-drafted theme set; the exception count and the instrument screen are the two checks that catch most of them.
- **K5**, at the judgement points this skill mandates: whether an exception is material, and how a pattern reads in a context you cannot access.

---
A Yazi Supplied Skill and resource.
