---
name: segment-comparison
description: >
  Compares groups honestly, whether they came from a segmentation, a screener or a
  natural split: base rules for every group, testing rather than eyeballing, the
  multiple-comparison problem, index scores and their misreading, the composition
  confound, and a reporting format that shows what is shared as well as what
  differs. Use for "compare the segments", "how do these groups differ", "which
  segment over-indexes", "cross-tab by segment", "is this difference real", "why
  does this group look so different", "compare users and non-users", "compare
  markets".
category: 09 Segmentation and Audience Understanding
ref: 09.04
tier: 1
inherits: [K2, K3, K4, K5]
---

# Segment Comparison

## 1. One-line description
Turns a comparison table into a defensible account of how groups differ, by requiring an adequate base for every group compared, testing differences rather than reading them, controlling the multiple-comparison problem that a segment-by-measure grid creates, checking whether an apparent segment difference is really a composition difference, separating a difference in level from a difference in structure, and reporting what the groups share alongside what divides them.

## 2. What this skill is used for

**The research problem it solves.** A segment-by-measure table is the most over-read artefact in research. It is generated automatically, it contains hundreds of cells, and every cell invites a comparison that nobody has tested. Four failures follow, and they compound. **Eyeballing:** differences are declared by looking, so a five-point gap on bases of 90 and 400 becomes a strategic distinction. **Multiplicity:** six segments compared on forty measures produces hundreds of implicit tests, and at a 95% threshold a substantial number of "significant" results are expected by chance alone, with no way to tell which. **The index trap:** an index of 180 against the total looks decisive and can rest on eleven respondents, and readers consistently treat a high index as a large group rather than as a ratio. **The composition confound:** two segments differ on digital channel use because one skews fifteen years younger, not because of anything the segmentation defines, and the report attributes it to the segment anyway. On top of these sits a quieter problem: comparison tables systematically overstate difference, because they only contain the measures on which groups were compared and only the rows that moved. Segments in most studies have far more in common than any comparison output suggests, and a business that reads only the differences ends up treating one market as five.

**Where it sits.** Analysis. It runs on any grouped data, whether the groups came from a segmentation, a screener, a customer database, a market split or a behavioural band. It sits between analysis and synthesis, and its output feeds insight development and reporting.

**Typical use cases.**
- Profiling segments from **09.01** against measures used and not used to build them.
- Comparing users and non-users, customers and prospects, or lapsed and active groups from a screener.
- Comparing markets, regions or channels in a multi-country or multi-site study.
- Comparing behavioural bands such as heavy, medium and light buyers.
- Auditing a segment comparison somebody else produced, particularly one built on index scores.
- Establishing whether an apparently distinctive group is distinctive, or simply younger, wealthier or more urban.

**Who uses it.** Quantitative analysts producing segment profiles and cross-tabulations; research directors checking what a comparison table actually supports before it becomes a strategy; client-side insight managers receiving segment tables from an agency; and anyone reviewing AI-generated segment comparisons, where every visible gap tends to be reported as a difference and the composition question is never asked.

## 3. When to use it

- You have groups and a set of measures, and you need to know which differences are real and which are worth acting on.
- A segmentation has been built and needs profiling, including on variables used to construct it.
- A comparison table has produced a long list of differences and somebody has to decide which ones to report.
- One group looks dramatically different and you suspect its demographic composition explains it.
- Index scores are in circulation and are being read as size rather than as ratio.
- Bases differ substantially between the groups being compared.
- Markets are being compared and the differences may be cultural, compositional or artefactual rather than substantive.
- You are auditing a comparison that declared differences without a stated test.

## 4. When NOT to use it

- **The groups do not yet exist and need to be built.** Constructing segments from data is **09.01 Audience Segmentation**. This skill compares groups that are already defined, by whatever means. Running a comparison to discover groups is how arbitrary splits acquire the appearance of a segmentation.
- **A group's base is too small to compare.** Below n=30 no percentage is reported at all, per K4 §7, and comparison on counts alone rarely supports a claim. Between 30 and 99 the comparison is directional only, and a difference will need to be large to be reliable. Report the base, say the comparison cannot be made at the strength requested, and name the sample that would allow it. A comparison the base cannot support is not improved by testing it.
- **The measures are not comparable across the groups.** Different questionnaire versions, different scale lengths, different fieldwork periods or different modes between the groups make a comparison arithmetic rather than meaningful. Multi-market work is the common case: an agreement scale used in two cultures with different response styles produces a difference that is partly a measurement artefact. Establish comparability first, per **02.07 Scale and Measurement Selection**, and where it cannot be established, say so rather than compensating with a caveat at the back.
- **The question is why the groups differ.** This skill establishes that a difference is real, is not compositional, and is or is not material. Explaining it is **08.01 Finding to Insight Development**, and a comparison output that arrives already explaining itself has crossed the K2 interpretation boundary without a signal word.
- **The comparison is being run to find something, anything, that differentiates a group.** Scanning hundreds of cells for whatever reaches a threshold, then reporting those cells, is a guaranteed producer of false findings under multiplicity and is cherry-picking under K4 §4.2. If the analysis is genuinely exploratory, say so, adjust for multiplicity, and label the results as hypotheses per K3 §4.3.
- **The difference is real and the business cannot act on it.** A two-point gap that survives testing on large bases is a statistical result and frequently not a finding worth anyone's attention. Reporting every real difference is a different failure from reporting unreal ones, and it produces a document nobody can prioritise. Materiality is a K5 §2.1 judgement and must be marked rather than assumed.
- **The groups are defined by the outcome being compared.** Comparing heavy and light buyers on how much they buy, or comparing a segment on the battery used to construct it, is circular. It is sometimes still worth showing, as a description of what defines the groups, and it must be labelled circular in the output. It is never evidence that the groups are meaningful.
- **The data has not been prepared or the bases are not established.** Comparison inherits every base error in the underlying tables. Run **05.01 Descriptive Analysis** first, so that every figure has a base description in words before any two of them are placed side by side.

## 5. Required inputs

**Required.** Without these the skill cannot run. If absent, ask.

- **The group definitions and how the groups were created.** Segmentation output, screener question, database flag, behavioural band or geography. This determines what the comparison can claim and which measures are circular.
- **The base size and base description for every group**, on every measure. Bases vary by measure because of routing, and a comparison run on the wrong base is worse than no comparison.
- **The measures, with their wording and scales**, so that comparability can be assessed and so that scale direction cannot be misread.
- **For segmentation-derived groups: the basis variable list**, without which circular comparisons cannot be identified.
- **The decision the comparison informs**, which sets the materiality threshold. Without it, every real difference has equal claim on the reader's attention.

**Optional, and what each one adds.**

- **Demographic and structural profile of each group** (age, market, tenure, channel, life stage, firm size): allows the composition confound to be checked, which is not optional in practice but is frequently unavailable in a supplied table.
- **Weighting variables and the effective base per group:** determines whether the effective base rather than the raw n governs the small-base rules, per **04.05**.
- **A pre-registered analysis plan or hypothesis list:** converts a fishing expedition into a set of planned comparisons, which changes the multiplicity treatment fundamentally and is the single strongest thing available here.
- **Prior wave comparisons:** allow a difference to be checked for persistence, which is stronger evidence than any single-wave test.
- **Total-sample figures for every measure:** required for index calculation and for showing what the groups share.
- **Qualitative evidence from the same groups:** allows a difference in level to be distinguished from a difference in kind, which numbers alone often cannot do.

## 6. Questions to ask before starting

1. **How were these groups created, and which measures were used to create them?** Determines which comparisons are circular and must be labelled. *Default if unanswered:* treat every measure that resembles the group definition as potentially circular and flag it.
2. **What is the base of every group on every measure to be compared?** Determines what can be claimed and which comparisons are impossible. *Default:* produce a base grid first and mark every cell below 100 and below 30.
3. **Were these comparisons planned, or is this exploratory?** Determines the multiplicity treatment and the confidence language. *Default:* treat as exploratory, adjust or disclose, and label the results as hypotheses.
4. **What decision does the comparison inform, and how big a difference would change it?** Sets the materiality threshold before the numbers are seen, which is the only point at which it can be set honestly. *Default:* report differences with their size and flag materiality for researcher judgement per K5 §2.1.
5. **How do the groups differ in composition?** Determines whether an apparent segment effect is an age, market or tenure effect. *Default:* profile the groups on demographics and structure before interpreting any behavioural or attitudinal difference, and state that composition was not controlled if it could not be.
6. **Are the measures comparable across the groups?** Determines whether any comparison is legitimate. *Default:* check mode, wording, scale and fieldwork period, and disclose any difference.
7. **Is this a level comparison or a structure comparison?** Determines the analysis. *Default:* look at both, because groups frequently agree on rank order while differing in level, and that distinction usually changes the implication.

## 7. Step-by-step methodology

**The position this method takes.** A comparison table shows arithmetic differences. Turning those into claims requires four separate things to be true, and each has to be established rather than assumed: the difference is reliable, it is not an artefact of who is in each group, it is not circular, and it is large enough to matter. Most comparison outputs establish none of the four. The method below runs them in order, and the order matters: testing an unreliable base wastes effort, and interpreting a compositional difference is worse than not interpreting it.

**1. Build the base grid before comparing anything.** One row per measure, one column per group, each cell containing the base for that group on that measure. Bases move with routing, so the base of Segment A on Q7 is not the base of Segment A on Q4. Mark every cell below 100 and every cell below 30. **A comparison is only as strong as its weakest group**, so a measure where one group has n=41 is a directional comparison for every group on that row, not just for the small one. Where a group is below 30, that group is not compared on that measure at all. *Correct result:* a base grid that determines, before any analysis, which comparisons are available, which are directional and which are impossible.

**2. Decide what makes a comparison meaningful rather than merely available, and write the list.** A segment-by-measure grid offers thousands of comparisons. Meaningful ones share three properties: the measure bears on the decision the comparison informs; a difference on it would change what someone does; and there is a reason to expect the groups might differ that predates seeing the data. Write the planned list before looking. Everything else is exploratory, and is reported as exploratory. This step is what separates a profile from a fishing expedition, and it costs nothing except the discipline of doing it in the right order. *Correct result:* a planned comparison list, dated before the analysis, plus an explicitly separate exploratory set.

**3. Set the materiality threshold now, before the numbers are visible.** Ask what size of difference would change the decision. In many commercial contexts the answer is far larger than the differences a table produces: a five-point gap in claimed preference will not change a range decision, whereas a fifteen-point gap in penetration will. Setting the threshold after seeing the data means setting it to whatever the data produced. *Correct result:* a stated threshold per decision, against which differences are later judged.

**4. Identify and label the circular comparisons.** Any comparison of segments on the variables used to build them is circular: the segments differ on those variables because they were constructed to. This is worth showing, because it describes what defines the groups, and it is not evidence of anything. Label it explicitly in the output: "these are the variables used to construct the segments; differences here are by construction". The same applies to screener-defined groups compared on the screener criterion and to behavioural bands compared on the behaviour that defines them. **The uncomfortable version of this rule is that a segmentation whose only large differences are circular has not been shown to be meaningful**, which is a finding about the segmentation, per **09.01**. *Correct result:* every comparison marked circular or independent, before any is interpreted.

**5. Test, per 05.02, rather than reading the table.** Ordinary comparative language for observed differences is permitted with both bases shown, per K4 §3.1, but any claim that a difference is real requires a test, the test named, and the threshold stated. Test the planned comparisons. Report the test used, the result and the threshold in the same place as the difference. Do not use "significantly", "notably", "clearly" or "markedly" for anything untested, and do not let those words enter a summary where the test result has been left behind. *Correct result:* every claimed difference carrying its test, its threshold and both bases.

**6. Handle multiplicity explicitly, and choose a treatment rather than ignoring it.** Six segments compared pairwise on forty measures generates six hundred comparisons. At a 95% threshold, roughly one in twenty of the null cases will pass by chance, so a table of that size can be expected to produce dozens of spurious "significant" results that are indistinguishable from real ones by inspection. Three legitimate responses, and one illegitimate one. **Restrict:** test only the planned comparisons from step 2, which is the strongest response and the cheapest. **Adjust:** apply a correction across the family of tests, accepting that it reduces sensitivity and will lose some real differences. **Disclose:** report the number of comparisons made and the number expected by chance, and label the whole exploratory set as hypothesis-generating per K3 §4.3. The illegitimate response is to run everything at 95% and report the winners as findings. State which treatment you used, once, prominently. *Correct result:* a stated multiplicity treatment, and an exploratory set that is labelled rather than presented as established.

**7. Check the composition confound before interpreting any difference.** This is the step that most often changes the conclusion. Two groups can differ on a measure because the groups differ on something else that also relates to that measure: age, market, tenure, life stage, urbanity, firm size, product holding. A segment that is fifteen years younger will differ on digital channel use, media consumption, price sensitivity and dozens of other measures, and none of those differences is caused by whatever the segmentation was built on. For every difference you intend to report, ask what else differs between these groups, and check the difference within levels of that variable: does the gap survive when you compare like with like? Where it survives, the comparison is about the group. Where it disappears, **the honest finding is that the two groups differ in composition, and the measure difference follows from that**, which is a real and useful finding stated correctly. Where the subgroup bases are too small to check, say that composition was not controlled and cap confidence at moderate per K3 §3.5. *Correct result:* every reported difference either checked against the main compositional candidates, or explicitly flagged as unchecked.

**8. Separate a difference in level from a difference in structure.** Two groups can give different absolute answers while agreeing entirely on rank order and relative importance, which means they want the same things and one group is simply more positive, more engaged or more scale-generous. That is a level difference, and it usually implies a different intensity of the same strategy. A structure difference is when the ordering itself changes: group A ranks price first and convenience fourth, group B the reverse. That implies genuinely different strategies. Level differences are far more common and are routinely reported as if they were structural, which produces differentiated plans for groups that want the same thing. Check by comparing rank orders and relative gaps, not only absolute values, and be alert to scale-use differences, which manufacture level differences that are pure measurement, particularly across markets. *Correct result:* every set of differences classified as level or structure, with the rank-order comparison shown.

**9. Use index scores carefully, and never alone.** An index expresses a group's figure as a ratio to the total, and it is useful for spotting concentration in a large table. It also misleads in three specific ways. **It hides the absolute value:** an index of 200 on a measure where the total is 3% means 6%, which is nearly nobody. **It hides the base:** a spectacular index can rest on a handful of respondents, and small bases produce extreme indices as an arithmetic consequence of instability. **It hides the size of the group:** a strongly over-indexing segment of 4% of the population contains fewer people than a slightly under-indexing segment of 40%. Rule: never present an index without the underlying percentage and the base in the same view, never index on a base below 100, and never rank opportunities by index alone, because that systematically selects small groups on rare measures. *Correct result:* an index table in which every index is accompanied by its percentage, its base and the group's size.

**10. Report what the groups share as prominently as what divides them.** After the differences are established, run the reverse analysis: on which decision-relevant measures do the groups not differ? In most studies this list is far longer than the difference list, and it is more actionable than it looks, because it tells the business which parts of the proposition, the product and the service can be common. A comparison output containing only differences produces a management team that believes it has five markets, and a cost base to match. *Correct result:* a shared-ground section, built from the same measure set, appearing in the main output and not an appendix.

**11. Apply the materiality threshold from step 3, and separate real from important.** Sort the surviving differences into three groups: real and above the threshold, real and below it, and untested or unreliable. Report the first as findings, the second as a short list explicitly marked as real but immaterial, and the third as hypotheses or not at all. **A difference that is statistically real and practically irrelevant is a common output of large samples and is not a finding.** Marking it as such prevents the reader from having to make that judgement themselves on forty rows. *Correct result:* a prioritised difference list with materiality stated, and a K5 §2.1 marker where materiality is genuinely a business judgement.

**12. Write the comparison record.** For each reported difference: the measure, both figures, both bases, the test and threshold, the composition check and its result, whether it is circular, whether it is level or structure, and its materiality. This is the artefact a research director inspects, and it is what allows a difference to be defended six months later when somebody asks whether the segments really differ. *Correct result:* a record from which any reported difference can be reconstructed without reopening the dataset.

## 8. Analytical framework

The five gates a difference passes before it is reported as a finding:

    Base → Independence → Reliability → Composition → Materiality

| Gate | Question | Fails if |
|---|---|---|
| **Base** | Do both groups have an adequate base on this measure? | Either group is below 30. Directional only between 30 and 99 |
| **Independence** | Was this measure used to define the groups? | Circular. Show it if useful, label it, claim nothing from it |
| **Reliability** | Has the difference been tested, with the test and threshold stated, and multiplicity handled? | Untested, or one of hundreds of unadjusted comparisons |
| **Composition** | Does the difference survive comparing like with like on the main structural variables? | The gap disappears within age, market or tenure bands. Report the composition finding instead |
| **Materiality** | Is the difference large enough to change the decision? | Real but below the threshold. Report as real and immaterial |

**Level versus structure.**

|  | Level difference | Structure difference |
|---|---|---|
| What it looks like | Same rank order, different absolute values | Different rank order or different relative weights |
| Common cause | Engagement, familiarity, scale use, sample composition | Genuinely different priorities |
| Frequency | Common | Uncommon |
| Implication | Same strategy, different intensity | Different strategies |
| Failure | Reported as structural, producing differentiated plans for groups that want the same thing | Missed because only absolute values were compared |

**Against the K2 chain**, this skill produces Findings and stops. "Segment B over-indexes on branch use, and the difference survives an age control" is a finding. "Segment B values human contact" is an interpretation, requires the signal language of K2 §3.2, and belongs to **08.01**.

## 9. Output format

**1. Comparison design statement.** How the groups were defined, which comparisons were planned and which are exploratory, the multiplicity treatment used, the materiality threshold and where it came from, and the tests and thresholds applied.

**2. Base grid.**

| Measure | Group A n | Group B n | Group C n | Total n | Lowest base flag |
|---|---|---|---|---|---|

**3. Circular comparison block.** The measures used to construct the groups, shown together and labelled as differences by construction.

**4. Difference table**, for independent measures.

| Measure | A % (n) | B % (n) | Total % | Difference | Test and threshold | Composition check | Level or structure | Material |
|---|---|---|---|---|---|---|---|---|

**5. Index table**, where used: index, underlying percentage, base and group size in the same row. No index reported on a base below 100.

**6. Shared ground.** Decision-relevant measures on which the groups do not differ, with the same evidence standard applied.

**7. Real but immaterial.** Differences that passed testing and fell below the materiality threshold, listed so the reader does not have to rediscover them.

**8. Exploratory findings.** Labelled as hypotheses per K3 §4.3, with the number of comparisons made and the number expected by chance.

**9. What could not be compared.** Measures where a base was too small, groups that could not be compared, and comparisons blocked by non-comparable measurement.

**When the evidence is thin.** The comparison shrinks. A group of 24 does not appear in the difference table; it appears in section 9 with its base. A measure where composition could not be controlled is reported with that stated and its confidence capped. An exploratory set stays labelled exploratory even where a result is striking, because striking results are exactly what multiplicity produces. Per K4 §1, a column in a comparison table is not a reason to fill it.

## 10. Quality checks

Run before anything is presented. These sit on top of K4 §8.

1. Does every comparison show the base for every group involved, in the same view as the figures?
2. Is any group below n=30 being compared on percentages, or any group between 30 and 99 being compared without a directional flag?
3. Are the comparisons that use the segmentation's own basis variables labelled circular?
4. Does every claimed difference carry a named test and a stated threshold?
5. Is the multiplicity treatment stated, and does the output distinguish planned from exploratory comparisons?
6. Has every reported difference been checked against the main compositional variables, or explicitly flagged as unchecked?
7. Where a difference disappeared under a composition control, is the composition finding reported rather than the original difference quietly dropped?
8. Is every set of differences classified as level or structure, with rank orders compared?
9. Does every index appear with its underlying percentage, its base and the group's size?
10. Is any index reported on a base below 100, or any opportunity ranked by index alone?
11. Is the shared-ground section present, in the main output rather than an appendix?
12. Are real but immaterial differences listed as such rather than mixed in with material ones?
13. Was the materiality threshold set before the figures were seen?
14. Has any evaluative word ("significantly", "notably", "markedly") been used without a test behind it?
15. Are the measures comparable across the groups in wording, scale, mode and fieldwork period?
16. Does any interpretation of why the groups differ appear in this output rather than being handed to **08.01**?

## 11. Common failure modes

| Failure | How to recognise it | How to prevent it |
|---|---|---|
| **Eyeballing** | Differences declared from a table with no test anywhere in the document | Test the planned comparisons; ordinary comparative language only, with both bases, for the rest |
| **Multiplicity ignored** | A long list of "significant" differences from a large grid, with no mention of how many tests were run | State the treatment: restrict, adjust or disclose. Label the exploratory set |
| **The composition confound** (the costliest failure) | A segment differs on a dozen digital measures and is fifteen years younger | Check within levels of the structural variable before interpreting anything |
| **Unequal bases ignored** | A group of 62 compared to one of 900 as though the comparison were symmetric | Base grid first; the weakest group governs the row |
| **Index read as size** | "Segment D over-indexes at 210, so target them", where the segment is 4% of the market | Index, percentage, base and group size in the same row, always |
| **High index on a low base** | Spectacular indices concentrated in the smallest group | Never index below n=100; small bases produce extreme ratios mechanically |
| **Circular comparison as validation** | "The segments differ hugely on these attitudes", where those attitudes built the segments | Label circular comparisons; validation requires independent measures, per **09.01** |
| **Level read as structure** | Differentiated strategies for groups with identical rank orders | Compare rank orders and relative gaps, not only absolute values |
| **Scale-use difference read as attitude** | One market lower on every item in a battery | Check whether the whole battery shifted; standardise within respondent where appropriate |
| **Real and irrelevant reported as a finding** | A two-point tested difference on a base of 4,000 leading a slide | Materiality threshold set before the figures were seen |
| **Difference-only reporting** | A twenty-row table of differences and no statement of what is shared | Shared-ground section, built from the same measure set |
| **AI: reporting every visible gap** | Every non-identical pair of numbers described as a difference | Gates in order: base, independence, reliability, composition, materiality |
| **AI: explaining while comparing** | The comparison table arrives with reasons attached | This skill stops at Finding. Explanation is **08.01**, and needs a signal word |

## 12. AI guardrails

Skill-specific only. K4 applies in full and is not repeated here.

1. **Never describe a difference as real, significant or notable without a test, its threshold, and both bases in the same place.**
2. **Never report a set of comparisons without stating how many were made and how multiplicity was handled.** Reporting only the ones that passed is cherry-picking under K4 §4.2.
3. **Never interpret a difference between groups before checking whether the groups differ in composition** on the obvious structural variables. Where the check was impossible, say so and cap confidence.
4. **Never present an index without its underlying percentage, its base and the group's size**, and never index a base below 100.
5. **Never rank opportunities or priorities by index score**, which systematically selects small groups on rare measures.
6. **Never present a comparison on the variables used to construct the groups as evidence that the groups are meaningful.** Label it as circular in the output.
7. **Never compare a group whose base is below 30 on percentages**, and never compare across groups without flagging the weakest base on that row.
8. **Never describe a level difference as a difference in priorities.** Compare rank orders before making any claim about what a group values.
9. **Never omit the shared-ground section.** A difference-only output systematically overstates how different the groups are.
10. **Never report an exploratory result in the language of an established finding**, however striking, per K3 §4.3.
11. **Never compare measures that were not asked in the same way**, in the same mode and in the same period, without disclosing the difference next to the comparison.

## 13. Best-practice principles

- **A comparison table is a list of candidates, not a list of findings.** Everything in it is arithmetic until it has passed the five gates.
- **The weakest base governs the row.** A comparison is only as good as its smallest group, and a strong base on one side does not compensate.
- **Ask what else is different about these people, every time.** The composition confound is the most common cause of a segment difference that dissolves under scrutiny, and it is invisible in the table that produced it.
- **A difference that disappears under a control is still a finding.** "These groups differ on channel because one is much younger" is more useful than a wrong attribution, and it points at a different intervention.
- **Most differences are level differences.** Groups usually want the same things in the same order at different intensities, and treating that as a structural difference produces expensive, unnecessary differentiation.
- **An index is a ratio, and readers hear it as a size.** Show the percentage, the base and the group size, or the index will be misread every time.
- **Testing everything is not rigour.** Six hundred tests at 95% produce dozens of false positives that look exactly like findings. Restricting the comparison set is stronger than any correction.
- **Set the materiality threshold before you see the numbers.** It is the only moment at which the threshold can be honest.
- **Report the sameness.** It is usually the larger and more actionable half of the answer, and it is the half a comparison table structurally cannot show you.
- **Circularity is not a small technical point.** A segmentation whose differences are all circular has not been shown to describe anything, and the comparison output is where that becomes visible.
- **Persistence beats significance.** A difference that appears in two independent waves is stronger evidence than one that passes a test once.
- **Comparison stops at what differs.** Why it differs is a separate skill with separate standards, and merging the two is how an artefact becomes an explanation.

## 14. Worked example

*Fictional scenario, used to demonstrate method. The organisation, figures and findings below are invented.*

**INPUT.** A mutual savings provider has a four-segment needs-based segmentation and wants a profile comparison across 38 measures to inform a channel investment decision. Sample 2,050. Segment bases: A n=720, B n=610, C n=480, D n=240. The supplied agency table declares 46 "significant" differences.

**PROCESS.**

*Steps 1 to 3.* The base grid shows that on the branch-experience measures, which were routed to branch users only, Segment D falls to n=71 and on one measure to n=26. Those rows become directional or unavailable. The planned comparison list is written against the channel decision: nine measures, chosen before looking. The materiality threshold is set with the client at ten percentage points on penetration measures, because anything smaller will not change the channel footprint.

*Step 4, circularity.* Eleven of the 46 declared differences are on the needs battery used to construct the segments. They are moved into a labelled circular block. They describe the segments; they validate nothing.

*Step 6, multiplicity.* The remaining 35 differences come from 38 measures across six pairwise segment comparisons, so 228 tests at 95%, with roughly eleven false positives expected among the null cases. The treatment adopted is to restrict: the nine planned comparisons are tested and reported as findings, and the rest are reported as exploratory with the count of tests and the expected chance rate stated.

*Step 7, the composition check, and the judgement call.* The headline claim in the supplied table is that Segment D is a digital segment: 71% use the mobile channel versus 44% total, index 161. Segment D's mean age is 34 against 52 overall. Comparing within age bands, the gap collapses: among under-40s, Segment D is at 78% and the rest of the sample at 74%, a four-point difference that does not survive testing on those bases. **The difference is an age difference wearing a segment's clothes.** The judgement call was whether to report this at all, since it removes the most quotable line in the deck. It is reported, as a composition finding: the segment is young, and the channel behaviour follows from age rather than from needs, which means the channel decision cannot be made by targeting the segment and must be made by age or by observed channel behaviour directly.

*Step 8, level versus structure.* On the seven-item importance battery, Segments A and B differ on absolute scores by six to nine points on every item. Rank orders are identical, and the relative gaps between items are within a point. This is a level difference and almost certainly reflects engagement and scale use, not different priorities. The supplied table had presented it as two different sets of needs.

*Steps 9 to 11.* Indices are re-presented with the underlying percentage, base and segment size. Two of the largest indices in the supplied table sat on bases of 58 and 71 and are withdrawn. The shared-ground section identifies 21 of the 38 measures on which no segment differs materially, including every measure of security expectation and service recovery, which becomes the basis for a common service standard rather than four differentiated ones. Six differences survive all five gates.

**OUTPUT.** A comparison design statement; a base grid; a labelled circular block; six findings with tests, thresholds, composition checks and materiality; a composition finding replacing the digital-segment claim; a level-versus-structure classification for the importance battery; a corrected index table; a 21-measure shared-ground section; an exploratory list labelled as hypotheses with the test count stated; and a **researcher decision required** marker on whether the channel investment should be planned by age or by segment, given that the two are confounded in this sample, per K5 §2.1.

## 15. Advanced usage

**Comparing many groups without a segmentation.** The same five gates apply to market comparisons, site comparisons, cohort comparisons and wave-by-segment comparisons. The multiplicity problem grows with each additional dimension, so a segment-by-market-by-wave table is a near-guaranteed false-positive generator unless the comparison set is restricted in advance.

**Controlling composition properly.** Where subgroup bases allow, comparing within levels of the confounder is transparent and easy to explain. Where they do not, a multivariable model estimating the segment effect with the structural variables included is the stronger tool, with the caution that it produces adjusted associations rather than causal effects, per **05.06**. Report both the raw and the adjusted difference, because the change between them is itself the finding.

**Comparing structures rather than levels formally.** Where the question is whether groups have different priority structures, compare the rank orders and the relative importance patterns directly rather than comparing item scores one at a time. This is more robust to scale-use differences and is usually what the decision actually needs.

**Multi-market comparison and response style.** Differences in how cultures use rating scales are large, systematic, and easily mistaken for differences in attitude. Where a whole battery shifts in one direction for one market, suspect response style first. Standardising within respondent removes most of it and removes some real variance with it, so report both versions and route the cultural reading to a human per K5 §2.2.

**Auditing a supplied comparison.** Run the five gates against someone else's table. In most supplied comparisons the base grid alone removes a substantial share of the claimed differences, the circularity check removes more, and the composition check changes the interpretation of several of the survivors. Produce a marked-up version showing which differences survive each gate, which is more persuasive than a rebuild and gives the client a reusable standard.

## 16. Skill chain

**Recommended previous skills:**
- **05.01 Descriptive Analysis.** Hands over correctly based figures with base descriptions in words, without which no comparison is trustworthy.
- **05.03 Cross-Tabulation.** Produces the grouped tables this skill interrogates, with the base register attached.
- **09.01 Audience Segmentation.** Hands over the group definitions and, critically, the basis variable list that identifies which comparisons are circular.
- **04.05 Weighting and Base Management.** Hands over effective bases per group, which govern the small-base rules where data is weighted.

**Recommended next skills:**
- **05.02 Statistical Testing.** Runs the tests this skill requires, with the multiplicity treatment agreed here.
- **05.06 Correlation, Regression and Causal Claim Control.** Handles composition control where subgroup bases will not support a within-level comparison.
- **08.01 Finding to Insight Development.** Takes established, composition-checked differences and explains them, which this skill deliberately does not do.
- **09.02 Persona Development.** Takes the differences that survived all five gates, which are the only ones a persona should assert.

**Runs well alongside:**
- **13.03 AI Output Verification**, run against any supplied or AI-generated comparison table, with the base grid and composition check as its focus.
- **13.04 Bias Detection**, where the comparison set was chosen after the data was seen.
- **K3 §3.2 and §4.3**, for the base thresholds and the hypothesis language the exploratory set requires.

---
A Yazi Supplied Skill and resource.
