---
name: behavioural-profiling
description: >
  Builds an evidence-based picture of what people actually do, distinguishing
  observed behaviour from claimed behaviour, working in repertoires and occasions
  rather than single-brand loyalty, and correcting the standing error of profiling
  the heavy user as though they were the market. Use for "what do our customers
  actually do", "usage and attitude analysis", "purchase behaviour profiling",
  "how often do they really buy", "who are our heavy users", "behavioural
  segments", "share of category", "brand repertoire", "occasion analysis", "claimed
  versus actual behaviour".
category: 09 Segmentation and Audience Understanding
ref: 09.03
tier: 1
inherits: [K2, K3, K4, K5]
---

# Behavioural Profiling

## 1. One-line description
Establishes what an audience actually does, by ranking the available behavioural evidence on reliability, treating claimed frequency as a weak estimate rather than a measurement, describing category behaviour as repertoires and occasions rather than as loyalty to one brand, and reporting the light-buyer majority that heavy-user profiling routinely conceals.

## 2. What this skill is used for

**The research problem it solves.** Most descriptions of customer behaviour are descriptions of what customers said about their behaviour, produced by asking a question of the form "how often do you usually...". People answer such questions by constructing an estimate from a general sense of themselves, and the estimate is systematically wrong in known directions: regular behaviours are over-reported, irregular ones are under-reported, socially approved ones are inflated, and recent or vivid instances are pulled forward in time. The resulting profile is not noise around a true value, it is biased, and the bias is largest exactly where the business is most interested. Three further errors compound it. Behaviour is described as loyalty to one brand, when in most repeat-purchase categories people hold a repertoire and allocate share across it. Behaviour is attributed to the person, when it frequently belongs to the occasion, so the same individual appears in two contradictory rows of the same table. And the heavy user is profiled as though the heavy user were the market, when in most measured categories the majority of buyers buy rarely and collectively account for a large share of volume. This skill sorts the evidence by reliability, states plainly which parts are observed and which are claimed, and builds the profile at the unit of analysis where the variance actually lives.

**Where it sits.** Analysis. It sits between data preparation and the sense-making skills, and it supplies the behavioural basis variables that segmentation and needs work depend on. It also runs on its own where the deliverable is a usage and behaviour picture.

**Typical use cases.**
- Building a usage and behaviour profile of a customer base or a category population.
- Reconciling survey-claimed behaviour with transactional or telemetry records that disagree with it.
- Establishing category repertoire and share of category rather than single-brand loyalty measures.
- Identifying the occasions in which a category is used, and how they differ from each other.
- Deriving behavioural groups from actual behaviour, as an input to segmentation.
- Testing whether a change in reported behaviour reflects changed people or changed circumstances.
- Correcting a strategy built on heavy-user characteristics that the light-buyer majority does not share.

**Who uses it.** Quantitative researchers and analysts building usage and attitude studies; category and brand strategists who need a defensible behavioural baseline; customer analytics teams reconciling survey and transactional views; UX and product researchers comparing claimed usage with telemetry; and anyone reviewing a behavioural profile assembled by an AI system from a survey, where "claimed" and "observed" are routinely merged without comment.

## 3. When to use it

- You need to describe what an audience does, and the description will be used to make a resource decision.
- Both survey and behavioural or transactional data exist, and they do not agree.
- The category is bought or used repeatedly, so repertoire and frequency matter more than a single choice.
- The business believes it has loyal customers and the evidence for that belief is a self-reported loyalty question.
- A strategy is being built around heavy users, and nobody has checked what share of the category the light buyers represent.
- Reported behaviour has changed between waves and nobody can say whether people changed or circumstances did.
- Segmentation is planned and behavioural basis variables need to be established first.
- Personas are planned and their behaviour section should be observed rather than claimed.

## 4. When NOT to use it

- **The only available evidence is a single claimed-frequency question, and it will be reported as a measurement.** This is the point at which the skill's honest output is a caveat rather than a profile. A "how often do you usually" item on a long recall period produces an estimate with known directional bias, and dressing it as a behavioural profile gives a business a precise number about a quantity nobody measured. Either report it explicitly as claimed frequency with the bias direction stated, or obtain a better measure. Do not convert it into an annual volume figure, because that multiplies the bias.
- **The question is whether two groups differ.** Comparing behaviour across groups is **09.04 Segment Comparison**, which supplies the testing, the base rules and the composition checks. This skill describes behaviour; it does not establish that a difference between groups is real.
- **The question is why people behave this way.** Motivation, need and job to be done are **09.05**, and explanation of a behavioural finding is **08.01 Finding to Insight Development**. A behavioural profile that arrives already explaining itself has crossed the K2 interpretation boundary without a signal word.
- **You are being asked to attribute behaviour to attitude.** Attitudinal correlates of behaviour are legitimate and useful; causal claims about them are not licensed by a cross-sectional design, per K4 §3.2. Where the request is "show that our brand attitudes drive purchase", say what the design can support (association, direction unestablished, and frequently the reverse direction is at least as plausible because people report better attitudes towards brands they already buy) and route the causal question to **05.06 Correlation, Regression and Causal Claim Control**.
- **The behavioural record is a partial view presented as a complete one.** A single retailer's transaction data describes purchases at that retailer, not category behaviour; a single app's telemetry describes use of that app, not the need it serves. Where the record has a coverage boundary, either state it on every figure or do not produce category-level statements from it, per K3 §6.
- **The category is genuinely one-off or very long-cycle.** Repertoire, frequency and share of category are meaningless for a purchase made once a decade. Describe the decision process instead, per **06.05 Customer Experience and Journey Analysis** or **09.05**.
- **The behavioural data identifies individuals and the purpose is not covered by consent.** Linked transactional and telemetry records are personal data, and analysis that profiles individuals for a purpose participants did not agree to is an ethics question before it is an analysis question. Go to **13.05 Research Ethics and Consent Design**.
- **The data has not been prepared.** Behavioural datasets carry duplicate transactions, test accounts, returned purchases, household versus individual identity confusion and gaps in coverage. Run **04.01** and **04.02** first, and reconcile the population in the behavioural file to the population the study is about, because they are rarely the same people.

## 5. Required inputs

**Required.** Without these the skill cannot run. If absent, ask.

- **At least one behavioural evidence source**, with its collection method documented: what was recorded, by what mechanism, over what period, for whom.
- **The population definition and the coverage boundary of each source.** Whose behaviour is in this file, and whose is missing. A transaction file covers card-paying customers of one channel; a panel covers panel members; telemetry covers logged-in sessions. This determines every claim the profile may make.
- **The period each source covers**, and whether that period was ordinary. A behavioural picture built on an atypical period describes the period, not the audience.
- **The category definition in use.** What counts as a purchase, a use, a session or an occasion. Behavioural figures are almost entirely determined by this definition, and it is usually undocumented.

**Optional, and what each one adds.**

- **A second, independent behavioural source:** allows claimed and observed to be compared directly, which converts the claimed-versus-observed gap from an assumption into a measured quantity for this specific audience and this specific question.
- **Occasion-level or in-the-moment data** (diary, experience sampling, timestamped records): allows occasion to be used as the unit of analysis, which is the single largest improvement available in most categories.
- **Competitive or category-wide behaviour**, from a panel or a category survey: allows share of category to be calculated rather than share of the client's own customer base, which is the only version of the number that means anything strategically.
- **Longitudinal or repeat-wave data on the same individuals:** allows the distinction between a changed person and a changed situation to be tested rather than assumed.
- **Attitudinal measures on the same respondents:** allows behaviour to be linked to attitude as association, and to be described rather than explained.
- **Household or account structure:** prevents a household's behaviour being attributed to one individual, which inflates every per-person frequency in the file.

## 6. Questions to ask before starting

1. **What behaviour, exactly, and over what period?** "Buys regularly" is not a behaviour. Determines the measure, the recall window and whether the data can support it. *Default if unanswered:* define the behaviour explicitly in the output, state the definition on every figure, and flag that the definition was set by the analyst.
2. **Which sources are observed and which are claimed?** Determines the reliability ranking and the reporting standard. *Default:* treat every survey-reported behaviour as claimed and label it so, without exception.
3. **What is the coverage boundary of each source?** Determines what population the profile describes. *Default:* state the boundary on every figure derived from that source, and do not make category-level claims from a single-provider record.
4. **Is the person or the occasion the right unit?** Determines the whole structure of the analysis. *Default:* if occasion-level data exists, check whether within-person variance exceeds between-person variance before choosing; if it does not exist, profile at person level and state that occasion variation could not be examined.
5. **Is the category one where people hold a repertoire?** Determines whether loyalty framings are appropriate at all. *Default:* assume a repertoire in any repeat-purchase category and measure share of category rather than reported loyalty.
6. **What share of buyers and of volume sits with light buyers?** Determines whether a heavy-user framing is defensible. *Default:* calculate and report the full buyer distribution before any heavy-user analysis is presented.
7. **Was the period ordinary?** Determines whether the profile is about the audience or about the moment. *Default:* check for known disruptions in the period and disclose any, per K4 §7.

## 7. Step-by-step methodology

**The position this method takes.** A behavioural profile is a claim about what people did. Almost everything that goes wrong follows from treating a claim about what people said they did as though it were the same thing. The method therefore begins by ranking sources rather than by analysing them, and it keeps the observed and claimed labels attached to every figure through to the last slide, because that is where they get lost.

**1. Inventory the sources and rank them on reliability, in writing.** The ranking is not absolute, but the ordering is consistent and worth stating explicitly. **Observed and transactional records** are the strongest: they record what happened, with no recall involved, and their weakness is coverage rather than accuracy. **In-the-moment capture** (diary entries, experience sampling, timestamped logging) is next: recall distance is minutes rather than months, but the act of recording changes some behaviour and participation is selective. **Short-window recall** ("did you do this yesterday", "in the last seven days") is workable, with the caveat that a single short window is a poor estimate of a person's rate. **Long-window recall and frequency claims** ("how many times in the last year", "how often do you usually") are the weakest, and are estimates rather than counts. **Claimed future behaviour** is not behavioural evidence at all and is filed separately. For each source record its coverage boundary, its period, its unit (person, household, account, session) and what it cannot see. *Correct result:* a source inventory with a reliability rank and a coverage statement per source, written before any figure is produced.

**2. Establish the definitions that will determine every number.** What counts as a purchase, a use, a visit, a session, an occasion, a category. Whether a repeat purchase within an hour is one event or two. Whether a household card is one buyer. Whether an abandoned session counts as a use. These decisions move behavioural figures by more than any analytical choice made later, and they are almost never documented, which makes comparison across studies and across waves unreliable. *Correct result:* a written definition set that a colleague could apply to the raw records and reproduce your figures.

**3. Treat claimed frequency as an estimate, and state the direction of its error.** People do not count; they estimate, using a rate heuristic ("about twice a month, so about 24 a year") or by recalling instances and extrapolating. Four consequences are predictable and should be stated wherever claimed frequency is reported. **Regular behaviours are over-reported** because the rate heuristic ignores the weeks it did not happen. **Irregular behaviours are under-reported** because instance recall misses events that were not distinctive. **Telescoping** pulls older events into the recall window, inflating counts over long periods. **Social desirability** inflates approved behaviours (exercise, saving, healthy eating, recycling) and deflates disapproved ones (alcohol, snacking, screen time, borrowing). The longer the window, the worse all four become, so a twelve-month frequency claim carries materially less information than a seven-day one. **Never multiply a claimed frequency up to an annual volume**, because that compounds the bias into a number that then enters a business case. *Correct result:* every claimed-frequency figure reported with the word "claimed" attached and the likely direction of bias stated once for each measure.

**4. Where both claimed and observed data exist, measure the gap rather than choosing between them.** Cross-tabulate claimed against observed at individual level where the data links, or at aggregate level where it does not. Report the size and direction of the gap, and who it is largest for, because the gap is often patterned: heavier users under-estimate, light users over-estimate, and the two errors together compress the true distribution towards the middle. **The gap is a finding, not a data quality problem to be resolved.** It tells you how the audience understands its own behaviour, which is directly relevant to how they will respond to communication about it. Where the two sources disagree, per K2 §4.4 report both and name the disagreement; do not average them and do not silently prefer the survey because it has more variables attached. *Correct result:* a measured claimed-versus-observed gap, with its pattern described, reported as evidence rather than reconciled away.

**5. Choose the unit of analysis by testing where the variance is.** In many categories the same person behaves differently on different occasions, and the difference between occasions within a person is larger than the difference between people. Where occasion-level data exists, compare within-person variation with between-person variation on the key behaviour. Where within-person variation dominates, **the person is the wrong unit**, and a person-level profile will describe an average that fits nobody and will produce an unstable segmentation downstream. Profile occasions instead, and describe people by their occasion repertoire. Where occasion data does not exist, say that the unit could not be tested and that person-level figures may be averaging across distinct situations. *Correct result:* an explicit, evidenced choice of unit, with the variance comparison shown where the data allowed it.

**6. Describe the category as a repertoire, and calculate share of category.** In most repeat-purchase categories, buyers use several brands and allocate volume across them; exclusive users of any one brand are a minority, and are often a minority of that brand's own buyers. Reporting "loyalty" from a self-reported main-brand question therefore describes a framing rather than a behaviour. Instead: report **repertoire size** (how many brands used in the period), **share of category requirements** (what proportion of a buyer's category volume goes to each brand), and **penetration** (what proportion of category buyers bought the brand at all). Penetration and share of requirements together tell a far more useful story than any loyalty index, because they separate two different growth routes: more buyers, or more share from the buyers you have. Where only single-brand data exists, say that repertoire could not be measured, and do not infer exclusivity from the absence of competitor records. *Correct result:* a repertoire and share-of-category picture, or an explicit statement that the data cannot support one.

**7. Report the whole buyer distribution before anyone profiles the heavy user.** In most measured repeat-purchase categories the buyer distribution is strongly skewed: a large majority of buyers buy infrequently, a small minority buys often, and because the light buyers are so numerous they collectively account for a substantial share of volume, frequently more than the heavy minority. Two errors follow from skipping this step. The first is **profiling the heavy user as the market**: heavy users are unrepresentative by construction, and a proposition designed around them addresses the group least in need of persuading. The second is **assuming volume growth must come from heavy users**, when the arithmetic in most categories favours reaching more light buyers. Produce the distribution, report the share of buyers and the share of volume in each band, and only then present any heavy-user analysis, clearly framed as a description of a minority. *Correct result:* a buyer distribution table showing buyers and volume by frequency band, appearing before any heavy-user content.

**8. Derive behavioural groups from behaviour, not from self-description.** Where behavioural groups are wanted, build them from recorded behaviour: frequency, recency, repertoire, share of category, occasion mix, channel mix, basket composition. Do not build them from a self-classification item ("would you describe yourself as a heavy user"), which measures self-image. Check that the resulting bands are not artefacts of the definition set from step 2, and check stability across periods, since a heavy buyer in one quarter is frequently a medium buyer in the next, purely through regression to the mean. **A behavioural group that does not persist across periods is a description of a period, not of a group of people**, and it will not sustain a targeting strategy. *Correct result:* behavioural groups with their persistence across at least two periods measured and reported.

**9. Separate a changed person from a changed situation.** When behaviour shifts between waves, three explanations compete and are routinely conflated: the same people changed their behaviour, different people are in the sample, or the same people are in a different situation (a price change, a supply problem, a seasonal effect, a life event). Only panel data on the same individuals can separate the first from the second. Where the sample is cross-sectional, say so, and say that the shift is at the aggregate level with the composition of the sample uncontrolled. Where individual-level data exists, decompose the change: how much comes from existing buyers changing rate, how much from buyers entering or leaving the category. This decomposition is usually more actionable than the headline movement. *Correct result:* a stated account of what kind of change this is, with the design's limits named.

**10. Link behaviour to attitude as association, and never as cause.** Attitudinal correlates of behaviour are worth reporting and are constantly over-claimed. Report the association, its strength and its base, and state explicitly that the direction is not established by this design, per K4 §3.2. In behavioural work the reverse direction is not a technicality: people report more favourable attitudes towards brands they already use, so an attitude-behaviour correlation is at least as consistent with usage producing attitude as with attitude producing usage. Where the causal question matters to the decision, name the design that would answer it and hand to **05.06**. *Correct result:* association language throughout, with the reverse-causation possibility stated in the same passage.

**11. Write the observed-versus-claimed statement into the output at the level of the figure.** Not once at the front, where it will be detached from the numbers on the first slide that travels. Every behavioural figure carries its source type, its base, its period and its coverage boundary. This is the discipline that survives contact with a summary deck, and it is the one most reliably lost. *Correct result:* a table in which a reader can tell, per row, whether the number came from a record or from a respondent's estimate.

**12. State what the behavioural evidence cannot see, and mark the judgement points.** Every behavioural record has blind spots: purchases in other channels, use by other household members, cash transactions, sessions while logged out, the category consumed but not bought by the respondent. Name them. Then mark for human decision, per K5: whether a behavioural difference is commercially material (§2.1), whether the category and occasion definitions match how the market is managed (§2.7), and whether linked behavioural analysis is within the consent given (§2.4). *Correct result:* a blind-spot list and review markers placed where the judgement occurs.

## 8. Analytical framework

The chain for every behavioural claim:

    Source → Coverage → Definition → Unit → Measure → Observed or claimed → Base and period

**Source.** Which record or which question, ranked on the step 1 reliability ordering.
**Coverage.** Whose behaviour this source can and cannot see.
**Definition.** What counts as an event, written before counting.
**Unit.** Person, household, account, session or occasion, chosen from where the variance lives.
**Measure.** Penetration, frequency, repertoire, share of category, recency, occasion mix.
**Observed or claimed.** Attached to the figure, not to the appendix.
**Base and period.** Per K2 §4.1, travelling with the number.

**The reliability ordering, stated once and applied everywhere.**

| Rank | Source type | Main weakness |
|---|---|---|
| 1 | Observed transactional or system records | Coverage boundary, identity resolution, no context |
| 2 | In-the-moment capture: diary, experience sampling, timestamped logs | Reactivity, participation selectivity, burden |
| 3 | Short-window recall, up to about seven days | One window is a poor estimate of a rate |
| 4 | Long-window recall and frequency claims | Rate heuristics, telescoping, social desirability. An estimate, not a count |
| 5 | Claimed future behaviour | Not behavioural evidence. Filed separately, never merged with measures 1 to 4 |

**The four behavioural measures that carry most of the meaning.**

| Measure | What it answers | Common misuse |
|---|---|---|
| Penetration | How many people bought or used at all | Confused with loyalty; ignored in favour of frequency |
| Frequency and its distribution | How often, and how unevenly | Reported as a mean over a skewed distribution |
| Repertoire and share of category | How the person spreads their behaviour | Replaced by a self-reported loyalty item |
| Occasion mix | When and in what situation | Collapsed into a person-level average that fits nobody |

**Against the K2 chain**, behavioural profiling stops at Finding. "Light buyers account for 54% of volume" is a finding. "Light buyers are the growth opportunity" is an interpretation and requires the signal language of K2 §3.2 plus a confidence level.

## 9. Output format

**1. Source inventory and reliability ranking.**

| Source | Type | Reliability rank | Coverage boundary | Period | Unit | What it cannot see |
|---|---|---|---|---|---|---|

**2. Definitions block.** What counts as a purchase, use, session, occasion and category, written so figures can be reproduced.

**3. Behavioural profile table.** Every row carries the measure, the figure, the source type marked observed or claimed, the base description and size, the period, and the coverage boundary.

**4. Buyer distribution.** Buyers and volume share by frequency band, presented before any heavy-user analysis.

**5. Repertoire and share of category**, where measurable, with penetration alongside. Where not measurable, an explicit statement of why.

**6. Occasion profile**, where occasion-level data exists: occasions, their frequency, what differs between them, and how people distribute across them.

**7. Claimed versus observed comparison**, where both exist: the gap, its direction, its pattern by subgroup, reported as a finding.

**8. Behavioural groups**, where derived: their definition, their sizes, and their persistence across periods.

**9. Attitude associations**, with direction explicitly unestablished and the reverse-causation possibility stated.

**10. Blind spots and what could not be established**, per K3 §5.2.

**When the evidence is thin.** The profile shrinks rather than being extrapolated. A claimed-frequency item is reported as a claimed frequency with its bias direction, not converted into an annual volume. A single-provider transaction file produces statements about that provider's customers, not about the category. Where repertoire cannot be measured, the output says so rather than inferring exclusivity from missing competitor data. And where only a self-classification item exists, no behavioural groups are derived, per K4 §1.

## 10. Quality checks

Run before anything is presented. These sit on top of K4 §8.

1. Does every behavioural figure state whether it is observed or claimed, at the figure rather than in a preamble?
2. Does every figure carry its coverage boundary, its period and its base?
3. Are the event and category definitions written down in a form that would reproduce the figures?
4. Has any claimed frequency been multiplied up into an annual or population volume?
5. Is the likely direction of recall bias stated for every claimed-frequency measure?
6. Where both claimed and observed data exist, has the gap been measured and reported rather than resolved?
7. Was the unit of analysis chosen by examining where the variance lies, or assumed to be the person?
8. Has the full buyer distribution been reported before any heavy-user analysis appears?
9. Is any strategic statement being made about heavy users without the light-buyer volume share alongside it?
10. Is loyalty being reported from a self-report item rather than from share of category?
11. Has exclusivity been inferred anywhere from the absence of competitor data?
12. Are behavioural groups derived from recorded behaviour rather than self-classification, and has their persistence been checked across periods?
13. Where behaviour changed between waves, is the changed-person versus changed-situation distinction addressed and the design's limits stated?
14. Is any attitude-behaviour relationship described in causal language, and is the reverse direction acknowledged?
15. Is a mean reported anywhere on a skewed behavioural distribution without the distribution beside it, per **05.01**?
16. Are the blind spots of the behavioural record listed?

## 11. Common failure modes

| Failure | How to recognise it | How to prevent it |
|---|---|---|
| **Claimed reported as observed** (the signature failure) | A usage table with no indication that every figure came from a recall question | Source type on every row; the word "claimed" survives into the summary |
| **Frequency claim scaled to volume** | "Users buy 2.4 times a month, so 28.8 times a year, so a market of X units" | Never multiply a claimed rate. State it as claimed frequency and stop |
| **Mean frequency on a skewed distribution** | An average buying rate almost no buyer is near | Report the distribution; use median plus bands; per **05.01** |
| **Heavy user as the market** | A whole strategy built on the top decile, with no mention of the rest | Buyer distribution first, always, with volume share by band |
| **Loyalty from a self-report item** | "68% say we are their main brand" presented as a behavioural fact | Share of category from repertoire data; label the self-report as self-image |
| **Exclusivity inferred from missing data** | A single-retailer file used to claim customers buy nowhere else | Coverage boundary on every figure; state what the source cannot see |
| **Person-level averaging of occasion behaviour** | A profile that contradicts itself, or segments that will not stabilise | Test within-person against between-person variance; profile occasions where it dominates |
| **The reconciled gap** | Claimed and observed disagreed, and one was quietly dropped | Report both, measure the gap, treat it as a finding per K2 §4.4 |
| **Changed people read as changed behaviour** | A cross-sectional wave shift described as customers changing habits | Name the design limit; decompose only where panel data exists |
| **Behavioural groups that do not persist** | Heavy buyers in Q1 who are medium buyers in Q2, targeted as though fixed | Measure persistence across periods before deriving a strategy |
| **Attitude promoted to driver** | "Brand trust drives purchase frequency", from cross-sectional data | Association language; state the reverse direction; hand to **05.06** |
| **AI: merging sources of different reliability** | One clean table combining transaction data and recall claims with no distinction | The reliability rank stays attached to every figure through every transformation |
| **AI: inventing plausible behavioural rates** | A frequency or share figure that fits the surrounding numbers and was not calculated | Recalculate from source; per K4 §2.1, never fill a behavioural cell with a plausible value |

## 12. AI guardrails

Skill-specific only. K4 applies in full and is not repeated here.

1. **Never present claimed behaviour as observed behaviour.** The label travels with the figure into every downstream document, including one-line summaries.
2. **Never convert a claimed frequency into a volume, an annual figure or a market size.** The bias compounds and the result enters a business case as a fact.
3. **Never merge sources of different reliability into a single figure** without stating the composition and the weaker source's limitations.
4. **Never make a category-level claim from a single-provider record**, and never infer exclusivity or loyalty from the absence of competitor data.
5. **Never present heavy-user analysis before the full buyer distribution**, and never describe heavy users in language that implies they represent the category.
6. **Never report loyalty from a self-report item as though it were a measure of behaviour.**
7. **Never average occasion-level behaviour to person level** without stating that the average may describe no actual occasion.
8. **Never resolve a claimed-versus-observed disagreement by preferring one source silently.** Measure the gap and report it.
9. **Never use causal language for an attitude-behaviour association**, and where the association is reported, state that usage plausibly produces attitude as well as the reverse.
10. **Never describe a between-wave shift as people changing their behaviour** unless the same individuals were measured.
11. **Never omit the coverage boundary and the period from a behavioural figure.** Both determine what the number is a number about.

## 13. Best-practice principles

- **What people say they do is evidence about how they see themselves.** It is genuinely useful, and it is not a count. Treated as a count it produces confident, precise and biased business inputs.
- **The recall window is the single biggest determinant of quality in claimed behaviour.** Seven days beats a year by more than any other design improvement available in a questionnaire.
- **The gap between claimed and observed is a finding.** It says how the audience understands its own behaviour, which is exactly what communication has to work with.
- **Behaviour usually belongs to the occasion, not to the person.** Where within-person variance dominates, person-level profiling averages away the structure and everything built on it becomes unstable.
- **Most categories run on repertoires.** Exclusive buyers are a minority almost everywhere, and a loyalty framing quietly redefines the strategic question towards defending share and away from reaching more buyers.
- **Penetration and share of requirements are two different growth routes, and reporting only one hides the choice.**
- **The light-buyer majority is where most of the volume usually is, precisely because there are so many of them.** A strategy addressed only to heavy users addresses the people least in need of addressing.
- **The heavy user is unrepresentative by construction.** Profiling them tells you about heavy users, which is a legitimate and much smaller question than the one usually being asked.
- **Behavioural bands regress.** Today's heavy buyers include people having an unusual quarter, and next period they will look average. Check persistence before targeting.
- **Coverage is the weakness of good data and recall is the weakness of cheap data.** Knowing which one you are living with determines every caveat in the report.
- **An attitude measured after behaviour is partly a report on the behaviour.** People like what they use. This is not a technicality; it inverts most attitude-to-behaviour narratives.
- **Definitions move behavioural numbers more than analysis does.** Write them down first, or no figure will be comparable to any other figure ever produced about this category.

## 14. Worked example

*Fictional scenario, used to demonstrate method. The organisation, figures and findings below are invented.*

**INPUT.** A soft drinks manufacturer wants a behavioural profile of its flagship brand's buyers to inform a range and pack-format decision. Available: a survey of 1,800 category buyers with claimed frequency, claimed main brand and an attitudinal battery; a continuous purchase panel covering 3,200 households with scanned category purchases over 52 weeks; and a seven-day drinking-occasion diary from 420 panel members.

**PROCESS.**

*Steps 1 and 2.* Sources ranked: panel records first, occasion diary second, survey claims fourth. Coverage stated: the panel covers take-home grocery purchasing only, so on-the-go and hospitality purchases are invisible to it, which matters directly to a pack-format decision. Definitions written: a purchase occasion is a shopping trip, a drinking occasion is a distinct consumption event, and multipacks are converted to units.

*Steps 3 and 4, the gap.* Survey respondents claim a mean of 3.1 category purchases per month. The panel records a median of 1.4 and a mean of 2.2 for the same households. The gap is largest among self-declared heavy users, who over-report by roughly a factor of two, and smallest among light buyers, who over-report slightly. **Judgement call:** the commercial team wants one number. Averaging the two would produce a figure describing nothing. The resolution is to use panel figures for all volume and frequency statements, use survey figures only for what people believe about their own behaviour, and report the gap itself as a finding, because it says the brand's most engaged buyers substantially overestimate their own loyalty, which reframes a planned loyalty campaign.

*Step 6, repertoire.* The survey said 68% name the brand as their main one. The panel says the average category buyer purchases 4.3 brands in 52 weeks, and the flagship brand's share of category requirements among its own buyers is 31%. Exclusive buyers are 6% of its buyer base. The main-brand claim is reported as self-image and clearly labelled; the behavioural picture replaces it as the basis for the decision.

*Step 7, the distribution.* Buyers in the lightest two frequency bands are 71% of the brand's buyers and account for 46% of its volume. The heaviest decile accounts for 24%. The pack-format proposal had been designed around the heaviest decile's basket. This table, placed before any heavy-user content, changed the direction of the project.

*Step 5, the unit.* The occasion diary shows that within-person variation in what is drunk, and why, exceeds between-person variation: the same individual drinks the category differently at breakfast, at work and socially. Person-level "brand preference" is therefore an average across situations that do not resemble each other. The profile is rebuilt at occasion level, and people are described by their occasion repertoire.

*Step 10.* The attitudinal battery correlates with purchase frequency. It is reported as an association, with the base, and with the explicit statement that people report warmer attitudes towards brands they already buy frequently, so the direction is not established by this design.

**OUTPUT.** A source inventory with coverage boundaries; a definitions block; an occasion-level behavioural profile with a person-level repertoire summary; a buyer distribution showing the light-buyer volume share ahead of any heavy-user content; repertoire and share-of-category figures replacing the loyalty claim; a measured claimed-versus-observed gap reported as a finding; attitude associations with direction unestablished; and a blind-spot list naming on-the-go and hospitality purchasing, with a **researcher decision required** marker on whether the pack decision can be made at all from take-home data, per K5 §2.7.

## 15. Advanced usage

**Occasion segmentation.** Where occasion data exists and within-person variance dominates, segment occasions rather than people and hand the result to **09.01**. Occasion segmentations are frequently more stable than person segmentations in the same category, because they are built on the level at which the behaviour actually varies, and they produce a different and usually more actionable deliverable: the business targets moments, formats and contexts rather than people.

**Decomposing a volume change.** Where panel data allows, break a change in brand volume into its components: change in penetration, change in purchase frequency among retained buyers, change in share of requirements, and the contribution of buyers gained and lost. Most volume changes turn out to be penetration movements, which points at a completely different intervention from the one a frequency framing suggests.

**Reconciling survey and telemetry in digital products.** The same claimed-versus-observed gap appears in product research, usually larger: claimed session frequency and duration are systematically overstated, and the gap is patterned by engagement. Where both exist, the honest deliverable includes the gap and its pattern, and the design implication often lies in the gap rather than in either number.

**Behaviour under an unusual period.** Where the period includes a disruption, do not adjust the data to remove it. Report the period behaviour as period behaviour, and where a prior comparable period exists, show both. A profile silently smoothed to look normal is a profile of a market that did not exist.

**When the standard approach does not fit.** Where only claimed data is available and no observed source can be obtained, the profile is still worth producing, provided it is labelled throughout as claimed, the bias directions are stated per measure, and no derived volumes are calculated from it. Cap confidence at moderate for relative statements and at low for absolute levels, per K3 §3.4, and name the observed source that would fix it.

## 16. Skill chain

**Recommended previous skills:**
- **04.01 Data Validation** and **04.02 Data Cleaning.** Hand over behavioural records with duplicate transactions, test accounts, returns and identity resolution addressed and logged.
- **04.04 Data Transformation and Dataset Preparation.** Hands over derived behavioural variables (frequency bands, recency, repertoire counts, share of category) with their construction rules, which determine every figure downstream.
- **05.01 Descriptive Analysis.** Hands over distributions, which is where the skew that governs the buyer distribution first becomes visible.
- **03.04 Fieldwork Monitoring and Response Quality**, where diary or in-the-moment capture was used and participation selectivity needs assessing.

**Recommended next skills:**
- **09.01 Audience Segmentation.** Takes behavioural basis variables, the occasion structure and the evidenced answer to whether the person is the right unit.
- **09.05 Needs, Motivation and Jobs-to-be-Done Analysis.** Takes the occasions and behaviours and works out what people are trying to achieve in them, which this skill deliberately does not do.
- **09.04 Segment Comparison.** Takes behavioural measures and compares groups on them properly, with testing and composition checks.
- **09.02 Persona Development.** Takes observed rather than claimed behaviour for the current-behaviour element, together with the observed-versus-claimed labelling the persona must carry onto the artefact.
- **08.01 Finding to Insight Development.** Takes behavioural findings and explains them, including the claimed-versus-observed gap, which is often the most explanatory material in the study.

**Runs well alongside:**
- **05.06 Correlation, Regression and Causal Claim Control**, wherever an attitude-behaviour relationship is being pushed towards a causal reading.
- **05.05 Trend and Tracker Analysis**, for the changed-person versus changed-situation question across waves.
- **13.05 Research Ethics and Consent Design**, wherever individual-level behavioural records are linked to survey responses.

---
A Yazi Supplied Skill and resource.
