---
name: customer-experience-and-journey-analysis
description: >
  Analyses experience and journey data: mapping the journey as measured rather
  than as designed, separating touchpoint-level from relationship-level
  measurement, identifying the moments that disproportionately shape overall
  evaluation without asserting causation, scoring pain points on frequency,
  impact and recoverability, distinguishing a dissatisfying moment from a
  churn-driving one, analysing effort and service recovery, and matching
  operational data to experience data. Use for "analyse the CX programme",
  "which touchpoint matters most", "why is satisfaction high but churn rising",
  "map the customer journey from the data", "find the pain points", "does fixing
  this reduce churn", or "link our operational data to the survey".
category: 06 Specialist and Advanced Analysis
ref: "06.05"
tier: 2
inherits: [K2, K3, K4, K5]
---

# Customer Experience and Journey Analysis

## 1. One-line description
Turns scattered experience measurement into a journey diagnosis that names where customers struggle, how badly, for whom, and what it costs, while holding the two lines that experience analysis routinely crosses: that a correlate of overall satisfaction is not a cause of it, and that satisfaction measured among people who stayed says nothing reliable about the people who left.

## 2. What this skill is used for

**The research problem it solves.** Experience programmes generate continuous measurement and very little diagnosis. Six failures recur. The journey analysed is the journey the organisation designed, with its stages named after internal processes, rather than the journey the customer experienced, which crosses channels, doubles back and includes steps the organisation does not know about. Touchpoint scores and relationship scores get mixed in one chart although they answer different questions on different bases at different moments. The moments that correlate with overall evaluation get called drivers and then get called causes, and a budget follows. Pain points get ranked by how loudly they were mentioned rather than by frequency multiplied by impact multiplied by whether the organisation can recover from them. The dissatisfying moment and the churn-driving moment are assumed to be the same, and they frequently are not: people tolerate a great deal in categories where switching is hard, and leave over things they rated adequately. And the whole programme measures people who are still customers, which is a survivorship problem that no sample size fixes. This skill supplies the structure that catches each of these.

**Where it sits in the research lifecycle.** After experience, operational and qualitative data have been collected, and before prioritisation of improvement work. It sits across the quantitative and qualitative boundary by nature, because a journey is described by numbers and explained by language, and it hands a prioritised, evidenced pain point set to insight and recommendation development.

**Typical use cases.**
- Building a journey map from measured evidence rather than from a workshop.
- Diagnosing which stage of a journey is losing people, and for which customers.
- Prioritising pain points on a defensible severity basis rather than on volume of mentions.
- Separating what makes an experience unsatisfying from what makes a customer leave.
- Analysing service recovery and whether a well-handled failure repairs the relationship.
- Linking operational records to experience responses, and handling what that linkage breaks.
- Analysing a journey where the customer uses several channels for one job.
- Auditing an experience programme that reports healthy scores while the business loses customers.

**Who uses it.** Customer experience and insight managers; service designers and UX researchers working at journey rather than screen level; operations and contact centre analysts; researchers running experience programmes for clients; and public sector and healthcare teams analysing service user journeys where the vocabulary differs and the method does not.

## 3. When to use it

- Experience data exists across several touchpoints and needs assembling into a journey view.
- A programme reports stable or improving scores while a commercial or operational indicator moves the other way.
- Pain points have been listed and need prioritising with something more defensible than mention volume.
- A specific moment is about to be funded as the priority because it correlates with the overall score.
- Operational data and survey data need joining, and the matching implications need thinking through.
- A journey crosses channels and the single-channel measurement is producing a partial picture.
- Service recovery performance needs assessing.
- An experience report is about to drive investment and needs auditing.

## 4. When NOT to use it

- **The question is which factors relate to an outcome, statistically.** Regression on satisfaction drivers, relative importance, controlling for confounders: that machinery belongs to **05.04 Driver Analysis** and **05.06 Correlation, Regression and Causal Claim Control**, which own the estimation and the causal language. This skill uses their output and adds the journey structure, the severity framework and the survivorship discipline. Where the whole request is "run the driver model", start there.
- **The question is whether an intervention improved the experience.** That is a causal question about a change, and it needs a design: **06.04 Experiment and A/B Test Analysis**. Comparing satisfaction before and after a service change, with no control, attributes to the change every other thing that happened, and experience metrics are unusually exposed to this because they move with season, weather, news, price changes and the composition of who responds.
- **There is measurement at one touchpoint only.** A single transactional survey describes that transaction. It does not describe a journey, and calling it one implies coverage that does not exist. Report it as touchpoint measurement and say which parts of the journey are unmeasured, which is usually the more useful output.
- **The evidence is entirely qualitative and the question is prevalence.** Where the input is twenty interviews and the request is how common a pain point is, the honest answer is that the sample cannot establish prevalence. Use **07.01 Thematic Analysis** for the structure and depth of the experience and be explicit that frequency claims require a different design.
- **The population that matters is not in the data.** Where the question is why customers leave and the sample is current customers, the answer is not in this dataset, and analysing it harder will not produce one. Say what population would answer the question (lapsed customers, rejected applicants, people who abandoned mid-journey, non-users) and treat the current-customer analysis as describing the retained population only.
- **The response rate is low and unknown in its composition.** Experience surveys frequently run at very low response rates, and responders differ systematically from non-responders on exactly the dimension being measured, since people with strong experiences respond more. Where the response rate is low and no non-response analysis is possible, report the scores as describing responders, per **K4 §7**, and do not present them as customer-level measures.
- **The request is to produce a journey map as a communication artefact.** A visually complete journey map with a smooth emotional curve is a design deliverable, not an analysis, and building one from thin evidence fills the unmeasured stages with plausible invention. Where evidence is absent for a stage, the map must show the gap, per **K4 §1**.
- **A proprietary experience score is to be reported against an external benchmark of unknown provenance.** Benchmark figures for experience metrics circulate widely with no methodology attached, and comparing a score collected one way against a benchmark collected another is not a comparison. Either the benchmark's method matches or it is not used.

## 5. Required inputs

**Required. Without these the skill cannot run. If absent, stop and ask.**
- **An inventory of what is measured, where, and how.** Every survey, its trigger, its question wording, its scale, its timing relative to the event, its base and its response rate. Experience programmes accrete over years and nobody holds the whole picture; building this inventory is the first real analytical act.
- **The unit and base of every metric.** Whether a score is per transaction, per interaction, per relationship or per customer, and what the denominator is. Mixing these is the commonest structural error in experience reporting.
- **The journey scope, defined as a customer job rather than an internal process.** "Renewing cover", "getting a diagnosis", "opening an account and using it for the first time". Where the scope is defined by department, the analysis will reproduce the organisation's structure rather than the customer's experience.
- **The population definition and its survivorship status.** Who is eligible to be surveyed, at what point, and whether people who left, abandoned or were rejected are included. This determines what the entire analysis can claim.

**Optional, and what each one adds.**
- **Operational records at the individual level.** Transform the analysis: they supply what actually happened (wait times, failure counts, contacts, resolution times) against what was reported, and they cover everyone rather than only responders. The matching problems they create are handled in Step 9.
- **Outcome data: churn, retention, spend, complaint escalation.** The only way to distinguish a dissatisfying moment from a consequential one.
- **Open-ended responses and verbatim contact records.** Supply the mechanism behind a score, which the score never contains. Hand large volumes to **06.06 Large-Scale Text Analytics** and smaller sets to **07.02 Open-Ended Response Coding**.
- **Journey-level qualitative work.** Reveals the stages nobody measures, including the ones the organisation does not know exist, and is usually where the real diagnosis originates.
- **Channel and interaction logs across channels.** Enable the multi-channel reconstruction in Step 10, without which a journey that crosses channels appears as several unrelated fragments.
- **Recovery case records.** Enable the recovery analysis in Step 8, which is otherwise not possible.
- **Non-responder characteristics.** Enable a non-response assessment, which converts an unknown bias into a bounded one.

## 6. Questions to ask before starting

1. **Whose journey, doing what job, over what period?** Determines the scope and prevents the analysis defaulting to the organisation's process map. Default if unanswered: define the job from the qualitative or operational evidence and state the definition used.
2. **Is the question about a moment or about the relationship?** These need different measurement and are frequently conflated. Default: analyse both separately and report the difference between them, which is often itself the finding.
3. **Are lapsed, rejected or abandoned customers in the data anywhere?** Determines whether any churn-related claim is possible. Default: state that the analysis describes retained customers and name the missing population.
4. **What is the response rate at each measurement point, and does it vary by stage or channel?** Determines comparability across touchpoints, since a stage with 3% response and one with 22% are not comparable measures. Default: report response rates on every chart and flag comparisons across very different rates.
5. **Can operational and experience records be joined at the individual level, and how reliable is the join?** Determines whether the strongest analysis is available. Default: work with the aggregate and state the loss.
6. **What can the organisation actually change?** Determines whether a pain point is actionable, which is part of severity. Default: score frequency and impact, and mark recoverability as unassessed pending an operational view.
7. **Is a proprietary or single-item score being used as the headline?** Determines how much interpretive weight it can carry. Default: report it as one measure among several, describe it generically by what it asks, and never treat it as a summary of the experience.

## 7. Step-by-step methodology

**Step 1. Build the measurement inventory before building anything else.** One row per measurement point: what triggers it, when it fires relative to the event, what it asks in exact wording, what scale, what base, what response rate, and which customers can receive it. This is unglamorous and it is where the analysis is won, because it reveals immediately that the programme measures four touchpoints heavily, three lightly and five not at all, and that two of the heavy ones use different scales and cannot be compared. A correct result is an inventory table plus a coverage statement naming the unmeasured stages.

**Step 2. Map the journey as measured, then against the journey as designed, and analyse the gap.** Lay the measurement inventory against the stages of the customer job. Three categories emerge: stages measured well, stages measured badly (wrong timing, tiny base, ambiguous wording), and stages not measured at all. Then compare this against the organisation's own journey map. The gaps are the finding. Organisations systematically over-measure the moments they own and control (the purchase, the call they answered) and under-measure the moments between them (the waiting, the second attempt, the conversation with a friend, the decision not to bother), which is where dissatisfaction accumulates. Where qualitative evidence exists, use it to name the unmeasured stages, since a customer will describe steps that appear in no process document. A correct result is a two-layer map: the journey as the customer does it, with the measurement coverage overlaid, and blank space shown as blank rather than filled.

**Step 3. Separate touchpoint-level from relationship-level measurement and keep them separate.** A touchpoint measure asks about a specific event, close to it in time, of the people who had it. A relationship measure asks about the overall connection, at an arbitrary moment, of a defined customer population. They differ in base (the transacting subset against all customers), in timing (immediate against retrospective), in what recall they draw on, and in what they predict. Three rules follow. Do not average them into a composite index: the result has no interpretation. Do not compare a touchpoint score against a relationship score and describe the difference as a change. And when they disagree, treat it as a finding rather than an inconsistency: uniformly good touchpoint scores with a weak relationship score usually means the individual interactions are fine and the accumulation is not, which points at effort, repetition or the gaps between touchpoints rather than at any single moment. A correct result is two separate views with an explicit reading of their relationship.

**Step 4. Identify the moments that disproportionately shape overall evaluation, and label the claim honestly.** The standard analysis relates touchpoint-level satisfaction to overall evaluation and reports which touchpoints have the strongest association. It is genuinely useful and it is not a causal result. Three specific hazards, beyond the general one. *Common-method association:* when the touchpoint rating and the overall rating come from the same person in the same survey, they share mood, response style and the halo of the overall attitude, which inflates the association independent of any real influence. *Reverse causation:* customers who are already satisfied rate every touchpoint more generously, so the association runs both ways. *Recency and peak weighting:* people's overall evaluation weights the most recent and most intense moments disproportionately, so a touchpoint's association reflects its position and salience as much as its importance. The defensible statement is that these moments are the strongest correlates of overall evaluation in this data, that the design does not establish direction, and that a change to one of them would need testing. Where operational data allows a moment to be related to a later outcome measured separately, the evidence is considerably stronger and should be preferred. A correct result is a ranked correlate list with the language discipline of **05.06** applied and a note of which hazards apply.

**Step 5. Score pain points on frequency, impact and recoverability rather than on mention volume.** Mention volume measures who complains, not what matters, and it is dominated by whichever channel is easiest to complain in. Score each pain point on three dimensions instead. *Frequency:* how many customers encounter it, expressed on a defined base, from operational data where possible rather than from survey mentions. *Impact:* how much it degrades the experience when it occurs, measured by the difference in the relevant score between those who encountered it and those who did not, with base sizes, or by severity of consequence for the customer. *Recoverability:* whether the organisation can put it right once it has happened, and how completely. This third dimension is what converts a pain point list into a prioritisation, because a frequent, high-impact, unrecoverable failure is categorically worse than a frequent, high-impact one that a good service response repairs, and the two need different responses (prevention against recovery capability). A correct result is a severity table with all three dimensions scored, the base for each, and the derivation stated. Do not combine them into a single index without saying how, since the weighting is a judgement and belongs to a human (**K5 §2.1**).

**Step 6. Distinguish the dissatisfying moment from the churn-driving one, with evidence.** These are different questions and the answer differs more often than programmes assume. People tolerate considerable friction in categories with high switching costs, and leave over things they rated adequately, sometimes long after the event. Establishing the difference requires outcome data joined to experience data: for each pain point or touchpoint, the subsequent retention or behaviour of those who encountered it against those who did not, with the base for each and the time window stated. Two cautions. The comparison is observational, so people who encounter a failure may differ from those who do not in ways that also affect retention, and this must be said. And the time lag matters: a churn effect may appear months later, so a window that is too short will show nothing. Where outcome data does not exist, the honest output is that the study measures dissatisfaction and cannot identify churn drivers, which is a considerably more useful sentence than a ranked list presented as if it could. A correct result is a two-column view, dissatisfaction rank against retention association, with the divergences highlighted, since the divergences are the finding.

**Step 7. Analyse effort as its own construct.** Effort measures how hard the customer had to work: steps taken, times they had to repeat themselves, channels they had to switch between, attempts before resolution. It is distinct from satisfaction, is often more strongly associated with subsequent behaviour, and is much better measured operationally than by asking. Where operational data exists, build the effort profile directly: contacts per resolution, repeat contact rate, channel switches per job, elapsed time from first contact to resolution against the customer's expectation. Where only survey data exists, use an effort question and be aware that it is subject to the same common-method issues as any self-report. The specific analytical value of effort is that it is actionable in a way satisfaction is not: nobody knows how to make a customer 8% more satisfied, and everybody knows how to remove a step. A correct result is an effort profile per journey stage with operational and survey measures reported separately.

**Step 8. Analyse recovery, and test rather than assume the recovery effect.** The claim that a well-handled failure produces a stronger relationship than no failure at all is widely repeated and is not universally true: it depends on the severity of the failure, the speed and completeness of the recovery, and whether the customer had to fight for it. Analyse it directly where the data allows: compare four groups on subsequent satisfaction and retention, being no failure, failure with no recovery attempt, failure with an attempted recovery, and failure with a resolved recovery, with the base for each. Report what the data shows rather than the general claim. Also analyse recovery reach: what proportion of failures are ever detected by the organisation, which is usually far lower than the failure rate, since most dissatisfied customers do not complain. That detection gap is frequently the largest single improvement opportunity in a CX programme and it is invisible to any analysis that starts from complaints. A correct result is a four-group comparison with bases, plus a detection rate estimate.

**Step 9. Join operational and experience data, and handle what the join breaks.** Individual-level linkage is the strongest move available in experience analysis, and it introduces five problems that must be handled explicitly. *Match rate:* not every survey response links to a record, and the unmatched are not random, so report the match rate and compare matched against unmatched on anything available. *Identity across channels:* the same person may appear as different identifiers in different systems, so a journey reconstructed by identifier may be several people or several fragments of one. *Timing alignment:* an operational event and a survey response have different timestamps, and joining the response to the wrong event is easy and silently wrong. *Definitional mismatch:* the operational definition of a resolved case and the customer's sense of resolution routinely differ, and where they differ, the operational record is not the ground truth people treat it as. *Consent and permitted use:* joining survey responses to operational records may exceed the consent given, and this is an ethics gate, not a data engineering one (**K5 §2.4**, and **13.05**). A correct result is a linked dataset with the match rate, the matching rule, the unmatched comparison and the consent position all documented, per **K2 §6**.

**Step 10. Reconstruct multi-channel journeys around the customer's job, not the channel.** Where a customer uses several channels to complete one task, channel-level measurement produces fragments that individually look fine and collectively describe a bad experience: a good web session, a good call, a good branch visit, for a job that took three attempts and should have taken one. Reconstruct by grouping interactions into job-level episodes using a defined rule (same identifier, same subject, within a time window), and then measure at the episode level: attempts per job, channels per job, elapsed time, resolution rate at first attempt. Report the episode-level distribution, not the average, because the story is in the tail: the 12% of jobs taking four or more contacts generate most of the cost and most of the damage. State the episode rule explicitly, because it is a modelling choice that changes the numbers. A correct result is an episode-level view with the grouping rule stated and the distribution shown.

**Step 11. Address the survivorship problem in writing, at the point of every claim it affects.** Satisfaction is measured among people who are still customers. Those most damaged by the experience have left and are not in the sample, so the measured distribution is truncated at the bad end, and the truncation gets worse the better the organisation's churn rate hides it. Three responses, in order of strength. Survey the lapsed population directly, which is the only real fix and is rarely done. Analyse the last measured score of customers who subsequently left, which is available in any programme with individual-level linkage and is frequently revealing, since departing customers often scored adequately shortly before leaving. And at minimum, state the limitation next to every satisfaction level and every trend, since a rising satisfaction score in a period of rising churn may simply be the departure of the dissatisfied. A correct result is an explicit survivorship statement wherever a level or a trend is reported, and where possible a comparison of leavers' final scores against stayers'.

**Step 12. Assemble the journey output in a form that can be acted on.** Three artefacts, in this order. The **evidenced journey map**: stages as the customer experiences them, with measurement coverage, the score or operational indicator at each measured stage with its base, and blank space where nothing is measured. The **prioritised pain point table**: frequency, impact, recoverability, the affected segment, the evidence reference for each, and the outcome association where available. And the **fix-and-check list**: for each prioritised pain point, what would need to change, what would be measured to know it had worked, and what the current baseline is. That third artefact is what distinguishes an experience analysis from an experience report, because it commits the finding to something checkable. A correct result is all three, with every number carrying its base and every gap shown as a gap.

## 8. Analytical framework

    Measured journey → Coverage gap → Moment against relationship
        → Correlates (not causes) → Severity (frequency, impact, recoverability)
            → Consequence (dissatisfaction against churn) → Fix and check

The framework's first two terms exist because the most common error in this field is analysing a journey nobody verified was the journey. Everything downstream inherits that scope, so a coverage gap at stage three means the severity ranking is a ranking of the stages that happened to be measured.

The fourth term carries the discipline. Experience data is almost entirely observational and almost entirely self-reported at the same moment as the outcome it is being related to, which makes it unusually prone to producing confident causal-sounding conclusions. The framework marks the transition explicitly: correlates are named as correlates, and the step to consequence requires outcome data measured separately, not a stronger adjective. The final term is the commitment: an experience finding that does not name what would be measured to know the fix worked will not survive contact with an operations team.

## 9. Output format

**The measurement inventory.**

| Measurement point | Trigger | Timing after event | Question wording | Scale | Base definition | Response rate | Customers eligible |
|---|---|---|---|---|---|---|---|

**The evidenced journey map.** Customer-defined stages across the top; rows for measurement coverage, key indicator with base, and evidence reference. Unmeasured stages appear as explicit blanks labelled `[not measured]`.

**The touchpoint against relationship view.** Two panels, never merged, with a written reading of any divergence.

**The correlate table.**

| Moment | Association with overall evaluation | Base | Same-survey measurement? | Outcome-data corroboration | Reading |
|---|---|---|---|---|---|

The fourth and fifth columns exist so a reader can see which associations are vulnerable to common-method inflation and which are corroborated independently.

**The pain point severity table.**

| Pain point | Frequency (base, source) | Impact (score difference or consequence, base) | Recoverability | Segment concentration | Outcome association | Evidence ref |
|---|---|---|---|---|---|---|

**The recovery table.** Four groups (no failure, failure unrecovered, recovery attempted, recovery resolved) by subsequent satisfaction and retention, with bases, plus the estimated detection rate.

**The episode view.** Distribution of attempts per job, channels per job and elapsed time, with the episode grouping rule stated.

**The fix-and-check list.** Pain point, proposed change, the measure that would show it worked, the current baseline, the owner.

**Where the evidence is thin**, use these forms:
- `[not measured: no instrument covers this stage]`
- `[base below 30 at this touchpoint: verbatim only, no score]`
- `[response rate 3%: describes responders, not customers]`
- `[correlate only: same-survey measurement, direction not established]`
- `[churn association not assessable: no outcome data joined]`
- `[recoverability unassessed: no operational view of resolution]`
- `[survivorship: measured among retained customers only]`

## 10. Quality checks

**K4 §8** runs anyway. These are specific to journey analysis.

1. Does the journey map show unmeasured stages as blank rather than inferred?
2. Is every score labelled as touchpoint-level or relationship-level, and are the two kept apart?
3. Does every metric carry its base, its denominator definition and its response rate?
4. Are touchpoints with very different response rates being compared without that being stated?
5. Is any correlate of overall evaluation described in causal language?
6. Is common-method inflation noted wherever the predictor and outcome came from the same survey?
7. Are pain points scored on frequency, impact and recoverability rather than on mention volume?
8. Is any composite severity index presented without its weighting shown?
9. Is the distinction between dissatisfying and churn-driving addressed explicitly, with evidence or with its absence stated?
10. Is the survivorship limitation stated at the point of every level and trend it affects?
11. Where operational and survey data are joined, are the match rate, the matching rule and the unmatched comparison reported?
12. Is the consent position for the data linkage documented?
13. Is the episode grouping rule stated wherever a multi-channel journey figure appears?
14. Are proprietary or single-item scores described by what they ask rather than treated as summaries of the experience?
15. Does the output include what would be measured to know a fix had worked?

## 11. Common failure modes

| Failure | How to recognise it | How to prevent it |
|---|---|---|
| **The organisation's journey, not the customer's** | Stage names matching internal department names | Scope defined as a customer job; qualitative evidence used to name the unmeasured stages |
| **Blank stages filled in** | A smooth emotional curve across stages with no measurement | Gaps shown as gaps (**K4 §1**) |
| **Touchpoint and relationship scores merged** | A composite experience index built from both | Two panels, never averaged; divergence read as a finding |
| **Correlate reported as driver, then as cause** | "Improving the onboarding call will lift satisfaction by 4 points" | **05.06** language discipline; the common-method column in the correlate table |
| **Mention volume as priority** | A pain point ranking that matches the complaint channel's volume | Frequency from operational data; three-dimension severity scoring |
| **Recoverability ignored** | Equal priority given to a repairable and an unrepairable failure of the same size | Third dimension mandatory in the severity table |
| **Dissatisfaction assumed to equal churn** | A retention programme built on the lowest-scoring touchpoint | Outcome data joined, or the inability to assess churn stated |
| **The survivorship blind spot** | Rising satisfaction during rising churn, reported as improvement | Survivorship statement at every level and trend; leavers' final scores analysed |
| **Low response rate treated as a sample** | Customer-level claims from a 3% response rate | Response rate on every chart; claims restricted to responders |
| **Channel fragments read as a journey** | Three good channel scores for a job that took three attempts | Episode-level reconstruction with the grouping rule stated |
| **Unmatched records ignored** | A linked analysis with no match rate reported | Match rate, matching rule and unmatched comparison required |
| **Benchmark of unknown provenance** | An external score comparison with no methodology behind it | Benchmarks used only where the method matches; otherwise not used |
| **AI: inventing the emotional arc** | Rich descriptions of how customers feel at stages with no data | Every stage claim carries an evidence reference or is marked unmeasured |
| **AI: composing a plausible verbatim for a stage** | A quote that captures a stage perfectly, with no participant ID | **K4 §2.3**; every quote carries an identifier |
| **AI: over-smoothing the severity index** | A single priority score with no visible weighting | Dimensions reported separately; any weighting is a stated human judgement |

## 12. AI guardrails

Universal prohibitions are inherited from **K4**. **K4 §3.2** (never convert correlation into causation) and **K4 §4.3** (never bury a limitation) govern this skill in full. The following are specific.

1. **Never populate a journey stage for which no evidence exists.** An unmeasured stage is shown as unmeasured, including in a visual journey map where the blank looks awkward.
2. **Never describe a touchpoint as driving, improving or causing an overall evaluation** where the association comes from a single self-report survey. It is a correlate, and the common-method issue is named alongside it.
3. **Never merge touchpoint-level and relationship-level measures into a composite**, and never compare one against the other as though a difference between them were a change over time.
4. **Never rank pain points by mention volume**, and never present a severity ranking without its three dimensions visible.
5. **Never present a dissatisfaction ranking as a churn-driver ranking** without outcome data joined and the time window stated.
6. **Never report an experience level or trend without the survivorship statement** where the population surveyed is retained customers.
7. **Never report a linked analysis without the match rate**, and never assume the unmatched are equivalent to the matched.
8. **Never present the operational record as ground truth for resolution.** Where the operational definition and the customer's differ, both are reported.
9. **Never quote an external experience benchmark whose collection method is unknown**, per **K4 §2.4**.
10. **Never describe a proprietary experience metric by its brand name or treat it as a summary of the experience.** Describe what it asks, on what scale, of whom.
11. **Never proceed with an individual-level data join without the consent position established**, which is an ethics gate under **K5 §2.4**.

## 13. Best-practice principles

1. **The measurement inventory is the analysis.** Half the findings in a mature experience programme come from discovering what it does not measure and why, and no amount of analysis of the measured parts substitutes for that.
2. **The gaps between touchpoints are where dissatisfaction lives.** Organisations measure the moments they control and the customer experiences the waiting, the uncertainty and the second attempt. Design the journey map so those spaces are visible.
3. **Effort is more actionable than satisfaction.** Nobody knows how to make someone 8% more satisfied. Everybody knows how to remove a step, stop asking for information twice, or answer at the first attempt.
4. **Recoverability changes the response, not just the priority.** An unrecoverable failure needs prevention. A recoverable one needs detection and a recovery capability, and the detection gap is usually the bigger problem, because most failures are never reported.
5. **A satisfied customer who leaves is the most informative case in the dataset.** Programmes that link experience scores to subsequent behaviour routinely find that leavers scored adequately, which tells you the instrument is not measuring what predicts the outcome.
6. **Beware the improving score in a shrinking base.** Retention itself filters the sample, so an experience metric can improve mechanically as the dissatisfied leave. Always read a satisfaction trend alongside the churn trend.
7. **Report distributions, not averages, for anything episode-level.** The average number of contacts per job is a number nobody has. The tail, the jobs taking four or more contacts, is where the cost and the damage are.
8. **Segment by behaviour before demographics.** In experience work, what someone was trying to do and how it went is far more explanatory than who they are, and demographic cuts of experience data are usually a distraction.
9. **The verbatim explains what the score cannot.** A score records that something was bad and never why. Pair every quantitative pain point with its language, and preserve the participant identifiers (**K2 §4.2**).
10. **Never let a single headline number carry the programme.** Any one-item score compresses a complex experience to a point, is highly sensitive to wording, scale and timing, and moves for reasons that have nothing to do with the experience. Report it as one measure with its base and read it alongside operational reality.
11. **Time the measurement to the customer's sense of completion, not the organisation's.** A survey fired when the case was closed in the system, but before the customer received what they were waiting for, measures a different thing entirely, and this mismatch is extremely common.
12. **A journey finding that names no measurable check has not finished.** The last column of the output is what makes the analysis auditable six months later, when someone asks whether the fix worked.

## 14. Worked example

**INPUT**

A fictional regional utility, Cape Meridian Water, runs an experience programme. It measures a transactional satisfaction rating after every contact centre call (response rate 11%), a post-installation rating after new connections (response rate 19%), and a twice-yearly relationship survey with a single-item recommendation-likelihood question among all account holders (response rate 4%). The relationship measure has improved for three consecutive waves. Complaints to the regulator have risen 22% over the same period. The board asks why the two disagree.

**PROCESS**

*Step 1, inventory.* The three instruments are documented. Two use a five-point satisfaction scale with different anchors; the relationship survey uses an eleven-point likelihood scale. None of them can be compared to the others directly. The contact centre survey fires when the agent closes the case in the system, which internal notes reveal happens at the end of the call even where a follow-up action is pending.

*Step 2, coverage.* The customer job is defined as "getting a supply problem resolved". Mapped as customers describe it in the existing qualitative work, the job has seven stages: noticing the problem, finding out whether it is theirs or the utility's, reporting it, waiting for an update, the visit, confirming resolution, and the bill afterwards. The programme measures two of the seven: the reporting call and, for new connections only, the visit. Waiting for an update, confirming resolution and the bill are unmeasured. This is the first finding.

*Step 3, moment against relationship.* Contact centre satisfaction averages 4.1 of 5 and has been stable. The relationship measure has improved. This apparent agreement conceals the structural problem: the contact centre measure is taken from the 11% who respond, immediately after a call, about the call, and the relationship measure from the 4% who respond, about the company. Neither covers the waiting stage, which the qualitative work identifies as the dominant source of frustration.

*Step 9, linkage. The judgement call.* Account-level linkage between the relationship survey and the operational records is technically possible. Checking the consent wording, the relationship survey's privacy statement covers analysis of responses but does not mention joining to service records. Resolved by not performing the individual-level join, running the analysis at aggregate area level instead, which is weaker, and flagging the consent wording for revision so future waves can support the stronger analysis. This is a **K5 §2.4** ethics gate and it constrains the whole analysis.

*Step 11, survivorship, and the answer.* Account holders cannot leave: the utility is a monopoly supplier in its area. So the survivorship problem takes an unusual form. Nobody exits, but the 4% who respond to the relationship survey are self-selecting, and analysis of response rate by area shows it has fallen from 6% to 4% across the three waves, with the steepest falls in the two areas where regulator complaints rose most. **The improving relationship score is consistent with dissatisfied customers ceasing to respond rather than with improving experience.** This is stated as the most plausible reading, with the alternative (genuine improvement plus an unrelated response decline) acknowledged, and the test that would separate them named: a non-response follow-up in the two affected areas.

*Steps 5 and 6, severity.* Complaint records to the regulator are coded and the dominant category is failure to receive a promised update, at 41% of complaints. Operationally, the utility has no measurement of update delivery at all, so frequency cannot be estimated from its own data, only from complaints, which measures who complained. Reported as `[frequency not estimable from available data]` with the specific operational record that would supply it named.

**OUTPUT**

> The two measures disagree because they measure different things on different populations, and neither covers the stage the complaints concern. The programme measures two of the seven stages of the supply-problem journey as customers describe it. Waiting for an update, confirming resolution and the follow-up bill are unmeasured, and the largest single category of regulator complaint (41% of complaints in the period) concerns the first of these.

> **The improvement in the relationship measure should not be read as an improvement in experience.** Response to that survey fell from 6% to 4% across the three waves, with the steepest declines in the two areas where regulator complaints rose most. A rising score on a falling and increasingly self-selected base is consistent with dissatisfied customers ceasing to respond. This is the most likely reading; genuine improvement alongside an unrelated response decline remains possible. A non-response follow-up in the two affected areas would separate the two, and until it is run, the trend should not be reported as an improvement.

> The contact centre survey fires when the agent closes the case, which internal process notes indicate occurs at the end of the call even where a follow-up action is outstanding. It therefore measures satisfaction with the call, not with the resolution, which is consistent with a stable 4.1 of 5 alongside rising complaints about what happened next.

> **What could not be established.** The frequency of the update failure, because the utility does not record whether promised updates were delivered; complaint counts measure who complained. Any individual-level relationship between experience scores and service history, because the relationship survey's consent wording does not cover joining to operational records.

> **Fix and check.** Measure update delivery operationally (baseline: not currently recorded). Move the contact centre survey trigger to the point of customer-confirmed resolution rather than system closure (baseline: 4.1 of 5 measured at call end, not comparable after the change, so the series restarts). Run a non-response follow-up in the two areas (baseline: response 4%, falling).

**Researcher decision required.** Changing the contact centre survey trigger will break the existing time series. Whether the continuity of a measure that fires at the wrong moment is worth more than a measure that fires at the right one is a judgement about how the programme is used internally (**K5 §2.7**).

## 15. Advanced usage

**Sequence analysis on journey data.** Where interaction logs exist, journeys can be analysed as sequences rather than as sets of touchpoints, which reveals the common paths, the loops (the same step repeated) and the abandonment points that a touchpoint-level view cannot see. Clustering sequences into a small number of typical paths, and then comparing outcomes across paths, is one of the strongest available uses of operational data, and it needs the episode rule from Step 10 to be defined carefully because it determines what a sequence is.

**Modelling the relationship measure from touchpoint history.** Where individual-level linkage and consent exist, model the relationship score as a function of the touchpoints a customer actually experienced, with the operational record rather than the self-report as the predictor. This removes the common-method problem entirely and is the single largest available improvement in the credibility of driver-style findings. Hand the estimation to **05.04** and **05.06**, and keep the interpretive discipline here.

**Testing a fix rather than inferring it.** Where an improvement can be rolled out to some customers and not others, do that: it converts a correlational journey finding into a causal one and calibrates the whole programme's understanding of effect sizes. See **06.04**.

**Lapsed and rejected populations.** The most under-researched populations in experience work, and the only source of evidence about what the experience does when it fails badly. A modest study of lapsed customers is frequently worth more than another wave of the main programme, and it is the direct answer to the survivorship problem in Step 11.

**Combining journey analysis with large-scale text.** Where contact records, chat transcripts or reviews run into thousands, classify them against the journey stage frame rather than a generic topic list, which turns unstructured text into stage-level frequency evidence. Hand to **06.06 Large-Scale Text Analytics**, and require the classification accuracy to be reported alongside any stage frequency derived from it.

**Public service and healthcare journeys.** The method transfers directly, with two differences worth naming. The customer frequently cannot exit, so dissatisfaction accumulates rather than resolving through churn, and the survivorship problem inverts into a response-bias problem. And the consequences of a failed journey are often severe and unequal, which raises the ethical stakes of the analysis and of who is not in the sample. Both are **K5 §2.4** considerations.

## 16. Skill chain

**Recommended previous skills**
- **04.04 Data Transformation and Dataset Preparation.** Hands over the linked, episode-structured dataset that journey-level analysis depends on.
- **07.01 Thematic Analysis** and **07.03 Interview and Transcript Analysis.** Hand over the customer-defined journey stages and the mechanisms behind the scores, without which the map is the organisation's own process diagram.
- **05.04 Driver Analysis.** Hands over the association estimates between touchpoint measures and overall evaluation, which this skill re-labels and bounds.
- **06.06 Large-Scale Text Analytics.** Hands over stage-classified volumes from contact and review text at scale, with its accuracy figures attached.

**Recommended next skills**
- **08.05 Insight Prioritisation and Sizing.** Takes the severity table and sizes the opportunity, which requires the frequency and impact figures this skill produces.
- **08.04 Recommendation Development.** Takes the fix-and-check list into owned actions with the **K5 §2.5** sign-off.
- **06.04 Experiment and A/B Test Analysis.** Takes any prioritised fix that can be tested rather than assumed.

**Runs well alongside**
- **05.06 Correlation, Regression and Causal Claim Control**, which governs every association between a touchpoint and an outcome.
- **13.05 Research Ethics and Consent Design**, which governs the individual-level linkage in Step 9.
- **05.05 Trend and Tracker Analysis**, where experience measures are tracked and a movement needs distinguishing from response-composition drift.

---
A Yazi Supplied Skill and resource.
