---
name: driver-analysis
description: >
  Identifies which measured attributes are most strongly associated with an
  outcome such as satisfaction, loyalty, preference or intention, chooses an
  approach that survives correlated attributes, and reports the result as
  association rather than causation. Use for "what drives satisfaction", "run a
  key driver analysis", "which attributes matter most", "derived importance",
  "importance versus performance", "what should we fix first", "why does stated
  importance disagree with the model", or "which of these attributes actually
  moves loyalty".
category: 05 Quantitative Analysis
ref: "05.04"
tier: 1
inherits: [K2, K3, K4, K5]
---

# Driver Analysis

## 1. One-line description
Establishes which of a set of measured attributes are most strongly associated with an outcome, using an approach chosen to survive the correlation between attributes that is normal in survey data, and states plainly what the technique's own name obscures: **driver analysis does not identify drivers.** It identifies statistical associations, and the causal reading is not licensed by the design.

## 2. What this skill is used for

**The research problem it solves.** Every organisation running a satisfaction, loyalty or brand study eventually asks the same question: of the twenty things we measure, which ones matter? Answering it badly is easy and common. Ask respondents directly and they will tell you everything is important, that price and reliability lead, and nothing will discriminate. Correlate each attribute with the outcome one at a time and the attributes will rank in an order that reflects how correlated they are with each other as much as anything else. Run a regression on twenty attributes that all correlate at 0.6 with one another and the individual coefficients become unstable to the point of arbitrariness: drop one attribute and the ranking reorders, add a hundred more respondents and a coefficient changes sign. Then the ranked list goes into a deck under a heading with the word "drivers" in it, and a business spends money on attribute three because it is third. This skill supplies the choice of approach, the diagnostics that say whether the model can carry the question, the conditions that invalidate it, and the reporting language that stops an association becoming a causal instruction.

**Where it sits in the research lifecycle.** After descriptive analysis and cross-tabulation have established the figures, and after the attribute battery and the outcome measure exist in the dataset. Before findings become recommendations. It depends on questionnaire decisions made much earlier, which is why it fails most often for reasons that have nothing to do with the analysis.

**Typical use cases.**
- Ranking service or product attributes by their association with satisfaction or loyalty.
- Deciding where to focus improvement effort when everything cannot be improved at once.
- Comparing what respondents say matters against what the data associates with the outcome.
- Building an importance and performance view for a service, journey or product.
- Testing whether a hypothesised relationship between an attribute and an outcome holds once other attributes are accounted for.
- Auditing a supplied driver model to see whether its rankings are stable enough to act on.

**Who uses it.** Quantitative researchers and analysts running customer experience, brand and satisfaction studies; research directors reviewing a ranked attribute list before it goes to a client; client-side insight and CX teams deciding improvement priorities; academic and public sector researchers relating attribute batteries to outcome measures.

## 3. When to use it

- An outcome measure and a battery of attribute ratings exist in the same dataset, and the question is which attributes relate most strongly to the outcome.
- Improvement effort has to be prioritised and the organisation cannot act on everything.
- Stated importance has produced a flat or implausible ranking and something better is needed.
- A hypothesis exists that a particular attribute matters, and it needs checking against the others rather than in isolation.
- A supplied driver model needs auditing for stability, collinearity and overreach.
- The attribute set is genuinely comprehensive for the decision at hand, so an omitted-variable problem is unlikely to dominate.
- Someone is about to make a causal claim from an attribute ranking and the claim needs bringing back within what the design supports.

## 4. When NOT to use it

- **The outcome has too little variance to model.** If 91% of respondents are in the top two boxes of the outcome, there is almost nothing to explain, and a model will fit noise and rank it. The same applies to a floor. Check the outcome distribution first (**05.01**); where variance is inadequate, say so and report attribute performance descriptively instead. A ranked driver list built on a ceiling is confident, precise and meaningless.
- **The attribute set is incomplete for the decision.** A driver absent from the questionnaire cannot appear in the model, and its association will be absorbed by whichever measured attributes correlate with it. If price was not asked and price is what matters, the model will attribute price's effect to value perception, or to fairness, or to whatever proxy was measured. Where the omission is known and material, the analysis is not run; where it is suspected, it is disclosed with the finding. This is the failure mode with the largest consequences and the least visibility, because the output looks identical either way.
- **A causal claim is what is actually wanted, and a decision depends on it.** Driver analysis is cross-sectional and observational: attributes and outcome are measured at the same moment from the same respondent, so temporal order is unknown, reverse causation is unexcluded, and common-cause confounding is untreated. If the question is "will improving X raise Y", the answer needs a design that can support it. See **05.06 Correlation, Regression and Causal Claim Control** for the confounding structures and the licensing designs, and **06.04 Experiment and A/B Test Analysis** where an experiment is feasible.
- **The question is about trade-offs between levels of an attribute rather than between attributes.** How much more people will pay for a faster service, or which feature bundle wins, is a trade-off question requiring a design that forces choices. That is **06.01 Conjoint and MaxDiff Analysis**. A rating battery cannot answer it, because respondents rate everything as important when nothing is scarce.
- **The base is too small to support the number of attributes.** A model estimating twenty coefficients on 120 respondents will fit the sample rather than the population, and its rankings will not replicate. A common working minimum is at least ten to fifteen respondents per attribute for stable estimation, considerably more where attributes are highly correlated. Below that, report bivariate associations with their bases and say the model was not estimable.
- **The attributes are so highly correlated that individual importance is not identifiable.** Where a battery of twenty items is essentially measuring three underlying dimensions, allocating importance to individual items is arbitrary, and different defensible methods will produce different rankings on the same data. The honest response is to model the dimensions rather than the items, and to say that item-level ranking cannot be supported. See Step 5.
- **The attributes and the outcome are not independent measurements.** An overall satisfaction outcome regressed on an attribute battery that includes an overall satisfaction item, or on a composite built from the same items, produces a strong result that is arithmetic rather than evidence. Check what every variable is made of before modelling it.
- **The ranking will be read as a budget allocation.** Where the receiving organisation intends to divide spend in proportion to importance scores, the relative-importance numbers will be used as though they were elasticities, which they are not. Either supply the framing that prevents this, or do not supply the numbers. Per **K4 §9**, name what would be needed for the question being asked.

## 5. Required inputs

**Required. Without these the skill cannot run. If absent, ask; if no answer is available and work must proceed, state the assumption at the point where it bites, per K5 §5.**

- **A defined outcome variable**, with its exact wording, scale and direction. Which outcome is being modelled changes the answer entirely: drivers of satisfaction, of stated repurchase intention and of recommendation are three different models with three different rankings.
- **The attribute battery as fielded**, with wording, scale and direction for every item. Attribute rankings are extremely sensitive to wording, and an item that combines two ideas ("staff are friendly and knowledgeable") cannot be interpreted whichever way it ranks.
- **A respondent-level dataset** with the outcome and all attributes on the same respondents and the same base.
- **The base description and base size for the modelling sample**, after listwise or other missing-data treatment, which is usually smaller than the achieved sample.
- **The decision the analysis informs.** Determines which outcome to model, how granular the attribute set should be, and whether item-level or dimension-level results are useful.

**Optional, and what each one adds.**

- **Stated importance ratings for the same attributes.** Allows the derived-versus-stated comparison, which is one of the most informative outputs this skill produces, and which is not available if only one of the two was asked.
- **Performance ratings on a comparable scale.** Allow the importance and performance view, and the prioritisation that follows from it.
- **Behavioural or transactional outcome data** (retention, spend, repeat purchase) linked at respondent level. Replaces a stated outcome with an observed one, which removes a large part of the common-method problem described in Step 3.
- **Previous waves of the same model.** Establish whether a ranking is stable or moves every wave, which is the only way to know whether a reordering is a finding.
- **Competitor or category benchmark performance.** Turns a performance score into a relative position, which is what a prioritisation actually needs.
- **A theoretical or prior model of the relationships.** Allows a structural approach in which attributes sit at different levels rather than competing in one flat list.

## 6. Questions to ask before starting

1. **Which outcome is being modelled, and why that one?** Determines the whole analysis, and the choice is a business decision rather than an analytical one. *Default if unanswered:* model the outcome named in the objectives; if several exist, model each separately and report the differences between the rankings rather than merging them.
2. **Is the attribute set complete for this decision, in the client's own judgement?** Determines whether an omitted-variable disclosure is required and how strongly. *Default:* ask what is missing, list the known gaps in the output, and state that the model can only rank what was measured.
3. **What will be done with the ranking?** Determines the level of granularity, whether relative-importance numbers should be supplied at all, and how firmly the causal firewall needs stating. *Default:* assume it will be used to prioritise improvement effort, and write the reporting language accordingly.
4. **Do the attributes correlate heavily with one another, and does the client need item-level or dimension-level answers?** Determines the approach and whether individual ranking is defensible. *Default:* inspect the correlation structure, and if items collapse to a few dimensions, report dimensions with item detail beneath.
5. **Was stated importance asked as well?** Determines whether the derived-versus-stated comparison is available. *Default:* if only derived is available, say so, because a client who has seen stated importance elsewhere will otherwise treat the disagreement as an error.
6. **Are the outcome and the attributes measured from the same respondent at the same moment?** Determines the size of the common-method problem and the strength of the causal disclaimer. *Default:* assume yes, and disclose it.
7. **Is this a single reading or part of a tracked series?** Determines whether the model specification can be chosen freely or must match a previous one for comparability. *Default:* match the previous specification, and report any better alternative alongside rather than in place of it.

## 7. Step-by-step methodology

**Step 1. Name the outcome, and say why it is the outcome.** Write the outcome variable, its wording, its scale and the reason this is the thing worth explaining. Different outcomes produce different rankings from the same attribute battery, and the difference is substantive rather than technical: attributes associated with how people feel about a service are not the same as attributes associated with whether they stay. Where more than one outcome matters, model each and compare, because an attribute that ranks first on one and eighth on another is a genuine finding about the business. *Correct result:* one sentence naming the outcome and the decision it serves, and a stated position on any other outcome not modelled.

**Step 2. Inspect the outcome distribution before anything else.** A driver model explains variation, so if there is little variation there is nothing to explain. Look at the full distribution. A pile-up at the top, common in relationship satisfaction batteries, means the model is being asked to explain the small minority who are not satisfied, which may be legitimate but changes what the ranking means and usually makes it unstable. Bimodality means two populations and one model describing neither. *Correct result:* a written reading of the outcome distribution, and an explicit decision to proceed, to model a recoded outcome (for example, detractors against everyone else), or to stop.

**Step 3. Inspect the attribute battery: distributions, missingness, and what the items actually measure.** Check scale direction on every item against the questionnaire rather than the variable name (**K4 §6.2**). Check for double-barrelled items, which cannot be acted on whichever way they rank. Check missingness: listwise deletion across twenty attributes can remove a quarter of the sample and change who the model describes. Then look at the correlation matrix between attributes, which is the single most informative diagnostic in this skill and the one most often skipped. Survey attribute batteries typically correlate at 0.4 to 0.7 with each other because of halo: a respondent who likes the brand rates every attribute higher. **Common-method variance** compounds this, because attributes and outcome come from the same respondent in the same instrument at the same moment, which inflates every association in the model relative to the real world. *Correct result:* an attribute inventory with direction verified, double-barrelled items flagged, the modelling base after missing-data treatment stated, and the correlation structure described.

**Step 4. Compare stated importance against derived importance, and treat the disagreement as information.** Where both exist, plot or tabulate them together. They routinely disagree, and the disagreement is the finding rather than a problem to resolve. Stated importance reflects what people believe matters, what is socially acceptable to say matters, and what is salient at the moment of asking. It over-reports price, reliability and safety, and under-reports anything people do not like to admit responding to, anything they are not conscious of, and anything that only becomes visible when it fails. Derived importance reflects what covaries with the outcome in this sample, and it over-reports whatever is most correlated with the general halo. Four patterns are worth naming. **High stated, high derived:** stated priorities, and usually already managed. **Low stated, high derived:** latent or unarticulated, frequently the most useful cell, and the reason derived importance is run at all. **High stated, low derived:** table stakes, where performance is uniformly adequate so the attribute no longer discriminates, which is not the same as the attribute not mattering. **Low stated, low derived:** genuinely peripheral, or measured badly. *Correct result:* a two-way comparison with each attribute placed, and a written note on the table-stakes cell, since a hygiene factor with no variance will look unimportant precisely because it is being done well everywhere.

**Step 5. Choose the approach from the correlation structure and the question, and state its assumptions.** Four families, described generically.

- **Correlation-based.** Each attribute's association with the outcome, one at a time. Transparent, needs no model, and is the honest floor. Its weakness is that it ignores the relationships between attributes entirely, so correlated attributes all score high together and nothing is separated. Useful as a first look and as a check on the other methods; never a defensible final answer on its own where attributes are correlated.
- **Regression-based.** All attributes entered together, so each coefficient reflects the association with the outcome that is not shared with the others. This is what "controlling for" means, and it is the natural answer to the question being asked. Its weakness is exactly the condition that is normal in this data: where attributes correlate strongly, the coefficients become unstable, standard errors inflate, signs can flip, and the split of shared association between two correlated attributes is close to arbitrary. Stepwise selection is not a remedy and makes it worse, because it selects on the same unstable estimates and produces a set whose composition is a function of noise.
- **Relative-importance methods.** Built specifically for the correlated case. They allocate the model's explained variance across attributes by averaging each attribute's contribution over all possible orders of entry, or by an equivalent orthogonal transformation of the predictors. The result is a set of non-negative shares summing to the model's explained variance, which is far more stable than raw coefficients and is directly interpretable as a share of what the model explains. Two things must be said whenever they are used: the shares are a decomposition of association within this model, not a measure of effect on the outcome; and the averaging makes the shares stable without making the underlying attribution any less shared. Where two attributes are near-duplicates, the method will split their common association between them evenly, which is defensible arithmetic and is not a discovery that each contributes half.
- **Structural approaches.** Where a theory exists about how the attributes relate to each other and to the outcome (touchpoint attributes feeding perceptions, perceptions feeding overall evaluation, overall evaluation feeding intention), a path or latent-variable model represents that structure rather than flattening it. It handles measurement error in the attributes, distinguishes direct from indirect association, and answers questions a flat model cannot. Its cost is that it requires the structure to be specified in advance and it will fit a wrong structure without complaint, so the model is only as good as the theory behind it. Fit statistics indicate consistency with the data; they do not confirm the causal ordering assumed.

Choose on two criteria: how correlated the attributes are, and whether a defensible structure exists. *Correct result:* a named approach with its assumptions written down and the reason it was chosen over the alternatives.

**Step 6. Diagnose multicollinearity explicitly, and act on what you find.** Report the correlation matrix, and a variance-inflation measure per attribute. Rules of thumb vary and none is a law: values above about 5 warrant attention and above about 10 are usually taken as serious, but the substantive test is better than the threshold. **Perturb the model and see whether the ranking survives.** Re-estimate on bootstrap resamples and report how often each attribute lands in the top three; drop one attribute and see whether the rest reorder; split the sample in half and compare rankings. If the top three attributes are the top three in 90% of resamples, the ranking is usable. If attribute two is anywhere between second and ninth depending on the resample, **the ranking is not a finding and must not be presented as an ordered list**, however precise the point estimates look. Where collinearity is severe, the options are to group correlated items into dimensions on a substantive basis and model those, to use a relative-importance method and report shares rather than an order, or to report that individual attribute importance is not identifiable in this battery. *Correct result:* a stability statement attached to the ranking, in the form of how often each attribute holds its position, not an unqualified ordered list.

**Step 7. Report model fit honestly, and interpret it correctly.** State the share of variance in the outcome that the model accounts for, the modelling base after missing-data treatment, and the attribute set. Then read it properly. A model accounting for 55% of variance in a stated outcome from attributes measured in the same questionnaire is not evidence that the business understands 55% of what drives its customers; a substantial part of it is common-method variance and halo. A low figure is not automatically a failure: it may mean the outcome is driven by things not measured, which is itself a finding worth reporting. What matters most is what the fit implies about the omitted attributes, which is the largest uncertainty in the whole exercise and the one no diagnostic can size. *Correct result:* fit reported with the base and the attribute set, plus one line on what the unexplained share is likely to contain.

**Step 8. Build the importance and performance view, and use it correctly.** Plot derived importance against current performance, with the axes crossed at meaningful reference points (the mean of each, or an agreed benchmark, stated either way). The four regions read as: **high importance, low performance,** the priority region; **high importance, high performance,** protect and do not disturb; **low importance, low performance,** leave alone; **low importance, high performance,** possible over-investment, and the most frequently misread region. Three cautions govern its use. The framing invites the conclusion that moving an attribute from low to high performance will move the outcome by an amount proportional to its importance score, which the design does not support at all. Attributes near the crossing point are not distinguishable from each other, so the quadrant boundary must not be drawn as though it were a decision line. And a table-stakes attribute performing uniformly well will sit in the low-importance region for that very reason, so an instruction to divest from it can be catastrophic. *Correct result:* a quadrant view with reference points stated, boundary uncertainty shown, and hygiene attributes annotated rather than left to be read off.

**Step 9. Apply the causal firewall to the draft.** Read the output as a hostile reader would and remove every construction that asserts what the design cannot support. "Drives", "leads to", "improving X will raise Y", "the impact of X on Y" and "because customers found X, they did Y" are all prohibited (**K4 §3.2**). What is permitted: "X is the attribute most strongly associated with Y in this model", "where X is rated higher, Y tends to be higher", "X accounts for the largest share of the variance the model explains". Where the conventional output uses the word driver, one statement appears in the same document saying that these are statistical associations and that the causal direction has not been established by this design. Reverse causation is specifically live here: people who are satisfied overall rate every attribute more favourably, so a model of attribute ratings on overall satisfaction is partly measuring the halo running the other way. Say so. *Correct result:* a draft in which no sentence claims more than association, and the association statement appears once, prominently, rather than in a footnote.

**Step 10. Say what the model could not see, and hand over.** List the attributes not measured that a reasonable person would expect to matter, the outcome variance left unexplained, the stability of the ranking, the base, and the fact that everything was measured at one moment from one respondent. Then hand the ranking to interpretation with its qualifications attached. *Correct result:* a reader who receives only the ranking still knows what it is a ranking of, how firm it is, and what it does not license.

## 8. Analytical framework

    Outcome → Variance → Attribute set → Correlation structure → Approach → Stability → Association → Priority

**Outcome.** Named, with the decision it serves and the reason it was chosen over other outcomes.
**Variance.** Whether there is enough variation in the outcome to explain at all.
**Attribute set.** What was measured, and, explicitly, what was not.
**Correlation structure.** How much the attributes share, which determines whether item-level importance is identifiable.
**Approach.** Chosen from the structure and the question, with its assumptions stated.
**Stability.** How often the ranking survives perturbation. An unstable ranking is not a finding.
**Association.** The result, stated in associational language, with fit and base.
**Priority.** Importance set against performance, with the boundary uncertainty visible.

Against the **K2** evidence chain, this skill produces Analysis and Findings. The step to "therefore fix this first" is Implication, and it crosses two gaps the model does not cover: from association to causation, and from causation to what this organisation can actually change. Both belong outside this skill, the first to **05.06** and a design that licenses it, the second to **08.03 Implication Development** and a human.

## 9. Output format

**1. Model specification block.** Outcome variable and wording; attribute list with wording; modelling base after missing-data treatment, with the achieved sample for comparison; approach used and why; software-independent statement of the estimation; weighting status.

**2. Attribute inventory.**

| Attribute | Wording | Scale and direction | Mean or top-box | % missing | Correlation with outcome | VIF |
|---|---|---|---|---|---|---|

**3. Derived importance table.**

| Rank | Attribute | Importance measure (state which) | Share of explained variance | Stability: % of resamples in top 3 | Performance score | Base |
|---|---|---|---|---|---|---|

**4. Stated versus derived comparison**, with the four-cell reading and the table-stakes attributes annotated.

**5. Importance and performance view**, with axis reference points stated and boundary uncertainty shown.

**6. Model fit and interpretation note.** Share of outcome variance accounted for, base, attribute set, and one line on what the unexplained share is likely to contain.

**7. The association statement.** One paragraph, prominent, not a footnote, saying that these are statistical associations from a cross-sectional design, that attributes and outcome were measured from the same respondent at the same moment, that reverse causation and common-cause confounding are unexcluded, and that no attribute has been shown to cause any movement in the outcome.

**8. What the model could not see.** Attributes not measured, the omitted-variable risk, the unexplained variance, and any subgroup for which the model was not estimated.

**Where the evidence is thin**, the format does not get filled. An unstable ranking is reported as a grouped set ("these four attributes are consistently in the leading group; their order is not determinable at this base") rather than as an ordered list with ranks. Where the outcome lacks variance, the importance table is not produced at all, and the output says why. Where an attribute battery collapses to three dimensions, item-level ranks are omitted and dimension-level results reported. Per **K4 §1**, a column headed "rank" is not evidence that a rank exists.

## 10. Quality checks

Run before any driver result is presented. **K4 §8** runs anyway; these are specific to this task.

1. Was the outcome distribution examined, and is there enough variance to model?
2. Is the modelling base stated after missing-data treatment, alongside the achieved sample?
3. Has scale direction been verified against the questionnaire for the outcome and every attribute?
4. Is any attribute double-barrelled, and if so is it flagged as uninterpretable whichever way it ranks?
5. Is the correlation matrix between attributes reported, with a collinearity measure per attribute?
6. Has the ranking been perturbed (resampled, split-half, or attribute-dropped) and is the stability result reported alongside it?
7. Is any ordered rank presented that the stability check does not support?
8. Is model fit reported with the base and the attribute set, rather than a bare percentage?
9. Does the output state which attributes were not measured, and name the omitted-variable risk?
10. Does every importance statement use associational language, with no instance of "drives", "leads to", "impact of" or "improving X will"?
11. Does the association statement appear once, prominently, in the same document as the ranking?
12. Where stated importance exists, is the disagreement with derived importance reported as informative rather than reconciled away?
13. Are table-stakes attributes annotated on the importance and performance view rather than left to read as low priority?
14. Is the possibility of reverse causation (halo from the outcome onto the attribute ratings) stated?
15. If the same attribute battery was modelled against more than one outcome, are the differing rankings reported rather than merged?

## 11. Common failure modes

| Failure | How to recognise it | How to prevent it |
|---|---|---|
| **The name taken literally** | A deck headed "key drivers" recommending investment on the basis that attribute three is third | Step 9 firewall; the association statement in the same document; associational language throughout |
| **Unstable ranking presented as ordered** | Precise ranks with no stability information, on a battery whose attributes correlate above 0.6 | Step 6 perturbation; grouped sets rather than ranks where stability fails |
| **Omitted driver absorbed** | A model in which a proxy attribute ranks implausibly high, and the obvious real cause was never asked | Step 3 attribute audit; explicit "what was not measured" section; disclose rather than model where a material attribute is missing |
| **Ceiling outcome modelled anyway** | An importance ranking built on an outcome where 90% are in the top boxes | Step 2 distribution check; recode the outcome or stop |
| **Stated and derived reconciled away** | One of the two quietly dropped because they disagreed | Step 4; the disagreement is reported as the finding |
| **Table stakes read as unimportant** | A recommendation to reduce investment in an attribute everyone performs well on | Hygiene annotation on the quadrant; the low-variance explanation stated beside the low importance score |
| **Stepwise selection** | An attribute set chosen by the algorithm, differing between waves with no substantive reason | Enter the specified set; selection on unstable estimates compounds the instability |
| **Fit read as understanding** | "The model explains 61% of satisfaction" presented as a statement about the business | Step 7; common-method variance named; unexplained share described |
| **Quadrant boundary as decision line** | Two attributes on opposite sides of the mean treated as different priorities when they are 0.1 apart | Reference points stated, boundary uncertainty shown, near-boundary attributes grouped |
| **Same-source inflation ignored** | Associations reported as though attributes and outcome were independently measured | Disclosed in the association statement; behavioural outcome used where available |
| **AI: producing a ranking without the model** | An ordered attribute list with no coefficients, no fit, no base and no method named | Nothing enters the table without the estimation that produced it; per **K4 §2.1** an unproducible number is `[not available]` |
| **AI: causal language reappearing in the summary** | Associational language in the analysis, "drives" in the executive summary | Firewall applied to every downstream document, not only the analysis; caveats travel (**K3 §7**) |
| **AI: importance shares read as elasticities** | "A 1-point improvement in X yields a 0.3-point rise in Y" from a variance decomposition | Shares are a decomposition of association within the model, and this is stated wherever they appear |

## 12. AI guardrails

Universal prohibitions are inherited from **K4** and are not repeated here. **K4 §3.2** governs this skill in full, including its specific provision that driver analysis produces correlates, not drivers, whatever it is called.

1. **Never use causal language about a driver result.** No "drives", "leads to", "causes", "the impact of", "improving X will increase Y", "because customers experienced X". The permitted constructions are association, covariation and share of explained variance. This applies to headings and chart titles as much as to prose, and a heading is where the violation usually survives.
2. **Never present an ordered ranking without a stability result.** If perturbation reorders the list, report grouped sets and say the order is not determinable.
3. **Never report importance without the model fit, the modelling base and the attribute set** in the same place. An importance score detached from what was in the model is uninterpretable.
4. **Never omit the statement of what was not measured.** A driver absent from the questionnaire cannot appear in the model, and a ranking that does not say what it could not have found is misleading by construction.
5. **Never reconcile stated and derived importance into a single number.** They measure different things. Report both and the pattern between them.
6. **Never treat an importance share as a rate of return.** A variance decomposition does not say what happens to the outcome if the attribute moves, and must never be converted into a projected gain.
7. **Never run a driver model on an outcome without checking its variance**, and never present a ranking built on a ceiling or floor outcome without saying so.
8. **Never select attributes by an automated procedure and present the survivors as the important ones.** Enter the specified set; if a reduction is necessary, do it on substantive grounds and say what was removed.
9. **Never model an outcome against an attribute set that contains the outcome, a component of it, or a composite built from the same items.**
10. **Never present a quadrant position as a decision** where the attribute sits close to the reference line. Group near-boundary attributes and say they are not distinguishable.
11. **Never carry a driver ranking into a recommendation without a human.** What the organisation can actually change, at what cost, is a **K5 §2.1** and **K5 §2.3** judgement the model contains no information about.

## 13. Best-practice principles

1. **The technique is misnamed, and the misnaming is the main risk.** Everything else in this skill is ordinary analysis. The word "driver" does the damage, because it arrives in the client's language already meaning causation. Say what the analysis is once, clearly, and then stop apologising for it.
2. **The model can only rank what was asked.** The most consequential decision in a driver analysis was made months earlier, when the attribute battery was written. An analyst inheriting a battery inherits its blind spots and cannot detect them from the data.
3. **Stability matters more than order.** A ranking that reorders on a bootstrap resample is not a finding. Clients act on ranks, so the burden is on the analyst to establish that the ranks exist.
4. **The disagreement between stated and derived is the output, not an obstacle.** Where they agree you have confirmation, where they disagree you have something to explain, and the explanation is usually more useful than either ranking alone.
5. **A hygiene factor looks unimportant because it is working.** Uniformly good performance destroys the variance a model needs to detect an attribute, so the attributes most dangerous to neglect systematically rank lowest. Annotate them every time.
6. **Correlated attributes share their association, and no method can unshare it.** Relative-importance methods make the split stable and defensible; they do not make it real. Two near-duplicate items that each receive half the importance have not each been shown to matter half as much.
7. **Halo runs both ways.** People who are satisfied rate everything higher, so part of every attribute-to-outcome association is the outcome colouring the attribute rating. This is not a small effect and it cannot be removed by any analysis on the same data.
8. **A model that explains little may be telling the truth.** Low explained variance usually means the outcome depends on things the questionnaire did not measure, which is a real finding about the study rather than a defect in the analysis.
9. **Model each outcome separately and compare.** An attribute that ranks first for satisfaction and eighth for retention is telling the business something important, and averaging the two rankings destroys it.
10. **Prefer an observed outcome to a stated one wherever it can be linked.** Modelling attributes against actual retention or spend removes a large part of the common-method problem and changes the rankings more often than analysts expect.
11. **Report the group, not the rank, when the base is modest.** "These four are consistently leading" is honest and usable. "1, 2, 3, 4" on a base of 200 with correlated attributes is neither.
12. **Structure beats a flat list when a structure exists.** Attributes are not usually parallel competitors: some are components of others, and some operate through others. A model that says so is more useful than one that ranks fifteen items in a line.

## 14. Worked example

**INPUT**

A fictional retail bank, Marrow Savings, runs an annual customer study with 1,850 respondents. It measures overall satisfaction on an 11-point scale, stated importance and performance on 16 service attributes, and stated likelihood to remain a customer. The brief is: "run a key driver analysis and tell us which three things to fix so satisfaction goes up".

**PROCESS**

*Step 1, outcome.* Two candidate outcomes exist. The brief asks about satisfaction; the business decision is about attrition. Both are modelled separately, and the difference between the two rankings is expected to be the substantive output.

*Step 2, variance.* Overall satisfaction has a workable spread with a moderate top-end concentration. Stated likelihood to remain is severely piled at the top: 84% give the top two points. Modelled as-is it would be explaining an 16% minority. Recoded to a binary of "anything other than the top two boxes" and modelled as such, with the recode disclosed and the reason stated.

*Step 3, attributes.* Scale direction verified against the questionnaire; two items are reverse-worded and were correctly coded. One item, "staff are helpful and well informed", is double-barrelled and flagged: if it ranks high the bank will not know which half to act on. Listwise deletion across 16 attributes removes 213 respondents, taking the modelling base from 1,850 to 1,637, and the removed group skews toward the least engaged customers, which is disclosed. The attribute correlation matrix runs from 0.38 to 0.79, with a cluster of five branch-experience items correlating above 0.7 with each other.

*Step 4, stated versus derived.* Stated importance ranks security first, fairness of charges second and branch staff friendliness fourteenth. Derived importance puts security eleventh. **This is the table-stakes pattern:** security performance is 9.1 out of 10 with almost no variance, so it cannot covary with anything. The output says that security ranks low in the model because it is uniformly well delivered, not because it does not matter, and that a decline in it would be expected to matter a great deal.

*Step 5, approach.* Bivariate correlations are computed as a floor and put nine attributes within 0.05 of each other, which separates nothing. A plain regression is estimated and shows the expected symptom: two of the five branch items take coefficients of opposite sign, one of them substantial. A relative-importance method is used for the reported result, and the five branch items are additionally modelled as one dimension for a second view.

*Step 6, stability, and the judgement call.* Bootstrap resampling shows "problems resolved first time" in the top three in 96% of resamples and "digital service works when I need it" in 91%. The third position is contested: three attributes, all from the branch cluster, occupy it between them, none holding it in more than 44% of resamples. **Judgement call:** the client asked for three things to fix. Reporting a third rank would be answering the question asked and would not be supportable. Resolution: report two attributes as individually identified, report the branch cluster as a single third leading factor at the dimension level with its five components listed and no order between them, and state explicitly that item-level ranking within the branch cluster is not determinable at this base with these correlations. The client receives three priorities, one of which is a dimension rather than an item, and the reason is on the page.

*Step 7, fit.* The model accounts for 48% of the variance in overall satisfaction on a base of 1,637. Reported with the note that a substantial part of that reflects the general favourability running through all ratings from the same respondent, and that the remaining 52% includes attributes not measured, of which the client's own list names two: mortgage decision speed and fee transparency at the point of sale.

*The two outcomes compared.* Against satisfaction, branch experience leads. Against the recoded retention outcome, "problems resolved first time" leads by a wide margin and branch experience falls to sixth. This is reported as the most useful finding in the study: what makes customers feel good about the bank and what keeps them are not the same, and a programme built on the satisfaction ranking would spend most of its money on the wrong thing.

*Step 9, firewall.* The draft headline reads "branch experience drives satisfaction". Rewritten to "branch experience is the factor most strongly associated with overall satisfaction in this model". The association statement is placed on the same page as the ranking, not in the appendix.

**OUTPUT**

A model specification block; an attribute inventory with directions verified, one double-barrelled item flagged and collinearity reported; two derived importance tables, one per outcome, with stability percentages against each attribute; a stated-versus-derived comparison with security annotated as table stakes; an importance and performance view with the reference points stated and the three contested branch items shown as a single grouped marker; a fit note; the association statement; and a list of what the model could not see.

**Researcher decision required.** Whether the improvement programme is built on the satisfaction ranking or the retention ranking is a business judgement about what the bank is trying to achieve, and the two produce materially different priorities (**K5 §2.1**, **K5 §2.3**). The analysis cannot make this choice and should not appear to.

## 15. Advanced usage

**Segment-level driver models.** Running the model separately by segment frequently shows that different attributes lead in different groups, which is more actionable than a single ranking. It also multiplies the estimation problem: each model needs its own base, and comparing two rankings across segments is itself a comparison requiring care. Do not conclude that an attribute matters more in one segment because it ranked higher there; that is an interaction claim and needs testing as one (**05.02**, Advanced usage).

**Non-linear and asymmetric relationships.** Many service attributes are asymmetric: poor performance damages the outcome far more than excellent performance improves it. A linear model averages the two and reports a middling association for an attribute whose real behaviour is a cliff on one side. Test by splitting the attribute into performance bands and looking at the outcome across them, or by modelling the negative and positive deviations separately. Where asymmetry is present, the improvement implication changes completely, from "raise this" to "never let this fail".

**Dimension reduction before modelling.** Where a battery clearly measures a few underlying dimensions, reducing the items to dimension scores on a substantive basis, checked against the correlation structure, produces a model that is far more stable and answers a question the business can act on. The cost is that the answer is at dimension level, which must be said, and the reduction rule must be documented (**04.04**).

**Structural models with mediation.** Where touchpoint attributes are believed to work through an intermediate perception, a path model estimates direct and indirect associations separately, which a flat model cannot. This is genuinely more informative and genuinely more assumption-laden, because the paths are specified rather than discovered, and a fitting model does not confirm the direction assumed. Mediation and its confounding structures are treated in **05.06**.

**Longitudinal and linked-behaviour extensions.** Where the same respondents are measured across waves, or attribute ratings can be linked to subsequent observed behaviour, temporal order becomes establishable and the analysis moves closer to something a causal claim could rest on. It is not automatic: confounders must still be addressed. The conditions are in **05.06**, and the design conversation belongs to **01.04 Research Method Selection**.

**When the standard approach does not fit.** Where the attribute set is small and highly correlated, no method will separate the items and the honest deliverable is a dimension-level result plus a statement that the battery cannot support item-level prioritisation. Where the outcome is rare (a small churn group), consider modelling the binary outcome directly rather than a continuous proxy, and report on a scale the client can read.

## 16. Skill chain

**Recommended previous skills**
- **05.01 Descriptive Analysis.** Hands over the outcome and attribute distributions that determine whether the model is estimable at all, and the base definitions the modelling sample is drawn from.
- **05.03 Cross-Tabulation.** Hands over the subgroup structure and the bivariate picture, which is the first check on whether a modelled association behaves the way the tables do.
- **02.07 Scale and Measurement Selection.** Determines, long before analysis, whether the attribute battery and outcome scales can carry a driver model at all.
- **04.03 Missing Data Handling.** Determines the modelling base, since listwise deletion across a long battery can remove a substantial and non-random part of the sample.

**Recommended next skills**
- **05.06 Correlation, Regression and Causal Claim Control.** Takes the causal question this skill refuses, supplies the confounding structures, and names the designs that would license the claim the client wants.
- **08.01 Finding to Insight Development.** Takes the associations and does the interpretation work, including the step from "associated with" to a business explanation, which is where **K2 §2.1** applies.
- **08.05 Insight Prioritisation and Sizing.** Takes the importance and performance view and adds the cost, feasibility and reach information the model does not contain.

**Runs well alongside**
- **05.05 Trend and Tracker Analysis**, wherever a driver model is repeated across waves and a reordering has to be judged against the model's own instability before it is read as change.
- **06.01 Conjoint and MaxDiff Analysis**, where the real question is a trade-off between levels rather than a ranking of attributes.
- **06.04 Experiment and A/B Test Analysis**, which supplies the design that can answer the causal question this skill is repeatedly asked.
- **13.03 AI Output Verification**, which audits a finished deck for causal language reintroduced downstream of the analysis.

---
A Yazi Supplied Skill and resource.
