---
name: conjoint-and-maxdiff-analysis
description: >
  Analyses and interprets trade-off studies: conjoint and discrete choice
  utilities, derived attribute importance, share-of-preference simulation, and
  MaxDiff scores. Audits the design that produced the data before trusting any
  estimate, and holds the line that a trade-off model predicts choices among the
  alternatives shown, under the conditions shown, and nothing else. Use for "run
  the conjoint", "what are the utilities", "which attribute is most important",
  "simulate the share for this product", "what price should we charge from the
  conjoint", "rank these messages with MaxDiff", "why does importance change when
  we change the levels", or "can we use this as a volume forecast".
category: 06 Specialist and Advanced Analysis
ref: "06.01"
tier: 2
inherits: [K2, K3, K4, K5]
---

# Conjoint and MaxDiff Analysis

## 1. One-line description
Estimates and interprets preference models from trade-off data, converts them into importance and simulated share, and refuses the readings the technique invites but does not support, above all the reading of a simulated share as a market forecast.

## 2. What this skill is used for

**The research problem it solves.** Direct questioning about preference produces answers that are almost useless for decisions, because respondents rate everything desirable as important and everything expensive as too expensive. Trade-off techniques solve this by forcing a choice: to have the better warranty, you must accept the higher price. The resulting model is genuinely powerful, and it is also the most over-read output in commercial research. Four failures recur. Utilities are compared across attributes as though they were on a common scale. Importance is quoted as a property of the attribute rather than of the levels the designer happened to pick. A share of preference is presented to a board as a volume forecast with a currency sign attached. And a design that could never have supported the question, because the levels were unbalanced, the prohibitions gutted the efficiency, or respondents were asked to do twenty-eight tasks and stopped reading at task nine, is analysed anyway because the data exists. This skill supplies the design audit that comes before estimation, the correct reading of each output, and the specific language that keeps a simulation inside its evidence.

**Where it sits in the research lifecycle.** After fieldwork and data preparation, when a choice-based or trade-off dataset is in hand. It assumes the design decisions have already been made and often badly, so a large part of the work is establishing what the design can carry. It hands finished, correctly caveated outputs to insight development and to pricing work.

**Typical use cases.**
- Estimating part-worth utilities from a choice-based conjoint and deriving attribute importance.
- Building a preference simulator and running product configurations through it.
- Assessing which features earn their cost in a product specification.
- Testing a portfolio or line-up rather than a single product, including cannibalisation between own items.
- Ranking a long list of messages, features or claims by MaxDiff when a rating scale would flatten them.
- Auditing a trade-off study someone else ran, before its conclusions are acted on.
- Telling a stakeholder why the conjoint cannot answer the question they are asking.

**Who uses it.** Quantitative research managers and directors running or commissioning trade-off work; insight managers client-side receiving a simulator and needing to know what it does and does not license; product and pricing analysts; UX researchers prioritising features; academic researchers using discrete choice experiments in health, transport or environmental valuation, where the same estimation applies under stricter reporting conventions.

## 3. When to use it

- A choice-based conjoint, adaptive conjoint, full-profile ratings conjoint or discrete choice experiment has been fielded and needs estimating and interpreting.
- A MaxDiff exercise has been fielded and needs scoring, and someone is about to read the scores as absolute importance.
- A simulator exists and configurations need to be run through it, with the results correctly bounded.
- Someone is asking which attribute matters most, and the honest answer depends on levels they have not been told about.
- A trade-off study is being used to support a pricing, feature or portfolio decision and the evidence needs auditing before the decision.
- Two waves of a trade-off study need comparing, and someone should check the designs are comparable before the utilities are.
- A finished conjoint deck contains a share number that has quietly become a forecast, and it needs correcting.

## 4. When NOT to use it

- **The design was not a trade-off design.** A battery of importance ratings, a set of monadic concept scores, or a "rank these nine features" exercise does not produce part-worths and cannot be estimated as though it did. Analyse it as what it is, with **05.01 Descriptive Analysis** and **05.03 Cross-Tabulation**. Retro-fitting conjoint machinery onto rating data produces numbers with no referent.
- **The question is about absolute demand at a price rather than relative preference among shown alternatives.** A conjoint tells you how choice moves as attributes change within the set displayed. It does not tell you how many people will buy, because it never observed a purchase and never presented the full market. Where the question is willingness to pay in level terms, or a demand curve, hand to **06.02 Pricing Research Analysis**, which owns the honest framing of stated-preference pricing, and which will hand the choice-based part back here.
- **The attribute set omits something that dominates real choice.** If availability, brand trust, the salesperson, the existing contract or the installation hassle drives the category and none of it is in the design, the model will attribute that variance to whatever is in the design. The output will be internally consistent and externally wrong. The correct response is to report the omission as a limitation on every derived number, and where the omission is central, to decline the simulation entirely.
- **The design is too degraded to estimate.** Where prohibitions have removed so many combinations that attributes are partially confounded, where a level appears in a handful of tasks, or where the design was not generated to be efficient and cannot be reconstructed, estimation will produce coefficients whose standard errors are not interpretable. Report the design failure. Do not produce utilities with a note.
- **Respondent effort collapsed.** Where a material share of respondents show non-trading (always choosing on one attribute), straightlining across MaxDiff sets, completion times implying they did not read, or a fit statistic indistinguishable from random choice, the aggregate model is being partly fitted to noise. Diagnose first, decide on exclusions with a human (**K5 §2.7**), and if the collapse is widespread, the study failed and should be reported as having failed.
- **The claim being sought is causal in the real world rather than within the experiment.** A conjoint is an experiment about stated choices among hypothetical profiles. It licenses causal statements about the effect of a shown attribute level on a stated choice inside the exercise. It does not license "reducing delivery time to two days will increase our sales by 6%". That step crosses from the experiment into the market and is not supported. See **05.06 Correlation, Regression and Causal Claim Control**, and **06.04 Experiment and A/B Test Analysis** for what a real-world causal claim requires.
- **A MaxDiff is being asked whether anything on the list actually matters.** MaxDiff produces a relative ordering of the items supplied. Every item can be trivial, or every item can be critical, and the scores look identical in both cases. If the question is absolute importance or a go/no-go threshold, either an anchored variant must have been fielded or a different instrument is needed. Do not answer the absolute question from unanchored data.
- **The result is wanted to confirm a decision already taken.** Where the configuration to be recommended is known in advance and the simulator is being run to produce a supporting number, the analysis has become advocacy. Report what the model says about the full sensitivity range, not the point that supports the case, per **K4 §4.2**.

## 5. Required inputs

**Required. Without these the skill cannot run. If absent, stop and ask.**
- **The design specification.** The attribute list, every level of every attribute, the design type (choice-based, adaptive, full-profile, MaxDiff), the number of tasks per respondent, the number of alternatives per task, whether a None or opt-out was offered, and any prohibited combinations. Without this, nothing in the data can be interpreted, because a utility has no meaning without the level it belongs to.
- **The respondent-level design and response data.** Which profiles each respondent saw in each task, in the order shown, and what they chose. Aggregate choice counts are not sufficient for individual-level estimation and are barely sufficient for anything else.
- **The base description and sample definition.** Who was eligible, how they were recruited, what quotas applied. A trade-off model is a model of the preferences of the people in it.
- **The exact wording shown to respondents.** The attribute descriptions, the level wording, the choice instruction. Level wording is data. "2 year warranty" and "2 year warranty, parts and labour, no claim limit" are different levels and will produce different utilities.

**Optional, and what each one adds.**
- **Holdout tasks not used in estimation.** The single most valuable optional input. They convert internal fit into a check of predictive accuracy, and they are the only honest evidence that the model predicts anything at all.
- **A stated competitive frame or market structure.** Lets the simulator be built with real competitors at real specifications rather than as an abstract set of profiles, which changes every share number the simulator produces.
- **Known current market shares or volumes.** Allows a calibration step, and more importantly allows the base case to be checked against something observable. A simulator whose base case cannot reproduce a known reality should not be used to project an unknown one.
- **Cost data by level.** Turns preference into a value-versus-cost view, which is where feature decisions are actually made.
- **Segment membership or profiling variables.** Allows preference heterogeneity to be described in terms a business can act on, rather than as unlabelled latent classes.
- **A previous wave with the same design.** Enables comparison over time, which is legitimate only where attributes, levels and wording are identical.
- **Timing data and task-level response times.** Enables the effort diagnostics in Step 4 to be run properly rather than approximated.

## 6. Questions to ask before starting

1. **What decision will this output be used for, and by whom?** A feature prioritisation, a price decision, a portfolio line-up and a communications ranking need different outputs from the same data, and the last of these frequently should not come from a conjoint at all. Default if unanswered: produce utilities, derived importance and a sensitivity view, and withhold any single-configuration recommendation.
2. **Who chose the attribute levels, and on what basis?** Because importance is a function of the range of levels tested, this determines whether importance can be reported as a finding or must be reported as conditional on the design. Default: report importance with the level range shown alongside every figure, and state the dependency explicitly.
3. **Was a None or opt-out included, and what was it worded as?** "None of these" and "I would keep what I have now" and "I would not buy at these prices" are three different questions producing three different models. Default: if a None exists but its wording is unavailable, report shares both including and excluding it and flag the ambiguity.
4. **What competitive set will the simulator represent?** Determines whether shares are interpretable at all. Default: build the base case from the levels most closely matching current market offerings, document the mapping, and label every share as share of preference within the defined set.
5. **Are holdout tasks available?** Determines whether predictive validity can be evidenced. Default: report internal fit only, state that predictive accuracy was not assessed, and cap confidence at moderate for any simulated result.
6. **How many tasks did each respondent complete, and how long did the exercise take?** Determines the fatigue risk and whether later tasks should be examined separately. Default: run the fatigue diagnostic in Step 4 and report it regardless.
7. **For MaxDiff, how was the item list constructed and is it exhaustive of the domain?** Determines whether a low-scoring item is unimportant or merely less important than the others present. Default: state that scores are relative to the list supplied and that the list is not a census of the domain.

## 7. Step-by-step methodology

**Step 1. Reconstruct the design before touching the responses.** Write out the attribute and level matrix in full, in the wording respondents saw. Count the levels per attribute. Record the design type, tasks per respondent, alternatives per task, whether the None was present, and the prohibitions. A correct result at this step is a one-page design sheet that someone who was not on the project could read and understand exactly what was shown. If this sheet cannot be produced, estimation cannot proceed, per **K4 §6.2**: a coefficient attached to a level you cannot describe is not a finding.

**Step 2. Audit the design preconditions and record which ones failed.** Six checks, each with a consequence for what may be reported.

*Attribute independence and realism.* Attributes must be things a respondent can trade off independently. Where two attributes are effectively the same construct, or where one only makes sense conditional on another, the utilities will be unstable. *Level balance.* Each level of an attribute should appear roughly equally often across the design, and each level should co-occur with the levels of other attributes roughly equally. Serious imbalance means some levels are estimated on far less information than others, which shows up as inflated standard errors on those levels specifically, not on the attribute as a whole. *Number of levels.* An attribute with more levels tends to attract more derived importance for purely mechanical reasons, because a wider range of estimated part-worths is more likely from more estimated points. This number-of-levels effect is a known artefact and is a reason to compare importance across attributes with matched level counts wherever possible, and to caveat where not. *Prohibitions.* Every prohibited combination removes information and introduces correlation between the attributes involved. A handful is normal. Many, especially between the two attributes the study is about, can make the pair partly inseparable. *Design efficiency.* Where a design efficiency measure is available from the generation process, record it. Where it is not, check the pairwise co-occurrence counts directly: a near-zero cell for a pair of levels means that pair was never or rarely tested and any interaction between them is unestimable. *Task count and cognitive load.* Complexity is roughly attributes times levels times alternatives times tasks. Long exercises produce simplification: respondents narrow to one or two attributes and ignore the rest, which biases importance toward whatever is easiest to process, typically price and brand.

A correct result is a short design audit with a verdict per check and an explicit statement of what each failure prohibits in the output. Where the audit is materially bad, say so now, before any number exists to argue about.

**Step 3. Run respondent-level quality diagnostics and decide exclusions before estimation.** Four diagnostics. *Non-trading:* respondents who always chose the alternative with the best level of one attribute regardless of everything else. Some non-trading is real preference, especially on price; widespread non-trading on a mid-list attribute usually means the task was not being read. *Speed:* median time per task, and the distribution. Flag respondents in the fastest decile and check whether their choices are distinguishable from random. *Consistency:* where the design includes a repeated task, the repeat-choice agreement rate. *Fit:* the individual-level fit statistic from estimation (for choice models, a root likelihood or equivalent), compared against the value expected from random choice given the number of alternatives, which is one divided by the number of alternatives. Respondents at or near that value contributed nothing but noise. Also look at the None rate per respondent: someone who selected None in almost every task has told you they would not buy any of it, which is information, but they contribute little to the estimation of trade-offs between the alternatives. A correct result is a numbered exclusion decision with counts, the criterion for each, and a statement of how the headline outputs change with and without the exclusions. This is a **K5 §2.7** decision point where the exclusion changes a reported number.

**Step 4. Choose the estimation approach and state what it buys and costs.** Aggregate multinomial logit fits one set of utilities to everyone. It is fast, it is stable, and it is wrong wherever preferences differ across people, because the average of two opposed preferences is a preference nobody holds. Latent class estimation fits a small number of preference groups and assigns probabilistic membership; it is the right choice when the question is whether distinct preference structures exist and how large each is, and the class count is a judgement made on fit criteria plus interpretability, not on fit alone. Hierarchical Bayes estimates a utility set per respondent, borrowing strength from the population distribution; it is the standard default for choice-based work because it captures heterogeneity while remaining estimable from the small number of tasks any one respondent completes. Its cost is that individual-level utilities are shrunk toward the population mean, so the spread of individual preference is understated, and its stability at the individual level for an attribute a respondent effectively ignored is poor. State which was used and why. A correct result names the estimator, the specification (main effects only, or which interactions), and the reason.

**Step 5. Read the utilities correctly, which mostly means refusing three readings.** Part-worth utilities are on an arbitrary interval scale with an arbitrary origin. Three consequences follow, and all three are routinely broken. First, **a raw utility value has no meaning on its own.** Only differences within an attribute are interpretable. Second, **utilities cannot be compared across attributes**, because each attribute's utilities are typically zero-centred within itself, so a value of 40 on colour and 40 on price mean nothing in common. Third, **utilities cannot be compared across respondents or across studies** without rescaling, because the scale factor absorbs choice consistency: a respondent who chooses very consistently gets larger utilities everywhere, which reflects their decisiveness, not the strength of their preferences. Zero-centred differences within an attribute, per respondent, are the interpretable quantity. A correct result is a utility table with the scale convention stated, and no cross-attribute comparison anywhere in it.

**Step 6. Derive importance, and report it with its dependency attached.** Attribute importance is computed per respondent as the range of that attribute's utilities (best level minus worst level), divided by the sum of ranges across all attributes, expressed as a percentage; then averaged across respondents. Averaging individual importances is correct; computing importance from already-averaged utilities is not, and understates the importance of attributes where respondents disagree about direction. The critical property, and the one that must appear in the output rather than in a footnote, is that **importance is a property of the range tested, not of the attribute.** If price was tested from 100 to 120, price will look unimportant. If it was tested from 100 to 300, it will dominate. Nothing about the market changed between those two studies. The correct reporting form is always importance plus range: "price, tested from 100 to 180, accounts for 34% of derived importance". A correct result is an importance table where every row carries its level range, and a sentence stating that a different range would produce different importance.

**Step 7. Establish whether the model predicts anything, using holdouts and face validity.** Internal fit tells you the model reproduces the data it was fitted to, which is not evidence of prediction. Where holdout tasks were fielded, compute the hit rate (share of holdout tasks where the model's highest-utility alternative was the one chosen) and compare it against the chance rate for that number of alternatives, and against the share predicted by a naive constant-choice rule. Also compute the aggregate share prediction error on the holdouts, which matters more than the hit rate when the output is a share. Where holdouts do not exist, use face validity: for any attribute with a natural order, check that the estimated utilities are monotonic in the expected direction. A price attribute whose middle level is preferred to the cheapest, in aggregate, is a signal that something is wrong: the level wording, the design, or the sample. A correct result is a validity statement with numbers, or an explicit statement that predictive validity was not assessed, which caps confidence at moderate for everything downstream (**K3 §3.4**).

**Step 8. Build the simulator base case and check it against something known.** A share of preference is computed by applying a choice rule to the total utility of each alternative in a defined set. The set is the entire content of the number. Define it explicitly: which products, at which levels, mapped from which real-world offerings, with the mapping documented item by item. Then check the base case: does it reproduce the current market ordering, and where market shares are known, does it land within a defensible distance of them? A base case that says the market leader has 8% share is telling you the simulator does not represent the market, and no amount of scenario running fixes it. Where the base case is off, calibration by adjusting alternative-specific constants is legitimate, provided it is disclosed, its size is reported, and it is not presented as though the model had produced the calibrated result unaided.

**Step 9. Choose the simulation rule and handle the IIA problem explicitly.** Three rules are in common use. *First choice* assigns each respondent entirely to their highest-utility alternative; it is simple and it exaggerates share differences. *Share of preference* (logit rule) assigns each respondent proportionally to the exponentiated utilities; it is the standard, and it carries the independence of irrelevant alternatives property, which is the central problem in the technique. IIA says the ratio of shares between two alternatives is unaffected by what else is in the set. That is the red bus, blue bus problem: add a product identical to one already present, and the logit rule splits share proportionally across all alternatives, so the two near-identical products together gain share that in reality they would take almost entirely from each other. This makes plain logit simulation actively misleading for exactly the questions clients most often ask, which are line extensions and portfolio additions. *Randomised first choice* addresses this by adding random error to attribute levels and to alternatives before applying a first-choice rule, which produces realistic differential substitution between similar products. Where the simulation involves any product similar to another in the set, use a rule that handles similarity, and say which rule was used and why. A correct result is a share table with the rule named on it.

**Step 10. Run sensitivity rather than points, and report the shape.** A single simulated share is the least informative output the model can produce. Run each attribute across its full tested range holding others at base, and report the resulting share curve. This shows where the response is steep and where it is flat, which is the thing a product or pricing decision actually turns on, and it exposes the range boundaries honestly: the model cannot say anything about a level outside the range tested, and the curve should stop where the design stopped. Extrapolation beyond the tested range is prohibited, because the model has no information there and the functional form will happily produce a number anyway. A correct result is a set of sensitivity curves with the tested range marked and no line drawn outside it.

**Step 11. For MaxDiff, score correctly and then bound the reading.** Estimate at the individual level where the design allows (hierarchical Bayes on the best-worst choices), producing a score per item per respondent. Rescale for reporting: zero-centred differences, or the probability-scaled form where scores sum to 100 across items, which is easier for stakeholders and carries an important property to disclose, namely that adding or removing an item changes every other item's score. Two limits govern all reporting. **Scores are relative to the list supplied**: an item scoring 2.1 is less preferred than one scoring 8.4 among these items, and that is the entire claim. **Everything on the list may be important, or nothing may be**, because the exercise never asked whether any item mattered in absolute terms. Where an absolute reading is needed, it requires an anchored design (a direct rating or a threshold task fielded alongside), and if that was not fielded, the absolute question is unanswerable from this data. Also report the spread: a MaxDiff where the top and bottom items are close is telling you the list is undifferentiated, which is itself the finding.

**Step 12. Assemble the outputs with their boundaries attached.** Every share number carries the competitive set, the simulation rule and the phrase that it is a share of preference among the alternatives specified, not a market share forecast. Every importance figure carries its level range. Every utility table carries its scale convention. Every simulated scenario states which levels sit outside the tested range (none should) and what the base case validation showed. A correct final result is a set of outputs from which a reader cannot extract a number that means more than it should, because the qualifier travels in the same line as the figure (**K2 §7**).

## 8. Analytical framework

    Design audit → Data quality → Estimation → Utilities → Importance
        → Validation → Simulation → Sensitivity → Bounded claim

Read it as a gate sequence in which each gate can stop the sequence entirely, and in which the boundary of every downstream claim is set by the earliest gate that was passed with a qualification. A design that failed level balance on one attribute produces utilities for that attribute with wider uncertainty, an importance figure that cannot be compared with the others, and a simulation in which moving that attribute is the least trustworthy movement. That chain must be visible in the output, not resolved silently.

The framework's centre of gravity is the second-to-last gate. Simulation is where a model of stated choices among displayed alternatives becomes, in the reader's mind, a statement about the world. Everything upstream is technique. That transition is the place where the discipline is either applied or lost, and it is why every step above ends by saying what the step licenses rather than only what it produces.

## 9. Output format

**The design sheet.** Attribute, every level in shown wording, level count, prohibitions involving it, tasks per respondent, alternatives per task, None wording. One page, produced before any estimate.

**The design audit.** One row per check.

| Check | Result | Verdict | What this prohibits downstream |
|---|---|---|---|
| Level balance | Levels 1 to 4 appear 22 to 26% each; level 5 appears 4% | Fail on level 5 | Level 5 utility reported with wide uncertainty, excluded from importance comparison |

**The utility table.** Attribute, level, mean zero-centred utility, standard deviation across respondents, and the share of respondents for whom that level is the best in its attribute. The last column is what stops a stakeholder reading an average as a consensus.

**The importance table.** Attribute, derived importance percentage, the level range tested, and the number of levels. All four columns are required; importance without its range is not reportable.

**The simulation output.** Scenario name, the full specification of every alternative in the set, the simulation rule, share of preference per alternative, and the base case comparison. Every table carries this line:

> These are shares of preference among the alternatives specified, produced by a model of stated choices in a designed exercise. They are not market share forecasts and do not account for awareness, distribution, sales effort, competitor response or anything outside the attributes tested.

**Sensitivity curves.** One per attribute, share on the vertical axis, level on the horizontal, plotted only across the tested range, with the base case marked.

**The MaxDiff output.** Item, rescaled score, rank, and a spread statistic across the list. Carrying this line:

> Scores are relative to the items in this list. They indicate order and relative distance among these items only, and do not establish that any item is important or unimportant in absolute terms.

**Where the evidence is thin**, the format must not manufacture a result. Use these forms:
- `[not estimated: design confounds attributes A and B, see audit row 4]`
- `[importance not comparable across attributes: level counts differ 2 to 6]`
- `[simulation not run: no defensible competitive set could be specified]`
- `[predictive validity not assessed: no holdout tasks fielded]`
- `[outside tested range: model has no information at this level]`
- `[MaxDiff: absolute importance not answerable from unanchored design]`

## 10. Quality checks

Run before any output leaves. **K4 §8** runs anyway.

1. Is every attribute level in the output written in the wording respondents saw?
2. Does every importance figure appear with its tested level range in the same table row?
3. Has any raw utility been compared across attributes anywhere in the output?
4. Are utilities compared across respondents or waves without a rescaling step and a statement of it?
5. Does every share number name its competitive set and its simulation rule?
6. Does every share table carry the not-a-forecast statement?
7. Where a line extension or similar product was simulated, was a rule used that handles similarity, rather than plain logit?
8. Was the base case checked against a known reality, and is the result of that check reported?
9. If calibration was applied, is its size disclosed?
10. Does any sensitivity curve extend beyond the tested range?
11. Are the exclusion decisions logged with counts and the effect on headline numbers?
12. Is the estimator named, with heterogeneity handling stated?
13. For MaxDiff, is the relativity statement present, and is any absolute claim made from an unanchored design?
14. Does the output distinguish what the model predicts about choices in the exercise from what anyone is inferring about the market?
15. Is the None option's exact wording reported wherever a share including None is quoted?

## 11. Common failure modes

| Failure | How to recognise it | How to prevent it |
|---|---|---|
| **Share of preference read as market share** | A share number with a currency or unit conversion beside it, or a revenue projection built on it | The not-a-forecast line travels on every share table; revenue conversion refused without external volume data and stated assumptions |
| **Importance quoted as a property of the attribute** | "Price is the most important attribute" with no range given | Importance never reported without its level range in the same row |
| **The number-of-levels artefact** | The attribute with the most levels appearing most important, across several unrelated studies | Match level counts where possible; caveat where not; never rank importance across attributes with very different level counts |
| **Cross-attribute utility comparison** | A chart with all attributes' utilities on one axis | Zero-centre within attribute and never plot attributes on a shared utility scale |
| **IIA ignored in a line extension** | A simulated portfolio where adding a near-duplicate grows total share implausibly | Use a similarity-aware rule; sanity check that the new item's share comes mostly from its nearest neighbour |
| **Extrapolation past the tested range** | A price point in the simulator that was never shown to anyone | Hard boundary in the simulator; curves stop at the range |
| **The missing dominant attribute** | A model that fits well and predicts market behaviour badly | Named as a limitation in Step 2 and repeated on the simulation output |
| **Fatigue treated as preference** | Later tasks showing narrower attribute use than earlier ones | Compare importance estimated on early versus late tasks; report the difference |
| **Averaged utilities used to compute importance** | Importance figures that look suspiciously flat | Compute importance per respondent, then average |
| **MaxDiff read as absolute** | "Item 9 is unimportant to customers" from an unanchored exercise | The relativity line on every output; absolute claims require an anchored design |
| **AI: producing utilities without the design** | A tidy part-worth table where no level wording was ever supplied | Nothing is estimated until the design sheet exists; a level you cannot name is not a level |
| **AI: filling a simulator scenario the model cannot support** | A share for a configuration containing a level outside the tested set | Range check on every simulated profile before it runs |
| **AI: smoothing the caveat into the appendix** | A clean share table on the summary page, the not-a-forecast statement on page 40 | Caveats travel with their claim (**K2 §7**) |
| **Calibration presented as estimation** | A base case that matches market shares suspiciously well, with no note | Calibration disclosed, sized and separated from the model's own output |

## 12. AI guardrails

Universal prohibitions are inherited from **K4**. **K4 §2.1** (never invent data) and **K4 §3.3** (never overgeneralise beyond the sample) govern this skill in full. The following are specific to trade-off analysis.

1. **Never produce a utility, importance figure or share without the design that generated it.** A part-worth is defined only relative to a named level. If the level wording is unavailable, the correct output is a statement that the design specification is missing, not a plausible table.
2. **Never describe a share of preference as a market share, a volume, a forecast or a projection**, and never convert one to revenue without external volume data, a stated conversion assumption, and a statement that the conversion is an assumption rather than a result.
3. **Never report attribute importance without the level range it was derived from**, in the same line. Importance stripped of its range is a number that will be quoted for years as a property of the category.
4. **Never compare utilities across attributes, respondents or studies without stating the rescaling applied**, and never at all where no rescaling was applied.
5. **Never simulate a profile containing a level outside the tested range**, and never interpolate between tested levels of a categorical attribute as though the attribute were continuous.
6. **Never run a portfolio or line-extension simulation under a plain logit rule without stating the IIA problem** and what it does to the result.
7. **Never present a MaxDiff score as evidence that an item is important or unimportant in absolute terms** where the design was unanchored.
8. **Never exclude respondents on quality grounds without logging the criterion, the count, and the effect on the headline figures**, per **K4 §4.4**.
9. **Never report a model as validated on internal fit alone.** Fit to estimation data is not prediction, and saying so is not optional.
10. **Never resolve a design failure by proceeding with a caveat where the failure removes the basis for the estimate.** A confounded pair of attributes produces coefficients that cannot be assigned, and the honest output is that they cannot be separated.

## 13. Best-practice principles

1. **The design decides what the analysis can say, and the analysis cannot repair it.** Almost every irrecoverable problem in trade-off work was created before fielding: the wrong attributes, ranges too narrow to matter or too wide to be believed, prohibitions that gutted the design. Audit first, and be willing to conclude that the study answers a different question from the one asked.
2. **Ranges are the hidden lever in every conjoint.** Whoever set the level ranges set the importance ordering. This is not a flaw to be hidden; it is a property to be reported, and it is the reason two studies of the same category can disagree completely while both being correctly analysed.
3. **A simulator is a persuasion device as much as an analytical one.** It produces confident-looking numbers on demand, in front of stakeholders, with no visible uncertainty. Build the uncertainty in: show sensitivity ranges rather than points, show the base case validation, and never hand over a simulator without a written statement of what it does not know.
4. **Heterogeneity is usually the finding.** The average utility hides that a third of the market will not trade on one attribute at all. Report the distribution of individual preference, not only its mean, and look at whether the classes or the spread map onto anything the business can target.
5. **Check monotonicity as a matter of course.** Where an attribute has a natural direction (price, speed, warranty length), the aggregate utilities should follow it. When they do not, the cause is nearly always a level-wording problem or a sample problem, and finding it is worth more than any downstream analysis.
6. **The None option is a measurement, not a nuisance.** Its rate tells you something about the acceptability of the whole offer set, and its wording determines what. Report it rather than removing it to tidy the shares, and where it is removed, show both.
7. **Prefer relative statements to absolute ones, everywhere.** The technique is strong at ordering, comparing and finding trade-off ratios. It is weak at levels. Structure the entire output around the questions it is strong at.
8. **A hit rate is only meaningful against a benchmark.** Sixty percent sounds good and is poor when three alternatives were shown and a constant-choice rule would have got fifty. Always report the chance rate and the naive rule alongside.
9. **Do not let the simulator become the report.** The most useful outputs of a trade-off study are usually the importance structure, the heterogeneity and the trade-off ratios (how much of one attribute is worth how much of another), all of which are more robust than any share number.
10. **Watch for the attribute that was included because it could be, not because it matters.** Every attribute in a design consumes respondent attention and design efficiency. A study with four attributes that matter beats one with nine, of which five were added for completeness.
11. **Comparison across waves requires identical designs, not similar ones.** A changed level, a reworded attribute or a different task count breaks comparability, and the utilities will move for reasons that have nothing to do with the market. Where the design changed, say the comparison cannot be made.
12. **When a stakeholder asks for one number, give the range and the condition.** The request for a single share is a request to strip the evidence trail, and it is the most common route by which a correctly executed conjoint produces a wrong decision.

## 14. Worked example

**INPUT**

A fictional public transport authority, Harbour Regional Transit, wants to know how to configure a new express bus service. A choice-based conjoint has been fielded to 900 commuters recruited from an online panel with quotas on corridor and travel frequency. Five attributes: fare (four levels, from 2.00 to 5.00), journey time (three levels, 25, 35, 45 minutes), frequency (three levels, every 10, 20, 30 minutes), guaranteed seat (yes, no), and real-time arrival information (yes, no). Twelve tasks per respondent, three alternatives plus a "I would use my current option instead" opt-out. Two holdout tasks were fielded. The transport planning team asks for "the share of commuters who will switch to the express service at a 3.50 fare".

**PROCESS**

*Steps 1 and 2, design.* The design sheet is reconstructed. Two problems appear. First, guaranteed seat and real-time information are both two-level attributes while fare has four, so derived importance across the five is not directly comparable and the number-of-levels effect will favour fare. Second, a prohibition was applied preventing the 25-minute journey time from appearing with the 2.00 fare, on realism grounds. Checking co-occurrence, the pair appears zero times, so the interaction between fare and journey time at the fast, cheap corner is unestimable. Neither problem is fatal. Both constrain the output, and both go in the audit table.

*Step 3, quality.* Median task time 14 seconds. The fastest decile averages 4 seconds and their individual fit statistics sit close to the 0.25 expected from random choice across four options. Forty-one respondents (4.6%) are flagged. Non-trading on fare appears for 96 respondents (10.7%): they always chose the cheapest alternative. This is plausible real behaviour for commuters on a tight budget, not inattention, and their fit statistics are good. *The judgement call.* Exclude the 41 speeders, retain the 96 fare-only choosers, and report both decisions with the effect: removing the speeders moves derived fare importance by under one point, and removing the fare-only group would move it by nine. That second figure is the reason they are retained and reported rather than removed, and it is disclosed.

*Steps 4 to 6, estimation and importance.* Hierarchical Bayes, main effects. Derived importance: fare 39% (range 2.00 to 5.00), journey time 27% (25 to 45 minutes), frequency 19% (every 10 to 30 minutes), guaranteed seat 9% (yes/no), real-time information 6% (yes/no). Every figure is reported with its range, and a line states that had fare been tested from 3.00 to 3.60, its importance would have been a fraction of this. Monotonicity holds on all three ordered attributes.

*Step 7, validity.* Holdout hit rate 61% against a chance rate of 25% and a naive always-choose-cheapest rule at 44%. Aggregate share prediction error on holdouts, 3.1 points. Adequate, and reported as such.

*Steps 8 and 9, simulation.* The base case is built with the current service (fare 2.00 equivalent, 45 minutes, every 30 minutes, no seat guarantee, no information) plus the opt-out. It reproduces the known current usage split within four points, which is reported as the validation. The requested scenario, an express at 3.50 with 25 minutes and every 10 minutes plus seat guarantee, is run. Because the scenario introduces a service similar in some respects to the existing one, randomised first choice is used rather than plain logit, and the difference is material: plain logit gives the express 47%, the similarity-aware rule gives 41%, with most of the difference coming from the existing service rather than from the opt-out.

*Step 10, sensitivity.* The fare curve is flat between 2.00 and 3.00 and steepens sharply above 4.00. The most decision-relevant finding in the study is that the fare can move a full unit with little share cost, and that the cliff sits between the third and fourth tested levels, where the model's resolution is coarsest and where the next study should place more levels.

**OUTPUT**

> At a 3.50 fare, 25-minute journey time, 10-minute frequency and a guaranteed seat, the express service takes 41% share of preference among the three alternatives specified (existing service, express service, opt-out), under a similarity-aware simulation rule (n=859 after 41 quality exclusions; holdout hit rate 61% against a 25% chance rate). This is a share of preference in a designed choice exercise. It is not a forecast of the share of commuters who will switch, and it does not account for awareness of the new service, timetable reliability in practice, employer travel schemes, or anything else outside the five attributes tested.

> Fare accounts for 39% of derived importance across the range tested (2.00 to 5.00), journey time 27% (25 to 45 minutes), frequency 19% (10 to 30 minutes), guaranteed seat 9% and real-time information 6%. Importance is a function of the ranges tested and would change with different ranges. The two-level attributes are not directly comparable with the four-level fare attribute.

> The fare sensitivity curve is flat from 2.00 to 3.00 and falls steeply above 4.00. Between 3.00 and 4.00 the model's resolution is limited by having only these levels tested.

**Researcher decision required.** The planning team's question was phrased as a switching forecast. The model cannot answer it. The decision is whether to report the share of preference with the boundary attached, or to commission the observational work (a pilot corridor, or a matched-control comparison per **06.04**) that could answer the switching question. This is a **K5 §2.3** judgement about what the authority will do with the answer.

## 15. Advanced usage

**Trade-off ratios instead of shares.** Dividing the utility difference between two levels of one attribute by the utility difference per unit of another gives an exchange rate: how many minutes of journey time a commuter will trade for a unit of fare. These ratios are more robust than shares, do not depend on the competitive set, and are frequently the most transferable output of the study. Where the denominator attribute is money, this is the willingness-to-pay reading, and it must carry the caveats in **06.02**.

**Interactions.** Main-effects designs assume the effect of one attribute does not depend on another. Where an interaction is theoretically important (brand and price is the classic case), it must be designed for, because a design efficient for main effects is usually poor at estimating interactions. Test the specific interaction rather than adding them all: the degrees of freedom cost is real and the risk of fitting noise is high.

**Latent class as a segmentation input.** Preference classes from a choice model are a stronger basis for segmentation than attitudinal clustering in categories where choice is the behaviour of interest, because the classes are defined on the thing that matters. Hand to **09.01 Audience Segmentation** with the class definitions, sizes and profiling, and be honest that class count is a judgement.

**Portfolio and cannibalisation work.** The natural extension of the simulator, and the place where the IIA problem bites hardest. Model the full line-up, and read the source of switching (which alternative each new item's share came from) rather than only the totals. A line extension that adds three points of total share while taking eight from the flagship is a different decision from one that takes eight from a competitor.

**Comparing a conjoint against revealed behaviour.** Where transaction data exists for a comparable choice, compare the model's predicted ordering against what people actually did. Agreement raises confidence considerably; disagreement is more useful still, because the gap between stated and revealed choice is usually concentrated in specific attributes and identifying which ones tells you where the exercise was unrealistic.

**Reporting conventions in academic and regulated settings.** Discrete choice experiments in health economics, transport appraisal and environmental valuation carry stricter reporting requirements: the full design generation method, the model specification and its statistical output, tests of the IIA assumption where a logit model is used, and often a sensitivity analysis across model specifications. Where the output will be published or submitted, meet those conventions and hand to **15.xx** academic reporting standards.

## 16. Skill chain

**Recommended previous skills**
- **01.07 Analysis Plan Development.** Hands over what the study was supposed to answer, which is the standard against which the design audit is judged.
- **04.02 Data Cleaning** and **04.04 Data Transformation and Dataset Preparation.** Hand over a respondent-level dataset in the long format that choice estimation requires, with task, alternative and choice correctly structured.
- **03.03 Stimulus and Concept Preparation.** Hands over the exact level wording and any visual stimulus used, without which the design sheet cannot be built.
- **05.01 Descriptive Analysis.** Hands over the sample profile and the base structure the model's population represents.

**Recommended next skills**
- **06.02 Pricing Research Analysis.** Takes the price-related output where the question is a price decision, and supplies the stated-preference caveats and the demand-curve framing this skill deliberately stops short of.
- **08.01 Finding to Insight Development.** Takes the importance structure, the heterogeneity and the sensitivity shape and asks what they mean, which is where a utility becomes an argument.
- **09.01 Audience Segmentation.** Takes latent classes or individual utilities as a preference-based segmentation input.

**Runs well alongside**
- **05.06 Correlation, Regression and Causal Claim Control**, which supplies the language discipline for the transition from an in-experiment effect to a real-world claim.
- **13.03 AI Output Verification**, which audits a finished conjoint deck for unranged importance, unbounded shares and extrapolated simulations.
- **K3 Confidence and Uncertainty Protocol**, which supplies the register for a model validated on holdouts versus one validated on nothing.

---
A Yazi Supplied Skill and resource.
