---
name: stimulus-and-concept-preparation
description: >
  Prepares concepts, advertising, packaging, prototypes and other material for
  testing, so that respondents react to the thing being tested rather than to its
  execution. Use when someone says "prepare the concepts for testing", "write the
  concept statements", "how finished should the stimulus be", "monadic or sequential
  monadic", "how many concepts can one person see", "should we rotate", or asks why
  one concept keeps winning.
category: 03 Fieldwork and Data Collection
ref: 03.03
tier: 2
inherits: [K2, K3, K4, K5]
---

# Stimulus and Concept Preparation

## 1. One-line description
Prepares the material respondents will react to, at a fidelity that matches the decision, written and presented at controlled parity across the set, so that the differences in the results are differences in the ideas rather than differences in how they were dressed.

## 2. What this skill is used for

**The research problem it solves.** Concept testing is unusually easy to run and unusually easy to invalidate, and the invalidation is silent. The single commonest failure is parity failure: one concept in the set is written with more benefit specificity, more enthusiasm, a stronger reason to believe or simply more words than its rivals, and it wins. The result is real, repeatable and about the writing. Nobody sees it, because there is nothing in the data that says "this concept had 40 more words and an extra proof point". A close second is fidelity mismatch: a finished, art-directed execution is put in front of respondents to test an idea, and they react to the photography, the colour, the model, the typeface. The debrief then reports that consumers rejected the idea, when what they rejected was one rendering of it. Both failures produce clean-looking data that answers a question nobody asked.

**Where it sits in the research lifecycle.** After the objectives and the design are set and before fieldwork. It sits alongside instrument design rather than after it: what the stimulus is determines what the questionnaire can ask, and the exposure design determines the sample structure, so preparing stimulus after the questionnaire is finished tends to produce a mismatch.

**Typical use cases.**
- Preparing a set of new product or service concepts for quantitative screening.
- Preparing advertising or communications material for pre-testing at a chosen finish level.
- Preparing packaging, on-shelf or in-context, with the competitive set that surrounds it.
- Preparing a prototype, journey or interface for evaluation, deciding how much of it needs to work.
- Preparing pricing or proposition material where the price has to be presented consistently.
- Auditing an existing concept set that has already been written, usually because one option is winning suspiciously.

**Who uses it.** Research managers and directors preparing or reviewing stimulus; client-side insight and innovation managers who receive concepts from an agency and have to judge whether they are testable; product and UX researchers preparing prototypes; brand and communications researchers running pre-tests.

## 3. When to use it

- Concepts, ads, packs or prototypes exist in some form and have to be made ready to put in front of people.
- The material is arriving from different authors (an agency, an internal team, a founder) and is visibly uneven in style and finish.
- A decision has to be made about how finished the stimulus should be, and the cost of finishing it is material.
- The design requires a choice between monadic, sequential monadic and comparative exposure.
- More than about three items will be shown to one respondent and order effects have to be managed.
- The category is one where the competitive context changes the reading, which is most categories.
- A previous test produced a result nobody believed, and the stimulus is the suspect.

## 4. When NOT to use it

- **The idea is not yet articulated.** A concept test needs concepts. Where the proposition is still being formed, testing a set of half-written ideas produces a ranking of drafting quality. Use qualitative exploration (**02.02 Discussion Guide Design**) to build the vocabulary and the propositions, then return to test them.
- **The objective is a trade-off, a price point or a preference share.** Rating a set of concepts on a scale does not produce willingness to pay, optimal feature bundles or share. Those need a designed choice exercise from category 06. Presenting concept ratings as though they forecast share is one of the most expensive misreadings in commercial research.
- **The decision is about execution and the stimulus is a rough concept, or vice versa.** Rough stimulus cannot adjudicate between executions, and finished stimulus cannot cleanly adjudicate between ideas. Where both questions are live, they are two studies or two stages, not one set of stimulus asked to do both.
- **The concepts cannot be brought to parity.** Where one idea genuinely requires more explanation than another (a novel mechanism versus a familiar one), forced parity distorts the comparison and unforced disparity invalidates it. Say which effect you are accepting, and consider testing the harder concept monadically against its own norm rather than in the set.
- **You are designing the questionnaire around the stimulus.** Question wording, scale choice and the diagnostic battery belong to **02.01 Survey Questionnaire Design** and **02.07 Scale and Measurement Selection**. This skill delivers the stimulus and the exposure design; it does not write the response instrument.
- **The result will be used as a market forecast.** Concept test scores rank options against each other reliably and predict absolute market outcome poorly, because the test removes almost everything that determines real-world performance: availability, price in context, competitive response, media weight, repeat purchase and the fact that nobody in the market will read the concept. If a volumetric forecast is required, it needs a calibrated forecasting design and external data, not a concept score.
- **The material is legally or regulatorily constrained and the constraints are unresolved.** Health claims, financial promotions, comparative advertising, medicines information. Testing a claim that cannot be made produces a result nobody can act on and can create a record that causes problems later. Resolve the claim status first.
- **Fieldwork has begun.** Changing stimulus mid-field splits the sample into non-comparable groups. If a stimulus defect is found in field, it is a **03.04 Fieldwork Monitoring and Response Quality** decision about pausing, and any change is a logged amendment that partitions the data.

## 5. Required inputs

**Required. Without these the skill cannot run.**
- **The decision the test informs**, specifically: which concepts go forward, whether this execution runs, whether to build the feature, whether the pack change is safe. The decision determines fidelity, exposure design and how many items can be tested.
- **The raw material for each item**, in whatever state it exists: written propositions, scripts, layouts, sketches, wireframes, pack designs.
- **What is intended to vary across the set, and what is intended to be held constant.** This is the experimental structure and without it the study is not a test, it is a display. If it is not stated, ask, and if no answer is available, stop.
- **Target audience and their category familiarity**, because the amount of context a concept needs is a property of the reader, not of the concept.

**Optional, and what each one adds.**
- **The competitive set**, real and current. Allows realistic context of exposure, which usually lowers scores and always improves their usefulness.
- **Price, and whether it is part of what is being tested.** Price presence changes appeal ratings substantially, so its inclusion or exclusion must be deliberate and consistent, and its status recorded.
- **Brand assets and whether branding is being tested.** Branded and unbranded stimulus answer different questions. Adding a strong brand to a weak concept can rescue it in the test and not in the market.
- **Norms or previous results from comparable tests.** Give the scores a reference point. Only usable where the stimulus format, exposure design and audience match, and that match should be checked rather than assumed.
- **Real-world exposure conditions**: where the material will actually be seen, for how long, in what medium. Lets the test approximate reality rather than a reading exercise.
- **Regulatory or claim substantiation status**, which determines what may be written on the stimulus at all.

## 6. Questions to ask before starting

1. **What exactly is under test: the idea, the expression, the execution, or the whole package?** Every subsequent decision follows from this. Default if unanswered: assume the idea, prepare at low fidelity, and state that execution effects are excluded by design.
2. **What decision follows from the result, and what score would change it?** A screening decision needs relative ranking and tolerates rough stimulus. A go/no-go on a build needs more. Default: treat the output as relative ranking only, and say so.
3. **Is price part of the concept?** A concept without a price is a wish, and appeal ratings on price-free concepts are systematically higher than the same concepts priced. Default: exclude price, apply that consistently across the set, and disclose that appeal was measured without price.
4. **How familiar is this audience with the category?** Determines how much explanatory context each concept needs, and therefore how hard parity is to achieve. Default: assume moderate familiarity and check comprehension in pilot.
5. **Will respondents see one concept or several?** This is a sample structure decision, not a presentation choice, and it must be settled before base sizes are agreed. Default: sequential monadic with rotation, which is the common compromise, and state its limitations.
6. **Where will the real thing be encountered, and for how long?** A pack seen for two seconds on a shelf and a pack studied for forty seconds on a screen are different objects. Default: state the intended exposure condition and note where the test departs from it.
7. **Who wrote these, and did the same person write all of them?** Different authors produce different levels of persuasiveness, and mixed authorship is the leading cause of parity failure. Default: assume parity failure until the audit in Step 3 shows otherwise.

## 7. Step-by-step methodology

**Step 1. Separate the layers and name the one under test.** Any stimulus a respondent reacts to has four layers: the **idea** (the underlying proposition), the **expression** (the words and framing chosen to convey it), the **execution** (the visual and material rendering), and the **exposure context** (what surrounds it and how long they have). The response is to all four, always. The design's job is to hold three constant and vary one. Write down which layer is under test and, for the others, what will be held constant and how. A correct result is a one-line statement of the form "this test varies the idea; expression is standardised by a common template, execution is held at a uniform low fidelity, and context is held constant with no competitive frame". Studies that skip this step usually end up varying idea and expression together and attributing the whole difference to idea.

**Step 2. Choose fidelity to match the decision, not the budget or the enthusiasm.** Rough stimulus (text-only propositions, sketches, wireframes, greyscale layouts) isolates the idea, is cheap enough to test many options, and invites respondents to imagine, which tends to produce generous and unstable ratings. Finished stimulus (art-directed, produced, working) tests something closer to what will exist, produces more grounded reactions, and confounds idea with execution because a respondent cannot separate a good idea rendered badly from a bad idea. Two rules govern the choice. **Fidelity must be uniform across the set**: a set containing one finished item and three roughs is not a test of ideas, it is a test of finish, and the finished one wins. **Fidelity must match the decision**: a screening decision among twelve ideas is served by rough parity stimulus; a decision to commit production money is not. Where fidelity has to rise between stages, treat it as a new test rather than a continuation, because scores do not carry across fidelity levels.

**Step 3. Write the set to parity, and audit it item by item.** This is the core discipline of the skill and the point at which most concept tests are won or lost before anyone is recruited. Build a single template that every concept fills, with the same components in the same order: an insight or context line, the proposition, how it works, the reason to believe, and the benefit. Then audit the set on six dimensions, comparing across items rather than reading each on its own.

- **Length.** Word counts within a narrow band. Longer concepts carry more information and more chances to find something appealing.
- **Benefit specificity.** "Saves you time" and "cuts twenty minutes off your evening" are not the same claim strength. Either all concepts quantify or none do.
- **Enthusiasm and adjective load.** Count the evaluative words. A concept written with "delicious", "effortless" and "brilliant" against rivals written flatly will win, and the finding will be about copywriting.
- **Reason to believe.** Either every concept has proof, or none does. A single concept carrying a credibility anchor its rivals lack has an advantage unrelated to the idea.
- **Concreteness.** Named ingredients, named features, specific occasions and specific numbers all raise appeal. Hold the level of concreteness constant.
- **Framing of the problem.** A concept that opens by describing a painful problem primes need before it offers the solution. If one does it, all do.

The audit is done by someone who did not write the concepts, and it is done by reading the set in a column, dimension by dimension, rather than concept by concept. A correct result is a parity table with a row per concept and a column per dimension, and either equivalence on every dimension or an explicit, documented decision to accept a disparity because it is intrinsic to the idea (see Section 4). Where an agency or a stakeholder wrote the set, expect the strongest-written concept to be the one they already prefer, and expect to have that conversation.

**Step 4. Choose the exposure design, and know what each one supports.** Three options, and the choice determines the sample structure.

- **Monadic**: each respondent sees one concept only, with respondents randomly allocated. This is the cleanest design. Scores are uncontaminated by comparison, they can be compared to monadic norms, and it is the only design that approximates how a real person encounters a proposition, which is on its own. It costs the most base, since every cell needs a full sample, and it produces no within-respondent preference data.
- **Sequential monadic**: each respondent sees several concepts one at a time, rating each fully before the next appears. Efficient, gives within-respondent comparison, and supports both rating and ranking. Its cost is that the second and later concepts are not judged in isolation: respondents calibrate against what they have already seen, later ratings are pulled toward the first item's anchor, and diagnostic differentiation degrades as fatigue accumulates. Rotation controls the systematic part of this, not the effect on absolute levels.
- **Comparative or side-by-side**: all concepts shown together and compared directly. It is the most sensitive design for detecting small differences and the least realistic, because it creates a choice that does not exist in the market and invites respondents to look for differences they would never otherwise notice. Useful for optimisation between close variants (two pack designs, two headlines); misleading as a measure of appeal.

The rule of thumb worth stating plainly: **monadic for absolute reads, sequential monadic for efficient relative reads, comparative for fine discrimination between close alternatives.** State the design and its limitation in the output.

**Step 5. Set the number of items per respondent from attention, not from arithmetic.** In sequential monadic exposure, quality falls with each successive item: reading time drops, open-end length drops, and ratings compress toward the middle. Three to five items is a common working range for full evaluation with diagnostics, fewer where each item is long or complex, and this is a design convention rather than a measured constant, so it should be calibrated in pilot by inspecting reading time and open-end length by position. Where more concepts must be covered than a respondent can evaluate, use an incomplete block design: each respondent sees a subset, subsets are balanced so every concept appears equally often and equally often in each position, and every pair of concepts co-occurs a similar number of times. This preserves comparability at the cost of a larger total sample, and it must be planned with the sampling strategy rather than improvised.

**Step 6. Control order and position.** Rotate the order of concepts across respondents so that position effects distribute evenly rather than accumulating on one item. Rotate, do not merely randomise, where the number of items is small enough for balanced rotation, since randomisation with small cells can leave real imbalance. Record the position each respondent saw each concept in, as a variable, so that a position effect can be measured rather than assumed away. In comparative exposure, rotate left-to-right and top-to-bottom placement, because visual position has its own advantage. Where a monadic design is used, allocation to cell must be random and the cells must be checked for composition equivalence before results are compared, since an imbalance in who saw what is a confound that looks exactly like a concept difference.

**Step 7. Build the exposure context to match reality, deliberately.** Decide and specify: whether the item is seen alone or among competitors, on what surface and at what size, for how long, and whether re-exposure is permitted. Each of these changes the result. Competitive context lowers appeal scores, usually substantially, and improves their diagnostic value, because the real question is never "do you like this" but "does this pull you away from what you do now". Time-limited exposure is essential for anything whose real encounter is brief (packaging on shelf, an outdoor poster, a scrolling feed) and misleading when applied to material that is genuinely studied (a service proposition, a policy document). Re-exposure, allowing the respondent to look again before answering diagnostics, produces more considered and less spontaneous responses; permit it or forbid it consistently and say which. A correct result is an exposure specification precise enough that another researcher could reproduce the conditions.

**Step 8. Decide branding and price, and apply the decision to every item.** Both are powerful, both are common sources of accidental confounding, and both must be uniform across the set unless one of them is the variable under test. Branded stimulus measures the idea plus the brand's permission to make it. Unbranded stimulus measures the idea, and overstates what an unknown entrant could achieve. Priced concepts score lower and more realistically than unpriced ones. There is no universally correct choice; there is only a choice that must be made once, applied everywhere, and disclosed.

**Step 9. Assemble the stimulus pack and its specification.** Produce the final material plus a specification sheet recording, for each item: fidelity level, word count, the parity audit result, branding and price status, exposure condition, rotation position pool, and the version number. Version control matters more than it appears to: stimulus is edited late, often by people outside the research team, and a study fielded on version 3 while the report describes version 2 is a failure that survives every quality check aimed at data.

**Step 10. Pilot for comprehension before fielding.** Show the stimulus to a small number of people from the target audience and ask them to say back, in their own words, what is being offered, who it is for and what is new about it. This finds the failure that ratings cannot: a concept that is not understood scores in the middle rather than badly, because respondents rate what they think it might be. Check also that no concept contains a term the audience does not use, that the promise is understood as intended rather than more strongly, and that no item is read as an existing product already on the market. Fix the stimulus, re-audit parity after fixing it, and version it again. Comprehension failures found in pilot are cheap; found in analysis they are indistinguishable from rejection.

## 8. Analytical framework

Every response to stimulus is a response to a stack:

    Idea → Expression → Execution → Exposure context → Response

Read downward, this is the design test: what varies across the set at each layer, and was that variation intended? Whatever varies uncontrolled becomes part of the measured difference and cannot be separated from it afterwards. A set that varies at the idea layer and, accidentally, at the expression layer produces a result that is a mixture of the two in unknown proportion, and no analysis recovers the split.

Read upward from a result, it is the diagnostic: before concluding that concept B beat concept A on the idea, check the expression layer (was B written with more benefit specificity or more enthusiasm), the execution layer (was B rendered more finished, more attractively, more legibly) and the context layer (was B seen first more often, or alone, or without price). Only when those are ruled out is the difference attributable to the idea.

The framework also fixes the reporting rule. Because the exposure context in a test is not the exposure context in the market, and because the market layer is absent altogether, **concept test results support comparison within the set and do not support absolute prediction of market outcome.** The ranking is the finding. The score is a measurement in an artificial condition, useful against norms collected in the same condition and useful for nothing else.

## 9. Output format

**1. Stimulus brief.** What is under test, at which layer, what is held constant and how, and the decision the test informs.

**2. Fidelity decision and rationale**, including what the chosen fidelity can and cannot adjudicate.

**3. The stimulus set**, final, versioned, one item per page or file, in a consistent template.

**4. Parity audit table.**

| Item | Words | Benefit specificity | Evaluative words | Reason to believe | Concreteness | Problem framing | Verdict |
|---|---|---|---|---|---|---|---|

Any accepted disparity is recorded with its reason and the direction of the advantage it confers.

**5. Exposure design specification.** Monadic, sequential monadic or comparative; items per respondent; rotation or block scheme; position variable to be captured; cell allocation and base per cell.

**6. Exposure condition specification.** Alone or in competitive context, medium and size, duration, re-exposure permitted or not, branding status, price status.

**7. Comprehension pilot results**, with the changes made and the re-audit after change.

**8. Version log.** Every stimulus version, its date, what changed and who authorised it.

**9. Reporting caveats, drafted in advance.** The sentences that will appear in the report about what these results do and do not support, including the relative-versus-absolute caveat and the exposure-condition caveat.

**10. Review points** per **K5 §3**.

**When the inputs are thin**, an item whose parity cannot be established because it arrived in a different format is marked `parity not established` and its result is flagged as non-comparable, rather than being scored anyway. Missing price or brand decisions are recorded as `[not specified]` and raised, since defaulting them silently makes the whole set incomparable to anything.

## 10. Quality checks

Run before stimulus is released to field. These sit on top of **K4 §8**.

1. The layer under test is stated, and the other three layers have a stated control.
2. Fidelity is uniform across the set, and matches the decision the test informs.
3. The parity audit has been completed by someone who did not write the concepts.
4. Word counts sit within a narrow band, and no item carries a reason to believe that others lack.
5. Evaluative language has been counted, not judged impressionistically, and is balanced across the set.
6. Branding status and price status are identical across items, or the difference is the variable under test.
7. The exposure design is named, and its limitation is written into the reporting caveats.
8. The number of items per respondent has been set against attention and will be checked in pilot by position.
9. Rotation or a balanced block design is in place, and the position each respondent saw is captured as a variable.
10. In monadic designs, cell allocation is random and composition equivalence across cells will be checked before comparison.
11. The exposure condition (context, duration, re-exposure) is specified precisely enough to reproduce.
12. Comprehension has been tested with the target audience and the stimulus re-audited after any change.
13. Every item carries a version number, and the version that will field is the version described in the specification.
14. The reporting caveat on relative versus absolute reading is written and agreed before results exist.

## 11. Common failure modes

| Failure | How to recognise it | How to prevent it |
|---|---|---|
| **The best-written concept wins** | One item leads on every measure including irrelevant ones, and reads more fluently | Parity audit by an independent reader, on the six dimensions, before fielding |
| **Fidelity mismatch within the set** | One finished item among roughs, or one with a photograph | Uniform fidelity, enforced at pack assembly, not at briefing |
| **Testing execution while reporting on idea** | The debrief says the idea failed; the stimulus was a single art-directed route | State the layer under test in the brief and repeat it in the report |
| **Order effects unmanaged** | Whichever concept is shown first scores highest | Rotate, capture position, and test for a position effect before interpreting |
| **Too many concepts per respondent** | Open ends shorten and ratings compress after item three | Set items per respondent from pilot evidence on reading time and open-end length |
| **Comparative design read as appeal** | Small differences reported as meaningful preference | Reserve comparative exposure for close-variant optimisation and say so in the report |
| **Silent price or brand asymmetry** | One concept mentions a price or a brand and the others do not | Branding and price decided once, applied to every item, recorded in the specification |
| **Context-free testing** | Everything scores well and nothing differentiates | Add the competitive frame; expect scores to fall and diagnostics to improve |
| **Comprehension failure read as rejection** | Middling scores with vague or contradictory open ends | Comprehension pilot with playback in the respondent's own words |
| **Version drift** | The report describes stimulus that differs from what fielded | Version every item; the specification names the fielded version |
| **AI: fabricated stimulus content** | Generated concepts containing features, ingredients, prices or claims nobody supplied | **K4 §2.1**. Substantive content comes from the client; missing elements are labelled placeholders |
| **AI: uneven persuasiveness** | Generated concept sets where one is noticeably more vivid because the model found it more interesting | Generate to a fixed template with a word budget, then run the parity audit on the generated set as strictly as on a human-written one |
| **AI: score prediction** | A statement about how a concept is likely to perform | **K4 §2.1** and **§3.5**. No prediction without data; concept scores are not forecastable from text |
| **The forecast slide** | A concept score converted into a volume or share estimate | Written caveat agreed before fielding; forecasting requires a calibrated design and external data |

## 12. AI guardrails

Skill-specific only. **K4** applies in full and is not repeated here.

1. **Never invent the substantive content of a concept.** Features, ingredients, prices, claims, technologies, availability and brand names come from the client. Where an element is required by the template and not supplied, output a labelled placeholder and name what is needed (**K4 §2.1**).
2. **Never write one concept in a set more persuasively than the others.** Generation runs to a fixed template with a word budget and a fixed count of evaluative words, and the generated set is parity-audited before use.
3. **Never predict how a concept will score, rank or perform in market.** Text does not license a performance estimate, and a plausible one will be quoted back as a benchmark.
4. **Never present a concept score as a market forecast, a share estimate or a demand figure**, and never convert one into another by any arithmetic.
5. **Never compare scores across fidelity levels, exposure designs or exposure conditions** without stating that the conditions differ; norms apply only where the format matches.
6. **Never alter stimulus without a version increment and a log entry**, including small copy edits requested late by a stakeholder.
7. **Never write a claim onto stimulus that has not been confirmed as substantiable**, particularly in health, financial and comparative advertising contexts. Flag it for confirmation instead.
8. **Never describe a concept test as validating a proposition.** It ranks options under artificial exposure. State what it supports, in the words drafted in Step 9 of the output format.

## 13. Best-practice principles

1. **Respondents cannot separate the idea from how it is written, and neither can the data.** Everything in this skill follows from that. Parity is not tidiness, it is the experimental control.
2. **The concept nobody had to explain is not necessarily the best idea.** Familiar propositions are easier to write, easier to understand and score better. A genuinely new idea is systematically disadvantaged in a concept test, and this should be said out loud when the results are read.
3. **Add price and add competitors, and expect scores to fall.** They fall because the test got closer to the world. A set that scores lower and differentiates more is a better test than one that scores high and does not.
4. **Rough stimulus gets generous ratings.** People fill in the gaps with their own best case. Treat high scores on rough stimulus as evidence of interest in the space, not of the specific proposition.
5. **A middling score with confused open ends is a comprehension failure, not a rejection.** Read the verbatims before believing the number, every time.
6. **Rotation converts a systematic bias into noise, which is survivable.** An unrotated set carries a fixed advantage in a fixed position and there is no analytical repair.
7. **The set is the unit of design, not the concept.** Review the items side by side, in a column, dimension by dimension. Reading them one at a time is how parity failures survive.
8. **Whoever wrote the concepts has a favourite, and it is usually the strongest-written one.** This is not bad faith, it is attention. It is also exactly why the parity audit is done by someone else.
9. **Fidelity is a decision about which question you can answer, not about how good the research looks.** Spending money to finish stimulus buys realism and costs you the ability to attribute the result to the idea.
10. **Write the caveat before the result exists.** A caveat drafted in advance is a design statement. The same caveat added after a disappointing result reads as an excuse and gets removed.
11. **Version control the stimulus like code.** Late edits are normal, undocumented late edits are how a report ends up describing something that was never fielded.
12. **Concept tests rank well and forecast badly.** Say it once, clearly, in every report. The ranking is worth having, and it is the part that survives contact with reality.

## 14. Worked example

**INPUT**

A fictional chilled foods manufacturer, Ashgrove Kitchens, has five ideas for a new range of evening meals and wants to take two forward to development. Material supplied: three written concepts from a brand agency, one from the innovation team, and one that exists as a single art-directed pack visual with a headline. Audience: main grocery shoppers who buy chilled ready meals at least fortnightly. Budget supports 600 completes.

**PROCESS**

*Step 1, layers.* The decision is which ideas to develop, so the layer under test is the idea. Expression will be standardised by template, execution held uniform at low fidelity, exposure context held constant.

*Step 2, fidelity.* The art-directed visual is the problem. Included as supplied, it would have carried a large and unmeasurable advantage. Two options existed: finish all five to the same standard (cost and time not available) or reduce the fifth to the same text template as the others. The second was taken and recorded, with the consequence stated: the test will not tell the client anything about whether that pack design works, which is a separate question needing a separate test.

*Step 3, parity audit.* Run by a researcher who had not seen the briefs. The three agency concepts ran 95 to 110 words; the innovation team's ran 172. Evaluative word counts were 6, 5, 7, 2 and 4. Two concepts quantified their benefit ("ready in 12 minutes"), three did not. One opened with a problem frame ("by seven o'clock, nobody wants to start cooking") and the others did not.

*The judgement call.* The 172-word concept was the one with the genuinely novel mechanism, and cutting it to 110 words removed the explanation that made it comprehensible. Forcing parity would have tested whether an unexplained new idea appeals, which is not the question. The resolution: the concept was cut to 128 words, still above the band, the residual disparity was documented in the parity table with its direction of advantage noted, and a comprehension check was added specifically for that item in pilot. The alternative considered and rejected was to test it monadically outside the set, which the base size did not support.

All five were then rewritten into a common template with an equal count of evaluative words, a quantified benefit in each (the client supplied real figures for the three that lacked them, and one was left as `[not available]` and raised, since inventing it was not an option), and the problem frame removed from the one concept that had it, since adding it to all five would have primed need across the whole test.

*Step 4 and 5, exposure.* Sequential monadic, three of five per respondent in a balanced incomplete block design, giving 360 evaluations per concept. Monadic was rejected on base size (120 per cell would not have supported the subgroup read the client wanted). Comparative was rejected because the concepts were not close variants.

*Step 6, rotation.* Balanced blocks with each concept appearing equally often in each of the three positions, position captured as a variable.

*Step 7, context.* Competitive frame added: each concept shown with a short line describing what the respondent currently buys, drawn from their own earlier answer, so the comparison is to their actual repertoire rather than to nothing. Price included at a stated indicative level, identical in format across all five.

*Step 10, pilot.* Twelve comprehension interviews. Four of twelve read the novel-mechanism concept as a product already sold by another brand, which would have depressed its "new and different" rating for a reason unrelated to the idea. One line was added to distinguish it, the parity table was re-audited, and the stimulus went to version 3.

**OUTPUT**

Five parity-audited concepts at uniform low fidelity, one documented disparity with its direction stated, a balanced incomplete block design with 360 evaluations per concept, a specified exposure condition including price and personal competitive frame, a version log ending at version 3, and two review points: whether the residual length disparity on concept 4 should be reflected in how its result is read, and whether the abandoned pack design should be scheduled as a separate execution test rather than being treated as answered by this study.

## 15. Advanced usage

**Testing idea and execution in one design.** Where both questions are live and budget allows, cross them: two ideas by two executions, monadically allocated, giving four cells. This is the only way to separate the two effects, it quadruples the base requirement, and it needs the cells to be balanced on composition. Anything less than a full crossing leaves the two confounded, and reporting a partial design as though it separated them is worse than not testing execution at all.

**Norms and their honest use.** A norm is only usable where fidelity, exposure design, exposure condition, price status, brand status, audience and country match. Because that is a demanding list, most norm comparisons are looser than they appear, and the safest use is directional: this concept sits in the upper part of the range for tests conducted this way, in this category. Where the match is imperfect, say which conditions differ and in which direction they would bias the comparison.

**Iterative optimisation.** Where a concept is close but flawed, a second round comparing variants of the same concept (one element changed at a time) is more informative than re-testing the whole set. Change one element per variant, hold everything else identical, and treat the result as diagnostic rather than as a new appeal score.

**Prototypes and interfaces.** For interactive stimulus, fidelity has a second dimension: how much actually works. A clickable prototype that fails on a path the respondent chooses produces a reaction to the failure, not to the design, so define the supported paths, script the exposure to stay within them, and record every departure. Where the respondent goes off-path, that is data about expectation and it should be captured rather than corrected.

**When the standard approach does not fit.** Highly novel propositions, complex services and anything requiring behaviour change are systematically underrated in short-exposure concept tests, because the test measures immediate reaction and the proposition needs understanding. Consider extended exposure, a two-stage design with a gap, or qualitative evaluation before quantitative screening, and be explicit that a low score in the standard format is weak evidence for that class of idea.

## 16. Skill chain

**Recommended previous skills**
- **01.04 Research Method Selection.** Establishes whether a concept test answers the question at all, and whether a choice-based design is needed instead.
- **02.02 Discussion Guide Design.** Where the propositions are not yet articulated, produces the vocabulary and the raw material this skill turns into testable stimulus.
- **01.06 Sampling Strategy.** Supplies the base sizes that determine whether monadic exposure is affordable and how large a block design can be.

**Recommended next skills**
- **02.01 Survey Questionnaire Design.** Builds the response instrument around the stimulus and the exposure design specified here.
- **02.07 Scale and Measurement Selection.** Settles the appeal, uniqueness, relevance and intent scales, and their comparability to any norm being used.
- **03.04 Fieldwork Monitoring and Response Quality.** Watches for stimulus-specific problems in field, including position effects and comprehension failures visible in the open ends.

**Runs well alongside**
- **06.01 Conjoint and MaxDiff Analysis** and **06.02 Pricing Research Analysis**, wherever trade-off, price sensitivity or share estimation is the real objective.
- **12.03 Research Report Compilation**, which publishes the exposure specification and the reporting caveats drafted here.
- **08.01 Finding to Insight Development**, downstream, which inherits the constraint that the result is relative rather than absolute.

---
A Yazi Supplied Skill and resource.
