---
name: scale-and-measurement-selection
description: >
  Chooses how a construct is measured: scale type, number of points, labelling,
  midpoints, single versus multi-item, and the statistic it will produce. Use when
  someone says "how many points should this scale have", "should there be a
  midpoint", "agree or disagree or something else", "5 point or 10 point",
  "should we use an existing scale", "can we change the scale on the tracker", or
  "should we report a net score".
category: 02 Instrument Design
ref: 02.07
tier: 1
inherits: [K2, K3, K4, K5]
---

# Scale and Measurement Selection

## 1. One-line description
Decides how each construct in an instrument is measured, from what the construct actually is through to the statistic that will be reported, so that the scale supports the analysis the study has promised rather than the one that happens to be convenient.

## 2. What this skill is used for

**The research problem it solves.** Scale decisions are usually made by habit and discovered at analysis. Someone reaches for a five-point agreement battery because the last study had one, and by the time anyone examines it the data has three properties nobody chose: a heavy acquiescence component, a distribution with 80% of respondents in two points, and a mean that cannot be compared with anything. Underneath the habit sit real decisions that were never made. Whether the thing being measured is one construct or several, and therefore whether one item can carry it. Whether the scale runs from nothing to a lot or from one pole to its opposite, which is a different question and gets a different answer. How many points, which is contested and usually settled by preference rather than by what the analysis requires. Whether there is a genuine neutral position or the midpoint is a parking space. Whether an agreement frame is being used because agreement is the construct or because it was easy to write. And what statistic will eventually be reported, which is the decision that should have driven all the others and is almost always made last. Two further failures are more expensive than any of these: changing a scale in a tracker without bridging it, which destroys a series that took years to build, and reducing a distribution to a net or difference score, which discards most of the information and produces a number that moves for reasons nobody can decompose.

**Where it sits in the research lifecycle.** During instrument design, after the objectives and the analysis plan exist and before the questionnaire is finalised. It also runs as a repair pass on an existing instrument, and as the decision process when a tracker's measurement has to change.

**Typical use cases.**
- Choosing the response frame for each construct in a new questionnaire.
- Fixing an instrument with four different scale lengths and two label sets across comparable questions.
- Deciding whether a construct needs a validated multi-item battery or whether one item will do.
- Settling the midpoint, point-count and labelling arguments that recur in every design review.
- Deciding whether and how to change a scale in a tracker, and how to bridge the change.
- Assessing whether a proposed net, index or difference score is a defensible summary of the data.

**Who uses it.** Research managers and directors designing or reviewing instruments; analysts who will have to work with whatever the scale produced; client-side insight managers deciding what their organisation's standard measures should be; academic and UX researchers choosing between writing a measure and adopting an existing one.

## 3. When to use it

- A questionnaire is being drafted and the response frames have to be chosen rather than inherited.
- An instrument contains inconsistent scales and someone has to decide the convention before it is fielded.
- A construct matters enough that the study's conclusion depends on measuring it well.
- The analysis plan requires means, correlations, regression, factor analysis or segmentation, all of which constrain the scale.
- A tracker's scale is being questioned, or a change is being proposed.
- Cross-market comparison is planned, and response-style differences will sit inside every scale mean.
- A headline metric is being defined for an organisation, which will then be reported for years.
- A net score, index or gap score is being proposed and nobody has examined what it discards.

## 4. When NOT to use it

- **The construct has not been defined.** A scale cannot be chosen for something nobody has named precisely. "Engagement", "trust", "loyalty" and "satisfaction" are labels for families of constructs, and choosing a scale before deciding which member of the family is being measured produces a precise number about an unspecified thing. Go back to the objectives, and to **01.07 Analysis Plan Development** for what the measure has to support.
- **The objective needs a trade-off, a price or a preference share.** Rating scales cannot produce these, and no number of points fixes it. Stated importance ratings in particular are among the least discriminating measures in commercial research, because nearly everything is rated important. This needs a designed exercise specified in category 06, not a better scale.
- **The measurement should be behavioural.** Where behaviour is recorded in transaction, usage, log or administrative data, a scale measuring recalled or claimed behaviour is a worse measurement at higher cost. Reserve scales for what cannot be observed.
- **You are auditing an instrument rather than designing its measurement.** For a systematic bias review, including scale balance and missing options, **02.04 Question Bias Detection**. This skill decides what the scale should be; that one detects that the existing one is unbalanced.
- **The instrument's problem is length, routing or population.** A better scale does not rescue an instrument that is too long (**02.01**), routes people into the wrong bases (**02.05**), or is being answered by the wrong people (**02.06**).
- **A tracker is in field and the wave is committed.** Changing a scale mid-wave produces two incomparable halves of one dataset. The correct output is a change plan for a future wave, with a bridging design, not an improvement applied now.
- **The scale is externally mandated.** Regulatory instruments, statutory reporting measures, funder-specified questions and sector-standard collections are frequently poorly designed and are not yours to fix. Document the defect, report its direction, and measure it as specified.
- **Nobody can say what statistic will be reported.** If the answer to "what will you do with this number" is not available, the scale decision cannot be made correctly and should not be guessed. Ask, and if no answer comes, choose the most analytically flexible option and record that the choice was made without the requirement.

## 5. Required inputs

**Required. Without these the skill cannot run.**
- **The construct, defined precisely enough to be measured.** Not the topic. What exactly is being measured, of what object, over what period, in whose judgement. Without this, stop and ask.
- **The planned analysis and the statistic to be reported.** Whether the output is a top-box percentage, a mean, a distribution, an input to a regression or factor model, a segmentation variable or a tracked series. This determines the level of measurement required, and it is the requirement that most often goes unstated.
- **Mode and device profile.** Point count, label style and layout are all constrained by whether the scale is read aloud, seen on a small screen or presented in a grid.

**Optional, and what each one adds.**
- **Previous wave or comparable instrument.** Establishes what is locked for comparability, which usually outweighs improvement.
- **External benchmarks the study must match.** A norm set, a sector collection or a client's own historical series constrains the scale to a specific form, and that constraint should be visible rather than discovered.
- **Existing validated instruments for the construct.** Lets the study inherit established reliability and validity evidence instead of manufacturing a new measure with none.
- **Expected distribution or previous data.** Predicts whether a scale will discriminate. A measure on which 92% choose the top two points is not measuring anything useful, and this is knowable in advance where prior data exists.
- **Markets and languages.** Determines whether scale means will be comparable, and whether label intensity survives translation.
- **Literacy, age and accessibility profile of the population.** Constrains point count, label complexity and the viability of numeric-only scales.

## 6. Questions to ask before starting

1. **What statistic will be reported from this question, and by whom?** The governing question. A top-box percentage, a mean, a distribution, a driver-model input and a tracked series impose different requirements, and choosing the scale without knowing is the origin of most measurement regret. Default if unanswered: choose the form that supports the widest range of analysis, prefer more points over fewer within the usable range, and record that the requirement was not supplied.
2. **Is this one construct or several?** Determines whether one item can carry it. Default: treat any construct that a respondent could reasonably answer differently about different aspects as multi-dimensional, and either measure the dimensions separately or name the compromise.
3. **Does the construct have a genuine opposite, or only an absence?** Determines unipolar versus bipolar, which is the decision most often made wrongly and least often noticed. Default: unipolar, because absence is the more common structure and a false bipolar scale forces respondents to choose a side they do not hold.
4. **Will this be compared to anything?** A previous wave, another market, an external benchmark, another question in the same instrument. Comparability constrains the scale more tightly than quality does. Default: assume within-instrument comparison at minimum, and fix one convention across comparable questions.
5. **Is there an existing validated measure for this construct?** Default: check before writing one. Inheriting a validated measure inherits its evidence; writing a new one means the study must establish its own.
6. **What is the mode, and what is the expected literacy and language profile?** Default: assume small-screen self-completion, and cap aural presentation at about four or five points with short repeatable labels.
7. **Where will the distribution sit?** A measure everyone answers at the top does not discriminate, however well constructed. Default: assume a positive skew on any evaluation of a service the respondent chose, and design for discrimination at the top end.

## 7. Step-by-step methodology

**Step 1. Define the construct, then test whether one item can carry it.** Write what is being measured in a sentence that names the object, the aspect and the period: not "satisfaction" but "satisfaction with the outcome of the most recent contact, at the time it concluded". Then apply the dimensionality test: could a reasonable person hold different positions on different aspects of this construct at once? If yes, a single item forces them to average, and the average is uninterpretable because the analyst cannot tell which aspect produced it. Overall evaluations of a whole relationship, abstractions like trust and value, and anything the client's internal model treats as having components, all fail this test. Where the construct is multi-dimensional, the honest options are to measure the dimensions separately, to adopt a validated multi-item battery, or to use a single item and record explicitly that it is a global judgement rather than a composite. A correct result is a written construct definition with a stated dimensionality decision.

**Step 2. Look for an existing validated measure before writing your own.** A published, validated multi-item measure brings reliability and validity evidence, comparability to other studies, and a defence against the criticism that the measure was invented to produce a result. Writing your own brings none of these and imposes an obligation to demonstrate that the items hang together. Use an existing instrument when the construct is one that has been studied (attitudinal constructs, wellbeing, workload, usability, organisational climate), when comparability to a literature matters, or when the finding will be scrutinised. Write your own when the construct is specific to the client's context, when no validated measure exists for it, or when the validated instrument is too long for the study and truncating it would break the validation anyway. Three disciplines apply when adopting one. Use it in full and unmodified where possible, because a shortened or reworded battery is a new instrument with the original's reputation. Where modification is unavoidable, say so and stop claiming the original's properties. And check that the validation population resembles the study population, because a measure validated on undergraduates in one country is not automatically valid on retired people in another.

**Step 3. Decide unipolar or bipolar, from the construct rather than from habit.** A unipolar scale runs from none to a maximum of one thing: not at all useful to extremely useful, never to always, none to a great deal. A bipolar scale runs from one pole through a neutral centre to its opposite: very dissatisfied to very satisfied, much worse to much better, strongly disagree to strongly agree. The test is whether the construct has a genuine opposite or merely an absence. Usefulness has an absence, not an opposite: something can fail to be useful without being harmful, so a bipolar usefulness scale invites respondents to endorse a position that does not exist. Satisfaction genuinely has an opposite. Getting this wrong has a specific and detectable consequence: on a mis-specified bipolar scale, respondents who hold the "none" position cluster on the midpoint or the first negative point, and the resulting distribution has a shoulder that looks like mild negativity and is actually absence.

**Step 4. Set the point count against the analysis, within the evidence.** The evidence supports a range rather than a single answer, and where it is contested this should be said rather than resolved by preference. Reliability rises with the number of points and plateaus in the region of five to seven, with little gain beyond that on most constructs; discrimination and sensitivity to small change continue to improve modestly with more points, which matters for trackers and for correlational work; and respondent burden and error rise once the number of points exceeds what someone can hold in mind, particularly when the scale is read aloud or presented on a small screen. Practical consequences. For a construct that will be reported as a distribution or a top-box percentage, five to seven points is usually sufficient and easier to label meaningfully. For a construct feeding correlation, regression or factor analysis, more points give more variance to work with, and seven to eleven is defensible. For aural administration, four or five points with short labels, because a longer scale becomes a memory test. For small screens, fewer points, presented vertically. For tracked measures where the study needs to detect small movements, more points help, and the decision is worth more attention because it will not be revisited for years. Longer scales, including ten and eleven point formats, are frequently chosen for benchmark comparability rather than for measurement reasons, and that is a legitimate reason as long as it is the stated one. State the reasoning in the design rationale, because the point count is the decision most likely to be challenged and least likely to have been justified.

**Step 5. Decide the midpoint deliberately.** A midpoint belongs where a genuine neutral or "neither" position exists and some respondents hold it: satisfaction, agreement, comparison against a benchmark. It does not belong where the construct is unipolar, because there is no middle between none and a lot, only a quantity. Removing a midpoint where one genuinely exists does not manufacture opinion; it pushes the neutral respondents to whichever adjacent point is more socially comfortable, usually the positive one, and inflates the top-box. Including one where none exists gives satisficing respondents a resting place and creates a spurious mass in the centre. The separate decision is whether the neutral respondents and the no-opinion respondents need to be distinguished, and they usually do: a midpoint labelled "neither satisfied nor dissatisfied" is a position, and "don't know" or "not applicable" is the absence of one, so they are different options with different placements, and collapsing them makes both uninterpretable.

**Step 6. Label the points, and check label intensity as carefully as point count.** Three options. Fully labelled scales, where every point carries a verbal label, give the most consistent interpretation across respondents and are the default for five to seven points. Endpoint-labelled numeric scales, where only the extremes are named, are usual above about seven points because inventing eleven distinct verbal labels produces overlapping and non-equidistant meanings, and they carry a cost: respondents treat the numbers as if they were equally spaced, which is an assumption rather than a fact. Partially labelled scales, with endpoints and midpoint named, are a middle position. Whichever is chosen, check the intensity symmetry of the labels, which is a more frequent defect than an unbalanced point count and far less visible: "excellent, very good, good, fair, poor" has an even number of points on each side of nothing and is still badly unbalanced, because three of the five labels are positive and the negative end stops at "poor" rather than reaching the intensity of "excellent". Symmetry means the distance from the centre to each end is comparable in both direction and force. Also fix the direction convention once for the whole instrument (positive on the left or positive on the right, and consistent), because mixed directions produce respondent error that is indistinguishable from real variance.

**Step 7. Prefer construct-specific frames over agreement frames.** An agree/disagree battery measures agreement, and agreement carries an acquiescence component: a general tendency to agree with statements, stronger in some populations than others, which inflates every item in the same direction and correlates the whole battery so that it looks like a construct. The construct-specific alternative asks about the construct directly, with a scale built for it: "how easy or difficult was it to..." rather than "the process was easy: agree or disagree"; "how often does this happen" rather than "this happens frequently: agree or disagree". Construct-specific frames outperform agreement frames on interpretability, on comparability, and on acquiescence resistance, at no cost in length. Where an agreement battery is unavoidable, because it is validated or tracked, mitigate: reverse the direction of some items and check at analysis that reversed items behave as expected, keep the battery short, and avoid running several same-direction batteries in sequence. Where cross-population comparison is planned, treat acquiescence as a design-level risk, since differences in response style across groups can exceed the substantive differences being measured.

**Step 8. Handle frequency scales and vague quantifiers.** Vague quantifiers ("often", "sometimes", "rarely", "regularly") are interpreted differently by different respondents, and the interpretation shifts with the base rate of the behaviour, so "often" for a daily act and "often" for an annual one describe different frequencies while producing the same data point. Prefer absolute frequencies with defined units and a defined window ("in a typical week, on how many days did you..."), which are interpretable and comparable. Two cautions. Absolute frequency questions inherit all the recall problems of counting a routine low-salience behaviour, so shorten the window until the respondent can reconstruct it, and where they cannot, ask about the most recent occasion instead. And the response ranges themselves anchor: respondents infer the normal range from the options offered and adjust toward the middle of it, so a set of ranges centred too high or too low biases the reported frequency. Where vague quantifiers are unavoidable, define them numerically in the question.

**Step 9. Decide behavioural versus attitudinal measurement for each construct.** Where a construct can be measured either way, prefer the behavioural, because behaviour is checkable, more stable, and more predictive of behaviour than attitude is. "Have you recommended this to anyone in the past six months, and to whom" is a different and better measure than "how likely are you to recommend". Attitudes earn their place where behaviour is unobservable, where the attitude precedes a behaviour the study wants to anticipate, where the diagnostic value is in the reasoning rather than the act, or where behaviour is too rare to measure in the sample. State which is being measured, because reports routinely present an attitudinal measure as though it described behaviour, and per **K3 §3.4** the distance between stated intention and behaviour is a confidence cost that must be carried into every downstream claim.

**Step 10. Match the scale to the analysis, explicitly, before finalising.** Walk through the planned analysis and check the scale supports it. Means require an assumption of interval spacing that fully labelled ordinal scales do not strictly meet, which is widely accepted in practice and should be stated rather than hidden, and it fails harder on short scales and on scales with unevenly spaced labels. Top-box and top-two-box percentages need enough points that the box is a meaningful subset, and they discard the rest of the distribution, so report them alongside the distribution rather than instead of it. Correlation, regression and factor analysis benefit from more points and suffer badly from floor and ceiling effects, so a measure everyone answers at the top will produce a null result that is an artefact of the scale. Significance testing on scale means requires that the scale is identical across the groups compared. Segmentation on scale items is affected by response style, and unstandardised scale data can produce segments that differ mainly in how enthusiastically people use scales. Where the scale cannot support the planned analysis, change one of them now, and say which.

**Step 11. Interrogate any net, index or difference score before agreeing to it.** Composite summaries are attractive to stakeholders and expensive analytically, and the costs should be stated once, plainly. A net (positive category minus negative category) discards the middle of the distribution and the intensity within categories, so two very different distributions produce the same net, and a movement in the net cannot be decomposed without going back to the underlying data. Nets are also more volatile at small bases than the components they are built from, because two proportions each carry sampling error and the net compounds both. A difference score (importance minus performance, expectation minus experience, ideal minus actual) inherits the measurement error of both components and typically has lower reliability than either, and it is frequently dominated by whichever component varies more, which is usually not the one the client cares about. An index summing several items assumes the items belong together and are equally weighted, and neither should be assumed without evidence. The rule that follows is not "never use them" but: always report the underlying distribution alongside the composite, never let a composite be the only number that travels, state the base at which it becomes unstable, and where a composite is tracked, watch for movements driven by one component while the headline stays flat, which is the failure mode that most often goes undetected.

**Step 12. Fix the conventions across the instrument, and document the specification.** One direction, one point count per construct type, one label set per scale family, one convention for non-substantive options. Questions that will be compared with each other must use identical scales, and questions that will not be compared should be allowed to differ if the construct requires it, rather than being made uniform for tidiness. Then write the scale specification: for each construct, the definition, the scale form, the point count, the labels verbatim, the direction, the midpoint decision with its reason, the non-substantive options, the statistic to be reported, and the source if the measure was adopted from elsewhere. This document is what stops the instrument drifting during stakeholder review, and it is what the analyst reads when deciding how to treat the variable.

**Step 13. If this is a tracker, treat any change as a break and bridge it.** Changing a scale changes the series, and the movement introduced by the change is indistinguishable from a real movement. Establish first whether the change is worth it: the gain from a better measure is usually smaller than the loss of a comparable trend, and the honest answer is often to keep the flawed measure and disclose its flaw. Where a change is genuinely necessary, bridge it. The defensible approach is a parallel run: field both versions in the same wave to random halves of the sample, which allows the relationship between them to be estimated on a single sample at a single point in time, with the two split cells comparable in every other respect. Report both series for at least the bridging wave, and preferably longer. Do not retro-convert history using a conversion factor unless it was calibrated on a parallel run, because a conversion estimated from two different waves confounds the scale change with real change. Document the break in the trend documentation, so that a reader in three years knows why the series has a seam, and per **K5 §2.7** the decision to break a trend is the researcher's, made with the cost visible.

## 8. Analytical framework

Every measurement decision runs down one chain, and it must be readable in both directions:

    Construct → Dimensionality → Measurement mode → Scale form → Points and labels → Level of measurement → Analysis operation → Reported statistic

**Forward** is the design test: this construct has this many dimensions, is best measured behaviourally or attitudinally for this reason, in this scale form, with this many points labelled this way, producing data at this level of measurement, which supports this operation, which yields this reported statistic.

**Backward** is the discipline that catches the errors, and it should be run first in practice: this number will appear on this chart, produced by this operation, which requires data at this level, which requires this many points in this form, which measures this construct. Working backward from the reported statistic prevents the most common failure in the whole skill, which is a scale chosen from habit and then asked to support an analysis it cannot carry.

Two rules govern the chain. **The scale is chosen by the analysis, not by convention**, and where convention wins (because a benchmark or a trend requires it), that is a comparability decision and should be recorded as one. And **every step of reduction loses information that cannot be recovered**: a distribution becomes a mean, a mean becomes a net, a net becomes a single arrow on a slide, and per **K2 §7** each reduction is a traceability step that must be reversible back to the distribution it came from.

## 9. Output format

**1. Measurement summary.** The constructs measured, the conventions fixed for the instrument, and the constraints applied (mode, benchmarks, tracked measures).

**2. Scale specification table**, the core deliverable.

| Construct | Definition (object, aspect, period) | Dimensionality | Behavioural or attitudinal | Scale form | Points | Labels verbatim | Direction | Midpoint and reason | Non-substantive options | Statistic to be reported | Source if adopted |
|---|---|---|---|---|---|---|---|---|---|---|---|

**3. Design rationale notes.** Short justifications for the decisions a reviewer will query: why this point count, why a midpoint is present or absent, why an agreement frame was or was not used, why an existing instrument was or was not adopted.

**4. Analysis compatibility check.** For each construct, the planned analysis and confirmation that the scale supports it, or a note of the mismatch and how it was resolved.

**5. Composite measures assessment**, where any net, index or gap score is proposed: what it discards, its stability at the study's base sizes, and what must be reported alongside it.

**6. Comparability register.** What is locked to a previous wave or an external benchmark, what is free, and every deliberate break with its bridging plan.

**7. Open items and review points**, per **K5 §3**.

**When the inputs are thin**, the format is not filled in anyway. A construct whose definition has not been agreed is recorded as unresolved rather than given a plausible scale. A statistic that nobody has specified is recorded as `[not supplied]`, and the flexible choice made in its absence is stated as such. Where an existing validated instrument is referenced, it is cited only if verified per **K4 §2.4**; an unverified recollection of a measure is not a source.

## 10. Quality checks

Run before the instrument is finalised. These sit on top of **K4 §8**.

1. Every construct has a written definition naming the object, the aspect and the period.
2. Every construct has a stated dimensionality decision, and every single-item measure of a multi-dimensional construct is flagged as a global judgement.
3. Unipolar or bipolar has been decided from the construct, not inherited.
4. The point count is justified against the planned analysis and the mode, and the justification is recorded.
5. The midpoint decision is deliberate, and neutral is distinguished from no-opinion where both exist.
6. Label intensity is symmetric either side of the centre, not just the point count.
7. Scale direction is consistent across the instrument.
8. Every agreement battery has been challenged, and the ones that remain are justified by validation or trend, with acquiescence mitigation applied.
9. Frequency questions use defined units and windows, or any vague quantifier is defined numerically in the question.
10. Behavioural and attitudinal measures are distinguished, and no attitudinal measure is labelled as behaviour.
11. Every scale supports the statistic it will produce, checked by walking the chain backwards from the chart.
12. Every proposed net, index or difference score has been assessed for information loss and base stability, and the underlying distribution is reported alongside it.
13. Questions that will be compared use identical scales; questions that will not are allowed to differ where the construct requires it.
14. Every tracked measure is unchanged, or the change carries a bridging plan and is logged as a break.
15. Any adopted validated instrument is used in full and unmodified, or the modification is disclosed and the original's properties are no longer claimed.

## 11. Common failure modes

| Failure | How to recognise it | How to prevent it |
|---|---|---|
| **Scale by habit** | Every construct measured on the same five-point agreement battery | Choose from the construct and the analysis. Uniformity is not consistency |
| **The unnoticed unipolar construct** | A bipolar scale on something with no opposite, and a strange shoulder in the distribution | Step 3's test: absence or opposite |
| **Midpoint by default** | A midpoint present or absent with no reason recorded, and no separate no-opinion option | Step 5. Decide it, record the reason, and keep neutral separate from don't know |
| **Asymmetric labels** | "Excellent, very good, good, fair, poor" passing review because it has five points | Check intensity, not just count. This is the most common surviving scale defect in commercial research |
| **Ceiling effect** | 92% in the top two boxes and no variance to analyse | Predict the distribution from prior data, and design for discrimination where the mass will sit |
| **Analysis-blind scale** | A driver model built on a three-point scale, or a mean reported from a two-point item | Walk the chain backwards from the reported statistic before finalising |
| **Vague quantifiers** | "Often", "regularly", "sometimes" in a frequency question | Absolute frequency with a defined window, or a numeric definition in the question |
| **Net as the only number** | A single composite on the slide, with the distribution nowhere in the deck | Always report the distribution alongside, and state the base at which the composite becomes unstable |
| **Gap score reliability** | An importance-minus-performance score treated as a precise diagnostic | Difference scores compound the error of both components. Report both components |
| **Silent tracker change** | A scale improved between waves and the movement reported as a market change | Step 13. Parallel run, dual reporting, documented break |
| **Retro-conversion** | Historical data converted to a new scale with a factor estimated across waves | A conversion not calibrated on a parallel run confounds the scale change with real change |
| **Truncated validated battery** | A ten-item validated measure cut to four, still described by the original's name | A shortened battery is a new instrument. Disclose it and stop claiming the original's properties |
| **AI: symmetric batteries** | Every construct given the same agree/disagree grid because it looks consistent | Construct-specific frames by default; agreement only where validated or tracked |
| **AI: invented psychometric claims** | Statements that a scale is reliable, valid, or standard, with no source | **K4 §2.4**. Reliability and validity are properties demonstrated on a population, not attributes of a format |
| **AI: fabricated norms and benchmarks** | "Scores above X are considered good", with a number attached | **K4 §2.4, §2.5**. Benchmarks come from a named, verified source or are not stated |
| **AI: plausible reconstruction of a validated instrument** | Items from a known measure reproduced from memory | **K4 §2.4**. A partially remembered instrument is not the instrument. Cite the verified source or mark as unavailable |

## 12. AI guardrails

Skill-specific only. **K4** applies in full and is not repeated here.

1. **Never reproduce the items of a validated instrument from memory.** A partially recalled battery is a new, unvalidated measure carrying a validated measure's name, which is the worst of both. Cite a verified source or state that the instrument must be obtained (**K4 §2.4**).
2. **Never claim reliability, validity or norm status for a scale without a verified source.** These are properties demonstrated on a specific population, not attributes that attach to a scale format (**K4 §2.4, §2.5**).
3. **Never state a benchmark, threshold or "good score" for any metric.** Thresholds are population-specific and sector-specific, and asserting one converts a measurement into a judgement nobody has evidenced.
4. **Never recommend a scale change to a tracked measure without stating the trend cost and the bridging requirement.** The loss of comparability is usually the larger effect, and it is the researcher's decision (**K5 §2.7**).
5. **Never present a net, index or difference score without stating what it discards and where it becomes unstable.** A composite delivered without its distribution is a number that cannot be interrogated by anyone downstream.
6. **Never assume the direction, polarity or point count of a scale carried over from another instrument.** Confirm it or mark it unverified (**K4 §6.2**), because a reverse-coded item treated as forward-coded inverts a finding silently.
7. **Never choose a scale to produce a more attractive number.** Selecting a point count, a midpoint or a box definition because it improves a headline is instrument-level cherry-picking (**K4 §4.2**).
8. **Never present a mean from an ordinal scale without acknowledging the interval assumption where it materially matters**, particularly on short scales, on scales with unevenly spaced labels, and in any cross-market comparison where response style differs.

## 13. Best-practice principles

1. **Choose the scale from the statistic, working backwards.** The chart title, then the operation, then the level of measurement, then the scale. Forward from the topic produces measures that cannot be analysed.
2. **Ask whether the construct has an opposite or only an absence.** This one question resolves most unipolar and bipolar arguments, and getting it wrong produces a distribution that misleads quietly.
3. **Prefer item-specific frames to agreement.** "How easy or difficult was it" beats "it was easy: agree or disagree" on acquiescence, interpretability and comparability, at no cost in length.
4. **Label intensity matters more than point count and gets less scrutiny.** A five-point scale with three positive labels is unbalanced no matter how symmetric the arithmetic looks.
5. **Use an existing validated measure when one exists and fits.** Inheriting evidence is cheaper than generating it, and a measure with a literature behind it survives scrutiny that a bespoke one does not.
6. **Consistency where things are compared, difference where constructs differ.** Uniformity applied for tidiness forces the wrong scale onto half the instrument.
7. **A measure everyone answers at the top is not measuring.** Predict the distribution before fielding, and design for discrimination where the mass will actually sit.
8. **Every reduction discards information permanently.** Distribution to mean to net to arrow. Each step is a choice, and each should be visible and reversible back to the data.
9. **Nets and gap scores are less stable than the numbers inside them.** Two proportions each carry error, and the composite compounds both, which is why nets on small bases move for no reason.
10. **In a tracker, the flawed measure you can compare usually beats the better measure you cannot.** Improve at the point of a redesign with a bridge, not opportunistically.
11. **Behaviour beats attitude wherever both are available.** Stated intention is a measure of a different thing, and the report must say which one was measured.
12. **Write the scale specification down.** The document is what stops a scale drifting during stakeholder review and what the analyst reads when deciding how to treat the variable three months later.

## 14. Worked example

**INPUT**

A fictional university, Ashgrove University, runs an annual student experience survey used for internal quality review and for a published summary. The current instrument measures teaching quality, workload manageability, support services and belonging using a single 22-item agree/disagree battery on a five-point scale, all items worded positively. A single "would you recommend Ashgrove to a friend" item on an eleven-point endpoint-labelled scale is reduced to a net of the top categories minus the bottom ones and reported as the headline. The academic board has asked for the instrument to be improved. Nine waves of history exist.

**PROCESS**

*Step 1, constructs.* Four constructs were named and defined. Three failed the dimensionality test immediately: support services covers academic advice, wellbeing provision and administrative support, which any student could reasonably rate differently, so a single agreement item forces an average that cannot be interpreted and cannot be acted on, since the review committee needs to know which service is the problem. Belonging is also multi-dimensional and is the construct with the most established measurement literature.

*Step 2, validated measures.* For belonging and for workload, established multi-item measures exist in the higher education literature. The recommendation was to adopt one for belonging in full and unmodified, with the source verified before use rather than reproduced from recollection, on the grounds that the construct is well studied, the finding will be published, and a bespoke measure would have no evidence behind it. For workload, the shortlisted instruments were longer than the survey could carry, and truncating one would have produced an unvalidated measure trading on a validated name, so a bespoke behavioural measure was proposed instead: hours in a defined recent week, plus a single manageability judgement, with the compromise recorded.

*Step 7, the agreement battery.* All 22 items positively worded and same-direction is the highest-acquiescence structure available. Recommended conversion of the diagnostic items to construct-specific frames ("how clear or unclear were the assessment criteria" rather than "the assessment criteria were clear: agree or disagree"), which also improved interpretability for the committee, since a proportion answering "unclear" is directly actionable and a mean agreement score is not.

*Step 4 and 5, points and midpoint.* The five-point agreement scale had a genuine neutral, so the midpoint was retained where agreement frames remained. On the new construct-specific items a seven-point fully labelled scale was proposed, because the analysis plan includes a regression of the diagnostics onto the overall measure, and five points across a positively skewed distribution left too little variance for that model to be informative. A separate "not applicable" option was added to the support services items, distinct from the midpoint, since a student who has never used wellbeing provision is not neutral about it, and the old instrument had been quietly counting those students in the middle of the distribution for nine years.

*Step 11, the headline net, and the judgement call.* The published headline was the net on the recommendation item. Three problems were documented. It discards the middle of an eleven-point distribution, so the same net arises from very different student bodies. It compounds the sampling error of two proportions, which made it visibly volatile in the smallest faculties, where it had twice been reported as a change and was within the range of noise. And it is an attitudinal proxy standing in for the construct the board actually discusses, which is whether students would choose the university again. The recommendation was not to abolish it, because it is published and comparable across nine waves, but to report the full distribution alongside it in every internal output, to state a minimum base below which the net is not reported at faculty level, and to add a behavioural item ("have you recommended Ashgrove to anyone considering university in the past year, and to whom") as a companion measure with no reporting history and no obligation to match.

*Step 13, bridging.* Changing the diagnostic items breaks nine waves of internal series. A parallel run was designed: in the next wave, half the sample randomly receives the existing agreement battery and half the new construct-specific items, with both halves comparable on every other variable, allowing the relationship between the two to be estimated on one sample at one point in time. Both series reported for that wave and the next. Explicitly rejected: retro-converting nine years of history with a factor estimated by comparing wave nine to wave ten, which would confound the scale change with whatever genuinely changed in that year. The headline recommendation item was left untouched, on the reasoning that its trend value exceeds its measurement value and its defects are now documented rather than hidden. Marked **RESEARCHER DECISION REQUIRED** per **K5 §2.7**, since breaking an internal series that feeds a published quality process is an institutional judgement.

**OUTPUT**

A scale specification covering four constructs with definitions, dimensionality decisions, forms, point counts, verbatim labels, midpoint reasoning and the statistic each produces; one adopted validated instrument with its source to be verified before use and a note that it must be used unmodified; one bespoke behavioural measure with its compromise recorded; a converted diagnostic battery with acquiescence mitigation; a composite measures assessment of the published net with a minimum reporting base and a required companion distribution; a comparability register naming what is locked, what is breaking and the parallel-run bridging design; and one **K5** decision point on breaking the nine-wave diagnostic series.

## 15. Advanced usage

**Cross-market measurement.** Response styles differ systematically across cultures: extreme responding, midpoint avoidance and acquiescence vary enough that the difference between two markets on a scale mean can be larger than any substantive difference the study is measuring. Three responses, in order of preference. Design around it: prefer behavioural measures and construct-specific frames, which are less exposed to response style than agreement scales. Design for it: use enough points that within-respondent standardisation is possible, and include measures that allow style to be estimated. And analyse for it: plan the standardisation approach in the analysis plan rather than discovering the problem in the cross-tabs. Also check translated labels for intensity, because the second point of a five-point scale is rarely equivalent across languages, and a scale that is symmetric in one language is frequently not in another.

**Building a headline metric for an organisation.** A measure adopted as an organisational headline will be reported for years, embedded in objectives, and eventually attached to incentives, at which point it will be optimised. Design for that from the outset: choose a construct that is worth optimising, prefer a measure whose components are visible so that movement can be decomposed, avoid composites that cannot be decomposed, state the minimum base for reporting at every level of the organisation, and specify at the outset what the measure does not capture. Where the metric will be used comparatively between units, check that the units face comparable populations, because a measure that differs by customer mix will rank units on their customers rather than on their performance.

**Measurement invariance and multi-item batteries.** Where a multi-item measure is compared across groups, waves or markets, the assumption that the items mean the same thing to each group is a testable one, not a given. Where the study's conclusion depends on the comparison, the test belongs in the analysis plan, and where it fails, comparing the composite means is not defensible even though it is easy. This matters most in academic work and in any published or regulated comparison.

**When the standard approach does not fit.** Populations with low literacy, young children, or cognitive impairment need shorter scales, fewer points, concrete anchors and often visual or verbal analogue formats, and the validation evidence for a measure developed on a general adult population does not transfer to them. Voice and conversational modes remove the visual scale entirely and cap the usable point count at what someone can hold in working memory while the anchors are read out. In both cases the constraint is the population and the mode rather than the ideal measure, and the honest output is a simpler scale with its limitation stated rather than a sophisticated one nobody can answer.

## 16. Skill chain

**Recommended previous skills**
- **01.07 Analysis Plan Development.** Hands over the tables, models and statistics the measurement must support, which is the input this skill cannot proceed properly without.
- **02.01 Survey Questionnaire Design.** Hands over the constructs, the draft questions and the scale decisions made at design level, which this skill settles.
- **02.04 Question Bias Detection.** Where an audit has flagged unbalanced or inappropriate scales, hands over the defects for redesign.

**Recommended next skills**
- **02.05 Survey Logic and Flow Review.** Tests the instrument once the scales are fixed, including any split-cell design carried for a parallel run.
- **05.01 Descriptive Analysis** and **05.05 Trend and Tracker Analysis**, downstream, which inherit the scale specification and the comparability register.

**Runs well alongside**
- **02.01 Survey Questionnaire Design**, iteratively, since a scale decision frequently changes the question and the length budget.
- **05.06 Correlation, Regression and Causal Claim Control**, where the analysis depends on scale properties the measurement has to deliver.
- **10.01 Literature Review and Desk Research**, when identifying and verifying an existing validated instrument for the construct.

---
A Yazi Supplied Skill and resource.
