---
name: trend-and-tracker-analysis
description: >
  Analyses change over time in a repeated study: checks the comparability
  preconditions before any wave-on-wave claim, separates real movement from
  sampling noise and seasonality, establishes a normal range of fluctuation for
  each metric, and refuses to narrate noise. Use for "has it gone up", "what
  changed since last wave", "wave-on-wave comparison", "is this movement real",
  "read the tracker", "why has the trend line moved", "we changed the question,
  can we still compare", "rebasing the tracker", or "write the tracker report".
category: 05 Quantitative Analysis
ref: "05.05"
tier: 1
inherits: [K2, K3, K4, K5]
---

# Trend and Tracker Analysis

## 1. One-line description
Establishes whether a measure has actually changed between two points in time, by first proving that the two readings are comparable at all, then separating real movement from sampling variation, seasonality, sample source drift and fieldwork conditions, and by treating "no change" as a legitimate and frequently correct headline rather than a failure to find a story.

## 2. What this skill is used for

**The research problem it solves.** A tracker exists to detect change, which creates a standing pressure to report change, and the pressure operates every wave regardless of whether anything happened. The result is a genre of research output in which a 3-point movement on a base of 500 is given a cause, a 2-point movement in the opposite direction the following wave is given a different cause, and nobody notices that the measure has been oscillating within its ordinary range for four years. Underneath that sit two more expensive failures. The first is comparability: a question was reworded, a scale was relabelled, the routing changed, the sample source shifted, or the weighting scheme was updated, and the resulting movement is an artefact of the study rather than a change in the world. These artefacts are usually larger than the real movements they hide, and they are invisible in the data. The second is seasonality: two adjacent waves straddling a seasonal boundary are compared as though timing were neutral, and the comparison measures the calendar. This skill supplies the preconditions that must hold before any wave-on-wave claim, the arithmetic that distinguishes movement from noise, and the discipline of not narrating noise.

**Where it sits in the research lifecycle.** After descriptive analysis has produced the wave's figures on the same base definitions as previous waves, and alongside statistical testing, which supplies the arithmetic for whether a movement exceeds sampling variation. Before the wave report is written, because the comparability check can invalidate the whole comparison and has to happen first.

**Typical use cases.**
- Reading a new wave of a tracker and deciding what, if anything, has moved.
- Distinguishing a real shift in a metric from the measure's ordinary fluctuation.
- Handling a necessary change to a tracked question, scale, sample or weighting scheme.
- Interpreting a movement that coincides with an external event during fieldwork.
- Establishing the normal range of variation for a metric so that a genuine shift is recognisable.
- Auditing a trend line that has moved, to check whether the movement is in the world or in the method.
- Writing a tracker report where the honest headline is that nothing changed.

**Who uses it.** Research executives and managers running continuous and wave-based studies; research directors reviewing tracker reports before they go out; client-side insight teams who receive a tracker every quarter and have to decide what to act on; public sector and NGO analysts running repeated population surveys where methodological continuity is a formal requirement.

## 3. When to use it

- A new wave of data has arrived and the question is what has changed since the last one.
- A movement in a metric needs to be classified as real, ordinary variation, or artefact.
- A change to the questionnaire, scale, sample source or weighting is proposed and someone has to decide what it costs the series.
- A trend line is being presented and the reader needs to know which of its movements are readable.
- An external event fell during or near fieldwork and its effect on the numbers has to be assessed.
- A stakeholder is asking why the number moved, and the honest first answer may be that it did not.
- A tracker is being designed or re-specified and the comparability rules need writing down before the first wave.

## 4. When NOT to use it

- **The two readings are not comparable, and the difference cannot be bridged.** Where question wording, scale, routing, base definition, sample source, weighting scheme or mode has changed materially, there is no valid wave-on-wave claim to make, and running the comparison anyway produces a confident number about nothing. The correct output is a statement that the series is broken at this point, what broke it, and what the options are. This is the most important refusal in the skill and it is the one most often overridden.
- **The comparison is between two studies rather than two waves of one.** Two surveys run by different teams on different samples with similar-sounding questions are not a trend, whatever the dates on them. Comparing them is a synthesis problem with a comparability assessment attached, and it belongs to **10.03 Multi-Source Research Synthesis**, not here.
- **The question is whether an intervention caused the change.** A before-and-after reading establishes that a number differs at two times. It does not establish why, because everything else also changed between the two dates. A causal claim needs a design that supports it: see **05.06 Correlation, Regression and Causal Claim Control** for what licenses one, and **06.04 Experiment and A/B Test Analysis** where a control group exists. A tracker with no control is the classic case of temporal order without confounder control.
- **The base is too small to detect anything the client would care about.** On bases of 150 per wave, movements smaller than roughly 15 points are invisible, which means most real change in most metrics is undetectable. Running the comparison produces a series of null results that will be read as stability. Compute the minimum detectable movement first (**05.02**, Step 11) and say what the tracker can and cannot see, before the wave is reported rather than after.
- **The metric is a composite or index whose construction changed.** An index rebuilt with different components, different weights or a different base is a different index. Report it as a new series, or reconstruct the previous waves on the new definition and show both.
- **Only two waves exist.** With two points there is no normal range of fluctuation, so a movement cannot be judged against the measure's own behaviour, only against sampling error. Say so: a difference between waves one and two is a difference, not a trend, and the word "trend" should not be used until there are enough points to see a direction that persists.
- **The purpose is to find a movement to talk about.** Where nothing has changed and the report must nonetheless contain news, no test and no caveat rescues the output. The honest deliverable says the measures were stable, quantifies what stability means for this series, and reports what the tracker did detect, including the absence of movement where movement was expected.
- **The wave's own data quality has not been checked.** A movement can be produced by a change in respondent quality, completion behaviour, or fieldwork supplier performance. Run **03.04 Fieldwork Monitoring and Response Quality** and **04.01 Data Validation** first; a data quality shift is a comparability failure like any other.

## 5. Required inputs

**Required. Without these the skill cannot run. If absent, ask; if no answer is available and work must proceed, state the assumption at the point where it bites, per K5 §5.**

- **The current wave's figures, on stated bases**, produced under the tracker's own conventions.
- **The corresponding figures for previous waves**, with their bases, ideally for as many waves as exist rather than only the last one.
- **The questionnaire as fielded in each wave being compared**, not a summary of it. Wording changes are the largest single source of false movement and they are invisible in a data table.
- **The base definitions and routing for each wave.** A base that quietly widened or narrowed produces a movement that is a redefinition.
- **Fieldwork dates for each wave**, which determine seasonality and whether an external event fell inside the field period.
- **Sample source and method for each wave**, including any change of supplier, panel composition, recruitment route, incentive or device mix.
- **The weighting scheme for each wave**, including target sources and any change to the targets themselves.

**Optional, and what each one adds.**

- **The full history of the metric, wave by wave.** Allows the normal range of fluctuation to be established, which is the single most valuable thing in this skill and cannot be substituted for by a significance test on the latest pair.
- **A parallel run of both question versions in one wave.** Where a change was unavoidable, this is what makes bridging possible rather than guesswork. It is the difference between a series that continues and a series that restarts.
- **Effective bases and design effects per wave (04.05).** Allow correct testing of movements on weighted data, and reveal a change in weighting efficiency that can itself produce apparent movement.
- **A record of external events during each field period.** Turns an unexplained spike into a documented one, and prevents an event being invented retrospectively to explain noise.
- **Category, market or population benchmarks over the same period.** Distinguish a movement specific to the subject from one affecting everything, which is frequently the more important reading.
- **Behavioural or operational data covering the same period.** Provides an independent check on whether a stated movement corresponds to anything observable.
- **The tracker's original design documentation.** States which metrics were designed to move and which were designed to be stable, which changes how a flat reading should be interpreted.

## 6. Questions to ask before starting

1. **Has anything changed about how this measure was produced since the last wave?** The precondition question, covering wording, scale, routing, base, sample source, weighting, mode and fieldwork conditions. *Default if unanswered:* compare the questionnaires and the base definitions directly rather than relying on an assurance, and where documentation is absent, state that comparability could not be verified and report the movement as provisional.
2. **How much does this metric normally move when nothing has happened?** Determines whether the current movement is news. *Default:* compute the wave-on-wave changes across the full history and report the current movement against that distribution; if fewer than about six waves exist, say the normal range is not yet established.
3. **When was each wave in field, and is there a seasonal pattern in this metric?** Determines whether adjacent waves are comparable at all and whether year-on-year is the right comparison. *Default:* compare like month with like month where the history allows it, and flag any comparison that crosses a known seasonal boundary.
4. **What happened externally during or just before fieldwork?** Determines whether a movement has a documented candidate explanation, and prevents one being invented later. *Default:* record events found in the field period, state that the association with the movement is untested, and do not attribute.
5. **What movement would matter to the decision this tracker informs?** Sets the materiality threshold, without which every movement gets equal narrative weight. *Default:* state that no threshold was agreed, report the movement with its interval, and flag the materiality judgement as a **K5** review point.
6. **Is the reading point-in-time or cumulative?** Determines what a movement means, because a rolling figure and a discrete wave figure move differently and cannot be compared with each other. *Default:* report discrete waves, and where a rolling figure is used, state the window and never place the two on one chart unlabelled.
7. **What is the minimum movement this wave could have detected?** Determines whether a null result is informative. *Default:* compute it and state it beside every "no change" statement.

## 7. Step-by-step methodology

**Step 1. Run the comparability check before looking at any number.** Seven preconditions must hold before a wave-on-wave claim is admissible. Check each against documentation, not memory.

- **Question wording identical**, including the preamble, the examples, and any change in the order of the response list, which affects selection independently of wording.
- **Scale identical**: same number of points, same labels, same labelled endpoints, same direction, same presence or absence of a midpoint and a don't know.
- **Routing identical**: the same respondents reach the question by the same path, and no filter was added or removed upstream.
- **Base definition identical**: the same people are in the denominator, with the same treatment of don't know and item non-response.
- **Sample source identical**: same recruitment route, same supplier, same panel composition, same incentive structure, same device and mode mix.
- **Weighting scheme identical**: same variables, same targets, same source for the targets, and no change in the target data itself.
- **Fieldwork conditions comparable**: same length of field period, same reminder pattern, no material difference in response rate or completion behaviour.

Record each as holding, changed, or unknown. **Unknown is not the same as holding**, and a precondition that cannot be verified is disclosed as such. *Correct result:* a seven-line comparability record for every metric taken forward, with any failure named and its likely direction of effect stated where that can be reasoned.

**Step 2. Where a precondition has failed, decide what the series can still support.** Three outcomes, and choosing between them is the substantive judgement in this step. **The comparison is invalid and is not made.** State that the series is broken, at which wave, and by what. Report the new reading as a new baseline. **The comparison is made with a stated caveat**, where the change is minor, its likely direction is known, and the movement is far larger than the change could plausibly produce. The caveat travels with the number into every downstream document (**K3 §7**). **The comparison is bridged**, where a parallel run exists. Bridging is treated in Step 10. What is never acceptable is comparing across a known change silently, because the reader has no way to detect it and the artefact will be given a business explanation. *Correct result:* an explicit disposition for every broken precondition, recorded in the tracker documentation as well as the wave report.

**Step 3. Establish the normal range of fluctuation for each metric before judging the current movement.** Take every wave-on-wave change in the metric's history and look at the distribution of those changes. A metric whose ordinary wave-on-wave movement has ranged between minus 4 and plus 5 points for sixteen waves has a noise floor of about 5 points, and a 4-point movement this wave is not news whatever a test says about it, because the same test would have flagged half the previous waves. Plot the series with a band showing that ordinary range, so that a genuine departure is visible as a departure. Two cautions: the historical range includes any real movements that occurred, so it is a conservative estimate of noise; and a metric that has genuinely trended will show a wide range for reasons that are not noise, so read the range alongside the shape of the series rather than as a number on its own. *Correct result:* a stated normal range per metric, and every current movement placed against it before anything else is said about it.

**Step 4. Test the movement rather than eyeballing it.** A difference between two waves is a difference between two samples, and it carries sampling variation like any other comparison. Test it, using the design that applies: independent samples where each wave is a fresh sample, and a paired test where the same respondents are re-interviewed, which is a different test and is chosen wrongly more often than any other selection in tracker work (**05.02**, Step 3). On weighted data use effective bases. Report the movement, its confidence interval and the test, not a flag. The interval is what makes a tracker readable: "down 3 points, interval minus 8 to plus 2" says clearly that the measure may not have moved at all. *Correct result:* every reported movement carrying its two bases, its interval and its test, and every unreported movement recorded as tested and not significant rather than silently dropped.

**Step 5. Count the comparisons this wave generates, and control them.** A tracker with 40 metrics tested wave-on-wave runs 40 tests, of which roughly two will flag at a 5% threshold with nothing happening anywhere; with a banner attached, the count runs into the hundreds. The standing pattern in tracker reporting is that the two chance flags become the story. Pre-specify the small set of headline metrics the tracker exists to monitor and test those at the nominal threshold; treat everything else as screening, correct within the family or label it exploratory (**05.02**, Step 6). *Correct result:* a stated count and a stated regime, with headline metrics separated from screening before the results are seen.

**Step 6. Check seasonality before attributing any movement.** Many measures move with the calendar: usage of a service, category purchasing, staff and customer satisfaction around annual cycles, public attitudes around budget or election timing, anything affected by weather or holidays. Two adjacent waves that straddle a seasonal boundary are not comparable in the way the reader assumes. Three rules. **Compare like period with like period** where the history allows: the year-on-year comparison is the seasonally clean one. **Establish the seasonal pattern from the history** rather than assuming it, by looking at the same calendar position across years. **Where the history is too short to establish seasonality, say the movement is confounded with timing** rather than choosing an explanation. A quarterly tracker whose Q4 reading is always the highest of the year has not improved every Q4. *Correct result:* every wave-on-wave movement checked against the same calendar position in previous years, and the year-on-year comparison reported alongside the wave-on-wave one wherever seasonality is plausible.

**Step 7. Read the three time comparisons as three different things.** **Wave-on-wave** is the noisiest and the most over-read, because it is one sample against one sample with everything that changed between two dates folded in. **Year-on-year** removes seasonality and is usually the more reliable single comparison, at the cost of being slower to detect a real shift. **The trend line** across all available waves is where actual direction lives, and it is what a tracker was built to provide. The discipline is to look at the line before looking at the latest point. A single-wave movement inside a flat line is noise; the same movement as the fourth consecutive move in one direction is a trend even if no individual step is significant, because consecutive moves in one direction are unlikely by chance. Read the shape: gradual drift, step change, oscillation, or a return to a previous level after an excursion, each of which implies something different. *Correct result:* every headline movement described in terms of what the line is doing, not only what the last two points did.

**Step 8. Check for sample source drift, which produces trends that are artefacts of recruitment.** This is the failure that most reliably produces a long, smooth, entirely false trend, and it is invisible unless looked for. Panels age, refresh unevenly, and change composition as recruitment sources shift; supplier mixes change; device and mode mixes change; incentive structures change. Because weighting corrects only the variables it targets, drift on anything not weighted passes straight through into the series. Diagnostics: track the unweighted profile of the sample wave by wave on demographics, on tenure with the panel where available, on device and mode, on completion time, and on any question that should be stable in the population. **A measure that ought not to move over eighteen months and has moved steadily is evidence of drift, not of change.** Track the weighting efficiency too: rising weight variability means the incoming sample is further from the targets, and it reduces effective base while leaving nominal base intact. *Correct result:* a sample composition trend chart maintained alongside the metric trends, and any metric movement checked against it before it is reported.

**Step 9. Record external events, and do not attribute without saying so.** List anything in or near the field period that could plausibly affect the measure: a policy announcement, a service failure, a price change, a competitor action, a news cycle, weather, a public event. Then apply the discipline. A coincidence of timing is a hypothesis, not a finding, and per **K4 §3.2** the language must not smuggle a causal claim in. Where an event is a plausible explanation, say so as a candidate, name what would test it (a split of the sample by whether they were interviewed before or after the event, if fieldwork spans it; the same metric in an unaffected market; a comparable benchmark), and run the test where the data allows. A within-wave split by interview date is often available and is the strongest evidence a single tracker can produce on its own. *Correct result:* a dated event log kept per wave, and every attribution either tested or labelled as an untested candidate explanation.

**Step 10. Handle a necessary change by bridging, and run both versions in parallel for a wave.** Sometimes a question must change: a response list is out of date, a scale is being harmonised, a sample source is closing. The professional answer is to run the old and new versions in the same wave on split samples, measure the difference between them on the same population at the same moment, and use it to state the size of the step. Two things follow. The series can be presented as continuous with the break marked and the step quantified, or the historical waves can be adjusted to the new basis with the adjustment documented, and the choice is a client decision. The parallel measurement is itself an estimate with sampling error, so the bridge carries uncertainty that must be shown rather than presented as a fixed correction. Where a parallel run was not possible, **do not construct a bridge from judgement**: state the break, restart the series, and keep the historical series available separately. *Correct result:* a documented break point, a quantified step with its uncertainty where a parallel run exists, and a series presentation that shows the break rather than smoothing over it.

**Step 11. Distinguish point-in-time from cumulative reading, and never mix them.** A discrete wave figure describes the people interviewed in that field period. A rolling or moving figure describes a window, smooths noise, lags real change by roughly half the window, and cannot be compared with a discrete figure. Cumulative figures across a period (total reach, total incidence to date) answer a different question again. State which is being reported, keep one convention per chart, and where both are useful, put them on separate charts. A rolling series is much less noisy than a discrete one, which is a presentational advantage and an analytical trap: consecutive rolling points share respondents and are not independent, so they must not be tested against each other as though they were. *Correct result:* one stated convention per chart, and no test run between overlapping windows.

**Step 12. Write the wave report without narrating noise, and let "no change" be the headline where it is true.** Report the headline metrics against their normal ranges and their intervals. State plainly which measures moved beyond ordinary variation, which did not, and what the wave could have detected. **A wave in which nothing moved is a finding**: it means the market is stable, or an intervention has not yet landed, or a decline has stopped, and each of those is actionable. The standing pressure runs the other way, and the specific failure to avoid is assigning a business explanation to each of the small movements in the table, which manufactures an account of a world that did not change. Where a stakeholder needs something to act on and nothing moved, the useful material is elsewhere: the level rather than the movement, the structure inside the sample, the gap against a benchmark, the measures that were expected to move and did not. *Correct result:* a report in which every movement narrated is a movement that survived Steps 1 to 8, and every metric that did not move is reported as stable with the detectable-movement figure beside it.

## 8. Analytical framework

    Comparability → Normal range → Movement → Significance → Seasonality → Trend shape → Source check → Claim

Read it as gates. **Comparability** is first and is absolute: a movement that fails it never reaches the second gate, whatever its size. **Normal range** places the movement against the measure's own history before any test is run. **Movement** is the raw change with both bases. **Significance** is the test, with its interval. **Seasonality** asks whether the calendar explains it. **Trend shape** asks what the line is doing, which is the question the tracker was built to answer. **Source check** asks whether the sample changed rather than the world. **Claim** is what may be written, and it is capped by the weakest gate.

Against the **K2** evidence chain, this skill produces Findings about change and stops there. Why a metric moved is Interpretation, and a causal account of the movement requires a design this one does not have. The commonest failure is the sentence that runs the movement into its explanation with no signal word between them (**K2 §3.2**): "satisfaction fell 4 points after the new system launched" asserts a connection that the tracker cannot establish.

## 9. Output format

**1. Comparability record**, per metric, per wave pair.

| Metric | Wording | Scale | Routing | Base definition | Sample source | Weighting | Fieldwork conditions | Disposition |
|---|---|---|---|---|---|---|---|---|

Each cell: holds, changed (with description), or unknown. Disposition: compared, compared with caveat, bridged, or not compared.

**2. Wave comparison table.**

| Metric | This wave % (n, n_eff) | Last wave % (n, n_eff) | Change | 95% CI on change | Test and p | Normal range for this metric | Year-on-year change | Verdict |
|---|---|---|---|---|---|---|---|---|

Verdict is one of: moved beyond ordinary variation; within ordinary variation; not comparable; not detectable at this base.

**3. Trend series**, per headline metric: all available waves, with the normal-range band shown, break points marked, and the reading convention (discrete or rolling, with window) stated on the chart.

**4. Sample composition trend.** Unweighted profile by wave on the key demographics plus device, mode, completion time and weighting efficiency, so that drift is visible alongside the metrics.

**5. Event log** for the field period, dated, with any within-wave split test that was run.

**6. Detectability statement.** For every metric reported as unchanged, the minimum movement this wave could have detected at its base.

**7. Break and bridge documentation**, where a change occurred: what changed, at which wave, whether a parallel run was done, the measured step and its uncertainty, and how the series is presented as a result.

**8. Conventions block**, carried forward every wave: base definitions, don't know treatment, scale conventions, weighting scheme, testing threshold and regime, rolling window if used, and the list of headline metrics pre-specified for testing.

**Where the evidence is thin**, the format is not filled anyway. A metric whose comparability failed appears with its disposition and no change figure at all, not with a change figure in smaller type. A wave where nothing moved produces a short report saying so with the detectability figures, not a long one built from the largest random movements. Per **K4 §1**, a column headed "change" is not evidence that anything changed.

## 10. Quality checks

Run before any tracker output is presented. **K4 §8** runs anyway; these are specific to this task.

1. Has the seven-point comparability check been completed and recorded for every metric being compared?
2. Is any precondition recorded as unknown being treated as though it held?
3. Does every reported movement carry both waves' bases, unweighted and effective?
4. Is every movement tested, with the interval on the change reported rather than a flag?
5. Is the design correct for the test: independent samples for fresh samples, paired for re-interviewed respondents?
6. Is every movement placed against the metric's own normal range of fluctuation, not only against a significance threshold?
7. Has seasonality been checked against the same calendar position in previous years, and is the year-on-year comparison reported where seasonality is plausible?
8. Is the trend line shown and read before the latest two points are discussed?
9. Has sample composition been checked for drift this wave, including weighting efficiency?
10. Is every external event a candidate explanation only, with no causal construction, and was a within-wave split run where fieldwork spanned the event?
11. Is the number of metrics tested stated, with the testing regime and expected chance positives?
12. Does every "no change" statement carry the minimum movement the wave could have detected?
13. Are discrete and rolling figures kept apart, with no test run between overlapping windows?
14. Is any break in the series shown as a break rather than smoothed over, and is any bridge adjustment documented with its uncertainty?
15. Has any movement within the normal range been given a business explanation anywhere in the output?

## 11. Common failure modes

| Failure | How to recognise it | How to prevent it |
|---|---|---|
| **Narrating noise** | Every small movement in the table has a sentence and a cause; next wave the causes reverse | Normal range established at Step 3; only movements clearing it are narrated; "no change" permitted as a headline |
| **Silent comparability break** | A movement larger than anything in the series history, coinciding with a wave in which something about the study changed | Seven-point check at Step 1, run against documentation rather than memory, recorded every wave |
| **Seasonal comparison read as change** | Q4 always highest, reported as improvement every Q4 | Same calendar position compared; year-on-year reported alongside wave-on-wave |
| **Sample source drift as trend** | A long, smooth movement on a metric that should be stable, with a changing unweighted profile underneath | Sample composition trend maintained; a stable-by-nature question tracked as a control |
| **Paired data tested as independent** | A re-interviewed panel compared wave-on-wave with a two-sample test | Design established before test selection (**05.02**) |
| **Chance flags become the story** | Two significant movements out of forty metrics, both unexpected, both explained | Headline metrics pre-specified; screening results labelled; test count stated |
| **Rolling and discrete mixed** | A trend line that is smooth for part of its length and noisy for the rest | One convention per chart, stated; no test between overlapping windows |
| **Bridge invented without a parallel run** | An adjustment applied to historical waves with no measurement behind it | Parallel run or no bridge; otherwise the series restarts and both are kept |
| **Null read as stability** | "No change in satisfaction" on bases of 150 where 15 points would have been needed to detect one | Detectability statement beside every no-change claim |
| **Event attribution after the fact** | An explanation located once the movement was seen, not recorded during fieldwork | Event log kept per wave in advance; within-wave split test where possible; candidate language only |
| **Index redefined mid-series** | A composite whose components or weights changed, presented as continuous | Treated as a new series or reconstructed historically, with the change documented |
| **AI: fluent narrative over a flat series** | A well-written account of movements that are all inside the noise band | Movements checked against normal range before any prose is written |
| **AI: causal connective smuggled in** | "Satisfaction fell after the launch", "following the campaign, awareness rose" | **K4 §3.2**; temporal coincidence stated as coincidence; the licensing design named if a causal claim is made at all |
| **AI: comparing figures it was not told were incomparable** | Two waves compared where the questionnaire changed, because both numbers were in the supplied table | Comparability record required before any comparison; unknown treated as unknown, not as holding |

## 12. AI guardrails

Universal prohibitions are inherited from **K4** and are not repeated here. **K4 §3.1** and **K4 §3.2** govern this skill in full.

1. **Never compare two waves without completing the comparability check**, and never treat an unverifiable precondition as satisfied. Where documentation is absent, say comparability could not be verified and label the comparison provisional.
2. **Never report a movement without both waves' bases and a confidence interval on the change.** A change figure alone is not reportable.
3. **Never describe a movement inside the metric's normal range of fluctuation as a change**, and never give it a business explanation.
4. **Never attribute a movement to an event because the timing coincides.** State the coincidence, name the test that would separate them, and run it where the data allows.
5. **Never state "no change" without the minimum movement the wave could have detected.** A null on a small base is not stability.
6. **Never construct a bridge between question versions without a parallel measurement.** An adjustment nobody measured is a fabricated number under **K4 §2.1**.
7. **Never present a series across a known break without marking the break**, whatever it does to the shape of the chart.
8. **Never mix discrete and rolling readings on one chart or in one comparison**, and never test two overlapping rolling windows against each other.
9. **Never let the requirement for a story determine what is reported.** Where nothing moved, report that nothing moved, per **K4 §9** offering the strongest honest alternative rather than manufacturing movement.
10. **Never report only the metrics that moved.** The count of metrics tested and the list of those that did not move are part of the result, and omitting them is cherry-picking under **K4 §4.2**.
11. **Never carry a previous wave's figure forward without re-checking it against its source.** Figures degrade as they are copied between decks (**K2 §7**).

## 13. Best-practice principles

1. **Comparability is the whole job, and it is decided outside the data.** No amount of analysis detects a rewording. The tracker's documentation is the instrument, and maintaining it wave by wave is what makes the series worth having.
2. **Know the noise floor before you read the number.** A metric's own history tells you how much it moves when nothing happens. Without that, a significance test on the latest pair is being asked to do a job it cannot do, because it says nothing about how often the same test has flagged before.
3. **Consistency beats correctness in a running series.** A better question, a better scale or a better weighting scheme introduced mid-series costs more than it gains. Improve at a planned break, with a parallel run, or not at all.
4. **The trend line is the finding; the latest wave is a data point.** Read the line first. Most requests to explain the latest movement dissolve once the line is on the screen.
5. **Consecutive small movements in one direction are stronger evidence than one large movement.** Four steps of 2 points the same way is a pattern; one step of 6 points is a data point. Noise does not usually walk in a straight line.
6. **Watch the sample, not only the metrics.** The most convincing false trends in tracking come from the sample changing rather than the world changing, and they are only visible if the sample profile is tracked as carefully as the results.
7. **"No change" is a legitimate and often correct headline.** A stable measure is information: it says the market did not shift, the campaign has not landed, or the decline has stopped. Reports that cannot say it end up saying something false instead.
8. **Every explanation offered for a movement is a hypothesis until tested.** A tracker establishes that a number differs at two times. It does not establish why, and the confidence with which explanations are offered in tracker reporting is out of all proportion to the evidence behind them.
9. **A tracker measures the world and the instrument at the same time, and cannot separate them unaided.** Building in a control question that should not move, and monitoring it, is one of the cheapest and most useful additions to any continuous study.
10. **Design the break before you need it.** Every long-running tracker will eventually need to change. Deciding in advance how a change will be bridged, and budgeting for the parallel wave, is far cheaper than losing a series.
11. **Report the level as well as the movement.** When nothing has changed, where the measure sits, how it compares with a benchmark and how it varies inside the sample are all still available and are frequently more useful than the change would have been.
12. **The pressure to find movement is structural, and it does not go away.** Name it, plan for it, and put the normal range on the chart so that everyone in the room can see what an ordinary wave looks like.

## 14. Worked example

**INPUT**

A fictional national licensing agency, the Vehicle and Driver Registry, runs a quarterly service tracker with 800 respondents per wave, weighted to the population of recent service users. It has run for 17 waves. The headline metric is the proportion rating their most recent interaction as good or very good. Wave 17 comes in at 68%, against 73% in wave 16. The agency's communications team has asked for an explanation of the 5-point fall, and has a candidate: a new online portal launched six weeks before wave 17 fieldwork.

**PROCESS**

*Step 1, comparability.* Six of the seven preconditions hold. The seventh does not: the sample source changed for wave 17, because the previous fieldwork supplier's contract ended and recruitment moved to a different route with a different device mix. Nothing about the questionnaire changed. This is recorded as a failure on sample source, with the disposition undecided pending Step 8.

*Step 3, normal range.* Wave-on-wave changes across waves 1 to 16 range from minus 4 to plus 5 points, with a standard deviation of about 2.6 points. A 5-point fall is at the edge of the ordinary range but not outside it: three previous waves moved by 4 points or more, and none was reported as news at the time.

*Step 4, testing.* Bases 800 nominal, effective 631 this wave and 664 last wave. The 5-point fall gives an interval on the change of minus 9.9 to minus 0.1 points, p=0.048. It clears the threshold by a very small margin, and the interval includes movements that would be trivial.

*Step 6, seasonality.* Wave 17 is a Q1 field period. Looking at the same calendar position in previous years, Q1 has been the lowest quarter in three of the four years available, by an average of about 3 points, which is consistent with the annual renewal peak putting more pressure on the service in that quarter. The year-on-year comparison is 68% against 70% last Q1, a 2-point fall with an interval spanning zero.

*Step 7, trend shape.* The series over 17 waves oscillates between 67% and 74% with no direction. Wave 17 sits inside that band, at its lower edge. The line does not show a step.

*Step 8, sample drift, and the judgement call.* The unweighted profile shows the supplier change moved the device mix substantially, with mobile completion rising from 44% to 61%, and completion times falling. A control question that should not move over three months, self-reported vehicle ownership category, moved by 4 points. Weighting efficiency also fell, with the effective base dropping from 664 to 631 on the same nominal base. **Judgement call:** the comparability failure on sample source is now not merely formal; there is direct evidence that the incoming sample differs on characteristics the weighting does not target. The 5-point fall cannot be separated from the supplier change. Resolution: the wave-on-wave comparison is reported as not valid, with the reason and the evidence shown; the level for wave 17 is reported as a new baseline; the year-on-year comparison is reported with the same caveat since it also crosses the supplier change; and a recommendation is made to run the previous supplier's route in parallel for one wave if the budget allows, which is the only way to size the step.

*Step 9, the portal.* Fieldwork spans four weeks and the portal launched before it began, so a within-wave split by interview date cannot separate the two. Respondents who used the online channel can be compared with those who used other channels, and they rate their interaction 3 points lower, on bases of 291 and 509, untested and confounded with who chooses the online channel. Recorded as a candidate hypothesis, explicitly not attributed, with a note that a channel-level question added to wave 18 would test it directly.

*Step 12, writing.* The communications team's request cannot be answered as asked. The report says so plainly, and offers what the wave can support.

**OUTPUT**

A comparability record showing six preconditions holding and sample source broken; a wave comparison table in which the headline metric carries the disposition "not compared, sample source changed" rather than a change figure; a 17-wave trend chart with the normal-range band shown and a break marked at wave 17; a sample composition chart showing the device shift, the control question movement and the fall in weighting efficiency; an event log carrying the portal launch as an untested candidate; and a detectability statement.

The headline reads: "Wave 17 records 68% rating their interaction good or very good (n=800, effective base 631). This wave is not comparable with wave 16, because the sample source changed and the incoming sample differs on characteristics the weighting does not correct: mobile completion rose 17 points and a control measure that should be stable moved 4 points. The 5-point difference cannot be separated from that change. Wave 17 is reported as a new baseline. Across waves 1 to 16 the measure oscillated between 67% and 74% with no direction, and the current reading is inside that band."

**Researcher decision required.** Whether to fund a parallel run of the previous recruitment route in wave 18 is a decision about the value of the series against its cost (**K5 §2.7**). Without it the tracker restarts at wave 17 and four years of history become a separate series, which is a material loss the analysis can describe but not price.

## 15. Advanced usage

**Continuous fielding and rolling windows.** Where a study fields continuously rather than in discrete waves, the choice of window is an analytical decision with consequences: a wider window is less noisy and slower to detect a real shift, and consecutive points share respondents so they are not independent. Report the window on every chart, and where a shift needs detecting quickly, run both a wide and a narrow window and say what each is for. Testing between overlapping windows is invalid.

**Formal control limits.** Where a metric has a long, stable history, control-chart logic gives a more disciplined version of the normal range: limits set from the historical variation, with signal rules for a point outside the limits and for runs of consecutive points on one side of the centre. This converts "does this movement matter" from a judgement into a pre-agreed rule, which is particularly valuable where the tracker reports to a stakeholder group every wave.

**Decomposing a movement.** Where a headline metric moves, the useful question is usually which part of the sample moved. A movement concentrated in one segment is a different finding from one distributed evenly, and a headline can be flat while two segments move in opposite directions. Decompose by the standard banner (**05.03**) and check the composition of the sample as well, because a headline can move purely because the mix of respondents changed.

**Cohort reading in panel trackers.** Where the same respondents are re-interviewed, individual-level change becomes available and is far more informative than the aggregate: the proportion who improved, worsened and stayed the same, and who moved in which direction. It also brings panel conditioning, where being repeatedly interviewed changes how people answer, and attrition, where those who leave the panel differ from those who stay. Both must be assessed rather than assumed away.

**Benchmarks and category movement.** A metric falling while everything comparable falls further is a relative improvement, and a tracker read only against itself will miss it. Where a category or population benchmark covering the same period exists, report the metric against it, since the movement in the benchmark absorbs a large part of the general conditions that would otherwise be attributed to the subject.

**When the standard approach does not fit.** A metric that is genuinely new, or a series with too few waves to establish a normal range, should be reported as a level with its interval and an explicit statement that trend reading is not yet available. A series broken beyond bridging should be presented as two series with the break shown, never as one line with a discontinuity smoothed away.

## 16. Skill chain

**Recommended previous skills**
- **05.01 Descriptive Analysis.** Hands over the wave's figures on base definitions matching previous waves, and the conventions that must not change between waves.
- **04.05 Weighting and Base Management.** Hands over the weighting scheme, effective bases and design effects per wave, and is where a change in weighting efficiency is detected.
- **03.04 Fieldwork Monitoring and Response Quality.** Hands over the fieldwork conditions and response-quality picture for the wave, which is one of the seven comparability preconditions.
- **01.07 Analysis Plan Development.** Hands over the pre-specified headline metrics that are tested at the nominal threshold, separating them from the screening set.

**Recommended next skills**
- **05.02 Statistical Testing.** Supplies the test, the interval and the minimum detectable movement for every wave-on-wave comparison, and the multiplicity control across a wide metric set.
- **05.03 Cross-Tabulation.** Decomposes a movement by subgroup, and supplies the banner whose definitions must be held constant across waves.
- **08.01 Finding to Insight Development.** Takes a movement that survived the gates and does the interpretation work, which this skill deliberately leaves undone.
- **12.05 Executive Research Reporting.** Takes a wave result, including a stable one, and presents it for a stakeholder audience without reintroducing the pressure to narrate noise.

**Runs well alongside**
- **05.06 Correlation, Regression and Causal Claim Control**, whenever a movement is about to be attributed to an intervention.
- **06.04 Experiment and A/B Test Analysis**, which supplies the design that could answer the attribution question a tracker cannot.
- **13.03 AI Output Verification**, which audits a wave report for narrated noise, missing comparability records and causal connectives.

---
A Yazi Supplied Skill and resource.
