---
name: fieldwork-monitoring-and-response-quality
description: >
  Diagnoses and fixes a study while it is still in field: soft launch checks, drop-off
  and completion analysis, length in practice, quota fill rates and composition drift,
  speeders, straight-lining, open-end quality, within-respondent contradictions, device
  and channel effects, and fraud signals. Use when someone says "check the soft launch",
  "the survey is in field and something looks wrong", "should we pause fieldwork",
  "how do I set a speeder threshold", "respondents are dropping out at Q12", "the
  quotas are filling unevenly", or "these open ends look machine-generated".
category: 03 Fieldwork and Data Collection
ref: 03.04
tier: 1
inherits: [K2, K3, K4, K5]
---

# Fieldwork Monitoring and Response Quality

## 1. One-line description
Reads the diagnostic signals a live study emits, works out whether each one is an instrument defect, a sourcing problem or respondent behaviour, and makes the pause, amend or continue decision while it can still change the outcome.

## 2. What this skill is used for

**The research problem it solves.** Most data quality problems are discovered after fieldwork closes, at which point every option is bad. A question that half the sample misread can be excluded, not fixed. A quota that filled with the wrong composition can be weighted, badly. A source that delivered inattentive respondents can be cleaned out, at the cost of base size and with no way to replace them. Fieldwork is the only window in which a study can still be repaired, and it is routinely treated as a waiting period between launch and data delivery. The purpose of in-field monitoring is not reassurance and it is not a report. It is a set of decisions, made under time pressure, on evidence that arrives incomplete, with a clock running on how much of the sample is still unspent.

**Where it sits in the research lifecycle.** It starts at soft launch and runs until fieldwork closes, alongside everything else. It is the enforcement point for decisions made in sampling, sourcing, instrument design and incentive design, and the last place any of them can be changed. Its handover is to data preparation: every flag raised in field becomes a variable on the dataset, and every intervention becomes an entry in a log that data preparation and analysis both need.

**Typical use cases.**
- Reading a soft launch and making the go, hold or fix decision before the rest of the sample is released.
- Diagnosing a drop-off spike at a specific question and deciding whether to amend.
- Monitoring quota fill rates and correcting the composition drift that comes from cells filling at different speeds.
- Setting and defending a speeder threshold that was not simply invented.
- Detecting straight-lining, duplicate identity, and open-end text that no respondent wrote.
- Deciding whether to pause fieldwork, and documenting what changed if it resumes differently.
- Monitoring a qualitative or AI-moderated study for depth, abandonment and escalation signals.

**Who uses it.** Research executives and managers running live fieldwork; research directors making the pause or continue call; client-side insight managers who want to know what is being checked on their study; data and operations specialists implementing the checks.

## 3. When to use it

- A study has entered soft launch and someone has to decide whether to release the rest.
- Fieldwork is live and any part of it is behaving unexpectedly: pace, composition, length, completion.
- The instrument contains new questions, complex routing, or stimulus that has not been fielded before.
- The sample comes from a source with limited history, or from more than one source.
- The audience is high value or low incidence, so every wasted complete matters.
- Incentives are high enough relative to the population that misqualification and fraud are plausible.
- The study is a wave of a tracker and any deviation from the previous wave's field profile needs explaining before it is read as a change in the market.
- Fieldwork has finished but the question is whether the collection process was sound, in which case this skill supplies the diagnostics and the boundary rules below still apply.

## 4. When NOT to use it

- **Fieldwork has closed and the task is cleaning the dataset.** Removing cases, recoding, imputing and validating a closed dataset belongs to **04.01 Data Validation**. The boundary is exact: **this skill acts while data is still being collected and can change what is collected; 04.01 acts on the collected data and cannot.** This skill hands over flags, thresholds and an intervention log; it does not delete cases from a finished file.
- **The task is preparing or analysing qualitative material.** Transcript structuring, per-participant reading and cross-participant coding belong to **07.03 Interview and Transcript Analysis** and **07.01 Thematic Analysis**. In-field monitoring of qualitative work is about depth, abandonment and escalation while sessions are still being run.
- **The instrument is fundamentally wrong.** Monitoring can catch a defective question and it cannot rescue a study whose objectives were never measurable. If soft launch shows the instrument is not answering the objectives, that is a design failure and the correct response is to stop and return to **02.01 Survey Questionnaire Design** or **02.05 Survey Logic and Flow Review**, not to tune thresholds.
- **The sample source is the problem and the study is nearly complete.** Where monitoring shows a source is delivering unusable respondents late in fieldwork, adding a different source to finish the sample creates two non-comparable halves. That is a **03.01 Recruitment and Sample Sourcing** decision with analysis consequences, and it needs a researcher, not an operational fix.
- **The intervention would be undocumented.** Any mid-field change to instrument, quota, source or incentive splits the sample into a before group and an after group that are not exchangeable. If the change cannot be logged and carried into analysis as a variable, it should not be made. An unlogged amendment is worse than the defect it corrects, because the defect is at least visible in the data.
- **The purpose is to remove respondents whose answers are inconvenient.** Quality thresholds set after looking at the results are not quality thresholds. If the criteria were not defined before the data was seen, they cannot be applied to it now without disclosure of exactly that fact.
- **The study is small enough that every case is inspected individually.** For a 20-interview qualitative study or a 40-respondent specialist survey, statistical quality diagnostics have no basis. Read the material. Thresholds derived from distributions require distributions.
- **No one is available to act on what monitoring finds.** Monitoring that produces no decision is an expense. If there is no route to pausing, amending or escalating within the fieldwork window, say so at launch rather than generating diagnostics nobody can use.

## 5. Required inputs

**Required. Without these the skill cannot run.**
- **Live field data with per-respondent metadata**: status (complete, partial, screened out, over quota), timestamps, question-level timing where available, entry point and device. Without timing and status data, most of the diagnostics below are unavailable and that limitation must be stated.
- **The instrument, with routing and expected paths.** Drop-off cannot be diagnosed without knowing what respondents were asked and in what order.
- **The quota structure and targets**, so fill rates can be read against plan rather than against each other.
- **The design assumptions to be tested**: expected length, expected incidence, expected completion rate, and the source or sources in use. Monitoring compares reality to a stated expectation; without the expectation there is nothing to compare to.
- **A decision route.** A named person who can authorise a pause or an amendment, and their availability. Without it, monitoring is observation.

**Optional, and what each one adds.**
- **Previous wave or comparable study field metrics.** Convert thresholds from judgements into calibrations, which is the difference between a defensible speeder cut and an arbitrary one.
- **Pilot or cognitive test findings.** Tell you which questions were already suspected of being difficult, so a drop-off there confirms rather than surprises.
- **Source variable on every respondent.** Lets every quality signal be read by source, which is the fastest way to distinguish a sample problem from an instrument problem.
- **Open-end text as it arrives**, rather than in a final export. Text quality degrades visibly and early, and it is the most sensitive single indicator of respondent effort.
- **Fraud and duplication signals available from the collection process**, whatever form they take. Raise the ability to detect the same person completing twice.
- **The analysis plan.** Tells you which measures matter, so that the effect of a quality decision can be evaluated against the numbers the study exists to produce.

## 6. Questions to ask before starting

1. **What was expected, on length, incidence, completion rate and composition?** Every diagnostic in this skill is a comparison to an expectation. Default if unanswered: derive expectations from the design (estimated length from the question inventory, incidence from the sourcing plan) and state that they are planning estimates, not measurements.
2. **What are the quality rules, and were they written before the data was seen?** Thresholds set in advance are quality control; thresholds set afterwards are selection. Default: define the rules now, before opening the data, and record the timestamp.
3. **How much sample is still unspent, and how fast is it going?** Determines whether an intervention is worth making. A defect found with 90% of the sample collected is a disclosure, not a fix. Default: check field pace before diagnosing anything.
4. **Who can authorise a pause, and how fast?** Determines what the monitoring plan can promise. Default: assume a same-day decision route and design checks that can be acted on within it.
5. **Is anything about this study comparable to a previous wave?** If so, a mid-field change costs comparability twice: within this wave and against the last one. Default: assume trend comparability is required and weight the amendment decision accordingly.
6. **What is the cost of a false positive on each quality rule?** Removing genuine respondents biases the sample toward whoever answers slowly and writes at length, which is not neutral. Default: prefer flagging to excluding, and let analysis decide with the flag visible.
7. **Where would fraud or misqualification pay?** High incentive, low incidence, or specialist audiences with easily guessed screening criteria. Default: assume any study with a valuable incentive and an obvious screener will attract some misqualification and design the checks accordingly.

## 7. Step-by-step methodology

**Step 1. Write the monitoring plan before launch, including the thresholds.** The plan names, for each check, what will be measured, at what point, against what expectation, what result triggers an action, and what the action is. Thresholds defined in advance can be applied without argument; thresholds defined after seeing results are indistinguishable from choosing the answer, and a reviewer is entitled to treat them that way. The plan also names who decides and how quickly. A correct result is a one-page document, timestamped before the first respondent, that a colleague could execute in your absence.

**Step 2. Read the soft launch as a set of specific questions, not as a look at the data.** Release a small share of the sample (a common convention is around 10%, or enough completes to see a distribution on the key measures, whichever is larger) and hold the rest. Then inspect, in this order, because the earlier items invalidate the later ones. **Did routing work**: is any question answered by people who should not have seen it, and is any question answered by nobody. **Is incidence as expected**: a screen-out rate far above plan means the feasibility calculation was wrong and the cost and timeline change now, not later. **Is length as expected**: median completion time against the design estimate. **Where do people leave**: drop-off by question. **Do the open ends make sense**: read every one, at this stage, not a sample. **Do the key measures have usable distributions**: a variable with 95% in one category will not support the analysis it was written for, and this is the last moment to add a question. **Do the answers cohere**: pick five respondents and read their whole record end to end, which catches things no aggregate does. The output of soft launch is one of three decisions: release, release with a named fix, or hold. A soft launch that produces no observations has not been read properly.

**Step 3. Analyse drop-off by question and diagnose the pattern, because each pattern means something different.** Plot completion as a curve across the instrument and look at both the level and the shape.

- **A spike at one question** means that question. Look for a required answer respondents cannot give, an option list missing their case, a question that reads as intrusive, a grid too wide for a small screen, or a technical failure. Read the last-seen data of the people who left.
- **A spike immediately after a question** usually means the question was answerable but unwelcome: the respondent answered, then reconsidered the study.
- **A steady gradient with no spike** is length and fatigue. It cannot be fixed by editing a question and it can be fixed by removing a block.
- **A step at a section boundary** points to a transition problem: a new topic, a new format, or a progress indicator revealing how much remains.
- **Early loss before the first substantive question** points to the invitation, the consent text or a mismatch between what was promised and what arrived.
- **Loss concentrated in one subgroup or one device** is a coverage problem masquerading as a quality problem: check drop-off by device, by source and by demographic before concluding the question is at fault.

A correct diagnosis names the mechanism, not just the location, and states what evidence separates it from the alternatives.

**Step 4. Compare length in practice to length in design, and act on the gap.** Compute the median, not the mean, since a small number of respondents who leave a session open for hours will destroy the mean. Then look at the distribution rather than the central value: a long tail is normal, a bimodal distribution is not and usually indicates two different behaviours (for example, engaged completion and rapid clicking). Compare against the design estimate and, where available, against the previous wave. If actual length materially exceeds the estimate, the consequences are already accruing (drop-off, satisficing, incentive inadequacy) and the options are to cut a block, accept a lower completion rate, or raise the incentive, each with its own cost. Where length is materially shorter than estimated, that is a signal too: it often means routing is sending respondents past sections they should have seen.

**Step 5. Track quota fill rates and, more importantly, the composition drift they cause.** Quota cells fill at different speeds, always, because the easy cells contain the people most available and most willing. Two consequences follow and only the first is widely noticed. The obvious one is that the hard cells finish last and may not finish at all. The subtle and more damaging one is **time confounding**: if young respondents fill in two days and older respondents take three weeks, then the young sample was collected in one period and the older sample in another, and any event during fieldwork affects the groups unequally. Monitor fill rate as a percentage of target per cell per day, project each cell's completion date against the field close, and act early, because interventions late in field are more disruptive. The correct interventions, in order of preference: throttle the fast cells so all cells run through the whole window (which preserves temporal comparability and is the option most often forgotten), boost recruitment to slow cells within the same source, then relax a cell target with a documented decision, and only then add a source, which is a sourcing change with its own consequences (**03.01**).

**Step 6. Set a speeder threshold you can defend, and understand what it costs.** A speeder is a respondent who could not have read the questions. Three approaches exist, and the honest position is that they are conventions rather than validated rules, so the requirement is to state which was used and why. **Absolute floor**: a fixed minimum time derived from the instrument itself, ideally from a reading-speed calculation over the actual word count, which is defensible because it is derived rather than borrowed. **Relative threshold**: a fraction of the median completion time (one third and one half are both in common use), which is simple and adapts to the instrument, and which becomes circular if the sample is already full of speeders, since the median is contaminated. **Section-level timing**: identifying respondents who are fast on the sections that require reading (grids, stimulus, open ends) rather than fast overall, which is the most precise approach and needs question-level timing to be captured. Whatever is chosen: define it before looking at results, apply it consistently, prefer flagging to deleting in field, and check what removal does to the composition of the sample. Speeding correlates with device, with age, with practised respondents and with source, so a speeder cut is never a random subtraction, and its effect on the achieved composition should be reported alongside the count removed.

**Step 7. Detect pattern responding, and distinguish it from genuine consistency.** Straight-lining is identical or near-identical answers down a battery. It is a strong signal of inattention and a weak one on its own, because some respondents legitimately hold the same view of every item, particularly on short batteries and on batteries where the items are genuinely similar. Strengthen it three ways. **Use reverse-coded items**: a respondent who agrees with an item and its reverse is not answering the content. **Look at the pattern across batteries**: straight-lining once is plausible, straight-lining every battery in the instrument is not. **Combine with timing**: a straight-lined grid completed faster than it could be read is a different finding from a straight-lined grid completed slowly. Also look for other pattern signatures: alternating answers, diagonal patterns, always selecting the first option, always selecting the last option, and identical response strings across supposedly independent respondents, which is a duplication signal rather than an inattention one.

**Step 8. Read the open ends, because they are the most sensitive indicator of effort you have.** Classify every open end into: substantive, minimal but real, non-answer (punctuation, single characters, "n/a"), off-topic, duplicated (identical or near-identical text appearing in more than one respondent's record, or repeated by the same respondent across different questions), copied from the question stem, and machine-generated. Machine-generated text has recognisable properties: it is fluent, well-structured, longer than typical respondent text, generic about the specific product or experience, and it often answers the question more completely than a person would while containing no concrete detail. Two cautions matter here and both are frequently ignored. First, fluency is not proof: articulate respondents exist, and excluding them because they wrote well biases the sample toward the inarticulate. Second, no detection method is reliable enough to justify silent removal, so text suspected of being machine-generated is flagged for human review (**K5 §2.6**), not deleted by rule. What can be applied by rule is duplication, non-answers, and text copied from the stem.

**Step 9. Run within-respondent consistency checks.** Contradiction is often a stronger signal than speed, because it does not depend on timing data and it is harder to fake. Build the checks into the instrument at design stage where possible: a factual item asked twice in different forms and separated by distance, a category usage claim that must be consistent with a later frequency answer, an age or tenure that must be consistent with a date. In field, look also for logical impossibilities the routing did not catch (a respondent who has never used the category rating their satisfaction with it), for claimed usage of an item that does not exist where a fictional decoy has been included, and for demographic answers inconsistent with the screener. Report the count of respondents failing each check separately, because a check that flags a quarter of the sample is a defective check, not a finding about respondents.

**Step 10. Read every diagnostic by device, by source and by entry point before attributing it to respondents.** This is the step that most often changes the conclusion. Small-screen respondents take longer on grids and shorter on open ends, and both look like quality problems and are mode effects. One source producing all the speeders is a sourcing finding, not a respondent finding. A spike in completions from an unusual entry point at an unusual hour is an access problem. The discipline is simple and rarely applied: **before removing respondents, cross every quality flag by device, source and entry point.** If a flag concentrates in one of them, the fix is upstream and removing individuals treats a symptom while biasing the sample.

**Step 11. Look for duplication and fraud signals, proportionate to the incentive and the audience.** The signals available vary with the collection method, and the point is not to enumerate techniques but to look at the right level. **Identity duplication**: the same person completing more than once, visible as repeated identifying details, identical open-end text, or improbably similar full response strings. **Coordinated completion**: an implausible burst of completes in a short window, unusual uniformity in demographic answers, or a cluster arriving through one entry route. **Misqualification**: qualification rates far above the expected incidence, which is the single most informative fraud signal there is, since a screener that suddenly passes 40% of a population known to be 8% is not finding more of them. **Inconsistent identity**: screener answers that conflict with profile data or with later answers. Escalate rather than resolve unilaterally: fraud findings have commercial and contractual consequences and the decision to reject a batch belongs to a person (**K5 §2.5**).

**Step 12. Make the pause, amend or continue decision explicitly, and log everything.** Four options, in increasing cost. **Continue**: the signal is within tolerance, or is a real property of the population rather than a defect. Record the observation and the decision to continue. **Continue with flagging**: capture the flag as a variable, hand it to analysis, decide later. This is the default for most respondent-level quality signals and it is almost always better than in-field deletion, because it preserves the option. **Amend**: change the instrument, quota, source or incentive. This is the expensive one, because **an amendment partitions the sample into a pre-change group and a post-change group that are not exchangeable**, and every affected variable must carry that partition into analysis. Amend only where the defect is material, where enough sample remains for the change to matter, and where the partition can be handled analytically. **Pause**: stop collection while the decision is made. Pausing costs field time and preserves sample, and it is the right call when continuing would spend sample on data that will not be usable.

Every one of these produces a log entry, and the log has a fixed form: timestamp, what was observed, on what evidence, what was decided, by whom, what changed, and which respondents are affected (specifically, the ranges of respondent IDs before and after). This log travels to **04.01** and its summary belongs in the methodology. An amendment that is not logged is indistinguishable from data that behaved strangely for unknown reasons.

## 8. Analytical framework

Every in-field signal runs through the same five-step diagnosis, and the middle step is the one that matters:

    Signal → Candidate causes → Discriminating evidence → Attribution → Intervention decision → Log

There are three families of candidate cause for almost every signal, and the whole discipline is deciding which one is in play:

| Family | What it means | Typical discriminating evidence |
|---|---|---|
| **Instrument** | The question, routing, layout or stimulus is at fault | The signal concentrates at a specific question or section, and appears across all sources and devices |
| **Sample and channel** | The source, device, entry route or invitation is at fault | The signal concentrates in one source, device or entry point, and is spread evenly across the instrument |
| **Respondent** | Individual inattention, misqualification or fraud | The signal concentrates in individuals who show it on multiple independent indicators, across the whole instrument |

The discrimination is almost always achieved by the same move: **cross the signal by question, by source and by device, and see where it concentrates.** A signal that concentrates by question is an instrument problem. One that concentrates by source or device is a sampling or mode problem. One that concentrates in individuals across everything is a respondent problem. Only the third family justifies acting on individuals, and it is the family researchers reach for first.

The second rule of the framework is about cost asymmetry. **A false negative leaves bad data in a flagged variable; a false positive removes a real respondent permanently and non-randomly.** Because in-field removal cannot be undone and because every quality criterion correlates with something substantive (age, device, literacy, source), the default is to flag and let analysis decide with the flag visible, reserving in-field removal for cases where continuing to collect from that respondent or that route is itself the problem.

## 9. Output format

**1. Monitoring plan**, timestamped before launch: checks, timing, expectations, thresholds, triggers, actions, decision owner.

**2. Soft launch report.**

| Check | Expected | Observed | Verdict | Action |
|---|---|---|---|---|

Covering routing, incidence, length, drop-off, open-end quality, key measure distributions, and the five full-record reads. Ends in one of: release, release with named fix, hold.

**3. Field status dashboard**, updated at an agreed cadence: completes against target overall and by cell, projected completion date per cell, median length, completion rate, drop-off curve, and quality flag counts.

**4. Drop-off diagnosis.**

| Question | Drop-off | Pattern type | Diagnosed cause | Evidence | Action |
|---|---|---|---|---|---|

**5. Quality flag summary.**

| Flag | Definition and threshold | Threshold set on (date) | Count | % of completes | Concentration by source / device | Recommended treatment |
|---|---|---|---|---|---|---|

Every flag records whether it is a flag or an exclusion, and no flag lacks a threshold-setting date.

**6. Composition monitor**: achieved versus target composition, and achieved composition by date, so time confounding is visible rather than inferred.

**7. Intervention log.** Timestamp, observation, evidence, decision, decision owner, change made, respondent IDs affected before and after. This is a required output even when the log is empty, in which case it says so.

**8. Handover to data preparation**: the flag variables on the dataset, their definitions, the thresholds used, and the interventions that partition the sample.

**9. Review points** per **K5 §3**, with the pause, amend and exclusion decisions marked.

**When the evidence is thin**, a diagnosis that cannot be made is written as `cause not established` with the candidate causes listed, rather than as the most plausible one stated confidently. Where timing data is unavailable, the speeder section says so rather than substituting a proxy without saying it is one (**K4 §2.5**).

## 10. Quality checks

Run on the monitoring itself, before any intervention. These sit on top of **K4 §8**.

1. Every threshold has a date showing it was set before the data it is applied to was seen.
2. Every quality flag has been crossed by question, source and device before being attributed to respondents.
3. Median, not mean, is used for completion time, and the shape of the distribution has been inspected, not just its centre.
4. Drop-off is diagnosed by mechanism, not just located by question number.
5. Quota fill is monitored as projected completion against field close, not as a snapshot percentage.
6. Composition by collection date has been checked, so time confounding in the quota cells is visible.
7. The effect of every proposed exclusion on achieved composition has been computed and reported, not assumed to be neutral.
8. Straight-lining is corroborated by at least one independent signal (reverse items, timing, or repetition across batteries) before it is treated as inattention.
9. Suspected machine-generated open ends are flagged for human review, never removed by rule alone.
10. Any check flagging an implausibly large share of the sample has been re-examined as a defective check before being believed.
11. Every intervention has a log entry naming the affected respondent ID ranges, so the pre- and post-change partition is reconstructable.
12. The decision to continue is logged as explicitly as the decision to amend, so that silence is not mistaken for the absence of a signal.
13. Where the study is a tracker wave, any field-profile deviation from the previous wave is recorded before results are compared.
14. The handover to data preparation lists every flag variable and its definition, so no threshold has to be reconstructed later from memory.

## 11. Common failure modes

| Failure | How to recognise it | How to prevent it |
|---|---|---|
| **Soft launch as a formality** | A soft launch runs, nothing is inspected, full launch proceeds | Specify the seven soft launch checks in advance and require a written verdict |
| **Thresholds set after seeing results** | The speeder cut appears once the topline is known | Timestamp the monitoring plan; disclose any threshold set later as what it is |
| **Removing the symptom** | Speeders deleted when they all came from one source | Cross every flag by source and device before touching individuals |
| **Single-signal exclusion** | Respondents removed on straight-lining alone | Require corroboration from an independent indicator |
| **Composition damage from cleaning** | Sample gets older, slower and more literate after quality removals | Compute composition before and after every proposed exclusion and report it |
| **Quota speed blindness** | Cells filled, composition fine, and each cell collected in a different fortnight | Monitor composition by collection date; throttle fast cells rather than only boosting slow ones |
| **Mean completion time** | A reported average length inflated by abandoned open sessions | Use the median and inspect the distribution |
| **The unlogged amendment** | A question changes mid-field and the dataset shows an unexplained discontinuity | No amendment without a log entry naming the affected ID ranges |
| **Monitoring without a decision route** | Detailed dashboards, no authority to act, fieldwork completes anyway | Confirm the decision owner and their availability at launch |
| **AI: invented benchmarks** | "Typical completion rates for this type of study are..." | **K4 §2.4**. Compare to the design estimate or a named comparable study, or state there is no benchmark |
| **AI: confident fraud attribution** | Respondents labelled fraudulent on a pattern, without escalation | Fraud findings are escalated to a person; they carry contractual consequences (**K5 §2.5**) |
| **AI: over-detection of machine-generated text** | Articulate respondents flagged because their answers are well written | Fluency is not evidence. Flag for review, require concrete-detail absence plus another signal, and never delete by rule |
| **Flag inflation** | Nine flags, a third of the sample carrying at least one, nobody acting on any | Keep flags few, defined, and tied to a stated treatment; a flag with no treatment is noise |
| **Late diagnosis** | A defect found at 85% of sample, when nothing can be done | Front-load the checks; the value of monitoring decays with every complete collected |

## 12. AI guardrails

Skill-specific only. **K4** applies in full and is not repeated here.

1. **Never set or adjust a quality threshold after seeing the substantive results.** If a threshold has to be set late, disclose that it was, and report the effect of applying it on the key measures.
2. **Never exclude a respondent in field on a single indicator.** Corroboration from an independent signal is required, and flagging is preferred to exclusion in every case where analysis could make the decision later with more information.
3. **Never state a completion rate, drop-off rate, incidence or length as an industry benchmark.** Compare to the design estimate or to a named comparable study, or state that no comparison is available (**K4 §2.4**).
4. **Never label text as machine-generated on fluency alone**, and never delete it by rule. Flag it, state the indicators, and route to human review (**K5 §2.6**).
5. **Never make a mid-field change without producing the log entry**, including the respondent ID ranges before and after. The log entry is part of the change, not documentation of it.
6. **Never attribute a quality signal to respondents before crossing it by question, source and device.** The most common wrong answer in this whole area is "the respondents are bad" when the true answer is "question 14 is broken on small screens".
7. **Never report a cleaned achieved sample without reporting what cleaning did to its composition.** Removals are non-random by construction.
8. **Never declare fraud.** Report the signals, their strength and what they are consistent with, and escalate the determination to a named person.
9. **Never present an in-field observation as a finding about the market.** A distribution seen in soft launch on a partial, non-randomly ordered sample is a process check, not a result, and it must not be shared as an early topline.

## 13. Best-practice principles

1. **Monitoring exists to produce decisions.** If a check cannot lead to an action within the fieldwork window, it is reporting, and it belongs after fieldwork or nowhere.
2. **The value of every diagnostic decays with each completed interview.** Front-load. The first hours of soft launch are worth more than the whole rest of the fieldwork period put together.
3. **Read whole respondent records, not only aggregates.** Five complete records read end to end will find things that no cross-tab of quality flags will, because incoherence is visible in sequence and invisible in summary.
4. **Every quality criterion selects on something substantive.** Speed relates to device, age and practice. Open-end length relates to literacy and language. Straight-lining relates to fatigue, which relates to position in the instrument. There is no neutral cleaning rule, so report the composition consequence every time.
5. **Prefer flagging to deleting.** A flag preserves the decision for analysis, which will know more than field does. Deletion in field is irreversible and made with less information.
6. **A check that flags a quarter of the sample is a bad check.** Suspect the rule before suspecting the respondents.
7. **The most informative fraud signal is a qualification rate far above the expected incidence.** People do not become more common because you started asking about them.
8. **Throttle the fast quota cells.** It is the least-used intervention and the one that best protects the study, because it keeps every subgroup collected across the same period and removes a confound that is otherwise permanent.
9. **An amendment splits the study.** Treat every mid-field change as creating two samples, and decide in advance how analysis will handle them. If the answer is "we will not mention it", do not make the change.
10. **The absence of a signal is worth logging too.** A recorded decision to continue is what makes it possible, later, to show that a check was run and found nothing.
11. **Fieldwork is the last cheap moment.** Everything found here can be fixed. Everything found after close can only be disclosed, excluded or lived with.
12. **Field metrics belong in the methodology.** Completion rate, median length, drop-off, exclusions and interventions tell a reader how much to trust the data, and they cost half a page.

## 14. Worked example

**INPUT**

A fictional retail bank, Calder Trust Bank, is running wave 4 of its quarterly customer experience tracker. Target 1,500 completes, quotas interlocked on age band, region and product holding. Instrument is stable from wave 3 except for a new eight-item battery on digital service, added at question 18. Sample drawn from the bank's customer records, single source. Design estimate 11 minutes. Soft launch of 140 completes.

**PROCESS**

*Step 2, soft launch.* Routing correct. Incidence not applicable (census-style customer frame). Median length 15.5 minutes against an estimate of 11. Drop-off curve flat until question 18, then a 9% loss at that single question and a further gradient afterwards. Open ends fine in the first half and noticeably shorter after question 18. Key measure distributions usable. Five full records read: two showed the new battery answered in an alternating pattern, which is a signal that the grid was being clicked rather than read.

*Step 3, drop-off diagnosis.* The spike is at one question, which points to the instrument rather than the sample. Crossing by device settled it: drop-off at question 18 was 3% on large screens and 17% on small ones, and 71% of respondents were on small screens. The new battery had been designed as an eight-by-five matrix, which renders as a wide grid requiring horizontal scrolling on a phone. The diagnosis is a mode failure in a new question, not respondent inattention, and the alternating-pattern responses in the two records read at Step 2 are consistent with that.

*Step 4, length.* The median of 15.5 minutes against 11 explains the post-18 gradient and the shortened open ends: the instrument is now long enough that the back half is being satisficed. Wave 3's median was 10.8, so the addition alone accounts for most of the increase.

*The judgement call.* Three options. Continue, and accept a wave with a known small-screen defect and a break in comparability caused by satisficing in the back half. Amend the battery to a single-item-per-screen sequence, which fixes the mode failure but splits wave 4 into 140 respondents on the matrix version and 1,360 on the sequential version, with a known format effect on exactly the new measure. Or pull the battery, field wave 4 as wave 3 was fielded, and reinstate the battery properly at wave 5.

Escalated as **RESEARCHER DECISION REQUIRED**, since the decision turns on knowledge the monitoring does not have: whether the digital service measure is needed this quarter or can wait. The bank's insight lead confirmed it was wanted for a board paper in two quarters, not this one. The battery was pulled. The 140 soft launch respondents were retained for the unaffected questions, flagged with a version variable, and excluded from any analysis of the pulled battery. Logged with the respondent ID range.

*Step 5, quotas.* Projected fill showed the 65-plus cell reaching target in 19 days against a 14-day window, while the 25 to 34 cell would fill in 4. Rather than only boosting the older cell, invitations to the fast cells were throttled to spread them across the full window, so that all age groups were collected over the same period. This mattered on this study specifically: a competitor's service outage was reported in the press during fieldwork, and had the younger sample been collected entirely before it and the older sample entirely after, the age difference on trust measures would have been uninterpretable.

*Step 6, speeders.* Threshold set in the pre-launch plan as one third of the wave 3 median (3.6 minutes), a relative rule calibrated to a prior wave rather than to this wave's contaminated median. Applied as a flag, not an exclusion. Twenty-two respondents flagged, concentrated slightly on small screens; composition effect of removing them computed and found to shift the age profile by under a point, and the decision on exclusion handed to **04.01** with the flag attached.

**OUTPUT**

A soft launch verdict of hold-and-fix, a diagnosed mode failure attributed to instrument rather than respondents on device-crossed evidence, one battery pulled with a logged intervention and an ID range, a throttling intervention that preserved temporal comparability across quota cells, a speeder flag calibrated to the previous wave and handed forward unexecuted, and two review points: whether the 140 pre-change respondents should carry a version flag into the trend, and whether the digital battery should be cognitively tested on a phone before wave 5.

## 15. Advanced usage

**Monitoring qualitative and AI-moderated fieldwork.** The signals differ but the framework does not. Watch section-level sufficiency rates rather than completion, abandonment point rather than drop-off question, probe counts actually fired, transcript length by section, and escalation counts. The same discrimination applies: a section failing sufficiency for everyone is an instrument problem in the probe tree, a section failing for one recruitment source is a sample problem, and a participant giving nothing anywhere is a respondent problem. Interventions to a probe tree partition the sample exactly as an instrument amendment does, and must be versioned and logged (**03.02**).

**Multi-market monitoring.** Monitor every market separately and never pool the diagnostics, because a pooled drop-off curve hides a single market where a translation failed. Expect legitimate differences in length, completion rate and open-end length between markets and languages, and calibrate thresholds per market rather than applying one global speeder rule, which will otherwise remove disproportionately from whichever language is most compact.

**Tracker discipline.** Field metrics are part of the trend. Median length, completion rate, drop-off shape and source composition should be tabulated wave on wave, because a change in any of them is a candidate explanation for a change in the results, and it is a much more common explanation than a change in the market. Where a wave differs on a field metric, that goes in the report next to the movement it might explain.

**Continuous and always-on studies.** Where fieldwork does not close, monitoring becomes a control chart problem: establish the normal range for each metric from history, and alert on deviation rather than inspecting every day. The risk is drift, which is invisible day to day and obvious over quarters, so include a periodic full re-read of the diagnostics rather than relying solely on alerts.

**When the standard approach does not fit.** For very small or specialist samples, abandon distributional thresholds and read every case. For studies where timing data is not collected, say so and lean on consistency checks and open-end quality, which do not require it. For high-incentive specialist studies, invert the priority order and put misqualification detection first, since one fraudulent respondent in a base of 60 does more damage than a hundred speeders in a base of 3,000.

## 16. Skill chain

**Recommended previous skills**
- **03.01 Recruitment and Sample Sourcing.** Hands over the source variable, the predicted source effects and the expected incidence, which are the expectations monitoring tests against.
- **02.05 Survey Logic and Flow Review.** Hands over the expected paths, so routing failures at soft launch can be recognised rather than discovered.
- **02.06 Screener and Quota Design.** Supplies the quota structure and targets against which fill rates and composition drift are read.

**Recommended next skills**
- **04.01 Data Validation.** Receives the flag variables, thresholds, exclusion decisions and the intervention log, and makes the final cleaning decisions on the closed dataset.
- **07.03 Interview and Transcript Analysis.** Receives the transcript-level quality records and exclusions from qualitative and AI-moderated fieldwork before per-participant reading begins.
- **12.03 Research Report Compilation.** Publishes the field metrics, the exclusions and any amendment.

**Runs well alongside**
- **03.02 AI-Moderated Interview Design**, whose probe tree and sufficiency rules are monitored in field by this skill.
- **03.03 Stimulus and Concept Preparation**, since comprehension failures and position effects surface first in field diagnostics.
- **03.05 Incentive and Participation Design**, because completion, drop-off and misqualification signals are frequently incentive signals.

---
A Yazi Supplied Skill and resource.
