---
name: screener-and-quota-design
description: >
  Designs the screening questions and quota frame that decide who takes part in a
  study. Use when someone says "write the screener", "design the quotas", "how do
  we find these people", "what should the quota frame be", "interlocking or not",
  "how do we stop professional respondents", "the incidence came back lower than
  expected", or when a sample definition has to become questions a recruiter can ask.
category: 02 Instrument Design
ref: 02.06
tier: 1
inherits: [K2, K3, K4, K5]
---

# Screener and Quota Design

## 1. One-line description
Turns a target population definition into observable screening criteria and a defensible quota frame, ordered to eliminate cheaply, disguised so it cannot be gamed, and built so that the sample it produces supports the analysis the study has promised.

## 2. What this skill is used for

**The research problem it solves.** No amount of analysis fixes a bad screener. Every other failure in a research project has a remedy: a biased question can be caveated, a routing fault can sometimes be rebased, a weak analysis can be redone. A sample of the wrong people is terminal, and it is terminal quietly, because the data looks exactly like data. The study runs, the tables populate, the deck gets written, and the findings describe a population nobody wanted to know about. Three mechanisms produce this. The first is definitional: a target described in the brief as "decision makers", "regular users" or "people who care about sustainability" is not a criterion, because it names an internal state rather than something a person can be asked and verified on, so the recruiter fills the quota with whoever says yes. The second is transparency: a screener that reveals what it wants teaches motivated respondents the answer, and incentivised populations contain people who are good at this. The third is the quota frame itself, which is usually inherited from the last study, is rarely justified against the analysis plan, and silently determines what the study can and cannot compare for the rest of its life. On top of these sits the feasibility problem: incidence estimates are guesses, cost and timing are built on them, and a wrong guess by a factor of three is common and turns a profitable study into a failing one.

**Where it sits in the research lifecycle.** After the sampling strategy is agreed and before recruitment or fielding. It converts a population definition into questions a recruiter can ask and a quota grid a field team can work to, and it is the last point at which "who is in this study" can still be decided rather than discovered.

**Typical use cases.**
- Turning a target population definition into a working screener for a quantitative or qualitative study.
- Building a quota frame from an analysis plan, and deciding what to interlock.
- Designing screening for a low-incidence or specialist population, where every additional criterion multiplies the cost.
- Adding fraud, duplication and professional-respondent controls to an incentivised study.
- Estimating incidence and feasibility before a proposal is priced.
- Diagnosing a study where the recruited sample does not match the intended one, or where screen-out rates are far from expectation.

**Who uses it.** Research executives and managers writing screeners; research directors pricing and scoping; fieldwork and operations teams working to a quota grid; client-side insight managers who need to know exactly who was interviewed; UX and product researchers recruiting their own participants without a field team.

## 3. When to use it

- A sample definition exists and someone has to turn it into questions.
- The target is specialist, low-incidence or professionally defined, and feasibility is genuinely in doubt.
- The study is incentivised and open to self-selection, which is the condition under which misrepresentation pays.
- The analysis plan requires subgroup comparison and the quota frame has to deliver reportable bases for each.
- Competitor, industry, agency or media exclusions apply, or the client's own employees must be kept out.
- A previous study returned a sample that turned out not to be the intended population.
- A tracker is being repeated and the screener must stay identical, or a deliberate change has to be assessed.
- A proposal is being priced and the incidence assumption drives the cost.

## 4. When NOT to use it

- **The sample has not been designed.** Screening criteria implement a sampling decision; they do not make one. How many, from where, with what representativeness claim, and what the minimum reportable base is, belong to **01.06 Sampling Strategy**. Writing a screener first tends to produce a sample defined by whatever was easy to ask.
- **The population is the whole of a general public and no subgroup reporting is required.** A screener that screens nothing adds length, teaches respondents what the study is about, and produces a false impression of rigour. Ask only for what determines eligibility or a quota.
- **The variables in question are for analysis rather than eligibility.** Anything you want to cut the data by, but which does not determine who takes part, belongs in the classification section of the instrument, not the screener. This is the single commonest source of over-long screeners, and it costs money on every screen-out.
- **The routing is the problem.** How screening questions interact with quotas, terminates and the rest of the instrument, and whether the paths work, is **02.05 Survey Logic and Flow Review**. This skill designs the criteria and the frame; that skill tests that they behave as intended.
- **The problem is where the sample comes from.** Panel versus client list versus intercept versus specialist recruitment, the sourcing implications for coverage and the trade-offs between them, belong to **03.01 Recruitment and Sample Sourcing**. A screener cannot correct for a frame that does not contain the population.
- **The intended sample cannot exist at the intended cost.** Where honest incidence arithmetic shows the study is not feasible, the output is that finding, not a looser screener. Loosening criteria to make a quota fillable changes the population without telling anyone, and it is the single most common way a study fails while appearing to succeed.
- **The population is defined by an attitude nobody can verify.** "People who value quality", "the environmentally conscious", "early adopters". These can be measured inside a study and used analytically. As screening criteria they are self-report on a desirable trait under an incentive, which is a recipe for a sample of people who like agreeing.
- **Recruitment is already complete.** The output then is a sample-composition assessment for the report: who was actually recruited, how that differs from the intention, and what it does to each finding. Say that plainly rather than delivering a screener nobody can use.

## 5. Required inputs

**Required. Without these the skill cannot run.**
- **The target population definition and the subgroups that must be reportable**, with the minimum base for each, from **01.06**. Without the subgroup requirement, a quota frame cannot be built, only guessed at. Stop and ask.
- **The research objectives and what the study is for.** Screening criteria are only defensible against a purpose: whether "recent purchaser" means three months or twelve depends on what decision the study feeds, not on convention.
- **Method, mode and incentive structure.** Determine how much screening can be done, how it is administered, and how strong the motivation to misrepresent is. A high-incentive specialist study needs controls that a low-incentive general population study does not.

**Optional, and what each one adds.**
- **Incidence data, from a previous wave, a client database, published statistics or an omnibus question.** Converts feasibility from a guess into an estimate, and it is the input with the largest effect on cost accuracy. Without it, incidence is stated as an assumption and the risk is flagged, not hidden.
- **The analysis plan (01.07).** Determines what the quota frame must guarantee, and which variables must be interlocked because they will be compared jointly.
- **Population benchmarks for the quota variables.** Census, industry or client data. Distinguish a representative quota (matching a known distribution) from a design quota (guaranteeing a reportable base), which are different things frequently confused.
- **Previous wave screener.** For a tracker, the screener is part of the trend and changing it changes the population.
- **Known fraud experience with this audience.** Specialist and high-incentive audiences attract misrepresentation at rates that differ by orders of magnitude; knowing which you have determines how many controls are worth their false-positive cost.
- **Client exclusion lists.** Employees, competitors, agencies, recent participants, embargoed accounts.

## 6. Questions to ask before starting

1. **What observable fact makes someone eligible, and how would you verify it if you had to?** The governing question. If the honest answer is "we would take their word for it", the criterion is self-report on an incentivised claim and needs either a supporting question or a lower confidence attached to it. Default: convert every attitudinal criterion into a behavioural one, and say so.
2. **What is the incidence of each criterion, and where does that number come from?** Determines feasibility, cost and timing. Default: state the assumption explicitly, flag it as unverified per **K4 §2.1**, and recommend a soft launch measurement before the full sample is committed.
3. **Which subgroups must be reportable, and at what base?** This is what the quota frame exists to guarantee. Default: apply the study's stated minimum reportable base, and if none has been agreed, flag any cell under 100 as directional and under 30 as unreportable as a percentage.
4. **What would a respondent gain by lying, and what would they need to know to do it well?** Determines disguise and controls. Default: assume the incentive is meaningful to some part of the sample and design the screener so that the qualifying answer is not guessable.
5. **Who must be excluded, and on what grounds?** Competitors, the client's own staff, agencies and media, anyone who took part recently, anyone in a household with a previous participant. Default: apply an industry exclusion covering the client's sector, market research and marketing, and a recent-participation exclusion, and confirm the window.
6. **Is this screener repeating a previous wave?** If so, wording, order and criteria are part of the trend. Default: assume comparability matters and flag any change as a break.
7. **What happens when a quota cell will not fill?** Deciding this in advance, in writing, prevents it being decided at 5pm on the last day of fieldwork by whoever is holding the phone. Default: recommend a documented replacement and relaxation policy agreed before fieldwork starts.

## 7. Step-by-step methodology

**Step 1. Translate the population definition into observable criteria.** Take each element of the target definition and rewrite it as something a person can be asked, can answer accurately, and would answer the same way tomorrow. "Decision maker" becomes a specific act: who signed off the last purchase in this category, or who would have to approve it, tested by asking what their role was in the most recent decision rather than whether they are a decision maker. "Regular user" becomes a frequency over a defined window, with the window set by the category's natural rhythm rather than by habit. "Interested in the category" becomes a behaviour that only an interested person performs. Three tests for each criterion: is it observable rather than internal; is it answerable accurately by the respondent, which routine low-salience behaviour often is not; and is it stable, so that the same person qualifies next week. A correct result is a criteria table with one row per element of the definition, each showing the original wording from the brief, the observable form, the question that will test it, and the confidence attached to the answer.

**Step 2. Separate eligibility criteria from analysis variables, ruthlessly.** Ask of every proposed screening question: does a wrong answer to this change whether the person takes part, or which quota cell they fall in? If neither, it does not belong in the screener. Everything else goes in the classification section of the main instrument, where it is asked once of qualifying respondents rather than of everyone who is screened. This matters commercially as well as methodologically, because in a low-incidence study most of the people who answer the screener will never enter the study, and every question asked of them is paid for and discarded. It also matters for quality: a long screener with several attitudinal or descriptive questions teaches the respondent what the study wants long before they reach anything that counts.

**Step 3. Order the screening questions by cost and elimination power.** The rule is cheapest and most-eliminating first, and it is arithmetic rather than preference. Put the criterion that removes the largest share of the population at the front, provided it is quick to ask, so that the remaining questions are asked of a much smaller group. Put anything slow, sensitive or expensive last, when it will be asked of few people. Where two criteria are similar in elimination power, ask the cheaper one first. Then check the order for a second property: it must not teach. A screener that opens by naming the category and asking whether the respondent uses it has both signalled the target and asked the one question a motivated respondent will answer strategically. A correct result is an ordered list with each question's estimated pass rate and the cumulative expected incidence after each step, which is also the input to feasibility in Step 8.

**Step 4. Disguise the target.** A screener should never let a respondent work out what the qualifying answer is. Five devices. **Embed the criterion in a list**, so the category of interest sits among plausible others and selecting it carries no signal. **Ask behaviour, not category membership:** "which of these have you done in the last month" rather than "are you a frequent shopper". **Avoid single yes/no eligibility questions**, which have a fifty per cent guess rate and an obvious right answer; use a frequency scale or a list instead, and set the qualifying range afterwards. **Neutralise the framing:** the invitation and the introduction should describe the subject broadly enough that they do not name the qualifying behaviour. **Vary what looks important:** ask about several categories with the same seriousness, so that the one that matters is not the one with the follow-up questions. A screener that passes this step can be shown to a motivated respondent without telling them how to qualify, and that is the test to apply.

**Step 5. Build the exclusions.** Four families, each for a different reason. **Industry exclusions:** employment, or a household member's employment, in the client's sector, in market research, in advertising or marketing, and in journalism. The reason is not only confidentiality; category professionals answer as professionals, and their responses are systematically different in ways that look like unusual insight. **Competitor exclusions**, where commercial sensitivity or exposure to internal information applies. **Recent participation:** a stated window, applied to the same category or the same client, because repeat participants become fluent in research and their answers become more articulate and less representative. **Relationship exclusions:** a household member who has already taken part, which matters most in small-sample qualitative work and in any study with a high incentive. State the window and the rationale for each exclusion, because they cost incidence and someone will eventually ask why the sample was harder to find than expected.

**Step 6. Design the fraud and professional-respondent controls, and price their false positives.** In incentivised research the qualifying answer has a cash value, and part of any open sample will be people optimising for it, including some doing it at scale. Layered controls work better than any single one. **Internal consistency:** the same fact asked in two forms at a distance, for example a birth year early and an age band later. **Impossible or implausible combinations:** a set of holdings, roles or behaviours that almost nobody genuinely has at once. **A knowledge check**, for specialist audiences, on something a genuine member of the population would know without effort and an outsider would have to search for. **An open-ended verification question**, which is the strongest single control for specialist samples because it is expensive to fake and cheap to evaluate: ask for a short description of something only a real member could describe, and read the answers. **Behavioural signals** from the fielding process: implausible completion speed, duplicate identifiers, and mismatches between claimed and observable location, all applied as flags rather than automatic removals. Then price the controls. Every control has a false-positive rate, and a legitimate respondent removed is not merely a lost interview: if the false positives are correlated with anything (low literacy, second-language respondents, older respondents, people who answer carefully and slowly), the control has introduced a sample bias while appearing to improve quality. State the intended action for each control, and decide in advance whether a flag terminates, deducts, or is recorded for review, per **K5 §2.4**.

**Step 7. Write the trap questions carefully, or not at all.** A trap question is a deliberate test with a known correct answer: an attention instruction embedded in a grid, an implausible item in a list of things the respondent might have done, or a fictitious brand in an awareness list. They work, and they are more dangerous than most researchers assume. Three disciplines. **Make the failure unambiguous:** an item that a careless reader could plausibly select for an honest reason is not a trap, it is a coin toss with consequences. **Use very few:** one or two in a screener, and a fictitious item only where the list is long enough that its presence is not conspicuous, since a respondent who spots the trap now knows the researcher is testing them and changes how they answer everything afterwards. **Never terminate on a single trap alone** unless the failure is unambiguous, and log every trap result so the false-positive rate can be estimated. A fictitious brand in an awareness list has an additional cost that is often forgotten: it contaminates the awareness measure it sits in, so it belongs in the screener rather than in a tracked awareness question.

**Step 8. Estimate incidence and test feasibility before the study is priced.** Incidence multiplies down the criteria chain, and the multiplication is where feasibility usually dies. Take each criterion's estimated pass rate and multiply, being explicit that criteria are rarely independent (people who do one thing in a category often do another, so multiplying independent probabilities usually understates incidence, while assuming full overlap overstates it). Produce a range rather than a point estimate, state the source of every input, and mark unsourced inputs as assumptions per **K4 §2.1**. Then convert: expected screen-outs per completed interview, expected screening cost, expected fieldwork duration, and the effect on the incentive required. Show the sensitivity, because it is the sensitivity that persuades: an incidence assumption of 8% that turns out to be 3% roughly triples the screening cost and can turn a viable study into an unviable one, and that arithmetic is far more useful to a research director than the point estimate. Where incidence is genuinely unknown, recommend measuring it: a short omnibus question or a soft launch on a small share of the sample buys a real number before the budget is committed.

**Step 9. Build the quota frame from the analysis plan, not from precedent.** A quota is only justified for two reasons, and they are different. A **representative quota** matches a known population distribution, and it requires that the distribution is actually known from a defensible source, and that it is the right one for the population being studied. A **design quota** guarantees a minimum base on a subgroup the analysis must report, whether or not it is proportional. Confusing them produces the common error of a "representative" sample that cannot report on the small segment the study was commissioned to understand. For each proposed quota variable ask: what analysis requires it, what happens if it is left free, and what is the cost of controlling it. Variables that fail all three come out of the frame, because every quota variable makes fieldwork harder, slower and more expensive, and quotas are not free precision.

**Step 10. Decide interlocking cell by cell, and do the arithmetic.** Non-interlocking (marginal) quotas control each variable independently: the sample matches on age and matches on region, but nothing guarantees the joint distribution, so a sample can be perfectly correct on both margins and contain almost no young people in the smallest region. Interlocking quotas control the combination, which delivers the joint distribution and multiplies the number of cells: three age bands by four regions by two user types is 24 cells, and if the smallest is 2% of the population, a sample of 800 gives it 16 respondents. Two rules. **Interlock only where the joint distribution matters to the analysis**, which is usually where subgroups will be compared jointly or where a known interaction exists between the variables. **Check the smallest cell before agreeing the frame**: calculate every cell's expected size at the planned sample and flag any that fall below the reportable minimum, because a cell too small to report is a cell that made fieldwork harder for no analytical return. Where a frame is unworkable, the options are a larger sample, fewer interlocked variables, a partially interlocked frame (interlocking the pair that matters and leaving the rest marginal), or a boost on the small cell with the analysis consequences stated.

**Step 11. Set the over-recruitment and replacement policy in writing, before fieldwork.** Qualitative studies over-recruit against no-shows at a rate that should be stated, not assumed, and should reflect the audience: high-value professionals cancel more than general consumers. Quantitative studies over-recruit against removals for quality. Both need a replacement rule agreed in advance: is a replacement drawn from the same cell, from the same source, at the same point in the fieldwork window, and how many replacements are permitted before the composition of the sample has materially changed. Replacement is not neutral. Replacing a hard-to-reach respondent with an easy-to-reach one from the same cell keeps the quota correct and changes the sample, and if that happens repeatedly the study drifts toward the most available members of every cell, which is a coverage problem the quota frame is specifically unable to detect. Record every replacement, so the drift is visible at analysis rather than invisible.

**Step 12. State the analysis consequences of the frame you have built.** This is the step that is almost always missed, and it is where the screener's choices become the study's limits. Four consequences to write down explicitly. **Quota variables cannot be studied as outcomes:** if the sample was quota-controlled to 50% users, the study cannot report the proportion of users in the population, and someone will eventually try. **Quotas are not weights:** a quota controls who is in the sample, not how they are counted; where the achieved sample deviates or where the frame is marginal rather than interlocked, weighting may still be required, and that belongs to **04.05 Weighting and Base Management**. **The frame determines what can be compared:** subgroups guaranteed a base can be compared, and everything else is whatever fieldwork happened to deliver. **The screening criteria define the population every finding applies to:** the report's population statement is written from the screener, not from the brief, and per **K4 §3.3** a finding about "recent purchasers in the last three months who are not category professionals" must not be reported as a finding about consumers.

**Step 13. Pilot the screener on live sample before committing the budget.** Soft launch to a small share and inspect four things: the actual incidence against the estimate, the distribution of screen-out reasons (which reveals whether the criteria are doing what was intended, and which single criterion is doing the eliminating), the trap and consistency flag rates against expectation, and the fill rate by quota cell, which shows early which cells will be the constraint. A screen-out reason distribution that differs sharply from the plan is the most useful early warning available: it means either the incidence assumption was wrong or the criteria are being read differently by respondents than by the author.

## 8. Analytical framework

Every element of the screener is built and checked on one chain, readable in both directions:

    Population definition → Observable criterion → Screening question → Eligibility rule → Quota cell → Analysis base → Population statement in the report

**Forward** is the design test: this element of the target is observable in this way, asked by this question, evaluated by this rule, which places the respondent in this cell, which delivers this base, which supports this comparison.

**Backward** is the discipline that keeps a study honest: this table is reported on this base, which comes from this cell, guaranteed by this quota, filled by respondents who passed this rule, which tested this criterion, which is what "the target population" actually meant in practice. The chain most often breaks at the second link, where an internal state from the brief was never converted into anything observable, and at the last, where the report describes a broader population than the screener admitted.

Two rules govern the whole frame. **A criterion that cannot be verified is a claim, not a fact**, and its confidence should be recorded per **K3 §3.4**. And **the population statement in the final report is written from the screener**: whatever the brief said, the study is about the people who passed these questions, in this window, from this source.

## 9. Output format

**1. Sample definition summary.** The target as stated in the brief, the observable translation, the source and the mode, and the exclusions applied.

**2. Criteria table.**

| Element of definition | Observable form | Screening question | Qualifying answer | Verifiable? | Confidence | Estimated pass rate and source |
|---|---|---|---|---|---|---|

**3. The screener**, in field order, with question text as the respondent will meet it, response options, terminate points, quota assignment points, and the disguise rationale for any question whose form is deliberate.

**4. Exclusion list**, with the window and the reason for each.

**5. Fraud and quality controls**, each with its trigger, its intended action (terminate, flag, review), its expected false-positive cost, and who decides.

**6. Incidence and feasibility estimate.**

| Criterion | Estimated pass rate | Source | Cumulative incidence |
|---|---|---|---|

With expected screen-outs per complete, a sensitivity range, and every unsourced figure marked as an assumption.

**7. Quota frame.** The grid, with target and minimum per cell, marked as representative or design quotas, showing what is interlocked and what is marginal, and flagging every cell whose expected size falls below the reportable minimum.

**8. Over-recruitment and replacement policy**, including the rule for when a cell cannot fill.

**9. Analysis consequences.** What this frame guarantees, what it prevents, which variables cannot be reported as population estimates, and the population statement as it should appear in the report.

**10. Open items and review points**, per **K5 §3**.

**When the inputs are thin**, the format is not filled in anyway. An incidence with no source is written as an assumption with its basis stated, never as a figure. A criterion that cannot be made observable is recorded as unresolved rather than replaced with a plausible proxy, per **K4 §1**. A quota frame that cannot deliver its smallest cell is reported as unworkable rather than presented with an optimistic target.

## 10. Quality checks

Run before the screener goes to field. These sit on top of **K4 §8**.

1. Every criterion is observable, answerable and stable, and every attitudinal criterion has been converted or explicitly justified.
2. Every screening question determines eligibility or a quota; nothing that belongs in classification is being asked of screen-outs.
3. Questions are ordered cheapest and most-eliminating first, with the cumulative incidence shown.
4. No question reveals the qualifying answer, and no eligibility rests on a single yes/no question.
5. Every exclusion has a stated window and reason.
6. Fraud controls are layered rather than single, and each has a stated action and an acknowledged false-positive cost.
7. Trap questions, if used, have unambiguous failures, are few, are logged, and do not contaminate a measure that will be reported.
8. Incidence is estimated with a range, every input has a source or is marked as an assumption, and the cost sensitivity is shown.
9. Every quota variable is justified by a named analysis requirement, and marked as representative or design.
10. Interlocking decisions are made cell by cell, and every cell's expected size is calculated against the reportable minimum.
11. The smallest cell is achievable at the planned sample and from the intended source.
12. An over-recruitment and replacement policy exists in writing, with a rule for a cell that will not fill.
13. The analysis consequences are stated, including which variables cannot be reported as population estimates.
14. The population statement for the report has been drafted from the screener, not from the brief.
15. For a repeat wave, the screener is identical to the previous wave, or every change is logged as a break in comparability.

## 11. Common failure modes

| Failure | How to recognise it | How to prevent it |
|---|---|---|
| **The unobservable criterion** | "Decision makers", "regular users", "people who care about X" pass into the screener unchanged | Step 1. Convert to an act, a frequency or a role in a specific recent event |
| **The transparent screener** | A respondent can tell within two questions which answer qualifies them | Step 4. Embed in lists, ask behaviour, avoid single yes/no eligibility |
| **Screener as questionnaire** | Attitudinal and descriptive questions asked of everyone, including screen-outs | Step 2. If a wrong answer does not change eligibility or cell, it is not a screening question |
| **Optimistic incidence** | A point estimate with no source, and a proposal priced on it | Step 8. Range, sources, sensitivity, and a soft launch measurement before commitment |
| **Precedent quotas** | The frame matches last year's study and nobody can say which analysis needs it | Step 9. Every quota variable justified by a named analysis requirement |
| **Marginally right, jointly wrong** | The sample matches on every variable separately and contains no members of a key combination | Step 10. Interlock where the joint distribution matters, and calculate the smallest cell |
| **The unfillable cell** | Fieldwork stalls at 94% with one cell short, and criteria get quietly relaxed at the deadline | Calculate cell sizes in advance; agree the relaxation and replacement policy in writing before fieldwork |
| **Silent criterion drift** | A recruiter widens a window or accepts a near-miss to fill a quota, and nobody records it | Every relaxation logged as a change to the population, with the report's population statement updated |
| **Over-trapped screener** | Six quality checks, a high termination rate, and a sample skewed toward fast, confident, literate respondents | Price the false positives. Controls that correlate with a respondent characteristic introduce the bias they were meant to prevent |
| **Replacement drift** | Every hard-to-reach respondent replaced by an easy one, and the quota still looks perfect | Log replacements; a quota frame cannot detect this by construction |
| **Quota variable reported as a finding** | "62% of the market are users", from a sample quota-controlled to 60% users | Step 12, written into the output and repeated in the report's method note |
| **AI: invented incidence and feasibility** | Confident pass rates, screen-out ratios and industry incidence figures with no source | **K4 §2.1, §2.4**. Mark every unsourced input as an assumption and show the sensitivity |
| **AI: the plausible quota frame** | A tidy, symmetrical grid matching population proportions nobody supplied | **K4 §2.5**. Population distributions come from a named source or the frame is marked as unverified |
| **AI: comprehensiveness over economy** | A twenty-question screener covering everything the brief mentions | Step 2. Screener length is paid for on every screen-out, and every extra question teaches the respondent more |
| **AI: fraud controls without cost** | Layers of traps and consistency checks proposed with no false-positive discussion | Every control carries an action and a stated cost, and the decision to terminate on a flag is a human one |

## 12. AI guardrails

Skill-specific only. **K4** applies in full and is not repeated here.

1. **Never state an incidence, a pass rate, a screen-out ratio or a population distribution as a fact without a named source.** These figures determine whether a study is commercially viable, and a plausible-sounding number here is more damaging than almost anywhere else in the library (**K4 §2.1, §2.4, §2.5**).
2. **Never loosen a criterion to make a quota fillable.** A criterion that cannot be filled is a feasibility finding. Changing it changes the population, and doing so silently makes every finding a claim about a group nobody defined (**K5 §2.7**).
3. **Never present a self-reported screening answer as verified.** Record which criteria are verifiable and which rest on the respondent's word, and attach confidence accordingly (**K3 §3.4**).
4. **Never recommend terminating on a single automated flag without stating the false-positive cost and who decides.** Removing legitimate respondents in a way correlated with literacy, language or age is a sample bias created by a quality control (**K5 §2.4**).
5. **Never invent an exclusion list, a competitor set or an industry list.** These come from the client. Where required and not supplied, output a labelled placeholder naming what is needed (**K4 §2.1**).
6. **Never present a quota frame as delivering representativeness without naming the source distribution and the population it represents.** A quota that matches an unnamed benchmark represents nothing (**K4 §2.5, §3.3**).
7. **Never carry the brief's population wording into the output as though it were the achieved sample.** The population statement is written from the screening criteria actually applied, and it belongs in the report.
8. **Never treat a quota-controlled variable as measurable in the population.** Where the frame fixes a proportion, the study cannot estimate that proportion, and the output says so.

## 13. Best-practice principles

1. **No amount of analysis fixes a bad screener.** This is the governing principle. Every other stage has a remedy; the wrong people is terminal, and the data will not look wrong.
2. **Screen on what people do, not on what they are.** Behaviour is checkable, stable and hard to fake convincingly. Identity and attitude are none of these under an incentive.
3. **A screener that can be gamed will be gamed, proportionally to the incentive.** Design so that reading the question does not reveal the qualifying answer, then test that by showing it to someone and asking them to try to qualify.
4. **Eliminate cheaply and early.** Screener economics are dominated by the questions asked of people who will never enter the study.
5. **Every criterion costs incidence, and incidence costs money and time.** A criterion nobody can justify against the objectives is a tax on the study, paid in fieldwork.
6. **Multiply the criteria before you price the study, not after.** Incidence chains collapse fast, and the arithmetic takes ten minutes.
7. **Quota only what the analysis requires, and know why each one is there.** Precision on a variable nobody will report costs fieldwork and buys nothing.
8. **Marginal quotas can be perfectly correct and jointly absurd.** Interlock where the combination matters, and always calculate the smallest cell before agreeing the frame.
9. **Quotas are not weights and do not correct for anything.** They control who is in the sample, not how the sample is counted.
10. **Every quality control has a false-positive rate, and if the false positives are correlated with anything, the control has introduced a bias.** Price them, do not just add them.
11. **Replacement is never neutral.** Repeatedly replacing hard-to-reach respondents drifts the sample toward the most available members of every cell, and the quota frame is structurally unable to see it.
12. **The report's population statement is the screener, written out.** Whatever the brief said, the study is about the people who passed these questions, in this window, from this source.

## 14. Worked example

**INPUT**

A fictional medical device manufacturer, Wenlock Medical, is planning a study of infusion pump usability ahead of a redesign. The target as written in the brief: "senior nurses and procurement decision makers with experience of infusion pumps in acute hospital settings". Required sample: 200 quantitative interviews with reportable subgroups by role (clinical, procurement) and by hospital size (large, small), plus 12 depth interviews. Incentive per completed interview is high, as is standard for specialist clinical audiences. No incidence data supplied.

**PROCESS**

*Step 1, observable criteria.* Three of the four elements of the definition were unusable as written. "Senior nurses" became a combination of current role title, years in post, and whether the respondent personally operates infusion pumps in their current role, because seniority claimed in the abstract is inflated and job titles are not comparable between hospitals. "Procurement decision makers" became a specific act: whether the respondent had a defined role in the most recent purchase or renewal decision for this equipment category, with the role selected from a list including "no involvement". "Experience of infusion pumps" became a frequency question with a defined window. "Acute hospital settings" was the only element that survived directly, as a setting selection from a list.

*Step 2, separation.* Six of the eleven proposed screening questions were moved to classification, including specialty, years qualified and shift pattern. None of them determined eligibility, and asking them of every screen-out in a low-incidence specialist study is expensive.

*Step 3, order.* Setting first (fast, and removes the large share of clinical respondents who do not work in acute settings), then equipment contact frequency, then role, then the purchase-involvement question, which is the slowest and most specific and is therefore asked of the fewest people. Cumulative incidence estimated after each step.

*Step 4, disguise, and the judgement call.* The client's first draft opened with "Do you use infusion pumps in your work?" as a yes/no. At this incentive level, that question tells a motivated respondent both the target and the answer. It was replaced with a frequency grid covering six categories of clinical equipment, of which infusion pumps were one, with no visual emphasis and no follow-up structure that marked it out. The client objected that this made the screener longer for everyone. The trade-off was recorded rather than argued: the extra length is paid on every respondent, and the protection is paid for once, and in a high-incentive specialist study with an open sourcing route the risk of a self-selected sample of people who guessed right is the larger cost. Recorded in the design rationale so the decision is visible.

*Step 6, fraud controls.* Four layers, because a high incentive and a specialist claim is the highest-risk combination in commercial research. A consistency check pairing years qualified in the screener against qualification year in classification. An impossible-combination check across the equipment frequency grid, where claiming frequent personal use of every category listed is implausible for any single role. A knowledge question on routine clinical practice that any genuine member of the population answers without thinking. And a short open-ended verification asking the respondent to describe, in their own words, what they do when a specific common alarm condition occurs, which is expensive to fake and quick to read. The open-ended was designated as the primary control and marked for human review rather than automated scoring, per **K5 §2.4**, because rejecting a genuine clinician on an automated text rule is both a sample bias and a reputational risk with a professional audience.

*Step 7, traps.* A fictitious device brand was proposed for the awareness list. It was retained in the screener only, with the reasoning recorded: had it sat in the tracked awareness question, it would have contaminated a measure the client reports.

*Step 8, incidence and feasibility.* No incidence data existed. Pass rates were estimated with sources named where any existed and marked as assumptions where none did, producing a cumulative incidence range rather than a point estimate. The sensitivity was the finding that changed the plan: at the optimistic end the study was straightforward; at the pessimistic end the screening cost roughly tripled and the fieldwork window doubled. Rather than pricing the midpoint, the recommendation was to measure incidence first with a short screening exercise on a small share of the sample before committing the full budget, and the proposal was structured in two stages.

*Steps 9 and 10, quota frame.* Role by hospital size, fully interlocked, gives four cells. At 200 interviews the procurement-in-small-hospitals cell calculated to well below the study's reportable minimum, because procurement involvement is concentrated in larger institutions. Three options were presented with their costs: increase the total sample, boost that cell and report it as a boosted subgroup that cannot be aggregated into a total without weighting, or drop the hospital-size split for procurement and report procurement on total only. Marked **RESEARCHER DECISION REQUIRED**, since it trades budget against a comparison the client asked for.

*Step 12, analysis consequences.* Written out: the study cannot report the proportion of acute-setting clinicians who use these devices, because that proportion is fixed by the screener; the role split is a design quota and not representative of the professional population; and the report's population statement reads as the screening criteria, not as "nurses and procurement professionals".

**OUTPUT**

A criteria table converting four brief elements into five observable criteria with confidence attached; a five-question ordered screener with disguise rationale; an exclusion list covering device manufacturers, market research, and recent participation with stated windows; four layered fraud controls each with an action, a false-positive cost and a named decision owner; a two-stage incidence estimate with a sensitivity range and every input sourced or marked as an assumption; an interlocked quota frame with one cell flagged as unworkable and three costed options; a written over-recruitment and replacement policy for the depth interviews; an analysis consequences section; and two **K5** points, on the unfillable cell and on human review of the open-ended verification.

## 15. Advanced usage

**Very low incidence populations.** Below roughly one or two per cent, screening inside the main instrument becomes the dominant cost and the study should usually be restructured: a very short standalone screening instrument, fielded broadly and cheaply, with qualifiers re-contacted for the main survey. This changes the design in ways that need planning rather than improvising: re-contact consent must be collected at screening, the delay between screening and interview introduces attrition and a possible change in status, and the two-stage structure needs its own quality controls because the incentive to misrepresent now sits at the cheap stage. Where a client list or an administrative frame exists, using it will usually beat any amount of screening, and that is a sourcing decision for **03.01**.

**B2B and professional audiences.** Role titles are not comparable between organisations, so screen on decision authority and on specific recent acts rather than on titles. Firmographic criteria (size, sector, structure) are frequently answered wrongly by respondents who do not know their own organisation's figures, so prefer bands, offer "don't know", and verify against an external source where one exists. Professional audiences also have a specific fraud profile: the impersonation is more skilled, the incentive is higher, and knowledge-based verification works better than behavioural flags.

**Longitudinal and tracker screening.** The screener is part of the trend. Changing a window, a role list or a frequency threshold changes the population while every other number stays comparable-looking, which is the most expensive undetected error in tracker work. Where a change is unavoidable, it is a parallel-run decision. In panel-based longitudinal work, add a re-qualification step at each wave, because eligibility changes: people move roles, leave categories and change circumstances, and a panel that was screened correctly two years ago is not screened correctly now.

**Qualitative screening and articulacy.** Qualitative screeners carry an extra function beyond eligibility: they select for people who can talk. This is legitimate and it is also a bias, and it should be recorded rather than pretended away. An articulacy question ("tell me briefly about the last time you...") is defensible when the sample is small and the session depends on the participant sustaining a conversation, and indefensible if it is used to select participants who will say interesting things about the subject, which is selecting on the dependent variable. State which is being done. Group composition requirements, which come from **02.02**, become screening criteria here and frequently interlock, and a group frame is usually more constrained than a quantitative one because each session is a cell of its own.

## 16. Skill chain

**Recommended previous skills**
- **01.06 Sampling Strategy.** Hands over the population definition, the sample size, the subgroups that must be reportable and the minimum base for each, which is what the quota frame is built to guarantee.
- **01.07 Analysis Plan Development.** Hands over the comparisons the study has promised, which determine what must be quota-controlled and what must be interlocked.
- **02.01 Survey Questionnaire Design** or **02.02 Discussion Guide Design.** Hand over the screener block and, for qualitative work, the group composition requirements that become screening criteria.

**Recommended next skills**
- **02.05 Survey Logic and Flow Review.** Tests that the screening criteria, terminates and quota assignment points behave as intended once programmed.
- **03.01 Recruitment and Sample Sourcing.** Takes the criteria and the feasibility estimate and decides where the sample comes from.
- **03.04 Fieldwork Monitoring and Response Quality.** Watches incidence, screen-out reasons, quality flag rates and cell fill against the estimates made here.

**Runs well alongside**
- **04.05 Weighting and Base Management**, downstream, which inherits the quota frame and decides whether weighting is needed on top of it.
- **13.05 Research Ethics and Consent Design**, wherever screening touches health, vulnerability or special-category data, or where a quality control could exclude people unfairly.
- **03.05 Incentive and Participation Design**, because the incentive level directly determines how hard the screener has to work.

---
A Yazi Supplied Skill and resource.
