---
name: recruitment-and-sample-sourcing
description: >
  Chooses and specifies where research participants will come from, and states what
  that choice does to the findings. Use when someone says "where will we get the
  sample from", "can we recruit these people", "is this feasible", "what incidence
  should we assume", "should we use the client's customer list", "we need to blend
  sources", "how do we reach this hard-to-find audience", "how do we recruit B2B
  respondents", or asks how to describe the sample source in the methodology.
category: 03 Fieldwork and Data Collection
ref: 03.01
tier: 1
inherits: [K2, K3, K4, K5]
---

# Recruitment and Sample Sourcing

## 1. One-line description
Matches the sample frame agreed in the design to the sources that can actually reach it, and makes explicit what each source adds to and subtracts from the evidence before a single respondent is recruited.

## 2. What this skill is used for

**The research problem it solves.** Sourcing is routinely handled as a procurement question: who can supply 800 of these people, by Friday, for this budget. Treated that way it is invisible in the report, and its effects surface later disguised as findings. A sample recruited from a client's own customer list produces high satisfaction. A sample recruited from people who volunteer for research produces high category engagement and high stated purchase intent. A sample blended across two sources produces a difference between the halves that gets read as a segment difference. None of these are recruitment problems at the point they appear. They are analysis problems, and by then they cannot be fixed, only disclosed. The source is a design decision with analysis consequences, and the moment to make it deliberately is before fieldwork, not in the limitations section afterwards.

**Where it sits in the research lifecycle.** After the sampling strategy has defined the target population, the frame and the required base sizes, and after the instrument's screening requirements are known. Before fieldwork begins and before incentive structure is settled, since incentive interacts with source. It is the bridge between a sample defined on paper and a sample that exists.

**Typical use cases.**
- Deciding where a study's participants will come from and documenting why.
- Assessing feasibility before a proposal is priced or a timeline is committed.
- Sourcing a low-incidence or specialist audience where the obvious source will not reach them.
- Deciding whether to use a client's customer or user base, and what that constrains.
- Designing a blended-source approach and the controls that make it analysable.
- Specifying sourcing across several markets where the same nominal approach is not the same approach.
- Writing the sample-source disclosure that belongs in a methodology statement.

**Who uses it.** Research managers and directors scoping and specifying studies; client-side insight managers assessing what a supplier proposes; UX and product researchers recruiting from their own user base and needing to know what that costs them analytically; academic and public sector researchers working with non-probability sources under scrutiny.

## 3. When to use it

- A sampling strategy exists and the sample now has to be obtained from somewhere real.
- Feasibility is uncertain: the audience may be too rare, too busy or too dispersed for the intended approach.
- The client has offered their own customer, member or user list and someone has to decide whether to accept it.
- The audience is specialist, professional, clinical or senior, and general-population approaches will not reach them.
- More than one source will be used, whether by plan or because the first one will not deliver.
- The study runs in several markets and the sourcing approach has to be specified per market.
- The study will be published, peer-reviewed, submitted to a regulator or used in a public claim, so the source description will be read critically.
- A previous wave used a particular source and this wave's comparability depends on matching it.

## 4. When NOT to use it

- **The target population and frame are not yet defined.** Sourcing cannot settle who the study is about. Recruiting first and defining the population afterwards produces a study about whoever answered. Go to **01.06 Sampling Strategy**, which owns the population definition, the frame, the required base sizes and the quota structure. This skill takes those as given and finds a route to them.
- **The question needs a probability sample and no probability frame exists.** Where the objective is a population estimate with a quantified margin of error (official statistics, prevalence estimation, a regulatory submission), no amount of careful sourcing from volunteer or opt-in populations produces it. The honest output is that the design cannot support the claim, and either the claim changes or the study does. Do not source a convenience sample and describe the result as representative.
- **The decision has already been made and the sample is being assembled to support it.** Sourcing chosen to produce a particular answer (recruiting only recent buyers, only engaged users, only the client's advocates) is not a sampling decision. Say so and stop.
- **The population is defined by something the source cannot verify and the study cannot afford to get wrong.** Clinical diagnosis, regulated professional status, verified income, security clearance, actual purchase of a specific product. Where the claim rests on the respondent genuinely being that person and self-report is the only check available, the study's validity rests on an unverified screener answer. Either add verification, change the claim, or do not field.
- **Quota and screener detail is what is actually needed.** The internal logic of the screener, the quota cell structure, the interlocking and the fallback order belong to **02.06 Screener and Quota Design**. This skill decides where to knock; that skill decides what to ask at the door.
- **Fieldwork has started and the question is quality in field.** Speeding, straight-lining, duplicate identity, drop-off and quota drift while a study is live belong to **03.04 Fieldwork Monitoring and Response Quality**. Sourcing decisions taken mid-field are almost always damage control, and they must be logged as changes to the design, not made quietly.
- **The participants are vulnerable, dependent on the recruiting organisation, or reachable only through a gatekeeper with power over them.** Patients recruited by their clinician, employees by their manager, service users by the service being evaluated, children through a school. The sourcing route creates a consent problem that a sourcing skill cannot resolve. Route to **13.05 Research Ethics and Consent Design** before recruiting.
- **Nothing about the source will be disclosed.** If the study's methodology statement will not say where the sample came from, the sourcing decision cannot be assessed by anyone downstream and the finding cannot be properly read. A refusal to disclose the source is a reason not to proceed, not a formatting preference.

## 5. Required inputs

**Required. Without these the skill cannot run.**
- **Target population definition and sample frame**, from the sampling strategy: who counts, who does not, and what the frame is a list of. Without a definition, stop and ask, because everything below is an assessment against it.
- **Required base sizes, in total and for every subgroup that must be reportable.** Feasibility is arithmetic and it needs the numerator. Without this, stop and ask.
- **The screening criteria that define eligibility**, at least in outline, because they determine incidence and therefore cost, timeline and the depth of the funnel.
- **Mode and fieldwork window.** Self-completion online, telephone, face to face, in-home, mobile diary, moderated session. Mode restricts sources more than most other constraints, and the window determines whether slow sources are usable at all.

**Optional, and what each one adds.**
- **Known or estimated incidence in the general population**, from a previous wave, published statistics or client data. Turns feasibility from a guess into a calculation and sets the screening load. Without it, the funnel has to be estimated and the estimate labelled as such.
- **A client customer, member or user list, with its own provenance.** Makes an otherwise impossible audience reachable, at the cost of a known and describable selection. What matters is not the list but how it was built: all customers, or those who opted into marketing, or those who left an email address at a service event.
- **Previous wave sourcing detail.** Determines whether trend comparison is legitimate. A tracker whose source changes between waves has a break in it whether or not anyone documents it.
- **Market-level context for multi-market work.** Internet penetration, device profile, language distribution, literacy, the structure of the local research market. Determines whether the same nominal source is the same population in each market.
- **Verification options available.** Professional registration checks, purchase records, employer verification, referral from a known contact. Raises what the study can claim about who its respondents actually are.
- **Ethical approval conditions or regulatory constraints.** Some sourcing routes are prohibited in some sectors and jurisdictions. This changes the option set rather than the preference order.

## 6. Questions to ask before starting

1. **What is the population this study is supposed to describe, and what will be said about it in the report?** The strength of the sourcing requirement follows from the strength of the claim. A study that will say "among our customers who responded" needs much less than one that will say "consumers in this market". Default if unanswered: assume the weakest defensible claim and say so in the disclosure.
2. **What is the incidence of the eligible group in the source population, and how confident is that number?** Determines feasibility, cost, screening burden and whether the study is possible at all. Default: treat any incidence figure without a source as an assumption, state it, and build the plan to be robust if it is out by a factor of two.
3. **Does the client have a list, and how was that list built?** A list built from all account holders and a list built from newsletter subscribers are different populations wearing the same name. Default: assume any client list over-represents the engaged, and plan the disclosure accordingly.
4. **Is this comparable to anything?** A previous wave, a norm set, a parallel market, a published benchmark. Comparability usually constrains the source harder than quality does, because a better source that breaks the comparison may be a net loss. Default: assume standalone and flag where a convention exists.
5. **How much does it matter that respondents are genuinely who they say they are, and what verification is available?** Sets the identity-assurance level, which is a cost and timeline decision, not a technical one. Default: for anything specialist, professional or high-value, assume self-report alone is insufficient and say what verification would add.
6. **What happens if the sample cannot be filled?** Reduce the base, extend the window, relax a criterion, add a source, or stop. Deciding this in advance turns a fieldwork emergency into a documented plan. Default: assume the sample will under-deliver in the hardest quota cell and agree the fallback order before launch.
7. **Who will read the methodology, and how critically?** An internal working study, a public claim, a peer-reviewed submission and a regulatory filing tolerate very different sourcing. Default: write the disclosure to the most critical plausible reader.

## 7. Step-by-step methodology

**Step 1. Restate the frame as something reachable.** Take the population definition and rewrite it as an operational specification: the observable characteristics a person must have, how each one will be established, and by whom. "Small business decision makers" is a population. "People who own or hold budget authority in a business with 5 to 49 employees, self-reported, in one of four named sectors, who have purchased in the category in the last 12 months" is reachable. Every criterion in this restatement is a filter, and every filter costs incidence. A correct result is a specification in which each criterion is either verifiable, self-reportable, or flagged as unestablishable, with the last group listed explicitly because that list is a limitation whatever source is used.

**Step 2. Do the feasibility arithmetic before discussing sources.** Work the funnel from the required completes backwards: completes required, divided by the expected completion rate among those who start, divided by the qualification rate (incidence after all screening criteria are applied together, not each one separately, since criteria are usually correlated), divided by the expected response rate to an invitation. The product is the reach required. Then check it against the plausible size of each candidate source. Two failure signatures show up here and both are common. The first is compound incidence: five criteria each passing 60% do not pass 60%, they pass roughly 8% if independent, and the honest estimate depends on a correlation assumption that should be stated. The second is subgroup feasibility: a total of 800 is easy and 150 in the smallest of six interlocked cells is not, and it is the cell that governs. A correct result at this step is a funnel with every rate labelled as measured, estimated from a comparable study, or assumed, and a clear statement of which single number the plan is most sensitive to.

**Step 3. Profile each candidate source on coverage, selection and behaviour, not on convenience.** For each source available, write three things: who it can reach and who it structurally cannot (coverage), how a person came to be in it (selection mechanism), and what that implies about how they will answer (behaviour). The honest general statement is that no source is superior in the abstract, and each is a different trade of coverage against control against speed. The main families, and what each does to the research:

- **Standing research panels.** People who have agreed in advance to be contacted for research. They deliver speed, profiled variables that reduce screening load, and quota control. They exclude anyone unwilling to join such a panel, which is not a random exclusion: joiners tend to be more category-engaged, more comfortable online, and more willing to give opinions to strangers than the population they are drawn from. Tenure matters (Step 4).
- **Client customer, user or member lists.** They deliver a real, verifiable relationship with the subject matter and often a rich behavioural record to link to. They cannot reach non-customers, lapsed customers who were removed, or customers who opted out of contact, and the last group is the one most likely to be dissatisfied. Any absolute measure from this frame (satisfaction, advocacy, brand perception) is a measure of the surviving, contactable, consenting customer base, which is not the customer base.
- **River, intercept and open-link recruitment.** Respondents are recruited at the point of an unrelated online activity, or by a link circulated without a closed frame. It reaches people who would never join a panel, which is a genuine coverage advantage, and it gives up almost all control over who arrives. There is no denominator, so response rate is undefined, and quota management has to happen at the door.
- **Social and paid digital recruitment.** Fast, broad, and able to target narrow interests that no panel variable captures. The targeting mechanism that makes it work is also a selection mechanism nobody can fully describe, and the campaign's own optimisation will drift the sample toward whoever responds most cheaply, which is a systematic and undocumented bias unless the delivery is controlled.
- **Specialist recruiters working from their own networks.** Highest touch, most suitable for depth qualitative, senior and clinical audiences, and studies requiring real verification. They are slower, more expensive, and carry a network effect: a recruiter's contacts resemble each other, and repeated use of one recruiter narrows the range of people the research ever meets.
- **Snowball and referral.** Often the only route into closed, stigmatised or highly specialist populations, and the referral chain is the sampling mechanism, so the achieved sample reflects the social structure of the seeds. Findings describe a connected network, not a population, and this must be stated rather than treated as an inconvenience.
- **In-person community and intercept sampling.** Reaches people no digital route reaches, including those with low connectivity or literacy, and produces location-bound samples whose composition is determined by who is at that place at that time. Time of day and day of week are sampling parameters and should be treated as such.
- **An existing user base recruited in-product.** Excellent for questions about the product experience, and structurally unable to answer questions about people who never adopted, churned or never heard of it. In-product intercepts also over-sample heavy users, because heavier use means more chances to see the invitation, which biases the sample toward the people whose experience is least representative of the median.

A correct result at this step is a short table, one row per source, with coverage, exclusion, selection mechanism and expected directional effect on the key measures.

**Step 4. Assess professional respondent risk and tenure effects.** Where a source contains people who participate in research repeatedly, three distinct effects apply and they should be separated rather than lumped as "professional respondents". **Conditioning**: experienced respondents learn how instruments work, recognise screener patterns and category batteries, and answer more consistently and less spontaneously. **Motivation drift**: where participation is driven mainly by the incentive, effort declines and the incentive becomes a reason to qualify rather than a reason to answer well. **Deliberate misqualification**: a small share of respondents will claim eligibility they do not have, and this rises with incentive value and with how obvious the screening criteria are made. These are different problems needing different responses: conditioning is addressed by tenure and recency-of-participation limits and by not relying on unaided awareness from long-tenured respondents; motivation drift by incentive design (**03.05**); misqualification by non-obvious screening, factual verification items, and consistency checks placed apart in the instrument. State which controls are in place, and where none is available, state that too.

**Step 5. Decide single source or blend, and if blending, design it as an experiment you are running on yourself.** A single source gives an internally consistent sample whose biases, whatever they are, are at least uniform. A blend improves coverage and reduces dependence on one source's peculiarities, and it introduces a systematic difference between subsamples that will show up in the data and be misread as a real difference. If blending is chosen, four controls make it analysable: (a) record source as a variable on every respondent, without exception, (b) allocate quotas within source rather than across it, so composition is comparable, (c) hold the invitation, screener and instrument identical across sources, so the only difference is the source, and (d) plan a source-effect check into the analysis, comparing key measures by source before any substantive analysis runs. If a source difference is found, it is reported as a source effect and the source variable is carried into the analysis, not averaged away. Blending decided mid-field without these controls is the single commonest way a study becomes uninterpretable.

**Step 6. Design the recruitment approach itself, because it selects too.** The invitation is part of the sampling mechanism. A subject line naming the topic recruits people interested in the topic, which is the most common self-inflicted bias in the whole process. Keep invitations neutral about the subject and honest about the burden. Place the qualifying criteria so that eligibility is not inferable from the invitation or from the first screener question. Over-recruit against the hardest cells first, since easy cells fill themselves and the study ends up over-weighted toward them. Where quotas will be enforced, decide in advance what happens to over-quota respondents, because turning people away after they have answered ten questions damages the source and the ethics of it are not neutral.

**Step 7. Handle low-incidence and hard-to-reach populations by changing the structure, not the effort.** Where incidence is low enough that screening cost dominates, four structural moves are available: run a short standalone screening survey and re-contact qualifiers (which needs consent for re-contact obtained at screening), screen against pre-profiled variables where they exist and are recent enough to trust, recruit through organisations and communities where the population is concentrated (accepting the coverage limit that creates), or accept a smaller base and change the claim from measurement to exploration. Where the population is hidden rather than merely rare (stigmatised, undocumented, or defined by something people do not disclose), referral-based approaches may be the only route, and the design should then treat the referral structure as data rather than as noise.

**Step 8. Treat B2B and specialist sourcing as a different problem.** Firmographic screening is compound (sector, size, role, budget authority, purchase involvement, category usage) and incidence collapses fast. Verification matters more, because respondents claiming senior roles they do not hold is a known failure and the incentive to misqualify is higher on high-value studies. Access is often through professional networks or specialist recruiters rather than open sourcing, and in some professions and some organisations, participation is restricted or prohibited by employer policy, which is a coverage constraint that biases the achieved sample toward smaller and less regulated employers. Plan longer fieldwork, expect scheduling rather than immediate completion, and set base expectations accordingly: a specialist study of 60 verified respondents may support more than 400 unverified ones.

**Step 9. Specify sourcing per market, and never assume the same nominal source is the same population.** In multi-market work the composition of any given source type varies by market, following local connectivity, the maturity of the local research market, urbanisation, language and how research is regarded. The consequence is specific and damaging: cross-market differences in the data will partly reflect differences in who the source reached, and they are indistinguishable from real market differences without a design that separates them. Mitigations, in order of strength: use the same source structure in each market and document where it was not possible, hold the frame definition constant even where the route to it differs, and quota to the same population targets in each market so composition is at least aligned on observables. State, in the output, which cross-market comparisons are safe and which are confounded with source.

**Step 10. Write the source disclosure before fieldwork, not after.** Draft the paragraph that will appear in the methodology: population, frame, source or sources with a description of how people entered each, screening criteria and how each was established, incidence achieved, quota structure and targets, whether the sample is probability or non-probability, and the specific inferential limit that follows. Writing this before fielding has an unusual property: a source that cannot be described defensibly in three sentences is usually a source that should not be used, and this is easier to act on before the money is spent.

## 8. Analytical framework

Sourcing is assessed along one chain, and every arrow loses people non-randomly:

    Target population → Frame → Source coverage → Invited → Responded → Qualified → Completed → Achieved sample

At each arrow, name the loss mechanism and its likely direction:

- **Population to frame**: definitional exclusion. Who is out of scope by design.
- **Frame to source coverage**: coverage error. Who cannot be reached by this source at all, and how they differ from those who can. This is the error no sample size fixes.
- **Coverage to invited**: selection into contact. Who the source chose to approach, or who happened to see the invitation.
- **Invited to responded**: non-response bias. Who chose to take part, and what interest, incentive or availability distinguishes them.
- **Responded to qualified**: screening. Both genuine ineligibility and misqualification in either direction.
- **Qualified to completed**: attrition. Length, difficulty and topic sensitivity remove people non-randomly, and 03.04 monitors this in field.

The achieved sample is what survived all six. The methodological question is never "is this sample perfect" (none is) but **"which of these six mechanisms could plausibly have produced the finding I am about to report, and can I rule it out?"** A finding that survives that question is reportable with its source described. A finding that does not is a finding about the sourcing.

Applied forward, the chain is the feasibility plan. Applied backward from a result, it is the first diagnostic when a number looks surprising: an unexpectedly high engagement score is more often an artefact of arrow two or four than a fact about the market.

## 9. Output format

**1. Sourcing summary.** Population, frame, required bases, mode, fieldwork window, and the claim the study intends to make about the population.

**2. Feasibility funnel.**

| Stage | Rate | Basis (measured / comparable study / assumed) | Implied reach required | Sensitivity |
|---|---|---|---|---|

Every rate is labelled. An assumed rate is written as assumed, never as a benchmark.

**3. Source assessment table.**

| Source | Reaches | Structurally excludes | Selection mechanism | Expected effect on key measures | Verification available | Speed and cost profile |
|---|---|---|---|---|---|---|

**4. Sourcing decision and rationale.** Which source or blend, why, what was rejected and why. Where a blend is used, the four blend controls in Step 5 are stated as commitments, not intentions.

**5. Quota and allocation plan by source**, showing which cells fill from where, and the over-recruitment plan for the hardest cells.

**6. Verification and identity-assurance plan**, proportionate to the claim, naming what will and will not be verified.

**7. Fallback plan.** Agreed in advance: the order in which base size, criteria, timeline or source will be relaxed if delivery falls short, and who decides.

**8. Draft source disclosure paragraph** for the methodology, in final wording.

**9. Limits of inference.** The specific claims this sample supports, and the specific claims it does not. Written as sentences a reader can check against, not as a generic caveat.

**10. Review points** per **K5 §3**, at the point of the decision.

**When the inputs are thin**, incidence is written as `[not established]` with the range the plan assumes, not as a figure. A source whose selection mechanism cannot be described is recorded as `selection mechanism not established`, which is itself a finding about that source. Do not populate the assessment table with plausible-sounding characterisations of a source nobody has described (**K4 §2.5**).

## 10. Quality checks

Run before the sourcing plan is agreed. These sit on top of **K4 §8**.

1. Every screening criterion is marked verifiable, self-reported or unestablishable, and the unestablishable list is disclosed.
2. The feasibility funnel uses compound incidence, not the product of separately quoted single-criterion rates presented as independent, and the correlation assumption is stated.
3. The smallest required subgroup cell, not the total, has been used to test feasibility.
4. Every source in the plan has a written selection mechanism, in a sentence, and a stated coverage exclusion.
5. The expected direction of source bias on each key measure has been stated in advance, so it cannot be rationalised afterwards.
6. Where the sample is blended, source is captured as a variable, quotas are set within source, and a source-effect check is in the analysis plan.
7. The invitation wording has been checked for topic disclosure that would recruit interest rather than a cross-section.
8. Where the study is a repeat wave, the source matches the previous wave, or the break is documented and its effect on trend stated before fielding.
9. Multi-market plans state which comparisons are confounded with source composition and which are not.
10. Professional respondent controls are specified where the source contains repeat participants, or their absence is disclosed.
11. The sample is correctly labelled probability or non-probability, and no margin of error is attached to a non-probability sample (**K3 §6**).
12. The draft disclosure paragraph exists, and describes the source specifically enough that a critical reader could assess it.
13. The fallback plan names a decision owner and a threshold, not just an intention to review.
14. Where recruitment runs through a gatekeeper with authority over participants, the ethics route has been taken (**13.05**) before the plan is finalised.

## 11. Common failure modes

| Failure | How to recognise it | How to prevent it |
|---|---|---|
| **Sourcing as procurement** | The plan names a supply route and a price and says nothing about coverage | Require the selection mechanism sentence for every source before cost is discussed |
| **The client list assumed representative** | Absolute satisfaction or advocacy figures reported as market measures | State the frame in the finding: "among contactable customers who responded" |
| **Compound incidence ignored** | Feasibility looks fine on each criterion and collapses in field | Calculate incidence on all criteria applied together, with the correlation assumption stated |
| **Cell blindness** | Total base delivered, key subgroup at 40, analysis plan unusable | Test feasibility on the smallest reportable cell from the start |
| **Silent blending** | A second source is added mid-field to fill quotas and never recorded | Source is a mandatory variable from launch, whether or not a blend is planned |
| **Source drift in a tracker** | A wave moves without explanation and the source changed too | Sourcing is part of the tracker specification and changes are treated as a break in series |
| **Topic-disclosing invitation** | High engagement, high category interest, unusually low "don't know" | Neutral invitations; check the subject line as part of the sampling design |
| **Verification theatre** | A screener asks respondents to confirm their profession and this is called verification | Distinguish self-report from verification in writing; claim only what was checked |
| **Multi-market false differences** | One market is an outlier on several unrelated measures at once | Compare source composition first; treat a broad outlier pattern as a sourcing signal until ruled out |
| **AI: invented incidence and response rates** | Confident figures for how common the audience is or how many will respond | **K4 §2.4**, **§2.5**. These are measured or assumed, never recalled. Label the basis of every rate |
| **AI: generic source characterisation** | Fluent paragraphs about a source's strengths and weaknesses with nothing study-specific | Require coverage and exclusion to be written against this study's frame, not in general |
| **AI: representativeness language by default** | The phrase "nationally representative" appears without a probability design or a stated quota basis | Reserve the term for a specified meaning (quota-matched on named variables to a named population statistic) and state which |
| **The unspoken fallback** | Fieldwork underdelivers and criteria are quietly relaxed | Fallback order agreed and logged before launch; any relaxation is a documented amendment |

## 12. AI guardrails

Skill-specific only. **K4** applies in full and is not repeated here.

1. **Never state an incidence rate, response rate, completion rate or feasibility estimate as a known figure.** Each is either measured from a named source, taken from a named comparable study, or assumed. Label which, every time (**K4 §2.4**).
2. **Never describe a sample as representative, nationally representative or population-matched** unless the basis is stated: probability design, or quota-matched on named variables to a named population source. Absent that, describe it as a non-probability sample and say what it is quota-controlled on.
3. **Never characterise a source's composition from general knowledge.** If the way people entered a given list, network or pool has not been described, record it as not established rather than inferring it from the source type.
4. **Never recommend a source as generally best.** Sources trade coverage, control, speed and verifiability differently for different populations. Where a recommendation is made, it is made for this population and this claim, with the trade named.
5. **Never quietly blend sources or accept a blend without the four controls** in Step 5. A blend proposed without source capture and within-source quotas is reported as a threat to interpretability.
6. **Never treat a self-reported screener answer as verification**, and never write a methodology sentence that implies a characteristic was checked when it was self-declared.
7. **Never propose a recruitment route through an authority figure over the participant** (employer, clinician, teacher, caseworker) without flagging the consent implication and routing to **13.05**.
8. **Never omit the source from the methodology output.** Where the researcher does not wish to disclose it, state the consequence for interpretability rather than producing a methodology paragraph that reads as complete.

## 13. Best-practice principles

1. **Sourcing is a measurement decision.** Where the sample comes from changes the level of every absolute measure in the study, usually more than question wording does, and almost always in a direction that can be predicted in advance if anyone bothers to.
2. **Predict the bias before fielding and write it down.** A directional prediction made in advance is a test. The same observation made afterwards is a rationalisation, and everyone can tell the difference.
3. **Relative comparisons survive bad sourcing better than absolute levels do.** A biased frame that biases all cells equally still supports comparison between them. This is why a source that cannot support "45% are satisfied" may still support "satisfaction is 12 points lower in the north".
4. **The coverage question is who cannot possibly be here.** Non-response is visible and worried about. Coverage exclusion is invisible, and it is the error that no sample size, weighting scheme or quota structure repairs.
5. **Do not confuse control with quality.** A source with tight quota control produces a sample matching your targets on the variables you specified, and says nothing about the variables you did not.
6. **Incentive and source interact.** Raise the incentive on a source that contains repeat participants and you change who qualifies, not just how many. Decide sourcing and incentive together (**03.05**).
7. **A referral chain is a finding, not a nuisance.** In snowball recruitment, record who referred whom. The structure tells you what kind of network the sample describes, and it is the only honest basis for saying how far the findings travel.
8. **In-product and customer-list samples over-represent frequency.** More usage means more exposure to the invitation. Unless usage is quota-controlled, the sample will skew heavy, and heavy users are the least representative respondents in the base.
9. **Fewer, verified respondents often beat more, unverified ones.** For specialist and B2B work, the base size that supports the claim is the one where you know who answered. Say what verification was done and let the reader price it.
10. **Write the methodology paragraph first.** The discipline is cheap and it catches the sourcing decisions that cannot be defended, at the only point where changing them costs nothing.
11. **The invitation is part of the instrument.** It selects, it primes, and it sets expectations about length and topic. It should be reviewed with the same care as the first question.
12. **When the source cannot support the claim, change the claim.** Reducing the scope of what a study says is a legitimate, professional, and frequently correct response. Manufacturing representativeness in the write-up is not.

## 14. Worked example

**INPUT**

A fictional national pharmacy regulator, the Meridian Pharmacy Council, wants to understand why community pharmacists are under-reporting minor medication errors through its voluntary reporting system. Objectives: understand the reporting decision, identify barriers, and size how common each barrier is. Required: 300 completes among practising community pharmacists, reportable by three practice settings (independent, small chain, large chain) and by years in practice. Mode: online self-completion, 15 minutes. Window: five weeks.

**PROCESS**

*Step 1, operational frame.* Practising community pharmacist, registered on the national register, dispensing in a community setting in the last three months. Registration is verifiable against a public register number. Practice setting and years in practice are self-reported. Recency of dispensing is self-reported and unestablishable independently, and is flagged as such.

*Step 2, feasibility.* Community pharmacists are a small fraction of the adult population, and the three-setting requirement means the smallest cell (independents) governs. The funnel showed that a general-population source would need to screen an implausible number of people to reach 300, and the independents cell would fail regardless. Recorded as: general-population sourcing not feasible, not merely expensive.

*Step 3, sources.* Four candidates assessed. A professional-body membership route reaches a defined and verifiable population but excludes non-members, who are known to skew toward independents, exactly the hardest cell. Specialist professional recruiting reaches non-members but is slow against a five-week window. Social and professional-network recruitment reaches both but cannot verify registration and is vulnerable to misqualification, which matters here because the topic (error reporting) is one an unqualified respondent could plausibly answer. A regulator-held register mailing reaches everyone in scope with full verification, and carries the problem that the invitation comes from the body that receives the reports.

*The judgement call.* The regulator's own register gives perfect coverage and verified identity, and it is the worst possible source for this specific topic. Pharmacists asked by their regulator why they do not report errors to that regulator have an obvious reason to answer conservatively, and the direction of the bias is toward under-stating barriers that reflect on the respondent (fear of sanction, uncertainty about what is reportable) and over-stating neutral ones (time, system usability). Recorded as a **K5 §2.4 and §2.7** review point rather than decided unilaterally.

The resolution taken to the researcher was a split design: recruit from the register, so coverage and verification hold, but field through an independent research provider with an explicitly stated separation, no identifiers returned to the regulator, and results reported only in aggregate, with that separation described in the invitation rather than buried in a privacy notice. The residual risk, that respondents do not believe the separation, was accepted and disclosed, and one direct question was added asking how confident respondents were that their answers could not be traced, so the residual effect could at least be described rather than assumed away.

*Step 5, blend.* Rejected. A second source would have introduced a difference between register-recruited and network-recruited respondents on precisely the measure of interest (willingness to disclose), which would have been uninterpretable.

*Step 6, invitation.* The first draft named error reporting in the subject line. Rewritten to reference professional practice and workload more broadly, because a topic-naming invitation would have recruited the pharmacists most engaged with reporting, which is the opposite of the population of interest.

*Step 10, disclosure.* Drafted before fielding: census invitation to the full national register, non-probability achieved sample, verified registration, quota-monitored (not enforced) on practice setting, with the regulator-as-sponsor effect stated as an unquantified limitation in a stated direction.

**OUTPUT**

A single-source plan with verified identity, a documented sponsor-effect limitation with a direction, a rewritten neutral invitation, a fallback ordering (extend window, then relax the independents cell from 100 to 70 and report it as a small base per **K4 §7**, then accept), and two review points: whether the sponsor separation is credible enough to field, and whether independents can be boosted through a professional-body route in the final week without creating a within-cell source difference.

## 15. Advanced usage

**Sourcing for tracked studies.** Treat the source as part of the locked instrument. Where a source must change, run a parallel wave on both sources rather than switching cleanly, quantify the shift on the tracked measures, and either adjust or restate the series with a documented break. Switching without a parallel wave converts a trend into two unconnected numbers, and no amount of subsequent analysis recovers the join.

**Source effects as a designed comparison.** Where a study will be repeated and the sourcing is contested, deliberately field a small parallel subsample from an alternative source, with identical instrument and screening. This converts an argument about representativeness into a measurement of source effect on the specific measures that matter, which is worth more than any general claim about which source is better.

**Weighting and its limits.** Weighting corrects composition on the variables used to weight and nothing else. Where the selection mechanism is correlated with the outcome variable independently of the weighting variables (research-willing people being more opinionated, for instance), weighting to demographic targets leaves that bias intact while making the sample look corrected. Weight where composition is off on variables known to relate to the outcome, state the scheme and the effective base (**K4 §7**), and do not present weighting as a repair for coverage error.

**Very small specialist populations.** Where the eligible population is in the hundreds, sampling theory matters less than access and the analysis shifts toward census logic. Report counts rather than percentages below the thresholds in **K3 §3.2**, treat non-response as the dominant risk, and consider whether a well-conducted set of 25 verified depth interviews answers the question better than 60 thin survey responses would.

**When no adequate source exists.** The professionally correct output is sometimes that the study as scoped cannot be fielded. Say which claim fails, offer the reduced claim the available sourcing does support, and name what would be needed to support the original. This is a **K5 §2.7** judgement and it belongs to the researcher.

## 16. Skill chain

**Recommended previous skills**
- **01.06 Sampling Strategy.** Hands over the target population, the frame, the required total and subgroup bases, and the quota structure this skill has to find a route to.
- **01.04 Research Method Selection.** Establishes the mode, which constrains the source set more than most other decisions.
- **02.06 Screener and Quota Design.** Supplies the criteria whose combined incidence drives the whole feasibility calculation.

**Recommended next skills**
- **03.05 Incentive and Participation Design.** Takes the source decision and sets an incentive appropriate to that population and burden, since incentive and source interact.
- **03.04 Fieldwork Monitoring and Response Quality.** Takes the sourcing plan, the quota structure and the predicted source effects and watches for them in field.
- **04.01 Data Validation.** Inherits the source variable and the verification status of every respondent.

**Runs well alongside**
- **13.05 Research Ethics and Consent Design**, wherever recruitment runs through a gatekeeper, involves vulnerable participants, or uses a client list whose consent basis is unclear.
- **12.03 Research Report Compilation**, which publishes the source disclosure drafted here.
- **02.01 Survey Questionnaire Design**, since screening burden and instrument length are traded against each other.

---
A Yazi Supplied Skill and resource.
