---
name: weighting-and-base-management
description: >
  Designs, applies, interrogates and discloses survey weights, and governs how
  bases are reported once a dataset is weighted: what weighting can and cannot
  fix, target selection and the authority of the source, rim and cell approaches,
  weight efficiency and effective base size, capping and the trade-off it makes,
  the point at which a sample should be rejected rather than weighted, consistency
  across tracker waves, weighted versus unweighted base conventions, and subgroup
  analysis on weighted data. Use for "weight this data", "should I weight",
  "the sample is skewed", "what is the effective base", "these weights look
  extreme", "how do I report weighted results", "rim weighting", "weight capping".
category: 04 Data Preparation
ref: "04.05"
tier: 1
inherits: [K2, K3, K4, K5]
---

# Weighting and Base Management

## 1. One-line description
Decides whether a sample should be weighted at all, builds and interrogates the weights against targets whose authority is stated, reports effective base size rather than nominal n wherever weighting reduces precision, and makes the whole scheme visible in the output, on the principle that weighting adjusts for known composition differences and does nothing about anything else.

## 2. What this skill is used for

**The research problem it solves.** Weighting is the most misunderstood step in data preparation, and the misunderstanding runs in one direction: people believe it fixes more than it does. A sample recruited badly is weighted and then described as representative. A response rate nobody measured is weighted away. A weight variable ranging from 0.2 to 11 is applied without anybody calculating what it does to precision, and a subgroup reported as n=180 is carrying the statistical weight of about 70 interviews. A tracker changes its weighting scheme between waves and the resulting movement is reported as a market shift.

The correction is a single sentence that has to be stated plainly and repeatedly: **weighting adjusts the composition of the achieved sample to match known targets on the variables used, and does nothing else.** It does not fix coverage error, because a group the sample frame never reached cannot be up-weighted from zero. It does not fix non-response bias on characteristics that were not measured, because the adjustment can only act on what it can see. And it does not convert a non-probability sample into a representative one: matching a population on age, gender and region says nothing about whether the people recruited are like the population on the attitudes being measured, which is the thing the study exists to find out.

The second failure is the reporting one. Weighted percentages arrive with unweighted counts, or with weighted counts presented as numbers of interviews, or with no effective base at all, and every downstream small-base rule is then applied to a number that overstates the precision available. **A heavily weighted sample is evidence of a recruitment problem, not a solution to one**, and the diagnostics that reveal it should be read as a finding about the fieldwork rather than as a technical footnote.

**Where it sits.** The last step of preparation, after cleaning, missing-data decisions and transformation, and before any figure is reported.

**Typical use cases.**
- Deciding whether a study should be weighted at all, and on what.
- Building a weighting scheme against defensible population or customer-base targets.
- Diagnosing an existing weight variable somebody else produced.
- Calculating and reporting effective base size for the total and every reported subgroup.
- Deciding whether to cap extreme weights, and disclosing the trade-off.
- Judging whether a sample is too distorted to weight and should be rejected or topped up.
- Maintaining a consistent weighting scheme across tracker waves.
- Writing the weighting disclosure a report, a client audit or a regulator requires.

**Who uses it.** Quantitative researchers and data managers applying weights; research directors deciding whether a sample can carry the claims a brief expects; client-side teams interrogating a supplier's scheme; anyone who has been handed a weighted file and needs to know what the bases actually mean.

## 3. When to use it

- The achieved sample composition differs from the population or customer base the study describes, on characteristics that are known and measured.
- Quotas were set and not fully met, leaving a skew that will affect total-level figures.
- A weight variable exists in a file and its provenance, design and effect are unknown.
- Reported figures must state an effective base and none has been calculated.
- Weights look extreme and somebody needs to decide whether to cap, re-target or reject.
- A tracker wave needs weighting consistently with previous waves.
- Subgroups will be reported from weighted data and the base conventions must be settled.
- A disclosure of the weighting scheme is required for a methodology section, an audit or a tender response.

## 4. When NOT to use it

- **The sample missed a group entirely.** Weighting multiplies the people you have. It cannot create people you never reached, and up-weighting three respondents to represent a population segment produces a figure driven by three people wearing the authority of a weighted estimate. Where coverage is the problem, the answer is more fieldwork, an explicit statement that the segment is not covered, or a decision not to report it. This is a coverage error, not a composition error, and no weight fixes it.
- **The intention is to make a non-probability sample look representative.** Matching a population profile on demographics is a composition adjustment. It is not evidence that the sample resembles the population on anything else, and per K4 §3.3 a weighted non-probability sample still does not support population estimates or a margin of error. State the sample type, state what the weighting adjusted, and do not let the two be read as the same claim.
- **The imbalance was caused by cases being removed.** If cleaning exclusions have skewed the sample, the first question is whether the exclusions were correct, not how to weight around them. Return to **04.02 Data Cleaning** and confirm the exclusion accounting before adjusting for its consequences.
- **The weights would be so extreme that the effective base collapses.** There is a point at which weighting stops being an adjustment and becomes an assertion that a handful of respondents represent a large group. When the effective base falls to a fraction of the nominal, or when the weight range spans an order of magnitude, the honest output is to report the sample as unable to support total-level estimates, and to recommend a top-up or a restriction of the claims, per K5 §2.7. Weighting a badly broken sample produces figures with the appearance of correction and none of the substance.
- **The target source cannot be defended.** A weight is only as good as the target it matches. Targets taken from an unverifiable source, a source measuring a different population, or a source too old to describe the current one, produce weights that move the data confidently in the wrong direction. Where no defensible target exists, report unweighted and describe the achieved composition, per K4 §2.4 on unverified sources.
- **The problem is item non-response rather than composition.** Weighting adjusts who is in the sample. It does not fill gaps inside a respondent's record. That is **04.03 Missing Data Handling**, and applying a composition weight to a variable with heavy item non-response produces a weighted figure on an unstated and biased base.
- **The question is whether a difference between weighted subgroups is real.** Weighting changes both the estimate and the precision, and significance testing on weighted data requires the effective base rather than the nominal one. That is **05.02 Statistical Testing**; this skill supplies the effective bases it needs and stops there.
- **The client wants weighting applied to reach a target result.** Choosing weighting variables, targets or a scheme because of the effect on a headline is result-shaping, per K4 §4.2. Weighting variables are chosen for their relationship to the measures of interest and for the availability of defensible targets, before the weighted results are seen.

## 5. Required inputs

**Required.** Without these the skill cannot run. If absent, ask. Where the work must proceed, weight nothing, report the achieved composition, and say what would be needed, per K5 §5.

- **The prepared dataset**, post-cleaning, post-missing-data decisions and post-transformation, with its logs.
- **The achieved sample composition** on every candidate weighting variable, and the intended composition or quota plan.
- **The population or universe definition**: exactly who the study is meant to describe. Weighting to the wrong universe is worse than not weighting, because it applies a confident correction toward the wrong thing.
- **The target figures, with their source, date and the population they describe.** A target without a citable source cannot be used, per K4 §2.4. Where targets come from a client's own records, they are labelled as client-supplied per K2 §6.
- **The sample design and type** (probability, quota, non-probability, river, panel-sourced, mixed), which determines what claims the weighted data can support regardless of how good the weighting is.

**Optional, and what each one adds.**

- **The previous waves' weighting specifications and achieved efficiencies:** required for any tracker, since a scheme change between waves creates movement that is not change and cannot be separated from real movement afterwards.
- **Design weights or selection probabilities** (household size, multiple selection routes, disproportionate stratification): these must be applied before any post-stratification adjustment, and omitting them produces a scheme that corrects composition while leaving the design distortion in place.
- **Auxiliary variables correlated with the key measures** (behavioural or attitudinal, not only demographic): allow a scheme that reduces bias on the things being measured rather than only on the things being profiled, which is the difference between a cosmetic weight and a useful one.
- **Historic weight efficiency for comparable studies:** provides a benchmark for judging whether this scheme's efficiency is normal or a warning.
- **The analysis plan and the reporting structure:** identifies which subgroups will be reported, which is what determines whether the scheme must preserve subgroup precision as well as total-level composition.
- **Any client, sector or tender requirement on weighting and disclosure:** determines the required documentation format and sometimes the permitted variables.

## 6. Questions to ask before starting

1. **What population is this study meant to describe, precisely?** Determines the targets and therefore everything else. "Consumers" is not a universe; "adults aged 18 and over in the three metropolitan markets, in the fieldwork period" is. *Default if unanswered:* do not weight; report the achieved composition and state that the universe was undefined.
2. **Does the composition skew actually matter for the measures being reported?** A skew on a variable unrelated to the outcomes changes almost nothing and adds variance for no gain. *Default:* check the relationship between candidate weighting variables and the key measures; where the relationship is negligible, report unweighted and say why.
3. **What is the target source, how old is it, and who published it?** Determines whether the weights are defensible and how they must be disclosed. *Default:* do not weight to an unciteable target; where the only available target is client-supplied, label it as such per K2 §6.
4. **Which variables should the scheme use, and were they chosen before the weighted results were seen?** Determines whether the scheme is a design decision or a result decision. *Default:* choose variables on their relationship to the key measures and on target availability, record the choice with its date, and report any post-hoc addition separately.
5. **What subgroups will be reported, and what will their effective bases be?** Surfaces the collision between weighting and reportability before the scheme is built. *Default:* calculate effective base for every planned subgroup and flag any that fall below the K4 §7 thresholds after weighting.
6. **Is this a tracker, and what scheme did previous waves use?** A scheme change is indistinguishable from real movement. *Default:* replicate the previous scheme exactly; if a change is necessary, run the wave both ways, publish the difference, and version the scheme.
7. **Will results be reported unweighted anywhere?** Determines the base conventions and prevents the common accident of a weighted headline and an unweighted subgroup in one deck. *Default:* weighted percentages with unweighted counts and effective bases shown throughout, and every table labelled.

## 7. Step-by-step methodology

**1. Decide whether to weight at all, and record the decision either way.** Weighting is not automatic. It is justified when the achieved composition differs materially from the universe on variables that are related to the measures being reported, and when defensible targets exist. It is not justified when the skew is small, when the skew is on a variable unrelated to the outcomes, when no defensible target exists, or when the resulting precision loss exceeds the bias reduction. Test the third condition directly: cross the candidate weighting variables against the key measures. If a variable does not differentiate the outcome, weighting on it moves the headline negligibly and costs precision for nothing. *Correct result:* a written decision, with the composition comparison and the variable-to-measure relationships behind it, whichever way the decision goes.

**2. Define the universe before choosing the targets.** Write it as a sentence with four elements: who, where, what period, and on what basis they qualify. Then check that the achieved sample is drawn from that universe, because a sample of a firm's customers cannot be weighted to a general population however good the demographic targets are: it would produce a general-population-shaped estimate of a customer-only phenomenon. *Correct result:* a universe statement that a reader could use to say whether a given person is in scope, and a confirmation that the sample and the targets describe the same universe.

**3. Select and document the targets with their authority.** For each weighting variable record the target distribution, the source, its publication date, the population it describes, and its status (official statistic, client record, industry estimate, prior study). Then check three things. Currency: a target several years old may no longer describe the population, and demographic drift matters most on the fastest-moving variables. Definitional match: the target's category boundaries must match the questionnaire's, and an age band that differs by a year at the boundary will misweight everyone near it. Population match: a national target applied to a regional sample corrects toward the wrong thing with full confidence. Where a target is client-supplied, it is labelled as such and not vouched for as though it were yours, per K2 §6. *Correct result:* a target table where every row carries its source and date, and every category boundary has been checked against the questionnaire.

**4. Choose between cell and rim approaches on the structure of the problem, not by habit.** Cell weighting (weighting to the joint distribution of the variables together, so that each combination of age, gender and region hits its own target) preserves the interlocking structure and is the stronger correction where the variables interact. Its cost is cells: with three variables of four, two and five categories there are forty cells, and a sample of 1,000 gives an average of 25 per cell with the smallest far below that, producing extreme weights from thin cells. Rim weighting (iteratively adjusting to each variable's marginal distribution in turn until all converge) needs only the marginals, tolerates more variables, and produces more moderate weights. Its cost is that the joint distribution is not controlled: every margin can be correct while an interaction is wrong. Choose cell weighting where the interlock matters and the cells are populated; choose rim where there are several variables or thin cells; and where rim is used, check the resulting joint distribution on the two or three interactions that matter most rather than assuming they came out right. *Correct result:* an approach chosen with its reason stated, and for rim schemes, a convergence check plus an inspection of the key joint distributions.

**5. Compute the weights, then interrogate them before applying them.** Produce the full weight distribution: minimum, maximum, mean (which should be 1 if the weights are scaled to the sample size, or the population size if grossing), standard deviation, and the ratio of maximum to minimum. Then look at who carries the extreme weights. A large weight means one respondent standing for many people, and the identity of those respondents is informative: extreme weights concentrated in one demographic cell mean that cell was badly under-recruited, which is a fieldwork finding. Check how many respondents carry a weight above 3 and what proportion of the weighted total they account for, because a small group carrying a large share of the weighted estimate is a fragile estimate whatever the nominal base says. *Correct result:* a weight distribution table, an identification of who holds the extremes and why, and a statement of the weighted share carried by the most heavily weighted cases.

**6. Calculate weight efficiency and effective base size, and report the effective base everywhere.** Weighting increases the variance of estimates, so a weighted sample of 1,000 does not carry the precision of 1,000 interviews. Effective base size expresses what it does carry: it is derived from the variability of the weights, falling as the weights spread. Weight efficiency is the effective base as a proportion of the nominal. **The effective base, not the nominal n, is the base that governs precision, small-base thresholds and every downstream claim.** Calculate it for the total and separately for every subgroup that will be reported, because efficiency varies by subgroup and is usually worst exactly where the recruitment was weakest, which is usually the subgroup somebody most wants to report. As a rough calibration: efficiency above roughly 90% is unremarkable; 70 to 90% is normal for a study correcting a real skew; 50 to 70% indicates a substantial correction that should be visible in the disclosure and prompts a look at the recruitment; below 50% means half the sample's statistical value has been spent on the correction and the sample should be questioned rather than weighted. Treat those bands as a prompt for judgement, not a standard. *Correct result:* an effective base and an efficiency figure for the total and for every reported subgroup, carried into every table.

**7. Decide on capping deliberately, and disclose the trade-off.** Capping limits the maximum weight, which reduces variance and raises the effective base, at the cost of leaving the sample not fully matched to its targets. That is the whole trade: **capping buys precision with bias.** Neither uncapped nor capped is automatically right. Decide by comparing the two: the effective base with and without, the residual deviation from target with and without, and the movement in the key measures between the two. Set the cap before seeing the effect on the headline where possible, choose a cap level with a stated rationale rather than a convenient one, and report which was used. Where a cap materially changes a headline figure, report both. And note what a cap conceals: capping does not remove the reason the weights were extreme, it hides it, so a capped scheme should always be accompanied by the uncapped diagnostics. *Correct result:* a capping decision with a with-and-without comparison on effective base, target deviation and key measures, plus the retained uncapped diagnostics.

**8. Judge whether the sample should be rejected rather than weighted.** This step exists because it is routinely skipped. The signals that a sample should not be weighted: an effective base far below the nominal; a weight range spanning an order of magnitude; a demographic cell filled by a handful of respondents who are being asked to represent a large population group; a target the sample cannot reach even with extreme weights; and a subgroup whose effective base falls below reportable thresholds after weighting despite an adequate nominal base. Where these appear, the options are to top up the sample in the deficient cells, to restrict the claims to what the sample supports, to report unweighted with the composition disclosed, or to decline to report the affected estimates. **A heavily weighted sample is evidence of a recruitment problem, not a solution to one**, and the diagnosis belongs in the report rather than in a private technical note. Mark it as a decision for the researcher, per K5 §2.7, because it turns on cost, timing and client tolerance that this skill does not know. *Correct result:* an explicit reject-or-weight judgement with its evidence, marked as a decision point wherever the honest answer is "this sample should not be carrying these claims".

**9. Fix the base reporting conventions and apply them everywhere.** Four rules, applied without exception. Percentages are weighted where the study is weighted, and every table says so. Counts shown alongside are unweighted, because an unweighted count is a number of real interviews and a weighted count is not: **a weighted count is never presented as a number of interviews**, and where a grossed-up weighted figure is genuinely required for a volume estimate, it is labelled as a population estimate and never as a base. Effective base is shown wherever a small-base rule or a precision claim is in play. And weighted and unweighted figures never appear in the same table without labels, nor in the same deck without a stated convention. *Correct result:* a convention statement plus a table template that makes the conventions structural rather than dependent on the analyst remembering them.

**10. Handle subgroup analysis on weighted data explicitly.** Three problems recur. First, weights designed to correct the total do not necessarily correct a subgroup, and a subgroup can be internally skewed while the total is perfectly balanced; check the composition of each reported subgroup, not only the total. Second, effective base within a subgroup can be far below its nominal base, so a subgroup reported as n=200 may carry the precision of 90; apply small-base thresholds to the effective base. Third, where subgroups are the primary analytical unit, consider whether the scheme should be built to preserve subgroup precision even at some cost to total-level fit, which is a legitimate design choice that must be stated. Comparisons between weighted subgroups require the effective base for testing, which is handed to **05.02 Statistical Testing**. *Correct result:* a subgroup table showing nominal base, effective base, efficiency and composition check for every reported subgroup.

**11. Maintain consistency across tracker waves, and version the scheme.** The scheme is a specification: variables, targets, target source and vintage, approach, cap, and the rule for updating targets when a new source is published. Apply the same specification every wave. Where a change is unavoidable (a new census, a changed universe, a scheme that has become indefensible), make it at a declared break, run the affected wave under both schemes, publish the difference, and record the version in the data. A weighting change coinciding with a real movement is unrecoverable afterwards, because the two effects cannot be separated. *Correct result:* a versioned scheme specification, the version recorded per wave in the data, and a documented dual-run at every change point.

**12. Write the weighting disclosure, and attach it to the figures.** Per K4 §7, weighted data states its scheme and shows its effective base. The disclosure covers: whether the data is weighted; the variables and targets used; the source, date and status of each target; the approach; whether weights were capped and at what level; the weight range and mean; effective base and efficiency for the total and each reported subgroup; the sample type and what that permits (including, for non-probability samples, that no margin of error is reported); and an explicit statement of what the weighting does not correct, namely coverage and non-response bias on unmeasured characteristics. *Correct result:* a disclosure block a report writer can lift verbatim, which makes the limits of the adjustment as visible as the adjustment.

## 8. Analytical framework

The weighting record uses the category's shared log structure, with the diagnostic step that weighting specifically requires:

    Original composition issue → Target and its authority → Scheme → Diagnostics → Impact → Residual limitation → Disclosure

**Original composition issue.** Achieved versus intended versus universe, variable by variable, with the size of each deviation.
**Target and its authority.** The figures, the source, the date, the population described, and the status of the source.
**Scheme.** Variables, approach, cap, and the rationale for each choice, with the date the choices were made.
**Diagnostics.** Weight range, mean, distribution, who holds the extremes, efficiency, effective base for total and subgroups.
**Impact.** What the weighting changes in the key measures, reported as weighted against unweighted for every headline.
**Residual limitation.** What remains uncorrected: coverage gaps, non-response on unmeasured characteristics, sample type constraints. This field is never empty, because there is always something weighting did not fix.
**Disclosure.** What travels with every reported figure.

Against the K2 chain, weighting sits between Evidence and Analysis and changes the value of every finding downstream. Its most dangerous property is that it improves the appearance of a sample without necessarily improving its validity, so the residual-limitation field is the one that protects the reader.

## 9. Output format

**1. Weighting disclosure block.** Weighted or not; variables and targets; source, date and status of each target; approach; capping and level; weight range and mean; effective base and efficiency, total and by subgroup; sample type and what it permits; and what the weighting does not correct.

**2. Composition table.**

| Variable | Category | Achieved n | Achieved % | Target % | Source and date | Deviation | Post-weight % |
|---|---|---|---|---|---|---|---|

**3. Weight diagnostics.**

| Metric | Value |
|---|---|
| Minimum weight, maximum weight, ratio | |
| Mean, standard deviation | |
| Respondents with weight above 3, and their share of the weighted total | |
| Cells carrying the extreme weights, and why | |
| Weight efficiency | |
| Effective base, total | |

**4. Effective base by subgroup.**

| Subgroup | Nominal n | Effective base | Efficiency | Composition check | Reportable under K4 §7 |
|---|---|---|---|---|---|

**5. Capping comparison**, where capping was considered: effective base, residual target deviation and key measure values, with and without.

**6. Weighted versus unweighted headline comparison.** Every headline measure, both figures, both bases. This is the equivalent of the sensitivity check in 04.02 and is published rather than held.

**7. Reject-or-weight judgement.** The evidence, the decision, and where the decision is the researcher's, the K5 marker with what turns on it.

**8. Scheme specification and version**, for trackers, with the wave range each version covers.

**9. What the weighting does not fix**, stated explicitly, per study.

**Where the evidence is thin**, the format does not get filled. Where no defensible target exists, the output is an unweighted report plus a composition description, not a weighting scheme built on an unciteable target. Where the effective base for a subgroup falls below reportable thresholds, the subgroup is reported at its effective base with the K4 §7 treatment, not at its nominal base. Where the sample cannot support the estimates requested, the output says so, per K4 §1, rather than producing weighted figures that look like an answer.

## 10. Quality checks

Run before any weighted figure is presented. These sit on top of K4 §8.

1. Is the universe defined precisely enough that a reader could say who is in scope?
2. Does every target carry its source, its date and the population it describes, and is any client-supplied target labelled as such?
3. Do the target category boundaries match the questionnaire's exactly?
4. Were the weighting variables chosen before the weighted results were seen, and is any post-hoc addition reported separately?
5. Has the relationship between each weighting variable and the key measures been checked, so that no variable is being weighted on for no analytical gain?
6. Have design weights been applied before post-stratification, where the design requires them?
7. Is the full weight distribution reported, including the maximum-to-minimum ratio and the share of the weighted total carried by the most heavily weighted cases?
8. Is effective base calculated and reported for the total and for every reported subgroup?
9. Are small-base thresholds applied to the effective base rather than the nominal n?
10. Is any weighted count presented anywhere as a number of interviews?
11. Are weighted and unweighted figures ever in the same table or deck without labels and a stated convention?
12. Was the capping decision made with a with-and-without comparison, and are the uncapped diagnostics retained?
13. Has the reject-or-weight judgement been made explicitly, rather than skipped by proceeding?
14. For a tracker, does the scheme match the previous wave, and is any change dual-run and versioned?
15. Does the disclosure state what the weighting does not correct, and does it travel with the figures?

## 11. Common failure modes

| Failure | How to recognise it | How to prevent it |
|---|---|---|
| **Weighting presented as representativeness** | A non-probability sample described as representative because it was weighted | State sample type and what weighting adjusted, separately; no margin of error on a non-probability sample |
| **Nominal base reported as the base** | A weighted subgroup at n=200 with no effective base shown | Effective base calculated for every subgroup and used for every threshold |
| **Weighted count as interviews** | A table reporting a weighted n as though it were people interviewed | Unweighted counts alongside weighted percentages, always; grossed figures labelled as population estimates |
| **Coverage error weighted away** | A segment represented by a handful of respondents carrying large weights | Identify who holds the extreme weights; coverage gaps are reported, not weighted |
| **Target with no provenance** | A target distribution with no source or date | Every target carries source, date and population; no citable source means no weighting |
| **Definitional mismatch** | Target age bands that differ from the questionnaire's by a year at a boundary | Check category boundaries against the questionnaire before building the scheme |
| **Rim scheme with wrong interactions** | Every margin correct, a key joint distribution wrong | Inspect the two or three interactions that matter after convergence |
| **Capping as a fix** | Weights capped, effective base improved, and the recruitment problem no longer visible | Retain uncapped diagnostics; disclose the trade-off; report both where the headline moves |
| **Scheme drift across waves** | A trend movement coinciding with a weighting change | Versioned specification, dual-run at any change, version recorded in the data |
| **Post-hoc variable selection** | A weighting variable added after the results were seen | Record the variable choice with its date; report post-hoc additions separately |
| **Subgroup skew inside a balanced total** | A perfectly weighted total with a badly composed subgroup | Composition check per reported subgroup, not only for the total |
| **AI: weighting because a weight variable exists** | Weights applied without a decision to weight ever being made | Step 1 is a decision recorded either way; an inherited weight variable is interrogated before use |
| **AI: reporting efficiency without effective base** | A percentage figure for efficiency with no base derived from it | Efficiency and effective base are reported together; the base is what governs thresholds |
| **AI: silent grossing** | Weighted totals that sum to a population figure presented in a base row | A grossed figure is labelled a population estimate and never appears as a base |

## 12. AI guardrails

Skill-specific only. K4 applies in full and is not repeated here.

1. **Never describe weighted data as representative on the strength of the weighting.** Say what the weighting adjusted and what it did not, and state the sample type separately.
2. **Never report a weighted figure without its effective base**, and never apply a small-base threshold to a nominal n on weighted data.
3. **Never present a weighted count as a number of interviews.** Unweighted counts accompany weighted percentages; grossed figures are labelled population estimates.
4. **Never weight to a target whose source, date and population cannot be stated**, and never present a client-supplied target as though it were an independent one, per K2 §6.
5. **Never select or change weighting variables, targets or the cap after seeing the effect on a headline.** Where that has happened, report it and show both.
6. **Never apply an inherited weight variable without interrogating it.** Provenance, range, efficiency and effective base are established before use, or the file is reported unweighted.
7. **Never weight a coverage gap.** A group the sample did not reach is reported as not covered, not up-weighted from a handful of cases.
8. **Never proceed past the reject-or-weight judgement silently.** Where the diagnostics indicate the sample should not carry the claims, say so and flag it per K5 §2.7.
9. **Never omit the residual-limitation statement.** Every weighting disclosure says what remains uncorrected, because something always does.
10. **Never change a tracker's weighting scheme without a dual run and a published difference.**
11. **Never let weighted and unweighted figures appear together without labels and a stated convention.**

## 13. Best-practice principles

- **Weighting corrects composition on measured variables and nothing else.** Every other claim made for it, in either direction, is an error. Say what it did, and say what it did not.
- **A heavily weighted sample is evidence of a recruitment problem.** The efficiency figure is a fieldwork diagnostic before it is a technical statistic, and it belongs in the report rather than in a private note.
- **The effective base is the base.** Once weights vary, the nominal n overstates precision, and every threshold, every small-base rule and every test should run on the effective base instead.
- **Weight on variables related to the outcome, not on variables that are merely available.** Weighting on a demographic unrelated to the measures adds variance and moves nothing, which is the worst of both.
- **Targets are evidence and need provenance.** A weight is an assertion that the population looks like the target. If the target cannot be cited, dated and matched to the universe, the assertion cannot be defended.
- **Capping trades bias for precision, and both sides of the trade must be shown.** A capped scheme that reports only its improved effective base has hidden half of what it did.
- **The joint distribution is where rim weighting fails.** All margins correct is not all correct, and the interactions that matter should be inspected rather than assumed.
- **Subgroups need their own diagnostics.** A perfectly weighted total can contain a badly composed and thinly effective subgroup, and that subgroup is usually the one somebody wants to talk about.
- **Consistency across waves outranks improvement within one.** A better scheme introduced mid-tracker produces a movement nobody can interpret. Change at a declared break, dual-run, and publish the difference.
- **Weighting improves the appearance of a sample faster than it improves its validity.** That gap is precisely where its misuse lives, and the residual-limitation statement is what closes it.
- **The decision not to weight is a decision.** Record it, with the composition comparison behind it, so that the absence of weights is visible as a judgement rather than as an oversight.

## 14. Worked example

**INPUT.** A fictional not-for-profit housing body surveys tenants across a national portfolio to inform a service investment decision. Achieved n=1,150, quota-controlled on region and property type, no quota on tenant age. Objectives concern satisfaction with repairs, understanding of tenancy rights, and demand for a digital service channel. The organisation's own tenancy records provide the universe.

**PROCESS.**

*Step 1.* Composition comparison against tenancy records shows region and property type close to target (within 2 points), and age badly skewed: tenants aged 65 and over are 31% of the tenancy base and 14% of the sample. Cross-checking age against the key measures shows a large relationship with digital channel demand (a 34-point gap between youngest and oldest bands) and a small one with repairs satisfaction (4 points). Weighting is therefore justified for the digital channel measure in particular. Decision to weight, recorded with the evidence.

*Steps 2 to 3.* Universe defined as all current tenants of the portfolio at the fieldwork date. Targets taken from the organisation's own tenancy records, dated two months before fieldwork, labelled client-supplied per K2 §6 rather than presented as an independent statistic. Age band boundaries in the records are 16 to 34, 35 to 64, 65 plus; the questionnaire used 18 to 34, 35 to 54, 55 to 64, 65 plus. The 16 to 17 mismatch is immaterial (tenancies are adult), the 35 to 64 split is collapsible, and the mapping is documented.

*Step 4.* Three variables, with 4, 3 and 4 categories, giving 48 cells for a cell scheme on 1,150 cases. Several cells would hold fewer than 10 respondents, and the oldest-age by one-property-type cell holds 4. Rim weighting chosen, with the reason recorded. After convergence, the age-by-region joint distribution is inspected: acceptable. The age-by-property-type joint distribution remains skewed, and that is stated rather than assumed away.

*Steps 5 to 6.* Weight range 0.41 to 6.8, ratio 16.6. Mean 1.00. Nineteen respondents carry a weight above 3, and together account for 8% of the weighted total. All nineteen are aged 65 plus in two property types, which is precisely the under-recruited cell. Weight efficiency 71%, effective base 817 against a nominal 1,150. **The efficiency figure is reported as a fieldwork finding**: the age quota that was not set is the cause, and the recommendation for the next wave is to set one.

*Step 7, and the judgement call.* Capping at 4 is considered. With the cap, efficiency rises to 82% and effective base to 943, and the residual deviation from the age target rises from 0 to 3.4 points among the 65 plus group. The digital channel demand figure moves from 38% uncapped to 40% capped, against 47% unweighted. **Resolution:** the cap is not applied, because the age variable is the one most strongly related to the primary measure and capping reintroduces exactly the bias the weighting exists to remove. The comparison is published so the reader can see the trade rather than being told the outcome, and the uncapped diagnostics stay in the disclosure.

*Step 8.* Reject-or-weight judgement. Efficiency of 71% is a substantial but not disqualifying correction; the total-level estimates are supportable. However, the oldest-age by one-property-type cell holds 4 respondents carrying an effective weight equivalent to a much larger group. That cell will not be reported at all, and the disclosure says so rather than allowing a reader to assume it was covered.

*Step 10.* Subgroup effective bases calculated. The 65 plus subgroup has a nominal base of 161 and an effective base of 94, because that subgroup carries most of the weight variance. Under K4 §7 it is reported with a small-base flag on the effective base, not the nominal. A draft that flagged it as n=161 and unflagged is corrected. Two other subgroups are unaffected.

*Step 12.* Disclosure written: weighted on age, region and property type by rim to client tenancy records dated two months pre-field; no capping, with the capping comparison published; weight range 0.41 to 6.8; efficiency 71%; effective base 817 total; subgroup effective bases tabled; sample is a quota sample of a defined tenant population, so no margin of error is reported; and an explicit statement that the weighting does not correct for tenants unreachable by the contact method used, whose profile is unknown.

**OUTPUT.** A weighted file; a composition table with client-supplied targets labelled and dated; full weight diagnostics identifying who carries the extremes and why; efficiency of 71% reported as a recruitment finding with a quota recommendation for the next wave; a published capping comparison with the decision and its reason; one cell excluded from reporting with the exclusion disclosed; subgroup effective bases governing every small-base flag; a weighted-versus-unweighted comparison on all headline measures; and a residual-limitation statement naming the coverage gap the weighting cannot touch.

## 15. Advanced usage

**Interrogating an inherited weight.** Where a file arrives weighted and undocumented, reconstruct what can be reconstructed: the weight distribution, the efficiency, the effective base, and, by comparing weighted and unweighted composition, which variables the scheme appears to have used and what targets it appears to have hit. What cannot be recovered is the target source and its vintage, which is exactly what determines whether the weights are defensible. Report the reconstruction, mark the provenance as unknown, and treat the weighted figures as unverified until the scheme is supplied, per K4 §2.5.

**Weighting on behavioural rather than demographic variables.** Demographic weighting is conventional and frequently weak, because demographics may be poorly related to the measures. Where reliable behavioural targets exist (product holding, usage frequency, tenure, channel use from client records), weighting on them usually reduces bias on the outcomes far more than a demographic scheme does. The constraint is target quality, and behavioural targets from client records need the same provenance treatment as any other. This is also where the trade-off between total-level fit and subgroup precision usually has to be made explicitly.

**Longitudinal and panel weighting.** Where the same respondents are re-interviewed, attrition creates a second composition problem on top of the original one, and the two need separating: a wave weight corrects the wave's composition, while an attrition adjustment corrects for who dropped out. Applying one and describing it as the other overstates what has been fixed. Coordinate with **05.05 Trend and Tracker Analysis** on which weight applies to which comparison.

**Multi-country studies.** Weight within market to market targets first, then apply a between-market weight if a global total is reported, and keep the two weights as separate variables. A single combined weight makes it impossible to report a market on its own basis, and market-level reporting is usually what clients actually use. State clearly whether a global figure is population-weighted (each market contributing in proportion to its population) or equally weighted (each market contributing equally), because the two produce different global figures from identical data and the choice is substantive.

**When the standard approach does not fit.** Where targets exist for a variable the questionnaire did not capture, the variable cannot be used, and the honest response is to note the uncorrected dimension in the residual-limitation statement. Where a sample is small enough that any weighting produces unstable weights, report unweighted with the composition disclosed. Where the client requires weighting that the data cannot defensibly support, explain what would be produced and what it would mean, offer the strongest honest alternative, and record the decision, per K4 §9.

## 16. Skill chain

**Recommended previous skills:**
- **04.02 Data Cleaning.** Hands over the post-exclusion achieved sample, which is the composition the weighting is calculated against, plus the exclusion accounting that explains any skew the cleaning introduced.
- **04.03 Missing Data Handling.** Hands over the diagnosis of which non-response relates to observed characteristics, which is the only kind of non-response weighting can address, and the base conventions weighted figures must respect.
- **04.04 Data Transformation and Dataset Preparation.** Hands over the subgroup structure and the derived variables the scheme's targets and diagnostics are defined against.
- **01.06 Sampling Strategy.** Hands over the design, the intended composition and any design weights that must be applied before post-stratification.

**Recommended next skills:**
- **05.01 Descriptive Analysis.** Takes the weight variable, the effective bases and the disclosure block, which every reported figure must carry.
- **05.02 Statistical Testing.** Takes the effective bases, without which any test on weighted data overstates its own precision.
- **05.05 Trend and Tracker Analysis.** Takes the versioned scheme specification and any documented scheme break, which constrains what movements can be interpreted.
- **12.06 Research Report QA**, which checks that the weighting disclosure actually travelled into the deliverable rather than stopping at the analysis file.

**Runs well alongside:**
- **K4 §7**, which sets the disclosure obligations this skill operationalises, and **K3 §6**, which distinguishes sampling, measurement and coverage uncertainty, none of which weighting removes.
- **03.01 Recruitment and Sample Sourcing**, where a poor efficiency figure is a recruitment finding that should feed the next study's design.
- **13.01 Research Quality Review**, where an extreme weighting scheme raises a question about whether the study can carry its claims.

---
A Yazi Supplied Skill and resource.
