---
name: bias-detection
description: >
  Audits a research project for bias across its whole span: the question itself,
  sampling and coverage, non-response, instrument, interviewer and mode effects,
  analysis choices, interpretation, reporting and chart framing, cultural and
  linguistic bias in multi-market work, and the bias created by who commissioned
  the study. Use when someone says "is this research biased", "check this for
  bias", "the findings look too convenient", "did we lead the respondents",
  "why does this chart look like that", "our sample skews", "the client wanted
  this answer", or "audit this study for confirmation bias".
category: 13 Research Quality, Ethics and Governance
ref: "13.04"
tier: 1
inherits: [K2, K3, K4, K5]
---

# Bias Detection

## 1. One-line description
Finds the systematic distortions in a research project, from the framing of the question to the framing of the final chart, and returns a bias register naming each mechanism, its likely direction of effect, and whether it can be corrected, must be disclosed, or invalidates the claim.

## 2. What this skill is used for

**The research problem it solves.** Bias is not error. Error is random and it averages out; bias is systematic and it accumulates, always in the same direction. A study can be executed impeccably at every stage and still deliver a confidently wrong answer, because the question presupposed its conclusion, or the frame excluded the people who disagreed, or the subgroup that worked was the subgroup that got reported.

What makes bias hard to find is that each individual decision is defensible. The frame was the best available. The question was worded the way the client's team words it. The subgroup was reported because it was interesting. The chart's axis was scaled to make the movement visible. The quotes chosen were the articulate ones. None of these is misconduct, and any of them defended alone will survive the defence. **The audit only works when the decisions are assessed together and by direction.** Six defensible choices that all push the same way are not six small compromises; they are one large bias with six components, and the project team will be the last people able to see it, because they made each choice for a reason they still believe.

The hardest part is not technical. It is that the largest single source of directional pressure in commercial research is the fact that somebody paid for the study and would prefer a particular answer. **This must be named in the register rather than politely omitted.** Leaving it out does not make it inoperative; it makes it undocumented, and an undocumented pressure is one nobody can compensate for.

**Where it sits.** Cross-cutting, at any stage. It runs against instruments before fielding, against the analysis before reporting, and against the report before it ships, and it is most valuable early, because question-level and sampling bias cannot be repaired downstream at any price.

**Typical use cases.**
- Auditing a whole project before delivery, particularly where the findings are convenient to somebody.
- Checking an analysis where subgroup results are being reported selectively and nobody has counted what else was run.
- Reviewing a report's charts, quotes and emphasis for framing that the numbers do not support.
- Assessing non-response and coverage in a study whose sample fell short or filled unevenly.
- Multi-market work where a construct, a scale or a response style does not travel between markets.
- Qualitative work where the researcher's own position shapes what is heard and what is written down.
- Establishing, at proposal or design stage, what the commercial and organisational pressures on this study are, and recording them before they can operate invisibly.

**Who uses it.** Research directors and quality leads auditing their own or their team's work; client-side insight managers assessing supplier output; methodologists; anyone who has noticed that a study's findings agree suspiciously well with what the person who commissioned it said at the kick-off.

## 3. When to use it

- The findings are convenient to the commissioner, the team, or a decision already taken.
- Subgroup differences are being reported and it is not clear how many were examined.
- A sample fell short, filled unevenly, or came from a frame that does not obviously cover the population.
- Response rates were low, unknown, or differed sharply between groups.
- A report's charts, quote selection or emphasis are doing rhetorical work the underlying numbers may not support.
- Work spans markets, languages or cultures, in which case bias is present by default and the question is only how much and where.
- Qualitative interpretation is carrying material weight and one researcher did most of it.
- A hypothesis existed before the study and the study appears to have confirmed it.
- Any point at which someone says the findings are "exactly what we expected", which is a prompt to check rather than a reassurance.

## 4. When NOT to use it

- **The task is auditing question wording before fielding.** Leading questions, loaded terms, double-barrelled items, unbalanced scales, assumption-carrying phrasing, order and context effects within an instrument: that is **02.04 Question Bias Detection**, which does it at the level of detail an instrument needs. **02.04 audits the instrument before it is fielded. This skill covers the whole project and runs at any stage**, including after fielding, when the instrument can no longer be changed and the question becomes which findings its wording compromised. Where this audit reaches the instrument, it hands the item-level work to 02.04 and keeps the project-level judgement: whether a flagged item is carrying a headline.
- **The concern is methodological soundness generally.** Whether the design could answer the question, whether the sample supports the claims, whether the analysis was executed correctly: **13.01 Research Quality Review**. A quality review will find gross bias as a by-product. It is not built to find the systematic, directional kind, and its severity framework asks a different question.
- **The concern is whether an AI output represents its inputs faithfully.** **13.03 AI Output Verification**. There is a real overlap (dropped contradictions appear in both) and the taxonomies differ: 13.03 asks whether the claim is in the data, this skill asks whether the distortion has a direction.
- **The concern is whether sources are real or say what they are cited for.** **13.02 Source and Citation Verification**. Selection bias in which sources were used is in scope here; whether the sources exist is not.
- **The concern is fairness to participants rather than accuracy of findings.** Whether a group is exposed by the reporting, whether consent covered this use, whether the research should be conducted at all: **13.05 Research Ethics and Consent Design**. These sit next to each other and are not the same. A study can be unbiased and unethical, and biased and permissible.
- **The audit is being commissioned to discredit findings somebody dislikes.** Where the instruction is to find bias in a study whose conclusion is unwelcome, and no equivalent audit is being run on the studies that agree, the exercise is itself the bias. Say so, per K4 §4.2, and offer to run the same audit on both.
- **Nothing about the study can change and nothing will be disclosed.** A bias register whose findings will be neither corrected nor disclosed is an unread document that creates a record of known distortion. Establish before starting that at least disclosure is available. If it is not, that is a governance problem to raise rather than an audit to perform.

## 5. Required inputs

**Required.**
- **The research question and the brief as agreed**, in original wording. Question-level bias is invisible in a paraphrase, and the paraphrase is usually where it was removed.
- **The instrument as fielded**, with routing, and the sample and quota plan against the achieved sample.
- **The analysis outputs, including what was run and not reported.** This is the input teams most often fail to supply and the one selective-reporting findings depend on entirely. A register built from the report alone cannot assess analysis bias.
- **The report or output under audit**, including its charts, quote selection and structure.
- **The commissioning context:** who commissioned the study, what decision it feeds, what position they held before it, and what answer would be commercially or organisationally convenient. If this is not supplied, ask for it directly. It is not an accusation and treating it as one is itself a tell.

**Optional, and what each one adds.**
- **Response and completion records by group, and fieldwork timing.** Make non-response bias assessable rather than merely acknowledgeable, through late-responder and wave comparisons.
- **The analysis plan agreed before fielding (01.07).** The single strongest defence against post-hoc hypothesis fitting, and the only way to tell a confirmed hypothesis from a fitted one.
- **Interviewer, moderator or field-source identifiers on each case.** Allow interviewer and source effects to be tested rather than assumed absent.
- **Source-language instruments, translations and back-translations.** Necessary for any real assessment of construct and linguistic bias in multi-market work.
- **The researcher's own position statement in qualitative work.** Who conducted and interpreted, their relationship to the topic, the participants and the client.
- **Prior waves or comparable studies.** Establish whether a distortion is new, which usually locates its cause.

## 6. Questions to ask before starting

1. **Who commissioned this, and what answer would be convenient?** Determines where to look hardest and what the register must name. *Default if unanswered:* infer from the brief and the decision, state the inference explicitly as an inference, and ask again.
2. **Was a position, hypothesis or expected finding stated before the study?** Determines whether confirmation can be assessed at all, and separates a confirmed prediction from a fitted one. *Default:* ask. An unrecorded prior belief still operates.
3. **At what stage is this audit, and what can still change?** Determines whether findings are corrections, disclosures or design notes for next time. *Default:* establish explicitly; a register aimed at the wrong stage is unactionable.
4. **What was analysed but not reported?** Determines whether selective reporting can be assessed. *Default:* ask directly for the full analysis output. If it is unavailable, record that analysis bias could not be assessed rather than concluding it is absent.
5. **How many markets, languages and modes are involved?** Determines whether construct, translation, response-style and mode bias are in scope. *Default:* if more than one of any, they are in scope.
6. **Who interpreted the qualitative material, and what is their relationship to the topic and the client?** Determines whether a position statement is needed. *Default:* ask, and note that the absence of a position statement is itself a finding in interpretive work.

## 7. Step-by-step methodology

**Step 1. Record the pressure field before you read the findings.** Write down, first and in writing: who commissioned the study, what decision it feeds, what position the commissioner held beforehand, what answer is commercially convenient for them, what answer is commercially convenient for whoever conducted it, and what happens to the relationship if the answer is unwelcome. Then add the researcher's own prior expectations.

Do this before reading the results, because afterwards the assessment is contaminated: once you know the answer, every pressure appears to have been resisted. This is the step that gets skipped for social reasons rather than methodological ones, and skipping it is why most bias audits find only technical bias. **Naming a pressure is not an allegation.** It is the record that makes compensation possible, and a study whose commercial context is documented is more defensible than one where it went unmentioned. *Correct result:* a short pressure statement, written before the findings are known, that a reader could use to predict where distortion would appear.

**Step 2. Audit the research question itself, because nothing downstream can fix it.** A biased question produces a biased study however well the instrument, the sample and the analysis are executed. Four tests:

- **Presupposition.** Does the question assume the thing it purports to establish? "Why do customers find the new process confusing?" cannot return "they do not". "How much has the campaign improved consideration?" has fixed the direction before measurement.
- **Asymmetry.** Is only one alternative available? A study of the barriers to adoption with no equivalent examination of the reasons non-adoption is rational will produce a barriers list, which will then be read as an explanation.
- **Ownership of framing.** Whose language is the question in, and what does that language make invisible? A question framed in the organisation's internal vocabulary (its product names, its segment labels, its theory of the customer) will produce findings inside that vocabulary and can produce nothing outside it.
- **Foregone conclusion.** Is the decision already taken, with the research supporting rather than informing it? Test by asking what finding would cause the decision to change. If none would, the research is documentation, and the register says so.

**Step 3. Work the collection chain: coverage, self-selection, non-response.** Three distinct mechanisms, routinely conflated.

**Coverage bias** is in the frame: who could never have been sampled. An online frame excludes those without reliable access; a customer list excludes the lapsed and the rejected; a store intercept excludes those who stopped visiting. State who is missing and what is likely to be systematically different about them, in the direction it matters for this study.

**Self-selection bias** is in who agreed: people with strong views, more time, an existing relationship, or an interest in the outcome. Incentive structure shapes this too, and a large incentive changes who participates as well as how many.

**Non-response bias** is the one that actually needs testing, because its size depends on whether non-response relates to the answer. A low response rate with non-response unrelated to the outcome costs precision; a high response rate with non-response related to the outcome produces bias. Three practical tests, in order of strength: compare the achieved profile against known population figures on variables correlated with the outcome; compare late responders against early ones, since late responders resemble non-responders more closely; compare the responder profile against the sample frame's own records where the frame holds data. Where none is possible, non-response bias is unassessed and the register says unassessed rather than absent. **And check quota-filling dynamics:** a cell filled in the final two days, often by a different route, produces a subgroup that is not comparable with the others and it will be compared with them anyway.

**Step 4. Reach the instrument, then hand it over and keep the project judgement.** Item-level wording work belongs to **02.04**: leading phrasing, loaded terms, double-barrelling, unbalanced scales, assumption-carrying stems, order and context effects. Run it, or check that it was run. What this audit keeps is the project-level question: **which findings rest on flagged items, and do any of them carry a headline?** A leading question on a peripheral measure is a note; the same defect under the report's central claim is a material bias with a known direction. Also check what the instrument could not capture: a fixed list with no meaningful other-specify constrains the answer to the frame the designer already had, and an unprompted question asked after ten prompted ones is no longer unprompted.

**Step 5. Test the delivery: interviewer, moderator, source and mode.** **Interviewer and moderator effects** are real, measurable and rarely measured: compare results by interviewer or moderator on the sensitive and the attitudinal items, since these are where the effect concentrates. A moderator who believes a hypothesis probes for it, follows it up more, and gets more of it, without any impropriety. **Source effects** appear where sample came from more than one route: compare the routes on the key measures, because differences here are frequently larger than the subgroup differences being reported as findings. **Mode effects** are systematic and directional: self-completion produces more admission of socially undesirable behaviour than interviewer-administered modes, scales are used differently between visual and aural presentation, and telephone and face-to-face produce more acquiescence than self-completion. Where modes or routes are mixed, mode is confounded with whatever else differs between them, and any comparison across them carries that confound.

**Step 6. Audit the analysis, which requires knowing what was run and not reported.** Four mechanisms.

**Selective subgroup reporting.** Count the subgroup comparisons available and the number reported. Reporting the ones that reached a threshold, from a banner generating thousands of comparisons, is a selection procedure that guarantees findings whether or not anything is there (05.02). The register needs the ratio: reported over examined.

**Post-hoc hypothesis fitting.** A pattern found in the data, then explained, then presented as though the explanation preceded it. The tell is an explanation that fits the result exactly and would have fitted the opposite result equally well. Test against the analysis plan; without one, all such findings are exploratory and must be labelled so.

**The flattering comparison.** The same number is favourable or unfavourable depending on what it is set against: a different competitor, a different period, a different segment, a different baseline. Ask why this comparator and not the obvious alternatives, and check whether the alternatives were computed.

**Choices that moved the answer.** Base and filter definitions, exclusions, outlier handling, missing-data treatment, weighting targets, scale collapsing (whether top-two or top-three box was chosen, and whether the choice was made before or after seeing both). Each is defensible individually. The question is whether every choice went the same way, and that is answered by listing them and looking at the direction column, not by evaluating them one at a time.

**Step 7. Audit interpretation, including the researcher's own position.** **Client confirmation bias** is visible as a finding that matches a stated prior belief and receives less scrutiny than findings that do not, and as contrary evidence handled at greater length and dismissed. Test by comparing how the supporting and contradicting evidence are treated: the asymmetry is usually plain once looked for. **Researcher confirmation bias** is the same mechanism inside the team, and it strengthens in proportion to how good the emerging story is.

Two structural checks. **The alternative explanation:** for each major interpretation, write the strongest competing reading and record why it was rejected. Where no competing reading was ever considered, that is the finding. **The disconfirmation search:** ask what evidence in the study would have contradicted this interpretation, then go and look for it rather than waiting to encounter it.

In qualitative work, the researcher's position operates directly on what is heard, what is followed up, what is coded and what is quoted. The mechanisms are specific: shared background with some participants and not others, a relationship with the client that makes certain findings uncomfortable, and prior domain familiarity that makes an unusual account sound like a misunderstanding. The mitigation is a written position statement plus independent coding of a sample, not an assurance of objectivity. **The absence of a position statement in interpretive work is a finding**, and where one researcher coded, interpreted and wrote alone, that is recorded with its direction of likely effect.

**Step 8. Audit the reporting layer, where bias enters last and unnoticed.** By this stage the analysis may be sound and the document still distorted.

**Quote selection.** Compare the quotes used against the distribution of views in the source. Articulate participants are over-selected, extreme statements are more quotable than typical ones, and the quote that expresses the theme perfectly is often the outlier that expresses it too well. Check whether the counter-position has a quote at all: a theme reported with three supporting quotes and its counter-theme reported in a paraphrase is asymmetric even where both prevalences are stated.

**Chart framing.** The recurring mechanisms: a truncated or non-zero axis making a small movement look large; scales differing between charts that a reader will compare; ordering by convenience rather than by value or by a stable order; colour coding valence so the favoured option is read as good before the numbers are; dual axes manufacturing an apparent relationship; a time window starting at a flattering point; and bases omitted where they would undercut the visual. Each is a decision with a direction, and each is worth checking against the alternative rendering.

**Emphasis.** What is in the executive summary against what is in the appendix; what gets a page against what gets a clause; whether the caveat travels with its claim or is left behind (K3 §7). Emphasis is the most powerful and least visible framing device in any report, because nothing in it is untrue.

**Step 9. In multi-market work, treat non-comparability as present until shown otherwise.** Four mechanisms, all directional. **Construct bias:** the concept does not exist in the same form in every market, so an identical instrument measures different things and produces a comparison that looks valid and is not. **Translation and linguistic bias:** intensity words, politeness and negation do not map cleanly; a back-translation confirms lexical equivalence and not conceptual equivalence, which is the thing that matters. **Response-style bias:** systematic differences in acquiescence, extreme responding and midpoint use between cultures, which contaminate any cross-market ranking and are frequently larger than the differences being reported. Where this is suspected, check whether within-market standardisation was considered and what it does to the ranking. **Sampling non-equivalence:** the same quota specification produces differently composed samples in markets with different underlying distributions, so "urban professionals" is not one population. Market differences that appear on attitudinal batteries and disappear on behavioural measures are the classic signature of response style rather than substance, and the register says so rather than reporting a cultural finding.

**Step 10. Build the register: mechanism, direction, magnitude, disposition.** Every entry names four things, and it is the second that makes the register useful.

**Mechanism:** the specific route by which distortion entered, stated precisely enough to be argued with. Not "sampling bias" but "the frame is the active customer file, excluding the 14% who closed accounts in the period, who are the group most likely to report dissatisfaction".

**Direction:** which way the finding is pushed. Inflates, deflates, or unknown. **Unknown is a legitimate and common entry**, and it is more useful than a guess, because an unknown-direction bias cannot be argued away as conservative.

**Magnitude:** estimated where estimable, from a reanalysis, a profile comparison or a sensitivity check, and otherwise stated as unestimated. Where a bias can be bounded ("recomputing on the alternative base moves the headline from 62% to 57%"), do it, because a bounded bias becomes a manageable disclosure rather than an unbounded doubt.

**Disposition:** one of four. **Correct it**, where a reanalysis, a reweight or a re-cut is available; this is always preferred and is more often available than teams assume. **Disclose it**, where correction is impossible and the finding survives with the disclosure attached, per K4 §4.3, disclosed next to the claim rather than in an appendix. **Restrict the claim** to the population, period or condition the evidence actually supports. **Invalidate**, where the bias is large enough and directional enough that the claim cannot be made at all.

Then do the step that the individual entries cannot do: **sum by direction.** Sort the register by direction and count. Several small biases all pushing the same way are a single large bias, and this is how a study made entirely of defensible decisions arrives at a confidently wrong answer. Where the sum is material, it is stated as its own register entry with its own disposition, and it is usually the entry that matters most.

## 8. Analytical framework

The bias chain, which is the order the audit runs in and the order distortion enters a project:

    Question → Frame → Response → Instrument → Delivery → Analysis
        → Interpretation → Reporting
         ↑                                          ↑
    Unfixable downstream.                Fixable cheaply, and
    Correction here means                the last place anyone
    a different study.                   looks.

And the register entry, applied to every finding:

    Mechanism → Direction (inflates / deflates / unknown)
        → Magnitude (estimated / bounded / unestimated)
            → Disposition (correct / disclose / restrict / invalidate)

**Applying it.** Run the chain forwards, because a bias early in the chain changes what the later stages even mean: there is no point assessing analysis bias on a sample the frame never covered. Then apply the summation rule, which is the part that distinguishes this from a list of caveats. **Individually defensible, collectively directional** is the pattern the whole skill exists to catch, and it can only be seen when every mechanism carries a direction and the directions are counted.

**Direction is the load-bearing column.** A register without directions is a list of things that could theoretically have gone wrong, which is unactionable and reads as hedging. A register with directions supports two things nothing else does: a statement of which way the study's answer is likely to be wrong, and a sensitivity assessment of whether the conclusion survives if the biases are real.

## 9. Output format

**A. Pressure statement.** Written first, per step 1. Commissioner, decision, prior positions, convenient answers on both sides, and the researcher's own expectations. Retained in the output even where nothing was found to have acted on it, because its value is that it was recorded.

**B. The bias register.**

| Ref | Stage | Mechanism (specific) | Evidence for it | Direction | Magnitude | Claims affected | Disposition | Owner |
|---|---|---|---|---|---|---|---|---|

**C. Directional summary.** The register sorted by direction, with the count and the combined assessment. This is the section a reader acts on.

**D. What could not be assessed**, with the reason: analysis outputs not supplied, no response data, no source-language instrument, no position statement. **Unassessed is not absent**, and conflating them is the commonest way a register understates.

**E. Sensitivity check.** For the two or three largest entries, what happens to the headline conclusion if the bias is real at its estimated magnitude. A conclusion that survives its plausible biases is considerably stronger than one that was never tested against them.

**F. Required disclosures**, drafted as the sentences that will appear next to the affected claims rather than as instructions to disclose.

**When the evidence is thin.** Do not manufacture entries: a register padded with theoretical biases nobody has evidence for trains readers to ignore it, and it is the bias-audit equivalent of over-flagging (K5 §4). Where a mechanism is plausible but unevidenced, it goes in section D as unassessed, with what would settle it. Where an audit finds a well-controlled study, it says so and lists what was examined. And where the commissioning pressure is the largest entry in the register and nothing else was found, that entry stands alone and is still the most useful thing in the document.

## 10. Quality checks

Run on the audit. K4 §8 runs anyway.

1. Was the pressure statement written before the findings were read?
2. Does every register entry name a specific mechanism rather than a category of bias?
3. Does every entry carry a direction, including the honest entry of unknown?
4. Was the directional summation performed, and is the combined effect stated as its own entry where material?
5. Was the research question itself audited, or does the register begin at sampling?
6. Was selective reporting assessed against what was actually run, and if the analysis outputs were not supplied, is that recorded as unassessed?
7. Was non-response tested by at least one of the three comparisons, or recorded as unassessed?
8. Were interviewer, source and mode effects tested rather than assumed absent?
9. Were the charts checked against their alternative renderings, particularly axis, scale, ordering and comparator?
10. Was quote selection compared against the distribution of views in the source, including whether the counter-position has a quote?
11. In multi-market work, were construct, translation, response-style and sampling equivalence each addressed separately?
12. Is the commissioning pressure named in the register rather than omitted for comfort?
13. Does every entry carry a disposition, and were corrections preferred over disclosures wherever a reanalysis was available?
14. Is anything in the register present to demonstrate thoroughness rather than because there is evidence for it?

## 11. Common failure modes

| Failure | How to recognise it | How to prevent it |
|---|---|---|
| **Auditing one decision at a time** | Every entry is defended successfully and the register closes with nothing found | Sum by direction. The pattern is collective, not individual |
| **The polite omission** | A register with no entry for the commissioner's preferred answer | Step 1, in writing, before the findings. Naming a pressure is not an allegation |
| **Starting at the sample** | The question was never examined, so its bias is inherited by every finding | Step 2 before step 3. Question-level bias cannot be fixed downstream |
| **Direction left blank** | A list of possible biases with no assessment of which way they push | Direction is mandatory. Unknown is a valid entry; blank is not |
| **Unassessed reported as absent** | "No evidence of non-response bias" where no response data existed | Section D. Say what could not be assessed and why |
| **The register nobody reads** | Twenty entries of equal weight, mostly theoretical | Evidence required per entry. No padding, and a directional summary at the front |
| **Selective reporting assessed from the report** | Reported subgroups all look reasonable; nobody counted what was run | Require the full analysis output, or record analysis bias as unassessed |
| **Response style read as culture** | Market rankings on attitudinal batteries that vanish on behavioural measures | Step 9. Check behavioural measures and consider within-market standardisation |
| **Back-translation as equivalence** | Translation quality confirmed; construct equivalence never examined | Back-translation establishes lexical, not conceptual, equivalence |
| **The reflexivity paragraph** | A position statement that asserts objectivity rather than naming a position | A position statement names relationships and expectations, and pairs with independent coding of a sample |
| **AI: theoretical bias inventory** | A textbook list of every named bias, applied without evidence from this study | Each entry cites evidence from this project at a specific location |
| **AI: symmetric hedging** | Every finding described as possibly biased in both directions, which cancels to nothing | Force a single direction or an explicit unknown, with the reasoning |
| **AI: dropping the commissioner** | A technically thorough register that never mentions who paid or what they wanted | The pressure statement is a required section, not a discretionary one |

## 12. AI guardrails

Universal prohibitions are inherited from K4. K4 §4.2 (never cherry-pick supporting evidence) and its rule that a stakeholder's preferred answer is information about the stakeholder govern this skill directly.

1. **Never omit the commissioning pressure from the register**, whoever the audit is for, and whether or not evidence was found that it operated. Its absence from a register is itself a distortion of the record.
2. **Never enter a bias without evidence from this study at a named location.** A general mechanism that could apply to any project belongs in a methods textbook, not in this register.
3. **Never leave direction blank, and never state both directions to avoid committing.** Symmetric hedging cancels to no information. Where the direction is genuinely unknown, that is the entry, with the reason.
4. **Never treat unassessed as absent.** An untested mechanism goes in the could-not-assess section with what would settle it, and never into a clean conclusion.
5. **Never grade a bias by how uncomfortable it is to raise.** The entries that are awkward (the commissioner's preference, the team's prior belief, the researcher's own position) receive the same treatment as the technical ones.
6. **Never audit only the findings that are inconvenient to somebody.** The audit runs across the whole study, including the findings everyone agrees with, which have received the least scrutiny.
7. **Never propose disclosure where correction is available.** A reweight, a re-cut on the correct base or a reanalysis on the full comparison set fixes the problem; a disclosure only tells the reader it exists.
8. **Never write a disclosure that names a bias without its direction and, where estimable, its magnitude.** "Findings may be affected by sample composition" tells a reader nothing they can use.
9. **Never present a register as complete when inputs were missing.** State what was not supplied and what could not therefore be assessed.

## 13. Best-practice principles

1. **Bias is directional, and that is what distinguishes it from error.** The question is never only whether a decision was defensible, but which way it pushed, and whether the others pushed the same way.
2. **Individually defensible, collectively directional.** This is the pattern. Every serious bias in commercial research is assembled from choices that each survive their own defence.
3. **Name the money.** Who paid, what they wanted, and what happens if the answer is unwelcome. Recording it costs a paragraph and it is the entry most often missing.
4. **Write the pressure statement before you read the findings.** Afterwards you will conclude the pressures were resisted, and you will be unable to tell whether that is true.
5. **Question-level bias is the only unfixable kind.** No sample, instrument or analysis can rescue a question that presupposes its answer, which is why the cheapest bias audit is the one run at design stage.
6. **Ask what finding would have changed the decision.** If the honest answer is none, the study is documentation and everything downstream should be read that way.
7. **Non-response rate is not non-response bias.** What matters is whether non-response relates to the answer, and that is testable more often than teams assume.
8. **Count the comparisons examined, not the ones reported.** A ratio of reported to examined is the single most informative number in an analysis-bias audit.
9. **Check the counter-position's quote.** Whether the minority view is represented by a verbatim or a paraphrase tells you more about a report's framing than the prevalences do.
10. **In multi-market work, assume non-equivalence and prove comparability.** The default assumption in the other direction produces confident cultural findings that are artefacts of response style.
11. **Prefer correction to disclosure, and disclosure to silence.** Most biases teams disclose could have been corrected by a reanalysis nobody attempted.
12. **A register with three evidenced entries beats one with twenty theoretical ones.** Over-flagging destroys the register's authority as surely as under-flagging destroys its purpose.
13. **Audit the findings everyone likes.** They got the least scrutiny on the way in, and they are where the collective direction shows.

## 14. Worked example

*Fictional scenario, used for illustration only. The organisation, markets, figures and quotes below are invented for the purpose of demonstrating method.*

**INPUT.** A financial services group commissions a four-market study on appetite for a new long-term savings product. The internal sponsor has been advocating the product for eighteen months and a build decision is scheduled for the following quarter. The delivered report concludes that appetite is strong and consistent across markets, with Market C notably ahead. The audit is requested by the group's insight director before the decision paper is written.

**PROCESS.**

*Step 1, before reading the findings.* Pressure statement: the sponsor has an eighteen-month public position; the decision is scheduled and a negative finding would be expensive to absorb; the supplier is in a competitive review and an unwelcome answer carries relationship risk; and the insight director requesting the audit was privately sceptical of the product, which is a pressure in the opposite direction and is recorded as such.

*Step 2, the question.* The brief asks: "What is driving appetite for long-term savings products among mass-affluent customers, and what features would maximise take-up?" Both clauses presuppose appetite. No objective anywhere asks whether appetite exists or why customers might rationally decline. Test applied: what finding would have changed the decision? Nothing in the design could have produced one. This is the register's first entry, and it inflates, at the level of the whole study.

*Step 3, coverage and non-response.* The frame is existing customers holding at least one savings product, so those who hold none, the group most likely to lack appetite, could never have been sampled. Response records show 11% completion with no non-response analysis performed. A late-responder comparison is possible from the timing data and is run: late responders score 9 points lower on the appetite measure than early ones, which suggests non-responders sit lower still. Direction: inflates. Magnitude: bounded at roughly 5 to 12 points on the headline, from the late-responder gradient.

*Step 4, instrument.* Routed to 02.04, which flags three items. Two are peripheral. The third is the headline appetite measure, which describes the product with four benefits and no costs, no lock-in period and no alternative use of the money, then asks how interested the respondent would be. The finding carrying the report's central claim rests on this item. Direction: inflates. Disposition: the claim is restricted to stated interest under a benefits-only description, which is a materially different sentence from "appetite is strong".

*Step 5 and 6, delivery and analysis.* Two sample routes were used in Market C and nowhere else. Comparing routes: the second route scores 14 points higher on appetite, which is larger than the cross-market difference the report treats as its Market C finding. Analysis outputs are requested and supplied: 62 subgroup comparisons were examined and 6 reported, all favourable, with no pre-specified comparison set. Base choice: the headline is computed on those expressing any familiarity with long-term savings products, excluding 22% of respondents, a decision made after the first tabulation.

*Step 7, and the judgement call.* The Market C result is the report's most quoted finding and it has two candidate explanations: a genuine market difference, or the second sample route. The team's interpretation, formed before the route difference was noticed, is a cultural one about savings orientation, and it is written persuasively. The audit's judgement is that the route explanation is at least as consistent with the evidence and was never considered, which makes this an alternative-explanation failure rather than a wrong conclusion. It cannot be resolved from the existing data. Disposition: restrict, with both explanations stated, and a note that a single-route re-run in Market C would settle it. Recorded as `RESEARCHER DECISION REQUIRED` per K5 §3.1, since the choice between delaying the decision paper and reporting an unresolved confound is the sponsor's, not the auditor's.

*Steps 8 and 9, reporting and markets.* Chart audit: the appetite chart's axis runs 40 to 80, making a 7-point cross-market spread fill the plot area; redrawn from zero it is visibly small. Quote audit: nine quotes support appetite, none represents the 31% expressing low interest, whose position appears once as a paraphrase. Multi-market: the four markets differ on attitudinal batteries and are near-identical on the behavioural measures (existing savings balances, contribution frequency), which is the classic response-style signature. Within-market standardisation was not considered; applied as a sensitivity check, it removes most of the Market C advantage.

*Step 10, the summation.* Twelve entries. Nine inflate, one deflates (a translation issue in one market that understates a secondary measure), two unknown. **The directional summary is the finding:** the question, the frame, the non-response gradient, the benefits-only stimulus, the post-hoc base exclusion, the selective subgroup reporting and the axis all push the same way. No single entry invalidates the study. Together they mean the headline is an upper bound produced by a chain of choices that each went the same direction, and the sensitivity check shows the central conclusion does not survive their combined plausible magnitude.

**OUTPUT.** A pressure statement written before the findings; a twelve-entry register with mechanism, evidence, direction, magnitude and disposition; a directional summary stating that nine of twelve entries inflate; four corrections available without new fieldwork (recompute on the full base, report all 62 comparisons or none, redraw three charts from zero, apply within-market standardisation to the cross-market claim); three required disclosures drafted as the sentences to appear next to their claims; one restriction (stated interest under a benefits-only description, not appetite); one unresolved confound in Market C with the re-run that would settle it; and a section on what could not be assessed, naming interviewer effects, since no interviewer identifiers were retained.

## 15. Advanced usage

**Auditing at design stage, which is where this skill pays for itself.** Run steps 1, 2, 3 and 4 on a proposal or a draft instrument. At that point the question can be rewritten to admit a negative answer, the frame can be extended to the group it excludes, a cost-and-benefit stimulus can replace a benefits-only one, and a comparison set can be pre-specified. The same four findings after fielding are permanent limitations. A design-stage bias audit takes a fraction of the time and prevents most of what a delivery-stage audit can only document.

**Bias in tracking studies.** Directional bias in a tracker is less damaging to the level and more damaging to the trend, and the mechanisms are different: a frame that ages relative to the population, panel conditioning, a mode migrating across waves, and a question reworded for good reasons in wave 6. Audit the version history rather than the current wave, and assess each change for whether it introduced a step in the series.

**Where the client is the source of the pressure and also the audience for the register.** Write the pressure statement in neutral, factual language: positions held, decisions scheduled, convenient answers, on all sides including the researcher's own. A register that names the researcher's incentives alongside the client's is read as rigour rather than accusation, and it is also more complete.

**When the standard approach does not fit.** Where the bias is structural in the whole programme rather than in this study (the same frame, the same sponsor and the same question shape across five years of research), a per-study register will keep finding the same entries and changing nothing. Escalate it as a programme-level finding to whoever owns the research function, with the series of registers as the evidence. That is the case where a single audit cannot fix what repeated audits have documented.

## 16. Skill chain

**Recommended previous skills:**
- **02.04 Question Bias Detection.** Hands over the item-level instrument audit, which this skill uses at step 4 and extends into the project-level judgement of which findings the flagged items carry.
- **01.06 Sampling Strategy** and **04.05 Weighting and Base Management.** Hand over the frame, quota and weighting decisions the coverage and analysis entries are assessed against.
- **13.03 AI Output Verification.** Hands over dropped contradictions and selective reporting found in AI-produced work, which this skill assesses for direction.

**Recommended next skills:**
- **13.01 Research Quality Review.** Takes the register as an input to severity and to its cumulative assessment.
- **05.03 Cross-Tabulation** and **05.02 Statistical Testing.** Where the disposition is correction, these run the reanalysis on the full comparison set and the correct base.
- **12.03 Research Report Compilation.** Applies restrictions, corrections and drafted disclosures into the report rather than appending them.
- **13.05 Research Ethics and Consent Design**, where an audit finding turns out to concern fairness to participants rather than accuracy of findings.

**Runs well alongside:**
- **10.02 Evidence Synthesis** and **10.03**, where selection bias across sources is the same mechanism operating on a different evidence type.
- **K3 §7**, whose failure modes (the disappearing caveat, confidence by repetition) are reporting-stage bias by another name, and **K5 §2.2**, for the cultural interpretation judgements that multi-market entries generate.

---
A Yazi Supplied Skill and resource.
