---
name: research-quality-review
description: >
  Reviews the methodological integrity of a whole research project, from the
  question through design, sample, instrument, fieldwork, data handling,
  analysis, interpretation and reporting, and returns a severity-graded findings
  report with an overall judgement that can be negative. Use when someone says
  "review this study", "is this research any good", "can we stand behind these
  numbers", "peer review this project", "the client is challenging our method",
  "check this before it goes out", "audit the methodology", "sign off the
  design", or "I inherited this project and I do not trust it".
category: 13 Research Quality, Ethics and Governance
ref: "13.01"
tier: 1
inherits: [K2, K3, K4, K5]
---

# Research Quality Review

## 1. One-line description
Assesses whether a piece of research is capable of supporting the claims made from it, stage by stage across the research lifecycle, and reports what is wrong, how serious it is, what must be done about it, and whether the study is fit for the purpose it is being used for.

## 2. What this skill is used for

**The research problem it solves.** Research fails quietly. A study is designed to answer one question and reported as answering another. A sample is drawn from a frame that never covered a third of the population, and the finding is written up as though it did. A quota fills late with a different kind of respondent and nobody re-derives what the sample can now support. A cleaning decision removes 8% of cases and moves the headline four points, and the decision lives in a colleague's memory rather than a log. None of these announce themselves. Each of them produces a report that reads perfectly well.

What makes this dangerous is that the pressure to accept research runs in one direction. A project has a budget, a timeline, a client who has already told people the answer is coming, and a team who have worked hard. A finding that is convenient will be scrutinised less than one that is not. The purpose of a quality review is to apply the same scrutiny to both, and it is the only step in the research process whose job is explicitly to find fault. **This skill is not here to agree with the researcher. It is here to make the research stronger, which frequently means being unwelcome.** A review that never returns a negative judgement is not performing a review; it is performing reassurance, and a reassurance function has no information value.

**Where it sits.** Cross-cutting. It runs against a whole project rather than a stage of one, and it runs at two natural points: at design sign-off, when almost everything is still changeable and correction is cheap, and before delivery, when almost nothing is changeable except the claims and the disclosures. It also runs on demand, when a project is inherited, challenged, or being reused for a purpose it was not designed for.

**Typical use cases.**
- Signing off a design before fieldwork commits budget to a study that cannot answer its question.
- Reviewing a finished study before it goes to a client, a board, a regulator or publication.
- Auditing an inherited project when the original team has moved on and nobody can vouch for the decisions.
- Responding to a challenge: a stakeholder, a competitor's agency or an internal analyst disputes a number or a method.
- Assessing whether an old study can be reused for a new question, which is a fitness question rather than an analysis question.
- Reviewing a supplier's or partner's work, where the reviewer has the documents but not the working knowledge.
- Peer review inside a research team, as a standing discipline rather than an escalation.

**Who uses it.** Research directors accountable for what leaves the building; quality leads and methodologists; client-side insight managers assessing supplier work; anyone asked to put their name to a study they did not run. It assumes methodological literacy: this skill audits method, and a reviewer who cannot recognise a wrong test cannot be rescued by a checklist.

## 3. When to use it

- A design is about to be signed off and money is about to be committed to fieldwork.
- A report is about to be delivered, published, submitted, or used to support a decision with real consequences.
- A finding is surprising, highly convenient, or contradicts what is already known, and someone should establish whether it is real before it is acted on.
- A project has been inherited mid-flight or after completion and its decisions are not documented.
- A claim has been challenged and the team needs to know whether the challenge is right before responding to it.
- An existing study is being reused to answer a question it was not commissioned for.
- Several methodological compromises were made under time or budget pressure and nobody has assessed their cumulative effect, which is usually larger than the sum of the individual notes.
- A recurring study (a tracker, an annual survey, a rolling programme) has drifted over waves and no one has looked at the whole thing for years.

## 4. When NOT to use it

- **The question is about the document rather than the research.** Internal consistency, numbers matching between the summary and the chart, terminology, base labels on charts, structure, spelling and whether every objective is addressed somewhere in the text are **12.06 Research Report QA**. That skill checks the document. **This skill checks the research.** A report can pass 12.06 completely and still describe a study that cannot support a word of it, and a study can be methodologically sound and produce a badly assembled report. Run both. Where the two overlap (a base size missing from a chart is a 12.06 defect; a base size too small to carry the claim it is under is a 13.01 defect), the test is whether fixing the document fixes the problem.
- **The output under review is AI-produced and the question is whether to trust it.** That is **13.03 AI Output Verification**, whose failure profile is different: fabricated specifics, plausible citations, silent gap-filling and uniform confidence. Run 13.03 first, then bring its verification report into this review as an input. A quality review that treats AI output like human output will look for the wrong errors.
- **The specific concern is bias.** Where the question is whose interest the study serves, whether the question itself was loaded, whether subgroups were selectively reported, or whether the commissioner's preferred answer shaped the work, that is **13.04 Bias Detection**, which runs a different taxonomy and produces a bias register. A quality review will find gross bias; it is not built to find the systematic kind.
- **The specific concern is the sources.** Where the doubt is whether cited material exists, says what it is claimed to say, or traces back to an origin, that is **13.02 Source and Citation Verification**.
- **The question is ethical rather than methodological.** Whether the research should have been done, whether consent covered this use, whether participants are identifiable, whether a finding could be used against the group it came from: **13.05 Research Ethics and Consent Design**. A study can be methodologically immaculate and ethically unacceptable, and this review will not catch it.
- **The required inputs do not exist.** A review needs the artefacts: the brief, the instrument as fielded, the sample records, the data or the analysis outputs, and the report. Reviewing a report alone produces an opinion about a report, not a judgement about a study, and it must be labelled that way. Where only the report is available, say so, restrict the review to what the report itself discloses, and record that the underlying method could not be assessed. Do not infer methodology from a methods paragraph.
- **The purpose is to justify a decision already taken.** Where a review is commissioned to produce a clean bill of health for a study that is going out regardless, or to discredit a study whose findings are unwelcome to the commissioner of the review, the exercise is not a review. Say so, per K4 §4.2, and either conduct a real one or decline. This is worth naming early, because the request usually arrives phrased as urgency rather than as instruction.
- **The study is fine and someone wants a document saying so.** Reviewing costs real time. Where there is no live decision, no challenge, no reuse and no consequence, the honest answer is that a review is not warranted. Over-reviewing devalues the practice in the same way over-flagging devalues a review marker, per K5 §4.

## 5. Required inputs

**Required.** Without these the review cannot run as a review. If absent, ask. If they cannot be produced, restrict the scope explicitly and say in the output what could not be assessed, per K5 §5.

- **The brief and the agreed objectives**, in their final form. Everything downstream is judged against what the study was for. A review with no statement of purpose can assess internal soundness but cannot assess fitness.
- **The design or proposal as agreed**, including the sampling plan, the intended analysis, and any documented deviations.
- **The instrument as fielded**: questionnaire with routing, discussion guide, screener, stimulus. Not the draft. The fielded version is the only one that generated data.
- **Sample records**: frame or source, quotas set, quotas achieved, response or completion rate where knowable, fieldwork dates, and the achieved profile against the intended one.
- **Analysis outputs with their bases**, and the data handling record: cleaning log, exclusions, recodes, missing-data treatment, weighting scheme.
- **The report or the claims being made**, since fitness is always fitness for a specific set of claims.

**Optional, and what each one adds.**

- **The raw dataset.** Converts the review from an inspection of outputs into an audit: a disputed figure can be recomputed at source rather than adjudicated between derived files, and distributional assumptions can be checked rather than taken on trust. It is the single most valuable optional input.
- **The analysis plan agreed before fielding (01.07).** Makes it possible to distinguish confirmatory findings from exploratory ones, which changes the severity of half the statistical findings a review produces.
- **Fieldwork records and interviewer or moderator notes.** Surface the things that never reach an analysis file: a question respondents kept misreading, a market where fieldwork ran late, a quota filled by a different route in the last two days.
- **Previous waves.** Establish the measure's own variation, without which no movement can be judged, and reveal drift in wording, base definition or segment naming across waves.
- **The decision the research informs, and its stakes.** Sets the review intensity honestly. A study supporting a pricing move across a portfolio warrants a different threshold from one supporting a creative choice.
- **Access to the original team.** Answers questions the documents cannot. Use it carefully: an explanation offered in conversation is not a record, and a decision that exists only in someone's memory is a documentation defect whatever the explanation turns out to be.

## 6. Questions to ask before starting

1. **What decision does this research support, and what would a wrong answer cost?** Sets the threshold. Materiality determines how severe a given defect is: an undisclosed base of 68 is a note in an exploratory study and a major defect under a capital allocation. *Default if unanswered:* assume the highest-stakes plausible use, review at that threshold, and say you have.
2. **Is this a design-stage or delivery-stage review?** Determines what the review can still change, and therefore what its findings should be aimed at. *Default:* establish it explicitly. Reviews aimed at the wrong stage produce findings nobody can act on.
3. **What authority does the review carry, and who receives it?** Whether the review can stop delivery, or only advise. This changes nothing about what you write and everything about the escalation path in step 10. *Default:* write it as though it can stop delivery, and name the escalation route.
4. **Which claims specifically are being made from this research?** Fitness is not abstract. The same study is fit to support one claim and unfit to support another. *Default:* extract the claim set from the report's headlines, executive summary and recommendations, and review against those.
5. **What is already known to be compromised?** Teams usually know their weak points. Asking directly is faster than finding them and it tells you whether the weakness was recognised and handled or recognised and hidden. *Default:* ask. A team that names three real limitations unprompted is a different review from one that names none.
6. **Am I reviewing my own work, or my own team's?** Determines whether the independence adjustments in step 9 apply. *Default:* assume yes if any of the reviewer's judgement is in the study, and apply them.
7. **Has an AI system performed any step of this work?** Determines whether 13.03 runs first and whether the disclosure obligations in K4 §7 and K5 §7 have been met. *Default:* ask explicitly rather than inferring from the output's style.

## 7. Step-by-step methodology

**Step 1. Fix the terms of the review before reading anything for content.** Write down, in four lines: the claims being reviewed, the stage (design or delivery), what the review can still change, and who receives the output. Then set the materiality threshold from the decision at stake. This takes ten minutes and it prevents the two commonest review failures: a review that reports twenty findings of equal weight, and a review that criticises things nobody can now fix. *Correct result:* a scope statement a recipient could disagree with before you begin.

**Step 2. Reconstruct what was actually done, not what was described.** Read the artefacts in a fixed order: brief, design, instrument as fielded, sample records, fieldwork records, data handling log, analysis outputs, report. **Read the report last.** Reading it first installs its narrative, and a reviewer who has absorbed the argument reads the evidence looking for confirmation rather than for problems. Build a one-page reconstruction: what was asked, what was designed, what was fielded, to whom, when, what was done to the data, what was analysed, what is claimed. **Then compare it with the report's method section line by line.** The gap between planned and actual is where most defects live, and it is almost never in the method section, because method sections are written from the proposal. *Correct result:* a reconstruction, plus a list of every divergence between planned and actual.

**Step 3. Test whether the design could have answered the question at all.** Take each research objective and ask the counterfactual: **if the true answer were the opposite of what was found, would this design have shown it?** A design that cannot produce a disconfirming result is not a study, it is a demonstration. Then check the type match. A descriptive design cannot license a causal claim, whatever the analysis technique is called (K4 §3.2). A stated-preference instrument does not measure behaviour. A cross-sectional design cannot establish change over time. A single-market study cannot establish a regional pattern. Where the question was "why", check whether anything in the design was capable of answering "why" rather than "what". This is the highest-yield step in the review and the one most often skipped, because by the time a study is under review its existence feels like evidence that it was the right study. *Correct result:* for each objective, a verdict of answerable, partly answerable with named restrictions, or not answerable by this design, with the reason.

**Step 4. Match the sample to each claim, one claim at a time.** Not the study to the sample: **the claim to the sample.** For every material claim, write the population it implies, then write the population the achieved sample represents, then look at the difference. A claim about "customers" resting on a sample of active account holders excludes the lapsed, who are usually the most informative group. A claim about a market resting on an online sample excludes those without reliable access, which in some populations is the finding. A claim about a subgroup rests on the subgroup's base, not the study's (K3 §3.2). Check the achieved profile against the intended one, check what changed to fill quotas late, and check whether weighting is carrying more than it can: a weight that triples a cell is not correcting a sample, it is manufacturing one, and the effective base tells you what is left (04.05). Where the response rate is unknown, the non-response risk is unquantified and must be stated, not waved past. *Correct result:* a claim-by-claim table of implied population against represented population, with the mismatch named where there is one.

**Step 5. Run the seven check classes across the whole study.** These are the recurring shapes of research failure. Each has a test, and the test is what makes the class useful rather than a label.

| Check class | What it looks like | The test |
|---|---|---|
| **Unsupported claims** | A statement in the report with no analysis behind it, or with an analysis that shows something adjacent | Trace it backwards per K2 §2. Name the evidence. If you cannot, it is unsupported, regardless of whether it is true |
| **Weak methodology** | A design or execution choice that materially reduces what the study can support: wrong mode for the population, a screener that selected on the outcome, a guide that prompted the theme it then reports | Ask what the choice makes impossible, not whether it was reasonable |
| **Questionable statistics** | Untested differences described as real, tests run on overlapping groups or weighted counts, an omnibus result read as specific, multiplicity unaddressed, a null read as equivalence | Per 05.02. Check the test against the design, the base against the claim, and the count of comparisons against the number reported |
| **Small bases** | A percentage on n=40; a subgroup claim on a base the report does not show; a segment defined on a cell of 22 | K4 §7 thresholds, applied to the base of the claim rather than the study |
| **Missing evidence** | An objective with no analysis; a question fielded and never reported; a contradicting stream that appears nowhere; an absent "what we could not establish" section | Work from the instrument, not the report. Every fielded question either appears, or is accounted for |
| **Selective reporting** | The subgroups shown are the ones that worked; one wave omitted; the quotes all point one way; the comparison chosen is the flattering one | Ask what else was run. Compare the analysis outputs against what reached the report |
| **Overinterpretation** | A finding stretched into an insight the data does not carry; causal language on associative evidence; a recommendation resting on a single item | Strip the claim of its evidence and ask what it still asserts (K2 §2.2) |

Work class by class rather than document by document, because a defect of one class recurs across a project and finding it once should find it everywhere.

**Step 6. Walk the lifecycle and name the stage-specific failure.** The check classes catch the recurring shapes; this catches the ones peculiar to a stage. At each stage, the characteristic failure to look for:

- **Question.** The question changed between brief and design and nobody agreed the change; or it cannot be answered by any research; or it is not the question the decision needs.
- **Design.** The design answers a different question well. No comparison where the claim is comparative; no pre-specification where the claim is confirmatory.
- **Sample.** Frame coverage, not sample size, is the usual defect. Also quotas set on the easy variables rather than the ones that matter, and an achieved sample nobody re-derived claims from.
- **Instrument.** A headline resting on a leading, double-barrelled or assumption-carrying item; unequal scale points treated as interval; order effects on the questions that matter most (02.04).
- **Fieldwork.** Mode effects between markets or waves; interviewer or moderator variance; the last-week quota rush; a field window containing an event that moves the measure; unknown or very low response.
- **Data handling.** Undocumented cleaning; exclusions that changed the answer; listwise deletion nobody noticed shrinking the sample differentially; weighting targets older than the population.
- **Analysis.** Undocumented deviation from the analysis plan; subgroup harvesting; a test that does not match the design; a derived measure whose construction is not written down.
- **Interpretation.** The alternative explanation never considered; the finding-to-insight leap with no argued step; confidence acquired in transit (K3 §7).
- **Reporting.** Caveats living in the appendix; bases absent from the page; headlines claiming more than their page; a study reported as answering objectives it did not answer.

*Correct result for steps 5 and 6:* a raw findings list, unsorted and unranked, with each entry naming the artefact and the location. Do not grade yet. Grading while finding causes the reviewer to soften the finding to fit the grade they expect to give it.

**Step 7. Separate defects from limitations. This is the judgement the whole review turns on.** Most research has limitations. Some research has defects. Confusing them produces either a review that condemns ordinary constraint or one that launders a real fault as a caveat.

A **limitation** is a known boundary of what the design can support, disclosed honestly, with the claims kept inside it: a non-probability sample, a single market, a base of 80 on a secondary subgroup. None is a fault. They are what the study bought with its budget.

A **defect** is an error within the design's own terms, or a claim that crosses the boundary the limitation sets: the wrong test, a margin of error attached to a non-probability sample, the base-80 subgroup reported as a finding with no base shown. The same underlying fact (n=80) is a limitation when disclosed and respected, and a defect when it is not.

**The test:** write the limitation on the page, next to the claim, in plain language. Does the claim still stand as written? If yes, it is a limitation to disclose. If the claim collapses, it is a defect, and the required action is to change the claim rather than to add the caveat. A caveat that would destroy the claim it is attached to is not a caveat; it is a correction being disguised as one.

**Step 8. Grade severity and state the required action.** Four levels. Every finding gets exactly one, and the level is set by consequence, not by how annoying the defect is.

| Severity | Definition | Required action |
|---|---|---|
| **Critical** | The research cannot support a central claim, and no disclosure fixes it. The claim is wrong, or unknowable from this study | Withdraw or rewrite the claim. Where the claim is the study's purpose, the study does not deliver, and that is the finding. Escalate per step 10 |
| **Major** | A specific claim is not supported as stated, though the study is sound. Or a defect whose effect on the results is unknown and could be material | Correct the claim, restrict it, or quantify the effect before delivery. Blocks sign-off until resolved |
| **Moderate** | The finding stands, but a limitation materially changes how it should be read and is not disclosed with it | Disclose next to the finding, per K4 §4.3. Does not block delivery once disclosed |
| **Minor** | A defensible choice poorly documented, an inconsistency with no effect on the result, a matter for the next study | Record. Do not clutter the review's front page with these |

Two rules stop severity inflation and severity deflation, which are equally common. **A finding is graded on what it does to the claim, not on how bad the practice was.** A textbook-wrong procedure that cannot have changed the answer is minor. **And cumulative severity is assessed separately:** four moderate findings that all push the headline the same way are collectively major, and a review that reports them individually and never sums them has missed the actual problem. State the cumulative assessment explicitly.

**Step 9. Adjust for the stage, and for whether the work is yours.** At **design stage** everything is still changeable: question, method, sample, instrument, analysis plan. Findings are written as design changes, and the review is worth a multiple of its cost. At **delivery stage** the design is fixed and only three things can change: the claims, the disclosures, and whether the study goes out. Design findings in a delivery review waste attention unless they are explicitly recorded for the next wave.

**Reviewing your own work is a different task needing different technique**, because the defect you cannot see is the one produced by a judgement you still believe. Four things help, and none is trying harder. **Review from the artefacts only:** if a decision is not in a document it is undocumented, whatever you remember agreeing. **Insert distance**, by time if available, otherwise by writing as a named external reviewer with no stake. **Start with the findings you like most**, because they received the least scrutiny on the way in. **Run the hostile read:** write the three sentences a competent critic wanting to discredit this study would say, then answer them from evidence rather than conviction. Independent review remains better, and where stakes are high the honest move is to say so and get one (K5 §2.8).

**Step 10. Write the judgement, and be prepared for it to be negative.** One overall judgement, in one of four forms: **fit for purpose**; **fit for purpose subject to the required corrections listed**; **not fit for purpose as it stands**, with what would make it so; or **not usable for this purpose**, where no correction available at this stage is sufficient. Fitness is always fitness for the stated claims, so a study can be unfit for the claim being made and fit for a narrower one, which is usually the most useful thing a review produces.

**The escalation path** exists because a reviewer's authority is often lower than the decision the review affects, and it is set before the review rather than during the argument. Moderate and below resolve between reviewer and project lead. A major finding the project lead does not accept goes to whoever owns the deliverable, in writing, with both positions recorded. An unaccepted critical finding goes to whoever owns the research function, and **the reviewer's disagreement is recorded in the project file whether or not it changes the outcome.** That record is the point: a review whose findings can be silently overruled provides no assurance, and the written record makes overruling a decision somebody owns. Where the disputed finding concerns participants, ethics or a public claim, it also goes to K5 §2.4 and §2.8, and the reviewer does not sign it off.

## 8. Analytical framework

Two frames run together. The first is the lifecycle spine, which determines where you look:

    Question → Design → Sample → Instrument → Fieldwork → Data handling
        → Analysis → Interpretation → Reporting

The second is the assessment chain, which determines what you do with each thing you find:

    Observation → Classification → Defect or limitation? → Severity → Required action
        → Cumulative assessment → Overall judgement

**Applying it.** The spine is traversed once, in order, and it is traversed forwards. A reviewer who starts at the reporting stage will assess claims against the report's own account of the method, which is the account most likely to be wrong. Each observation from the spine enters the assessment chain separately, and the chain is what stops a review becoming a list of complaints: an observation with no severity is an opinion, a severity with no required action is a criticism, and a set of required actions with no overall judgement leaves the recipient to decide for themselves whether the study is usable, which is the decision they asked the reviewer to make.

**The three questions the whole review reduces to**, and the order is deliberate:

1. **Could this design have answered this question?** If no, nothing downstream matters.
2. **Does the evidence support the specific claims being made?** Claim by claim, against the base of that claim.
3. **Is what is wrong disclosed, and would a reader who acted on this report be misled?** This is the fitness question, and it is the one the judgement answers.

## 9. Output format

**A. Review scope and terms.** The claims reviewed, the stage, the materiality threshold and where it came from, what was and was not made available, what could therefore not be assessed, and the reviewer's independence status.

**B. Overall judgement.** One of the four forms, in the first fifty words, with the two or three findings that drive it. A judgement buried after twelve pages of findings will be read as inconclusive.

**C. Findings register.** The core of the output.

| Ref | Stage | Check class | Finding | Evidence (artefact and location) | Defect or limitation | Severity | Effect on which claim | Required action | Owner |
|---|---|---|---|---|---|---|---|---|---|

Every row names a specific artefact and a specific location. "The sampling is weak" is not a finding. "The sample frame is the active customer file, which excludes the 14% who closed accounts in the review period, while claims C2 and C5 are about customers generally (proposal §3.2; report p.7, p.19)" is a finding.

**D. Cumulative assessment.** Where several findings act on the same claim or push the same direction, their combined effect, stated as its own severity.

**E. Claim-by-claim fitness table.**

| Claim | Where it appears | Evidence it rests on | Base | Supported as stated? | Restriction required |
|---|---|---|---|---|---|

**F. What the study can still support.** The constructive half, and the half that gets the review acted on. Where a claim fails, the narrower claim that survives is usually available and usually still useful.

**G. Review points and escalations.** Per K5 §3, each naming the decision, why it needs a human, and what turns on it. Any unresolved major or critical finding is recorded here with both positions.

**When the evidence is thin.** A review has its own version of the fabrication risk: a finding invented because the register has rows, or a severity raised because a review with only minor findings looks like an inadequate review. Neither is permitted (K4 §1). A review that finds a well-run study says so, briefly, and lists what it checked so the assurance is inspectable. Where an artefact was not available, the register carries an explicit `[not assessed: instrument as fielded not supplied]` row rather than an inference, and the overall judgement is qualified by what could not be seen. **A review is entitled to conclude that it could not determine whether the research is sound. That is a real and reportable outcome**, and it is more useful than a judgement manufactured from partial evidence.

## 10. Quality checks

Run on the review itself, before it is issued. K4 §8 runs anyway.

1. Was the reconstruction built from the artefacts before the report was read, and does it record every divergence between planned and actual?
2. Does every finding name a specific artefact and location, such that the project team can go to it without asking?
3. Is every finding classified as defect or limitation, using the on-the-page test rather than by feel?
4. Is every severity justified by consequence to a named claim, rather than by how poor the practice was?
5. Has the cumulative effect of the moderate findings been assessed separately from the individual ones?
6. Does the review distinguish what this study cannot support from what any study of this kind cannot support?
7. Is every required action something the recipient can actually do at this stage of the project?
8. Does the overall judgement follow from the findings, and would it survive the project lead disagreeing with the two findings they are most likely to dispute?
9. Has the review checked the objectives that were not answered, not only the claims that were made?
10. Has the review looked for what is absent (a fielded question never reported, a contradicting stream, a missing "what we could not establish") rather than only assessing what is present?
11. Where the reviewer contributed to the work, are the independence adjustments applied and the conflict declared in section A?
12. Is anything in the register present because a review is expected to find problems?
13. Would a competent researcher who disagrees with this review be able to see exactly what they are disagreeing with?
14. Is the escalation path stated, and is every unresolved major or critical finding recorded with both positions?

## 11. Common failure modes

| Failure | How to recognise it | How to prevent it |
|---|---|---|
| **The reassurance review** | Findings are all minor, the tone is collegial, the judgement is positive, and it took two hours | Set the materiality threshold and the claim set first. Ask what a hostile competent reader would find, and answer it from evidence |
| **Reading the report first** | The review assesses the method as described rather than as executed | Step 2's fixed reading order. Report last, always |
| **Severity inflation** | Everything is major; the register has thirty rows of equal weight | Grade on consequence to a named claim. A review that flags everything at maximum severity is as useless as one that flags nothing |
| **Severity deflation on a convenient finding** | The defect under the finding everyone likes is graded lower than the identical defect elsewhere | Grade by class before looking at which claim it hits. Then check the two gradings against each other |
| **Defect laundered as limitation** | A caveat added that, if read, would destroy the claim it sits under | Step 7's on-the-page test. If the claim collapses, correct the claim |
| **Limitation condemned as defect** | The review criticises a study for being the study it was budgeted to be | Ask whether the claims stayed inside the boundary. Constraint honestly disclosed is not a fault |
| **The cumulative blind spot** | Six moderate findings, all pushing the headline upward, reported as six moderate findings | Step 8's cumulative rule. Assess direction of effect, not just count |
| **Design findings in a delivery review** | Recommendations nobody can act on, and a team that stops reading | Step 9. State explicitly where a finding is recorded for the next wave rather than for this one |
| **Reviewing only what is there** | A clean review of a study that silently dropped two objectives | Work from the instrument and the brief, not from the report's contents page |
| **AI: fluent agreement** | A review that summarises the study accurately and finds nothing, or finds only presentational issues | The review's job is adversarial. Run the three counterfactuals in step 3 explicitly and record the answers |
| **AI: inventing a methodological standard** | A finding citing a threshold, convention or rule that does not exist, stated with confidence | Every methodological standard invoked must be nameable and defensible in general terms. If it cannot be, it is the reviewer's judgement and is labelled as such |
| **AI: severity by vocabulary** | Severity assigned from how serious the defect sounds rather than from its effect on a claim | Force the "which claim, and what happens to it" column to be populated before the severity column |
| **Self-review from memory** | "We did consider that" appears in the review as though it were a record | If it is not in a document, it is undocumented. That is itself a finding |

## 12. AI guardrails

Universal prohibitions are inherited from K4 and not repeated. K4 §4.1 and §4.2 govern this skill in full, since a review is precisely where suppression of the inconvenient finding does most damage.

1. **Never grade a finding's severity by how the finding was received.** Severity is set before the project team's response and does not change because the response was reasonable, senior or upset. It changes only on new evidence.
2. **Never soften, merge or reorder findings to make the review more acceptable.** Combining two majors into one moderate is falsification of a review, and it is the commonest way a review is neutralised.
3. **Never invent a finding to fill the register, and never raise severity so the review looks substantial.** A study with three minor findings gets a three-row register (K4 §1).
4. **Never assess a method you have not seen.** If the instrument as fielded was not supplied, the instrument was not reviewed, and the register says so rather than inferring from the report's method section (K4 §6.4).
5. **Never cite a methodological standard, threshold or convention you cannot state and defend in general terms**, and never attribute one to an authority. Where the standard is the reviewer's professional judgement, say so; where practice is genuinely divided, give both positions and say which this review applied.
6. **Never let the direction of a finding's commercial convenience enter the assessment**, in either direction. Findings that support the client's preferred answer get the same scrutiny as those that do not, and findings that embarrass the reviewer's own team get no extra.
7. **Never issue an overall judgement that the findings do not support**, in either direction. A register of critical findings cannot end in "fit for purpose with minor corrections", and a register of minor ones cannot end in "not usable".
8. **Never treat the absence of documentation as the absence of a problem, or as proof of one.** An undocumented decision is a documentation defect of known severity and an analytical defect of unknown severity, and the review says exactly that rather than assuming either way.
9. **Never review a report as though it were the research.** Where only the report was available, the output is labelled a report-level review throughout, including in its judgement.

## 13. Best-practice principles

1. **The review's job is to make the research stronger, not to agree with it.** These are not the same activity and the second one is what most reviews quietly become. A reviewer who has never delivered an unwelcome judgement should assume they are performing the second.
2. **Read the evidence before the argument.** The single highest-leverage habit in the skill. A reviewer who reads the report first is reviewing their memory of the report.
3. **Ask what would have shown the opposite.** Nearly every serious design defect is visible from this one question, and it takes a minute.
4. **Fitness is always fitness for a claim.** "Is this good research?" is not answerable. "Can this research support this sentence?" is, every time.
5. **Most research has limitations and only some has defects.** A review that cannot tell them apart will either destroy usable work or approve unusable work, and it will do both in the same document.
6. **Grade on consequence, not on offence.** The wrong test that cannot have changed the answer is minor. The undisclosed base under the headline is major. Professional distaste is not a severity level.
7. **Sum the moderates.** Individually tolerable compromises, all pointing one way, are how studies end up confidently wrong. The cumulative assessment is often the only place the real problem appears.
8. **The most convenient finding deserves the most scrutiny.** This holds whether the convenience is the client's, the team's, or the reviewer's own.
9. **Review your own work from the artefacts, never from memory**, and start with the findings you like best. Everything you remember agreeing is invisible to the review and to everyone downstream.
10. **A design review is worth several delivery reviews.** At design stage the finding costs a conversation; at delivery it costs the study. Where a team can afford only one review, this is where it goes.
11. **Say what the study can still support.** A review that only removes claims will be resisted and eventually ignored. The narrower claim that survives is nearly always available and is usually still worth having.
12. **Write the review so that disagreeing with it is easy.** Specific artefact, specific location, specific claim affected. Vague criticism cannot be answered, which means it cannot be resolved and will simply be outlasted.
13. **Record the overrule.** A review whose critical findings can be set aside without a written record provides no assurance to anyone, including the person who set it aside.

## 14. Worked example

*Fictional scenario, used for illustration only. The organisation, figures and documents below are invented for the purpose of demonstrating method.*

**INPUT.** A national employment services agency commissions an evaluation of a six-month job-readiness programme. The delivered report concludes that the programme "increases employment outcomes by 18 percentage points" and recommends national rollout. The review is requested at delivery stage, three weeks before a funding submission. Artefacts supplied: the brief, a proposal, the questionnaire as fielded, sample records, analysis tables, and the report. No raw data, no analysis plan.

**PROCESS.**

*Steps 1 and 2.* Scope fixed: three claims under review, the headline outcome claim and two supporting ones; delivery stage; what can change is claims, disclosures and whether it goes out; materiality is high because a funding decision rests on it. The reconstruction, built before reading the report, produces two divergences from the proposal. The proposal specified a comparison group of eligible non-participants drawn from the same regional waiting lists. The sample records show the comparison group was drawn instead from a general working-age population sample, because waiting-list contact details could not be released in time. And fieldwork ran eleven weeks rather than six, with 40% of comparison-group interviews completed in the final two weeks.

*Step 3, and the finding that decides the review.* The counterfactual: if the programme had no effect, would this design have shown it? Participants are people who applied to and completed a six-month programme. The comparison group is the general working-age population, which includes people who never sought work, people already employed, and people who would never have been eligible. The two groups differ on the thing most predictive of the outcome, which is motivation to seek work, before the programme does anything. The design compares self-selected completers against the general public, so an 18-point gap is exactly what would be observed if the programme did nothing at all. This is not a limitation. It is a defect in the design's own terms, because the design was specified with an eligible-non-participant comparison and that is what would have made the claim possible.

*Step 4.* The claim implies a population of eligible jobseekers. The participant sample represents programme completers, n=612, excluding the 23% who started and did not complete, whose outcomes are absent from the analysis entirely. Completion is itself likely to correlate with the outcome, so the exclusion compounds the same problem rather than being separate from it.

*Step 5 and 6, other findings.* Questionable statistics: the 18-point gap is reported as significant with no test named and no bases on the chart. Small bases: two of the four regional breakdowns rest on 47 and 38 participants and are reported as percentages without bases. Missing evidence: three of eleven fielded questions, including the one on prior employment history that would have supported a matched comparison, do not appear anywhere in the report or the analysis tables. Reporting: the fieldwork extension and the comparison-group substitution appear in neither the method section nor the appendix.

*Step 7, the judgement call.* The late fieldwork skew looked at first like a defect. Applying the on-the-page test: "40% of comparison interviews were completed in the final two weeks of an extended field period" written next to the finding does not by itself collapse the claim, and no evidence of a compositional shift is available without the raw data. It is graded a moderate limitation requiring disclosure, with a note that it could not be assessed further because the dataset was not supplied. Recorded as `[not assessed: raw data not available]` rather than inferred either way.

*Step 8.* One critical (comparison group cannot support a programme effect), one major (non-completers excluded, effect on the estimate unknown and plausibly large), three moderate (untested difference presented as significant, unbased small-base regional claims, undisclosed design change), two minor. Cumulative assessment: the critical and the major push in the same direction, both inflating the apparent effect, so the honest statement is that the reported effect is an upper bound of unknown looseness rather than an estimate.

*Step 10.* Judgement: **not fit for purpose as it stands**, for the outcome claim. The review then does the constructive half. The study does support three narrower claims: what completers report about the programme's usefulness, which parts they valued, and how their self-reported confidence changed, all as descriptive findings on the completer sample with bases shown. Those are worth having and are not in dispute. What would make the effect claim possible is named: a matched comparison using the prior-history question that was fielded and never analysed, which may be recoverable from the existing data, or a waiting-list comparison in a second phase.

**OUTPUT.** A findings register of seven rows with artefact locations; a claim-by-claim fitness table showing one claim unsupported, one requiring restriction and one supported; a cumulative assessment stating the direction of the combined bias; an overall judgement of not fit for purpose for the headline claim with the three surviving claims specified; and an escalation, since the funding submission is in three weeks and the project lead does not accept the critical finding. Both positions are recorded in writing to the research function owner, and the reviewer does not sign off. `RESEARCHER DECISION REQUIRED` per K5 §3.1: whether the submission proceeds on the narrower claims or is delayed is a decision for the accountable owner, not for the reviewer, but it must be made knowing that the 18-point figure cannot be defended.

## 15. Advanced usage

**Reviewing a programme rather than a project.** Trackers, rolling studies and multi-year programmes accumulate drift that no single-wave review will find: a base definition that moved in wave 4, a question reworded in wave 7, a segment redefined in wave 9, each documented at the time and none reassessed since. Review the programme as a whole once every few waves, working from the version history rather than the current wave, and treat comparability across the full series as the claim under review.

**Reviewing under adversarial conditions.** Where the review has been commissioned because a study is contested, separate three things that will otherwise be argued as one: whether the method is sound, whether the claims match the method, and whether the conclusion is one the commissioner likes. Report all three separately. A study can be methodologically sound and still be making claims it cannot support, and that distinction is the whole value of the review to a party who has assumed the argument is binary.

**Reviewing for reuse.** When an existing study is proposed as the answer to a new question, the review question is not whether the study was good but whether it is fit for the new claim: the population, the period, the wording and the definitions all have to hold for the new use. Most reuse failures are period and definition failures rather than quality failures, and this review is fast and high-value.

**When the standard approach does not fit.** Where the artefacts are so incomplete that no reconstruction is possible, do not produce a diluted review. Produce an assessability report instead: what exists, what is missing, what could be assessed if it were produced, and what claims are currently unverifiable. That is an honest and actionable output, and it frequently causes the missing material to appear.

## 16. Skill chain

**Recommended previous skills:**
- **13.03 AI Output Verification.** Where any part of the work was AI-produced, hands over a verification report identifying fabricated specifics, unverified citations and calibration failures, so this review assesses method rather than rediscovering them.
- **13.02 Source and Citation Verification.** Hands over a verification log with a status per source, which resolves the desk-research component of the evidence base before methodological review begins.
- **13.04 Bias Detection.** Hands over a bias register with mechanism and likely direction, which this review reads as an input to severity and to the cumulative assessment.
- **12.06 Research Report QA.** Hands over a document-level check, so this review is not spending its attention on inconsistencies that are presentational.

**Recommended next skills:**
- **01.01 Research Brief Interrogation.** Where the review traces a defect back to an unclear or changed question, this is where the correction belongs for the next project.
- **01.04 Research Method Selection** and **01.06 Sampling Strategy.** Where a design-stage review concludes the method or the sample cannot answer the question, these rebuild it.
- **12.03 Research Report Compilation.** Where claims must be restricted or removed, the corrected report is recompiled with its evidence map rather than patched.

**Runs well alongside:**
- **05.02 Statistical Testing**, which supplies the standard the "questionable statistics" check class is run against.
- **13.05 Research Ethics and Consent Design** and **13.06 AI Research Governance**, wherever a review finding turns out to be an ethics or governance question rather than a methodological one.
- **K3** and **K5**, which supply the calibration standard the interpretation and reporting stages are assessed against, and the escalation and sign-off discipline the review's judgement depends on.

---
A Yazi Supplied Skill and resource.
