---
name: research-rigour-audit
description: >
  Audits a study's methodological integrity independently of what it found:
  question, design, sampling, measurement, procedure, analysis and inference,
  plus alignment, reproducibility, reporting completeness and the signals of
  analytic flexibility. Use for "audit this study's methods", "is this study
  rigorous", "check the methodology of this paper", "could someone reproduce
  this", "were the statistics reported properly", "does the analysis look
  chosen after the data", "assess qualitative rigour", "methodological review
  before submission".
category: 15 Academic University Research
ref: "15.22"
tier: 3
inherits: [K2, K3, K4, K5]
---

# Research Rigour Audit

## 1. One-line description

A structured audit of whether a study was designed, executed, analysed and reported well enough for its conclusions to mean anything, conducted deliberately without regard to what the study found, and reported as located findings with severity ratings rather than as an overall verdict.

## 2. What this skill is used for

**Read this before you start.** Whether an AI system may be used on the material you are auditing is governed by institutional policy, and the material is frequently unpublished work belonging to an identified colleague or student, sometimes under review, sometimes containing participant data. Where the audit concerns a student's work it also engages assessment policy. Settle both questions before opening the file. Section 5 makes the policy check a required input; Section 12 sets the operating rules. Where the position is unknown, the conservative default is that the material is not uploaded, and the audit is confined to structuring and testing the auditor's own reading.

**The research problem it solves.** Methodological review is contaminated by results. A study reporting a striking finding is scrutinised for reasons to believe it; a study reporting nothing is scrutinised for reasons to dismiss it. The consequence is that method problems are found when the finding is unwelcome and missed when it is congenial, which is precisely backwards, because a strong finding produced by a weak method is the more dangerous object. Beyond that, review without a structure misses systematically: it catches the sampling problem because sampling is familiar, and does not notice that the analysis reported was not the analysis planned, that no exclusion criteria are stated, or that the reported statistics are incomplete in ways that make the effect size uncheckable. This skill separates the audit from the findings, imposes a layered structure so nothing is skipped, and requires located evidence and a severity rating for every finding.

**Where it sits in the research lifecycle.** After a study exists in reportable form. Before a thesis is submitted or examined, before a paper is sent out, before a body of work is built on, and as a standing quality function within a research group or a departmental review process.

**Typical use cases.**
- Auditing a study underlying a thesis before submission or as part of examination.
- Internal methodological review before a paper leaves a research group.
- Assessing whether a published study is sound enough to build on or to cite as a foundation.
- Checking that a completed study can be reproduced from what is written.
- Screening a body of literature for methodological quality in a systematic review.
- Reviewing a colleague's work at their request, or as a departmental quality process.
- Establishing, for a research group, whether analytic decisions were made before or after the data were seen.

**Who uses it.** Supervisors auditing a student's underlying study; examiners establishing the methodological facts before making a criterion judgement; research group leads and departmental quality panels; methodologists and statisticians asked for a second opinion; systematic reviewers appraising included studies; researchers auditing their own work before submission, which is the highest-value and least-practised use.

## 3. When to use it

- A study is complete and someone has to judge whether its conclusions are supportable.
- A finding is surprising, and the correct next question is whether the method could have produced it.
- A study is about to be built on, cited as a foundation, or used to justify a further programme of work.
- A thesis is being examined and the methodological facts need establishing before criterion judgements are made.
- A paper is coming back from review with methodological criticisms and you need to know which are correct.
- A group wants a standing pre-submission check that is applied the same way to every study.
- You need to know whether a study could be reproduced by someone with only the written account.
- Analytic decisions in a study appear to have been made after the data were seen, and the question needs asking properly rather than insinuated.

## 4. When NOT to use it

- **The question is whether the work meets a degree standard.** That is examination: criteria, calibration to a degree, an outcome. Go to **15.21 Thesis Examination and Marking**. An audit establishes methodological facts and issues no verdict on a candidate; the two are frequently used together, and confusing them produces an audit with an outcome attached, which is neither.
- **The concern is misconduct.** Fabrication, falsification and plagiarism are institutional matters with defined evidence standards, investigative procedures and rights of response. This skill identifies methodological concerns and does not make allegations. Where an audit finding is consistent with misconduct, it is still reported as a methodological concern, and the route is the institution. See §12.2 and Step 12.
- **The study is still in design.** Auditing an unrun study against execution criteria produces nothing useful. Design review is a different task, served by **01.04 Research Method Selection** and, for academic work, **15.08 Research Design and Methodology Chapter** and **15.16 Doctoral Methodology Justification and Rigour**.
- **What is wanted is a commercial deliverable review.** Where the object is a client report and the question includes whether the conclusions serve the decision, whether the narrative is supported and whether the deliverable is fit to present, that is **13.01 Research Quality Review**. This skill is narrower and stricter: it audits the study, not the report, it takes no view on usefulness, and it applies academic standards for reporting completeness and reproducibility that a commercial deliverable is not expected to meet. Where both are needed, run 13.01 for the deliverable and this skill for the study beneath it.
- **The method cannot be established from what is available.** Where the write-up reports too little to evaluate, the audit output is a reporting finding, not a methods verdict. Do not reconstruct the probable method. An audit that infers what the researchers presumably did is worse than no audit, because it launders a guess into an assessment.
- **You cannot assess the domain.** Field-specific method conventions, instruments and analytic standards vary enormously, and an audit conducted from general methodology against a specialist design will generate confident false findings. Audit what you can assess, mark the rest as requiring specialist review, and say so.
- **The purpose is to discredit the study.** An audit commissioned to find fault will find it, because every study has limitations. The audit is only worth anything if the same protocol would have been applied to a study you liked. If it would not, do not run it.
- **The auditor has a stake in the result.** A competing research programme, a co-authorship, a prior public position on the question. Declare it, and let whoever commissioned the audit decide.

## 5. Required inputs

**Required.**
- **A complete methodological account of the study**: what was done, to whom, with what instruments, in what order, and how it was analysed. A results section alone cannot be audited.
- **The stated research question or hypotheses.** Without them, alignment cannot be assessed, and alignment is where most of the value is.
- **Institutional policy on AI use for this material, and on the personal data it may contain.** If unknown, resolve before starting.
- **The purpose of the audit and who receives it.** A pre-submission internal check, a contribution to an examination, and a review of published work carry different obligations and different tones.

**Optional, and what each one adds.**
- **A protocol, pre-registration, analysis plan, ethics submission or approved proposal.** The single most valuable optional input. It converts the hardest question in the audit, whether analytic choices were made before or after the data were seen, from inference into comparison. Without it, that question can only be raised as a signal, never settled.
- **The dataset, or the analysis code and outputs.** Enables the reproducibility check to be run rather than assessed, and lets reported statistics be verified rather than checked for internal consistency only.
- **The instrument in full.** Question wording, response options, order and routing. Measurement problems are invisible in a methods summary and obvious in the instrument.
- **Ethics approval and the consent materials.** Both a compliance matter and a methodological one: what participants were told constrains what can be inferred.
- **Earlier drafts or conference versions.** Where an outcome, a hypothesis or a subgroup appears in one version and not another, that is evidence about the analytic history that no single version contains.
- **The authors' own stated limitations.** Tells you what they already know, so the audit adds rather than repeats, and a study that names its own weaknesses accurately is displaying a form of rigour worth recording.

## 6. Questions to ask before starting

1. **Can I be blinded to the findings, and if not, how do I manage that?** Determines the whole shape of the audit. Default: read methods before results, complete the design and analysis layers before reading any result, and record that the blinding was partial.
2. **What is the audit for, and who will read it?** A pre-submission check invites detail on everything; an audit feeding an examination must confine itself to methodological fact. Default: assume internal, constructive, and written for the authors.
3. **What standards apply in this field?** Reporting conventions, acceptable designs, expected completeness of statistical reporting, whether pre-registration is normal, and what qualitative rigour means here all vary by discipline and by journal. Default: audit against general methodological principles, state that you have done so, and mark anything that may be a field convention rather than a defect.
4. **Is a protocol or analysis plan available?** Determines whether analytic flexibility can be assessed or only flagged. Default: assume none exists, ask for one, and if none exists say so as a finding in its own right rather than treating its absence as evidence of anything.
5. **Is this quantitative, qualitative or mixed, and what would rigour mean for each strand?** Default: audit each strand on its own terms, and audit the integration separately, because mixed-methods work most often fails at the join.
6. **What is the study claiming?** The audit assesses whether the method supports the claims made, not whether it meets an abstract ideal. A modest claim on a modest design is not a finding. Default: locate the strongest claim in the abstract and the conclusion, and audit against that.
7. **Would I be applying this protocol to a study I agreed with?** Default: if the honest answer is no, stop.

## 7. Step-by-step methodology

**Step 1. Set the scope and the standard, and fix the reading order.**
Record what is being audited, what materials you have, what standard you are auditing against (general methodological principles, a field reporting guideline, or a protocol), the AI and data position, and any interest you hold. Then fix the reading order and do not deviate: methods first, in full; then the design and analysis layers audited; then results; then discussion and conclusions. This is the mechanism that delivers the audit's founding discipline, which is that you do not know and do not care what the study found while you are judging how it was done. *Correct result: a scope note written before reading, naming the standard and the reading order.*

**Step 2. Audit the question.**
Is a research question or hypothesis stated explicitly, in one place, in answerable form? Is it one question or several bundled into a sentence? Are the constructs defined, or named only? Is it the question the study went on to answer, or one rewritten to match what was found, visible as a question that fits the results exactly and that no design would have been built to answer? *Correct result: the question quoted with its location, its constructs listed, and a note of any bundling, vagueness or post hoc fit.*

**Step 3. Audit the design.**
Name the design in your own words from the description, not from the authors' label; where the two disagree, the procedure governs. Then ask what it can and cannot establish: association, difference, sequence, mechanism, causation. Check that the comparison, control or counterfactual the claim requires actually exists. Check the threats specific to it: confounding and selection in observational work, allocation and blinding in experimental work, attrition and testing effects in longitudinal work, and, for any design, whether the unit of analysis matches the unit of sampling. *Correct result: a design statement in your own words, a list of what it can support, and located threats.*

**Step 4. Audit sampling and recruitment.**
Population defined, frame described, selection mechanism stated, achieved sample characterised. Then the questions that matter more than sample size: who could not enter the sample at all, and how do they differ; what was the response or consent rate and how was non-response handled; who dropped out and were they different; and are exclusion criteria stated, were they stated in advance, and how many cases did each remove. Exclusions are audited carefully because they are the quietest place in a study for a result to be manufactured, and a legitimate exclusion rule and an outcome-serving one look identical when only the surviving n is reported. Check that the achieved sample supports the claims made about the population, which is separate from whether it is large (K4 §3.3). *Correct result: a sampling account with the exclusion cascade reconstructed as far as the report allows, and any point at which the numbers do not reconcile located.*

**Step 5. Audit measurement.**
For each construct in the question, identify what actually measured it: a validated instrument used as validated, a validated instrument modified, or items written for this study. Modification is common and legitimate and it forfeits the borrowed validity evidence, which is a finding whenever the study still cites that evidence. Check reliability where the design depends on it, check scale direction and anchoring, and check wording for the standard failures: double-barrelled items, leading framing, assumed knowledge, response options that do not span the range. Then ask the directness question (K3 §3.4): does the measure capture the construct in the claim, or something adjacent? *Correct result: a construct-to-measure map, with the distance between each claim and its measure stated.*

**Step 6. Audit procedure and execution.**
What happened, to whom, in what order, by whom, over what period. Look for the execution facts that change interpretation and are frequently omitted: who administered the instrument and whether they knew the condition, the setting, whether the fieldwork period contained an event that would affect responses, any protocol deviation and how it was handled, and how missing data arose and were treated. Missing data handling is a required disclosure and its absence is a finding, because deletion, imputation and carrying forward have materially different consequences. *Correct result: a procedure timeline, and the execution facts needed to interpret the results that are absent from the report.*

**Step 7. Audit the analysis.**
Identify each analysis performed, and for each: does it match the design and the data type; are its assumptions stated and checked; is the model specification given in full, including every covariate; is the unit of analysis correct, particularly where observations are nested or repeated. Count the comparisons actually made, including those implied by subgroup reporting, and check for any correction or acknowledgement of multiplicity. Then compare the analysis performed against the analysis the question required; where they differ, that is a finding regardless of cause. *Correct result: an analysis inventory, each entry marked appropriate, questionable or mismatched, with the reason.*

**Step 8. Audit the inference.**
Now read the results and the conclusions, and line up each conclusion against the specific result that supports it. Look for the four standard overreaches: causal language where the design licenses only association (K4 §3.2); generalisation beyond the sampled population, period or setting; a claim of no effect drawn from a non-significant result in an underpowered study, which is not evidence of absence; and a conclusion resting on a subgroup or secondary outcome while the primary outcome is reported quietly or not at all. Check that the abstract says what the results say, because that is where overreach concentrates and is the only part most readers will see. *Correct result: a conclusion-by-conclusion table, each with the result it rests on and a verdict of supported, overreaching or unsupported.*

**Step 9. Run the four alignment checks.**
The audit's highest-yield step, and it can only run once the layers are done. Each check is a single question with a yes or no answer and a located reason.

| Check | The question |
|---|---|
| **Design to question** | Could this design, executed perfectly, have answered the question as stated? |
| **Sample to claim** | Does the achieved sample support claims about the population the conclusions name? |
| **Analysis to design** | Does the analysis performed match the design and the structure of the data? |
| **Conclusion to result** | Does each conclusion follow from the results actually reported? |

A study can be locally sound at every layer and still fail an alignment check, and when it does, that failure is the study's central problem and everything else is detail. *Correct result: four answers, each located, with a one-line reason.*

**Step 10. Assess reproducibility.**
The operational question: could a competent researcher in this field repeat this study from what is written, without asking the authors anything? Work through what would be needed: the instrument or protocol in full, the sampling and recruitment procedure, the exclusion rules with their thresholds, the analysis specification including software and version where it matters, and access to data or the terms on which it is available. Mark each present, partial or absent. Reproducibility is a checklist, not an ideal to be praised in the abstract. *Correct result: a reproducibility table with the absent elements named.*

**Step 11. Check statistical reporting completeness.**
Reported statistics are audited for what is missing as much as for what is wrong, because the common omissions make a result uncheckable. Look for: a denominator for every proportion; the exact test and its assumptions; an effect size alongside any significance claim, since a p value alone says nothing about magnitude; the estimate's precision as an interval or standard error; the n entering each test; how missing data affected each n, which is why n often varies between analyses unexplained; and any significance claim made without a test (K4 §3.1). Then run the consistency checks the report permits: do subgroup ns sum to the total, do percentages sum, does the statistic correspond to the reported n. Inconsistency is reported as inconsistency, located, never as accusation. *Correct result: located omissions and inconsistencies, each with what it prevents a reader establishing.*

**Step 12. Assess analytic flexibility, and name questionable practices plainly.**
Some studies show signals that the analysis was chosen after the data were seen. Name the signals, never the motive: an outcome in the results that appears in neither the question nor the methods; a subgroup analysis carrying the headline with no prior rationale; exclusion criteria described after the analysis rather than with the sampling; an unusual analytic choice with no justification; a hypothesis phrased in a way that could only have been written knowing the result; covariates present in one model and not another; a reported n inconsistent with the sampling account. Four practices are named plainly because unnamed they go unexamined:

| Practice | What it looks like |
|---|---|
| **Selective reporting** | Outcomes, measures or analyses collected and not reported, or reported only where they reached a threshold |
| **Undisclosed flexibility** | Analytic decisions made among many possible ones, with only the chosen path described |
| **Hypothesising after results are known** | A hypothesis presented as prior that was formed once the data were seen |
| **Exclusion misuse** | Exclusion rules applied, changed or discovered after their effect on the result was visible |

**The distinction that governs this step: a questionable research practice is not misconduct.** Most instances arise from ordinary analytic freedom exercised without a protocol, in fields where none is expected, by researchers acting in good faith. Misconduct is fabrication, falsification and plagiarism, requires intent, and is established by an institutional process with evidence standards and rights of response. **This skill identifies methodological concerns and does not make allegations.** Every finding here is a located observation about the report, with the innocent explanation stated where one exists and the resolving material named. Where a protocol exists, compare and report differences factually; where none exists, the finding is that the question cannot be settled from the available material. `RESEARCHER DECISION REQUIRED` (K5 §2.4) on whether anything goes further than the audit report, which is for the academic and the institution. *Correct result: signals as located observations, with innocent explanations and resolving material named, and no statement of intent anywhere.*

**Step 13. Audit qualitative rigour on practice, not vocabulary.**
Qualitative work is routinely audited by scanning for the words: credibility, transferability, dependability, saturation, reflexivity, triangulation. That audits nothing, because the vocabulary is cheap and the practice is not. Audit the practice. Does the sampling strategy fit the purpose, and is the number of participants justified on grounds other than convention? Is analysis evidenced as systematic, with a described coding process and an audit trail of how codes became themes? If saturation is claimed, is there a stated criterion, or is the word doing the work alone? Do the data shown support the prominence claimed for each theme? Is disconfirming evidence presented, or does every participant agree, which is a finding about the analysis rather than the participants? Is reflexivity substantive, tracing a position's effect, or a paragraph declaring one exists? Is the named analytic approach actually followed? *Correct result: a rigour assessment in which every judgement cites a described practice.*

**Step 14. Assign severity, evidence each finding, and write the audit.**
Every finding gets a location, a statement of what it prevents, and a severity. Severity is assigned on consequence for the conclusions, not on how much it offends good practice.

| Severity | Definition |
|---|---|
| **Critical** | The stated conclusions are not supportable by this study as conducted. |
| **Major** | A specific conclusion requires substantial qualification, or a claim must be narrowed. |
| **Minor** | Reporting or execution weakness that does not change the conclusions but limits what a reader can check. |
| **Observation** | A defensible alternative choice, or a limitation the authors have already stated accurately. |

Then apply two discipline tests. Would this finding have been raised if the study had found the opposite? And is it a finding about the study, or about your preference, which the observation class exists to absorb? Finally, report what is strong, specifically, because an audit that lists only faults is uncalibrated and will be discounted. *Correct result: an audit report of located, severity-rated findings, with strengths recorded, and no overall score.*

## 8. Analytical framework

The audit is built on a layered stack, worked bottom to top, with alignment checks running across it:

    Question → Design → Sampling → Measurement → Procedure → Analysis → Inference

Each layer is audited on its own terms, and a defect at any layer caps everything above it: perfect analysis of data from a design that cannot answer the question buys nothing. This is why the order is fixed and why the audit is not permitted to start at the analysis, which is where methodological reviewers instinctively begin because it is the most technical layer and the easiest to be confident about.

Cutting across the stack are the four alignment checks, which are diagonal rather than vertical:

    Design ⟷ Question      Sample ⟷ Claim
    Analysis ⟷ Design      Conclusion ⟷ Result

Layer defects are local and usually fixable in the writing or by qualification. Alignment failures are structural, and a study can pass every layer and fail here. When it does, that is the audit's headline.

Each finding is then carried as:

    Finding → Location → What it prevents → Severity → What would resolve it

The fourth element is what makes the audit usable and the fifth is what makes it constructive. An audit that reports findings without saying what would resolve them is a critique, not an audit.

## 9. Output format

**1. Audit scope.** What was audited, what materials were available and what was not, the standard audited against, the reading order used and whether blinding to results was achieved, the auditor's interests, and the AI and data position.

**2. The study in the auditor's own words.** Question, design, sample, measures, analysis, in five sentences, without the authors' framing. Divergence between this and the study's own account is itself informative and is noted here.

**3. Alignment.** The four checks, with answers and located reasons. This section goes near the front because it usually carries the audit's most consequential content.

**4. Findings by layer.**

| ID | Layer | Finding | Location | What it prevents | Severity | What would resolve it |
|---|---|---|---|---|---|---|

**5. Reproducibility.** Element-by-element: present, partial, absent.

**6. Statistical reporting completeness.** Located omissions and inconsistencies.

**7. Analytic flexibility observations.** Signals as located observations, each with an innocent explanation where one exists and the material that would resolve it. This section carries a standing statement that it identifies methodological concerns and makes no allegation.

**8. Strengths.** Specific and located.

**9. Limitations of this audit.** What could not be assessed and why: absent materials, domain limits, partial blinding.

**No overall score, and no verdict.** The audit reports findings with severities and stops. A single quality rating collapses exactly the information the audit exists to produce, and it invites the reader to skip the findings. Where the commissioning process requires a summary judgement, the academic makes it from the findings and owns it (K5 §2.8).

**When the material is too thin to audit, the format must not manufacture an assessment** (K4 §1). Layers that cannot be assessed are marked "insufficient information to assess", which is a reporting finding of its own and frequently the most important thing the audit says. Do not infer the probable method.

## 10. Quality checks

Run before the audit is issued.

1. Were the methods audited before the results were read, or is the partial blinding disclosed?
2. Does every finding carry a location in the material?
3. Does every finding state what it prevents a reader from establishing, rather than only naming a deviation from good practice?
4. Is every severity rating justified by consequence for the conclusions rather than by how far the study departs from an ideal?
5. Has each finding been tested against the reversal question: would I have raised this if the study had found the opposite?
6. Is any finding a matter of legitimate alternative choice, and if so, is it in the observation class?
7. Has any part of the method been inferred rather than read? Remove it, and report the absence instead.
8. Is any statement about the researchers' intent, motive or good faith present anywhere? Remove it.
9. Does every analytic flexibility observation name the innocent explanation and the material that would resolve it?
10. Have the four alignment checks all been answered, with reasons?
11. Is the qualitative assessment based on described practice rather than on the presence of vocabulary?
12. Are strengths recorded specifically?
13. Are field conventions distinguished from defects, and is anything you could not assess marked as such?
14. Would the authors be able to act on this audit without a conversation?
15. Is there anywhere the audit implies an outcome, a verdict or a score it is not entitled to give?

## 11. Common failure modes

| Failure | How to recognise it | How to prevent it |
|---|---|---|
| **Results-contaminated review** | The methods look worse when the finding is unwelcome | Fix the reading order and audit before reading results (Step 1) |
| **Starting at the analysis** | A detailed statistical critique of a design that could not answer the question | Work the stack bottom to top (§8) |
| **Vocabulary auditing** | Qualitative rigour assessed by finding the word "saturation" | Audit described practice only (Step 13) |
| **The ideal-study standard** | Every study fails because none is perfect | Severity assigned on consequence for the stated conclusions (Step 14) |
| **The reconstructed method** | The audit describes what the researchers presumably did | Insufficient information is a finding, not a gap to fill (K4 §6.4) |
| **Sample size as a proxy for sampling** | n is discussed, the frame and non-response are not | Audit frame, coverage, response and exclusions before size (Step 4) |
| **Exclusion cascade unexamined** | Only the final n appears anywhere | Reconstruct the cascade and locate where it stops reconciling |
| **Insinuation** | Language implying the analysis was manipulated, without a located observation | Signals as observations, innocent explanations stated, no intent (Step 12) |
| **Allegation creep** | An audit finding restated as misconduct | The QRP and misconduct distinction is stated in the report itself |
| **Missing statistics unnoticed** | Effect sizes, intervals and denominators absent and not remarked on | Run the completeness list explicitly (Step 11) |
| **The abstract unchecked** | Conclusions audited, the abstract's stronger version not | Audit the abstract against the results separately (Step 8) |
| **Field convention treated as defect** | A discipline-normal practice reported as a problem | Mark uncertain items as possible convention, requiring specialist view |
| **AI-generated methodological criticism** | Findings that describe things the material does not contain | Every location verified in the document (§12.3) |
| **Uncalibrated audit** | Only faults, no strengths | Record strengths specifically (Step 14) |

## 12. AI guardrails

Skill-specific. The universal prohibitions in K4 apply in full and are not repeated.

1. **These skills assist an academic's judgement, they do not replace it.** Assessment, examination and progression decisions are the responsibility of the named academic and cannot be delegated to an AI system. The skill never assigns a mark, never produces feedback to be passed to a student unread, and never makes a progression or examination recommendation the academic has not independently reached. Institutional regulations and marking criteria take precedence over anything in this skill, and where a student's work may have integrity concerns, that is an institutional process and not a judgement this skill makes.

2. **This skill identifies methodological concerns and does not make allegations.** Never state or imply that a researcher fabricated, falsified, manipulated or concealed anything. Never characterise intent. Every analytic flexibility observation is written as a located observation about the report, with the innocent explanation stated where one exists and the material that would resolve it named. The distinction between a questionable research practice and misconduct is stated in the output, not assumed to be understood.

3. **Never audit a method you have not read, and never cite a location you have not verified.** Where materials are partial, the audit covers what is present and names what is absent. A methodological criticism of something the study does not contain destroys the audit and damages the auditor.

4. **Never reconstruct the probable method.** If the report does not say how cases were excluded, the finding is that exclusions are not reported. It is not an inference about what was likely done, however conventional the likely answer.

5. **Never recalculate a statistic and present the recalculation as the study's** unless the underlying data were supplied and the calculation was actually run (K4 §2.1). Internal consistency checks on reported figures are legitimate and are reported as consistency observations, not as corrected results.

6. **Settle the policy and data position before the material is opened.** Audited material is frequently unpublished, under review, belonging to an identified colleague or student, and may contain participant data. Institutional policy governs whether an AI system may touch it. Where the position is unknown, do not upload it.

7. **Never issue an overall score, grade, verdict or pass judgement.** The audit reports located findings with severities. Any summary judgement is made by the academic from those findings.

8. **Never let the study's findings enter the methodological assessment.** If a severity rating would change on learning what the study concluded, the rating is wrong. Where blinding was not achievable, disclose that in the scope note.

9. **Distinguish field convention from defect explicitly, and mark what you cannot assess** (K3 §6, coverage uncertainty). An audit conducted from general methodology against a specialist design must say so rather than generating confident findings about a domain it does not know.

10. **Never present the absence of a reported detail as evidence about the researchers.** It is evidence about the report. The two are conflated constantly and the conflation is where audits turn into accusations.

## 13. Best-practice principles

- **Audit the method before you know the result, and if you cannot, say so.** This one habit changes more audit outcomes than any technique here, because it removes the mechanism by which method problems get found selectively.
- **Work bottom to top.** A defect at a lower layer caps everything above it, and a beautiful analysis of data from an unanswerable design is worth nothing.
- **The alignment checks find what the layer checks miss.** Most studies that are locally sound and substantively wrong fail on design-to-question or conclusion-to-result.
- **Read the exclusion cascade like a balance sheet.** Numbers that do not reconcile are the most informative thing in many methods sections, and they are visible only if you add them up.
- **Missing statistics are findings.** An effect without a magnitude, an interval or a denominator is uncheckable, and uncheckable is a methodological property, not a formatting preference.
- **Name questionable practices plainly and never name intent.** Insinuation is worse than plain description. "The primary outcome named in the methods is not reported in the results" is plain, located and neutral.
- **The absence of a protocol is a fact, not a suspicion.** In many fields it is normal. Record it as the reason a flexibility question cannot be settled, and stop there.
- **Qualitative rigour lives in described practice.** The vocabulary can be produced by anyone; the coding process, the negative cases and the audit trail cannot.
- **Severity is about consequences for conclusions.** A study with many minor findings and no critical ones is a sound study reported imperfectly, and saying so is more useful than a long list.
- **Record strengths, specifically.** An audit that only finds fault is uncalibrated, and authors reasonably discount it.
- **Apply the reversal test to your own findings.** If you would not have raised it against a study you liked, it is not a finding.
- **The most valuable audit is the one you run on your own work before submission,** and the hardest, because blinding is impossible and the reversal test has to do all the work.

## 14. Worked example

Generic fictional scenario, academic, health services research.

**INPUT**

A departmental research quality panel asks a methodologist from another group to audit a completed study before it becomes the basis of a grant application. The study evaluates whether a redesigned appointment reminder process reduces missed appointments across four clinics. The auditor receives the full manuscript, the reminder protocol and the ethics submission. No pre-registration or analysis plan exists, and the field does not routinely require one for service evaluations.

**PROCESS**

*Step 1.* Scope note written: general methodological principles plus completeness of statistical reporting, reading order fixed, no prior involvement or competing programme. The abstract's conclusion is deliberately not read until Step 8, which requires covering it on the first page.

*Steps 2 to 3.* The question is stated once, at the end of the introduction, as whether the redesigned process "improves attendance". The construct is not defined: attendance could be the missed-appointment rate, the attended-appointment count, or cancellations converted to rebookings, and these move differently. The design is labelled a "controlled before-and-after study", but the described procedure is that the new process was introduced at two clinics and not the other two, with no randomisation, no matching, and one period before and one after. The label implies more control than the procedure delivers. What it can support is a difference in change between two non-equivalent groups of clinics, subject to confounding by anything else that differed between them.

*Step 4.* Sampling. The four clinics are the units of allocation; individual appointments are the unit of analysis, and clustering is not accounted for. Exclusions appear in the results rather than the methods: appointments booked less than 48 hours in advance were excluded, the number removed is not stated, and short-notice appointments are plausibly related both to the intervention and the outcome.

*Steps 5 to 7.* Measurement comes from the clinics' administrative records, which is appropriate and direct. Procedure has a gap: the before period includes a two-week clinic closure at one intervention site, mentioned once in a footnote. The analysis compares missed-appointment proportions before and after within each arm, with no adjustment for clustering, no covariates, and no interaction test between arm and period, which means the study's central comparison, whether the change differed between arms, is never tested.

*Step 8.* Results and conclusions read for the first time. The abstract states the redesign "reduced missed appointments by 22%". The results show a fall in the intervention arm and a smaller fall in the control arm, each tested separately.

*The judgement call.* The auditor's first instinct is to lead with the clustering error, the most technical item and the one a statistician's audit would centre. Step 9 reorders the report. Design-to-question is answerable with heavy qualification. Sample-to-claim fails on the undocumented exclusion. Analysis-to-design fails on clustering. But conclusion-to-result fails hardest and most simply: the claim of an effect rests on a between-arm difference that was never tested, and the 22% is a within-arm change presented as an intervention effect. That requires no statistical sophistication to see, and an audit beginning at the analysis would have missed it. Clustering becomes a major finding rather than the critical one, because even a correctly clustered version of the tests performed would not support the claim.

*Steps 10 to 12.* Reproducibility: the reminder protocol is present and detailed; the exclusion rule is present but unquantified; the analysis specification is partial. Reporting completeness: no effect size, no confidence intervals, denominators present, and the arm ns do not sum to the total appointments stated in the sample description, a discrepancy of about 3%, located and reported as an inconsistency requiring explanation. Flexibility: one signal only, the exclusion criterion introduced in the results section. The innocent explanation is stated, that a 48-hour rule is a routine operational decision in this setting, and the resolving material named, being the count removed per arm and period. No protocol exists to compare against, so the audit records that the timing of the decision cannot be established and does not speculate.

*Step 14.* One critical finding (conclusion not supported by the comparison performed), three major (clustering, undocumented exclusion, unexplained n discrepancy), four minor (construct undefined, closure not handled, no effect sizes or intervals, analysis specification partial), two observations. Strengths recorded: the intervention protocol is unusually well documented and directly reproducible, and the outcome measure is administrative rather than self-reported.

**OUTPUT**

An audit report with the four alignment answers at the front, ten located findings with severities and resolution conditions, a reproducibility table, a completeness list, one flexibility observation with its innocent explanation, and two named strengths. No score, and no recommendation about the grant application. `RESEARCHER DECISION REQUIRED` on whether the n discrepancy is pursued further, which is the panel's judgement and the institution's process (K5 §2.4).

## 15. Advanced usage

**Auditing at scale for a systematic review.** Hold the protocol constant and record the same fields for each study, so a quality distribution can be described rather than asserted. Resist collapsing to a single score for pooling decisions: carry the layer-level findings, because a literature uniformly weak at one layer, for example uniformly unclustered or reliant on one measure, is a finding about the field that a per-study score destroys.

**Auditing your own work before submission.** The highest-value application and the hardest, because blinding is impossible. Make the reversal test formal: for each finding you decide not to raise, write one sentence on why it would not have been raised against someone else's study. Better still, swap studies with a colleague, which restores partial blinding at no cost.

**Auditing when the protocol exists.** Step 12 becomes a comparison rather than a signal-scan: outcomes listed against outcomes reported, analyses planned against analyses performed, sample size intended against achieved, exclusion rules stated against applied. Report differences factually with both versions quoted. Differences are common and often legitimate; the finding is whether they are disclosed, not whether they occurred.

**Where the standard approach does not fit.** For mixed-methods work, run the stack separately on each strand and add a fifth alignment check on integration, which is where such studies most often fail. For secondary data analysis, the sampling and measurement layers become an audit of the source dataset's own properties and of the fit between what it measured and what is now claimed. For fields with mandatory reporting guidelines, audit against the guideline as well as the stack, reporting guideline items separately so the authors can act on them directly.

## 16. Skill chain

**Recommended previous skills:**
- **15.08 Research Design and Methodology Chapter** and **15.16 Doctoral Methodology Justification and Rigour.** Supply the standards a design and its justification are audited against.
- **15.10 Data Analysis Chapter Development.** Supplies the reporting completeness expectations applied in Step 11.

**Recommended next skills:**
- **15.21 Thesis Examination and Marking.** Takes the audit's methodological facts and converts them into criterion judgements and an outcome, which the audit itself does not make.
- **15.11 Discussion and Contribution Development.** Where the audit finds conclusions overreaching their results, this is where the claims are rebuilt to what the study supports.
- **15.24 Journal Article Development.** Where a pre-submission audit's findings are resolved before the paper goes out.

**Runs well alongside:**
- **13.01 Research Quality Review**, its commercial counterpart, which audits a deliverable and its usefulness where this skill audits the study beneath it.
- **13.03 AI Output Verification**, where part of the analysis or write-up was AI-assisted.
- **15.06 Systematic Literature Review** and **15.14 Advanced Systematic Review and Meta-Analysis**, where this skill supplies the quality appraisal of included studies.

---
A Yazi Supplied Skill and resource.
