---
name: advanced-systematic-review-and-meta-analysis
description: >
  Runs a systematic review to reporting-guideline standard with a registered
  protocol, dual independent screening, risk-of-bias assessment and a defensible
  decision about whether quantitative pooling is possible at all. Use for
  "systematic review protocol", "PRISMA flow diagram", "meta-analysis", "pool
  effect sizes", "fixed or random effects", "heterogeneity I-squared",
  "publication bias funnel plot", "risk of bias assessment", "can I meta-analyse
  these studies", "meta-ethnography", "thematic synthesis", "certainty of
  evidence", "dual screening kappa".
category: 15 Academic University Research
ref: "15.14"
tier: 3
inherits: [K2, K3, K4, K5]
---

# Advanced Systematic Review and Meta-Analysis

## 1. One-line description

A protocol-driven method for synthesising a body of primary research at doctoral and publication standard, covering reproducible searching, dual independent screening, risk-of-bias assessment, quantitative pooling where the evidence permits it, qualitative synthesis where it does not, and an explicit judgement about how much certainty the result deserves.

## 2. What this skill is used for

**The research problem it solves.** A systematic review is a study whose participants are studies, and it fails in the same ways a badly run primary study fails: an unstated protocol that lets decisions be made after seeing the results, a search that cannot be repeated, screening by one person whose judgement nobody checked, and a pooled estimate produced because the software would produce one. The most damaging failure is specific to this method: combining studies that measure different constructs, in different populations, with different comparators, produces a number with a narrow confidence interval that is precisely wrong, and it is more persuasive than any of the individual studies it came from. This skill installs the protocol first, keeps the search reproducible to the string, makes screening agreement measurable, assesses bias against the designs actually included, and treats the decision about whether to pool as the most consequential judgement in the review rather than a formality.

**Where it sits in the research lifecycle.** Usually as a standalone study, frequently as a doctoral chapter or a published paper in its own right, and often as the work that establishes what the rest of a thesis must add. It also runs before a primary study to establish what is already settled, and after one to place the result in the accumulated evidence.

**Typical use cases.**
- A doctoral literature chapter that must meet reporting-guideline standard rather than narrative standard.
- Establishing the pooled magnitude of an effect where enough comparable studies exist.
- Establishing that pooling is not possible, and saying so with the analysis that demonstrates it.
- Synthesising qualitative studies to build an interpretation that no single study reached.
- Providing an evidence base for a guideline, a policy submission or a funding case.
- Mapping heterogeneity in a literature as a finding about how a field measures its constructs.

**Who uses it.** Doctoral candidates in health, education, psychology, management, social policy and any field with an accumulating empirical literature; postdoctoral researchers and methodologists; review teams preparing submissions to guideline bodies. It assumes a working knowledge of research design and, for the meta-analytic sections, of effect sizes and sampling variance.

## 3. When to use it

- The review will be examined, peer reviewed, or used to support a decision by someone who did not run it.
- A body of primary studies exists that addresses a common question, and the question is whether they agree.
- You need to know the magnitude of an effect rather than only its direction.
- The literature contains apparently contradictory results and you need to establish whether the contradiction is real.
- A protocol must be registered before searching, because the field or the journal expects it.
- The review must be reproducible: someone else running your search should retrieve your set.
- You are synthesising qualitative studies and need a defensible method rather than a summary of summaries.
- A guideline, commissioner or examiner requires a formal certainty-of-evidence judgement.

## 4. When NOT to use it

- **The question is exploratory, the timeline is short, or the purpose is orientation.** A purposive, documented review answers "what is broadly known" far more efficiently and must not be dressed as systematic. Use **10.01 Literature Review and Desk Research**. Retrofitting the systematic label onto a purposive search is a misrepresentation an examiner or reviewer will detect from the log.
- **Boundary with 15.06 Systematic Literature Review.** 15.06 runs a structured review at masters standard: a documented search, explicit inclusion criteria, transparent screening and a narrative synthesis, normally by a single researcher. This skill takes over where the review must meet publication or doctoral standard: a registered protocol, a peer-reviewable search strategy reported in full, two independent screeners with agreement reported, formal risk-of-bias instruments matched to the included designs, quantitative pooling, and a certainty judgement. In the other direction, this skill does not replace 15.06 for a review that will not be pooled, will not be registered and is not being examined as a methodological artefact; the additional machinery costs months and buys nothing there. If you are unsure which applies, the deciding question is whether a reader must be able to reproduce the review, not whether the topic is important.
- **Fewer than a handful of comparable studies exist.** A pooled estimate from two or three small studies is unstable, and its confidence interval understates the uncertainty because between-study variance cannot be estimated from that few. Report them individually with their designs and limitations, and say why pooling was not attempted.
- **The studies measure different constructs under the same label.** This is the central prohibition of the method. Where "engagement", "resilience", "adherence" or "wellbeing" is operationalised incompatibly across the set, pooling produces a precise answer to no question. The correct output is the heterogeneity itself, reported as a finding about the field's measurement practice.
- **The evidence base is dominated by studies too poorly reported to appraise.** Risk-of-bias assessment requires enough method reporting to judge. Where most studies cannot be assessed, the review's conclusion is about the state of reporting in the field, and no pooled estimate should carry weight.
- **The question needs primary data.** Synthesis cannot answer what nobody has measured. Where the absent studies are the ones you need, the output is a gap specification and a primary research design. See **10.04 Research Gap Identification** and **01.04 Research Method Selection**.
- **Causal inference is required and the included designs are observational.** Pooling observational studies yields a more precise estimate of an association, not evidence of causation (K4 §3.2). Confounding does not average out across studies; where all studies share a confounder, the meta-analysis inherits it with a narrower interval.
- **Academic integrity.** Per §12.1: this skill does not write review text for submission as the candidate's own unaided output, and it never generates study records, extracted data, effect sizes or citations. Where an institution prohibits AI assistance for a task, this skill must not be used for it.

## 5. Required inputs

**Required.**
- **A specific, structured review question** naming the population, the exposure, intervention or phenomenon of interest, the comparator where one exists, and the outcomes. A question that cannot be structured cannot generate eligibility criteria, and without eligibility criteria screening is opinion.
- **Access to the databases and sources appropriate to the field, and the ability to open full texts.** A review conducted on abstracts is not a review (K4 §6.4).
- **The reporting guideline and any registration requirement the field expects**, or an explicit statement that none applies. These vary by discipline and by review type.
- **At least one second screener, or an explicit statement that screening is single with the limitation declared.** Dual screening is a design feature, not a courtesy, and its absence changes what the review can claim.

**Optional, and what each one adds.**
- **A registered protocol.** Converts every later analytical choice into a pre-specified one, which is the difference between a subgroup analysis and a fishing expedition. Deviations are then reportable as deviations rather than invisible.
- **An information specialist or librarian review of the search strategy.** Catches the missing synonym, the wrong field tag and the misapplied controlled vocabulary term, any of which silently removes part of the evidence base.
- **Statistical software and the raw effect size data.** Needed for pooling, prediction intervals, meta-regression and sensitivity analyses. Without it the review is narrative regardless of intent.
- **Contact details for primary authors.** Missing standard deviations, unreported subgroup results and clarifications about overlapping samples are frequently obtainable by asking, and doing so is expected in high-standard reviews.
- **A prior review on the same question.** Establishes whether an update is what is needed, and lets you report what has changed rather than repeating what has not.
- **A theoretical framework for a qualitative synthesis.** Determines whether framework synthesis is appropriate and what the a priori framework is.

## 6. Questions to ask before starting

1. **Is the aim to estimate a magnitude, to establish whether an effect exists, to map a literature, or to build an interpretation?** These are four different reviews with different methods, and choosing the method before the aim is the commonest structural error. Default: derive the aim from the review question and state it explicitly.
2. **What would make pooling inappropriate here?** Ask before the data is in, so that the answer is not shaped by whether the pooled result is attractive. Default: pre-specify the minimum construct comparability, design comparability and study count required, and hold to them.
3. **Will screening be dual and independent, and how will disagreement be resolved?** Determines whether an agreement statistic can be reported and how much unmeasured selection sits in the included set. Default: dual screening on all titles and abstracts, with a named arbitration route.
4. **Which risk-of-bias instrument matches the designs likely to be included?** Randomised, non-randomised, observational, diagnostic and qualitative designs each need different appraisal, and applying one instrument to all of them produces meaningless scores. Default: select per design and state the mapping.
5. **What is the unit of analysis?** Studies, reports, samples, comparisons and effect sizes are different things. Multiple reports of one sample double-count; multiple outcomes within one study create dependency. Default: define the unit as the independent sample and record the report-to-study mapping.
6. **What date, language and publication-status limits apply, and what do they cost?** Every limit is a coverage decision with a bias consequence. Default: no language limit where translation is feasible, no exclusion of unpublished work, and both limitations declared where they are applied (K3 §6).
7. **What is the plan if heterogeneity is high?** Deciding afterwards invites the analysis that makes it disappear. Default: pre-specify the subgroup and sensitivity analyses, and commit to reporting heterogeneity as a result rather than treating it as a problem to remove.

## 7. Step-by-step methodology

**Step 1. Write and register the protocol before searching.**
The protocol states the question, the eligibility criteria in full, the information sources, a draft search strategy for at least the primary database, the screening process and number of screeners, the data to be extracted, the risk-of-bias instruments, the planned synthesis including whether meta-analysis is intended and under what conditions, the pre-specified subgroup and sensitivity analyses, and the certainty framework. Registration fixes the decisions before the results are visible; without it every later choice is unfalsifiably contaminated by the results it produced. Where the field has no registration route, deposit a dated protocol somewhere immutable and say where. *Correct result: a dated protocol that a stranger could execute, with eligibility criteria specific enough that two people would apply them the same way.*

**Step 2. Build the search so it can be re-run, not merely described.**
For each concept, assemble free-text synonyms including spelling variants, historical terminology and the terms used by adjacent disciplines, then add each database's controlled vocabulary with its hierarchy handled explicitly. Combine within concepts with OR and across concepts with AND, and avoid filters that trade recall for precision unless they are validated. Then supplement, because databases alone miss material systematically: reference-list checking, forward citation tracking, hand-searching key journals and proceedings, trial and protocol registries, thesis repositories, and contact with active researchers. Record for every source the exact string, the interface, the date run and the number of records. The test is not whether a strategy sounds thorough; it is whether a reader can paste it into the same database and retrieve the same number. *Correct result: full strategies per database reproduced verbatim in an appendix, plus a supplementary search record, with total records and duplicates removed.*

**Step 3. Screen in two independent passes and report the agreement.**
Two people screen titles and abstracts against the eligibility criteria independently and blind to each other's decisions, then reconcile. Report the agreement before reconciliation using a chance-corrected statistic, and report it honestly: low agreement is diagnostic rather than embarrassing, and it usually means the criteria are underspecified. Fix the criteria and re-screen rather than arbitrating case by case, since arbitration hides the underspecification and makes the review unrepeatable. Full texts are then screened by both, with the reason for every exclusion recorded in categories tied to the criteria. Full-text exclusions with reasons are reportable and expected; title-abstract exclusions are counted, not itemised. *Correct result: screening counts at each stage, an agreement statistic with its interpretation, and a full-text exclusion list with categorical reasons.*

**Step 4. Build the flow diagram and make the numbers reconcile.**
Records identified per source, duplicates removed, records screened and excluded, full texts sought, not retrieved, assessed, and excluded with reasons, then studies included and reports included. The last two are not the same count: one study may have five reports, and five reports may describe one sample. The numbers must add up in every direction, and the commonest published error in systematic reviews is a flow diagram that does not. Reconcile it before writing anything else, because a diagram that does not balance means a record has been lost or double-counted. *Correct result: a flow diagram whose arithmetic closes, and an explicit statement of the study-to-report mapping.*

**Step 5. Extract into a pre-specified form, in duplicate, from the paper.**
Pilot the form first; forms always change on contact with real papers, and changing one mid-extraction without going back is how inconsistent data enters. Extract study identification, design, setting, population and eligibility, sample size and attrition, the intervention or exposure and its comparator, the outcome and exactly how it was measured, the timing, the analysis, the results with their variability, and funding. Two extractors, independently, with discrepancies resolved against the paper rather than by discussion. Every figure carries a locator (K2 §4.3). Where a needed statistic is absent, record it as absent and contact the author; never derive it by assumption. *Correct result: a populated extraction table with per-cell locators, a discrepancy log, and a list of items sought from authors.*

**Step 6. Assess risk of bias against the design, per outcome, not per study.**
Choose the instrument that matches each design and apply it at the level it specifies. Judge domains, not a total score: summing bias domains implies they are exchangeable, which they are not, and a study with a fatal flaw in one domain is not rescued by strength in four others. Record each judgement's reason with the supporting text from the paper, so the assessment is inspectable. Risk of bias applies to the result being used, so a study can be low risk for one outcome and high for another, typically where one was objectively measured and another self-reported by unblinded participants. Two assessors, independently. *Correct result: domain-level judgements per study per outcome, each with a written justification, and a summary figure showing the pattern across the set.*

**Step 7. Decide whether to pool, and treat this as the review's central judgement.**
Pooling is appropriate when studies address the same question, in populations similar enough that a common effect is a meaningful concept, with comparable interventions or exposures, comparable comparators, and outcomes measuring the same construct even if on different scales. Test each substantively, before looking at statistics: heterogeneity statistics diagnose inconsistency in results, they do not license combining incomparable things. The most damaging error in this method is the review that pools studies of different constructs and reports a tight interval, because the precision is real and the meaning is not. Where pooling fails on any criterion, do a structured narrative synthesis, which is a method and not a fallback: group by comparability, tabulate direction and magnitude, and examine patterns against study characteristics. `RESEARCHER DECISION REQUIRED` where any comparability criterion is borderline (K5 §2.7). *Correct result: an explicit, argued pooling decision with the criteria applied one by one, made before any pooled estimate is computed.*

**Step 8. Extract and convert effect sizes on a common metric.**
Select the effect size that matches the outcome type and the question: a standardised mean difference where scales differ, an unstandardised difference where the scale is meaningful and shared, a risk ratio or odds ratio for binary outcomes with the choice stated and consistent, a correlation for association questions. Convert using established formulae and record every conversion with its inputs so it can be checked, applying small-sample corrections where they apply. Handle dependency explicitly: multiple outcomes from one sample, multiple time points and multi-arm trials all violate independence, and the options are a pre-specified selection rule, averaging within study, or a model accounting for the dependency structure. Choosing silently is the error, not choosing wrongly. *Correct result: one effect size and variance per independent unit, a documented conversion trail, and a stated rule for dependency.*

**Step 9. Choose the model on what it assumes, not on what it produces.**
A fixed-effect model assumes every study estimates one common true effect and that all variation between them is sampling error, which is defensible only for near-replicates. A random-effects model assumes true effects vary across studies and estimates the mean of that distribution, a different quantity from a common effect, and it weights small studies more heavily, which matters when small studies are systematically different. Choose from the substantive question and the design of the set, not from which gives a narrower interval. Report which was used and why, and where the choice is arguable report both. Report a prediction interval alongside a random-effects mean: it states the range within which a future study's true effect is expected to fall, is frequently much wider than the confidence interval, and is exactly the information a reader needs and rarely gets. *Correct result: a stated model with its assumption named, the pooled estimate with its confidence interval, and a prediction interval where random effects are used.*

**Step 10. Assess heterogeneity, and read it as information.**
Report the between-study variance and the proportion of observed variation attributable to real differences rather than chance, with its uncertainty, and inspect the forest plot rather than relying on a single index. Then interpret. High heterogeneity is a finding about the literature: the effect depends on something, and the useful work is identifying what. Investigate through pre-specified subgroup analyses and meta-regression where enough studies exist, remembering these are observational comparisons between studies and that a subgroup difference is a hypothesis, not a result (K4 §3.4). Never remove studies to reduce heterogeneity, and never present a pooled estimate from a highly heterogeneous set as though it described a real population. If the set is too inconsistent to summarise with one number, say so, and the review's contribution becomes the account of what the inconsistency depends on. *Correct result: heterogeneity quantified and interpreted, pre-specified investigations reported whether or not they explained anything, and no post-hoc exclusions.*

**Step 11. Assess reporting bias, and state the limits of the assessment.**
Funnel plot asymmetry and its associated tests are the standard tools and they are weak: they need a reasonable number of studies to have any power, asymmetry has causes other than publication bias including true heterogeneity and quality-driven small-study effects, and adjustment methods rest on assumptions that cannot be checked. Use them, and report them as one indirect line of evidence. The stronger evidence is structural: published results compared against registered protocols and registry entries, the presence or absence of grey literature in the included set, whether reported outcomes differ from pre-specified ones, and whether the search covered unpublished work at all. *Correct result: a reporting-bias assessment that names its own limitations and does not present a small-study test as a verdict.*

**Step 12. Run sensitivity analyses that could change the conclusion.**
Pre-specify them: excluding high risk-of-bias studies, excluding the largest and smallest, alternative effect metrics, alternative models, alternative handling of dependency, and alternative decisions where an extraction judgement was contestable. The purpose is to establish whether the conclusion survives reasonable alternative choices. Report every one, especially those that change the answer. A sensitivity analysis reported only when it confirms the main result is not a sensitivity analysis (K4 §4.2). *Correct result: a table of analyses against results, stating which conclusions are robust and which are contingent on a specific choice.*

**Step 13. Where the evidence is qualitative, choose the synthesis approach by its aim.**
These are not interchangeable. Meta-ethnography translates studies' concepts into one another to produce an interpretation beyond any single study, and suits conceptually rich, few studies. Thematic synthesis codes findings across studies inductively, building descriptive then analytical themes, and suits a defined question across a larger set. Framework synthesis maps findings onto an a priori framework, adapting it as the data requires, and suits applied work where a framework exists. Each carries different commitments about whether findings from different traditions can legitimately be combined, and that question has no agreed answer: some traditions hold that decontextualising an interpretation destroys it. State the position taken and why. Appraise included studies to inform the weight given to each rather than to exclude mechanically, since checklist-based exclusion of qualitative work is contested. *Correct result: a named approach, a justification tied to the aim and the studies, and a synthesis whose interpretive steps are visible and traceable to study findings.*

**Step 14. Judge the certainty of the evidence, per outcome.**
A pooled estimate is not the conclusion; the conclusion is what the estimate licenses. Rate each main outcome down for risk of bias in the contributing studies, inconsistency across them, indirectness of population, intervention, comparator or outcome relative to the review question, imprecision, and suspected reporting bias. Rate up, where the design permits, for a large effect magnitude, a dose-response gradient, or plausible confounding working against the observed effect. Record the reason for every rating change. Then write the conclusion at the certainty the rating supports, in K3 language, and cap any downstream recommendation at that level (K3 §3.6). *Correct result: a summary-of-findings table giving, per outcome, the effect, the number of studies and participants, the certainty rating and the reasons for it.*

## 8. Analytical framework

The review is a chain in which each link constrains everything after it:

    Question → Protocol → Search → Screening → Included set →
    Risk of bias → Comparability judgement → Synthesis (pooled or structured) →
    Heterogeneity → Reporting bias → Sensitivity → Certainty → Conclusion

Two links carry most of the risk. **Search to included set** determines what the review can possibly see: a set assembled by a search that missed a vocabulary or excluded unpublished work is biased in a way no downstream analysis repairs. **Comparability judgement** determines whether the synthesis means anything: everything after it inherits the assumption that these studies were about the same thing.

A second structure governs interpretation of a pooled result:

    Precision of the estimate  ≠  Certainty of the evidence

Precision comes from sample size and study count. Certainty comes from bias, consistency, directness and reporting completeness. A narrow interval on a set of high-risk, heterogeneous, indirect studies is a precisely estimated summary of unreliable evidence, and reporting it without the certainty rating is the way meta-analysis most often misleads. The summary-of-findings table exists to keep these two things visibly separate.

## 9. Output format

**1. Structured abstract**, including the registration identifier or a statement that none exists.

**2. Methods.** Protocol and registration, eligibility criteria, information sources with dates, full search strategy for at least one database with the rest in an appendix, screening process and number of screeners, extraction process, risk-of-bias instruments per design, effect measures, synthesis methods including the pooling decision rule, and the certainty framework.

**3. Flow diagram**, with reconciled counts and the study-to-report mapping.

**4. Study characteristics table.**

| Study | Design | Country and setting | N and attrition | Population | Intervention or exposure | Comparator | Outcome and measure | Timing | Funding |
|---|---|---|---|---|---|---|---|---|---|

**5. Risk-of-bias summary.** Domain-level judgements per study per outcome, with justifications available.

**6. Synthesis.** Either forest plots with pooled estimates, model stated, heterogeneity and prediction intervals; or a structured narrative synthesis grouped by comparability with direction and magnitude tabulated. Never both for the same outcome without explaining why.

**7. Heterogeneity investigation.** Pre-specified subgroups and meta-regression, reported whether or not explanatory, with post-hoc analyses labelled as post-hoc.

**8. Reporting bias assessment**, with its limitations stated.

**9. Sensitivity analyses**, all of them, with the effect of each on the conclusion.

**10. Summary of findings table.**

| Outcome | Studies (participants) | Effect (95% CI) | Certainty | Reasons for downgrade or upgrade | Plain interpretation |
|---|---|---|---|---|---|

**11. Discussion.** What the evidence establishes, at what certainty, for whom, and what it does not establish. Limitations of the review itself, distinct from limitations of the included studies.

**12. Deviations from protocol**, listed, with reasons and their likely effect.

**When the evidence is thin, the format must not force fabrication (K4 §1).** A review that includes four studies reports four studies. Empty cells stay empty and are labelled not reported, never imputed by inference. Where pooling was planned and abandoned, the abandonment and its reason are the result, presented in the synthesis section rather than hidden in limitations. A review that concludes the evidence base cannot answer the question, and shows exactly why, is a complete and publishable review.

## 10. Quality checks

Run before anything is presented. These sit on top of K4 §8.

1. Was the protocol dated and fixed before searching, and is every deviation from it listed?
2. Is the full search strategy reproduced exactly enough that a reader could re-run it and retrieve the same records?
3. Does the search cover adjacent-discipline vocabulary, unpublished work and non-database sources, or are those limitations declared?
4. Was screening independent and duplicated, and is the agreement statistic reported before reconciliation?
5. Do the flow diagram numbers reconcile in every direction, and is the study-to-report mapping explicit?
6. Was every included study read in full, and does every extracted figure carry a locator?
7. Is risk of bias judged at domain level per outcome, with written justifications, and are no bias scores summed?
8. Was the pooling decision made on substantive comparability before any pooled estimate was computed?
9. Is the model choice justified by its assumption rather than by its result, and is a prediction interval reported with any random-effects estimate?
10. Was any study excluded to reduce heterogeneity? If so, that is a defect, not a sensitivity analysis.
11. Are subgroup and meta-regression findings labelled as observational and hypothesis-generating?
12. Are all sensitivity analyses reported, including those that change the conclusion?
13. Does the reporting-bias assessment state the limits of its own tools?
14. Does every conclusion carry a certainty rating, and is no conclusion stated more strongly than its rating supports?
15. For a qualitative synthesis, is the approach named, justified by its aim, and its epistemological position stated?

## 11. Common failure modes

| Failure | How to recognise it | How to prevent it |
|---|---|---|
| **Protocol written after the results** | Analyses look perfectly targeted at what was found | Register and date before searching (Step 1) |
| **Unreproducible search** | The strategy is described in prose rather than reproduced | Report exact strings, interfaces and dates (Step 2) |
| **Single screening presented as dual** | No agreement statistic anywhere in the paper | Report pre-reconciliation agreement (Step 3) |
| **Flow diagram that does not add up** | Included studies plus exclusions does not equal records screened | Reconcile before writing (Step 4) |
| **Double-counted samples** | Two reports of one cohort entered as two studies | Map reports to independent samples (Step 5) |
| **Bias scores summed** | Studies ranked by a total quality number | Judge domains, never totals (Step 6) |
| **Apples pooled with oranges** | A tight interval across studies measuring different constructs | Apply comparability criteria first (Step 7) |
| **Model chosen by result** | Fixed effects used because random effects lost significance | State the assumption that justifies the model (Step 9) |
| **Heterogeneity suppressed** | Studies dropped until the index falls | Report heterogeneity as a finding (Step 10) |
| **Prediction interval omitted** | A random-effects mean presented as if it were a common effect | Report both intervals (Step 9) |
| **Funnel plot as verdict** | "No publication bias was found" on eight studies | State the test's power limits (Step 11) |
| **Selective sensitivity reporting** | Only the confirmatory analyses appear | Pre-specify and report all (Step 12) |
| **Precision mistaken for certainty** | A narrow interval on high-risk indirect studies stated confidently | Separate the two in the summary table (Step 14) |
| **Qualitative synthesis without a method** | Themes across studies with no named approach | Name and justify the approach (Step 13) |
| **AI-generated study records** | Plausible studies, plausible effect sizes, unverifiable | Nothing enters the table unopened. See §12 |

## 12. AI guardrails

Skill-specific. The universal prohibitions in K4 apply in full and are not repeated.

1. **Academic integrity, and it binds this whole skill.** These skills assist a researcher's thinking, structure and rigour. They do not produce work to be submitted as the student's own unaided output. The user must comply with their institution's AI use policy and its declaration requirements, which vary by institution and by assessment. Where an institution prohibits AI assistance for a task, this skill must not be used for it. The skill never writes a passage for submission as though the student wrote it; it interrogates, structures, critiques and teaches. Operationally here: the skill may build the protocol structure, check search logic, critique eligibility criteria, verify that flow numbers reconcile and interrogate a pooling decision. It does not screen studies on the candidate's behalf as though a human had, and any AI involvement in screening, extraction or coding is disclosed in the methods section with a statement of what a human verified (K4 §7).

2. **Never produce a study, a citation, a sample size or an effect size from model knowledge.** Every row of the extraction table comes from a paper that has been opened. A fabricated study in a meta-analysis contaminates the pooled estimate, the heterogeneity assessment and the conclusion simultaneously, and it is undetectable to every downstream reader.

3. **Never impute a missing statistic without saying so.** Deriving a standard deviation from a confidence interval or a standard error is a legitimate documented conversion. Assuming one from a similar study, from a typical value, or from a plausible range is fabrication (K4 §2.1). Where a statistic is unavailable, the study is reported without it, or excluded from that analysis with the exclusion counted.

4. **Never compute a pooled estimate on data you have not been given.** If the effect sizes are not supplied or extractable from supplied papers, no meta-analysis exists. Do not demonstrate the method on invented numbers unless they are labelled as an illustration in the same sentence and kept entirely out of the results.

5. **Never present a pooled estimate without the pooling decision that licensed it.** A number produced because the format expects a number is exactly the failure K4 §1 describes, and in this method it is more persuasive than in any other.

6. **Never describe a search you did not run.** Do not report databases consulted, hit counts, or date ranges searched unless they were. An unsearched source is logged as not searched, with the reason.

7. **Never resolve screening disagreement by choosing the more inclusive or the more convenient answer.** Disagreement is data about the criteria. Report it and fix the criteria.

8. **Never report heterogeneity as a problem that was solved.** If it was reduced, say what was removed and why; if it remains, report it. Presenting a homogeneous-looking result obtained by exclusion is a misrepresentation of the literature.

9. **Cap the certainty of any conclusion at the level the formal rating supports, and never let a precise interval upgrade it** (K3 §3.4). Where risk of bias is high across the contributing studies, no amount of pooled precision produces a high-certainty conclusion.

10. **Never state that a qualitative synthesis approach is the standard one.** These traditions disagree with each other, sometimes fundamentally. Present the options with their aims and commitments and let the researcher choose.

## 13. Best-practice principles

- **The protocol is the review's control condition.** Everything decided after the results are visible is unfalsifiably contaminated, and the only defence against that suspicion is a dated prior record.
- **A search is judged by reproducibility, not by effort.** "We searched five databases comprehensively" tells a reader nothing. The strings, the interfaces and the dates tell them everything.
- **Screening disagreement is a measurement, not a nuisance.** Low agreement almost always means the eligibility criteria are underspecified, and fixing the criteria is faster than arbitrating fifty cases.
- **The unit of analysis is the independent sample, not the paper.** Multi-report studies and multi-outcome papers are the most common source of silently inflated evidence.
- **Ask whether pooling is meaningful before asking whether it is statistically permissible.** Statistical heterogeneity measures inconsistency of results; it cannot tell you that two studies measured different constructs.
- **High heterogeneity is the interesting result.** It says the effect depends on something. A field that produces wildly inconsistent estimates of the same quantity has a measurement problem, and documenting it is a real contribution.
- **Report the prediction interval.** A confidence interval around a random-effects mean describes uncertainty about an average; readers almost always want to know what to expect in a new setting, and only the prediction interval answers that.
- **Small-study effects have several causes and publication bias is only one.** Smaller studies are also often lower quality, conducted in more selected populations and more intensively delivered. Asymmetry is a prompt to investigate, not a diagnosis.
- **Risk of bias attaches to a result, not to a study.** The same trial can be low risk for mortality and high risk for a self-reported symptom score.
- **Contact the authors.** Missing data is routinely obtainable, and reviews that ask get a materially more complete evidence base than reviews that do not.
- **A review that finds the literature cannot answer the question has answered a question.** It saves the field from a synthesis that would have looked authoritative, and it specifies what the next primary study must do.
- **Update rather than repeat.** Where a competent review exists, establish what has changed since, and report the difference. Re-running an unchanged review is work without contribution.

## 14. Worked example

Generic fictional scenario, academic, public health and education.

**INPUT**

A doctoral candidate wants to establish whether school-based sleep education programmes improve adolescent sleep duration. She has found 31 studies and her supervisor has asked for a meta-analysis. Her draft already contains a pooled estimate of a 22-minute improvement with a narrow interval.

**PROCESS**

*Steps 1 to 3.* The protocol is written retrospectively, which is noted as a limitation and registered before the final search is re-run so that at least the search and synthesis decisions are pre-specified. Eligibility criteria are tightened after the first dual-screening pass returns poor agreement: "school-based" had been left undefined, and the two screeners were treating after-school programmes differently. The criteria are amended, screening is repeated, and agreement rises to an acceptable level. The amendment and its reason are recorded as a protocol deviation.

*Steps 4 and 5.* The flow diagram initially does not reconcile. Investigation shows two pairs of papers reporting the same cohort at different follow-up points, counted as four studies. The corrected count is 29 reports of 27 independent samples.

*Step 6.* Risk of bias is assessed by design. Nineteen studies are non-randomised, most with no adjustment for baseline differences. The outcome domain matters: eleven studies measure sleep by self-report diary in unblinded participants who knew they had received sleep education, which is high risk for that outcome, while six use actigraphy, which is not.

*The judgement call.* The pooling decision. Sleep duration looks like a single construct, and the software will pool it happily. Applying the comparability criteria one by one shows it is not. Some studies measure school-night sleep, some average across the week including weekend catch-up sleep, which behaves differently. Some measure time in bed, some measure estimated sleep. Follow-up ranges from immediately post-programme to twelve months. The temptation is to pool everything and report subgroup analyses. The decision taken is to pool only the actigraphy-measured, school-night, post-programme outcomes, which is six studies, and to synthesise the rest structurally.

*Steps 9 to 12.* The six-study random-effects pooled estimate is roughly nine minutes with a wide interval crossing no effect, and a prediction interval wide enough to include a meaningful decrease. Heterogeneity remains substantial even in this restricted set. The original 22-minute figure came from pooling self-reported outcomes, which are systematically larger, and from including weekend-inclusive measures. Sensitivity analysis excluding the single largest study moves the estimate materially, which is reported.

*Steps 13 and 14.* Certainty for the primary outcome is rated low: downgraded for risk of bias in the wider set, for inconsistency, and for imprecision. The conclusion is written accordingly: objectively measured sleep duration shows a small and imprecisely estimated improvement immediately after school-based programmes, the larger effects reported in the literature come predominantly from unblinded self-report, and no study establishes maintenance beyond six months.

**OUTPUT**

A review whose headline is smaller and considerably more defensible than the draft, whose central contribution is the demonstration that the field's apparent effect is measurement-dependent, and whose gap statement specifies exactly what the next primary study should measure and for how long.

`RESEARCHER DECISION REQUIRED` was raised at the pooling decision and resolved by the candidate and supervisor jointly, and the decision and its reasoning are recorded in the methods (K5 §2.7).

## 15. Advanced usage

**Network meta-analysis.** Where multiple interventions have been compared in different combinations, indirect comparisons can be estimated across a connected network. It requires the transitivity assumption, which is that studies comparing A with B are similar enough to studies comparing B with C for the indirect comparison to be valid, and it requires checking consistency between direct and indirect evidence wherever both exist. It is a substantially harder method and should not be attempted because the network diagram looks impressive.

**Individual participant data synthesis.** Where the primary data can be obtained, it permits consistent handling of covariates, proper investigation of participant-level effect modification rather than the ecological subgroup comparison a study-level meta-regression provides, and standardised outcome definitions. It is the strongest form of synthesis and the most expensive in time and negotiation.

**Living reviews.** For fast-moving evidence bases, maintain the protocol and searches as a standing asset with scheduled re-runs, and report the diff: what is new, what is superseded, whether the pooled estimate has moved and whether any conclusion has changed. The discipline required is to pre-specify the updating rule, otherwise updating becomes selective.

**Mixed-methods synthesis.** Where both quantitative and qualitative evidence bears on the question, the two syntheses are conducted separately by their own standards and then integrated at the interpretation stage, typically by using the qualitative synthesis to explain heterogeneity in the quantitative one or to identify outcomes that matter to participants and were not measured. Do not merge the evidence streams before synthesis.

**Reviews of reviews.** Where multiple systematic reviews exist on overlapping questions, an overview can map their agreement, but it must handle overlapping primary studies explicitly, since the same trial appearing in five reviews creates a false impression of accumulation. Report the primary-study overlap matrix.

**When the standard approach does not fit.** In fields with small, conceptually diverse literatures, a full meta-analysis will rarely be available and the review's value lies in comparability mapping and in specifying what a synthesisable literature would require: agreed constructs, shared measures and reported variances. That specification is itself a methodological contribution.

## 16. Skill chain

**Recommended previous skills:**
- **15.13 Doctoral Research Positioning and Originality.** Hands over the contribution the review must make, and whether an update or a new review is warranted.
- **15.06 Systematic Literature Review.** Hands over a structured review that is being escalated to publication or examination standard.
- **10.01 Literature Review and Desk Research.** Hands over a scoping picture that establishes whether enough comparable primary studies exist to justify this method.

**Recommended next skills:**
- **15.15 Theory Building and Conceptual Contribution.** Takes a synthesis whose heterogeneity points to an unspecified moderator and builds the theory that explains it.
- **15.16 Doctoral Methodology Justification and Rigour.** Takes the gap the review establishes and defends the primary design that closes it.
- **15.18 Thesis Architecture and Chapter Coherence.** Places the review as a chapter and connects it to the empirical chapters.
- **15.24 Journal Article Development.** Takes the review to publication.

**Runs well alongside:**
- **15.03 Referencing and Citation Management**, which verifies the reference set a review of this size accumulates.
- **05.01 Statistical Analysis Planning**, for the effect size, model and dependency decisions.
- **13.02 Source and Citation Verification**, which audits the included set.

---
A Yazi Supplied Skill and resource.
