---
name: systematic-literature-review
description: >
  Runs a protocol-driven systematic literature review to academic standard:
  registered protocol, structured review question, explicit inclusion and
  exclusion criteria, documented multi-database search, two-stage screening
  with recorded reasons, a reconciling flow diagram, design-appropriate quality
  appraisal, structured extraction and defensible synthesis. Use for "systematic
  review", "do a SLR", "write my review protocol", "review flow diagram",
  "screening and extraction", "quality appraisal of included studies", "is a
  systematic review realistic for my dissertation", "scoping review versus
  systematic review", "what does the literature actually establish".
category: 15 Academic University Research
ref: "15.06"
tier: 3
inherits: [K2, K3, K4, K5]
---

# Systematic Literature Review

## 1. One-line description

A protocol-driven method for reviewing a body of literature so that the search, the screening, the appraisal and the synthesis are pre-specified, documented and reproducible, and so that the review reports what the literature establishes rather than what it repeats.

## 2. What this skill is used for

**The research problem it solves.** A narrative review is a reading of the literature. A systematic review is a study whose data are published studies, and it is judged as a study: a reader must be able to repeat the search and arrive at the same included set. Without a protocol, three failures are near certain. The search is shaped by what the reviewer already believes, because the inclusion decision is made after the study is read. The included set cannot be reconstructed, because the reasons for excluding forty of the fifty-two full texts were never written down. And the synthesis pools studies that measure different things on different populations, producing a confident conclusion no single included study supports. This skill installs the protocol, the audit trail, the appraisal step and the synthesis discipline. It also installs a proportionality gate, because the most expensive error at masters level is starting a full systematic review and executing it badly.

**Where it sits in the research lifecycle.** Early, and often as the study itself. Some masters dissertations are systematic reviews with no primary data collection. Where primary work follows, the review establishes what is known, defines the constructs, and specifies the gap the primary study addresses. It runs before **15.07 Theoretical and Conceptual Framework Development** and before **15.08 Research Design and Methodology Chapter**.

**Typical use cases.**
- A dissertation whose entire empirical contribution is a review of existing studies.
- Establishing, defensibly, that a gap exists rather than asserting it.
- Resolving a literature that appears to disagree, by appraising the studies rather than counting them.
- Building the evidence base for an intervention where primary trial work is out of reach.
- Producing a review chapter that will be examined and must survive a methodologist reader.
- Deciding, before committing, whether a full systematic review is feasible in the time available.

**Who uses it.** Masters candidates writing a review-based dissertation or a review chapter; supervisors checking a protocol before searching starts; students in health, education, management, engineering, social policy and information systems, where structured review methods are established and expected.

## 3. When to use it

- The review will be examined, and the examiner may ask how the included studies were selected.
- The research question is answerable from published studies, and the value is in the aggregation rather than in new data.
- The literature is large enough that unstructured reading would produce an arbitrary sample of it.
- Two or more bodies of work appear to contradict each other and the contradiction needs adjudicating.
- The department expects a documented search log, an inclusion protocol and a flow diagram, whatever it calls them.
- You need to state a gap in a proposal and be able to defend the claim that nobody has done this.
- A supervisor has asked for "a systematic review" and you need to establish whether that is what the timeline supports.

## 4. When NOT to use it

- **The timeline does not support it, which at masters level is common.** A defensible systematic review needs a pilot search, a protocol, dual screening of hundreds of records, full-text retrieval, appraisal and extraction. Where you have eight weeks, a single reviewer and no access to interlibrary loan, the honest options are a **scoping review** (broader question, no quality appraisal, maps what exists rather than what works), a **structured review** (protocol-driven search and explicit criteria, single reviewer, appraisal narrative rather than instrumented), or a **narrative review conducted transparently**. Each is respectable and each is examinable when named accurately. Attempting a full systematic review and delivering a partial one is penalised harder than delivering a well-executed scoping review. Decide at Step 1, not in week six. `RESEARCHER DECISION REQUIRED` (K5 §2.7).
- **The purpose is orientation rather than synthesis.** If you are still deciding what the topic is, a systematic search is premature because you cannot yet write inclusion criteria. Use **15.01 Academic Research Topic Selection**, then **15.02 Academic Literature Search Strategy** for the exploratory searching that precedes a protocol.
- **The task is the search itself rather than the whole review.** Building database-specific strings, controlled vocabulary, Boolean logic and a search log is **15.02 Academic Literature Search Strategy**. This skill consumes that output and adds protocol, screening, appraisal, extraction and synthesis. In the other direction: 15.02 stops at the retrieved record set and does not appraise or synthesise. Do not describe a 15.02 search as a systematic review.
- **The work is commercial or applied rather than examined.** Where the output feeds a business decision, exhaustiveness is usually the wrong trade and a purposive, documented review is better value. Use **10.01 Literature Review and Desk Research**, which appraises source quality on eight dimensions and is explicit that it is purposive, not systematic. In the other direction: 10.01 hands over to this skill wherever the work will be examined, peer reviewed, or must be reproducible by a third party, and 10.01's search log will not support a systematic claim retrofitted onto it.
- **You intend to pool results statistically.** Meta-analysis has its own requirements: effect size extraction, heterogeneity assessment, a defensible pooling model, publication bias analysis and sensitivity testing. That is **15.14 Advanced Systematic Review and Meta-Analysis**. This skill covers the review up to the point of synthesis and includes narrative and tabular synthesis. It does not license computing a pooled effect.
- **The evidence base is too small or too heterogeneous to synthesise.** Where the search returns four studies on four different populations with four different outcome measures, the correct output is a description of the state of the literature and the reasons it cannot be synthesised. Forcing a synthesis onto incommensurable studies is worse than reporting the fragmentation, which is itself a finding and often the strongest one available.
- **The topic is too new to have a literature.** Emerging technologies and recent policy changes often have conference material and preprints but no body of studies. A systematic review of six items is a report on absence dressed as a review. Say the literature does not yet exist, document the searches that establish this, and reconsider the design.
- **Your institution prohibits AI assistance for this task.** See §12 item 1. Where the assessment forbids it, this skill must not be used for it.

## 5. Required inputs

**Required.**
- **A review question specific enough to write inclusion criteria against.** "AI in education" is a topic. "Among secondary school teachers, what does empirical research report about the effects of automated feedback tools on marking workload?" is a question. If only a topic is supplied, ask. A review without a question has no inclusion criterion and therefore no protocol.
- **The institutional and disciplinary expectations for a review at this level.** Departments differ sharply on whether appraisal instruments are required, whether dual screening is expected, and what the review chapter is worth. Ask, or find the marking rubric and the departmental handbook. Do not assume a discipline.
- **Confirmed access to at least two bibliographic databases of the appropriate type, and to full texts.** A review that can retrieve records but not read them cannot proceed past screening.
- **A realistic date on which searching must stop and a realistic date on which the review must be complete.**

**Optional, and what each one adds.**
- **A second screener, even for a subset.** Turns an unverifiable single-reviewer judgement into a measurable one, and lets you report screening agreement. Where a full dual screen is impossible, a dual screen of 10 to 20 percent is a real quality claim and takes hours, not weeks.
- **A supervisor's sign-off on the protocol before searching.** Converts the protocol from a document you wrote into a commitment, which is the entire point of writing it first.
- **An existing published review on an adjacent question.** Gives you a tested search vocabulary, a candidate inclusion frame, and the studies it included, which are a seed set you can check your search against. If your search misses studies a published review found, your search is wrong.
- **Reference management with deduplication.** Deduplication by hand across four databases is the point at which the numbers stop reconciling.
- **A named appraisal instrument the department accepts.** Removes an argument in the viva about why you chose the instrument you chose.
- **A pre-specified list of grey literature sources.** Decided in advance, this is a scope decision. Decided later, it is a source of bias.

## 6. Questions to ask before starting

1. **Is this review the dissertation, or a chapter of it?** Determines the scale of everything. A review-as-dissertation justifies a large screening burden. A review chapter that must leave time for fieldwork does not. Default if unanswered: assume it is a chapter, and scope the protocol so it can be completed in a third of the available time.
2. **What review type does the department expect, and what does the rubric reward?** Systematic, scoping, integrative, rapid and structured reviews have different obligations. Default: name the type explicitly in the protocol and describe the method accurately rather than aspirationally.
3. **Is a second screener available, for all or part of the screening?** Determines whether you report agreement or declare single-reviewer screening as a limitation. Default: assume single reviewer, and build in a dual-screened subset.
4. **What study designs will the search return, and are they appraisable with one instrument?** A review that will include trials, cohort studies, surveys and qualitative work needs more than one instrument or an explicitly justified single approach. Default: expect design heterogeneity and plan for design-specific appraisal.
5. **Will the synthesis be narrative or quantitative?** This decides the extraction form. Extracting for a narrative synthesis and then deciding to pool means re-reading every study. Default: assume narrative, and record effect sizes anyway where they are reported, since capturing them costs nothing at extraction and cannot be recovered later.
6. **Are you registering the protocol, and does registration fit the timeline?** Prospective registration is strong practice and increasingly expected. Some registers have review queues measured in weeks. Default: write the protocol regardless, date it, lodge it with the supervisor, and register where the register and the timeline allow.
7. **What is the language and date window, and what is the justification?** These are inclusion criteria and they need reasons, not conveniences. "English only" is defensible when stated as a limitation with its coverage risk named, and indefensible when unstated. Default: state both, with the reason and the risk.

## 7. Step-by-step methodology

**Step 1. Run the proportionality gate before anything else.**
Estimate the workload from a pilot search, not from optimism. Run your draft question's core concepts in one database, unrestricted, and look at the number of records. Then apply this arithmetic: title and abstract screening runs at roughly 60 to 120 records an hour once criteria are stable, full-text retrieval and screening at 4 to 10 papers an hour, appraisal and extraction at 1 to 3 papers an hour. Multiply by two if dual screening. Compare the total against the hours you actually have alongside every other dissertation task. If the pilot returns 3,000 records and you have six weeks, the question is too broad or the review type is wrong. Narrow the question (tighter population, tighter outcome, tighter date window, each with a reason) or change the review type and say so. *Correct result: a written feasibility estimate with record counts, hourly rates, total hours and a decision. This paragraph belongs in the methodology chapter, because it is the justification for the scope.*

**Step 2. Write the protocol, in full, before the first real search.**
The protocol contains: the review question in structured form, the rationale, the inclusion and exclusion criteria, the sources to be searched, the draft search strategy for at least one database in full, the screening procedure and who does it, the appraisal instrument and how disagreements are resolved, the extraction fields, the planned synthesis approach, and the date. Register it where a suitable register exists for your field and the timeline allows; where it does not, date it, have the supervisor acknowledge receipt, and treat that as the record. The protocol's function is to make your later decisions checkable: a criterion applied after you have read the studies is indistinguishable from a preference. Deviations from protocol are permitted and normal. They are reported as deviations, with the reason and the date. *Correct result: a dated protocol document that a second reviewer could execute without asking you a question.*

**Step 3. Put the question into a structured form and derive the criteria from it.**
Use the structured question framework conventional in your discipline. Quantitative effectiveness questions typically decompose into population, intervention or exposure, comparator, outcome and study design. Qualitative and experiential questions decompose differently, typically into sample, phenomenon of interest, design, evaluation and research type. Context and time are usually added. Whichever frame you use, the point is that each element becomes an inclusion criterion with a decidable test. Write each criterion so that two people reading the same abstract reach the same decision: "adults" is not decidable, "participants aged 18 or over, or a study reporting adult subgroup results separately" is. Write the exclusions as the mirror of the inclusions plus the specific exclusions the topic needs (non-empirical commentary, conference abstracts without full papers, studies where the intervention is bundled and cannot be isolated). *Correct result: a criteria table with columns for element, inclusion, exclusion and the operational test, containing no criterion that requires judgement you have not defined.*

**Step 4. Build, pilot and freeze the search.**
For each concept, assemble free-text synonyms and the controlled vocabulary terms the database uses, since indexed vocabulary and author language diverge and searching only one loses studies silently. Combine synonyms with OR inside a concept, concepts with AND across the search, and use truncation and phrase operators as each database's syntax requires. Pilot the string against a seed set of three to six studies you know should be included: if the search does not return them, the search is wrong, and fixing it now is cheap. Adapt the string for each database rather than pasting it, because syntax and vocabulary differ and a pasted string usually silently returns almost nothing. Search a minimum of two, preferably three or more, databases of complementary type: a broad multidisciplinary citation index, at least one discipline-specific bibliographic database, and where the topic warrants it a source of theses, reports or trial registrations. Supplement with backward citation searching (the reference lists of included studies) and forward citation searching (what has cited them), both of which routinely surface studies the database search missed. Record for every database: the platform, the exact full string, any field and date limiters, the date run, and the number of records. Then stop searching. Re-running a search after screening has begun makes the numbers irreconcilable unless the re-run is documented as a separate dated search. *Correct result: a search appendix a reader could execute line by line, containing the full string for every database, not a summary of it.*

**Step 5. Deduplicate, then screen in two stages, recording reasons at the second.**
Import everything, deduplicate, and record the number removed as duplicates, because this number appears in the flow diagram. Stage one screens title and abstract against the criteria; screen liberally, since the cost of retrieving one unnecessary full text is far lower than the cost of losing an eligible study, and record only include or exclude at this stage. Stage two screens the full text, and here every exclusion carries a reason drawn from a short closed list derived from your criteria (wrong population, wrong intervention, wrong outcome, wrong design, not empirical, duplicate report of an included study, full text unobtainable, language). Categorical reasons are what make the exclusions inspectable: if 30 of 48 full texts were excluded as "wrong outcome", either your outcome criterion is too narrow or your search strategy is targeting the wrong concept, and you should know which. Where a second screener is available, screen independently, compare, and report agreement and how disagreements were resolved. Where one is not, dual-screen a random subset, report the agreement on that subset, and name single-reviewer screening as a limitation. *Correct result: a screening log with one row per record, the stage-two rows carrying a reason code, and an included set you have read in full.*

**Step 6. Reconcile the numbers and draw the flow diagram.**
The diagram records four stages: identification (records found per source, plus records from other methods), screening (records after duplicates removed, records screened, records excluded), eligibility (full texts assessed, full texts excluded with reasons and counts by reason), and inclusion (studies included, and separately reports included, since one study can generate several papers). The arithmetic must close at every stage, and the reason counts at full text must sum exactly to the full-text exclusions. Examiners check this, because it is the fastest available test of whether the review was actually done. The commonest breaks are duplicates removed in two passes and counted once, records added by citation searching that never enter the diagram, and the study-versus-report distinction being collapsed. Fix the numbers, do not fix the diagram. *Correct result: a flow diagram whose every number is traceable to a row count in the screening log, and which balances.*

**Step 7. Appraise quality with an instrument appropriate to the designs you actually included.**
Select the instrument after you know what designs are in the set, from the standard instruments used in your discipline: a risk-of-bias tool for randomised designs, a checklist appropriate to observational designs, an appraisal framework built for qualitative research, or a mixed-methods instrument where the set is mixed. Do not apply an instrument designed for trials to survey research, which is the most common appraisal error and produces uniformly low ratings that carry no information. Appraise per domain and report per domain: a single overall score hides the distinction between a study with one fatal flaw and a study with four minor ones. Two rules matter more than the instrument choice. First, appraisal informs the synthesis rather than the inclusion decision, unless your protocol said otherwise in advance. Second, appraise the study as reported: poor reporting and poor conduct are indistinguishable from outside, and the honest rating is "unclear", not "high risk". *Correct result: a per-study, per-domain appraisal table, with the instrument named, and a paragraph stating what the appraisal profile implies for how much weight the synthesis can carry.*

**Step 8. Extract into a piloted structured form, from the paper, with a locator.**
Design the form from the review question, pilot it on three papers, and revise it before extracting the rest, because a field you add at study twenty means re-reading nineteen. Minimum fields: full citation, country and setting, design, sampling and recruitment, sample size and characteristics, the intervention or phenomenon as the authors define it, the comparator, the outcome measures and how they were operationalised, the analysis method, the results relevant to your question with the numbers as reported, the authors' stated limitations, funding and conflicts, and your appraisal rating. Extract what the paper says, not what you remember of it, and note the page or section for every extracted claim. Where a study reports a construct under a different name from yours, record both names, because construct drift across an included set is invisible unless recorded. *Correct result: one row per study, complete or explicitly marked "not reported", with every substantive cell carrying a locator.*

**Step 9. Synthesise in the way the evidence supports, and say which way that is.**
Narrative or thematic synthesis is the default and is not a lesser option: it groups studies by the questions they answer, describes the pattern of findings, and accounts for variation by design, population, measure and quality. Tabular and vote-count summaries are permitted as description but never as conclusion, since counting how many studies found an effect ignores their sizes, samples and quality, and systematically favours the underpowered. Quantitative pooling is licensed only where the studies share a comparable population, intervention, comparator and outcome measure, and where heterogeneity has been assessed rather than assumed; if you intend it, this skill hands over to 15.14. Whatever the mode, synthesise by question and by claim, never study by study. A results section running "Smith found... Jones found... Patel found..." is an annotated bibliography and transfers the analytical work to the reader. For each claim, report how many studies support it, at what appraisal quality, in what populations, and what the dissenting studies found. *Correct result: a synthesis organised under the review sub-questions, in which every claim names its supporting studies, their quality profile and the dissent.*

**Step 10. Separate what the literature establishes from what it assumes and what it merely repeats.**
This is the analytical step that distinguishes a systematic review from a summary, and it is the step most often skipped. Sort every recurring claim in the literature into three categories. **Established**: tested empirically, in more than one independent study, with methods adequate to the claim. **Assumed**: treated as background across the literature, cited to prior work, but never tested within it; definitional claims and mechanism claims commonly sit here. **Repeated**: traceable to a single original study or a single non-empirical source that subsequent papers cite without re-testing, so that apparent consensus is one finding with good distribution. Test for repetition by tracing the citations behind a widely stated claim: where twelve papers assert it and eleven cite the twelfth, you have one study. Reporting this properly is often the review's strongest contribution, and it is the material from which a defensible gap statement is built. *Correct result: a three-way classification of the literature's main claims, with the citation traces that establish each classification.*

**Step 11. Report limitations of the review, as distinct from limitations of the included studies.**
These are two different sections and collapsing them is a marking criticism. Review limitations: databases not searched, languages excluded, date window, grey literature omitted, single-reviewer screening, no independent extraction check, publication bias unassessed. Evidence-base limitations: the quality profile of the included studies, their design distribution, their geographic and population concentration, and the outcomes nobody measured. Then state the confidence in each synthesised conclusion in K3 language, capped by the weakest link beneath it: a conclusion drawn from six studies all rated high risk of bias is low confidence however consistent they are. *Correct result: two limitation subsections, and a confidence level attached to every conclusion.*

## 8. Analytical framework

The review is built on one pipeline and one classification, and both must be visible in the output.

**The review pipeline.** Each arrow is a documented transition, and each is a place the review can fail:

    Protocol → Search → Records retrieved → Deduplicated →
    Title/abstract screened → Full texts assessed → Included studies →
    Appraised → Extracted → Synthesised → Conclusion + confidence

Applying it: every arrow carries a number, and the numbers reconcile in the flow diagram. Every arrow after "Records retrieved" carries a documented decision rule from the protocol. The critical arrow is *Full texts assessed → Included studies*, because that is where reasons are recorded and where a review that was actually shaped by preference becomes detectable.

**The claim classification.** Applied to the literature's assertions rather than to its studies:

    Established (tested, replicated, adequately)
      → Established but narrow (tested once, or in one population)
        → Assumed (background, cited, never tested in this literature)
          → Repeated (many citations, one origin)
            → Contested (comparable studies, incompatible results)

Applying it: a review's conclusion may rest on the first two levels. Claims in the third and fourth levels are reported as what the literature assumes, which is a different statement from what it shows, and they are where gaps and research questions come from. Contested claims are adjudicated by appraisal quality and design, never by counting papers.

## 9. Output format

**1. Review question and rationale.** The structured question, the sub-questions, and why the review is needed.

**2. Method.** Review type named accurately. Protocol status and date, and registration where applicable. Inclusion and exclusion criteria table. Databases and sources searched with dates. Screening procedure, number of screeners, agreement where measured. Appraisal instrument named. Extraction fields. Planned synthesis. Deviations from protocol, with reasons.

**3. Search appendix.** The full search string for every database, verbatim, with the platform, limiters, date run and record count.

**4. Flow diagram.** Four stages, with reconciling numbers and full-text exclusion reasons with counts.

**5. Characteristics of included studies.**

| ID | Citation | Country and setting | Design | Sample (n, population) | Intervention or phenomenon | Outcome measures | Analysis | Appraisal rating |
|---|---|---|---|---|---|---|---|---|

**6. Quality appraisal.** Per study, per domain, with the summary of what the profile means for the synthesis.

**7. Synthesis, organised by sub-question.** Each claim with supporting study IDs, quality profile, population coverage and dissenting findings.

**8. What the literature establishes, assumes and repeats.** The three-way classification with citation traces.

**9. Limitations.** Review limitations and evidence-base limitations, separately.

**10. Conclusion and implications**, each with a K3 confidence level, and the gap statement that follows.

**When the evidence is thin, the format must not force fabrication (K4 §1).** A review that includes four studies reports four studies, describes why the search returned so few (with the search log as evidence that looking was done), and states what cannot be concluded. The characteristics table shrinks; it is never padded, and a study is never included to fill it. If the synthesis is not supportable, the output is a description of the state of the literature plus a research gap statement, and that is a successful review.

## 10. Quality checks

Run before submission. These sit on top of K4 §8.

1. Was the protocol written and dated before the first search, and is every deviation from it reported with a reason?
2. Does the flow diagram's arithmetic close at every stage, and do the full-text exclusion reasons sum to the full-text exclusions?
3. Is the full search string reproduced for every database, in full, rather than summarised?
4. Does the search return the seed studies you knew about before you started?
5. Was every included study read in full, and is any study included on the basis of its abstract alone?
6. Does every full-text exclusion carry a reason from the pre-specified list?
7. Is the appraisal instrument appropriate to the designs actually included, and is it named?
8. Is appraisal reported per domain rather than collapsed into a single score?
9. Does every extraction row carry a locator, and is every "not reported" cell marked as such rather than left blank?
10. Is the synthesis organised by question and claim rather than study by study?
11. Has any conclusion been drawn by vote counting?
12. Has every claim asserted repeatedly in the literature been traced to its origin, so that repetition is not reported as convergence?
13. Are review limitations and evidence-base limitations reported separately?
14. Does every conclusion carry a confidence level capped by the quality of the studies beneath it?
15. Is the review described by its accurate type throughout, with no drift between "structured", "scoping" and "systematic" in different chapters?

## 11. Common failure modes

| Failure | How to recognise it | How to prevent it |
|---|---|---|
| **Retrofitted protocol** | The protocol is written after screening and matches the included set exactly | Date the protocol, lodge it, and report deviations (Step 2) |
| **Scope collapse** | The question narrows in week six to whatever the reading produced | Run the proportionality gate first, narrow with reasons (Step 1) |
| **Irreconcilable numbers** | The flow diagram does not add up, or citation-search finds appear from nowhere | Every number is a row count in the log (Step 6) |
| **Pasted search string** | The same string across four databases, with one returning almost nothing | Adapt syntax and vocabulary per database, pilot each (Step 4) |
| **Unrecorded exclusions** | "We excluded studies that were not relevant" | Closed-list reason codes at full text (Step 5) |
| **Wrong appraisal instrument** | A trials instrument applied to interview studies, all rated poor | Select after the design profile is known (Step 7) |
| **Vote counting** | "Seven of eleven studies found an effect, so the effect exists" | Synthesise by weight and quality, not by count (Step 9) |
| **Citation cascade read as consensus** | A claim everyone makes, one original study behind it | Trace claims to origin (Step 10) |
| **Annotated bibliography** | The synthesis is a sequence of study summaries | Synthesise by claim (Step 9) |
| **Construct drift** | Studies measuring different things pooled under one label | Record the authors' own construct names at extraction (Step 8) |
| **Method-label inflation** | A structured review described as systematic because it sounds better | Name the type accurately and consistently (§4, Step 2) |
| **AI-invented included studies** | A characteristics table row for a paper nobody retrieved | Nothing enters the table without a retrieved full text (§12) |
| **AI-invented search results** | Record counts or database hits that were never run | Never narrate a search that was not performed (§12) |
| **Silent full-text failure** | A study cited as included that could not actually be obtained | "Full text unobtainable" is an exclusion reason, and appears in the diagram |

## 12. AI guardrails

Skill-specific. The universal prohibitions in K4 apply in full and are not repeated. The human in the loop for every K5 marker in this skill is the candidate together with their supervisor.

1. **Academic integrity is a condition of use, not a footnote.** These skills assist a researcher's thinking, structure and rigour. They do not produce work to be submitted as the student's own unaided output. The user must comply with their institution's AI use policy and its declaration requirements, which vary by institution and by assessment. Where an institution prohibits AI assistance for a task, this skill must not be used for it. The skill never writes a passage for submission as though the student wrote it; it interrogates, structures, critiques and teaches. Operationally in this skill: it will draft protocol structure, criteria tables, extraction forms, screening reason lists, appraisal domain templates and synthesis scaffolds, and it will critique a draft the student has written. It will not write the review's prose findings, and it will not fill a characteristics table with studies the student has not retrieved and read.

2. **No study enters the included set unless the student has retrieved and read the full text.** Not the abstract, not a summary, not a search result snippet. A characteristics table populated from model knowledge is fabrication in bulk, and it is trivially detectable by an examiner who checks one row.

3. **Never generate a search result, a record count or a database hit.** Do not state that a database was searched, do not produce a number of records retrieved, and do not report that a search "returned approximately" anything. The student runs the search and reports the number. If asked for the counts, the correct response is to supply the log structure, empty.

4. **Never complete a reference, a DOI, a sample size, a p value or an effect size by inference.** Where a field is not in the paper in front of the student, it is "not reported". An invented DOI or a plausible-looking sample size is more harmful than a blank, because it will be checked and the whole extraction table will fall under suspicion.

5. **Never assign an appraisal rating to a study you have not been given.** Appraisal is a judgement about a specific document. Offer the instrument's domains and the questions each domain asks; the rating comes from the student reading the paper.

6. **Never assert that no study exists on a topic.** Absence claims rest on a search, and you did not run one. The permitted statement is that the searches recorded in the log did not identify one, naming where the student looked (K4 §6.4).

7. **Never upgrade the review type.** If the method executed was a structured single-reviewer review, it is described as one, in the methodology, the abstract and the viva. Systematic is a claim about procedure, not a synonym for thorough.

8. **Never resolve a disagreement between studies by averaging or by silent selection.** Report both positions, their appraisal quality, and the diagnosed reason for the difference. Where no adjudication is possible, say so (K4 §4.1).

9. **Never present a claim's citation count as evidence of its truth.** Where a widely cited claim traces to one source, that is the finding, and it must appear in the output rather than being smoothed into consensus.

10. **Where the student's protocol and their executed method have diverged, say so plainly rather than helping the document conceal it.** A reported deviation with a reason is normal practice. A concealed one is misconduct, and assisting with the concealment is not a service.

## 13. Best-practice principles

- **The proportionality decision is the highest-leverage decision in the whole review.** A well-executed scoping review is marked above a poorly executed systematic review, every time, and the choice is only cheap to make in week one.
- **The protocol is a commitment device, not paperwork.** Its value is precisely that it stops you from deciding what counts after you know what you found.
- **Pilot the search against studies you already know.** A search that cannot find the papers that prompted the review is broken, and this test takes ten minutes.
- **Screen liberally at abstract, strictly at full text.** Errors at stage one are unrecoverable; errors at stage two cost one wasted retrieval.
- **Reason codes are what make screening inspectable.** They also diagnose your own criteria: a pile of exclusions on one code usually means a criterion is mis-specified.
- **Appraise the reporting as well as the study, and say which you are appraising.** "Unclear" is an honest and common rating, and forcing it to "low" or "high" invents information.
- **Poor-quality studies still belong in the synthesis, weighted accordingly.** Excluding on quality after the fact, without a protocol rule, is where reviews acquire their conclusions.
- **The most valuable output is usually the establishes-assumes-repeats classification.** It is what a supervisor cannot get from reading the same papers quickly, and it is what makes the gap statement defensible.
- **Extract once, extract properly.** Re-reading forty papers because the form lacked a field is the single largest avoidable time loss in review work.
- **Keep the log live, not retrospective.** A screening log reconstructed at the end will not reconcile, and the numbers are the first thing an examiner checks.
- **Report the review's own limitations generously.** A review that names its coverage gaps reads as competent. A review with no limitations section reads as one whose author did not understand the method.
- **Write the method section while doing the method.** It is the only chapter that is easier to write in the present tense of the work than from memory two months later.

## 14. Worked example

Generic fictional scenario, academic.

**INPUT**

A masters candidate in a health-adjacent social science department plans a dissertation on peer support programmes for carers of people with long-term conditions. The supervisor has said "do a systematic review". The candidate has fourteen weeks in total, no second screener, and access to three databases through the institution.

**PROCESS**

*Step 1.* Pilot search on the two core concepts, unrestricted, returns roughly 4,100 records in the broad index alone. At 90 records an hour, title and abstract screening alone is 45 hours, before retrieval, appraisal or extraction, and before any of the other dissertation chapters. The candidate has perhaps 120 hours for the whole review. The gate fails as scoped.

*The judgement call.* Three options are on the table: narrow the question, change the review type, or proceed and hope. Proceeding is rejected. Narrowing is attempted first, since it preserves the systematic claim: the population is tightened to carers of adults with dementia specifically, the outcome is tightened from "wellbeing" to "carer burden or caregiver strain as measured by a named instrument", and the window is set at the last fifteen years with the reason that the service model changed after a policy reform in the candidate's country. The re-piloted search returns 610 records across three databases. At the same rates, screening is around 7 hours, roughly 55 full texts retrieved at 6 an hour is 9 hours, and appraisal plus extraction of an expected 15 to 20 included studies at 2 an hour is around 10 hours. Total is feasible. The review stays systematic, with single-reviewer screening declared and a 15 percent dual-screened subset arranged with a peer.

*Steps 2 and 3.* Protocol written and dated, lodged with the supervisor. The question is structured as population (unpaid carers of adults with a dementia diagnosis), intervention (peer support delivered by other carers, in any modality), comparator (usual support or none), outcome (carer burden or strain, measured), designs (any empirical design including qualitative, since the candidate wants experience as well as effect). Criteria are written to be decidable: "peer support" excludes professionally led groups unless a peer co-facilitates, and this test is written down before screening.

*Steps 4 to 6.* Searches adapted per database and run on one day. 610 records, 148 duplicates, 462 screened, 401 excluded, 61 full texts sought, 4 unobtainable, 57 assessed, 39 excluded (wrong population 14, professionally led intervention 11, no burden outcome 9, not empirical 5), 18 studies in 20 reports included. The diagram balances. The dual-screened subset shows agreement on 66 of 70 abstracts, with the four disagreements all on the peer-facilitation criterion, which is reported.

*Step 7.* The included set is 11 quantitative (7 uncontrolled pre-post, 3 controlled, 1 randomised) and 7 qualitative. Two instruments are used, one for the quantitative designs and one for the qualitative, both named. Ten of the eleven quantitative studies are rated unclear on allocation and blinding, mostly because they are single-group designs where those domains do not apply, which is recorded as a design characteristic rather than a bias rating.

*Steps 9 and 10.* Vote counting would say nine of eleven studies found reduced burden, so peer support reduces burden. The synthesis refuses this: seven of the nine are uncontrolled pre-post designs on small samples, in which regression to the mean and the natural trajectory of a crisis-point referral both predict improvement without any intervention effect. The one randomised study finds a small effect that does not reach the conventional threshold. The three controlled studies split. The honest synthesis is that the effect on measured burden is not established, while the qualitative studies converge strongly on a different outcome the quantitative studies do not measure: a reduction in isolation and self-blame that participants describe as the main benefit. The classification step finds that the widely repeated claim that peer support "reduces service utilisation" traces, across nine citing papers, to one economic evaluation of a different programme type in a different country.

**OUTPUT**

A review that reports the burden question as under-evidenced with the design reasons named, reports the qualitative convergence as the better-supported finding, identifies the outcome-measure mismatch between the two literatures as the central problem in the field, and traces one widely repeated claim to a single non-transferable original. The gap statement writes itself: a controlled study measuring the outcomes carers themselves nominate. Confidence is stated as low on the burden conclusion and moderate on the experiential one.

`RESEARCHER DECISION REQUIRED` at Step 1 on the scope narrowing (K5 §2.7), and `RESEARCHER SIGN-OFF REQUIRED` on the protocol before searching.

## 15. Advanced usage

**Living the protocol into the methodology chapter.** Write the protocol in the structure the methodology chapter needs, and the chapter is drafted as a by-product. The tense changes and the deviations section is added; nothing else needs rewriting. This alone saves a week.

**Reviews as instrument sources.** Extract the outcome instruments the included studies used, not only their findings. Where a primary study follows the review, using an instrument the literature already uses makes your results comparable, which is worth more than a marginally better instrument that compares to nothing. Hands over to **02.07 Scale and Measurement Selection**.

**Deliberate disconfirmation search.** After synthesis, run one additional search using the vocabulary a sceptic would use, including terms for null and negative findings and for the harms or costs of the intervention. Literatures are systematically biased toward the language of effect, and this search takes an hour. See **13.04 Bias Detection**.

**Sensitivity checks without meta-analysis.** Even in a narrative synthesis, re-state the main conclusion twice: once excluding the studies rated at highest risk of bias, and once excluding the largest study. If the conclusion changes, that fragility is a finding and belongs in the write-up.

**Where the standard approach does not fit.** In fields with little indexed literature, shift the source map toward theses, working papers, regulator and agency reporting and conference proceedings, lower the appraisal expectation explicitly, and describe the review as a scoping review of a fragmented evidence base. In fields where the literature is overwhelmingly qualitative, the appropriate synthesis methods are interpretive rather than aggregative, and the review's claim is about the range and depth of accounts rather than about effect. Say which you are doing.

## 16. Skill chain

**Recommended previous skills:**
- **15.01 Academic Research Topic Selection.** Hands over a topic narrow enough that inclusion criteria can be written against it.
- **15.02 Academic Literature Search Strategy.** Hands over the database selection, controlled vocabulary, Boolean strings and search log this skill screens from.
- **15.03 Referencing and Citation Management.** Hands over the reference management setup that makes deduplication and citation integrity survivable.

**Recommended next skills:**
- **15.07 Theoretical and Conceptual Framework Development.** Takes the constructs and relationships the review surfaced and builds the framework.
- **15.08 Research Design and Methodology Chapter.** Takes the gap statement and the instruments the literature uses, and designs the primary study.
- **10.04 Research Gap Identification.** Turns the establishes-assumes-repeats classification into a specification of what must be researched.
- **15.14 Advanced Systematic Review and Meta-Analysis.** Takes over where quantitative pooling is intended or doctoral-standard rigour is required.

**Runs well alongside:**
- **13.02 Source and Citation Verification**, which audits the reference set this review produces.
- **13.04 Bias Detection**, for the disconfirmation pass.
- **15.05 Academic Writing Structure and Argumentation**, for turning the synthesis into examinable prose.

---
A Yazi Supplied Skill and resource.
