---
name: meta-analysis-across-studies
description: >
  Answers a question from the research an organisation has already conducted,
  before commissioning anything new. Use for "what do all our studies say about",
  "can we answer this from existing research", "synthesise our past studies",
  "do we already know this", "pull together everything we have on", "what does
  the body of our research show", "before we commission another study", "have our
  findings changed over time".
category: 14 Activation and Knowledge Management
ref: "14.04"
tier: 2
inherits: [K2, K3, K4, K5]
---

# Meta-Analysis Across Studies

## 1. One-line description

A method for interrogating an organisation's own accumulated studies against a specific question, assessing honestly whether they are comparable enough to be combined, and producing either a state-of-knowledge statement with confidence per claim or an explicit verdict that the existing corpus cannot answer it.

## 2. What this skill is used for

**The research problem it solves.** Most organisations with a research history are sitting on the answer to the question they are about to spend money on, and cannot get to it. The instinct when a question arrives is to design a study, because designing a study is what a research function does, and because the existing corpus is not in a state that can be interrogated. So the same ground is re-covered, at cost, and the new study does not connect to what came before, which makes the corpus one study more fragmented than it was. **For an organisation with several years of research behind it, this is usually the highest-return skill in this library, and it is worth saying plainly: the cheapest study is the one you do not commission because you already had the answer.** The difficulty is that cross-study work has a specific and common failure mode. Studies with different objectives, samples, periods, methods and question wording are lined up, their headline numbers are compared, and a trend or a consensus is reported that is an artefact of the differences between the studies rather than a fact about the world. This happens quietly, it is very hard to detect downstream, and it is where most internal meta-analysis fails. This skill puts the comparability assessment before the synthesis, weights studies by quality rather than counting them, separates real change over time from methodological difference, and treats "the corpus cannot answer this" as a legitimate and frequently correct output.

**Where it sits in the research lifecycle.** Before design, as the first move on any new question in an organisation with a research history; and after activation, as the way accumulated work is made to pay. It closes the loop in this library: it consumes the curated corpus and produces either an answer or a specification for the next study, which returns to the start of the lifecycle.

**Typical use cases.**
- A question arrives and the honest first move is to establish what the organisation already knows.
- A study is being scoped and the brief must exclude what has already been established.
- Someone asks whether a measure has changed over time and the evidence sits across several non-identical studies.
- A recurring finding needs to be assessed for whether it genuinely replicates or has been asserted repeatedly from one source.
- A strategic review needs a state-of-knowledge statement on a subject the organisation has researched repeatedly.
- Two or more past studies appear to disagree and someone must establish whether they actually do.
- A funder, board or regulator asks what the accumulated evidence supports.

**Who uses it.** Research directors deciding whether to commission; insight leads asked what the organisation knows; strategists and policy analysts building a position from existing evidence; knowledge managers who curate the corpus and are asked to interrogate it; consultants inheriting a client's research history.

## 3. When to use it

- A question has arrived and existing studies plausibly bear on it.
- A new brief is being written and its scope should be set by what is already known.
- A trend claim is being made across studies that were not designed as a series.
- A finding is treated as established inside the organisation and nobody can say which studies support it.
- Past studies appear to conflict and the conflict needs diagnosing rather than adjudicating by recency.
- A body of research must be converted into a defensible position for a board, a funder or a regulator.
- The organisation is deciding whether to invest in more research on a subject and needs to know the marginal value of doing so.
- A repository has been built and its first real test is whether it can answer a question.

## 4. When NOT to use it

- **The corpus does not contain studies bearing on the question.** Adjacent evidence assembled into an answer is worse than no answer, because it looks like knowledge. Where the honest finding is that the organisation has never researched this, say so and route to **01.02 Business Problem to Research Question** and **10.04 Research Gap Identification**. This is a successful outcome of the skill, not a failure of it.
- **The evidence you need is external and published rather than the organisation's own.** That is **10.02 Evidence Synthesis**, which synthesises external published evidence, and **10.01 Literature Review and Desk Research**, which finds and appraises it. The boundary runs both ways: 10.02 synthesises evidence produced by others, where the analytic problem is source quality and provenance; **14.04 synthesises the organisation's own study corpus, where the analytic problems are comparability across studies you commissioned and the systematic bias in what you chose to research.** Where a question needs both, run them separately and integrate through **10.03 Multi-Source Research Synthesis**; do not pool internal and external evidence in one undifferentiated table.
- **Formal quantitative pooling is required to academic or regulatory standard.** Effect-size pooling, heterogeneity statistics, publication-bias assessment and protocol-standard reporting are a different method with different prerequisites. That is **15.14 Advanced Systematic Review and Meta-Analysis**. Do not describe the structured synthesis in this skill as a meta-analysis in the statistical sense, and do not report a pooled estimate the corpus cannot support.
- **The studies are too heterogeneous to combine, and the comparability assessment says so.** This is a real and common verdict. Where objectives, populations, wording and methods differ enough that no claim survives the assessment in step 3, report the studies separately with what each establishes, and stop. Forcing a synthesis produces false precision, which is the most damaging output this skill can generate.
- **There is only one relevant study.** Re-reading one study more carefully is secondary analysis, not meta-analysis. Go to the relevant analysis skill in Category 05 or 07, or simply return to the original report.
- **The decision requires current data and the corpus is old.** Where every relevant study predates a structural change (a policy shift, a product change, a market event, a change in the population itself), the corpus describes a world that no longer exists. The correct output is a dated picture labelled as historical, plus a specification for what must be measured now.
- **The provenance of the studies is unknown or unverifiable.** Where studies arrive without method documentation, sample descriptions or base sizes, they cannot be quality-weighted, and a synthesis that weights unassessable studies equally with assessed ones is arithmetic rather than analysis. Recover the documentation, exclude the studies that lack it and say so, or run **13.01 Research Quality Review** first.
- **The task is to assemble and structure the corpus rather than to interrogate it.** That is **14.03 Research Repository and Knowledge Curation**. 14.03 makes a corpus interrogable; this skill interrogates it. Do not build a repository in order to answer one question.
- **A conclusion has been fixed in advance.** Where the request is to demonstrate that the accumulated research supports a decision already made, this is not synthesis. Per K4 §4.2, say so, and offer either an honest synthesis or a clearly labelled assessment of the stated position against the corpus.

## 5. Required inputs

**Required.** Without these the skill cannot run. If absent, ask. If no answer is available and work must proceed, state the assumption at the point where it bites, per K5 §5.

- **The question, framed as something evidence could settle**, with the population, construct and period it concerns. A topic is not a question and cannot be interrogated against a corpus.
- **The candidate studies, with their methodology documentation**: objectives, sample definition and size, fieldwork dates, method and mode, and the instrument as fielded. **The instrument matters more than the report.** Question wording is where most false comparability hides, and a study whose wording cannot be recovered cannot be compared on that measure.
- **The findings with their bases and references** (K2 §4). A headline percentage with no base and no question reference cannot be weighted, compared or pooled.
- **An honest account of what is missing from the corpus**, which requires asking rather than inspecting: studies that were run and never written up, projects abandoned mid-flight, work held by an agency that was never handed over, and studies excluded by consent or contract.

**Optional, and what each one adds.**

- **Raw datasets rather than only reports**: transform the exercise. With raw data, questions can be re-analysed on a common base and a common definition, which removes the largest single source of false comparability. Without it, you are comparing other people's analytical choices.
- **A curated repository** (from **14.03**): reduces corpus assembly from weeks to hours and supplies the status, confidence and provenance fields this skill would otherwise have to reconstruct.
- **The original briefs**: reveal what each study was designed to do, which determines whether its finding on your question was a primary measure or an incidental one, and incidental measures are systematically weaker.
- **Quality review records**: allow weighting by assessed quality rather than by inference from method description.
- **Fieldwork and analysis notes**: hold the caveats that never reached the report, and frequently explain an anomaly that would otherwise be read as a real difference.
- **A record of what changed in the world between studies** (product, pricing, policy, competitive or regulatory events): converts the temporal analysis in step 6 from speculation into something checkable.

## 6. Questions to ask before starting

1. **What decision does this synthesis serve, and would a definitive answer change it?** Determines the required standard of evidence and whether the exercise is worth running. *Default if unanswered:* assume a commissioning decision, and report explicitly on whether the corpus is sufficient to avoid new fieldwork.
2. **What population, construct and period does the question concern?** These are the inclusion criteria, and without them the corpus cannot be screened. *Default:* propose them explicitly and state them as an assumption, since a synthesis run on an unstated construct pools studies measuring different things.
3. **What is missing from the corpus, and who would know?** Determines the honesty of the whole exercise. *Default:* ask the longest-serving researcher and the commissioning function directly, because unwritten and abandoned studies do not appear in any inventory and their absence is rarely random.
4. **Are raw datasets available, or only reports?** Determines whether re-analysis on a common base is possible, which is the difference between comparing measurements and comparing other people's analyses. *Default:* assume reports only, and cap the claims accordingly.
5. **What changed in the world between the earliest and latest study?** Determines whether a difference over time can be read as change at all. *Default:* build the event timeline before looking at any results, so it cannot be constructed to fit the pattern you find.
6. **Is there a belief inside the organisation about what the corpus shows?** Determines what to guard against, since a synthesis is unusually easy to steer. *Default:* record the belief before starting, so any convergence with it is visible rather than invisible.
7. **Who commissioned this synthesis, and is a particular answer expected?** Determines whether the exercise is honest. *Default:* ask directly, and note the answer in the output, per K4 §4.2.

## 7. Step-by-step methodology

**Step 1. Frame the question so studies can be interrogated against it, and write the interrogation criteria.** A cross-study question must name a population, a construct and a period, and must be answerable by evidence rather than by preference. "What do we know about customer loyalty" cannot be interrogated. "Among smallholder farmers in the two eastern regions, has the proportion reporting that they would recommend the programme changed between 2019 and 2025, and does it differ by farm size?" can. Then, before opening any study, write the criteria a study must meet to bear on the question: which population, which construct, what minimum method documentation, what base. Writing them first is what stops the corpus being selected around the answer, which is the easiest thing in the world to do without noticing. *Correct result: one to four interrogable questions, each with written inclusion criteria fixed before screening.*

**Step 2. Assemble the corpus, and state what is missing from it.** Screen every candidate study against the criteria and record the disposition of each: included, excluded with a categorical reason (wrong population, wrong construct, outside the period, insufficient method documentation, inaccessible, restricted by consent or contract). Then do the part that separates an honest synthesis from a plausible one: **establish what the corpus does not contain.** Four categories, none of which appears in any file listing. Studies that were run and never written up, which are common and are systematically the ones that produced inconvenient or inconclusive results. Projects abandoned mid-fieldwork. Work commissioned from an agency and never transferred. Studies excluded by consent or contract terms. Ask people, because none of this is discoverable by inspection. Record the missing set explicitly in the output, because a corpus with a known hole is interpretable and a corpus with an unknown hole is not. *Correct result: an inclusion table with a disposition for every candidate, and a named missing-studies register with what each would have contributed.*

**Step 3. Assess comparability, dimension by dimension, before comparing anything.** This is the hard part, it is where internal meta-analysis silently fails, and it must be completed before any result is looked at, because comparability judged after seeing the numbers is judged to fit them. Build a matrix with the studies as rows and these dimensions as columns. **Objective:** was this measure the study's primary purpose or an incidental item? Incidental measures get less design attention, sit in worse questionnaire positions and are weaker. **Population and sample frame:** who was eligible, how were they selected, and are the two populations the same people described differently? **Base:** what is the specific base for this measure, which is frequently not the study's headline sample. **Period:** fieldwork dates, not publication dates, and whether the period was unusual. **Method and mode:** self-completion and interviewer-administered produce systematically different answers on sensitive and socially desirable measures, and the difference can exceed the effect being studied. **Question wording and scale:** exact wording, scale points, labelling, and whether the scale ran the same direction. **Construct definition:** what the measure was taken to mean. **Preceding context:** what came before the question, since a preceding battery changes the answer to the one that follows. Then score each pair of studies on each dimension as comparable, comparable with caveat, or not comparable, and record the reason. **A single "not comparable" on wording or construct is disqualifying for that measure, regardless of how well everything else lines up.** The commonest single error in internal meta-analysis is comparing two questions that read alike and were not the same question: a five-point scale against a seven-point, "satisfied" against "very satisfied or satisfied", "in the last year" against "ever". *Correct result: a comparability matrix with a reason recorded against every non-comparable pair, and a set of measures cleared for comparison that is normally much smaller than the set that was hoped for.*

**Step 4. Choose the synthesis type honestly, and say why.** Two are available and they are not interchangeable. **Formal quantitative meta-analysis** pools effect sizes across studies, weights by precision, quantifies heterogeneity, and assesses publication bias. It requires studies asking the same question of comparable populations with reported variance and consistent measurement, which is a set of conditions that clinical trials and replicated experiments meet and that commercial, policy and organisational research corpora almost never do. The reason is structural rather than a matter of effort: studies commissioned to answer different business questions at different times are not designed to be pooled, wording is revised between waves precisely because someone wanted a better question, and variance is rarely reported. Attempting a pooled estimate on such a corpus produces a number with a decimal point and no meaning, and the decimal point is what makes it dangerous (K3 §4.4). **Structured qualitative synthesis** is what these corpora do support: findings assessed for convergence, divergence and replication; weighted by study quality; with the direction and approximate magnitude reported and the pattern across studies described. It is not a weaker version of the first. It is the correct method for heterogeneous evidence, and stating which one you are doing and why is a required part of the output. Where a subset of studies genuinely is comparable (a repeated tracking wave on identical wording, say), quantitative comparison is legitimate for that subset and is reported as such, with the rest handled qualitatively. *Correct result: an explicit statement of the synthesis type, the reason, and any subset handled differently.*

**Step 5. Weight by study quality, not by counting studies.** Three studies agreeing is not evidence if all three are small, incidental and poorly documented, and one well-designed study on an adequate base outweighs them. Score each included study on: base size for the specific measure; sample quality and frame coverage; method fit for the construct; instrument quality including whether the wording is neutral and the scale sound; whether the measure was primary or incidental; documentation completeness; and independence, meaning whether this study is genuinely separate evidence or a repeat of an earlier one by the same team using the same instrument, which corroborates less than it appears to. Do not average these into a single score, since a fatal weakness on one is not offset by strength on another; assign a tier and record the reasons. Then apply the weighting where it bites: **a claim's confidence is set by the strongest studies that support it, not by the number of studies mentioning it** (K3 §3). Counting studies is how a weak finding becomes established: it gets repeated across reports, each repetition is counted as support, and the count is mistaken for corroboration. *Correct result: a tiered study table with reasons, and an explicit rule stating that claims are assessed on the strength of supporting evidence rather than on its volume.*

**Step 6. Analyse the temporal dimension, and separate real change from methodological difference.** When a measure differs between an earlier and a later study, there are four possible explanations and only one of them is change in the world. It may be **methodological**: the wording, scale, mode, sample frame or base changed. It may be **compositional**: the population changed, so a different set of people is being measured under the same label, which is common where a customer base has grown or a programme has expanded. It may be **noise**: the difference is within what sampling variation would produce on those bases. Or it may be **real**. Work through them in that order, and note that the first three are collectively far more likely than the fourth. Test the methodological explanation by putting the two instruments side by side, not by comparing the reports. Test the compositional explanation by comparing sample profiles. Test noise by asking whether the difference is larger than the bases could produce by chance, and where no test has been run, say so rather than describing the difference as significant (K4 §3.1). Only where all three are excluded is a change claim licensed, and even then it is a change in a measure, not a demonstrated cause. Where you cannot exclude them, the honest output is that the corpus cannot establish whether the measure changed, which is itself a useful finding and a specification for what a proper time series would require. *Correct result: a temporal table listing each apparent change with each of the four explanations assessed and the verdict recorded, including the verdict "cannot be determined".*

**Step 7. Separate what replicates from what appears once.** Findings that hold across independent studies with different samples, and ideally different methods, are the corpus's most valuable output and can be stated with high confidence where the studies are of adequate quality. Findings that appear in one study only are not thereby wrong, but they are unreplicated, and the corpus's job is to say which is which. Check three things before recording a replication. Are the studies genuinely independent, or does a later one repeat an earlier instrument with the same team and the same frame, which is closer to a re-run than to corroboration? Does the finding hold in the same direction and with a similar magnitude, or only loosely in the same area? And has anything failed to replicate, which is the case nobody looks for and which caps confidence at moderate wherever it exists (K3 §3.5). *Correct result: every claim classified as replicated (with the studies named), single-source, or contested, and a register of failed replications.*

**Step 8. Assess the corpus's own bias, because it is systematic rather than random.** An organisation researches what it is interested in, what it has budget for, what a stakeholder championed, and what is easy to reach. The consequence is that the gaps in a corpus are not random and cannot be treated as noise. Look for four patterns. **Topic bias:** the subjects that attract repeated study, usually those close to a funded programme, against those never examined. **Population bias:** who has been researched repeatedly (usually current customers, engaged users, accessible participants) and who has never been researched at all (lapsed, refused, excluded, hard to reach, and the people the organisation does not serve). This is the most consequential and the most invisible: a corpus of current customers cannot answer any question about non-customers, however large it is. **Question bias:** whether the corpus has consistently asked about a thing in one framing, so that the same measurement artefact appears in every study and reads as a robust finding. **Outcome bias:** whether studies producing inconvenient results were written up and retained at the same rate as others, which connects directly to the missing-studies register from step 2. Report the bias assessment as part of the output, next to the conclusions it qualifies. *Correct result: a named bias assessment on all four dimensions, with an explicit statement of the questions this corpus is structurally incapable of answering.*

**Step 9. Build the state-of-knowledge statement, with confidence per claim.** Organise the output by claim rather than by study, since a study-by-study account is a reading list and pushes the analysis onto the reader. For each claim state: what the corpus establishes; the studies supporting it, by reference, with their bases and dates; the confidence level per K3 §2 with the reason it is not higher; the studies that dissent, if any, and the diagnosed cause; and what would raise the confidence. Keep the K2 boundary visible: what individual studies found is a finding, what you conclude from reading them together is an interpretation, and the join must be legible (K2 §3.3). Include the age of the evidence in every claim, since a high-confidence claim resting on studies from five years ago is a high-confidence claim about five years ago. *Correct result: a claim-level statement with confidence, supporting studies, dissent and the ageing of the evidence visible on every line.*

**Step 10. Deliver the verdict, including the verdict that the corpus cannot answer the question.** Classify the question as **answered** (the corpus establishes it at adequate confidence, and no new research is needed), **partially answered** (a component is established and a component is not, with the components named), **contested** (studies genuinely disagree after diagnosis, and the disagreement is the finding), or **unanswerable from this corpus** (no adequate evidence, or comparability fails). The last is frequent and it is a successful output. It routes to **01.02 Business Problem to Research Question** and **10.04 Research Gap Identification** with something far more valuable than a blank page: a specification derived from exactly why the corpus failed, whether that was population coverage, wording inconsistency, base size, ageing or the absence of any measurement at all. And where the answer is that the corpus does contain the answer, say so plainly and name the money not spent. *Correct result: a verdict per question, with either a stated answer and its confidence, or a specification for the study that would close the gap.* `RESEARCHER DECISION REQUIRED` where the verdict is that no new research is needed on a material decision (K5 §2.5, §2.7).

## 8. Analytical framework

The synthesis runs on a gate structure, and each gate can stop it:

    Question → Corpus assembly → [comparability gate] → Quality weighting
        → Convergence and replication → [temporal gate] → Bias assessment
            → State of knowledge + confidence → Verdict

**Applying it.** The two gates are what distinguish this from lining studies up in a table. The **comparability gate** removes measures that cannot legitimately be compared, and it normally removes more than expected: a corpus of nine studies frequently yields three that can be compared on the measure in question. The **temporal gate** removes apparent changes that are methodological, compositional or noise. A synthesis that passes everything through both gates unchanged has almost certainly not applied them.

**The four explanations for a difference between studies**, which is the single most useful diagnostic in the skill:

| Explanation | How to test it | Frequency |
|---|---|---|
| Methodological (wording, scale, mode, base, frame) | Put the two instruments side by side | Most common |
| Compositional (the population itself changed) | Compare sample profiles, not headline sizes | Common |
| Noise (within sampling variation on those bases) | Check the bases; test if the data allows, and say if untested | Common |
| Real change in the world | Only claimable once the other three are excluded | Least common |

**Weight, not count.** The governing rule of the whole skill. Evidence strength is a function of study quality, independence and base, not of how many documents mention a claim. A repeated claim in a corpus is very often one original finding with good internal distribution, which is the internal equivalent of citation chaining.

## 9. Output format

**1. Question and criteria.** The interrogable questions, the population, construct and period, and the inclusion criteria as written before screening.

**2. Corpus and its holes.**

| Study ID | Title | Date of fieldwork | Method | Sample and base for this measure | Included / excluded and why |
|---|---|---|---|---|---|

Plus the **missing-studies register**: studies known to exist and unavailable, run and never written up, abandoned, or restricted, with what each would have contributed.

**3. Comparability matrix.** Studies against dimensions (objective, population, base, period, method and mode, wording and scale, construct, preceding context), with comparable, caveated or not comparable recorded per pair and a reason on every non-comparable cell.

**4. Synthesis type.** Which method is being used, and why the corpus does or does not support quantitative pooling.

**5. Study quality tiers.** Tier per study with the reasons, and an explicit note of any studies that are not independent of each other.

**6. State of knowledge by claim.**

| Claim | Supporting studies | Bases and dates | Replication status | Dissent and diagnosed cause | Confidence | What would raise it |
|---|---|---|---|---|---|---|

**7. Temporal assessment.** Each apparent change, the four explanations assessed, and the verdict including "cannot be determined".

**8. Corpus bias assessment.** Topic, population, question and outcome bias, with an explicit statement of what this corpus structurally cannot answer.

**9. Verdict and recommendation.** Per question: answered, partially answered, contested, or unanswerable from this corpus, and either the answer with its confidence or the specification for the research that would close the gap.

**When the evidence is thin.** The output shrinks and the verdict section grows; it is never padded. A claim supported only by studies that failed the comparability gate is not reported as a claim, it is reported as a measurement problem. Where the whole synthesis returns "unanswerable from this corpus", that is the deliverable, and it is a short, valuable document: it prevents a study being commissioned to a brief that would have repeated the corpus's own blind spot, and it specifies the study that would actually work. Per K4 §1, no claim is manufactured because the output format has a row for it.

## 10. Quality checks

Run before anything is presented. Sits on top of K4 §8.

1. Were the inclusion criteria written before any study was screened, and is every excluded study recorded with a categorical reason?
2. Does the output name what is missing from the corpus, including studies never written up, abandoned or restricted?
3. Was the comparability matrix completed before any results were compared?
4. Has the exact question wording been checked between studies, from the instruments rather than from the reports?
5. Has any measure been compared where scale points, labels, direction, timeframe or base definition differ?
6. Is the synthesis type stated, with the reason the corpus does or does not support quantitative pooling?
7. Is any pooled estimate or averaged figure reported that the comparability assessment does not license?
8. Are claims weighted by study quality rather than by the number of studies mentioning them?
9. Has independence been checked, so that a repeated instrument by the same team is not counted as corroboration?
10. For every apparent change over time, have the methodological, compositional and noise explanations been assessed and recorded before change was claimed?
11. Is every claim's confidence capped by the weakest link beneath it, per K3 §3.6?
12. Does every claim carry the age of the evidence supporting it?
13. Are failed replications and dissenting studies reported, or only the converging ones?
14. Is the corpus bias assessment present, and does it state what this corpus cannot answer?
15. Is the boundary between what the studies found and what you conclude from reading them together visible?
16. Where the verdict is that no new research is needed, has that been flagged for human decision?

## 11. Common failure modes

| Failure | How to recognise it | How to prevent it |
|---|---|---|
| **Comparing questions that are not the same question** (the dominant failure) | Two similar-reading items with different scales, timeframes or bases lined up in one chart | Step 3, from the instruments, before results are seen |
| **Counting studies instead of weighting them** | "Five studies show..." with no mention of base, method or quality | Step 5. Confidence set by the strongest support, not the count |
| **False corroboration** | Several studies by the same team using the same instrument treated as independent | Check independence explicitly at step 5 |
| **Methodological difference read as a trend** | A change over time that coincides exactly with a mode or wording change | Step 6. Test the first three explanations before claiming the fourth |
| **The invisible corpus hole** | A confident synthesis with no statement of what is missing | Step 2. Ask people; unwritten studies do not appear in file listings |
| **Population blindness** | A corpus of engaged customers used to answer a question about lapsed ones | Step 8. Name what the corpus structurally cannot answer |
| **Pooling what cannot be pooled** (the signature AI failure) | An averaged percentage across studies with different bases and wording, carrying a decimal point | Step 4. State the synthesis type and refuse the pooled estimate (K3 §4.4) |
| **Recency as adjudication** | Two studies disagree and the newer one is simply adopted | Diagnose the disagreement; newer is not automatically better |
| **The established finding nobody can source** | A claim everyone repeats, traceable to one small study | Step 7. Classify every claim as replicated, single-source or contested |
| **Corpus selected around the answer** | The included set happens to exclude the studies that disagree | Criteria written first; every exclusion carries a categorical reason |
| **Synthesis as reading list** | The output is a sequence of study summaries | Organise by claim, not by study (step 9) |
| **Confidence inflation across the corpus** | Moderate findings in every study producing a high-confidence conclusion | K3 §3.6. Combination does not upgrade weak evidence unless it is genuinely independent |
| **Refusing the negative verdict** | A thin answer produced because the exercise was commissioned | "Unanswerable from this corpus" is a legitimate and frequent output (K4 §1) |

## 12. AI guardrails

Skill-specific only. Universal prohibitions are inherited from K4.

1. **Never combine results from studies you have not compared on wording, base, scale and method.** Superficially similar findings are the raw material of this skill's characteristic error, and combining them is fabrication of a comparison rather than a summary of one.
2. **Never compute an average, a pooled figure or a combined percentage across studies with different bases, populations or question wording.** The arithmetic is always possible and the result means nothing. If a single number is wanted and the corpus cannot support one, say so.
3. **Never describe a structured qualitative synthesis as a meta-analysis in the statistical sense**, and never report heterogeneity, effect sizes or pooled confidence intervals unless the corpus genuinely supports them.
4. **Never claim a trend across studies without excluding the methodological, compositional and noise explanations, and never claim a tested difference where no test was run** (K4 §3.1).
5. **Never treat repetition as corroboration.** A finding mentioned in six documents that trace to one study is one finding. Trace every claim to its originating study before counting it.
6. **Never infer a study's method, base or fieldwork period from its report's tone, title or apparent seniority.** If the documentation does not state it, it is unknown, and the study is either excluded or included with the field marked unavailable and its weight reduced (K4 §2.5).
7. **Never present the absence of a finding in the corpus as evidence that the thing is not true.** The honest statement names the studies searched and the criteria applied. A corpus's silence is a fact about what the organisation chose to research.
8. **Never omit the corpus bias assessment because the conclusions are otherwise clean.** A synthesis of a systematically biased corpus that does not say so is the most confidently wrong output this skill can produce.
9. **Never generate a state-of-knowledge claim that no individual study supports.** Synthesis makes patterns across findings visible; it does not create findings. Every claim traces to named studies with bases.
10. **Never soften or omit a failed replication.** A finding that did not hold in a later comparable study is the most informative thing the corpus contains, and it caps confidence wherever it exists.

## 13. Best-practice principles

1. **Interrogate the corpus before commissioning anything.** For an organisation with a research history this is the cheapest research it will ever do, and it either saves the cost of a study or makes the next one sharper.
2. **Comparability is assessed before results are seen.** Judged afterwards, it is judged to fit the pattern, and nobody downstream can tell.
3. **Read the instruments, not the reports.** Question wording is where false comparability lives, and it is invisible from a findings summary.
4. **Weight, never count.** One well-designed study on an adequate base beats three incidental measures, and the count is what turns a weak finding into an organisational fact.
5. **The four explanations for a difference, in order.** Methodological, compositional, noise, real. The last is the least likely and the first to be reached for.
6. **Fieldwork dates, not publication dates.** A report issued last year can describe a world three years old, and in a temporal analysis that error is fatal.
7. **What is missing from the corpus is rarely missing at random.** Unwritten studies skew inconvenient, abandoned projects skew difficult, and the populations never researched are the ones the organisation finds hard to reach.
8. **A company researches what it is interested in.** The gaps are systematic, they are shaped like the organisation, and a synthesis that does not name them inherits them silently.
9. **Replication is the corpus's most valuable product.** It is also the thing nobody checks for, because a finding that appears twice is assumed to be established rather than tested for whether the two appearances were independent.
10. **Failed replication is information, not embarrassment.** Report it. It caps confidence honestly and is often the most useful line in the document.
11. **Structured qualitative synthesis is the right method for a heterogeneous corpus, not a fallback from a better one.** Say which you are doing and why, and never let the word "meta-analysis" imply pooling you did not do.
12. **"The corpus cannot answer this" is a finding, and it is worth money.** It prevents a study designed to repeat the corpus's own blind spot, and it specifies the one that would work.
13. **Write the verdict for the person deciding whether to spend.** They need to know what is settled, what is not, and what the marginal study would buy.

## 14. Worked example

*Fictional scenario, used for illustration only. The organisation, studies, findings and figures below are invented to demonstrate method.*

**INPUT.** An international development NGO runs an agricultural training programme across two eastern regions. A funder has asked whether participant-reported benefit has improved since the programme was redesigned in 2022. The programmes director is preparing to commission a new evaluation survey. The knowledge lead is asked first whether the existing research can answer it. The corpus: nine studies between 2018 and 2025, comprising four annual participant surveys, two qualitative studies with field officers and participants, one baseline study, one externally commissioned mid-term evaluation, and one rapid assessment run during a drought year.

**PROCESS.**

*Steps 1 and 2.* The question is framed as: among enrolled smallholder participants in the two eastern regions, has the proportion reporting that the programme improved their household's food security changed between 2018 and 2025, and does any change coincide with the 2022 redesign? Inclusion criteria are written before screening. Six of the nine studies are included; the drought-year rapid assessment is excluded as an anomalous period but retained in the record because excluding it silently would remove the corpus's worst year. The missing-studies enquiry (step 2) surfaces two things no file listing shows: a 2021 participant survey that was fielded and never analysed after the field team's contract ended, and a 2023 qualitative study held by a contracted agency and never transferred. Both are recorded. The 2021 gap is directly on the redesign boundary and materially weakens any before-and-after claim, and saying so early is the most useful thing the exercise does.

*Step 3, and the finding that decides the project.* The comparability matrix is built from the instruments, not the reports. The food security item was asked in all four participant surveys, and it changed twice. In 2018 and 2019 it read "Has the programme improved your household's ability to feed itself?" with a yes/no answer. From 2020 it read "To what extent has the programme improved your household's food security?" on a five-point scale, reported as the top-two boxes. And in 2024 the base changed: earlier waves based the item on all respondents, the 2024 wave based it on those who had completed at least one training module, excluding 118 of 640. Three changes, any one of which is disqualifying for a direct comparison. The apparent time series everyone had been quoting internally (62%, 66%, 71%, 78%) is therefore not a time series. This is recorded as not comparable with the reason on every cell.

*Steps 4 to 6.* Synthesis type: structured qualitative. Quantitative pooling is not available and the reason is stated plainly, so that the funder understands it as a property of the corpus rather than a limitation of effort. Quality tiering finds the mid-term evaluation strongest (independent, documented, adequate base) and the 2020 survey weakest on this measure (the item was incidental and sat after a battery about programme problems, a preceding-context effect the report never mentions). The temporal analysis works the four explanations. The 2020 to 2024 rise of seven points on the top-two-box measure survives the wording test (wording constant across those waves) but fails the base test: recomputing 2024 on the earlier base is possible because the raw file exists, and doing so reduces the rise to two points on bases of 612 and 640, untested. The apparent improvement is largely a base artefact. Compositional analysis adds a second qualification: the participant profile shifted between 2020 and 2024 toward newer enrollees, whose reported benefit is lower, which if anything works against the measured direction.

*Steps 7 to 10, and the judgement call.* One finding does replicate: across both qualitative studies and the mid-term evaluation, participants attribute benefit to the input-supply component rather than the training component, with the training described as useful but not the reason for change. That is three independent sources, two methods, and it is stated at high confidence. The bias assessment names the corpus's structural limit: every study samples enrolled participants, and nobody has ever researched people who declined to enrol or dropped out, so the corpus cannot speak to programme reach or to why people leave, which is a more important question than the one asked. The verdict on the funder's question is **unanswerable from this corpus**: the instrument changed at the redesign boundary, the 2021 wave is missing, and the one comparable stretch shows a difference that is largely a base artefact. The temptation, and it is a real one because the funder is waiting and the internal numbers look like a rise, is to report the four-point series with a caveat. Per K4 §4.1 and K3 §4.4 it is not reported at all, because a caveated false time series is quoted without its caveat within a month. Instead the recommendation to the programmes director is that the planned survey be rescoped: hold the current wording constant from now on, restore the earlier base as a parallel measure for one wave so the series can be bridged, and add a non-participant sample, which the corpus shows has never existed. `RESEARCHER SIGN-OFF REQUIRED` on the report to the funder (K5 §2.5, §2.8).

**OUTPUT.** A synthesis reporting one high-confidence replicated finding about attribution to the input-supply component; an explicit verdict that the corpus cannot answer the change-over-time question, with the three specific reasons; a missing-studies register naming a survey never analysed and an agency study never transferred; a corpus bias assessment naming non-participants as a population never researched; and a rescoped brief for the planned survey that fixes the instrument, bridges the base and adds the missing population. The new survey still runs, and it now answers a question the corpus could not, rather than adding a fifth non-comparable point to a series that was never a series.

## 15. Advanced usage

**Standing state-of-knowledge statements.** For the five or six questions an organisation returns to, maintain the synthesis rather than rebuilding it. Each entry carries its supporting studies and confidence, and is revisited whenever a contributing study is added or changes status in the repository. This converts meta-analysis from an occasional project into the front end of every brief, and it is the mature state of the relationship between this skill and **14.03**.

**Where a subset genuinely is comparable.** A tracking series with stable wording, base and mode inside an otherwise heterogeneous corpus can be analysed quantitatively. Do it, report it as a bounded quantitative component, and keep it visibly separate from the qualitative synthesis around it, so the rigour of the subset is not read as applying to the whole.

**Re-analysis from raw data.** Where datasets survive, the strongest version of this skill re-analyses rather than compares reports: recompute each study's measure on a common base and a common definition, which removes the largest source of false comparability at a stroke. It is more work than most syntheses budget and it is often the difference between "cannot be determined" and an answer. Note that recomputation is a new analysis and carries its own documentation obligations (K4 §4.4).

**Building the next study so the corpus improves.** The most durable output of a meta-analysis is a set of instrument decisions: which wordings to freeze, which bases to hold, which populations to add. Hand these to **02.01 Survey Questionnaire Design** and **01.06 Sampling Strategy**. A corpus becomes interrogable because someone decided, at some point, to stop changing the question.

**When the standard approach does not fit.** Where the corpus is almost entirely qualitative, comparability turns on sample composition, discussion guide structure and analytical framework rather than on wording and base, and the synthesis is a comparison of themes across contexts rather than of measures. The gates are the same; the dimensions in step 3 change. Where the corpus spans markets, treat market as a comparability dimension rather than a subgroup, and have someone with local context read the findings before they are combined (K5 §2.2).

## 16. Skill chain

**Recommended previous skills:**
- **14.03 Research Repository and Knowledge Curation.** Hands over an assembled, provenance-complete corpus with status and confidence per finding. Without it, steps 2 and 3 take weeks instead of hours. 14.03 makes the corpus interrogable; this skill interrogates it.
- **13.01 Research Quality Review.** Hands over quality assessments that step 5 would otherwise have to infer from method descriptions.
- **10.01 Literature Review and Desk Research.** Runs the equivalent method on external published sources, and hands over the external evidence picture this synthesis sits beside.

**Recommended next skills:**
- **01.02 Business Problem to Research Question.** Takes the verdict, and where the corpus cannot answer the question, turns the diagnosed reason into the next brief.
- **10.04 Research Gap Identification.** Takes the corpus bias assessment and the unanswerable verdicts and turns them into a research agenda, which is where the loop closes back to the start of the library.
- **02.01 Survey Questionnaire Design** and **01.06 Sampling Strategy.** Take the instrument and sampling decisions the synthesis exposed, so the corpus becomes more comparable rather than less.
- **14.01 Insight Activation and Socialisation.** Takes the state-of-knowledge statement, which is frequently more useful to stakeholders than any single study report and needs the same activation discipline.

**Runs well alongside:**
- **10.02 Evidence Synthesis**, which synthesises external published evidence where this skill synthesises the organisation's own corpus, and **10.03 Multi-Source Research Synthesis**, which integrates the two when a question needs both.
- **13.04 Bias Detection**, run against the corpus assembly and the inclusion decisions, which are unusually easy to steer.
- **K3 §3**, which supplies the six factors that set confidence per claim, and **K2 §4.4**, which governs how divergence between evidence streams is reported rather than averaged away.

---
A Yazi Supplied Skill and resource.
