---
name: evidence-synthesis
description: >
  Combines findings from multiple studies into a single assessment of what is
  known, weighted by evidence quality rather than study count. Use for
  "synthesise these studies", "what does the evidence say overall", "combine
  these findings", "pull these reports into one view", "how strong is the
  evidence for this", "three studies say one thing and one says another",
  "weight of evidence", "strength of evidence statement".
category: 10 Desk Research and Evidence Synthesis
ref: "10.02"
tier: 1
inherits: [K2, K3, K4, K5]
---

# Evidence Synthesis

## 1. One-line description

A method for combining findings from multiple studies into a single assessment of what is known, organised claim by claim, weighted by evidence quality rather than by study count, with disagreement diagnosed rather than averaged and dissenting sources kept visible in the output.

## 2. What this skill is used for

**The research problem it solves.** A set of appraised sources is not an answer. The distance between "here are eleven studies" and "here is what is known, and how firmly" is where synthesis fails, and it fails in a small number of predictable ways. Studies get counted rather than weighted, so three weak sources agreeing appear to outrank one strong source. Findings that share a word get pooled although they measure different constructs, which is the commonest silent error in the whole discipline because nothing in the output reveals it. The evidence that happened to be published is treated as the evidence that exists. And disagreement is resolved by averaging, by quietly preferring the convenient result, or by omission. This skill supplies the working method: a claim-by-source matrix as the physical artefact, a commensurability test that runs before anything is combined, quality weighting, an explicit heterogeneity judgement about whether synthesis is legitimate at all, and a strength-of-evidence statement per claim that a reader can inspect.

**Where it sits in the research lifecycle.** Downstream of source discovery and appraisal, upstream of insight. It converts an appraised evidence set into a set of positions with confidence attached. It also runs late in a project, when a primary finding needs to be read against the wider evidence base, and at the start of a programme, where the synthesis decides what still needs measuring.

**Typical use cases.**
- Turning a completed literature or desk review into a defensible statement of what is established.
- Reconciling several studies of the same question that report different results.
- Producing the evidence base for a policy position, a business case or a clinical or operational guideline.
- Assessing how strong the case for an assumption really is before a decision rests on it.
- Combining several waves or several markets of a repeated study into one picture.
- Assembling the evidence section of a proposal, where the credibility of the argument depends on the weighting being visible.

**Who uses it.** Research directors and senior analysts building a point of view from an evidence set; policy and evaluation researchers; insight leads consolidating a fragmented internal evidence base; applied academic researchers writing a review that has to state strength of evidence, not just describe the literature.

## 3. When to use it

- You have two or more independent studies bearing on the same question and need one assessment rather than a list.
- Sources disagree and someone has to say what the evidence supports.
- A decision rests on an assumption and you need to know how well evidenced it actually is.
- A body of research has accumulated over time and nobody has ever consolidated it.
- You need to state confidence in a claim, not just report that studies exist.
- A stakeholder is citing a count ("five studies show...") and the count is doing work the quality of the studies will not support.
- A primary finding has landed and needs contextualising against what was already known.
- You are being asked whether the existing evidence is sufficient, and the answer requires more than an inventory.

## 4. When NOT to use it

- **The sources have not been found, verified and appraised yet.** This skill starts where that finishes. Locating sources, verifying they exist and say what they are claimed to say, tracing citation chains to their origin, and tiering sources on quality all belong to **10.01 Literature Review and Desk Research**. That skill finds and appraises sources; this one combines what they say. Running synthesis on an unappraised set produces a confident average of unknown material.
- **The evidence set is heterogeneous in kind rather than in result.** Client-supplied internal figures, operational and behavioural data, expert input and previous studies of different types cannot be pooled, and forcing them into one matrix creates false equivalence between sources of different epistemic status. That is **10.03 Multi-Source Research Synthesis**, which handles source types that will not combine. This skill assumes sources that are at least candidates for comparison.
- **Quantitative pooling is intended.** Calculating a summary effect size, weighting by inverse variance, testing statistical heterogeneity or producing a forest plot is meta-analysis, a different method with prerequisites this one does not impose. Hand over to **15.14 Advanced Systematic Review and Meta-Analysis**. Do not describe a narrative synthesis as a meta-analysis, and never present a hand-averaged figure as a pooled estimate.
- **All the sources trace back to one original.** Where citation-chain tracing shows five documents restating a single study, there is nothing to synthesise. Report the one study with its limitations. Presenting it as a converging body of evidence is the most damaging misuse of this method, because the false convergence is invisible downstream.
- **The studies measure different constructs and no defensible mapping exists.** If the commensurability test in Step 2 fails, the correct output is a set of separately reported findings with the incompatibility named, not a synthesis. Forcing it produces a claim about a construct that no study measured.
- **The decision needs a causal claim and the evidence is observational.** Synthesising twelve correlational studies does not manufacture causal warrant, however consistent they are (K4 §3.2). Consistency across studies that share the same confounder is consistency in the confounding. Where causality is what the decision turns on, see **06.04 Experiment and A/B Test Analysis** or design accordingly through **01.04 Research Method Selection**.
- **The task is turning synthesised claims into business meaning.** What the claims imply for an organisation, and what should be done, is **08.01 Finding to Insight Development** and the reporting skills in category 12. This skill stops at the assessment of what is known and how firmly.
- **The stakes are high and the evidence base is small, interested and homogeneous.** Where three sources exist, all commissioned by parties with the same interest, a synthesis lends them a collective authority none of them has individually. Report the evidence base as it is, and say what independent evidence would be needed. `RESEARCHER DECISION REQUIRED` (K5 §2.5).

## 5. Required inputs

**Required.**
- **An appraised source set,** with each source carrying its method, population, period, base and a quality tier. If quality has not been assessed, stop and run 10.01 §Step 6 first. Synthesis without appraisal is vote counting with extra steps.
- **The question or claim set the synthesis has to serve.** Synthesis is organised around claims, so the claims have to exist before the matrix can be built. If only a topic is supplied, derive a candidate claim set and get it confirmed.
- **The extracted findings themselves, at claim level, with bases and question wording where quantitative.** A source summary is not enough. You cannot test commensurability against a summary, because the summary is where the construct detail was lost.

**Optional, and what each one adds.**
- **The construct definition the decision depends on.** Converts commensurability from a judgement into a test with a criterion. Without it, you are comparing sources to each other rather than to the thing that matters.
- **Full method sections or technical appendices.** Let you distinguish methodological disagreement from genuine disagreement, which is usually impossible from headline findings alone.
- **Knowledge of who commissioned each study and whether a study series exists.** Makes reporting bias assessable rather than assumed. An unreleased wave in a commissioned series is a specific, checkable signal.
- **Registered protocols, trial registrations or pre-analysis plans where they exist.** Allow outcome reporting bias to be detected directly by comparing what was planned with what was reported.
- **Access to the raw data behind any source.** Turns a reported figure into a checkable one and occasionally reveals that two sources are the same data analysed twice.
- **The decision timeline and stakes.** Determines whether a moderate-confidence synthesis is sufficient or whether the honest answer is that the question is not yet settled.

## 6. Questions to ask before starting

1. **What claims does the decision actually need settled?** Determines the rows of the matrix and prevents synthesising everything the sources happen to discuss. Default if unanswered: derive three to six claims from the review questions and mark them as assumed.
2. **Is a directional assessment sufficient, or does a number have to survive scrutiny?** A directional synthesis tolerates construct approximation; a numeric one does not. Default: assume directional, and flag every place a number is being asked to carry more than its source set supports.
3. **Which sources, if any, are already believed inside the organisation?** These carry political weight regardless of their tier, and a synthesis that contradicts one without addressing it will be dismissed rather than argued with. Default: ask, because discovering this after the fact forces a rewrite.
4. **Is the evidence base likely to be complete, or is there a known reason studies would be missing?** Commissioned research in a commercial category, or an area where negative results are unwelcome, has a systematically visible half. Default: assume incompleteness and say so.
5. **Has the thing being studied changed over the period the sources cover?** If so, temporal heterogeneity is not noise to be smoothed but the finding itself. Default: order the evidence by date of data before combining anything, and look at it.
6. **What would the strongest opponent of this synthesis say?** Identifies where dissent must be retained rather than resolved. Default: run the disconfirmation pass in Step 10 regardless.
7. **Who reads the output, and will the confidence language survive the summary?** A synthesis whose caveats fall off at the executive summary has failed (K3 §7). Default: write the summary line for each claim yourself, at the correct confidence, rather than letting someone else compress it.

## 7. Step-by-step methodology

**Step 1. Fix the unit of synthesis as the claim, and write the claim set.**
Synthesis organised by source produces an annotated bibliography; synthesis organised by claim produces an assessment. Write each claim as a single testable proposition naming a population, a construct, a direction and, where relevant, a magnitude and a period. "Price matters to switchers" is not a claim. "Among domestic energy customers who switched supplier in the last two years, price was the most frequently cited reason for switching" is. Keep claims narrow enough that a source either speaks to them or does not; a claim that every source partially addresses is really three claims. Six to twelve claims is a workable set for most syntheses. Then, for each claim, write in advance what evidence would count as supporting it and what would count against it. Doing this before looking at the matrix is the single cheapest guard against fitting the criteria to the result. *Correct result: a numbered claim register, each claim with its supporting and disconfirming evidence specified in advance.*

**Step 2. Test commensurability before combining anything.**
This is the step that is most often skipped and does the most damage when it is. For every source against every claim, establish five things: the construct definition in use, the operationalisation (instrument, exact item wording, scale, response format), the referent (a brand, a category, a specific event, a general disposition), the population and frame, and the measurement period. Two sources both reporting "trust" may be measuring competence and benevolence respectively, which move independently. Two reporting "switching" may mean an intention, a completed transaction and a lapse in usage. Record the answer as **commensurable** (same construct, comparable operationalisation), **partially commensurable** (same construct, materially different operationalisation, comparable in direction but not in level), or **incommensurable** (different constructs, must not be combined). Where the verdict is partial, the synthesis may use direction but not magnitude, and this must be stated in the output rather than held in your head. *Correct result: a commensurability verdict recorded for every source-claim pair, with the specific difference named where it is not full.*

**Step 3. Build the claim-by-source matrix as the working artefact.**
Claims as rows, sources as columns. Each cell records: what the source says about the claim (supports, contradicts, mixed, does not address), the finding in the source's own terms with its base and locator, the commensurability verdict from Step 2, and the source's quality tier. The matrix is not a presentation object, it is the instrument. Filling it forces you to notice the things a written synthesis conceals: that the two sources you thought agreed are answering different questions, that the claim everybody accepts is supported by one column, that six of your eight columns are empty for the claim the decision actually turns on. Keep it live and keep it in the appendix of the final output, because it is what makes the synthesis checkable by someone who was not there. *Correct result: a complete matrix with no cell left ambiguous, and an immediate visual read of where the evidence is thick and where it is thin.*

**Step 4. Weight by quality and independence, never by count.**
Read each claim's row and ask what the evidence would look like if the strongest source were the only one present. A single well-designed study on an adequate base, measuring the construct directly, is stronger evidence than three small, thinly reported studies agreeing with each other, and considerably stronger than three that share a common weakness. Agreement between sources with a shared flaw is not corroboration: four studies using the same leading question wording will agree, and their agreement is evidence about the question wording. Practically: identify the strongest source per claim, ask what the weaker sources add beyond it (independent population, independent method, independent commissioning, or nothing), and record only genuine additions as corroboration. Where sources share a dataset, an author group, a sponsor or an instrument, note the dependency and count them as one line of evidence with multiple expressions. *Correct result: for each claim, a statement of the strongest evidence, what independently corroborates it, and what merely repeats it.*

**Step 5. Refuse the vote-counting shortcut explicitly.**
Vote counting is deciding by majority of studies: four found an effect, two did not, therefore the effect is real. It is wrong for three separate reasons and each has to be checked. First, it ignores size and precision, so a study of 3,000 and a study of 60 cast equal votes. Second, it treats a non-significant result as evidence of no effect, when in an underpowered study it is evidence of nothing at all; the correct reading of a null result requires knowing what effect the study could have detected. Third, it discards direction and magnitude, which is where the actual information lives: six studies all showing a small positive effect, only two of them individually significant, is a more coherent body of evidence than three large positive and three large negative results with four of them significant. Replace the vote with a distribution: for each claim, lay out every source's direction and magnitude alongside its base and quality, and read the shape. *Correct result: no claim resolved by a study count anywhere in the output, and every claim's evidence displayed as a pattern of direction and magnitude rather than a tally.*

**Step 6. Assess heterogeneity, and decide whether to synthesise at all.**
Separate three kinds. **Conceptual heterogeneity** is variation in populations, constructs, interventions or outcomes; it is diagnosed in Step 2 and is the kind that most often forbids synthesis. **Methodological heterogeneity** is variation in design, mode, sampling and quality; it usually permits synthesis with the variation reported as a moderator. **Result heterogeneity** is variation in the findings themselves beyond what the bases would lead you to expect. Then make an explicit decision, recorded with reasons: synthesise across the whole set; synthesise within defensible subgroups (by market, by period, by method type) and report the subgroups separately; or do not synthesise, and report the sources as a set of context-specific findings with the incompatibility explained. The third option is a legitimate and frequently correct outcome, and a synthesis that never reaches it is not exercising judgement. Where subgroup synthesis is chosen, the subgroups are defined before the results are inspected, not after, or the exercise becomes a search for a favourable cut. *Correct result: a heterogeneity verdict per claim with reasons, and any subgrouping defined in advance.*

**Step 7. Assess what is missing from the evidence base.**
The available evidence is not a random sample of the evidence produced, and treating it as one is an error of the same class as treating a self-selected sample as representative. Work through the mechanisms. **Publication bias**: results consistent with the expected direction are published more often and faster. **Outcome reporting bias**: within a study, measured outcomes that disappointed are not reported, which is detectable by comparing a registered protocol or a stated method section with what appears in the results. **Commissioning bias**: in commercial and advocacy research, studies whose results were unwelcome are simply never released, so the visible set is a filtered set with no trace of the filter. **Language and access bias**: the evidence in languages you did not search, or behind access you do not have, is not absent, it is invisible. **Citation bias**: findings that support a prevailing view are cited more and therefore found more. Then look for the diagnostic patterns: an evidence base where every source points the same way and every source has the same interest; a set where the smaller studies show larger effects than the larger ones; a series with a visible gap where a wave should be. Record the assessment as a domain in the strength rating rather than as a caveat sentence. *Correct result: an explicit statement of what would be missing if it existed, why, and what that does to each claim's confidence.*

**Step 8. Adjudicate disagreement by diagnosed cause.**
Where sources conflict, name the cause before judging the conflict, working through six in order: **definitional** (different constructs under the same word), **methodological** (stated importance against observed behaviour, different modes, different question wording, different analysis choices), **sampling** (different populations, frames or recruitment), **temporal** (different periods, and the thing genuinely changed), **geographic** (different markets or settings), and only then **genuine** (comparable studies, incompatible results). The first five account for the large majority of apparent contradictions, and each has a different consequence: a definitional conflict means the claim was wrong, not the sources; a temporal conflict is often the most valuable finding in the synthesis. Where the conflict is genuine, report both positions with their evidence profiles and say which is better evidenced and why. Never split the difference, and never resolve a conflict by preferring the source that suits the argument (K4 §4.1). Where you cannot adjudicate, say so and leave both positions standing; an honestly unresolved conflict is a usable output and a silently resolved one is not. *Correct result: every conflict named, assigned a cause, and either resolved with stated reasons or explicitly left open.*

**Step 9. Assign a strength-of-evidence rating to each claim.**
Rate the body of evidence, not the individual sources, on six domains: **quantity and independence** (how many genuinely independent lines of evidence), **quality** (the tier profile of the contributing sources), **consistency** (do they point the same way once heterogeneity is accounted for), **directness** (does the evidence measure the claim, or an adjacent construct, or a proxy), **precision** (are the bases adequate for the magnitude being claimed), and **risk of missing evidence** (Step 7). Start from the level the strongest contributing evidence would support alone, then downgrade for each domain that materially bites. Upgrade only for a large and consistent effect, a dose-response pattern, or convergence across genuinely different method types, and never by more than one level. Map the result to K3: **high** where multiple independent adequate sources converge with no material contradiction, **moderate** where one strong source carries it or convergence comes with a real caveat, **low or hypothesis** where the evidence permits the claim but does not establish it. A claim can be no more confident than its weakest materially biting domain, and directness is the domain most often waved through. *Correct result: a rating per claim with the domain that capped it named, so a reader can see why it is not higher.*

**Step 10. Write the synthesis with dissent visible, then run a disconfirmation pass.**
Each claim is written as a statement at its assigned confidence, in the K3 register for that level, with its supporting sources referenced per K2 §4.3, the dissenting sources named rather than omitted, and the reason for the dissent given. A claim whose dissent has been dropped between the matrix and the write-up has been silently promoted. Then run the disconfirmation pass: for each high-confidence claim, state what evidence would overturn it and check the set once more for anything approaching that. This takes under an hour and catches the synthesis that assembled itself around a prior. Finally, write the claims that failed: the ones where the evidence was too heterogeneous, too dependent or too thin. These belong in the output, not in a discard pile, because they are the specification for what to research next. *Correct result: a claim register in which every claim carries evidence, confidence and its dissent, and a companion list of what could not be established and why.* `RESEARCHER SIGN-OFF REQUIRED` where the synthesis will support a material decision (K5 §2.5).

## 8. Analytical framework

The synthesis is built on one chain, applied per claim, and one weighting rule applied within it.

**The claim chain:**

    Claim → Source set → Commensurability verdict → Independence check →
    Evidence profile (direction, magnitude, base, tier) → Heterogeneity verdict →
    Adjudicated position → Strength of evidence → Retained dissent → Residual gap

Every arrow is a place the synthesis can fail and each is separately inspectable. Two arrows carry most of the risk. *Commensurability verdict* is where sources that measure different things are either caught or silently pooled, and the failure leaves no trace in the output, which is what makes it the most dangerous step. *Independence check* is where repetition is either distinguished from corroboration or mistaken for it, and it is the step that determines whether the strength rating means anything.

**The weighting rule.** Evidence is weighted, not counted, and weight is a product rather than a sum:

    weight = directness × quality × independence

The formulation is conceptual, not arithmetic; do not compute it. Its use is that it is multiplicative, so a zero anywhere is a zero overall. A tier A source measuring the wrong construct contributes nothing, whatever its sample size. Three sources sharing one dataset have the independence of one. A perfectly relevant, perfectly independent source with an uninterpretable method contributes nothing either. This is why quality weighting cannot be done by assigning scores and adding them up: the domains are not compensatory, and any scheme that lets strength on one dimension offset a fatal weakness on another will eventually average a fabrication into a finding.

## 9. Output format

**1. Question and claim set.** The decision served, the claims assessed, and the scope boundary (populations, periods, markets, constructs).

**2. Evidence base description.** How many sources, of what types, from what period, with the tier profile. States plainly what the set is and how it was assembled, and whether it inherits a purposive or a systematic search.

**3. Claim register.** The core of the output.

| ID | Claim | Supporting sources | Dissenting sources | Independent lines of evidence | Commensurability | Strength of evidence | Capping domain |
|---|---|---|---|---|---|---|---|

**4. Claim-by-claim assessment.** For each claim: the statement at its confidence level in K3 language, the evidence profile behind it, the dissent and its diagnosed cause, and what would change the assessment.

**5. Disagreement register.** Every conflict, its diagnosed cause among the six, and its resolution or the reason it stays open.

**6. What is missing from the evidence base.** The reporting and publication bias assessment, the access and language limits, and what the pattern of available evidence suggests about the unavailable evidence.

**7. Claims that could not be synthesised.** Where heterogeneity, dependence or thinness prevented it, with the reason in each case.

**8. The claim-by-source matrix,** in full, as an appendix.

**When the evidence is thin, the format must not force fabrication (K4 §1).** A claim supported by one source is reported as supported by one source, at the confidence that warrants, with the register showing a single column populated. The matrix is never padded with sources that do not address a claim, and "does not address" is a legitimate and informative cell value. Where a claim has no adequate evidence, it stays in the register with an empty row and a note, because the empty row is the finding. A synthesis that concludes the evidence base cannot settle the question is a successful synthesis, and it is considerably more useful than a confident average.

## 10. Quality checks

Run before anything is presented. These sit on top of K4 §8.

1. Is every claim a single proposition with a named population, construct and direction, rather than a topic?
2. Has a commensurability verdict been recorded for every source-claim pair, and is every partial verdict reflected in how the claim is worded?
3. Has every apparently converging set been checked for shared datasets, shared authorship, shared sponsors and shared instruments?
4. Has any claim been resolved by counting studies, anywhere, including in a summary sentence?
5. Is any null result being read as evidence of no effect without regard to what the study could have detected?
6. Is the heterogeneity verdict recorded per claim, with reasons, and were any subgroups defined before results were inspected?
7. Does every strength rating name the domain that capped it?
8. Does any claim carry a confidence higher than its least direct piece of contributing evidence supports?
9. Are dissenting sources named in the write-up for every claim where they exist, not only in the matrix?
10. Has the missing-evidence assessment been done, and does it appear as a rating domain rather than a closing caveat?
11. Is every number in the synthesis present in a source, with no figure averaged, interpolated or reconciled into existence?
12. Would a reader be able to reconstruct any claim's rating from the matrix alone?
13. Does the output say which claims could not be synthesised, and why?
14. Has the confidence language survived compression into the summary at the same level it carries in the body?

## 11. Common failure modes

| Failure | How to recognise it | How to prevent it |
|---|---|---|
| **Vote counting** | "Most studies find..." with no reference to size or quality | Replace the tally with a direction-and-magnitude distribution (Step 5) |
| **Construct pooling** | Two figures combined because they share a word | Commensurability verdict before any combination (Step 2) |
| **False corroboration** | Several sources agreeing, all sharing a dataset, sponsor or instrument | Independence check per claim (Step 4) |
| **Null read as absence** | A non-significant result counted as evidence against | Ask what effect the study could have detected before reading its null |
| **Averaging disagreement** | A single figure with no trace of the conflict behind it | Adjudicate by diagnosed cause and report both positions (Step 8) |
| **Silent heterogeneity** | Studies from four markets and six years combined without comment | Explicit heterogeneity verdict with reasons (Step 6) |
| **Post hoc subgrouping** | A subgroup analysis that happens to isolate the favourable result | Define subgroups before inspecting results |
| **The published set treated as the whole set** | No mention of what might be missing | Missing-evidence assessment as a rating domain (Step 7) |
| **Compensatory scoring** | Quality scores summed, so a fatal flaw is offset by a strength | Weighting is multiplicative in effect, not additive (§8) |
| **Dissent lost in the write-up** | The matrix shows a contradicting source, the text does not | Dissent is a required field in the claim register (Step 10) |
| **Confidence inflation at summary** | Body says "appears to", summary says "is" | Write the summary line yourself at the assigned level (K3 §7) |
| **Synthesis of one** | A confident body-of-evidence statement resting on one original | Trace citation chains before counting (10.01 Step 5) |
| **The completed matrix** | Every cell filled, nothing marked "does not address" | Empty cells are findings; a full matrix is a warning sign |
| **Meta-analysis by hand** | An average of percentages presented as a pooled estimate | Averaging across studies is prohibited here; see 15.14 |

## 12. AI guardrails

Skill-specific. The universal prohibitions in K4 apply in full and are not repeated.

1. **Never produce a summary figure by combining figures across sources.** No mean of percentages, no midpoint of a range across studies, no weighted blend computed informally. A number in the output is a number in a source, attributed to that source. If the decision needs a pooled estimate, the answer is that pooling requires a method this skill does not contain (15.14), not an average.

2. **Never let a shared word license a combination.** Where two sources use the same term, the default assumption is that they mean different things until the operationalisations have been compared. State the construct check you performed. Where the source does not report enough detail to run the check, the verdict is unknown, not commensurable, and the claim is worded accordingly.

3. **Never present the count of supporting sources as the strength of the evidence,** in the text, in a table column, or in a summary sentence. If a count appears anywhere, it must appear beside the independence verdict and the tier profile, or it will be read as the finding.

4. **Never manufacture a source to balance a claim.** The impulse to show both sides can produce a plausible-sounding dissenting study that does not exist. Where no dissenting evidence was found, say that none was found in the sources assembled, and note whether the search would have surfaced it.

5. **Never infer a study's method, base or population when the source does not report it.** An unreported base is unreported. Do not estimate it from the reported percentages, do not describe a sample as representative because the source calls itself national, and do not fill a matrix cell by inference. Mark the cell insufficiently reported, which is itself a quality finding.

6. **Never resolve a conflict by preferring the source that fits the emerging story,** and never resolve one silently. Where adjudication is not possible, both positions stand. A synthesis with an unresolved conflict in it is doing its job.

7. **Never upgrade a partially commensurable comparison into a magnitude claim.** Direction may transfer where operationalisations differ; levels do not. "Both studies find price ranks above service" is available; "price matters to around 60% across both" is not.

8. **Never omit the missing-evidence assessment on the grounds that it cannot be quantified.** It is a qualitative judgement about a real mechanism, and leaving it out implies the evidence base is complete. State what you can establish about how the set was filtered, and where you cannot establish anything, state that.

9. **Where a strength rating rests on sources you could not fully appraise, cap the claim at low and name the appraisal gap** (K3 §3.4). A body of evidence cannot be rated higher than its constituent evidence can be inspected.

10. **Never let the register's structure generate content.** If a claim has no dissenting sources, the field says none identified, not a hedge invented to fill it. Format is not evidence (K4 §1).

## 13. Best-practice principles

- **The claim, not the source, is the unit of work.** Everything follows from this. Organising by source produces summaries; organising by claim produces an assessment, and forces the questions about commensurability and independence that source-by-source writing never raises.
- **Do the commensurability check before you look at the results.** Once you know which sources agree, the temptation to find them comparable is strong and largely unconscious. Sequence protects judgement more reliably than intention does.
- **Ask what the strongest source alone would support, then ask what the rest add.** This single question dissolves most vote-counting instincts, and it usually reveals that a large evidence base is one study with company.
- **Agreement between sources that share a flaw is evidence about the flaw.** Independence is a property of design, sponsorship, data and instrument, not of publication. Two reports from different organisations analysing the same public dataset are one line of evidence.
- **A non-significant result in a small study is not a negative finding.** Read every null against what the study was capable of detecting, and where that cannot be established, treat the study as uninformative for that claim rather than as evidence against.
- **Heterogeneity is information before it is a problem.** When results differ by market, by period or by method, the pattern of difference is frequently the most useful output of the synthesis, and smoothing it away discards the finding to preserve the format.
- **The absent evidence has a shape.** In commissioned research especially, what was never published is not random: it is the results that disappointed whoever paid. An evidence base where every study agrees and every study shares an interest should be reported as one source of evidence about that interest.
- **Report the dissent even when it is weak,** with its weakness stated. A reader who later finds the contradicting study you omitted will discount the whole synthesis, and they will be right to.
- **Confidence is capped by directness more often than by volume.** Large, consistent, well-conducted evidence about an adjacent construct is still evidence about an adjacent construct, and this is the domain most frequently waved through in practice.
- **Say which claims failed.** A synthesis that reports only what it could establish gives no information about the shape of the evidence base, and its silence will be read as coverage.
- **Keep the matrix and ship it.** It is the difference between a synthesis a colleague can check and a synthesis they have to trust, and it costs nothing once it has been built.
- **A short synthesis over strong evidence beats a long one over weak evidence.** Length signals thoroughness and is routinely mistaken for weight.

## 14. Worked example

Generic fictional scenario, public health.

**INPUT**

A regional health authority is deciding whether to fund a text-message reminder programme to improve attendance at routine screening appointments. Nine sources have been assembled and appraised through 10.01: four published evaluations of similar programmes in other regions, two academic studies, two internal evaluations of small local pilots, and one report from a national body reviewing attendance interventions. The team wants a single answer: does this work, and by how much.

**PROCESS**

*Step 1.* Four claims are written. (C1) Text reminders increase attendance at routine screening appointments among invited adults. (C2) The effect size is large enough to change service planning, defined in advance as a sustained increase above five percentage points. (C3) The effect persists beyond the first invitation cycle. (C4) The effect is similar across age and deprivation groups. Supporting and disconfirming evidence is specified for each before the matrix is opened.

*Step 2.* Commensurability is tested, and it fails in two places immediately. Three sources measure "attendance" as attendance at any point within a twelve-month window; two measure attendance at the specific appointment offered. These are different constructs, and the first is systematically higher. One further source measures reported intention to attend, which is incommensurable with both and is excluded from C1 and C2 entirely.

*Step 3.* The matrix is built. C1 has seven populated cells. C3 has two. C4 has one, from a source whose subgroup bases are not reported.

*Step 4 and 5.* The counting instinct says seven sources support C1. The independence check finds that the national body's report is a review of three of the four published evaluations, so it is not an eighth line of evidence, it is a restatement of three already in the matrix. That leaves five independent lines. The two local pilots are small, single-site, and used the same reminder wording drawn from the same national template, so their agreement carries less than it appears to.

*The judgement call.* One of the four published evaluations reports no significant effect, and the temptation is to record C1 as supported four to one. Instead the null is examined: the study had 240 invited adults and could not have detected an effect below roughly ten percentage points. It is not evidence against the claim; it is uninformative for it. This is recorded explicitly rather than being quietly dropped, because a reader who finds it later needs to see that it was considered.

*Step 6.* Heterogeneity. The evaluations span six years and four regions with different baseline attendance rates. A subgroup view defined in advance, splitting by baseline attendance above and below the median, shows the effect concentrating where baseline attendance is low. This is reported as a moderator, not smoothed away.

*Step 7.* Missing evidence. Both local pilots that were written up reported positive results, and the service ran four pilots. Two were never evaluated. This is recorded as a specific, named reporting gap rather than a general caveat, and it caps C2.

*Step 9.* Ratings. C1: **moderate to high**, five independent lines, consistent direction, capped by directness because the attendance construct differs across sources. C2: **low**, because the magnitude claim requires commensurable measurement that the set does not provide, and because the local evidence base is visibly filtered. C3: **low**, two sources, one of them the weaker construct. C4: **not established**, one source, subgroup bases unreported.

**OUTPUT**

A claim register stating that text reminders increase screening attendance, with the direction well supported and the magnitude not, that the effect is larger where baseline attendance is lower, that persistence beyond one cycle is essentially untested, and that equity of effect across deprivation groups is unknown. The funding recommendation is separated from the evidence assessment: the evidence supports proceeding, and does not support the business case's assumed uplift figure, which came from the single most favourable pilot.

`RESEARCHER SIGN-OFF REQUIRED` on the evidence assessment feeding the funding decision (K5 §2.5), and `RESEARCHER REVIEW RECOMMENDED` on the equity claim, where an absence of evidence in the synthesis must not be read as an absence of an equity effect.

## 15. Advanced usage

**Synthesis as a maintained asset.** Where a question recurs, keep the claim register and the matrix as living documents with a scheduled re-run. The value is the diff: which claims strengthened, which were overturned, which new claims appeared. A register that has been updated three times is worth considerably more than three separate syntheses, because it carries the history of how the assessment changed. See **14.03 Research Repository and Knowledge Curation**.

**Sensitivity analysis on the synthesis itself.** Re-run the strength ratings twice: once excluding every source below tier B, and once excluding every source with a declared interest in the result. Where a claim's rating survives both, say so, because that is a strong statement about robustness. Where it does not, the claim is resting on exactly the evidence a sceptic will attack first, and it is better to know before the meeting.

**Synthesising when the constructs are irreconcilable.** Where commensurability fails across the whole set, the productive move is to synthesise one level up, at the level of mechanism rather than measurement. Five studies measuring five different outcomes may still converge on the same causal story, and that convergence is reportable as long as it is labelled as a synthesis of mechanism, not of effect. This is a weaker claim, it is often the only honest one available, and it is far more useful than an incommensurable average.

**Where the standard approach does not fit.** For rapidly moving topics, weight recency explicitly as a domain and state a shelf life for the synthesis. For evidence bases dominated by commissioned commercial research, expect the missing-evidence assessment to be the most important section and word every claim as what the released evidence shows. For fields with an established measurement standard, commensurability is largely solved and the effort shifts to independence and reporting bias.

## 16. Skill chain

**Recommended previous skills:**
- **10.01 Literature Review and Desk Research.** Hands over the verified, appraised, tiered source set and the claim-level extraction table this skill reorganises. Without it there is nothing to weight.
- **01.02 Business Problem to Research Question.** Hands over the decision that determines which claims are worth synthesising.

**Recommended next skills:**
- **08.01 Finding to Insight Development.** Takes the claim register and works out what the established claims mean for the organisation.
- **10.04 Research Gap Identification.** Takes the claims that could not be established and turns them into a prioritised research specification.
- **12.04 Research Evidence Integration** and the reporting skills in category 12, which carry the claim register and its confidence levels into a deliverable.
- **15.14 Advanced Systematic Review and Meta-Analysis,** where quantitative pooling is warranted and the prerequisites are met.

**Runs well alongside:**
- **10.03 Multi-Source Research Synthesis,** where the evidence set includes source types that cannot be pooled.
- **13.04 Bias Detection,** which audits the synthesis for selection and confirmation effects.
- **13.03 AI Output Verification,** which checks the claim register against its sources.

---
A Yazi Supplied Skill and resource.
