---
name: multi-source-research-synthesis
description: >
  Combines fundamentally different kinds of evidence in one assessment: primary
  research, desk research, client-supplied internal data, operational and
  behavioural data, expert input and previous studies. Use for "pull all the
  evidence together", "combine our data with the research", "the client's
  numbers disagree with the survey", "triangulate these sources", "integrate
  internal data and primary research", "we have a survey, a data extract and
  some interviews", "which source do we believe".
category: 10 Desk Research and Evidence Synthesis
ref: "10.03"
tier: 1
inherits: [K2, K3, K4, K5]
---

# Multi-Source Research Synthesis

## 1. One-line description

A method for combining evidence of fundamentally different kinds, primary research, desk sources, client-supplied internal figures, operational and behavioural records, expert input and previous studies, into one assessment that keeps each source's epistemic status and provenance visible rather than flattening them into a single voice.

## 2. What this skill is used for

**The research problem it solves.** Real projects almost never rest on one kind of evidence. A typical assessment draws on a survey you ran, a data extract the client supplied, a set of transaction or usage records, two published reports, a study from three years ago, and the considered view of people who have worked in the category for a decade. These cannot be pooled, and the failure mode is not that people try to pool them statistically. It is subtler and far more common: they get written into one narrative in one voice, and by the second paragraph the reader can no longer tell which sentence rests on a measured finding, which on a figure somebody in finance pulled from a system, and which on an experienced person's impression. Definitions that never matched are quietly treated as matching. Time periods that differ by eight months are set beside each other. A client's internal number and a primary research finding disagree, and the conflict is resolved by whichever party is in the room. This skill installs a source inventory built before any synthesis, an explicit statement of what each source type can and cannot establish, an alignment step on definitions and periods, a triangulation test that distinguishes real convergence from coincidence, and a named procedure for the client-data conflict.

**Where it sits in the research lifecycle.** Late, at the point where separate workstreams have to become one answer, and again at the start of a programme, where taking stock of every available source determines what primary work is actually needed. It is the skill that stands between a set of workstreams and a single deliverable.

**Typical use cases.**
- Integrating a commissioned study with the organisation's own operational data.
- Building an assessment where the budget covered only part of the question and the rest has to come from what exists.
- Producing a programme or performance evaluation drawing on monitoring data, beneficiary research and staff knowledge.
- Reconciling a client's internal reporting with what your fieldwork found.
- Assembling the evidence base for a business case that mixes market data, internal financials and primary research.
- Bringing expert or stakeholder input into an evidence base without letting it acquire the authority of measurement.

**Who uses it.** Research directors and consultants running mixed-evidence projects; evaluation and impact researchers; insight leads inside organisations who hold both the research and the operational data; strategy teams assembling a case from whatever is available.

## 3. When to use it

- Your evidence base contains more than one kind of source and they cannot be pooled.
- A client or internal stakeholder has supplied figures that need to sit alongside research you produced.
- Operational or behavioural data exists for part of the question and primary research for another part.
- Two sources of different kinds disagree, and the decision needs a position.
- An evaluation must combine monitoring records, participant research and practitioner judgement.
- You are being asked to state what "all the evidence" says and the evidence is genuinely heterogeneous.
- Expert or stakeholder opinion is in the mix and needs a defensible place in the assessment.
- A previous study is being cited alongside current work and its comparability has not been examined.

## 4. When NOT to use it

- **The sources are comparable studies of the same kind.** Where the evidence set is several studies that are at least candidates for direct comparison, the correct method weights them against each other claim by claim, and that is **10.02 Evidence Synthesis**. That skill synthesises comparable studies; this one handles heterogeneous source types that cannot be pooled. Using this skill on a comparable set adds inventory overhead and loses the weighting discipline.
- **The sources have not been individually appraised.** A published source still needs verification, provenance tracing and quality appraisal before it enters the inventory, and that is **10.01 Literature Review and Desk Research**. A dataset still needs its own examination before its figures are usable. This skill integrates appraised sources; it does not appraise them.
- **The task is analysing one dataset well.** If the real question is what the survey shows, or what the transaction records show, run the analysis skill for that data type. Wrapping a single-source analysis in an integration frame gives it a spurious appearance of triangulation.
- **The sources genuinely cannot be aligned on definition or period, and the decision needs a single number.** Where a client's "active customer", a survey's "regular user" and a records system's "account with a transaction in 90 days" cannot be mapped onto each other, forcing them into one figure creates false precision that will be quoted for years. Report each definition's number separately, or state that the question cannot be answered to that precision. `RESEARCHER DECISION REQUIRED` (K5 §2.7).
- **The only sources are expert opinion and stakeholder belief.** A synthesis of what informed people think is a useful input to a design, and it is not an evidence base. Present it as expert consultation with its participants and its limits stated, and do not let the integration format lend it the standing of measurement. See **04.05 Expert and Stakeholder Interviewing** for how to gather it properly.
- **The client-supplied data cannot be interrogated.** If nobody can say what the field means, how the query was built, what was filtered out or whether the definition changed last year, the figure is not evidence and must not carry a claim. Record it as client-supplied and unverifiable per K2 §6, use it as context, and say plainly that it has not been checked (K4 §6.2).
- **Causal attribution is what the decision needs.** Convergence across four kinds of evidence is not a causal design, and a strong triangulated association remains an association (K4 §3.2). Where attribution is the question, an evaluation design capable of supporting it is required; see **01.04 Research Method Selection**.
- **The integration exists to make a predetermined conclusion look well evidenced.** Where a stakeholder wants "all the evidence" cited behind a position already taken, the honest output is an assessment that reports where the sources do not support it. Treat the stated preference as information about the stakeholder, not about the world (K4 §4.2).

## 5. Required inputs

**Required.**
- **The question or decision the integrated assessment must serve,** with enough specificity to test each source's relevance. Without it there is no basis for the authority mapping in Step 3, and every source appears equally applicable.
- **Access to each source itself, not a summary of it.** A number in a slide from another team is not a source; it is a claim about a source. Where the underlying source cannot be reached, it enters the inventory marked as such and does not carry a claim.
- **For every source: what it is, who produced it, when, and by what method.** A source whose provenance cannot be established cannot be given an epistemic status, and a source without a status cannot be integrated.

**Optional, and what each one adds.**
- **Data dictionaries, field definitions and query logic for supplied data.** Turns an internal figure from an assertion into something checkable, and is usually where the definitional conflict is found and dissolved.
- **The questionnaire or discussion guide behind any primary source.** Lets you compare constructs at the item level rather than at the label level, which is where alignment actually succeeds or fails.
- **A history of definition changes in the operational system.** A metric that changed definition eighteen months ago produces a trend that is an artefact, and nothing in the data reveals it.
- **The commissioning context of each external source.** Determines how much independent corroboration a convergence actually represents.
- **Direct access to the people who produced the internal data.** Twenty minutes with the analyst who built the extract resolves more definitional conflict than a week of inference.
- **The organisation's own prior conclusions from these sources.** Reveals which numbers are already load-bearing internally, so contradicting one is a deliberate act rather than an accident.

## 6. Questions to ask before starting

1. **What decision does the integrated assessment serve, and to what precision?** A directional answer tolerates imperfect alignment; a number that will be planned against does not. Default: assume directional and flag every place a precise figure is being drawn from imperfectly aligned sources.
2. **For each question, which source is closest to the thing being asked about?** This is the authority mapping, and asking it early prevents the default assumption that all sources speak to everything. Default: derive the mapping yourself in Step 3 and get it confirmed rather than assumed.
3. **Who produced the internal data, and can they be spoken to?** Determines whether definitional conflict can be diagnosed or only described. Default: treat undocumented internal figures as context, not evidence, and say so.
4. **Have any definitions or systems changed over the period the sources cover?** A definition change is invisible in the data and fatal to a trend. Default: ask explicitly, and where the answer is unknown, restrict comparisons to within-period.
5. **Are any of these sources already being used to report performance internally?** A figure that appears in a board pack has institutional weight, and contradicting it needs to be done deliberately and with the diagnosis attached. Default: identify them at the outset.
6. **Where would the client prefer the answer to land?** Not to accommodate it, but to know where the integration will be under pressure and to make sure those junctions are the best documented in the output. Default: assume there is a preferred answer and document the contested junctions hardest.
7. **What is not covered by any source?** Asked before synthesis, this shapes the inventory. Asked after, it becomes an apology. Default: draft the expected gap list at inventory stage and revise it.

## 7. Step-by-step methodology

**Step 1. Build the source inventory before any synthesis begins.**
Every source gets a row and a short source code, and every later reference in the project uses that code (K2 §6). Record: source code, name, type, who produced it, who commissioned or paid for it, the date of the data as distinct from the date of the document, the population or universe covered, the method, the base or volume, the definitions used for the key constructs, known limitations, and whether you have the source itself or a report of it. Building this first is not administration. It is the step that determines what the synthesis can honestly say, and it routinely reveals before any analysis that two sources everyone assumed were comparable cover different populations, or that the "current" figure is fourteen months old. Do not begin writing anything until the inventory is complete, because a source added later tends to be added in the register of the argument it supports. *Correct result: a complete inventory table with no unknown cells left blank; unknowns are written as unknown, which is itself an inventory finding.*

**Step 2. Assign an epistemic status to every source, and state what it can and cannot establish.**
Source types differ in what they are capable of showing, and the difference is not a matter of quality. Work through them explicitly. **Primary research you produced**: method known, chain intact, limits known precisely. **Primary research produced by others**: method reported rather than observed, comparability to be established. **Client-supplied internal figures**: derived for operational purposes rather than designed to answer a research question, definitions internal and often undocumented, provenance not yours to vouch for. **Operational and behavioural records**: strong evidence of what happened and weak evidence of why, complete within their coverage and blind outside it, subject to instrumentation and definition drift. **Expert and practitioner input**: informed judgement, valuable for mechanism and for interpreting anomalies, not measurement, and anchored to the individual's own experience in ways they cannot fully report. **Previous studies**: evidence about a period that has passed, comparable only if method and definition are shown to be comparable. Write the status against each source in the inventory, in the form of one sentence on what it establishes and one on what it cannot. **A client-supplied figure is never presented with the same authority as one you produced** (K2 §6): it is labelled client-supplied wherever it appears, including in the summary. *Correct result: each source carries an explicit capability statement, and no source is left to be read as general-purpose evidence.*

**Step 3. Establish which source is authoritative for which question.**
Do not treat sources as equal across the board. For each question in the assessment, name the source that is closest to the thing being asked about, and say why. Behavioural records are usually authoritative for what happened and to whom, within their coverage. Primary research is usually authoritative for why, for attitude, for the population the records do not see, and for anything the organisation does not instrument. Internal financial and operational data is usually authoritative for volume, cost and internal process. Expert input is usually authoritative for nothing on its own, and is frequently the best available source for how a market behaves in ways nobody measures. Published desk sources are authoritative for external context the organisation cannot observe. The map is per question, not per source: a source can be authoritative for one row and inadmissible for the next. Write it down as a table, because the discipline is in the writing; held in the head, it collapses back into treating every source as evidence for everything. *Correct result: a question-by-source authority map, with a named authoritative source per question and the reason, and secondary sources marked as corroborating or contextual rather than as equal.*

**Step 4. Align definitions, and record every place alignment fails.**
Definitions across sources almost never match, and the mismatch is usually invisible because the labels are the same. For each construct that appears in more than one source, write out the actual definition each source uses, at the level of the operational rule: what counts as a customer, what counts as active, what counts as a complaint, what counts as an attendance, what the denominator is. Then classify the pair as **aligned** (same rule, directly comparable), **mappable** (different rules, with a defensible transformation you can state), or **unalignable** (different constructs, comparable in direction at best). Where a mapping is applied, state the transformation and its assumption in the output, not in a working file. Where alignment fails, the sources are reported separately with their own definitions attached, and no combined figure is produced. Forcing unalignable definitions together is the most common route to false precision in mixed-evidence work, and the resulting number is durable: it will be quoted long after the caveat has been lost. *Correct result: a definitions table with a verdict per construct per source pair, and every mapping stated with its assumption.*

**Step 5. Align time periods and reference frames.**
Sources rarely cover the same window and almost never the same reference frame. Record for each: the period the data covers, the period a respondent was asked about, the point at which a record was written, and any seasonality in the underlying behaviour. A survey asking about the last three months, an operational extract for the last financial year, and a published report using data from two years earlier describe three different worlds, and setting their figures side by side implies a comparability that does not exist. Where periods differ materially, either restrict the comparison to an overlapping window, or state the offset explicitly beside every comparison, or drop the comparison. Watch particularly for recall windows against record windows: self-reported behaviour over 30 days and a system count over 30 days are not the same measurement even when the window matches, because recall error is directional and larger for frequent low-salience events. *Correct result: a period alignment note per comparison, and no side-by-side figure whose periods differ without the offset stated in the same place.*

**Step 6. Build the integration frame and populate it.**
Questions as rows, sources as columns, with the authority map from Step 3 marking which cell carries the answer. Each populated cell holds the source's finding, its definition, its period, its base or volume, and its epistemic status. Cells where a source does not speak to a question are marked as such, and this is informative rather than a gap to be filled. The frame is what keeps the synthesis honest under compression: when a paragraph has to be written, the frame shows exactly which source is doing the work and which are decoration. Keep it and ship it as an appendix. *Correct result: a populated integration frame in which every question has a named authoritative source or an explicit statement that none exists.*

**Step 7. Triangulate, and be precise about what convergence buys.**
Convergence between sources raises confidence only where the sources are genuinely independent and their errors are genuinely different. A survey and a set of operational records agreeing on a usage pattern is strong, because self-report error and instrumentation error are unrelated. Two internal reports drawing on the same underlying system agreeing is not convergence at all. An expert's view agreeing with your survey is weak corroboration where the expert has seen the survey, and worth very little where they were involved in the project. Practically: for each converging set, name the error each source is subject to, and ask whether an error in one would produce the same wrong answer in the other. Where the answer is no, the convergence is worth an upgrade in confidence; where it is yes, it is not. Divergence, equally, is not automatically a problem: sources measuring different aspects of the same phenomenon are expected to differ, and the interesting output is often the size and direction of the gap. *Correct result: every convergence claim accompanied by an independence statement, and every divergence characterised before it is treated as a conflict.*

**Step 8. Handle conflict between internal data and primary research with a named procedure.**
This is the delicate case, it is common, and it goes wrong in both directions: deferring to the client's number because it is theirs and the conversation is uncomfortable, or dismissing it because you did not produce it. Neither is defensible. Work through five checks in order and stop at the first that explains the gap. **Definition**: do the two measure the same construct under the same rule (Step 4)? Most conflicts dissolve here. **Population**: do they cover the same universe? Records see registered, transacting or served people; research often covers a wider population including lapsed and never-served. **Period and recall**: same window, same reference frame, and is recall error expected to be directional here? **Mode and measurement**: is one self-report and the other a record, and is there a known systematic gap between them for this behaviour, such as under-reporting of undesirable actions or over-reporting of routine ones? **Provenance of the internal figure**: who built the extract, what filters were applied, has the field definition changed, is the figure a system output or somebody's transformation of one? Where the conflict survives all five, do not adjudicate to a single number. State that both are correct within their own scope, name what each is authoritative for, quantify the gap, and put the choice in front of the researcher and the client with what turns on it. `RESEARCHER DECISION REQUIRED` where the decision changes depending on which figure is used (K5 §2.1, §2.5). *Correct result: every internal-versus-primary conflict diagnosed against the five checks in order, with the resolving check named, or explicitly carried forward unresolved with both scopes stated.*

**Step 9. Identify what no source covers.**
Read the integration frame for empty rows. A question that no source is authoritative for is a gap, and it must appear in the output as prominently as the answered questions, because a mixed-source assessment reads as comprehensive and its silences will be taken as coverage. Distinguish three kinds: questions no source addresses at all; questions addressed only by a source that is not authoritative for them, which is a weaker position than it looks; and questions where sources exist but their definitions could not be aligned, which is a gap in comparability rather than in evidence. Each has a different remedy and the difference matters to whoever plans the next piece of work. *Correct result: an explicit gap list, classified by kind, feeding forward to gap prioritisation.*

**Step 10. Write with provenance retained on every claim.**
Every claim in the output names its source by inventory code, its period, its definition where it differs from the default, and its epistemic status. Client-supplied figures are labelled at every appearance, including in the executive summary, which is where labels are most often dropped and where the damage is greatest. Where a claim rests on more than one source, all are named and the convergence or divergence is stated per K2 §4.4. Confidence per K3 is capped by the least authoritative source materially contributing, not by the average. And the finding-to-interpretation boundary stays visible: what the sources say is a finding, what you conclude from reading them together is an interpretation, and an integrated assessment is mostly the second. *Correct result: no claim in the output whose source, period and status a reader cannot identify from the claim itself.* `RESEARCHER SIGN-OFF REQUIRED` on any integrated assessment supporting a material decision (K5 §2.5).

## 8. Analytical framework

Two structures, applied together.

**The integration chain, per question:**

    Question → Source inventory → Epistemic status → Authority mapping →
    Definition and period alignment → Populated frame → Triangulation test →
    Conflict diagnosis → Integrated answer → Provenance retained → Residual gap

The chain runs per question, not per source, which is what stops the output becoming a tour of the evidence base. Two arrows carry most of the risk. *Authority mapping* is where the assumption that every source speaks to every question is either broken or preserved, and preserving it is how weak sources end up carrying strong claims. *Definition and period alignment* is where false precision is either caught or manufactured, and once manufactured it is very difficult to remove, because the resulting number is more quotable than its caveat.

**The epistemic status ladder.** Not a ranking of quality but of what a source is capable of establishing, and it is domain-specific:

    What happened, within instrumented coverage → operational and behavioural records
    Why, and among whom the records do not see → primary research
    External context the organisation cannot observe → appraised published sources
    Internal volume, cost and process → client-supplied operational data
    Mechanism, anomaly interpretation, and the unmeasured → expert and practitioner input
    A period that has passed → previous studies, comparable only where shown to be

Applying it: a source used inside its domain can carry a claim at full confidence for that claim; the same source used outside its domain carries nothing, however good it is. A behavioural record cannot establish motivation. A survey cannot establish what the system did. An expert cannot establish prevalence. The commonest failure in mixed-source work is not using a weak source, it is using a strong source outside the domain in which it is strong.

## 9. Output format

**1. Question set and scope.** The decision served, the questions the assessment answers, and the boundary.

**2. Source inventory.** The full table, near the front and not in an appendix, because the reader needs to know what the assessment is made of before reading its conclusions.

| Code | Source | Type | Producer and commissioner | Date of data | Population covered | Method | Base or volume | Key definitions | Epistemic status | Known limitations |
|---|---|---|---|---|---|---|---|---|---|---|

**3. Authority map.** Question by source, with the authoritative source named per question and the reason.

**4. Definition and period alignment note.** Each construct, each source's operational rule, the verdict (aligned, mappable, unalignable), and any transformation applied with its assumption.

**5. Integrated findings, question by question.** Each with its answer, the source carrying it, corroborating and diverging sources named, the triangulation statement including the independence assessment, and a confidence level in K3 language.

**6. Conflicts and their diagnosis.** Every conflict between sources, the check that resolved it or the statement that it stands, and what each source is authoritative for.

**7. What no source covers.** The gap list, classified by kind.

**8. Provenance and limitations.** What was verified and what was taken as supplied, per source.

**When the evidence is thin, the format must not force fabrication (K4 §1).** A question with no authoritative source is written as such, with the sources that touch on it named and their inadequacy stated. The integration frame is never padded to look complete: "does not address" is the correct and informative cell value, and an assessment where most cells are empty is telling the reader something true about the evidence base. A client-supplied figure is never promoted to fill a gap left by a source you could not obtain, and a gap is never closed with expert input presented as measurement.

## 10. Quality checks

Run before anything is presented. These sit on top of K4 §8.

1. Was the inventory completed before any synthesis was written, and does every source carry a code used consistently throughout?
2. Does every source carry an epistemic status with an explicit statement of what it cannot establish?
3. Is every source used only inside the domain its status supports?
4. Is every client-supplied figure labelled as client-supplied at every appearance, including in the summary?
5. Has the definition of every shared construct been written out per source at the level of the operational rule, not the label?
6. Is every mapping between definitions stated with its assumption, and is every unalignable pair reported separately rather than combined?
7. Does every side-by-side comparison state the periods, and is any recall window being compared to a record window without comment?
8. Does every convergence claim carry an independence statement naming the error each source is subject to?
9. Has every internal-versus-primary conflict been run through the five checks in order, with the resolving check named?
10. Has any conflict been resolved by preferring the source belonging to whoever is in the room?
11. Does the confidence on each integrated answer reflect the least authoritative materially contributing source rather than the average?
12. Is the gap list present, classified, and as prominent as the findings?
13. Does every claim name its source, period and status without the reader having to consult the appendix?
14. Is the boundary between what the sources say and what you conclude from reading them together visible on every page?

## 11. Common failure modes

| Failure | How to recognise it | How to prevent it |
|---|---|---|
| **The single voice** | A narrative in which every sentence carries the same authority | Provenance on every claim, status labels retained (Step 10) |
| **Client figure adopted silently** | An internal number in the summary with no label | Label at every appearance, per K2 §6 |
| **Label alignment** | Two sources combined because both say "active users" | Write out the operational rule per source (Step 4) |
| **Period collision** | Figures side by side from different windows | Period alignment note on every comparison (Step 5) |
| **Recall against record** | Self-reported frequency compared to a system count as though equivalent | Treat mode difference as a known directional gap (Step 8) |
| **False triangulation** | Two sources drawing on the same system reported as converging | Independence statement per convergence (Step 7) |
| **Expert input as measurement** | A practitioner view carrying a prevalence claim | Status ladder: expert input is authoritative for mechanism, not prevalence |
| **Source used outside its domain** | Behavioural data explaining motivation | Authority map per question (Step 3) |
| **Conflict resolved by seniority** | The client's number wins with no diagnosis recorded | Run the five checks in order and record the resolving one (Step 8) |
| **Conflict resolved by ownership** | Your number wins because you produced it | Same procedure, applied symmetrically |
| **Stale study treated as current** | A three-year-old figure in a current trend line | Date of data recorded in the inventory and carried into every claim |
| **Definition drift undetected** | A trend that inflects exactly at a system change | Ask for the definition-change history (Step 5) |
| **The comprehensive-looking gap** | An assessment covering six sources and silent on the key question | Gap list as a required output section (Step 9) |
| **Averaging across kinds** | A blended figure from a survey and a record system | Prohibited; report both with their scopes (§12) |

## 12. AI guardrails

Skill-specific. The universal prohibitions in K4 apply in full and are not repeated.

1. **Never combine figures across source types into a single number.** No average of a survey estimate and a system count, no blended rate, no reconciled figure. Where two sources of different kinds bear on one quantity, both appear with their definitions, periods and scopes. A reconciled number that appears in no source is a fabrication however reasonable the arithmetic (K4 §2.1).

2. **Never present a client-supplied or externally supplied figure as though you had produced it.** The label travels with the figure into every downstream document, including the executive summary and any single-page summary. Where you were given a number and could not see how it was derived, say so at the point of use, not in a methodology note.

3. **Never infer the definition behind a supplied figure.** Do not decide what "active" means from context, from the magnitude of the number, or from what would make the sources agree. An undocumented definition is recorded as undocumented, and a comparison that depends on it is not made (K4 §6.2).

4. **Never let a source carry a claim outside the domain its type supports.** Behavioural records do not establish motivation. Surveys do not establish what a system did. Expert input does not establish prevalence, however experienced the expert and however confident the phrasing. State the domain restriction where it bites, rather than relying on the reader to notice.

5. **Never describe sources as triangulating without stating what makes them independent.** The word implies mutually independent error, and using it for two sources that share a data feed, a definition, an analyst or a sponsor misrepresents the strength of the evidence.

6. **Never resolve an internal-versus-primary conflict by choosing.** Run the diagnosis, and where it does not resolve, report both with their scopes and put the decision to the researcher. The pressure to produce one number in this situation is exactly the pressure the guardrail exists to resist.

7. **Never fill a gap in the frame with the nearest available source.** An empty cell is a finding. Substituting a source that is not authoritative for the question, without saying so, converts a known gap into an unknown error.

8. **Never carry a previous study's figure forward as current.** Restate it with its date of data and its method attached every time it appears, and where it is being compared to current evidence, state the comparability assumption explicitly (K2 §6).

9. **Never let expert or stakeholder input be summarised into the same register as measured findings.** Attribute it, describe how it was gathered, and word it as judgement. The integration format is precisely what makes this easy to get wrong, because a frame with a cell for every source implies every cell is the same kind of thing.

10. **Where a source could not be obtained or interrogated, say so and let the gap stand** (K4 §6.4). Do not describe the contents of an extract you were sent a summary of, and do not characterise a system's data on the basis of what such systems usually hold.

## 13. Best-practice principles

- **The inventory comes first, always.** Sources added after the argument has started are added in the register of the argument. The inventory is also the single most persuasive part of the deliverable, because it shows the reader what the assessment is made of before it tells them what it concludes.
- **Ask what each source is capable of showing before asking what it shows.** The commonest failure in mixed-source work is not a weak source used weakly, it is a strong source used outside the domain in which it is strong.
- **Definitions are where the real conflict lives.** Most disagreements between internal data and research are definitional, and diagnosing that is the highest-value hour in the project, because it converts an argument into a clarification.
- **Twenty minutes with the person who built the extract is worth a week of inference.** Internal data has an author, and the author knows about the filter that nobody documented.
- **A client-supplied figure is evidence about what the client's systems record, not about the world.** That is often exactly what you need, and it is not the same claim.
- **Convergence only counts when the errors are unrelated.** Independence is a property of how sources were produced, not of who published them.
- **Divergence between a record and a self-report is usually information, not error.** The size and direction of the gap between what people do and what they say they do is frequently the most useful finding in the assessment, and reconciling it away destroys it.
- **Never let the number that is easiest to defend become the number in the summary.** Precision is not accuracy, and the operational figure is more precise than the research estimate about a different, narrower population.
- **Expert input earns its place by explaining anomalies,** not by voting. Use it where the data behaves oddly and someone knows why, and attribute it as judgement.
- **Say what nobody measured.** A mixed-source assessment reads as comprehensive by its structure, so its silences carry more weight than in any other deliverable.
- **Write the conflict section before the findings section.** The conflicts are where the thinking is, and drafting them first prevents a findings narrative that has already smoothed them away.
- **Keep the frame and ship it.** It is the difference between an assessment a colleague can audit and one they have to trust.

## 14. Worked example

Generic fictional scenario, international NGO.

**INPUT**

An NGO running a rural water programme needs to assess whether its handpump maintenance scheme is working, ahead of a funding renewal. Available evidence: a household survey it commissioned this year (S1); the programme's own maintenance service logs (S2); a set of technician inspection records held in a separate system (S3); structured input from six field staff with long experience in the districts (S4); an independent evaluation of the same programme carried out three years ago (S5); and a published sector study of handpump functionality across the region (S6).

**PROCESS**

*Step 1 and 2.* The inventory is built and immediately produces two findings before any analysis. S3's inspection records cover only pumps registered to the maintenance scheme, roughly two thirds of the pumps in the districts. S6's fieldwork is four years old although it was published last year. Epistemic statuses are written: S2 and S3 establish what the programme did and what technicians recorded, within registered coverage, and cannot establish what households experienced; S1 establishes household experience across the whole district including unregistered pumps, and cannot establish technical cause; S4 is authoritative for mechanism and for interpreting anomalies, and for nothing quantitative.

*Step 3.* Authority mapping. "Are pumps functional?" is assigned to S3 within registered coverage. "Do households have water?" is assigned to S1, because it is the only source covering households served by unregistered pumps. "How quickly are faults repaired?" is assigned to S2. "Why do repairs stall?" is assigned to S4, with S1's open responses as corroboration.

*Step 4.* Definition alignment. S3 defines a pump as functional if it delivers water at the moment of inspection. S1 asks households whether their main water point supplied water on every day of the last 30. S6 uses functional within the last seven days. These are three constructs, not three measurements of one. Verdict: unalignable on level, comparable in direction only.

*The judgement call.* S3 reports 91% functionality; S1 reports 68% of households with uninterrupted supply over 30 days. The programme director reads this as the survey being wrong. The temptation is either to accept that, because the inspection records are objective, or to reject it, because the survey is the commissioned research. Neither is done. The five checks in Step 8 are run in order. Definition explains most of the gap: a point-in-time inspection and a 30-day continuity measure will diverge by construction, and a pump that fails for four days a month passes one and fails the other. Population explains more of it: S3 covers registered pumps only, and unregistered pumps serve a third of households. Period explains a further part: inspections cluster in the dry season when the scheme is most active. After the three checks, the residual gap is small and no longer surprising.

*Step 7.* Triangulation is examined rather than asserted. S2 and S3 both draw on technician-entered records and share the same reporting incentive, so their agreement is not independent corroboration. S1 and S4 agree that repairs stall at the point of parts availability, and they are independent: a household survey and staff judgement are subject to entirely different errors, so this convergence does raise confidence.

*Step 9.* Gaps. No source covers households served by pumps outside the scheme with respect to repair times, because S2 and S3 cannot see them and S1 did not ask. That is precisely the population the funding case is about.

**OUTPUT**

An assessment stating that registered pumps are functional at inspection in the large majority of cases (S3, registered coverage only, point-in-time definition stated), that household experience of continuous supply is substantially lower and this is largely explained by definition, coverage and seasonality rather than by contradiction, that repair delay concentrates on parts availability with independent support from two sources, and that repair performance for unregistered pumps is unmeasured by any source. The headline figure the funding case had been using, 91%, is retained with its definition and coverage attached rather than removed, and is no longer presented as a household outcome.

`RESEARCHER DECISION REQUIRED` on which functionality definition the renewal case reports against, since the two produce materially different pictures and the choice is a reporting judgement, not an analytical one (K5 §2.1). `RESEARCHER SIGN-OFF REQUIRED` on the assessment as a whole (K5 §2.5).

## 15. Advanced usage

**Standing evidence inventories.** In organisations that commission repeatedly, maintain the source inventory as a permanent asset rather than rebuilding it per project. The value compounds: definitions get documented once, definition changes get logged as they happen instead of being reconstructed later, and the recurring question of what is already known becomes answerable in an afternoon. See **14.03 Research Repository and Knowledge Curation**.

**Deliberate design for triangulation.** Where primary research is still being designed, use the inventory to specify what the primary study should measure so that its errors are independent of the internal data's errors. A survey that replicates what the system already counts adds precision to nothing; a survey that measures the population the system cannot see adds an entire dimension. This turns the inventory into a design input and hands over to **01.04 Research Method Selection**.

**Quantifying the record-to-report gap.** Where both a behavioural record and a self-report exist for the same population and period, the gap between them is itself measurable and reusable. Establishing that a given behaviour is over-reported by a stable factor lets future self-report-only studies be interpreted with a known correction, stated as an assumption. This is valuable and it is easy to overreach with: the correction is specific to the behaviour, the population and the wording, and it must be re-established rather than carried across contexts.

**Where the standard approach does not fit.** Where internal data is undocumented and its authors have left, treat it as an external source of unknown provenance rather than as internal evidence, and appraise it as you would any unverifiable document. Where the evidence base is entirely retrospective and no primary work is possible, the assessment's main output is usually the alignment analysis itself, because establishing that three internal figures were never comparable is a more useful finding than any of the three.

## 16. Skill chain

**Recommended previous skills:**
- **10.01 Literature Review and Desk Research.** Verifies and appraises the published and secondary sources before they enter the inventory.
- **10.02 Evidence Synthesis.** Where a subset of the sources are comparable studies, synthesises them into a single claim before that claim enters this frame as one source.
- **01.01 Research Brief Interrogation.** Establishes the decision, the stakeholders, and which internal figures are already load-bearing.

**Recommended next skills:**
- **08.01 Finding to Insight Development.** Takes the integrated findings and works out what they mean for the organisation.
- **10.04 Research Gap Identification.** Takes the gap list and turns it into a prioritised research specification.
- **12.04 Research Evidence Integration** and **12.03 Research Report Compilation,** which carry the inventory, the provenance labels and the confidence levels into the deliverable.

**Runs well alongside:**
- **13.02 Source and Citation Verification,** for the external sources in the inventory.
- **13.04 Bias Detection,** which audits the integration for accommodation of stakeholder preference.
- **10.05 Competitive and Market Research,** where external market sources form part of the mix.

---
A Yazi Supplied Skill and resource.
