---
name: research-evidence-integration
description: >
  Gets heterogeneous evidence into a single research document with provenance
  intact: quantitative tables, qualitative themes, verbatim, images, video and
  audio, client-supplied figures, previous waves and secondary sources. Use for
  "how do I combine survey and interview evidence", "the numbers do not match
  between two files", "how do I label client data in the report", "can we
  compare this to the last wave", "which evidence supports this claim", "build
  the evidence map", "this quote has no participant ID".
category: 12 Report Design and Compilation
ref: 12.04
tier: 1
inherits: [K2, K3, K4, K5]
---

# Research Evidence Integration

## 1. One-line description
Handles the mechanics of getting many different kinds of evidence into one research document with provenance intact: inventorying and coding every input, matching evidence to claim rather than to convenience, reconciling figures that disagree, labelling what is not yours to vouch for, and leaving an audit trail a reviewer can walk.

## 2. What this skill is used for

**The research problem it solves.** Evidence degrades in transit. A number leaves an analysis workbook with a base, a question reference and a filter, and arrives in a report as a percentage in a chart title. A quote leaves a transcript with a participant identifier and a segment, and arrives in a chapter as a pull quote with a first name. A client sends a sales figure in a spreadsheet, and three weeks later it is a headline indistinguishable from data you analysed. A figure from a study two years ago becomes a trend line without anyone checking whether the question was worded the same way. And when two files give different numbers for the same quantity, the one that fits the story tends to win, quietly, with nobody deciding to do it. None of these are acts of dishonesty. Each is a small convenience under time pressure, and together they are the mechanism by which an honest project produces a document nobody can check. The problem is worse in mixed-method and multi-source work, which is most substantial research, because there is no single source of truth to fall back on and the streams measure different things on different people at different times. This skill is the mechanics of holding evidence together across those joins.

**Where it sits.** Reporting, running underneath compilation rather than before or after it. Analysis has produced outputs; the report does not exist yet; the evidence has to be got from one to the other without losing what makes it evidence.

**Typical use cases.**
- A mixed-method study where survey findings and qualitative themes must appear in one argument, converging in some places and diverging in others.
- A report that will draw on client-supplied operational data alongside primary research.
- A tracking wave where comparability with earlier waves has to be established rather than assumed.
- A project where two analysts' workbooks give different numbers for the same measure.
- A study using images, video or audio as evidence, which needs different handling from a table.
- A report inheriting claims from an earlier document whose sources have to be re-established.
- Any deliverable a client will interrogate claim by claim.

**Who uses it.** Research directors and senior researchers responsible for what the report asserts; project leads assembling several people's outputs; consultants integrating client data with primary research; anyone who will have to answer "where did this come from" in front of the people who commissioned it.

## 3. When to use it

- The report will draw on more than one evidence type, method or source.
- Client-supplied figures, secondary sources or previous waves will appear alongside your own analysis.
- Two or more files give different numbers for what should be the same quantity.
- Quotes, images or audio will be used as evidence and their identifiers and consent status must survive into the document.
- Comparability with a previous study is being claimed, or is about to be assumed.
- A claim in a working document has arrived without a source and somebody has to decide whether it can be used.
- The evidence points in different directions and the divergence has to be handled rather than averaged.
- An auditable trail is required, whether by a client, a regulator, a funder or your own quality standard.

## 4. When NOT to use it

- **The task is the compilation sequence and the narrative rather than the evidence mechanics.** Deciding what the report argues, prioritising the finding register, setting chapter order, building the conventions register and drafting are **12.03 Research Report Compilation**. The boundary runs in both directions and it is worth stating precisely. **12.03 owns the sequence and the story: what gets said, in what order, and how the report is assembled and written. This skill owns the evidence handling inside that sequence: inventorying and coding inputs, verifying them, matching claim to evidence, labelling provenance, reconciling conflicting figures to their origin, and maintaining the evidence map.** Compilation records a divergence between two streams and reports it. This skill establishes whether it is a genuine divergence, a base difference or a measurement artefact, before compilation has anything to record. Where compilation reaches a claim with no source, it comes back here; where this skill produces a reconciled evidence set, it goes forward to there.
- **The analysis has not been done.** Integration is not a substitute for analysis. If a theme has not been coded, a cross-tab has not been run, or a dataset has not been analysed at all, integrating what exists produces a report built from fragments. Finish the analysis (Category 06 and 07) first.
- **The question is how to combine two datasets analytically rather than how to report them together.** Merging files, aligning variables, joint modelling and integrated analysis of quantitative and qualitative data are analytical operations and belong to **07.06 Qualitative and Quantitative Integration**. This skill assumes the analysis outputs exist and handles their provenance in the document.
- **The question is whether a source is credible.** Assessing the authority, method and age of a secondary source, and verifying that a citation exists and says what it is claimed to say, is **13.02 Source and Citation Verification**. This skill records verification status; it does not perform source appraisal from scratch.
- **The evidence is fundamentally unreliable.** Where the dataset itself is in doubt, integrating it produces a well-documented trail to an unreliable conclusion. Run **13.01 Research Quality Review** first.
- **Nothing can be traced.** Where analysis tables arrive with no question references, quotes with no participant IDs and figures with no origin, integration would launder unverifiable material into an authoritative document. Send the inputs back for referencing, or integrate only the traceable subset and state in the report what was excluded and why. This is a stopping condition, not an inconvenience.
- **The task is to make two irreconcilable sources agree.** They may not agree. Where investigation cannot establish the origin of a difference, the correct output is both figures with their sources and the range, not a chosen number. A request to produce one figure regardless is a request to fabricate a reconciliation, per K4 §1.
- **A single-source report from one clean analysis.** One dataset, one method, one analyst: the coding, inventory and map overheads here do not earn their keep. Reference inline per K2 §4 and write it up.

## 5. Required inputs

**Required.** Without these the skill cannot run. If absent, ask. If no answer is available and work must proceed, state the assumption at the point where it bites, per K5 §5.

- **Every input the project will draw on**, including the ones you expect not to use, since an inventory that lists only what was used cannot show that the rest was looked at.
- **Source identifiers on each input**: question references and base descriptions on tables, participant IDs on quotes and transcripts, capture references on images and audio, file provenance and date on client material, full citation on secondary sources.
- **The method behind each input**: what it measured, on whom, when, and how. Two figures that look like the same measure are frequently not, and the method line is where that becomes visible.
- **The claims the report intends to make**, at least as a finding register, so evidence can be matched to claim rather than assembled and then interpreted.
- **The consent and permission status of any personal, visual or audio material**, which determines whether it can be used at all, independently of whether it is good evidence.

**Optional, and what each one adds.**

- **The raw dataset**: allows a disputed figure to be recomputed at source, which is the only way to settle a discrepancy between two derived files rather than adjudicating between them.
- **The questionnaire and discussion guide as fielded, with routing**: lets base descriptions be checked against actual routing and lets prompted material be distinguished from volunteered, which changes what a finding means.
- **The previous wave's questionnaire and technical report**: the only way to establish comparability rather than assume it. Without them, a trend claim rests on the assumption that nothing changed.
- **The analysts who produced the workbooks**: turn a three-hour discrepancy investigation into a five-minute one, and are frequently the only place a filter decision was ever recorded.
- **The client's data dictionary or documentation for supplied files**: converts an unverifiable figure into a labelled one with a known definition, which is the difference between usable and unusable client data.
- **A previous evidence map from an earlier wave or related project**: supplies the source-code taxonomy and prevents a new coding scheme being invented every project.

## 6. Questions to ask before starting

1. **What are all the inputs, and who vouches for each?** Determines the inventory and the verification status of every item. The question "who vouches for this" is more useful than "where did this come from", because it identifies the accountable party. *Default if unanswered:* treat anything you did not analyse as unverified until established otherwise, and label accordingly.
2. **Which inputs are client-supplied, and what do we know about how they were produced?** Determines labelling per K2 §6 and determines what they can be used to support. *Default:* label as client-supplied at every appearance and do not lead a chapter on one.
3. **Is comparability with previous waves or studies being claimed?** Determines whether wording, base definition, sample and method have to be checked before any trend statement exists. *Default:* assume comparability is unestablished until checked, because the reverse assumption produces false trends and they are hard to retract.
4. **What consent and permission covers the visual, audio and personal material?** Determines what may appear at all, and it is a use question rather than a quality question. *Default:* do not use identifiable material without confirmed consent for this use, and flag it per K5 §2.4.
5. **Where do we already know two sources disagree?** Determines the investigation list, and knowing before compilation begins is far cheaper than discovering it in QA. *Default:* run the cross-source figure comparison at step 6 rather than waiting for a discrepancy to surface.
6. **What will the client interrogate?** Determines where the audit trail must be strongest. Some claims will be checked line by line and most will not, and knowing which is which is a legitimate way to allocate effort. *Default:* assume every recommendation and every executive summary claim will be interrogated, per K2 §5.

## 7. Step-by-step methodology

**What integration is.** The work of getting evidence from where it was produced to where it is asserted without losing the properties that make it evidence: what it measures, on whom, when, by what method, and who vouches for it. Two artefacts do the work: the **evidence inventory** and the **evidence map** (K2 §5), and the compilation rules in **K2 §6** are the standard this skill implements.

**Step 1. Inventory and code every input.** List every input before reading any of it for content, and give each a short source code used in every subsequent reference. A workable taxonomy separates streams by prefix, for example `QNT` for quantitative analysis outputs, `QUAL` for coded qualitative outputs, `VBM` for transcript and verbatim sets, `VIS` for images and video, `AUD` for audio, `CLI` for client-supplied material, `SEC` for secondary sources, and `PRV` for previous waves and studies. The prefix matters because it makes the provenance of a claim visible at a glance in the map: a chapter resting entirely on `CLI` codes is visible as such without reading it. Record for each: what it is, who produced it, the date, the method behind it, what it measures and on whom, the base or sample, and whether it has been through your own analysis. Include items that will not be used, marked as such. Flag duplicates and near-duplicates immediately: two files carrying overlapping numbers in different states cause most integration errors, and they are cheapest to identify before either has been quoted. *Correct result:* a coded inventory in which no item is described as "the latest version" without a date, and every code resolves to one identifiable file.

**Step 2. Write the method line for every input.** For each item, write one sentence stating what it measured, on whom, when, and how. This looks clerical and is the highest-value step in the skill, because it is where the false equivalences become visible. "Satisfaction, 8-point scale, all customers who had contact in the last 6 months, n=812, online panel, March" and "Satisfaction, 5-point scale, all account holders, client CRM extract, rolling 12 months" are two things called satisfaction that cannot be compared, and nobody notices until the method lines sit next to each other. Where a method line cannot be written because the information is not available, that is itself the finding: the input's verification status is unverified and it cannot support a comparison. *Correct result:* a method line per input, and a shortlist of pairs that look comparable but are not.

**Step 3. Assign verification status.** Every input gets one of four statuses, and the status travels with every claim it supports. **Verified**: produced by your analysis from source data you hold, or independently checked against source. **Client-supplied**: provided by the commissioning organisation, definition possibly documented, provenance not yours to vouch for. **Third-party**: secondary or published, with its own method and age, appraised per **13.02**. **Unverified**: origin cannot be established, or the method line could not be written. The status determines what the input may do in the report. Verified evidence may carry a headline. Client-supplied evidence may support a finding, is labelled at every appearance per K2 §6, and does not lead a chapter, because a headline resting on an input you did not produce puts your name on somebody else's number. Third-party evidence carries its citation and date. **Unverified evidence supports nothing.** It is either verified before use, used with the claim explicitly marked as resting on an unverified input, or dropped, per K2 §6. There is no fourth option, and the common failure is the input that gets used because it was in the folder. *Correct result:* a status against every code, and a short list of inputs that must be verified or dropped before compilation.

**Step 4. Match evidence to claim.** For every claim the report intends to make, ask what evidence type would properly support it, and only then ask what is available. Doing it in the other order is the central failure of integration: the claim gets supported by the evidence at hand, which is how a behavioural claim ends up resting on a stated-intention question and a market claim ends up resting on a customer sample. The matching rules that do most of the work: a claim about **behaviour** needs observed or recorded behaviour, and stated behaviour is recall, which is weaker; a claim about **prevalence** needs a quantitative base, and a qualitative theme count is a count of participants, not an estimate of a population; a claim about **why** needs qualitative or experimental evidence, and a cross-tab shows association only, per K4 §3.2; a claim about **change** needs matched measurement at two points, not two studies that both exist; a claim about a **segment** needs an adequate base within that segment, not an adequate study; and a claim about the **market** needs a sample of the market, not a sample of one company's customers, per K4 §3.3. Where the right evidence does not exist, there are three honest outcomes: restate the claim at the level the available evidence supports, mark it as a hypothesis per K3 §4.3, or drop it. *Correct result:* a claim-evidence matrix in which every claim names the evidence type it requires and the codes that supply it, with mismatches marked.

**Step 5. Integrate each evidence type on its own terms.** Different evidence needs different handling, and flattening them into one format is where provenance is lost.

**Quantitative tables**: every figure carries question reference, base description, base size, weighting status and any filter, per K2 §4.1. A figure quoted without its filter is a different figure. Never transcribe by hand where the file can be read; where a figure is transcribed, it is recomputed at QA.

**Qualitative themes**: a theme carries its prevalence as a count of participants out of the total, never as a percentage, and it carries the counter-evidence within the same theme, per K4 §4.1. A theme name is a label and never a finding, per K2 §7.

**Verbatim**: every quote keeps its participant identifier end to end, and the identifier is checked against the segment it is attributed to. Permitted edits are those in K4 §2.3 and the convention is stated once in the report. A quote that arrives without an ID does not enter the report until the ID is recovered from source.

**Images and video**: each carries a capture reference, date, location or context, and consent status. The caption states what is shown, never what it proves: a photograph of an empty waiting room shows an empty waiting room at one moment, and the claim that the service is under-used is a separate assertion needing separate evidence. Faces and identifying detail need consent for this use specifically.

**Audio and voice notes**: transcribed before use, with the transcript treated as the evidence and the recording as the source. Tone and hesitation are observations, and where they carry weight, the researcher's observation is labelled as an observation rather than presented as something the participant said.

**Client-supplied data**: labelled at every appearance, with the extraction date and the definition where documented. Its strength is that it measures the business rather than the respondent, which is why it frequently outranks your survey on what customers actually did. Its weakness is that you cannot vouch for it. Both go in.

**Previous waves**: dated, with the method noted, and comparability established before any comparison, per step 6.

**Secondary sources**: full citation, publication and access dates, and the source's own method and base where the figure is quoted rather than the conclusion. A secondary figure quoted without its base is unusable, whatever its provenance.

*Correct result:* every input represented in the working set in a form carrying its own provenance requirements, with nothing normalised into a shape that drops them.

**Step 6. Reconcile figures that differ, by investigation.** Compare every figure that should be the same across sources, before compilation rather than during QA. Where two sources disagree, **investigate to origin**. The usual causes, in rough order of frequency: a different base (all respondents versus those who answered), a different filter, an excluded or included "don't know", a different rounding point, a different weighting state, a reworded question between waves, and a different time period. Each cause is itself reportable and several are more interesting than the figure. **Never resolve a discrepancy silently, and never adopt the more convenient figure.** The convenience is the tell: notice when one of the two numbers makes the story better, and treat that as a reason for more scrutiny rather than less. Where the raw data exists, recompute at source rather than adjudicating between derived files. Where the origin genuinely cannot be established, report neither as fact: state the range, name both sources, and flag it per K5 §3.1. For previous waves, run the comparability check before any trend claim: question wording identical, base definition identical, sample definition and method equivalent. Where any of the three differs, the comparison is not made, and the reason is reported, per K4 §2.5's prohibition on describing a study as having done what it did not. *Correct result:* a discrepancy log with an investigated cause against every entry, and a comparability verdict for every previous-wave comparison the report intends.

**Step 7. Handle convergence and divergence.** Where two independent streams address the same question, state the relationship explicitly, per K2 §4.4. **Converging** streams raise confidence and the multi-source reference names both. **Complementary** streams answer different parts of the question and are not a check on each other, which is the relationship most often mislabelled as convergence: a survey saying how many and interviews saying why do not corroborate each other. **Diverging** streams are reported as divergence and never averaged. Divergence is frequently the most useful thing in a study, and the discipline is to establish first whether it is real: check that both streams measure the same construct on comparable people at a comparable time (step 2's method lines), because most apparent divergence is a measurement difference. Where it is real, assess which source is stronger for this specific claim and show the reasoning, rather than deciding by stream type. *Correct result:* every multi-source claim labelled converging, complementary or diverging, with a note of what was checked before the label was applied.

**Step 8. Build the evidence map as you go.** The map (K2 §5) is the working artefact of integration, not a document produced afterwards for an auditor, and the difference is not administrative: built during integration it records the link between claim and source while that link is still in front of you, and built afterwards it records what someone remembers. Each row: reference, report location, claim, its level on the K2 chain, the source codes it rests on, base and question reference, verification status, and confidence per K3. Every recommendation and every executive summary claim has a row, per K2 §5. **A row with no source is a defect, not a gap to fill later.** Populate recommendation rows backwards from the recommendation to the findings beneath it; a recommendation that cannot be traced backwards is deleted at this point rather than defended at sign-off. *Correct result:* a map with no empty source cells, in which the provenance mix of any chapter is visible from the code prefixes alone.

**Step 9. Handle strong evidence that does not fit the narrative.** At some point in most projects, a well-evidenced finding sits outside or against the story. The pressure to drop it is real and it is rarely experienced as dishonesty: it feels like editing. Per K4 §4.1 it is reported. Four honest placements: **in the argument**, where it qualifies the main finding, which usually makes the report better and more credible; **as a separate finding**, outside the argument, in its own section; **in the limitations or "what we could not establish" section**, where it constrains how the main finding should be read; or **in the appendix**, only where it is genuinely secondary to the decision and never because it is inconvenient. The distinction between the last option used honestly and used dishonestly is whether you would be comfortable with the client finding it there. **The evidence that most wants dropping is usually the evidence the reader most needs**, because it is the thing that would have changed their mind. *Correct result:* a written disposition for every prioritised finding that does not fit, none of which is "omitted".

**Step 10. Close the audit trail.** Before handing to compilation, confirm the trail runs both ways: from any claim, the map names its codes, and every code resolves to a real file that a colleague could open; and from any input, you can say whether it was used, where, and if not, why not. Record the integration decisions that a reviewer would otherwise have to reconstruct: discrepancy causes, comparability verdicts, verification statuses, dropped inputs, and any claim restated because the right evidence did not exist. This record is what lets a reviewer trace any claim back to source in minutes rather than reconstructing your reasoning from the report. *Correct result:* an integration record accompanying the evidence map, and the test in Section 10 check 1 passing on a sample of claims chosen by someone else.

## 8. Analytical framework

The integration chain:

    Claim → evidence type required → evidence available → match or mismatch
        → provenance label → verification status → confidence → map row

**Applying it.** The chain runs from the claim, not from the evidence, and that direction is the whole discipline. Running it from the evidence produces claims shaped by what happened to be available, which is how a report ends up asserting behaviour from a stated-preference question. The mismatch step is where most of the value is realised: a claim whose required evidence type is not present is restated, marked as a hypothesis, or dropped, and that decision is recorded rather than absorbed.

**The provenance triple.** Every piece of evidence in a report carries three things, and losing any one of them makes it uncheckable:

    What it measures (and on whom) → Who produced it → How strongly it is held

K2 §4 governs the first, K2 §6 the second, K3 the third. A number with a base but no provenance, a client figure with provenance but no definition, and a confident claim with neither are three versions of the same defect.

**The four verification statuses and what each may do:**

| Status | May carry a headline | Labelled at every appearance | May support a finding |
|---|---|---|---|
| Verified | Yes | No label needed | Yes |
| Client-supplied | No | Yes | Yes |
| Third-party | No | Yes, with citation and date | Yes, with its own base |
| Unverified | No | Yes, explicitly | Only with the claim marked as resting on it, or not at all |

## 9. Output format

**A. Evidence inventory**

| Code | Input | Producer | Date | Method line (what, on whom, when, how) | Base or sample | Verification status | Used |
|---|---|---|---|---|---|---|---|

**B. Claim-evidence matrix**

| Claim | Evidence type required | Codes supplying it | Match, mismatch or gap | Resolution |
|---|---|---|---|---|

**C. Evidence map** (per K2 §5)

| Ref | Report location | Claim | K2 level | Source codes | Base and question ref | Verification status | Confidence |
|---|---|---|---|---|---|---|---|

**D. Discrepancy log**

| Figures in conflict | Sources | Investigated cause | Resolution | Reported as |
|---|---|---|---|---|

**E. Comparability record** (where previous waves or studies are used)

| Comparison intended | Wording match | Base definition match | Sample and method match | Verdict | If not comparable, what is reported instead |
|---|---|---|---|---|---|

**F. Integration decisions record**
Inputs dropped and why; claims restated because the required evidence type was absent; findings that did not fit the narrative and their disposition; consent and permission decisions on visual and audio material.

**When the evidence is thin.** The matrix shrinks; it does not get filled. A claim with no matching evidence is restated at a level the evidence supports, marked as a hypothesis with its validation named per K3 §4.3, or dropped, and the decision goes into the integration record. An unverified input that is the only support for a point either gets verified or the point goes. **Do not resolve thinness by substituting adjacent evidence**: supporting a behavioural claim with an attitudinal measure because the behavioural measure does not exist is the mismatch this skill exists to catch, and it is more damaging than an absent claim because it looks like evidence.

## 10. Quality checks

Run before evidence goes forward to compilation. Sits on top of K4 §8.

1. Pick five claims at random from the map: does each name source codes that resolve to real files, and can each be traced to source in under a minute?
2. Does every input appear in the inventory with a code, a method line and a verification status, including inputs not used?
3. Can a method line be written for every input, and where it cannot, is the input marked unverified and excluded from comparisons?
4. Does every claim name the evidence type it requires, and does the evidence supplied actually match that type?
5. Is any behavioural claim resting on stated intention or recall, any prevalence claim on a qualitative count, any causal claim on a cross-tab, or any market claim on a customer sample?
6. Does every quantitative figure carry question reference, base description, base size, weighting status and filter?
7. Does every quote carry a participant identifier, verified against source and against the segment claimed?
8. Does every theme carry a participant count out of the total rather than a percentage, and its counter-evidence?
9. Does every image, video and audio item carry a capture reference, a caption stating what it shows rather than what it proves, and a confirmed consent status for this use?
10. Is client-supplied material labelled at every appearance, not only the first, and does it lead no chapter or headline?
11. Does every previous-wave comparison have a recorded comparability verdict covering wording, base definition, sample and method?
12. Was every figure discrepancy investigated to its origin and logged, with no discrepancy resolved to the more convenient figure?
13. Is every multi-source claim labelled converging, complementary or diverging, with complementary streams not described as corroboration?
14. Does every recommendation and every executive summary claim have a map row populated backwards to findings?
15. Is there any row in the map with an empty source cell?
16. Does every strongly evidenced finding that does not fit the narrative have a written disposition other than omission?
17. Could a reviewer who was not on the project reconstruct why any given figure was chosen over its conflicting alternative?

## 11. Common failure modes

| Failure | How to recognise it | How to prevent it |
|---|---|---|
| **Evidence chosen for availability** | A behavioural claim resting on a stated-intention question; a market claim on a customer sample | Step 4. Ask what evidence the claim requires before asking what exists |
| **Provenance decay in transit** | A figure in the draft with no base; a quote with a first name and no ID | Steps 1 and 8. The code travels with the claim; never rewrite a sentence without carrying its reference |
| **Client input laundering** | A client-supplied figure carrying a headline, indistinguishable from analysed data | Step 3. Label at every appearance; never lead with it, per K2 §6 |
| **The convenient reconciliation** | Two figures existed, the report shows one, and it is the one that helps | Step 6. Investigate to origin. Notice when a number improves the story and scrutinise harder |
| **False comparability** | A trend line across waves where the question was reworded or the base redefined | Step 6 comparability check before any trend claim exists |
| **Complementary mislabelled as converging** | "Both the survey and the interviews confirm this", where one gives prevalence and the other gives reasons | Step 7. Corroboration requires two measures of the same construct |
| **The theme as a statistic** | "62% of participants said" from a qualitative study of 24 | Step 5. Counts out of the total, never percentages, per K4 §7 |
| **The caption that argues** | A photograph captioned with what it proves rather than what it shows | Step 5. Captions describe; claims need their own evidence |
| **The unverified input that got used** | An input in the folder with no known origin, now supporting a finding | Step 3. Unverified evidence supports nothing until verified or explicitly marked |
| **Silent quote repair** | Quotes that read fluently and match the report's register | Verify word for word against transcript. Permitted edits only, per K4 §2.3 |
| **The map built afterwards** | A tidy evidence map produced after the report, with plausible sources | Step 8. Build it during integration. Reconstructed traceability is memory dressed as record |
| **The inconvenient finding that quietly went to the appendix** | A well-evidenced finding against the story, findable but not visible | Step 9. Four dispositions, none of which is omission. Test: would you be comfortable with the client finding it there |
| **Secondary figure without its base** | A published percentage quoted with a citation and no sample information | Step 5. A figure needs its own base wherever it appears, whoever produced it |
| **AI-supplied plausible provenance** | A source reference, base or participant ID that looks right and was inferred rather than read | Guardrail 1 in Section 12. A reference is copied from source or it does not exist |

## 12. AI guardrails

Skill-specific only. Universal prohibitions are inherited from K4 and are not repeated here. Compilation traceability follows **K2 §6**; quote handling follows **K4 §2.3**.

1. **Never infer or reconstruct a source reference, base description, participant ID, capture reference or citation.** A reference is read from the source or it does not exist. An inferred reference is worse than a missing one, because it will be trusted.
2. **Never present an input as verified that you did not verify.** Verification status is a statement about what was actually checked, not an assessment of how reliable something seems.
3. **Never normalise evidence into a common format in a way that drops its provenance requirements.** A quote is not a row, a theme is not a percentage, and an image is not a data point.
4. **Never resolve two conflicting figures without an investigated cause**, and never choose the figure that suits the narrative. Where the origin cannot be established, report both with the range.
5. **Never claim comparability between waves or studies without having checked wording, base definition, sample and method.** Comparability is established, never assumed, and an unchecked trend claim describes a study as having measured change when it did not, per K4 §2.5.
6. **Never describe complementary evidence as corroborating.** Two streams answering different parts of a question do not confirm each other, and saying they do manufactures confidence that no evidence supports.
7. **Never let a client-supplied, third-party or unverified input carry a headline or lead a chapter**, and never let its label appear only at first mention.
8. **Never write an image, video or audio caption that asserts what the material proves.** Captions state what is shown. Everything beyond that needs its own evidence and its own confidence marker.
9. **Never use identifiable personal, visual or audio material without a confirmed consent status covering this use**, and flag the question per K5 §2.4 rather than deciding it.
10. **Never omit a well-evidenced finding because it does not fit the narrative.** Four dispositions exist at step 9 and none of them is omission, per K4 §4.1.
11. **Never fill a gap in the claim-evidence matrix with adjacent evidence.** Restate the claim, mark it as a hypothesis, or drop it.

## 13. Best-practice principles

1. **Evidence degrades in transit, and every degradation is a small convenience.** The discipline is not honesty, which most researchers have. It is refusing the twenty small conveniences that each look reasonable alone.
2. **Ask what evidence the claim requires before asking what evidence exists.** Reversing that order is the single most common integration failure and it is nearly invisible in a finished report.
3. **The method line is the highest-value clerical act in reporting.** What it measured, on whom, when, how. Two figures called the same thing are usually not the same thing, and the method line is where that becomes visible.
4. **Who vouches for this is a better question than where did this come from.** It identifies an accountable party rather than a filename.
5. **A discrepancy is information.** Two analysts disagreeing about a number is usually telling you something true about a base, a filter or an exclusion, and the cause is frequently more interesting than either figure.
6. **Notice when a number improves the story.** That is the moment to check harder, not less, and it is the only reliable defence against unconscious selection.
7. **Comparability is established, never assumed.** The cost of checking is twenty minutes. The cost of a retracted trend claim is the client's confidence in everything else in the report.
8. **Divergence is usually the most useful thing in a mixed-method study**, provided you have first ruled out that it is a measurement artefact.
9. **Client data measures the business; your data measures the customer.** Each outranks the other on its own territory, and neither should be quietly promoted onto the other's.
10. **A theme name is a label, not a finding, and a participant count is not a population estimate.**
11. **Build the map while you integrate.** Afterwards you are recording what you remember, and memory is wrong in exactly the places that matter.
12. **The evidence you most want to leave out is usually the evidence the reader most needs.** It is the thing that would have changed their mind, which is why it is uncomfortable.
13. **The trail must run both ways.** From claim to source, and from input to where it was used or why it was not. A trail that only runs one way cannot show what was left out.

## 14. Worked example

*Fictional scenario, used for illustration only. The organisation, inputs, figures and quotes below are invented for the purpose of demonstrating method.*

**INPUT.** A microfinance NGO commissions research into why loan repayment rates have fallen in three rural districts. The report will draw on: a survey of 640 borrowers with cross-tabs from one analyst; 26 depth interviews coded into a theme frame by a second analyst; 40 field photographs of collection points; 18 voice notes recorded by field officers; a repayment extract supplied by the NGO's operations team; a borrower survey from a different NGO published two years ago; and a baseline study the same NGO ran four years ago in two of the three districts.

**PROCESS.**

*Steps 1 to 3.* Twelve source codes assigned across seven prefixes. Method lines written. Two problems surface immediately from the method lines alone. The published survey (`SEC-01`) measured "repayment difficulty" on a self-report scale among borrowers of a different product with a different repayment schedule, so it is not comparable to this study's repayment measure and can only supply context, not a benchmark. And the field photographs (`VIS-01`) have no capture references or dates: the field officers took them on personal phones and the metadata was stripped when they were emailed. Status: unverified. They can illustrate the physical setting with a general caption, and they cannot support any claim about a specific district or time.

*Step 4, and the mismatch caught.* The report intends to claim that borrowers are prioritising other obligations over loan repayment. The required evidence type is behavioural: what borrowers actually paid and to whom. What is available is a survey question asking what borrowers would prioritise if money were short (`QNT-01`, Q22), which is stated intention. The operations extract (`CLI-01`) has actual repayment timing but nothing about competing obligations. The resolution is to split the claim: the survey supports a claim about stated priorities, marked as stated; the interviews support a claim about the sequence borrowers describe following, at 19 of 26 participants; and the behavioural claim about actual prioritisation is restated as a hypothesis per K3 §4.3, with the validation named (a payments diary in the next phase).

*Step 6, and the discrepancy.* The survey puts the proportion of borrowers more than 30 days in arrears at 31%. The operations extract puts it at 22%. The convenient figure depends on the audience: the higher one supports the case for intervention, the lower one flatters operations. Investigation to origin finds two causes, not one. The survey question asked about "any missed payment in the last six months", which is not the same construct as 30-day arrears. And the operations extract counts an account as current if any payment was received in the period, including partial payments, which the survey question does not distinguish. Neither figure is wrong and neither answers the other's question. Both are reported, each with its own definition on the page, and the definitional difference becomes a finding in its own right, because the NGO's own reporting uses the operations definition and had not noticed that it treats a partial payment as a payment.

*Step 6 continued, comparability.* The four-year-old baseline study covers two of the three districts. Wording of the repayment items is identical, base definition is identical, but the sampling frame changed: the baseline sampled from a branch list and this study sampled from the loan book. Verdict: not comparable for prevalence, comparable for within-study rank ordering. The trend chart is dropped; the rank comparison is retained with the frame change stated on the page.

*Step 7.* The survey and the interviews both address why repayment slipped. They are labelled complementary rather than converging: the survey gives prevalence of stated reasons, the interviews give the sequence and the reasoning, and neither is a check on the other. Where they do diverge (the survey ranks crop failure first, the interviews place it third behind school fees and a health event), the method lines show both measured comparable people at a comparable time, so the divergence is real and is reported, with the note that the survey offered a prompted list and the interviews did not, per K2 §4.4.

*Steps 9 and 10.* One well-evidenced finding sits against the narrative: repayment in the third district held steady, and the district differs from the other two in having a fortnightly rather than monthly collection schedule. It complicates a story about regional economic stress. It is placed in the argument as a qualifier, which strengthens the report, and flagged `RESEARCHER REVIEW RECOMMENDED` per K5 §2.1, because whether the NGO can change collection schedules is organisational knowledge the study does not contain. The integration record notes the dropped trend chart, the unverified photographs, the split claim, and the definitional discrepancy.

**OUTPUT.** A twelve-item coded inventory with method lines and verification statuses; a claim-evidence matrix with three mismatches resolved; an evidence map with no empty source cells; a discrepancy log recording the arrears investigation and its two causes; a comparability record dropping one trend comparison and qualifying another; and an integration record covering the unverified image set, the restated behavioural claim, and the disposition of the finding that did not fit.

## 15. Advanced usage

**Multi-market and multi-language projects.** Add a market prefix to every source code and keep the inventories separate until the method lines have been compared, because the same instrument fielded in three markets frequently produces three different constructs once translation and local convention are accounted for. Have a reviewer with local context read the qualitative evidence before findings are fixed, per K5 §2.2, and treat cross-market divergence as a finding requiring investigation rather than an inconsistency to be resolved.

**Longitudinal and tracking programmes.** Maintain the inventory and the source-code taxonomy as a standing asset across waves rather than rebuilding it each time, and version the comparability record. Every change to a question, base definition, segment or sampling frame is logged with the wave it entered, and is disclosed at the point of first comparison. Over several waves this record becomes more valuable than any individual wave's report, because it is the only thing that makes the series interpretable.

**Integrating behavioural or operational data at scale.** Where client systems supply large volumes, the labelling requirement does not relax but the mechanics change: define the extract once (definition, period, filters, refresh date), record it as one coded input, and reference that definition wherever figures from it appear. The most common failure is two extracts pulled weeks apart with different filters, both called "the sales data".

**When an input cannot be verified and cannot be dropped.** Occasionally a client insists on a figure whose origin nobody can establish. Do not verify it by assertion and do not quietly use it. State in the report that the figure is supplied and its derivation is not documented, keep it out of headlines and recommendations, and mark it `RESEARCHER DECISION REQUIRED` per K5 §3.1. If it must carry weight, the honest output is a report that says the conclusion depends on an undocumented figure.

**When two evidence streams cannot be reconciled at all.** Report both readings with the evidence for each and the conditions under which each would be right, and say plainly that the study cannot adjudicate. That is a more useful document than a false synthesis, and it is a legitimate output.

## 16. Skill chain

**Recommended previous skills:**
- **07.06 Qualitative and Quantitative Integration.** Hands over analytically integrated outputs, which this skill then handles for provenance in the document.
- **07.04 Quote and Evidence Extraction.** Hands over verified verbatim with participant identifiers intact, which is the precondition for using quotes at all.
- **13.02 Source and Citation Verification.** Hands over appraised secondary sources with verification status, which this skill records rather than re-establishes.
- **12.01 Research Report Architecture.** Hands over the chapter set and the main-document boundary, which determine where each evidence stream is needed.

**Recommended next skills:**
- **12.03 Research Report Compilation.** Takes the reconciled evidence set, the inventory and the map, and builds the argument and the document on top of them.
- **12.06 Research Report QA.** Takes the evidence map and checks the finished document against it, claim by claim.
- **12.05 Executive Research Reporting.** Takes the map where an executive artefact is owed, since every summary claim needs a row in it, per K2 §5.

**Runs well alongside:**
- **13.03 AI Output Verification**, run against a draft to confirm that no reference, base or identifier was inferred rather than read.
- **13.01 Research Quality Review**, where the reliability of the underlying evidence, rather than its handling, is in question.
- **K2 §5 and §6**, which define the evidence map and the compilation rules this skill implements, and **K5**, at the review points integration generates: consent for visual and audio material, disposition of a finding that does not fit, and use of an input whose origin cannot be established.

---
A Yazi Supplied Skill and resource.
