---
name: research-repository-and-knowledge-curation
description: >
  Structures a body of research so that a question can be answered from work
  already done: what to curate, how to tag it, how to keep it from going stale,
  and how to stop it dying of neglect. Use for "build a research repository",
  "organise our past research", "we keep re-researching the same thing", "nobody
  can find our old studies", "insight library", "how should we tag research",
  "our repository is out of date", "what taxonomy should we use for research".
category: 14 Activation and Knowledge Management
ref: "14.03"
tier: 2
inherits: [K2, K3, K4, K5]
---

# Research Repository and Knowledge Curation

## 1. One-line description

A method for structuring an organisation's accumulated research so that someone with a question can get an answer out of it, built on curated findings rather than stored documents, with a status convention that keeps ageing evidence from being read as current and a maintenance model that survives the first person who set it up leaving.

## 2. What this skill is used for

**The research problem it solves.** Organisations accumulate research and lose it. Not physically: the files exist, on a shared drive, in a folder structure organised by the year and the agency. What is lost is retrievability, because nobody searches a repository for a document. They arrive with a question ("do we know anything about why small business customers leave in the first year?") and the store can only offer them a list of study titles, one of which may contain the answer somewhere in eighty pages. The result is a research function that re-commissions work it has already done, a stakeholder population that believes the research team never has an answer ready, and a body of expensive evidence whose value decays to nothing about eighteen months after each project closes. The second failure is worse and less visible: a repository that does work becomes trusted, and a trusted repository serving a four-year-old finding as though it were current does more damage than no repository at all. This skill treats the repository as an answering system rather than a filing system, curates findings rather than reports, imposes a deliberately small controlled vocabulary, and installs the status and review discipline that ageing evidence requires. It also states plainly what most repository projects discover too late: repositories fail on maintenance, not on design.

**Where it sits in the research lifecycle.** After delivery, alongside and after activation. It is infrastructure rather than a project step: each study contributes to it, and it in turn feeds the front of the lifecycle by telling the next brief what is already known.

**Typical use cases.**
- A research function with several years of studies and no way to answer a question from them.
- The same question being commissioned twice, discovered only when someone recognises the brief.
- A team that has just lost the person who remembered where everything was, which is how most organisations discover they had no repository.
- A repository that exists, was well designed, and has not been updated for a year.
- A merger, restructure or agency change that leaves two bodies of research to reconcile.
- Preparation for cross-study work, where a corpus has to be assembled before any synthesis is possible.
- An audit or governance requirement to show what evidence supports a claim the organisation is making.

**Who uses it.** Insight and knowledge managers who own the repository; research directors deciding whether to build one and what it should cost; researchers contributing findings at the end of a project; strategists, designers and product managers who consume it; anyone asked "have we looked at this before".

## 3. When to use it

- There is a body of completed research, typically more than fifteen or twenty studies, and no reliable way to interrogate it.
- Questions arrive that existing research probably answers, and answering them takes days of manual searching by whoever has been there longest.
- Research is being re-commissioned on subjects already covered.
- Findings from different studies need to sit side by side, which requires them to be structured comparably.
- The organisation is about to run cross-study synthesis and the corpus must be assembled and standardised first.
- An existing repository has degraded, and the question is whether to fix it, rebuild it or retire it.
- Institutional memory currently lives in one or two people and the risk of losing it is now visible.
- Evidence needs to be traceable for governance, regulatory or client-audit reasons.

## 4. When NOT to use it

- **There is not enough research to warrant one.** Below roughly fifteen studies, a well-maintained index (a single table listing study, date, question, sample, method, headline findings and where the file lives) outperforms a repository, costs almost nothing and does not require a curator. Building a taxonomy for eight studies produces a taxonomy tested on eight studies, which then constrains the next eighty. Start with the index and let the categories emerge from real search behaviour.
- **Nobody will own maintenance.** This is the decisive test and it should be applied before any design work. An unmaintained repository is worse than no repository, because it looks authoritative while serving stale evidence, and because people stop contributing to it long before they stop trusting it. If no named person has time allocated to curation, do not build it. Say so, and propose the index instead.
- **The real problem is that current findings are not being used.** A repository does not fix activation failure; it gives unused findings a tidier place to be unused. If the last three studies produced no decisions, the problem is upstream and belongs to **14.01 Insight Activation and Socialisation**.
- **The task is to answer one specific question from existing studies.** That is **14.04 Meta-Analysis Across Studies**, and it can be done without a repository, though it is much cheaper with one. Do not build infrastructure to answer a single question. The boundary in the other direction: 14.03 makes a corpus interrogable; 14.04 interrogates it.
- **The requirement is raw data storage, participant records or personal data management.** That is data governance, with retention schedules, lawful basis and deletion obligations, and it is a different discipline with legal consequences. Route the participant-data dimension to **13.05 Research Ethics and Consent Design** before any personal data enters a searchable system, and never let a repository become the informal home for material that a retention schedule requires to be deleted.
- **Consent or contract does not permit retention or reuse.** Some studies cannot be curated: participant consent covered a single purpose, a client contract restricts reuse, or a data-sharing agreement has expired. These are excluded, and the exclusion is recorded so that a future searcher knows the study existed rather than concluding the subject was never researched.
- **What is actually wanted is a document store.** If the requirement is that people can find and download final reports, that is a well-organised file share with a naming convention, and it is legitimate and cheap. Say that plainly rather than building a curation programme around it. Curation earns its cost only where findings, not files, are the thing being retrieved.
- **The evidence base is not trustworthy.** Curating unreviewed findings distributes them with the repository's implied endorsement. Where the quality of past work is genuinely unknown, either curate with an explicit quality assessment per entry or run **13.01 Research Quality Review** on the material that will carry weight, before it is entered.

## 5. Required inputs

**Required.** Without these the skill cannot run. If absent, ask. If no answer is available and work must proceed, state the assumption at the point where it bites, per K5 §5.

- **The questions the repository must be able to answer.** Real ones, collected from real requests, not invented categories. A repository designed from a topic taxonomy answers taxonomy questions; a repository designed from the last thirty questions people actually asked answers those.
- **An inventory of the existing research**: study, date, commissioner, method, sample, market, and where the material lives. Studies that cannot be located are recorded as known to exist and unavailable, which is information (K4 §6.4).
- **A named curator with allocated time.** Not a volunteer and not a committee. If none exists, stop and say so, per Section 4.
- **The confidentiality, consent and contractual terms attached to each study**, or an explicit statement that they are unknown, which caps what may be entered.

**Optional, and what each one adds.**

- **A log of past research requests and how they were answered**: the single most useful input available, because it reveals the actual retrieval patterns, the vocabulary people use, and the questions the repository will be judged on.
- **Existing analysis outputs rather than only final reports** (theme records, code frames, cross-tab sets, verbatim files): allow findings to be curated with their bases and references intact, which is the difference between a curated finding and a quotation from a slide.
- **The decisions each study informed, and what was decided**: turns the repository from a store of findings into a record of what the organisation knows and what it did about it, which is what makes it useful in governance and in the next brief.
- **A record of which studies have been quality reviewed**: lets confidence and status be assigned per entry rather than assumed.
- **Search or usage analytics from any existing store**: shows what people looked for and did not find, which is the highest-value input to taxonomy design.
- **A budget and a platform decision already made**: determines what can be automated and what must be manual, though the method here is deliberately independent of any particular system.

## 6. Questions to ask before starting

1. **What questions must this answer, and who asks them?** Determines the unit of curation, the taxonomy and the success test. *Default if unanswered:* collect the last thirty requests before designing anything; if none are recorded, interview three heavy users and take their questions verbatim.
2. **Who is the curator, how much time do they have, and what happens when they leave?** Determines whether to build at all and how heavy the schema can be. *Default:* if there is no answer, flag `RESEARCHER DECISION REQUIRED` (K5 §2.1) and recommend the index instead. This question decides the project.
3. **How far back should curation go?** Determines the backfill cost, which is usually the largest single cost. *Default:* curate forward from today plus the studies that are actually being asked about, which is typically a small and identifiable set.
4. **What consent, contractual and confidentiality constraints apply, and to which studies?** Determines what may be entered, at what granularity, and who may see it. *Default:* treat consent as covering the original purpose only until confirmed otherwise, and exclude verbatim and participant-level material pending review (13.05).
5. **What is the review capacity for ageing entries?** Determines whether a status convention can be maintained, and therefore whether the repository can be trusted over time. *Default:* set review periods long enough to be honoured rather than short enough to look rigorous, and mark entries dated rather than removing them when a review is missed.
6. **Are there two or more existing stores, and will this replace them?** Determines migration and, more importantly, whether the old ones are closed. *Default:* two live repositories are worse than one bad one; get an explicit decision to retire the others.
7. **What does the organisation call things?** Determines the controlled vocabulary. *Default:* take the terms from the request log and the business's own documents, not from a research textbook or a previous employer's scheme.

## 7. Step-by-step methodology

**Step 1. Define the repository by the questions it must answer.** Write the question set first, from the request log or from user interviews, in the questioners' own words. Group them by shape rather than by topic: subject questions ("what do we know about churn among small business customers"), population questions ("what have we run with clinicians"), method questions ("have we ever used diary studies here"), decision questions ("what evidence supported the pricing change"), and provenance questions ("where did this number come from"). Each shape imposes a different requirement on the schema, and a repository designed only for subject questions fails the other four. Then write the acceptance test: three real questions that a competent colleague must be able to answer from the repository in under ten minutes without asking anyone. *Correct result: fifteen to thirty real questions, grouped by shape, and three named acceptance tests written before any structure is designed.*

**Step 2. Choose the unit of curation, and make it the finding.** The instinct is to curate studies, because studies are what the organisation produced. Nobody searches a repository for a study. The unit that answers a question is a **finding or an insight**: a single claim, stated in one or two sentences, that stands on its own with its evidence attached. A study becomes a parent record holding method, sample, dates, objectives and the file; the findings are the children, and they are what is searched, tagged, aged and retrieved. This is the highest-cost decision in the skill, because curating findings requires someone to read the study and extract them, whereas curating documents requires only that they be uploaded. It is also the decision that determines whether the repository works. Two rules make the unit workable. A finding is **atomic**: one claim, so that it can be tagged, dated and superseded independently. And a finding is **self-sufficient**: it carries its own base, population, period and confidence, so that it is safe read alone, which is how it will be read. Where a study's real contribution is a synthesised insight rather than a set of discrete findings, curate the insight and link the findings beneath it (K2 §2). *Correct result: a defined entry unit, with an atomicity test applied to a sample of ten findings from a real study and the failures rewritten.*

**Step 3. Write the entry schema, with provenance non-negotiable.** Each finding entry carries: the claim, in one or two sentences; the evidence reference in K2 §4 format (question or theme, base description, base size, weighting, test where a comparison is claimed, or participant identifiers and prevalence for qualitative); the confidence level per K3 §2 with the reason; the parent study; the population and market it applies to; the fieldwork period, which is not the publication date; the method; the status (step 5); the review date; the tags (step 4); the decision it informed, where known; and the curator and entry date. **The provenance fields are the load-bearing ones.** A repository entry that cannot be traced back to its study, its question and its base is a rumour with a search index, and it will eventually be quoted in a business case. Per K2 §7, the acts that break traceability (rounding for readability, dropping a base to fit a field, merging two findings into one entry) are exactly the acts curation invites, because curation is compression. Guard the schema against them by making base and reference mandatory fields that cannot be left blank, and by rejecting entries that fail. *Correct result: a schema in which every field has a definition and a filled example, and in which no entry can be saved without a reference and a base.*

**Step 4. Design a small controlled vocabulary, not a large free one.** Tagging fails in two directions. Free tagging produces "onboarding", "on-boarding", "new customer journey" and "activation" as four labels for one concept, and search returns a quarter of the relevant entries while appearing to work. Over-engineered taxonomy produces a hundred and forty terms nobody can navigate and contributors guess at, which produces the same result more expensively. **A small controlled vocabulary beats a large uncontrolled one, and the discipline is refusing to add terms.** Use a small number of independent facets rather than one deep hierarchy, since findings belong to several categories at once and a tree forces a false choice. The working facets: **topic** (aim for fifteen to thirty terms, drawn from the request log, not from the business's org chart, which reorganises); **audience or population** (the people the finding is about, using the organisation's own segment names); **method** (a short closed list, so method questions are answerable); **market or geography**; **date** (fieldwork period, as a field rather than a tag); **decision informed** (which links findings to what the organisation did, and is the facet that makes the repository useful in governance); **confidence** (per K3); and **status** (step 5). Write a definition for every term, including the boundary case that distinguishes it from its neighbour, because undefined terms drift within months. Set a rule for adding terms: a new term requires a stated definition, a check that no existing term covers it, and the curator's approval. Review the vocabulary twice a year and merge the terms that collected fewer than three entries. *Correct result: a facet scheme with defined terms, a documented process for adding one, and a test in which three people independently tag the same five findings and agree on at least four.*

**Step 5. Install the status convention and the ageing discipline.** Every finding ages, and a repository that does not say so will serve a four-year-old result to someone who assumes it is current. This is the failure that destroys trust in a repository permanently, because it is discovered by the person who acted on it. Four statuses, applied per finding rather than per study, since findings within one study age at different rates. **Current:** within its review period, no known invalidating event, safe to use. **Dated:** past its review period or predating a known change, still the best available evidence, and usable only with its date and the reason for the flag visible. **Superseded:** a later study has answered the same question, with a link to the entry that replaced it. Superseded entries are kept, not deleted: the historical record is what lets a change over time be seen at all, and deleting it makes the organisation's knowledge look static. **Retired:** no longer applicable, because the product, policy, market or population it described no longer exists. Kept, with the reason. Then set review triggers rather than relying on dates alone: a date trigger (a review period set by volatility, typically twelve months for behavioural findings in a fast-moving category, longer for structural or attitudinal ones); an event trigger (a product change, a pricing change, a regulatory shift, a market entry, a reorganisation of the population); and a contradiction trigger (a new study reports something incompatible, which fires the process in step 6). Ageing findings are re-statused, never quietly edited: changing a finding's wording destroys the record of what was believed and when. *Correct result: every entry carrying a status, a review date and at least one named event trigger, and a documented rule that a missed review downgrades to dated automatically rather than leaving the entry current by default.*

**Step 6. Handle duplication and contradiction explicitly.** These are different problems with opposite remedies. **Duplication** is the same finding entered twice, usually because two studies reported the same result or because a finding was re-entered on a later wave. Merge, keeping the strongest evidence reference and linking the second study as corroboration, which is more valuable than either entry alone because it records independent replication. **Contradiction** is two studies that disagree, and the instinct to resolve it is the thing to resist. Both entries stay, both stay discoverable, and they are linked by an explicit contradiction record that states: what each found, with its base and period; the diagnosed cause of the difference where one can be established (different population, different question wording, different period with real change, different method, or genuine disagreement); which is better evidenced and why; and what would settle it. A repository that shows only the most recent or most convenient of two conflicting findings has silently made an analytical judgement on behalf of every future searcher, and per K4 §4.1 that is exactly what must not happen. The contradiction record is frequently the most useful entry in a repository, because it is the honest state of knowledge and it is the specification for the next study. *Correct result: a merge log for duplicates and a contradiction register in which every pair carries a diagnosed cause or an explicit statement that the cause is unknown.*

**Step 7. Set access, confidentiality and participant-data rules before anything is loaded.** Searchability changes the risk profile of material that was safe in a report. A verbatim quote in an aggregated study is one thing; the same quote, tagged by market, segment and employer, retrievable by anyone in the organisation, is another, and a participant who consented to the first did not consent to the second. Decide, per class of material: what may be entered at all (findings, yes; participant-level records, usually not); what is entered but access-restricted (commercially sensitive findings, findings about identifiable individuals or small groups, anything covered by a client confidentiality term); what is redacted on entry (names, employers, locations, any detail that identifies in combination); and what retention period applies, since a repository is not exempt from a deletion schedule. Route the consent question to **13.05** rather than deciding it here. Record the constraint on the entry, so a searcher can see that restricted material exists rather than concluding nothing was found, which is a meaningful difference. *Correct result: an access model with named rules per material class, a redaction standard, and a retention position agreed with whoever owns data governance.*

**Step 8. Build the contribution workflow into project close, and make curation someone's job.** Repositories fail on maintenance. The design is done once by someone enthusiastic; the contribution happens weekly forever, by people whose project has finished and whose attention has moved. Three things make the difference. **Contribution is a step in project close**, not an afterthought: the study is not complete until its findings are entered, and this is stated in the project plan and honoured by the research lead first. **The contribution cost is kept low**, which is why the schema must be small: a project contributing five to fifteen findings is realistic, one contributing sixty is not, and a schema with twenty fields will not be filled. **A curator reviews every contribution** for atomicity, provenance, tagging against the controlled vocabulary, and status. The review is what stops vocabulary drift and unreferenced entries, and it takes minutes per finding. Without it the repository degrades invisibly for about a year and then visibly all at once. Set a maintenance rhythm: contributions reviewed weekly, ageing reviewed quarterly, vocabulary reviewed twice a year, and the acceptance tests from step 1 re-run annually. *Correct result: a written workflow with the curation step inside project close, a curator with allocated hours, and a maintenance calendar with named dates.*

**Step 9. Backfill selectively, not comprehensively.** Retrospective curation is where repository projects die: the team decides to enter six years of studies, spends four months on it, and never starts the forward workflow. Curate forward from today, and backfill only what is being asked about, identified from the request log. A study nobody has asked about in three years is unlikely to be asked about now; leave it in the index with enough metadata to be findable, and curate it if and when a question arrives. *Correct result: a forward workflow running from day one, and a named backfill list of typically ten to twenty studies, prioritised by request frequency.*

**Step 10. Measure use, and act on what it shows.** Track what is searched, what is found, what is searched for and not found, and what is cited. The most valuable of these is the failed search: it names the gap in the corpus, the term missing from the vocabulary, or the finding that exists but is untagged. Ask consumers to cite entry references when they use a finding, which is the cheapest instrumentation available and turns use measurement into a search (see 14.01 step 9). Report use as retrieval and citation, not as logins, and re-run the acceptance tests annually with someone who did not build the repository. *Correct result: a short quarterly report covering successful retrievals, failed searches with their diagnosis, contribution rate by team, and the proportion of entries overdue for review, which is the leading indicator of decay.*

## 8. Analytical framework

The repository is built on a two-level record with a status layer over it:

    Study (parent)                    Finding (child, the searchable unit)
    ├─ objectives, method             ├─ claim
    ├─ sample, fieldwork period       ├─ evidence reference, base, confidence  (K2, K3)
    ├─ markets, commissioner          ├─ population, period, method inherited
    ├─ files and outputs              ├─ tags: topic, audience, method, market, decision
    └─ consent and contract terms     └─ status: current | dated | superseded | retired
                                          └─ review date + event triggers
                                          └─ links: corroborates, supersedes, contradicts

**Applying it.** The three link types do most of the work and are what distinguish a repository from a list. **Corroborates** records independent replication, which raises confidence legitimately in a way that repeating the same study does not. **Supersedes** records that a question has been re-answered and preserves what was previously believed, without which no change over time is visible. **Contradicts** records honest disagreement and refuses to resolve it on the searcher's behalf.

**The retrieval test, which is the only real measure of design quality.** Take a real question. Can a colleague who did not run any of the studies find the relevant findings, see their bases, see how old they are, see whether anything contradicts them, and reach a defensible answer, in ten minutes, without asking a person? Every design decision in this skill is justified by that test and by nothing else. A taxonomy that is elegant and fails the test is a bad taxonomy.

**The decay curve.** Repositories do not fail at launch. They pass acceptance, then contribution rate falls as the founding projects close, then vocabulary drifts as unreviewed tags accumulate, then a stale finding is served as current and trust breaks, then people go back to asking the longest-serving researcher. The intervention points are contribution (step 8) and ageing (step 5), and both are maintenance rather than design, which is why the ownership question in Section 6 decides the project.

## 9. Output format

**A. Repository specification**

1. **Purpose and question set.** The questions it must answer, grouped by shape, and the three acceptance tests.
2. **Entry schema.** Every field, with definition, whether it is mandatory, and a filled example.
3. **Controlled vocabulary.** Each facet, its terms, and a definition per term including the boundary case.
4. **Status convention.** The four statuses, review periods by finding type, and the event triggers.
5. **Access and confidentiality model.** Rules per material class, redaction standard, retention position.
6. **Contribution workflow.** Where in project close it sits, who does it, what the curator checks.
7. **Maintenance calendar.** Review rhythms and named owners.

**B. The record structures**

Study record:

| Field | Content |
|---|---|
| Study ID, title, commissioner, date | |
| Objectives (as questions) | |
| Method, sample, fieldwork period, markets | |
| Quality review status | |
| Consent, contract and confidentiality terms | |
| Files and where they live | |

Finding record:

| Field | Content |
|---|---|
| Finding ID, parent study | |
| Claim (one to two sentences) | |
| Evidence reference, base description, base size, test | |
| Confidence and reason (K3) | |
| Population, market, fieldwork period | |
| Tags (topic, audience, method, market, decision informed) | |
| Status, review date, event triggers | |
| Links (corroborates, supersedes, contradicts) | |
| Access restriction, if any | |
| Curator, entry date | |

**C. Contradiction register**

| Pair | Finding A | Finding B | Diagnosed cause | Better evidenced, and why | What would settle it |
|---|---|---|---|---|---|

**D. Quarterly use report.** Retrievals, citations, failed searches with diagnosis, contribution rate by team, and entries overdue for review.

**When the evidence is thin.** An entry is never completed by inference. A study whose base sizes cannot be recovered is entered with the base field marked unavailable and its confidence capped at low, or it is not entered as a finding at all and remains only in the study index (K4 §2.1, §2.5). A finding whose fieldwork date cannot be established is entered as undated and is not eligible for current status, because status is a claim about age. Where a study's material cannot be located, the study record is created with the file field marked missing, which tells a future searcher the work exists rather than letting them conclude the subject was never covered.

## 10. Quality checks

Run on the design before launch, and on a sample of entries quarterly. Sits on top of K4 §8.

1. Can three real questions from the request log be answered from the repository in under ten minutes by someone who did not run the studies?
2. Does every finding entry carry an evidence reference, a base description and a base size, with no blanks filled by inference?
3. Is every finding atomic, so that it can be superseded or retired without affecting a claim it was bundled with?
4. Is every finding self-sufficient, meaning it is still accurate read alone with no other entry visible?
5. Does every entry carry its fieldwork period rather than only its publication or entry date?
6. Does every entry carry a confidence level with a stated reason, per K3?
7. Does every entry carry a status, a review date and at least one event trigger?
8. Do entries past their review date downgrade to dated automatically, rather than remaining current by default?
9. Are superseded and retired entries retained and linked, rather than deleted?
10. Is every contradiction between entries recorded, discoverable from both sides, and diagnosed or explicitly marked undiagnosed?
11. Do three people independently tagging the same findings agree, and has vocabulary drift been checked since the last review?
12. Has every entry been checked against consent, contract and confidentiality terms, with restricted material visible as restricted rather than invisible?
13. Is any personal or identifying participant data present that should not be, and is the retention position agreed with data governance?
14. Is the contribution rate holding, and what proportion of entries are overdue for review?
15. Is there a named curator with allocated time, and a stated succession if they leave?

## 11. Common failure modes

| Failure | How to recognise it | How to prevent it |
|---|---|---|
| **Curating documents rather than findings** | Search returns study titles; the searcher still has to read eighty pages | Step 2. The finding is the unit; the study is the parent record |
| **Death by maintenance** | Contribution rate falls after the founding projects close | Step 8. Curation inside project close, a curator with hours, a maintenance calendar |
| **The stale finding served as current** | A four-year-old result quoted in a business case with no date attached | Step 5. Status per finding, automatic downgrade on a missed review |
| **Vocabulary sprawl** | Four tags for one concept; search returns a quarter of what exists | Small controlled vocabulary, curator review of every contribution, term merges twice a year |
| **The over-designed taxonomy** | A hundred and forty terms, contributors guessing, inter-tagger agreement poor | Few facets, defined terms, the three-person tagging test |
| **The backfill that never ends** | Four months of retrospective entry and no forward workflow | Step 9. Forward from today; backfill only what is asked about |
| **Silent contradiction resolution** (the signature AI failure) | Two studies disagreed and the repository shows one, usually the more recent | Step 6. Both entries stay, linked, with a diagnosed cause (K4 §4.1) |
| **Provenance decay** | A claim in the repository with no base, no question reference, no date | Mandatory fields that cannot be saved blank; curator rejects entries that fail |
| **Compression that changes the claim** | The entry says "customers prefer X"; the study said "among lapsed users in one market" | Self-sufficiency test at step 2. The entry carries its population, always |
| **Merged findings** | One entry making two claims, so it cannot be superseded without losing the other | Atomicity test at step 2 |
| **Two live repositories** | Contributors ask which one to use; searchers check both and trust neither | Get an explicit retirement decision on the old store before launching the new one |
| **The invisible restriction** | A searcher concludes nothing exists when restricted material does | Show that restricted entries exist without showing their content |
| **Participant data drift** | Verbatim with employers and locations, searchable organisation-wide | Step 7 before loading anything. Redaction standard, access model, 13.05 |
| **The repository nobody uses** | Good design, high contribution, no retrievals | Step 1. It was built from a taxonomy rather than from the questions people ask |

## 12. AI guardrails

Skill-specific only. Universal prohibitions are inherited from K4. Curation is compression, and compression is where provenance is lost, so the traceability rules in K2 §7 apply with unusual force here.

1. **Never create a repository entry for a finding you have not read in its source.** An entry generated from a report summary, a slide title or another entry is a restatement presented as evidence, and it will be cited as though it were the study.
2. **Never fill a base, a date, a sample description or a confidence level by inference.** If the study does not state it, the field is marked unavailable and the entry's confidence is capped accordingly (K3 §3.2). A plausible base is worse than a blank one, because a blank invites checking.
3. **Never merge two findings into one entry to make it read better.** Atomicity is what allows a finding to be superseded, retired or contradicted independently. A merged entry corrupts the status layer permanently.
4. **Never rewrite a finding's claim during curation in a way that changes its population, period or scope.** Curation may shorten. It may not widen. "Among lapsed users in one market" does not become "customers".
5. **Never resolve a contradiction between two entries by selecting one.** Both remain discoverable and linked, with the cause diagnosed or explicitly marked unknown. Selecting is an analytical judgement made invisibly on behalf of everyone who will ever search.
6. **Never delete a superseded or retired entry.** The historical record is what makes change over time visible and what prevents the organisation believing it has always known what it currently knows.
7. **Never mark an entry current without checking its review date and triggers.** Status is a claim about the evidence's present applicability, and asserting it without checking is the fabrication that this skill's failure mode is built on.
8. **Never enter participant-level material, verbatim with identifying detail, or personal data without an explicit consent and access decision.** Searchability is a new purpose, and it is not covered by consent given for a report (13.05).
9. **Never generate tags outside the controlled vocabulary.** A plausible new tag is how vocabulary sprawl starts, and it is invisible until search quality has already degraded. Propose the term to the curator instead.
10. **Never assert that the repository contains no evidence on a subject.** The honest statement names the searches run and the terms used, per the same principle as K4 §6.4: absence of retrieval is a fact about the search, not about the corpus.

## 13. Best-practice principles

1. **Design from the questions people ask, not from a taxonomy of the subject.** Every failed repository was organised sensibly and answered nothing.
2. **Curate findings, not files.** Nobody searches for a PDF. The extra cost of extracting findings is the entire value of the exercise.
3. **A small controlled vocabulary beats a large uncontrolled one, and the skill is refusing to add terms.** Every new tag makes the existing ones slightly less reliable.
4. **Facets, not a tree.** A finding is about a topic, a population, a market and a decision simultaneously, and forcing it into one branch loses three of the four ways people will look for it.
5. **Status is per finding, not per study.** Findings inside one study age at different rates, and the structural ones outlive the behavioural ones by years.
6. **An out-of-date finding presented as current is worse than no repository.** Trust breaks once, and it breaks at the moment someone acts on it.
7. **Keep what has been superseded.** An organisation that deletes its old answers cannot see that anything changed, and change over time is often the most valuable thing a corpus holds.
8. **Two studies that disagree are an asset, not a defect.** Record the disagreement, diagnose it if you can, and let the searcher see both. It is also, usually, the next brief.
9. **Repositories fail on maintenance, not on design.** Spend the effort on the contribution workflow and the curation hours, not on the schema.
10. **Contribution must be part of project close or it will not happen.** After the project ends, the attention has gone and no reminder recovers it.
11. **Backfill only what is asked about.** Comprehensive retrospective curation is how repository programmes consume their budget and never launch.
12. **The failed search is the most valuable telemetry you have.** It names the gap, the missing term or the untagged entry, and it is free.
13. **Ask consumers to cite the entry reference.** It costs them nothing, it makes use visible, and it is what lets you defend the repository's budget.
14. **A repository that answers questions attracts contributions.** One that does not will not be fed, whatever the policy says, and no amount of governance overcomes that.

## 14. Worked example

*Fictional scenario, used for illustration only. The organisation, studies, findings and figures below are invented to demonstrate method.*

**INPUT.** A software company's design research team has run 63 studies over five years across two products: usability tests, diary studies, three surveys and a segmentation. Material sits in a shared drive organised by year and researcher. The team has lost two of its four researchers in the past year. The trigger is that a product manager commissioned a study into why users abandon the setup flow, and a departing researcher mentioned in her final week that this had been covered twice, in 2023 and 2024.

**PROCESS.**

*Step 1.* The request log does not exist, so the last four months of requests are reconstructed from the team's messages: 31 questions. Grouped by shape, they are 14 subject questions, 6 population questions ("have we ever spoken to administrators rather than end users"), 4 method questions, 5 decision questions ("what did the segmentation actually change"), and 2 provenance questions. The population and decision questions are a surprise, and they change the schema: a subject-only taxonomy would have failed a third of the real demand. Three acceptance tests are written, one of them the setup-flow question that triggered the project.

*Steps 2 and 3.* The unit is set as the finding. A pilot on one 2024 usability study extracts eleven findings, of which four fail the atomicity test (each bundles an observation with a recommendation) and two fail self-sufficiency (they say "users" where the study covered trial users on one product only). All six are rewritten. The pilot establishes the real cost: roughly ninety minutes per study for extraction and entry, which sets the backfill budget realistically for the first time.

*Steps 4 and 5.* The topic facet is drafted at 41 terms from the team's existing folder names and cut to 19 by merging near-synonyms and deleting terms that would carry fewer than three findings. Three researchers independently tag the same five findings; agreement is four of five, with the disagreement on a boundary between two topic terms, which is resolved by writing the boundary case into both definitions. Review periods are set by finding type: twelve months for findings about the interface, thirty-six for findings about user goals and workflow, on the reasoning that the interface changes quarterly and the workflow does not.

*Step 6, and the judgement call.* The two prior setup-flow studies are curated, and they disagree. The 2023 study (moderated usability, 12 participants, both products) found abandonment concentrated at the account-linking step. The 2024 study (unmoderated, 340 sessions, one product only) found abandonment spread across the flow with no clear peak. The convenient move is to enter the more recent, larger study and mark the older one superseded, which is what a simple recency rule would do. Diagnosis says otherwise: the studies used different methods, covered different products and ran either side of a change to the account-linking step. This is not supersession, it is two findings about different things. Both are entered as current, both linked by a contradiction record naming the diagnosed cause, and the record itself becomes the most useful entry in the repository for the product manager who triggered the project: it tells him the question is open, why, and specifically that the 2024 evidence does not cover his product's linking step. His new study is rescoped rather than cancelled, which is a better outcome than either cancelling it or running it as originally briefed.

*Steps 7 to 10.* Verbatim entry is restricted: quotes may be entered without employer, role detail or product-account identifiers, and clips are not entered at all pending a consent review routed to 13.05, because participant consent from the 2022 and 2023 studies covered internal report use and does not obviously cover a searchable system. Contribution is written into the team's project-close checklist and the research lead enters her own study first. Backfill is limited to 14 studies drawn from the request log, not 63. Curation is allocated at four hours a week to a named person, with a stated succession. `RESEARCHER DECISION REQUIRED` is flagged on the retention question for the 2021 studies, whose consent terms cannot be located (K5 §2.4).

**OUTPUT.** A repository specification built from 31 real questions; a finding-level schema with mandatory provenance fields; a 19-term topic facet with defined boundaries and four other facets; a status convention with review periods differentiated by finding type; a contradiction register whose first entry immediately rescoped a live study; an access model with verbatim restricted and clips withheld pending consent review; a forward contribution step inside project close; a 14-study backfill list; and a named curator with allocated hours and a succession plan.

## 15. Advanced usage

**From repository to standing synthesis.** Once a corpus is curated to a consistent schema, cross-study work becomes cheap rather than heroic, and that is the point at which **14.04 Meta-Analysis Across Studies** stops being a project and becomes a routine. A mature repository maintains standing state-of-knowledge statements on the organisation's five or six recurring questions, each linked to its supporting findings, each updated when a contributing finding changes status. That artefact, rather than the search interface, is what senior stakeholders actually use.

**Curating for gap identification.** The repository's silences are informative. Mapping findings against the business's decision areas exposes where the organisation has no evidence, and mapping them against populations exposes who has never been researched. Because a company researches what it is interested in, these gaps are systematic rather than random, and naming them is a direct input to **10.04 Research Gap Identification**.

**Merging two corpora.** After a merger, an agency change or a restructure, two repositories with different schemas must become one. Do not migrate field by field. Map both to the target schema, curate findings rather than converting records, and treat every incoming entry as unverified until its provenance is confirmed against the source study. Expect the older corpus to lose a third of its entries at the provenance gate, and record that as an exclusion list rather than absorbing unreferenced claims.

**Governance and audit use.** Where an organisation must show what evidence supports a public or regulatory claim, the decision-informed facet and the provenance fields turn the repository into an audit trail. This is worth designing for from the start, since retrofitting the link between findings and decisions is nearly impossible once the decisions have been made.

**When the standard approach does not fit.** For a small team with no curator, do not build a light version of this; build the index instead and revisit in a year. For an organisation whose research is almost all continuous tracking rather than discrete studies, invert the model: the parent record is the tracker and the findings are dated observations, with supersession as the normal state rather than the exception.

## 16. Skill chain

**Recommended previous skills:**
- **14.01 Insight Activation and Socialisation.** Hands over dated core findings with references, confidence levels and review triggers already attached, which is the cheapest possible source of curated entries.
- **14.02 Research Debrief and Workshop Design.** Hands over the decision record, so the repository can link findings to what the organisation actually decided, which is what makes the decision facet work.
- **12.06 Research Report QA** and **13.01 Research Quality Review.** Hand over an assessment of whether a study's findings are sound enough to be curated with the repository's implied endorsement.

**Recommended next skills:**
- **14.04 Meta-Analysis Across Studies.** Takes the curated corpus and interrogates it against a new question. This skill makes the corpus interrogable; 14.04 does the interrogating.
- **10.04 Research Gap Identification.** Takes the repository's mapped silences and turns them into a research agenda.
- **01.02 Business Problem to Research Question.** Takes what the repository shows is already known and scopes the next study to add rather than repeat.

**Runs well alongside:**
- **13.05 Research Ethics and Consent Design**, which owns every consent, retention and participant-data question a searchable system raises.
- **13.06 AI Research Governance**, where AI assists extraction, tagging or summarisation and that involvement must be disclosed per entry (K5 §7).
- **K2 §4 and §7**, which define the reference format every entry carries and name the acts of compression that curation must not commit, and **K3 §2**, which supplies the confidence levels the schema records.

---
A Yazi Supplied Skill and resource.
