---
name: ai-research-governance
description: >
  Governs how AI is used across a research project: what may and may not be sent
  to an AI system, what consent covers, disclosure to clients and participants,
  the audit trail that lets an AI-performed step be reconstructed, human sign-off
  records, model version and reproducibility, measured accuracy for AI coding and
  classification, confidentiality and intellectual property, and a governance
  policy a research team can actually adopt. Use when someone says "can we put
  this in an AI tool", "what do we tell the client about AI", "do we need an AI
  policy", "is this covered by the consent", "the model changed and the numbers
  moved", or "how do we prove a human checked this".
category: 13 Research Quality, Ethics and Governance
ref: "13.06"
tier: 1
inherits: [K2, K3, K4, K5]
---

# AI Research Governance

## 1. One-line description
Sets and operates the rules for AI use across a research project: what may be sent, under what consent, with what human check, disclosed how, recorded how, and reproducible how, and produces a governance policy a working research team can adopt without hiring anyone.

## 2. What this skill is used for

**The research problem it solves.** AI arrived in research practice faster than the arrangements around it, and the gap shows up as a set of small, ordinary decisions nobody owns. A researcher pastes a transcript into a system to get a summary. Nobody has decided whether that transcript may leave the project environment, whose consent covers it, whether the client's contract permits it, whether the input is retained by the provider, or what happens to it afterwards. A second researcher runs the coding through a different system three weeks later, and the code frame is different, and nobody can reconstruct why because neither step was recorded. A report goes to a client describing a thematic analysis, and the client reasonably assumes a person read the transcripts.

None of this is misconduct. It is the absence of a decision. And the cost is not usually a breach: it is that **the work becomes unreconstructable.** When a client asks how a theme was derived, or a finding is challenged, or a regulator asks what processed personal data, the answer has to be assembled from memory, and memory is not a record.

The other half of the problem is disclosure. Clients increasingly ask, contracts increasingly require it, and the standing professional obligation is not to present AI-assisted work as though it were not (K5 §7). But a disclosure that says "AI tools were used in the preparation of this report" tells a reader nothing they can act on. **A useful disclosure says what the AI did, what a human verified, and who that human was.**

This skill is deliberately operational. Aspirational AI principles documents are abundant and change no behaviour. What changes behaviour is a one-page classification a researcher can apply in ten seconds, a step register that takes a line per use, and a disclosure paragraph somebody has already written.

**Where it sits.** Across the whole project, from proposal to archive. It is set up once for a team and applied per project, and it runs before AI is used rather than after.

**Typical use cases.**
- Deciding whether a specific dataset, transcript or client document may be sent to an AI system.
- Establishing whether existing participant consent covers AI processing, and what to do when it does not.
- Writing the AI disclosure for a proposal, a report or a participant information sheet.
- Building the audit trail so an AI-performed step can be reconstructed months later.
- Governing AI-performed coding, classification or extraction, including the accuracy figure that must be disclosed.
- Handling a model or version change mid-project, and explaining why a rerun did not reproduce.
- Answering a client's or a procurement team's questions about AI use, confidentiality and intellectual property.
- Writing a team's AI governance policy for the first time, or replacing one nobody follows.

**Who uses it.** Research leads and directors accountable for delivery; operations and quality leads writing the team's policy; client-side commissioners setting supplier requirements; anyone who has to answer "what did the AI actually do here" after the fact.

## 3. When to use it

- Before AI is used on any project material, which is the only point at which most of these decisions are cheap.
- A project involves participant data, client-confidential material, or anything covered by a confidentiality obligation.
- A team is adopting AI in delivery work and has no written position.
- A client, procurement process or contract asks what AI is used, on what data, and with what safeguards.
- A deliverable must carry an AI disclosure, or a proposal must describe the method honestly.
- AI will perform coding, classification or extraction at a scale nobody will read in full.
- A study will be rerun, extended into a further wave, or repeated for comparison.
- A model, version or system changes during a project, or between waves of a tracking study.
- An output is challenged and the process behind it needs reconstructing.
- Consent was obtained before AI processing was contemplated, and a new use is proposed.

## 4. When NOT to use it

- **The question is whether a specific AI output can be trusted.** That is **13.03 AI Output Verification**, which audits an output against its inputs. Governance says an output must be verified and recorded; verification is the act. Do not use a governance record as evidence that an output is correct; it is evidence that someone checked.
- **The question is what participants were told and whether they understood it.** Consent design, participant information, distress protocols and identifiability are **13.05 Research Ethics and Consent Design**. The two meet at consent for AI processing: 13.05 designs it, this skill operates it across the project. Where an AI use is not covered by consent, the answer comes from 13.05, not from a policy exception.
- **The question is a definitive legal or contractual one.** This skill sets out the questions to ask of a system and a contract (retention, training use, sub-processing, transfer, ownership of outputs) and applies the general principles of applicable data protection law. **Requirements differ by jurisdiction and contracts differ by client, and legal advice is not a research judgement.** Where a specific processing operation's lawfulness, or a specific contract's meaning, is at stake, it goes to qualified advice.
- **The question is procurement or security assessment of a system.** Whether a given system meets an organisation's security standard is an information security and vendor assessment matter with its own specialists. This skill governs research use of whatever has been approved.
- **The question is methodological.** Whether AI-produced analysis is sound, whether the design supports the claim, whether the sample carries the finding: **13.01 Research Quality Review**. Governance does not make weak research strong, and a fully compliant AI process can produce a worthless analysis.
- **The policy will not be enforced or resourced.** A governance policy nobody applies is worse than none, because it creates a documented standard the team is visibly failing. Where there is no owner, no sign-off capacity and no willingness to slow a project down, say so and write the shortest possible policy covering only what will actually be done. **A three-rule policy that is followed beats a twelve-page one that is not.**
- **The intent is to construct a record after the fact.** Where AI steps have already been performed unrecorded and someone wants a governance trail written retrospectively, that is not an audit trail. Reconstruct what can be reconstructed, label it as reconstruction with its date, state what is unknown, and change the process going forward.

## 5. Required inputs

**Required.**
- **The project's data inventory**: what material exists, where it came from, and what it contains. Governance decisions are made per dataset, not per project.
- **The consent basis for any participant data**, in its actual wording. What consent covers is a matter of what it says, and it is usually narrower than the team assumes.
- **The client's contractual position on AI, confidentiality and subprocessing**, or an acknowledgement that it has not been established, which is itself a finding.
- **Which steps in the project will use, or have used, AI**, and for what. Without this the governance is generic and unenforceable.
- **A named accountable person.** Every governance arrangement needs someone who owns it. Where there is no name, there is no governance.

**Optional, and what each one adds.**
- **The organisation's approved-systems list and its security position.** Converts "may I use AI for this" into a decision with an answer rather than a discussion.
- **The system's data handling terms**: retention of inputs, whether inputs are used to improve the service, subprocessing, transfer locations, deletion mechanisms. These determine what classification of material may be sent at all.
- **Prior projects' AI step registers.** Establish what the team actually does, which is usually different from what its policy says, and it is the fastest route to a policy people will follow.
- **The model or version identifiers in use.** Necessary for reproducibility and for explaining a rerun that differs.
- **A human-coded comparison sample.** Produces the measured accuracy figure that K4 §7 requires to be disclosed for AI-performed coding.
- **Client or sector AI requirements.** Frequently stricter than the team's own position, and better discovered at proposal stage than at delivery.

## 6. Questions to ask before starting

1. **What is the most sensitive material this project holds, and where is it?** Determines the whole classification. *Default if unanswered:* inventory it before any AI use. Most teams underestimate what is in their free-text fields and their transcripts.
2. **What does the consent actually say, in its own words?** Determines whether participant data may be processed by AI at all. *Default:* if it does not mention it, it does not cover it, and the question routes to 13.05.
3. **What does the client contract say about AI, confidentiality and third-party processing?** Determines what is permitted regardless of what is technically possible. *Default:* if it is silent, ask the client before proceeding rather than after.
4. **Which steps will use AI, and what would happen if each were wrong?** Determines the human check required at each. *Default:* every step producing a specific that reaches a deliverable gets a verification step (13.03).
5. **Will this study be repeated, extended or compared?** Determines whether the version must be pinned and recorded. *Default:* record the version regardless; it costs one line and it is unrecoverable later.
6. **Who signs off, and on what?** Determines the K5 record. *Default:* the person whose name is on the deliverable, and if that is unclear, that is the first governance problem to fix.

## 7. Step-by-step methodology

**Step 1. Classify the material, in one table the team can apply in seconds.** Governance fails when the rule requires thought at the moment of use. Build a classification with a small number of tiers and a single rule for each, and put it where people work rather than in a policy document.

| Tier | What it covers | Rule |
|---|---|---|
| **Restricted** | Direct participant identifiers, special-category data, anything where consent does not cover AI processing, client material under a contractual bar | Not sent to any AI system. No exceptions at project level |
| **Controlled** | De-identified participant data (transcripts, open ends), client-confidential material where the contract permits processing | Approved systems only, under the project's recorded terms, with the AI step registered |
| **Internal** | The team's own working material, analysis outputs, drafts containing no participant or client-confidential content | Approved systems, registered where it touches a deliverable |
| **Open** | Published material, the team's own templates and methods | No restriction |

Two rules make the table work. **Classification attaches to the material, not to the task**, so a transcript does not become Internal because the task is trivial. And **the default for unclassified material is the most restrictive tier**, which is what prevents the classification from being quietly bypassed by anything nobody thought about.

Three things routinely sit in the wrong tier. **Free text**, which contains names, account numbers and third-party details whatever the survey asked for, so it is Controlled at best until screened. **Screening data** from people who were never recruited. And **client-supplied operational data**, which is frequently more commercially sensitive than the research and arrives with no classification at all.

**Step 2. Establish the consent boundary for every participant dataset.** Read the actual consent wording, not the study's description of it. Three outcomes. **Covered:** the consent describes AI or automated processing in terms a participant would recognise as this. **Not covered:** it does not mention it, in which case silence is not permission. **Ambiguous:** it refers to processing by the research team or by third parties in terms that could include it, which is the commonest and the most quietly abused case.

For not-covered and ambiguous data the options are, in order: **re-consent**, which is often practical for panels and ongoing studies and always cleanest; **restrict** the AI use to steps the consent does cover; **aggregate** to a level where the material is no longer personal data and the question does not arise; or **do not use AI on it**. What is not an option is proceeding because the data is de-identified, since the participant's interest is in what reads their words, not only in whether it can name them (13.05). Route the judgement to a named human per K5 §2.4 and record it. *Correct result:* a per-dataset consent position with a stated basis and a decision.

**Step 3. Register every AI step before it runs.** One line per use, and the register is the single artefact that makes everything downstream possible. Fields: the step, the material and its tier, the system and version, what it was asked to do, the date, who ran it, what human check was applied, by whom, and what the check found.

This takes about a minute per step and it is what turns "we used AI in the analysis" into a reconstructable account. Two things make it survive contact with a busy project. **It is filled at the time**, because a register completed at the end of a project is a reconstruction wearing a record's clothes. And **it lives with the project files**, not in a governance system nobody opens.

**Step 4. Set the human check for each registered step, and record who performed it.** Governance's contribution here is not to specify verification, which is 13.03's job, but to fix **who** and **at what point**, and to make the record exist. Use K5 §6's calibration: formatting gets a spot check, descriptive analysis gets figures verified against source, coding gets a human-approved code frame before scale plus an audited sample after, quote selection gets every quote verified, insight development gets human judgement on each one, and recommendations get full review and sign-off without exception.

**A sign-off record needs four things: what was signed off, by whom (a name, not a role), on what date, and on what basis.** "Reviewed by the research team" is not a sign-off. The basis matters most and is omitted most: "verified all quotes against transcript, spot-checked 20% of coded segments" is a record; "checked" is not.

**Step 5. Govern AI-performed coding, classification and extraction specifically, because it is where scale outruns checking.** Four requirements, and they are the practical core of this step.

**A human approves the code frame or classification scheme before it is applied at scale.** After application, changing it means recoding everything, which nobody does, so the frame becomes fixed by accident at the moment of first use.

**Accuracy is measured, not assumed.** A human independently codes a random sample and agreement is computed **per code, not overall**, because overall agreement hides the failure that matters: rare and boundary categories collapse while the aggregate stays high. Stratify the sample to over-represent rare codes.

**The measured accuracy is disclosed**, per K4 §7, alongside the finding it supports rather than in a methods appendix. And the figure disclosed is the one a reader should rely on, which is the worst per-code rate among codes that carry a finding, not the flattering average.

**Where accuracy cannot be measured, the output is labelled unvalidated** and may not carry a headline. This is the rule most often broken, usually because the sample coding was planned and never scheduled.

**Step 6. Pin the version, and be honest about reproducibility.** Record the system and version identifier for every AI step at the time it runs. Then hold two facts about what that does and does not buy.

**An analysis rerun on a different model version may not reproduce**, and the difference can be material rather than cosmetic: different themes, different boundary classifications, different emphasis. This is not a defect to be apologised for; it is a property of the method, and it belongs in the disclosure.

**A rerun is not a replication.** Running the same instruction twice on the same system and getting a similar answer tests stability, not correctness, and two agreeing AI outputs are not two pieces of evidence (K2 §4.4). Only comparison against source, or against independent human work, tests correctness.

The operational rules follow. **Pin one version for a project** where the work is comparative, and record it. **Where a version changes mid-project, do not silently continue:** either complete the affected stream on the pinned version, or re-run the earlier steps on the new one and re-verify, and record which was done. **In tracking studies, treat a version change like a question rewording:** assess whether it introduces a step in the series, and disclose it at the point of comparison. **And never present a rerun's agreement as corroboration.**

**Step 7. Establish the confidentiality and intellectual property position, by asking three questions of every system.** These are practical and answerable, and a team that cannot answer them for a system it is using has a governance gap rather than a policy debate.

**What happens to the input?** Is it retained, for how long, is it used to improve the service, can it be deleted on request, and where is it processed. This determines the highest tier of material that may be sent.

**Who else can see it?** Subprocessors, support access, and any human review of submitted content. Client confidentiality obligations frequently prohibit disclosure to third parties in terms that include this, and the obligation is not suspended by the fact that the third party is a system.

**Who owns the output, and what is it derived from?** The team's position on ownership of AI-assisted deliverables, and whether outputs may reproduce material the team has no right to publish.

Then check the **client's own position**, which frequently differs from the supplier's and is usually stricter. Establish it at proposal stage, in writing. Discovering at delivery that a client prohibits AI processing of their data is a contractual problem and a relationship problem, and it is entirely preventable by one question at the start.

**Step 8. Write the disclosure, and make it say something.** The standing obligation is not to present AI-assisted work as though it were not (K5 §7). The practical requirement is that a disclosure be specific enough to be useful. **A good disclosure has three parts: what the AI did, what a human verified, and by whom.**

Weak: "AI tools were used in the preparation of this report."

Useful: "Interview transcripts were transcribed automatically and coded by an AI system against a code frame approved by [name] before application. [Name] independently coded a stratified 15% sample; per-code agreement ranged from 0.71 to 0.94, and the two codes carrying headline findings agreed at 0.89 and 0.91. All quotes were verified word for word against transcript by [name]. Analysis, interpretation and recommendations are the researchers' own. Findings were reviewed and signed off by [name] on [date]."

Disclosure runs to three audiences and each needs a different form. **Clients** get the detail above, in the report's method section and in the proposal that preceded it. **Participants** are told at consent, in plain language, what will process their words and what a human does with the result (13.05). **Internal readers** get the AI step register. Where a client asks for no AI disclosure in a public-facing document, that is a decision for a named human and it does not change the obligation to the participants or the internal record.

**Step 9. Write the policy, and keep it short enough to be followed.** A team policy needs seven sections and no more: the **classification table** from step 1; the **approved systems list** with the tier each may receive; the **register requirement**, saying what is recorded and where; the **human check standard**, mapping task types to required checks and sign-off; the **disclosure standard**, with the paragraph templates pre-written; the **consent rule**, saying what to do when consent does not cover a use; and the **exception route**, naming who may authorise a departure and requiring that it be recorded.

Two design rules make it real. **Every rule names who decides**, because a policy without an owner is a preference. And **the exception route exists deliberately**, because a policy with no legitimate way to depart from it will be departed from illegitimately, and the departures will be invisible. Review the policy on a fixed date rather than when something goes wrong, since the systems, the client requirements and the team's actual practice all move.

**Step 10. Close the loop: audit the register against the deliverable.** Before delivery, check that every AI step that touched the output is in the register, that each has a recorded human check with a name, that the disclosure matches what was actually done, and that any measured accuracy figure in the disclosure matches the sample that produced it. **The commonest governance failure at this point is a disclosure that describes the intended process rather than the executed one**, which is a false statement about method however well-intentioned. Where the register and the deliverable disagree, the register is corrected and the disclosure is rewritten to match reality.

## 8. Analytical framework

The governance chain, applied per AI step rather than per project:

    Material → Tier → Consent basis → Permitted? → System and version
        → Instruction → Output → Human check (who, what, when)
            → Record → Disclosure

Each arrow is a gate that can stop the step, and the first three are where a step is stopped cheaply. A step that reaches "output" before anyone asks about tier or consent cannot be un-run.

And the three questions that any governance arrangement has to be able to answer, at any point, about any AI-performed step:

1. **Was it permitted?** Classification, consent, contract.
2. **Was it checked?** By a named human, in a specified way, with a recorded finding.
3. **Can it be reconstructed?** Inputs, instruction, system, version, date, output.

**Applying it.** These three are the test of whether a policy is real. A team that can answer all three for a step performed four months ago has governance. A team that can answer the first two but not the third has compliance without accountability, which fails at exactly the moment it is needed, which is when something is challenged.

**The proportionality rule.** Governance overhead should scale with consequence, not be uniform. A researcher summarising published material needs the register line and nothing else. AI coding participant transcripts that will carry a client's headline finding needs the full chain. A policy that applies the same weight to both will be abandoned for the first case and then, by habit, for the second.

## 9. Output format

**A. Data classification table.** Per step 1, on one page, held where the team works.

**B. AI step register.**

| Ref | Project step | Material and tier | System and version | Instruction summary | Date | Run by | Human check applied | Checked by | Finding |
|---|---|---|---|---|---|---|---|---|---|

**C. Consent position by dataset.**

| Dataset | Consent wording basis | Covers AI processing? | Decision | Authorised by | Date |
|---|---|---|---|---|---|

**D. Sign-off record.**

| What was signed off | By whom (name) | Date | Basis (what was actually checked) |
|---|---|---|---|

**E. Accuracy record**, for any AI coding, classification or extraction: sample size and how drawn, per-code agreement, the worst rate among codes carrying findings, and the disclosure text that follows from it.

**F. Disclosure text**, in three versions: client-facing method paragraph, participant-facing consent language, and internal record. Drafted from the register rather than from intention.

**G. The governance policy**, seven sections per step 9, with the classification table, approved systems, register requirement, human check standard, disclosure standard, consent rule and exception route.

**H. Exception log.** Every departure from the policy, with who authorised it, why, and what compensating check was applied. An empty exception log in a busy team usually means departures are going unrecorded rather than not happening.

**When the evidence is thin.** Where an AI step was performed without a record, the entry is completed as a reconstruction, marked as such, dated, and accompanied by what is unknown (typically the version and the exact instruction). **Do not write a reconstruction into the register as though it were contemporaneous.** Where a system's data handling terms cannot be established, the material tier that may be sent to it is Open, and that constraint is recorded rather than assumed away. Where accuracy was not measured, the disclosure says so in those words, and the affected output does not carry a headline.

## 10. Quality checks

Run before delivery, and again at policy review. K4 §8 runs anyway.

1. Does every AI step that touched the deliverable appear in the register, filled at the time rather than reconstructed?
2. Is every material tier assigned to the material rather than to the convenience of the task, with unclassified material defaulting to the most restrictive tier?
3. Was the consent wording read in its own words, and is the position recorded per dataset rather than per project?
4. Where consent did not cover AI processing, was the judgement routed to a named human and recorded?
5. Does every sign-off name a person, a date, and what was actually checked, rather than a role and the word "reviewed"?
6. Is the system and version recorded for every step, so a later difference can be explained?
7. For any AI coding or classification, was the frame human-approved before scale, and was agreement measured per code rather than overall?
8. Does the disclosed accuracy figure reflect the codes carrying findings rather than the flattering average?
9. Does the disclosure describe the process that was executed, not the process that was planned?
10. Does the disclosure name what the AI did, what a human verified, and by whom?
11. Were participants told what would process their words, in language they would recognise as this?
12. Has the client's contractual position on AI and third-party processing been established in writing?
13. Is there an exception route, is it being used, and are exceptions recorded with compensating checks?
14. Could a colleague who was not on the project reconstruct any single AI step from the register alone?

## 11. Common failure modes

| Failure | How to recognise it | How to prevent it |
|---|---|---|
| **The aspirational policy** | Principles, values and commitments, no rule anyone applies at a keyboard | Seven sections, a classification table, and a named decider per rule |
| **Retrospective register** | Completed in one sitting near delivery, all entries the same shape | Fill at the time. Reconstructions are marked as reconstructions |
| **Classification by task** | A transcript treated as low-risk because the request was trivial | Classification attaches to the material. Unclassified defaults to the most restrictive tier |
| **Free text underclassified** | Open ends processed as though they contained only opinions | Screen before processing. Free text is Controlled at best until it is screened |
| **Silence read as permission** | Consent does not mention AI, and processing proceeds because the data is de-identified | Silence is not permission. Re-consent, restrict, aggregate, or do not use |
| **Sign-off as a role** | "Reviewed by the insight team" | A name, a date, and what was actually checked |
| **Unmeasured accuracy** | AI coding at scale, no sample coded by a human, a confident prevalence table | Measure per code, disclose the worst rate among codes carrying findings, or label unvalidated |
| **Overall agreement disclosed** | A high aggregate figure hiding a rare code at 0.4 that carries a finding | Per-code agreement, stratified sample over-representing rare codes |
| **Version amnesia** | A rerun produces different themes and nobody can say what changed | Record system and version per step. Pin for comparative work |
| **Rerun as replication** | Two AI passes agreeing, presented as corroboration | A rerun tests stability, not correctness. Corroboration requires an independent source |
| **Disclosure of intent** | The method section describes the planned process; the register shows a different one | Step 10. Write disclosure from the register |
| **The unfindable contract position** | AI use is discovered at delivery to be prohibited by the client's terms | Establish it in writing at proposal stage |
| **No exception route** | A clean policy, and everybody knows people work around it | Build a route, name the authoriser, log the exceptions |
| **AI: generic policy text** | A policy that would fit any organisation and names no system, tier, person or check | Every rule names material, a system class, a decider and a record |
| **AI: asserting a legal or contractual requirement** | A confident statement about what a jurisdiction or a contract requires | State principles generically. Specific legal and contractual questions go to qualified advice |

## 12. AI guardrails

Universal prohibitions are inherited from K4. K4 §7's requirement to disclose AI-performed coding, classification and extraction, and K5 §7's standing disclosure obligation, are the two this skill operationalises.

1. **Never present AI-assisted work as though it were not**, and never accept an instruction to omit disclosure from a deliverable that carries a method statement. Where a client asks for it, the decision is a named human's and the obligation to participants and to the internal record is unchanged.
2. **Never send material to an AI system without a classification decision**, and never classify by how routine the task feels. Unclassified material takes the most restrictive tier.
3. **Never treat de-identification as consent.** The participant's interest includes what reads their words, not only whether it can name them.
4. **Never write a disclosure that describes an intended process rather than the executed one.** The disclosure is drafted from the register.
5. **Never disclose an overall agreement rate where per-code rates differ materially**, and never disclose an accuracy figure that is not the one a reader should rely on.
6. **Never present an AI output rerun, on the same or a different version, as corroboration of the first.** Agreement between two passes tests stability, not correctness.
7. **Never continue a comparative or tracking analysis across a model version change without recording it and assessing whether it introduced a step in the series.**
8. **Never complete a register entry from memory without marking it as a reconstruction**, with its date and what is unknown.
9. **Never assert a specific legal or contractual requirement.** State the questions to ask of a system and a contract, and route specific determinations to qualified advice.
10. **Never let an approved-systems list substitute for a per-step decision.** A system being approved for Controlled material does not make this material Controlled.

## 13. Best-practice principles

1. **Governance is what makes work reconstructable, not what makes it permitted.** Permission is the easy half. The value shows up months later when a finding is challenged and somebody can answer how it was produced.
2. **A three-rule policy that is followed beats a twelve-page one that is not.** Length is the enemy of adoption, and an unfollowed policy documents a standard the team is visibly failing.
3. **Make the rule applicable at the keyboard.** If a researcher has to think, or look something up, or ask, at the moment of use, the rule will be skipped. One table, where they work.
4. **Classification attaches to the material.** The single most common breach is a sensitive input treated casually because the task was casual.
5. **Silence in a consent is not permission.** This is the rule most often bent, and the bending is always justified by de-identification, which answers a different question.
6. **Record the version. It costs one line and it is unrecoverable afterwards.** The moment you need it is the moment you cannot get it.
7. **A rerun is not a replication.** Two AI passes agreeing is one piece of evidence produced twice, and treating it as two is a traceability failure as well as a governance one.
8. **Measure accuracy per code, and disclose the number that matters.** The aggregate is almost always flattering, and the codes carrying findings are almost always not the common ones.
9. **A sign-off is a name, a date, and what was checked.** Anything else is a formality that will not survive being asked about.
10. **Build the exception route deliberately.** A policy with no legitimate departure will be departed from invisibly, and invisible departures are the thing governance exists to prevent.
11. **Establish the client's position in writing at proposal stage.** It is one question, and asking it late converts a simple constraint into a contractual problem.
12. **Scale the overhead to the consequence.** Uniform governance is abandoned first for trivial cases and then, by habit, for the ones that mattered.
13. **Write the disclosure you would be comfortable reading as the client.** That test settles most arguments about how much detail is enough, and it settles them in the right direction.

## 14. Worked example

*Fictional scenario, used for illustration only. The organisation, arrangements and figures below are invented for the purpose of demonstrating method.*

**INPUT.** A software company's in-house user research team of six runs continuous discovery: roughly 40 customer interviews a quarter, a quarterly satisfaction survey with substantial free text, and usability sessions recorded on video. AI use has grown organically: automatic transcription, summarisation of sessions, and coding of survey open ends. There is no policy. The trigger is an enterprise customer's procurement team asking, in a contract renewal, what AI processes their staff's research data and what safeguards apply. The team has no answer.

**PROCESS.**

*Step 1, classification, and the first surprise.* The inventory finds four material types and two of them are in the wrong place. Video recordings of usability sessions show faces, screens containing real customer account data, and in two cases visible internal documents belonging to the customer's own organisation. These had been treated as ordinary session material. They are classified Restricted. Survey free text, treated as opinion data, is found on inspection to contain names, email addresses and in one case a partial payment reference; it is classified Controlled and gains a mandatory screening pass before any processing. Transcripts, de-identified at the point of transcription, are Controlled. The team's own analysis notes are Internal.

*Step 2, consent, and the judgement call.* The interview consent, unchanged for three years, says data may be "processed by the research team and its service providers for analysis". This is the ambiguous case. It could be read to include automated processing, and a participant would not have understood it that way. The team's instinct is to rely on it, since re-consenting three years of participants is impractical. The assessment separates two things: the existing archive and future collection. Future collection gets new consent wording naming automated transcription and coding in plain language, drafted with 13.05. The archive is restricted to the steps the old consent plainly covers, which excludes sending transcripts to a general-purpose system, and where analysis of the archive is needed it is done on aggregated material. Routed to and authorised by the research lead, recorded, per K5 §2.4.

*Step 3 and 4, register and checks.* The register is set up as a shared sheet with the project files, one line per use. Reviewing the last quarter retrospectively produces 31 entries, all marked as reconstructions with their unknowns (nine have no recoverable version identifier). The exercise is the argument for the register: nobody can say which system version coded the previous quarter's open ends.

*Step 5, coding accuracy, and the finding that changes a deliverable.* The most recent quarterly survey's 1,900 open ends were AI-coded to a 14-code frame and the results are already in a product decision paper. A validation is run for the first time: a researcher independently codes a stratified 200-response sample. Overall agreement is 0.86, which sounds reassuring. Per code, it ranges from 0.94 on the two largest codes to 0.52 on "pricing concern", which has 61 mentions and is the code carrying the paper's headline about price sensitivity. The disclosed figure becomes 0.52 for that finding, not 0.86 for the analysis, and the headline is downgraded to a directional observation pending a human recode of that code's segments. This is the moment the team stops treating governance as paperwork.

*Step 6, versions.* The satisfaction survey is quarterly and comparative, so the version is pinned for the series and recorded. A provider version change two months in is handled explicitly: the affected quarter is completed on the pinned version, the change is logged, and the next wave's comparison carries a note that the coding system version changed between waves, treated in the same way as a question rewording.

*Step 7, contract and confidentiality.* The three questions are asked of each system in use. One retains inputs for a period the team had not known about and permits human review of submitted content; it is removed from the Controlled tier and restricted to Open material. The enterprise customer's contract turns out to prohibit disclosure of their material to third-party processors without prior written approval, which the video recordings had breached in principle. Disclosed to the customer proactively, with the classification change and the deletion, rather than discovered by them.

*Steps 8 to 10, disclosure and policy.* A three-part disclosure is written from the register. The policy comes to two pages: classification table, approved systems with tiers, register requirement, check standard mapped to task type, disclosure templates, the consent rule, and an exception route naming the research lead as authoriser with a required log entry. Review date set at six months. The register is audited against the quarter's decision paper, and the method paragraph is rewritten because the original described the intended process rather than the executed one.

**OUTPUT.** A one-page classification table with two material types reclassified; a consent position separating archive from future collection, with new wording for the latter; a live AI step register; a per-code accuracy record that changed a product decision paper's headline; a pinned version for the comparative series with a logged mid-series change; one system moved out of the Controlled tier on its retention terms; a proactive disclosure to an enterprise customer of a contractual breach; a two-page policy with a named authoriser and an exception route; and an answer to the procurement question that is specific, true, and evidenced. `RESEARCHER SIGN-OFF REQUIRED` per K5 §3.1 on the reworked pricing finding before the decision paper is reissued.

## 15. Advanced usage

**Governing across an agency and its clients.** Where a team works for multiple clients with different AI positions, the classification table gains a client dimension and the approved-systems list becomes per-client. This is more overhead and it is unavoidable: the alternative is applying the strictest client's rules to everything, which teams say they will do and do not. Establish each client's position at proposal stage and hold it in one place where anyone can check it in seconds.

**Governing an AI-moderated study.** Where the AI is talking to participants rather than processing their data afterwards, governance and ethics merge and the requirements tighten: system disclosure to the participant, an always-visible exit, a designed escalation trigger and a named human reachable during fieldwork (13.05, 03.02). The register carries the moderation configuration and any mid-field change to it, because a change to the moderator during fieldwork is a change to the instrument.

**Reproducibility over a long series.** Trackers and rolling programmes run for years across multiple version changes, and a series where the processing changed silently three times is not comparable however consistent the questionnaire was. Keep a version history alongside the questionnaire version history, treat each change as a potential step in the series, and where a change is material, run a bridge: process one wave both ways and report the difference. That is the only way to know whether a movement is the market or the model.

**When the standard approach does not fit.** Where a team genuinely cannot meet the standard (no capacity to validate coding, no approved systems, no named owner), do not write the full policy. Write the honest minimum: what may not be sent anywhere, what must be disclosed, and who to ask. Then restrict AI use to the tiers the team can actually govern. **A narrow scope properly governed is defensible; a broad scope nominally governed is not**, and the second is what most first attempts produce.

## 16. Skill chain

**Recommended previous skills:**
- **13.05 Research Ethics and Consent Design.** Hands over the consent position, including whether AI processing was designed into the consent, which this skill operates across the project.
- **01.08 Research Proposal and Scope Development.** Hands over the proposal, which is where the client's AI position and the disclosure commitment should be established in writing.
- **01.05 Research Plan.** Hands over the project steps, from which the AI step register is built before work begins.

**Recommended next skills:**
- **13.03 AI Output Verification.** Performs the human check this skill requires and records, at the coverage the check standard specifies.
- **12.03 Research Report Compilation.** Carries the disclosure text and the accuracy figures into the deliverable at the point they bite rather than into an appendix.
- **13.01 Research Quality Review.** Reads the register and the sign-off record as inputs when assessing whether the study can support its claims.
- **14.03 Research Repository and Knowledge Curation.** Takes the register, consent position and version record into the archive, where they become the answer to reuse questions later.

**Runs well alongside:**
- **04.02 Data Cleaning** and **07.02 Open-Ended Response Coding**, which are the steps most often AI-performed and most often unrecorded.
- **03.02 AI-Moderated Interview Design**, where governance and participant-facing ethics meet.
- **K4 §7**, which requires the AI coding disclosure this skill produces, and **K5 §6 and §7**, which supply the review-intensity calibration and the standing disclosure obligation it operationalises.

---
A Yazi Supplied Skill and resource.
