---
name: thesis-examination-and-marking
description: >
  Assesses a thesis or dissertation against the stated criteria and the standard
  of the degree, with located evidence behind every judgement, and produces a
  defensible examiner's report and a bounded corrections list. Use for "examine
  this thesis", "write my examiner's report", "assess this dissertation against
  the criteria", "is this flaw fatal or correctable", "how do I mark this thesis",
  "the other examiner and I disagree", "draft a corrections list", "how do I judge
  an originality claim".
category: 15 Academic University Research
ref: "15.21"
tier: 3
inherits: [K2, K3, K4, K5]
---

# Thesis Examination and Marking

## 1. One-line description

A disciplined method for examining a thesis against the stated criteria and the standard of the degree rather than against the thesis the examiner would have written, evidencing every judgement with a location, separating fatal flaws from correctable ones and from matters of taste, and producing a report and corrections list that are specific, bounded and defensible.

## 2. What this skill is used for

**Read this before you start.** Whether an AI system may be used at any point in examining or marking a student's work is governed by your institution's policy, and by the appointing institution's policy where you are an external examiner. It also involves processing an identified person's personal data and unpublished, unprotected intellectual work, frequently under an examiner's confidentiality undertaking. Both questions are settled before the thesis is opened. Section 5 makes the policy check a required input; Section 12 sets the operating rules. Where the position is unknown, the conservative default is that the thesis is not uploaded to any external system, and that AI assistance is confined to structuring and stress-testing the examiner's own reading.

**The research problem it solves.** Examination fails in characteristic ways that are all failures of discipline rather than of expertise. The examiner assesses against the thesis they would have written, so a legitimate methodological choice becomes a criticism. They calibrate to the best thesis they have read rather than to the standard of the degree, and a sound thesis is treated as a weak one. Criticisms are stated as general impressions ("the literature review is superficial") that the candidate cannot answer and a board cannot audit. Corrections lists are unbounded and quietly demand a different thesis. And the hardest judgement of all, whether a problem is fatal, correctable or simply not to the examiner's taste, is made implicitly and defended after the fact. This skill installs the criteria before the reading, requires a located page reference behind every judgement, and forces the fatal-correctable-taste classification to be made explicitly.

**Where it sits in the research lifecycle.** At the end, at examination. It also runs at any formal assessment point with criteria and an outcome: an upgrade or transfer viva, a progression review, a taught-dissertation marking exercise, a second-marking or moderation pass.

**Typical use cases.**
- Examining a doctoral or masters thesis as internal or external examiner.
- Marking a taught dissertation against a rubric, and moderating across a cohort.
- Preparing for a viva or oral defence: the questions, and what turns on the answers.
- Deciding whether an identified weakness is fatal, correctable within a corrections period, or a matter of preference.
- Writing an examiner's report that will be read by the candidate, the supervisor and a board.
- Constructing a corrections list that is actionable, bounded and achievable in the time the regulations allow.
- Resolving a disagreement with a co-examiner before a joint recommendation is made.

**Who uses it.** Internal and external examiners; first and second markers on taught programmes; chairs and independent chairs of examination panels; moderators and external examiners of programmes; supervisors preparing a candidate by reading as an examiner would, though that use case is properly **15.19**.

## 3. When to use it

- You have been appointed to examine a thesis and have the criteria and regulations in front of you.
- You are marking a dissertation against a rubric and need consistency across a cohort.
- You have read a thesis, formed an impression, and need to convert it into judgements you can evidence.
- You have found a serious problem and must decide whether it is fatal to the degree or correctable.
- You need to write a report that a candidate can act on and a board can rely on.
- You are second-marking, moderating, or resolving a divergence between two markers.
- A viva is scheduled and the questions have to be built from the thesis rather than from general practice.
- The thesis is competent and adds little, or original and technically flawed, and you need to work out which outcome each of those warrants under the regulations.

## 4. When NOT to use it

- **The purpose is development, not assessment.** A supervisor commenting on a draft is doing a different job with a different discipline: prioritised, principle-led, teaching rather than judging. Go to **15.20 Supervision Feedback and Development**. Examination-register feedback on a mid-programme draft is discouraging and, worse, misleading, because it applies a final standard early.
- **You have not read the thesis in full.** Examination is a whole-document judgement. A report or recommendation built on a summary, a sample of chapters, or a description of the work is not an examination, and no method makes it one. See §12.1 and §12.3.
- **You have a conflict of interest.** Prior involvement in the work, a co-authorship with the candidate or supervisor, a personal relationship, a competing research programme, or a financial interest. Declare it to the institution and let the institution decide. This is a regulatory question, not a judgement to make privately.
- **The concern is academic integrity.** Suspected plagiarism, undeclared AI use, data fabrication or falsification. These have institutional processes with defined thresholds, evidence standards and rights of response. An examiner records what they observed and refers it; this skill does not detect, assess or allege misconduct.
- **The criteria or the regulations are not available.** Assessment without the criteria is assessment against the examiner's private standard, which is exactly the failure the method exists to prevent. Obtain them, or stop. The regulations also define the available outcomes, and an examiner who recommends an outcome the regulations do not contain has created work for everyone.
- **The thesis is outside your competence.** An adjacent field is not the field. Where the methodology, the theory or the technical content sits outside what you can assess, say so to the institution and ask for a third examiner or a specialist. A confident examination of work you cannot evaluate is worse than a declined appointment.
- **You are auditing the study rather than examining the thesis.** Where the question is whether a piece of research is methodologically sound, independent of whether the document meets a degree standard, that is **15.22 Research Rigour Audit**. The two overlap in method and differ in purpose: examination judges a candidate against a degree; an audit judges a study against methodological standards and issues no outcome.
- **The material is a paper under peer review, not a thesis.** Different standard, different criteria, different obligations to the author. See **15.26 Peer Review Response** for the candidate side; peer review itself is not this skill.

## 5. Required inputs

**Required.**
- **The full thesis, read in full by the examiner.**
- **The stated assessment criteria and the institution's regulations**, including the available outcomes and their definitions, the corrections period, and the report format required.
- **The degree being examined and the standard it sets.** A professional doctorate, a research doctorate, a masters by research and a taught masters dissertation set materially different bars, and the bars differ by country and institution. Never infer this.
- **Your institution's, and the appointing institution's, position on AI use in examination and on candidate data.** If unknown, resolve before starting.
- **Confirmation that no conflict of interest exists**, or a declaration made.

**Optional, and what each one adds.**
- **The examination model.** Whether there is a viva or oral defence, whether it is public, whether examiners report independently before meeting, whether an independent chair is present, and whether the report is written before or after the defence. This changes the report's function entirely: a pre-viva report is a set of questions and provisional judgements; a post-viva report is a recommendation. Never assume a model.
- **Your co-examiner's report or provisional view, where the regulations permit exchange.** Enables genuine calibration. Where the regulations require independent reports first, do not seek it.
- **Publications arising from the thesis.** Evidence that parts have survived external review, which is relevant to the contribution judgement but does not substitute for it, and which is weighted differently across disciplines.
- **Previous theses you have examined at this level, or exemplars at the boundary.** The single best defence against calibrating to the best thesis you have ever read.
- **The candidate's own statement of contribution**, where the format includes one. Gives you the claim to test rather than one you have inferred.
- **Ethics approval documentation.** Where human participants, animals or sensitive data are involved, its absence is a finding, and in most regulations a serious one.

## 6. Questions to ask before starting

1. **What outcomes do the regulations allow, and how are they defined?** Everything downstream is a choice among these. The categories, their names and their corrections periods vary widely by country and institution. Default: obtain them. Do not proceed on an assumed set of outcomes.
2. **What is the standard of this specific degree at this institution?** Not the standard of a doctorate in general, and certainly not the standard of the best thesis in your field. Default: read the regulations' own words for the degree, quote them into your working notes, and calibrate to them explicitly.
3. **Is there a viva, and does my report come before or after it?** Determines whether the document is a set of provisional judgements and questions or a final recommendation. Default: ask. If unavailable, write provisional judgements and mark them as such.
4. **Which criteria are threshold and which are quality?** A threshold criterion, if unmet, is fatal by definition; a quality criterion affects the level of the judgement. Conflating them is how examiners talk themselves into a fail on a matter of degree, or a pass on a missing requirement. Default: classify each criterion yourself and record the classification.
5. **What is the discipline's convention here, and is it mine?** Voice, structure, whether a separate methodology chapter is expected, how theory is used, how much of the literature belongs in chapter two, whether a monograph or a paper-based thesis is normal, and what counts as a contribution. Default: where the candidate's convention differs from yours and is legitimate in their field, it is not a criticism. Say so in your notes so you do not drift.
6. **What is the candidate claiming as the contribution, in their own words?** You assess the claim they made, not one you constructed. Default: locate it, quote it into your notes, and if you cannot find it stated anywhere, that absence is itself a finding.
7. **Have I examined at this level recently, and where does this sit against those?** Default: name two reference points before you start reading, not after you have formed a view.

## 7. Step-by-step methodology

**Step 1. Settle the regulatory, conflict and AI position before the thesis is opened.**
Confirm your appointment, the regulations, the available outcomes and the report format. Declare any conflict. Settle what may and may not be done with an AI system on the candidate's unpublished work under both institutions' policies, and record the decision. Examination material is normally confidential to the examination; treating it otherwise is a breach independent of any research-quality question. *Correct result: a one-paragraph note, written before reading, stating the outcomes available, the report format, the conflict position and the AI and data position.*

**Step 2. Write the standard down before you read, in the regulations' own words.**
Copy the criteria into a working document, each on its own line, and mark each as threshold or quality (Question 4). Then write one sentence, in your own words, describing what a bare pass at this degree looks like. This sentence is the calibration anchor, and writing it before reading is what prevents the commonest calibration failure: assessing against the best thesis you have read, which is not the standard and was never the standard. A thesis is not weak because it is not excellent. *Correct result: a criteria list with threshold and quality marked, and a written bare-pass description you would be willing to show a board.*

**Step 3. First pass: read the whole thesis for the argument, not for the detail.**
Read straight through. Do not verify a citation, do not check a statistic, do not annotate style. The single question of the first pass is: what does this thesis claim, and does the document as a whole get there? Keep only a running note of where the argument moves, where it stalls, and where you lost it. At the end, write, in three sentences and without looking back: the research question, the contribution claimed, and how the evidence is supposed to support it. A thesis whose spine you cannot state after a full read has a structural problem, and that finding will not be visible from any amount of detailed second-pass work. *Correct result: a three-sentence spine, or an explicit note of where the spine broke down, with chapter locations.*

**Step 4. Second pass: assess against each criterion, and evidence every judgement.**
Now work criterion by criterion rather than chapter by chapter, because chapter-by-chapter reading produces a chronicle and criterion-by-criterion produces an assessment. For each criterion, record the judgement, and behind it at least one located instance: page or section, and what specifically is wrong or right there. The rule is absolute and it is the operational core of this skill: **no judgement enters the report without a location.** "The literature review is superficial" is not a finding a candidate can answer or a board can audit. "The review covers the three main positions on this question at pp. 31 to 44 but does not engage the methodological objection raised in the field to the dominant position, which matters because the study adopts that position's design without addressing the objection" is. Where you find yourself unable to locate an instance for an impression you hold, the impression is not yet a finding, and it either becomes one on a further look or it is dropped. *Correct result: a criterion-by-criterion table where every cell carries at least one page reference (K2 §4.3).*

**Step 5. Test the contribution claim against what is demonstrated.**
Take the candidate's own statement of the contribution and treat it as a claim with evidence behind it. Three questions in order. Is the claim clearly stated, or must you construct it? Is it original in the sense the degree requires, which varies: a new finding, a new method, a new application to a new context, a new synthesis, a new theoretical account, or a substantial body of new evidence. And is it *demonstrated by this thesis*, as opposed to asserted in the introduction and conclusion. The commonest gap is between the claim in chapter one and the evidence in chapters four and five. Assess originality by testing the claim, not by searching your own memory of the field for prior work: a recollection that something similar exists is not evidence (K4 §2.4). Where you believe a claim is anticipated, locate that work and cite it, or do not make the criticism. *Correct result: a written verdict with the claim quoted, the type of originality identified, and the pages that do or do not demonstrate it.*

**Step 6. Assess the methodology on justification and execution, not on preference.**
Two questions, and only two. Was the choice justified against the research question and the plausible alternatives, at the point in the thesis where it was made? Was it executed competently and reported completely enough to be evaluated? A design you would not have chosen, competently justified and correctly executed, is a pass on this criterion, and saying otherwise is the commonest way an examiner substitutes their own thesis for the candidate's. Reserve criticism for four things: a design that cannot answer the stated question, an execution error that undermines the results, reporting too incomplete to evaluate, and a conclusion that outruns what the design licenses (K4 §3.2). Where you would have done it differently and the candidate's approach is defensible, note it as a viva discussion point. *Correct result: a methodology verdict framed as justified or not, executed or not, with genuine flaws distinguished explicitly from divergences from your preference.*

**Step 7. Classify every concern: fatal, correctable, or taste.**
This is the hardest judgement in examination and it must be made explicitly, in a table, before any recommendation forms. Apply these tests.

| Class | Test | Consequence |
|---|---|---|
| **Fatal** | The concern goes to a threshold criterion, and no amount of rewriting fixes it without new research. The data cannot support the claim; the design cannot answer the question; a threshold requirement is absent. | Determines the outcome. |
| **Correctable** | The thesis contains what is needed, and the problem is in its presentation, completeness, framing or analysis of existing material. Fixable within the corrections period the regulations allow. | Goes on the corrections list. |
| **Taste** | A legitimate alternative choice, a structural preference, a stylistic view, a body of literature you would have included. Defensible either way. | Does not appear as a criticism. May appear as a viva discussion or a suggestion for publication. |

Two discipline points. First, "fixable within the corrections period" is doing real work: a correction requiring six months of new data collection is not a correction, whatever it is called, and a corrections list that quietly requires a different thesis is a fail delivered dishonestly. Second, the taste category is under-used, and every item you move into it makes the report stronger, because a report free of preference-dressed-as-criticism is much harder to argue with. *Correct result: every concern classified with a written reason, and the fatal items identified by criterion.*

**Step 8. Handle the two hard cases directly.**
The competent-but-unoriginal thesis and the original-but-flawed thesis are different problems with different outcomes, and they are frequently confused.

*Competent but unoriginal.* Everything done properly and nothing new established. This is a contribution problem, and whether it is fatal depends entirely on how the degree defines contribution, which is why Step 2 matters. In many research doctorates the contribution criterion is a threshold and a thesis failing it fails regardless of technical competence; in many taught masters dissertations it is a quality criterion and the same thesis passes modestly. Do not import one degree's answer into another's.

*Original but flawed.* Something genuinely new, executed with a real problem in it. The question is whether the flaw undermines the specific claim the originality rests on. A flaw elsewhere is correctable; a flaw in the analysis supporting the novel claim is fatal to that claim, whatever its interest.

*Correct result: where either pattern is present, a named paragraph saying which it is, which criterion it engages, and why the outcome follows.*

**Step 9. Reach the recommendation, in the regulations' own categories.**
Work from the fatal list. If any fatal concern engages a threshold criterion, the outcome follows from the regulations, and you are not entitled to soften it out of sympathy. If none does, the outcome is determined by the volume and nature of the correctable items measured against the corrections period. Write the recommendation as a sentence naming the regulatory category, and beneath it a short chain: criterion → concern → classification → outcome. Then apply the reversal test: if the outcome were challenged, would this chain hold? `RESEARCHER SIGN-OFF REQUIRED` (K5 §2.5, §2.8). The recommendation belongs to the named examiner, is made after the viva where the model includes one, and this skill does not make it. *Correct result: a recommendation in the regulations' language, with a traceable chain from criteria to outcome.*

**Step 10. Build the corrections list: actionable, located, bounded, and closed.**
Every correction is one numbered item with a location, a statement of the problem, and what would resolve it. Do not write "improve the literature review". Write "at pp. 31 to 44, add engagement with the methodological objection to the dominant position, and state in one paragraph why the study's design is defensible in the light of it." Then bound the list: state that it is complete, that a candidate addressing every item has met the requirement, and who checks it. An unbounded list, or one inviting further changes at the checker's discretion, is unfair and produces a second round nobody planned. Count the list against the corrections period honestly; if it cannot be done in the time, the outcome is wrong, not the period. *Correct result: a numbered list, each item located with a resolution condition, and an explicit completeness statement.*

**Step 11. Calibrate against your own examining.**
Before submitting, compare this thesis against the two reference points named in Question 7 and against the bare-pass sentence from Step 2. Two checks: is this thesis being held to a standard you did not apply to the last one, and is the report's language proportionate, since a report whose tone implies a fail attached to a pass recommendation confuses everyone and the candidate will remember the tone. Across a cohort, check that the same criterion has been applied the same way to each, and that the volume of criticism tracks quality rather than reading order or fatigue. *Correct result: an explicit calibration note, and any adjusted judgements recorded with the reason.*

**Step 12. Handle examiner disagreement as a process, not a negotiation.**
Where co-examiners diverge, do not average and do not defer to seniority. Locate the disagreement first: examiners who appear to disagree about an outcome frequently agree on the facts and differ on classification (Step 7) or on the standard (Step 2). Exchange evidenced judgements rather than conclusions, and identify whether the divergence is about what is in the thesis, about whether a concern is fatal or correctable, or about the standard of the degree. Only the first is settled by re-reading; the second is settled by the regulations; the third requires the institution, and is what an independent chair exists for. Where genuine disagreement persists, the regulations almost always provide for separate reports or a third examiner. Use the provision rather than manufacturing a consensus neither examiner holds, because a concealed disagreement resurfaces at appeal. *Correct result: either an agreed evidenced position, or separate reports filed under the regulations, with the point of divergence stated plainly in both.*

## 8. Analytical framework

Examination runs on one chain, applied per criterion:

    Criterion → Located evidence → Judgement → Classification → Outcome

Each arrow is inspectable, and each is where examination fails in a distinct way. *Criterion → Located evidence* fails when an impression is not anchored, producing criticism the candidate cannot answer. *Evidence → Judgement* fails when the examiner's own preferred approach substitutes for the criterion. *Judgement → Classification* is the hardest arrow and the one that decides outcomes: fatal, correctable, or taste. *Classification → Outcome* fails when sympathy, fatigue or the perceived cost to the candidate moves the recommendation away from where the classification points.

Two calibration axes sit across the chain and must be held consciously, because both drift:

    The standard axis:    bare pass at this degree ←→ the best thesis I have read
    The ownership axis:   the thesis the candidate wrote ←→ the thesis I would have written

Every judgement should be placed on both. A criticism that only makes sense at the right-hand end of either axis is not a criticism, and belongs in the taste class or in a discussion at the viva.

## 9. Output format

Follow the institution's required report format where one exists; it takes precedence. Where it does not, or where you are producing working notes behind it:

**1. Examination details.** Thesis, degree, institution, examiner role, date, conflict declaration, and the AI and data position taken under Step 1.

**2. Summary of the thesis.** Three to five sentences in your own words: question, approach, findings, claimed contribution. This demonstrates the thesis was read as a whole and is the first thing a board looks for.

**3. Assessment against each criterion.**

| Criterion | Threshold or quality | Judgement | Located evidence (pp.) | Classification of any concern |
|---|---|---|---|---|

**4. The contribution.** The candidate's claim quoted, the type of originality, and the verdict on whether the thesis demonstrates it, with chapter locations.

**5. Concerns, classified.** Fatal, correctable and taste, each with location and the reason for the classification. The taste items appear here as discussion points, explicitly labelled as not criticisms.

**6. Questions for the viva**, where the model has one and the report precedes it. Each question tied to a location and to what the answer would change.

**7. Recommendation.** In the regulations' own category, with the criterion-to-outcome chain beneath it.

**8. Corrections list.** Numbered, located, each with a resolution condition, and closed with an explicit completeness statement.

**When the thesis is weak, the format must not manufacture balance, and when it is strong, it must not manufacture criticism** (K4 §1). A strong thesis produces a short report with few corrections, and that is a correct output, not a lazy one. Section 5 is not padded to look rigorous, and section 3 does not acquire a criticism per criterion for symmetry. Equally, a weak thesis does not acquire compensating praise: state what is sound, briefly and specifically, and let the criteria table carry the rest.

## 10. Quality checks

Run before the report is submitted.

1. Was the whole thesis read, by you, before any judgement was formed?
2. Were the criteria and the bare-pass sentence written down before reading, not after?
3. Does every judgement in the report carry at least one page or section location?
4. Is there any criticism that reduces to "I would have done it differently"? Reclassify it.
5. Has every concern been explicitly classified as fatal, correctable or taste, with a reason?
6. Does every fatal concern name the threshold criterion it engages?
7. Is any claimed prior work, anticipation or missing source one you have actually located and can cite?
8. Is the corrections list achievable within the regulations' corrections period, item by item?
9. Is the corrections list closed, so that a candidate addressing every item has met the requirement?
10. Does the recommendation use the regulations' own category names, and is it one the regulations offer?
11. Is the tone of the report consistent with the recommendation?
12. Has the thesis been calibrated against named reference points, and against the standard of this degree rather than the best you have read?
13. Where a discipline convention differs from your own and is legitimate, has that been removed from the criticisms?
14. Would the report be intelligible and answerable to the candidate, and auditable by a board, without you present to explain it?
15. Is anything in the report a statement about the candidate rather than about the thesis?

## 11. Common failure modes

| Failure | How to recognise it | How to prevent it |
|---|---|---|
| **Examining your own thesis** | Criticisms that name what you would have done | Test every criticism on the ownership axis (§8) |
| **Calibrating to the best** | A sound thesis described in the register of a weak one | Write the bare-pass sentence before reading (Step 2) |
| **The unlocated criticism** | "The analysis is superficial" with no page | No judgement without a location (Step 4) |
| **Chronicle instead of assessment** | The report walks through chapters in order | Assess criterion by criterion on the second pass |
| **Contribution assessed from memory** | "This has been done before", no citation | Locate the prior work or drop the claim (K4 §2.4) |
| **The unbounded corrections list** | Items like "strengthen chapter 3 throughout" | Each item located, with a resolution condition (Step 10) |
| **Corrections that require a new thesis** | A correction implying further data collection | Test each item against the corrections period |
| **Threshold and quality conflated** | A fail on a matter of degree, or a pass with a requirement missing | Classify each criterion in Step 2 |
| **Sympathy in the classification** | A fatal concern quietly relabelled as correctable | Classify before forming the recommendation (Steps 7, 9) |
| **Tone and outcome mismatched** | A harsh report attached to a clear pass | Calibration check before submission (Step 11) |
| **Averaged disagreement** | A joint report neither examiner would have written alone | Exchange evidence, not conclusions (Step 12) |
| **Fatigue drift** | Later chapters attract fewer or harsher comments | Assess by criterion, and re-check the last-read chapter |
| **AI-generated criticism of things not present** | A report citing a section, table or claim the thesis does not contain | Every location verified in the document (§12.4) |
| **Impression laundered into evidence** | A general view supported by a page reference chosen afterwards | The location must show the problem, not merely exist |

## 12. AI guardrails

Skill-specific. The universal prohibitions in K4 apply in full and are not repeated.

1. **These skills assist an academic's judgement, they do not replace it.** Assessment, examination and progression decisions are the responsibility of the named academic and cannot be delegated to an AI system. The skill never assigns a mark, never produces feedback to be passed to a student unread, and never makes a progression or examination recommendation the academic has not independently reached. Institutional regulations and marking criteria take precedence over anything in this skill, and where a student's work may have integrity concerns, that is an institutional process and not a judgement this skill makes.

2. **Settle the policy, confidentiality and data question before the work.** A thesis under examination is unpublished, confidential to the examination, and the personal data and intellectual property of an identified person. Institutional policy on AI in assessment governs whether this skill runs at all. Where the position is unknown, the thesis is not uploaded, and the output says what it was based on.

3. **Never produce an assessment of a thesis that has not been read in full by the examiner.** Where only part of the document is available, the output is confined to that part, names what it covers, and does not use the register of a whole-thesis judgement.

4. **Never cite a page, section, table, figure, source or quotation from the thesis that has not been verified in the document.** An examiner's report criticising something the thesis does not contain is fatal to the examiner's credibility and to the examination.

5. **Never assert that a contribution is anticipated by existing work without locating that work.** A recollection that something similar has been done is not evidence. Either produce the reference, verified, or do not make the criticism (K4 §2.4). This is the highest-consequence fabrication risk in examination, because it can cost a candidate a degree.

6. **Never state a recommendation, a mark, a grade or a classification.** Outcome language belongs to the named examiner and the board. The skill may lay out the criterion-to-outcome chain the examiner has built; it does not complete it.

7. **Never invent, complete or infer the regulations.** Outcome categories, corrections periods, report formats and the definition of the degree standard vary by institution and country. Where they are not supplied, say they are required and stop, rather than assuming a familiar model.

8. **Never make or imply an integrity allegation.** Do not comment on whether text appears AI-generated, plagiarised or fabricated. Where something concerns the examiner, the route is the institution's process, which carries evidence standards and rights of response this skill does not.

9. **Never smooth a disagreement between examiners.** Where two evidenced positions conflict, both are stated with their evidence. Do not generate a compromise position that neither examiner reached (K4 §4.1).

10. **Distinguish criticism from preference in the output itself, not silently.** Any observation that rests on the examiner's methodological or stylistic preference is labelled as such, and does not appear among the criticisms.

11. **Cap confidence where the thesis reports too little to evaluate** (K3 §3.4). Incomplete method reporting is a finding in its own right and is reported as one; it is not an invitation to reconstruct what the candidate probably did.

## 13. Best-practice principles

- **Write the standard down before you read.** Calibration set after reading is calibration set by the thesis in front of you, which is how a good thesis raises the bar for the next candidate and a poor one lowers it.
- **The first pass is for the spine.** Detail-level reading destroys the ability to see whether the thesis holds together, and structural failure is the finding with the largest consequence.
- **Every judgement carries a page number.** This one discipline improves examination more than any other, because it converts impressions into findings and forces the weak ones to fall out.
- **The taste category is your friend.** Every criticism you move out of the report because it is a legitimate alternative choice makes the remaining criticisms harder to dismiss.
- **A thesis is not weak because it is not excellent.** The standard is the degree's, and most theses that pass are neither remarkable nor deficient.
- **Assess the claim the candidate made.** Constructing a stronger contribution claim on their behalf and then finding it unsupported is a common and unfair error; so is assessing against a weaker claim than they made.
- **A correction that cannot be done in the corrections period is not a correction.** It is a different outcome, and calling it a correction moves a difficult judgement onto the candidate and the supervisor.
- **Read the methodology chapter for justification, not for agreement.** The question is whether the choice was reasoned against alternatives, not whether it matches yours.
- **Say what is good, specifically and briefly.** A report that identifies only problems tells the candidate and the board nothing about where the work stands, and it is usually read as an examiner who did not engage.
- **Examine the thesis, not the supervision.** Weak supervision is visible in many theses, and it is not the candidate's fault and not the examination's business. Raise it separately with the institution if it warrants raising.
- **Disagreement between examiners is normal and useful.** Concealing it produces a report neither examiner believes and an appeal nobody can answer.
- **The report is read by a person whose career it affects.** Precision is the courtesy that matters here; softening is not.

## 14. Worked example

Generic fictional scenario, academic, applied environmental science.

**INPUT**

An external examiner is appointed for a research doctorate. The thesis presents a two-year field monitoring study of a remediation technique at three sites, with an accompanying statistical model. The examiner's institution and the appointing institution both prohibit uploading examination material to external systems. There is a viva; independent reports are filed before it.

**PROCESS**

*Steps 1 and 2.* The examiner records the outcomes available under the regulations (five, with corrections periods of three and twelve months), the conflict position (none), and the AI position (working notes only, no thesis content shared). The criteria are copied out: two threshold (an original contribution to knowledge, and demonstrated capacity for independent research) and three quality. The bare-pass sentence is written: at this institution, a bare pass establishes something not previously established, by methods a competent researcher in the field would accept, reported well enough to be evaluated and built on.

*Step 3.* First pass, no annotation. Spine reconstructed in three sentences: the question is whether the technique performs under field conditions across contrasting site types; the approach is two years of monitoring at three sites plus a model relating performance to site characteristics; the claimed contribution is the first field-scale evidence of performance variability by site type. The spine holds. The examiner notes that chapter 6, the modelling chapter, felt disconnected from the field chapters, and marks it for the second pass.

*Step 4.* Criterion-by-criterion pass, fourteen located judgements. Among them: the field protocol at pp. 88 to 103 is unusually well documented and reproducible; the literature chapter engages the field's central disagreement at pp. 40 to 52; and at pp. 154 to 161 the model is fitted and evaluated on the same data, with no held-out validation and no discussion of overfitting, while chapter 7 uses its coefficients to generalise to site types not among the three studied.

*Step 5.* Contribution. The claim is quoted from p. 12. The originality type is new evidence rather than new method, and the field chapters demonstrate it. The generalisation in chapter 7 goes further than three sites can support.

*Step 6.* Methodology. The examiner would have used a different sampling interval and model family. Both of the candidate's choices are justified against alternatives at pp. 84 and 150 and executed correctly, so both go to the taste class, and the sampling interval becomes a viva discussion point rather than a criticism.

*Step 7 and the judgement call.* Two serious concerns, and the classification is not obvious. The modelling problem first reads as fatal: an unvalidated model used for generalisation. But the classification test asks whether the thesis contains what is needed. It does: two years of data at three sites supports a held-out validation by site or by period, and the raw data are in the appendices. The problem is in the analysis of existing material, not in the design or the data. It is therefore correctable, requiring re-analysis and a rewritten chapter 7, achievable within the twelve-month corrections period but not within three. The overreaching generalisation is a consequence of the first and is fixed with it. Neither engages a threshold criterion, because the contribution rests on the field evidence, not on the model.

Had the modelling been the contribution, the same technical problem would have been fatal, because the novel claim would have rested on an analysis that does not support it. The examiner writes that reasoning into the working notes, because it is what makes the classification defensible if challenged.

*Steps 9 to 11.* The report precedes the viva, so the recommendation is provisional and the modelling issue becomes the viva's central question, with a note of what the answer would change: if a held-out validation was done and omitted, the correction shrinks to a reporting fix; if not, it is a re-analysis. The corrections list has nine located items and a completeness statement. Calibration against two previously examined theses confirms the severity is proportionate, and two criticisms about referencing style are removed as taste.

**OUTPUT**

A provisional report with fourteen evidenced criterion judgements, one contribution verdict, a classified concerns table, five viva questions tied to locations, a nine-item bounded corrections list, and an explicitly provisional recommendation. `RESEARCHER SIGN-OFF REQUIRED` on the recommendation, which is finalised by the named examiners after the viva (K5 §2.5, §2.8).

## 15. Advanced usage

**Marking a cohort rather than examining one thesis.** The risk shifts from calibration against the best thesis to drift across the pile: order effects, fatigue, and the standard moving as the cohort's shape becomes apparent. Mark criterion by criterion across all scripts rather than script by script where volume allows, re-mark the first three at the end, and keep a boundary file of the scripts nearest each grade line so the boundary is defined by exemplars rather than memory. Where a rubric is used, record which descriptor each judgement matched, not only the band.

**Examining a paper-based or practice-based thesis.** The contribution assessment changes shape. For paper-based theses, published papers have passed external review, but the thesis-level contribution, the coherence of the collection and the candidate's share of co-authored work still require assessment, and the authorship statement is evidence to test rather than accept. For practice-based work, the artefact and the written component are usually assessed against different criteria and sometimes by different people, so run Step 2 twice.

**Using the classification table to prepare a viva.** Every correctable concern is a question with a knowable answer, and every taste item is a discussion. Building the viva from the table rather than from a general list of standard questions produces a defence that is about this thesis, and it lets you record in advance what each answer would change, which is what makes a post-viva recommendation defensible.

**Where the standard approach does not fit.** Where the thesis is in a sub-field you can assess only partly, mark each criterion with your confidence in your own judgement of it and say so to the institution rather than compensating with confident generality. Where the regulations do not define the degree standard usefully, calibrate against theses you have examined at that institution and say that is what you did.

## 16. Skill chain

**Recommended previous skills:**
- **15.22 Research Rigour Audit.** Where the methodological soundness of the underlying study needs establishing independently, before it is folded into a criterion judgement.
- **15.20 Supervision Feedback and Development.** The developmental work that precedes examination, and the record that shows what was advised.

**Recommended next skills:**
- **15.23 Student Progress and Milestone Planning.** Where the outcome is corrections or resubmission, and the candidate needs a plan against the period the regulations allow.
- **15.24 Journal Article Development.** Where the examination identifies material strong enough to publish, and the taste-class observations become useful suggestions.

**Runs well alongside:**
- **15.19 Examiner-Proofing and Defence Strategy**, which is the candidate-side counterpart and shows what a well-prepared candidate will have anticipated.
- **13.01 Research Quality Review**, whose appraisal logic informs the methodology criterion.
- **13.02 Source and Citation Verification**, where a thesis's referencing or a suspected anticipation claim needs checking properly rather than from memory.

---
A Yazi Supplied Skill and resource.
