---
name: sentiment-and-emotion-analysis
description: >
  Validity-controlled sentiment and emotion classification of research text, with
  measured agreement against a human-labelled sample, the reason paired to every
  score, and explicit refusal where the text or base cannot support it. Use for
  "run sentiment on this", "how positive are these responses", "sentiment
  analysis of reviews", "emotion in these verbatims", "is sentiment up or down
  this month", "score these comments", "sentiment tracker".
category: 07 Qualitative Analysis
ref: 07.05
tier: 1
inherits: [K2, K3, K4, K5]
---

# Sentiment and Emotion Analysis

## 1. One-line description
Classifies the valence, and where justified the emotion, of research text under explicit validity controls: measured agreement against human labels, the cause attached to every score, and a stated refusal where the material or the base cannot support the claim.

## 2. What this skill is used for

**The research problem it solves.** Standalone sentiment scoring is one of the weakest things routinely sold as analysis in commercial research, and the reason is simple: a sentiment score is a compression of text into a single number whose meaning nobody has defined, produced by a classifier nobody has validated, reported on a base nobody has checked, and narrated as movement when it moves within its own error. It answers a question the business did not ask ("how positive was the text?") in place of the one it did ("what is going wrong, and for whom?"). Two failure patterns dominate. The first is the decorative sentiment dashboard: a percentage positive with no cause attached, which cannot be acted on because it does not say what would have to change. The second is the sentiment tracker, where a two-point movement inside a noisy classifier gets a paragraph of explanation in the monthly report, and the organisation spends a year responding to noise. This skill states that position and then does the job properly, because sentiment done with controls is genuinely useful: it prioritises where to look, it segments a large corpus, and it makes a coded dataset navigable. What it is not, ever, is a finding on its own.

**Where it sits.** Analysis, downstream of content coding. It attaches valence to material whose subject matter is already known, and hands prioritised, cause-attached segments to insight development and reporting.

**Typical use cases.**
- Attaching valence to already-coded open ends so a code frame can be read as problems and strengths.
- Prioritising which of several thousand reviews or support comments a human should read first.
- Aspect-level sentiment: what customers feel about each part of a service, separately.
- Segmenting a large text corpus before qualitative analysis of a sampled subset.
- Monitoring a text stream for genuine step changes, with a noise band established in advance.
- Validating or challenging an existing automated sentiment feed that the business already relies on.

**Who uses it.** CX and insight teams running continuous feedback; analysts handed a text corpus and a dashboard requirement; researchers asked to explain a sentiment movement that may not be real; anyone who has been asked to "run sentiment on this" and suspects that is the wrong request.

## 3. When to use it

- The corpus is large enough that reading everything is impossible, and you need a defensible way to decide what to read.
- Content coding has already been done, and valence would make the coded frame more useful.
- The question is aspect-level: not "how do customers feel" but "what do customers feel about each specific thing, and why".
- A sentiment feed already exists in the business and its validity has never been tested.
- You need to detect genuine step changes in a text stream and want a noise band established before anything is narrated.
- The text is substantial enough to carry sentiment: full sentences with context, not one-word answers.
- A human-labelled sample can be produced, so agreement can be measured rather than assumed.

## 4. When NOT to use it

- **Sentiment is being asked for instead of coding.** Where the business wants to know what is wrong, valence is not the answer and content is. Run **07.02 Open-Ended Response Coding** or **07.01 Thematic Analysis** first. **This skill is downstream of coding, never a replacement for it**, and a sentiment output with no content codes behind it should not be produced at all.
- **The text is too short to carry sentiment.** One to five word responses, rating justifications like "fine" or "price", and single-clause fragments do not contain enough context for any classifier, human or machine, to be reliable. Report the content codes and decline the score.
- **The base is too small.** Below roughly 100 items in a reporting cell, a sentiment proportion is noise dressed as a metric, and below 30 no percentage is permitted at all, per K4 §7. This applies to every subgroup and every time period separately, not to the corpus total.
- **No human-labelled sample can be produced.** Without a reference set, agreement cannot be measured, and an unvalidated classifier's output is an opinion with a decimal point. Where labelling is genuinely impossible, the output is qualitative segmentation for triage only, explicitly not a measurement, and no percentage is published.
- **The material is translated and the wording carries the valence.** Translation systematically normalises tone: sarcasm, understatement, politeness formulas and intensifiers are the first casualties. Sentiment on translated text measures the translator as much as the participant, per K5 §2.2.
- **The corpus is dominated by the known failure conditions.** Where screening (step 3) shows heavy sarcasm, comparison, conditionals or domain vocabulary that inverts ordinary polarity, the classifier is measuring something other than sentiment. Report the screening result and stop, rather than publishing a score you know is wrong.
- **The deliverable is a single headline sentiment number for an organisation.** A whole-corpus sentiment percentage aggregates incommensurable things: a complaint about billing and a compliment about a shop assistant do not average into a meaningful quantity. Report by aspect and by segment, or do not report.
- **The request is to explain a small movement in an existing tracker.** Where the movement sits inside the classifier's measured error, the correct output is that the movement is not distinguishable from noise, per K4 §3.1. Producing an explanation for it is the most common way this skill gets misused.
- **Emotion labels are wanted for a clinical, diagnostic or wellbeing purpose.** Emotion classification from text has weak reliability even in ideal conditions and no diagnostic standing. Where participant welfare is involved, this is **13.05 Research Ethics and Consent Design**, not a text-analytics task.

## 5. Required inputs

**Required.** Without these the skill cannot run. If absent, ask. If no answer is available and the work must proceed, state the assumption at the point where it bites, per K5 §5.

- **The text, complete and untruncated**, with an identifier on every item, per K2 §4.2. Truncated text changes valence, because qualifications tend to come at the end.
- **The content codes**, from prior coding. Sentiment is attached to a subject; a score with no subject cannot be reported.
- **The question or context that produced the text.** Text elicited by "what went wrong?" is negative by construction, and its negativity is a property of the question, not of the customer base.
- **A human-labelled reference sample**, or the ability to produce one. Without it, no agreement figure exists and no percentage may be published.

**Optional, and what each one adds.**

- **A closed satisfaction, recommendation or rating measure on the same respondents**: provides an external criterion against which classified sentiment can be checked, which is stronger evidence of validity than agreement with human labels alone.
- **Aspect or entity annotations**: allow sentiment to be resolved to the thing it is about, which is the difference between a usable output and a dashboard number.
- **Prior period data with the same classifier and the same frame**: makes a noise band computable and a genuine step change detectable.
- **The source and characteristics of each item** (channel, prompt, respondent segment): sentiment distributions differ sharply by channel, and pooling them produces a number about the channel mix.
- **Domain vocabulary or a category glossary**: identifies the terms whose ordinary polarity does not apply here, which is a large share of the errors in specialist corpora.

## 6. Questions to ask before starting

1. **What decision would a sentiment score change?** *Default if unanswered:* treat sentiment as a triage and prioritisation tool only, produce no headline metric, and say why.
2. **What is the unit being scored: the item, the sentence, or the aspect?** Mixed sentiment within one response is normal, and item-level scoring destroys it. *Default:* score at aspect level where aspects are available, at clause level otherwise, and roll up with the mixed category preserved.
3. **Can a human-labelled sample be produced, and by whom?** Determines whether a percentage may be published at all. *Default:* require at least 200 items double-labelled, and if that is impossible, publish no metric.
4. **Is this a one-off read or a tracker?** Determines whether a noise band must be established before any movement is narrated. *Default:* on a tracker, no movement is reported until the band exists.
5. **What is the elicitation?** Complaint channels, review platforms, post-resolution surveys and unprompted feedback have wholly different baseline distributions. *Default:* report each channel separately and never pool without stating the mix.
6. **Is the text in the participants' first language, and is it translated?** *Default:* score original-language text only; where translated, mark the whole output as translation-dependent and restrict it to triage.
7. **Does the category invert ordinary polarity anywhere?** In some domains "aggressive", "sharp", "cheap", "addictive" or "brutal" are neutral or positive. *Default:* build a short domain glossary at step 3 and check it against the errors found in validation.

## 7. Step-by-step methodology

**The position this skill takes, stated in the output.** A sentiment score is a classification of expressed valence in a piece of text, produced by a model, under measurable error. It is not a measurement of how a person feels, not a measurement of satisfaction, and not comparable across corpora produced under different elicitations. Everything below is designed to make that statement true of the specific output, rather than leaving it as a caveat nobody reads.

**1. Gate the request before running anything.** Check four conditions: content coding exists; the text is long enough to carry valence; the reporting cells will clear the base thresholds; and a human-labelled sample is producible. If any fails, say which, and offer what can honestly be done instead. *Correct result:* either a documented go decision or a written statement of which gate failed and what the alternative is. **Most misuse of sentiment analysis is prevented here and nowhere else.**

**2. Define the construct and the unit before scoring.** Write down what a positive label means for this corpus: expressed positive evaluation of the subject, satisfaction with an outcome, positive affect, or absence of complaint. These are different constructs and they produce different distributions on identical text. Then fix the unit: item, sentence, clause or aspect. *Correct result:* a one-paragraph construct definition and a stated unit. Skipping this is why two analysts scoring the same corpus disagree by ten points and neither is wrong.

**3. Screen the corpus for the failure conditions, and profile them.** Draw a sample of 100 to 200 items and read them against a checklist, recording the incidence of each condition.

- **Sarcasm and irony.** "Brilliant, another outage." Surface polarity inverts. High incidence in complaint and social channels.
- **Negation and scope.** "Not bad at all" and "I would not say I was unhappy" defeat token-level polarity; negation scope errors are the classic short-text failure.
- **Comparison.** "Better than the old one" says nothing about absolute valence, and against a competitor it may be a criticism of someone else.
- **Conditionals and hypotheticals.** "If they fixed the app I would be delighted" is a complaint that reads positive.
- **Mixed sentiment in one item.** "The staff were lovely but I waited fifty minutes." Item-level scoring will pick one and discard the other; both are the finding.
- **Code-switching and mixed language.** Sentiment-bearing words in a second language pass through unclassified, and the item is scored on its neutral remainder.
- **Non-English and translated text.** Politeness conventions, understatement and intensifier norms differ by language and are not preserved by translation.
- **Domain vocabulary.** Where a negative-polarity word is neutral or positive in context ("this drug is aggressive" in oncology, "cheap" in value retail, "addictive" in games), general classifiers are reliably wrong.
- **Very short text.** Under about eight words, there is usually not enough context for any label to be defensible.
- **Reported speech and quotation.** "The agent said it was my fault" carries the agent's valence, not the respondent's.

*Correct result:* an incidence table for each condition and a domain glossary of inverted terms. Where any condition exceeds roughly 10 percent of the corpus, it is disclosed with the results and its likely direction of bias stated. Where the conditions collectively dominate, step 1's gate is reapplied.

**4. Score content first, valence second.** Sentiment is attached to already-assigned content codes, per **07.02**. This is not a sequencing preference; it is what makes the output a finding. "Sentiment on delivery is 62 percent negative, driven by the delivery-window code" is actionable. "Sentiment is 62 percent negative" is not, because nothing follows from it. *Correct result:* every scored item carries at least one content code, and the reporting unit is code-by-valence, not valence alone.

**5. Build the human-labelled reference set properly.** Draw 200 to 400 items at random, stratified across content codes and across text length, because error concentrates in short items and in specific codes. Have two humans label independently against the step 2 construct definition and a written label set that includes **mixed** and **no sentiment expressed** as first-class options, not as residuals. Adjudicate their disagreements and record where they disagreed, because human disagreement marks the genuinely ambiguous items and sets the realistic ceiling on machine performance. *Correct result:* an adjudicated gold set with a recorded human-to-human agreement figure. **A classifier cannot meaningfully be held to a standard the humans could not reach**, and reporting machine agreement without the human baseline makes it uninterpretable.

**6. Score the corpus at the chosen unit, preserving mixed.** Apply the classifier. Where an item contains opposing valences about different aspects, retain both rather than resolving to a net. **Never average opposing sentiments within an item into a neutral score**: "lovely staff, fifty-minute wait" is not a neutral experience and reporting it as one deletes both findings. *Correct result:* a scored dataset in which mixed items are identifiable and countable, and the proportion of mixed items is itself reported, because a rising mixed rate usually means the classifier or the frame is losing resolution.

**7. Measure agreement against the reference set, and report it per class and per segment.** Compare classifier output to the gold labels. Report: overall agreement, agreement per class (positive, negative, neutral, mixed), and where the errors go, because a classifier that is strong on positive and weak on negative will systematically flatter a corpus. Report agreement separately for the segments and text-length bands that will be reported, because **an overall accuracy figure hides that the classifier is near-random on the short items that make up a third of the corpus**. Name the statistic used and the threshold you regard as acceptable, before seeing the result. *Correct result:* an agreement table with per-class and per-segment figures, the human baseline alongside, and an explicit statement of which segments the classifier is not good enough for. Those segments are not reported as sentiment.

**8. Pair every score with its reason.** For every sentiment figure that will be published, produce the content codes driving it, ranked, with their own valence splits, and two or three verified verbatims per driver, selected per **07.04**. *Correct result:* no sentiment number appears anywhere without the cause beneath it. **Sentiment without cause is not a finding, it is a mood ring**, and this is the single rule that separates useful sentiment work from decorative sentiment work.

**9. Apply the reportability gates.** For each reporting cell, check: base at or above threshold; median text length above the minimum; classifier agreement in that segment above the stated threshold; and the failure-condition incidence within tolerance. Cells failing any gate are reported as counts with verbatims, or not at all. *Correct result:* a reportability table showing which cells passed and which were suppressed, with reasons. Suppressed cells are shown as suppressed, not omitted silently, because an absent cell in a dashboard reads as zero.

**10. State confidence per segment, not overall.** Confidence follows the weakest relevant factor per K3 §3, and the relevant factors differ by cell. A segment with 2,000 long items and 0.8 agreement supports high confidence; a segment with 90 short items and 0.5 agreement supports none. *Correct result:* a confidence column in the results table, per segment, with the reason. A single overall confidence statement is a K3 failure because it averages the strong and the weak, and readers act on the weak parts as though they were strong.

**11. Treat emotion classification as a separate, weaker instrument.** Where classification beyond valence is required (frustration, anxiety, delight, disappointment), state plainly that reliability is materially lower than for valence: the categories are less distinct, human agreement is lower, cultural expression varies, and the label set is a theoretical choice rather than a natural fact. Use a small set of three to six categories defined for this corpus, build a separate gold set for it, measure agreement separately, and do not publish any emotion category whose agreement falls below the valence threshold. Never infer an emotional state the text does not express: absence of an emotion word is not evidence of the emotion's absence, and its presence is a statement, not a diagnosis. *Correct result:* either an emotion output with its own measured agreement and an explicit reliability caveat, or a decision not to classify emotion, with the reason.

**12. On a tracker, establish the noise band before narrating any movement.** Compute the expected variation from the classifier's own error and from sampling, using the measured agreement and the period base. Publish the band as part of the tracker specification, before the next wave lands. Then apply one rule: **a movement inside the band is reported as no change, in those words, and is not explained.** A two-point move in a monthly sentiment feed is almost always noise, and narrating it teaches the organisation that the metric is meaningful at a resolution it does not have. Where the band is wider than any movement the business cares about, the correct conclusion is that this metric cannot serve as a tracker and should be replaced by coded content volumes, which are far more stable. *Correct result:* a stated band, a movement rule, and a tracker report that says "no change" when there is none. Any frame or classifier change is a declared break, run in parallel for one period, per the tracker discipline in **07.02**.

**13. Write the finding so that it cannot be over-read once it leaves the page.** Sentiment figures travel badly: a percentage positive lifted onto a board slide loses its channel, its base, its agreement figure and its cause in one move, and arrives looking like a measure of how customers feel about the company. Three drafting rules close that gap. First, **name the elicitation inside the sentence, not in a footnote**: "of app store reviews mentioning billing, 64 percent expressed negative evaluation" rather than "billing sentiment is 64 percent negative". Second, put the driver codes in the same sentence or the immediately adjacent one, so the number cannot be quoted without its cause. Third, state the negative case explicitly where a reader would otherwise infer it: the complement of negative sentiment is not satisfaction, it is everything else, including items expressing nothing. Then read the whole output back asking one question: if a reader saw only the headline sentences, would they believe anything the analysis does not support? *Correct result:* findings that survive being quoted alone, and an explicit statement in the front matter that this analysis measures expressed valence in a specific corpus under a specific elicitation, which is the position from the top of this section made concrete rather than left as a disclaimer.

## 8. Analytical framework

The chain the output is built on:

    Text → Content code → Aspect → Valence (with unit) → Validated score
        → Cause-paired finding → Reportability gate → Confidence per segment

**Applying it.** Valence enters only after content and aspect are fixed, and leaves only through the gate. The two arrows that are routinely skipped are the third and the seventh, and skipping either produces the same artefact: a number about everything, which is a number about nothing.

**What a sentiment score is and is not.**

| It is | It is not |
|---|---|
| A classification of expressed valence in text | A measurement of how a person feels |
| Comparable within a corpus and a classifier | Comparable across corpora, channels, languages or classifiers |
| A prioritisation and triage instrument | A key performance indicator |
| Informative when paired with its cause | Actionable on its own |
| Subject to measurable error | Precise because it has a decimal point |

**The five reportability gates.** All five must pass for a cell to be reported as a percentage.

| Gate | Threshold | If it fails |
|---|---|---|
| Base | n ≥ 100 in the cell (n ≥ 30 absolute floor) | Report counts and verbatims, no percentage |
| Text length | Median above the minimum set at step 2 | Report content codes only |
| Agreement in this segment | At or above the pre-stated threshold | Suppress and state why |
| Failure-condition incidence | Within the tolerance set at step 3 | Disclose with direction of bias, or suppress |
| Cause available | Content codes attached | Do not publish the number |

**The tracker noise band.** Movement is classified before it is described.

| Movement | Reported as |
|---|---|
| Inside the band | No change. No explanation offered |
| At the band edge | Directional, untested, watch next period |
| Beyond the band, one period | A change, with the cause codes that moved |
| Beyond the band, sustained two or more periods | A step change, reported with confidence per K3 |

## 9. Output format

**A. Sentiment note (front matter)**
The position statement on what this score is and is not; construct definition and unit; classifier used in general terms and the fact that AI performed the classification, per K4 §7; the elicitation and channel mix; the failure-condition screening profile; the reference set design and the human baseline agreement; thresholds set in advance; consolidated review points per K5 §3.3.

**B. Validation table**

| Segment / text-length band | n in gold set | Human baseline agreement | Classifier agreement | Agreement by class (pos / neg / neu / mixed) | Dominant error direction | Reportable (Y/N) |
|---|---|---|---|---|---|---|

**C. Sentiment by content code (the main result)**

| Content code | n | % positive | % negative | % neutral | % mixed | Base gate | Confidence | Top driver verbatims |
|---|---|---|---|---|---|---|---|---|

Never a single corpus-level figure without this table beneath it.

**D. Aspect table**
Aspect, valence split, the codes contributing, and the two or three verified verbatims per aspect.

**E. Reportability and suppression log**
Every cell not reported, with the gate it failed. Suppressed cells appear here; they are never silently absent.

**F. Tracker section, where applicable**
The noise band and how it was computed; the movement classification for each reported cell; every cell inside the band explicitly labelled no change; any declared frame or classifier break.

**G. Emotion section, where applicable**
Label set and why these categories; separate gold set and agreement; the reliability caveat, stated at the top rather than in a footnote; categories suppressed for low agreement.

**H. What this cannot tell you**
The constructs sentiment does not measure here; the segments suppressed and what would be needed to report them; the explicit statement that absence of expressed sentiment is not neutrality.

**When the evidence is thin.** No cell is filled to complete the table. Where a segment failed a gate, the cell reads suppressed with the gate named, not a percentage with a small-base asterisk. Where the whole corpus fails the length gate, the output is content codes and verbatims with no sentiment at all, and the front matter says why. Where agreement came in below threshold, the classification is reported as triage only and no percentage is published anywhere in the deliverable, including in the appendix, because appendix numbers get lifted.

## 10. Quality checks

Run before anything is presented. Sits on top of K4 §8.

1. Does every published sentiment figure have content codes and verbatims attached to it?
2. Was agreement measured against a human-labelled gold set, and is the human baseline reported alongside it?
3. Is agreement reported per class and per reported segment, rather than only overall?
4. Is any segment reported whose in-segment agreement fell below the pre-stated threshold?
5. Were the thresholds set before the results were seen, and are they in the output?
6. Is mixed sentiment preserved as a category rather than averaged into neutral?
7. Does the failure-condition screening profile appear in the output, with the incidence of each condition?
8. Are channels and elicitations reported separately rather than pooled?
9. Does every cell clear the base gate, and are suppressed cells shown as suppressed?
10. Is confidence stated per segment rather than once for the whole analysis?
11. On a tracker, is the noise band published, and is every movement inside it labelled no change with no explanation attached?
12. Is any sentiment figure compared across corpora, channels, languages or classifier versions?
13. Is emotion classification, if present, carrying its own agreement figures and its reliability caveat at the top?
14. Does the output state that absence of expressed sentiment is not neutrality and not satisfaction?
15. Is any translated text scored, and if so is the whole output marked translation-dependent?

## 11. Common failure modes

| Failure | How to recognise it | How to prevent it |
|---|---|---|
| **Sentiment without cause** (the defining failure) | A dashboard shows percentage positive and nothing about what drives it | Step 8. No number is published without its content codes and verbatims |
| **Noise narrated as finding** (the tracker failure) | A monthly report explains a two-point movement | Publish the noise band first. Inside the band is reported as no change, in words |
| **The single corpus number** | One headline sentiment figure for a whole organisation | Report by code, aspect and segment. A pooled figure averages incommensurable things |
| **Unvalidated classifier** | An accuracy figure is quoted from the model's documentation, not measured on this corpus | Build the gold set. Validation is corpus-specific or it is not validation |
| **Accuracy hiding a weak segment** | Overall agreement is respectable, short-item performance is near random | Report agreement per segment and per text-length band, and suppress what fails |
| **Mixed collapsed to neutral** | Items containing praise and complaint score as neutral, and neutral is large | Score at aspect or clause level. Preserve and count mixed |
| **Sarcasm inversion** | Complaint channels score unexpectedly positive | Screen at step 3. Disclose incidence and direction of bias |
| **Domain polarity error** | The classifier marks accurate technical description as negative | Build the domain glossary. Check it against validation errors |
| **Elicitation mistaken for population feeling** | "Customers are 70 percent negative", from a complaints inbox | Report the elicitation with every figure. Never generalise beyond it, per K4 §3.3 |
| **Emotion labels presented with valence-level confidence** | Six emotion categories reported with no separate validation | Separate gold set, separate agreement, caveat at the top, suppress what fails |
| **Silence read as neutrality** | Items with no expressed sentiment are counted as neutral and inflate the middle | "No sentiment expressed" is a distinct category and is reported as one |
| **Comparison across classifier versions** | A trend line spans a model upgrade | Declare the break. Run in parallel for one period. Never draw across it |
| **Sentiment substituted for coding** | The deliverable has valence and no themes | Sentiment is downstream of coding. Without codes, do not run it |

## 12. AI guardrails

Skill-specific only. Universal prohibitions are inherited from K4 and are not repeated here.

1. **Never publish a sentiment percentage without measured agreement against a human-labelled sample from this corpus**, reported alongside the human baseline.
2. **Never publish a sentiment figure without the content codes and verbatims that explain it.** Sentiment without cause is not a finding.
3. **Never average opposing sentiments within an item into neutral.** Mixed is a category, and it is counted.
4. **Never report a movement inside the noise band as a change**, and never offer an explanation for one.
5. **Never state a single confidence level for a whole sentiment analysis.** Confidence is per segment, set by the weakest factor in that segment.
6. **Never compare sentiment across corpora, channels, languages, elicitations or classifier versions** without declaring the break and treating the comparison as untested.
7. **Never treat absence of expressed sentiment as neutrality, satisfaction or agreement.**
8. **Never score text below the length threshold, or a cell below the base threshold**, and never fill a suppressed cell with an asterisked estimate.
9. **Never present emotion classification with the same confidence as valence**, and never infer an emotional or psychological state the text does not express.
10. **Never run sentiment as a substitute for content coding**, and never present a valence output as though it answered a question about causes.
11. **Never score translated text without marking the entire output translation-dependent**, and never make a claim resting on tone in translated material.

## 13. Best-practice principles

1. **The most useful sentiment output is a filter, not a metric.** Its best job is deciding what a human reads next, and it is very good at that.
2. **Valence is cheap; cause is the product.** Any competent system will give you positive and negative. Only the coding tells anyone what to do.
3. **Validate on your corpus or not at all.** A classifier's published accuracy was measured on somebody else's text, in another domain, at another length.
4. **The humans set the ceiling.** If two trained labellers agree on 78 percent of items, a machine reporting 91 percent agreement is not better than the humans; something is wrong with the comparison.
5. **Short text is where sentiment goes to die.** The items most likely to be misclassified are the ones a corpus has most of.
6. **Mixed is the interesting category.** Items containing both praise and complaint are usually the ones that explain the relationship, and they are the first thing item-level scoring destroys.
7. **Every corpus has a baseline, and it belongs to the question.** A complaints channel is negative because it is a complaints channel. Compare like elicitations or compare nothing.
8. **Set the thresholds before you see the results.** Thresholds chosen afterwards are conclusions with a procedure attached.
9. **A tracker whose noise band is wider than the movements the business cares about is not a tracker.** Say so, and offer coded content volumes instead, which move for reasons.
10. **Emotion categories are a theory, not an observation.** Choose a small set, define it for this corpus, and hold it to a higher evidential bar than valence, not a lower one.
11. **Report the suppressions.** A dashboard cell that is empty because it failed a gate tells a reader more than one filled with an unreliable number.
12. **When someone asks for sentiment, find out what decision it is for.** Usually the honest answer is a coded frame with valence attached, and they will prefer it once they see it.

## 14. Worked example

*Fictional scenario, used for illustration only. All figures, responses and findings below are invented for the purpose of demonstrating method.*

**INPUT.** A telecommunications provider runs a monthly text feedback programme: post-interaction survey comments, app store reviews and a complaints inbox, roughly 9,000 items a month combined. The existing dashboard shows one number, "customer sentiment", currently 58 percent positive, down from 61 percent. Leadership has asked for an explanation of the three-point fall. Content coding exists from a 41-code frame.

**PROCESS.**

*Step 1, and the first judgement call.* The gate is applied to the request as asked. Content coding exists, so that gate passes. The others do not survive contact: the three channels are pooled, and the channel mix moved between months (a service incident drove complaints inbox volume from 12 percent to 19 percent of the total). No validation has ever been run on this classifier. Median text length in the post-interaction channel is six words. **The requested deliverable, an explanation of the three-point fall, cannot honestly be produced**, because a substantial part of the movement is a change in what is in the corpus rather than a change in how customers feel. This is stated plainly, with the mix figures, and the work is rescoped to: validate the classifier, de-pool the channels, attach causes, and establish a noise band.

*Steps 2 to 3.* Construct defined as expressed positive or negative evaluation of the provider or the interaction. Unit set to aspect level, because screening finds 23 percent of review-channel items contain opposing valences. Screening profile on 200 items: sarcasm 14 percent in the complaints inbox and 2 percent elsewhere; comparison 9 percent, mostly against a competitor; conditionals 6 percent; mixed 23 percent in reviews; very short text 44 percent of the post-interaction channel. Domain glossary built: "unlimited", "capped", "throttled" and "fair use" carry technical meaning that the general classifier reads as negative even in neutral description.

*Steps 5 to 7, and the second judgement call.* Gold set of 320 items, stratified by channel, code and length, double-labelled. Human baseline agreement 0.79, with disagreement concentrated in short items, which is informative in itself. Classifier agreement overall 0.72, which reads acceptable. Broken out, it is 0.84 on reviews (long text), 0.74 on the complaints inbox, and **0.51 on the post-interaction channel**, which is close to chance on a three-class problem and is 44 percent of the corpus. Per-class analysis shows the errors are asymmetric: negatives are misclassified as neutral far more often than the reverse, so the corpus reads more positive than it is. Decision: **the post-interaction channel is suppressed for sentiment reporting entirely** and reported as content codes and verbatims only. The suppression is published, with the agreement figure, rather than the channel quietly disappearing from the dashboard.

*Steps 8 to 10.* Sentiment reported by content code within each remaining channel, with drivers and verified verbatims per **07.04**. The reviews channel shows negative sentiment concentrated in two codes, one of which (billing clarity) rose in both volume and negative share. Confidence stated per segment: high for reviews, moderate for complaints inbox, not reported for post-interaction.

*Step 12, and the third judgement call.* The noise band is computed from measured agreement plus sampling variation on the monthly base. For the reviews channel it comes out at plus or minus 4 points. **The original three-point fall sits inside the band**, and the pooled figure it was drawn from was not comparable month to month anyway. The output states, in those words, that there is no measurable change in overall sentiment, and that the substantive finding is elsewhere: the volume of billing-clarity comments doubled, which is a stable count rather than a noisy proportion, and it is the thing worth acting on. Recommendation flagged for researcher sign-off per K5 §2.5: replace the single sentiment headline with coded complaint volumes plus aspect-level sentiment on the two channels that support it.

**OUTPUT.** A validation table showing the classifier is unusable on 44 percent of the corpus; a suppression log; sentiment by content code for two channels with causes and verbatims attached; per-segment confidence; a published noise band of plus or minus 4 points; an explicit statement that the three-point fall is not distinguishable from noise and is partly a channel-mix artefact; and a recommendation to change what the dashboard reports.

## 15. Advanced usage

**Aspect-based sentiment done properly.** Resolve each sentiment-bearing clause to the entity or aspect it evaluates, rather than to the item. This is more work and it is the only version of sentiment analysis that reliably produces action, because it answers "what about the service is bad" rather than "is the service bad". It also handles mixed items natively, since opposing valences attach to different aspects rather than competing for one label.

**Validating an existing business feed.** Where an organisation already relies on a sentiment number, the highest-value use of this skill is to build a gold set and measure the feed against it, per segment. Report the result as a measurement of the instrument, not as a criticism of the team that built it, and pair it with what the feed can still be trusted to do, which is usually triage. Expect the weakest performance in exactly the segments the business watches most closely, because those are usually the shortest text.

**Sentiment as a sampling frame.** For a very large corpus, use classified valence to draw a stratified sample for genuine qualitative analysis: the strongly negative, the strongly positive, and, most usefully, the mixed and the ambiguous. Hand that sample to **07.01** or **07.02**. This uses the classifier for what it is good at and never publishes its numbers.

**Multi-language corpora.** Score each language separately with its own gold set and its own agreement figure, and never pool the results into a cross-market sentiment figure. Politeness norms alone will produce apparent market differences of ten points or more that have nothing to do with satisfaction. A native reviewer is required per K5 §2.2.

**Combining with a closed measure.** Where the same respondents gave a rating, cross classified sentiment against it. Agreement is a validity check; systematic disagreement is a finding. Respondents who rate highly and write negatively, or the reverse, are usually the most informative group in the dataset, and they are invisible in either measure alone. Hand this to **07.06**.

**When the standard approach does not fit.** Where the corpus is above roughly 50,000 items and the goal is structure rather than valence, **06.06 Large-Scale Text Analytics** is the correct instrument. Where the interest is the emotional experience itself rather than its distribution, this skill is the wrong tool and depth qualitative work is the right one.

## 16. Skill chain

**Recommended previous skills:**
- **07.02 Open-Ended Response Coding.** Hands over content codes, without which no sentiment figure can be reported. This skill is strictly downstream of it.
- **06.06 Large-Scale Text Analytics**, where the corpus needs structuring before either coding or classification.
- **04.02 Data Cleaning**, to ensure text arrives complete, untruncated and with identifiers intact.

**Recommended next skills:**
- **07.04 Quote and Evidence Extraction.** Supplies the verified verbatims that must accompany every published sentiment figure.
- **07.06 Qualitative and Quantitative Integration**, where classified sentiment is checked against a closed satisfaction or recommendation measure.
- **08.01 Finding to Insight Development**, which receives cause-paired sentiment rather than a valence number.

**Runs well alongside:**
- **05.05 Trend and Tracker Analysis**, for the noise-band discipline where sentiment feeds a continuous metric.
- **13.03 AI Output Verification**, run against the classified dataset and the validation table.
- **K5**, at the two review points this skill mandates: the decision to suppress a segment, and any recommendation to change what a business metric reports.

---
A Yazi Supplied Skill and resource.
