---
name: data-visualisation-and-chart-selection
description: >
  Chooses the right chart for the analytical question and builds it honestly:
  encoding matched to question type, data ordered by value, axes and scales that
  do not mislead, base sizes shown, direct labelling, colour that encodes
  meaning and works without colour vision, and the finding stated on the chart.
  Use for "what chart should I use", "which visual for this data", "is this
  chart misleading", "should this be a pie chart", "how do I show this
  finding", "clean up these charts", "the charts are hard to read".
category: 11 Reporting and Storytelling
ref: 11.04
tier: 1
inherits: [K2, K3, K4, K5]
---

# Data Visualisation and Chart Selection

## 1. One-line description
Selects and builds the visual form that answers a specific analytical question, so that the chart states a finding, carries its base, and cannot be read as claiming more than the data supports.

## 2. What this skill is used for

**The research problem it solves.** Charts in research reports are usually made from the wrong starting point. The researcher has a table, so a chart of the table gets made. The result is a report of visuals that are individually accurate and collectively useless: they display data instead of answering questions, they follow questionnaire order instead of value order, they carry legends the reader has to decode, and they hide their base sizes. The reader works harder than they would have with a sentence.

Underneath that is a more serious problem. **A chart is an argument made in a visual grammar, and the grammar can assert things the data does not.** A truncated bar axis claims a difference that is not there. A dual axis manufactures a relationship out of two scaling decisions. A circle sized by radius exaggerates a ratio quadratically. A stacked bar invites comparison of segments that share no baseline. None of these requires bad faith and all of them survive review, because the numbers underneath are right and the misleading part is the encoding. Chart honesty is therefore not a design topic. It is an evidence topic, and it belongs in the same protocol as base sizes and confidence language.

Selection and execution are treated here as one decision, because they are one decision. Choosing a bar chart and then ordering it by questionnaire number produces a different claim from choosing a bar chart and ordering it by value. The choice is not finished until the chart is built.

**Where it sits.** Reporting. It runs alongside narrative, writing, compilation and presentation rather than after them: the visual is decided when the claim is decided. It hands the finished visuals to design and to the deck.

**Typical use cases.**
- Deciding the visual form for each finding in a report or deck.
- Rebuilding a chart set that is accurate and unreadable.
- Auditing a set of charts for encodings that mislead.
- Presenting a segmentation, a journey or a qualitative theme structure, where the default chart types do not fit.
- Showing a finding on a small base without implying precision the base does not support.
- Cutting a chart set down when there are more charts than findings.

**Who uses it.** Researchers and analysts building their own visuals; insight leads reviewing a deck before it goes out; anyone who has been asked whether a chart is misleading and needs a principled answer rather than an instinct.

## 3. When to use it

- A finding needs a visual form and the default choice is not obviously right.
- A chart set exists and the reader cannot see the findings in it.
- A comparison, a trend or a composition has to be shown and the encoding will determine what the reader concludes.
- Base sizes are small or uneven across categories, so the chart must carry its own qualification.
- The report or deck has more charts than findings, and something has to be cut on principle.
- Someone has proposed a chart type and it needs testing rather than accepting.
- The audience includes people with colour vision deficiency, or the output will be printed in greyscale or read on a phone.
- A qualitative structure (themes, a journey, a set of relationships) needs to be shown and there is no numerical chart that fits.

## 4. When NOT to use it

- **The finding is not settled.** Building a visual for a claim that is still moving means rebuilding it, and a chart made early tends to fix the framing of a finding before the analysis has finished arguing about it. Settle the claim first (Category 05, 06 or 07, then **08.01**).
- **The question is one of visual styling, template or brand.** Typography, palette definition, grid, template application and the visual system across a document are **12.02 Research Report Design**. **The boundary, stated in both directions: this skill decides what is encoded and how, and 12.02 decides how it looks within the house system.** A chart that is beautiful and wrongly encoded is this skill's failure; a correctly encoded chart in the wrong typeface is 12.02's.
- **The visual answers no question.** This is the most common case and the correct response is deletion, not improvement. Where a chart exists because a section needed something on the page, it is the visual form of padding, per K4 §1. Delete it and write the sentence.
- **A sentence is stronger.** Two numbers, a single proportion, or a comparison the reader will grasp instantly do not need a chart. "Two thirds of enterprise accounts renewed, against a third of small business accounts (n=412 and n=1,190)" is faster to read than any bar chart of the same fact. A report where every finding has a chart has stopped choosing.
- **The base is too small to be shown as a proportion.** Below 30, report counts or verbatim, never percentages, per K4 §7. A chart of percentages on a base of 22 is a precise-looking claim about nothing, and the chart form makes the false precision harder to notice than the number would have been.
- **The task is exploratory rather than communicative.** Exploratory plots made during analysis serve a different purpose and follow different rules: many of them, quickly, ugly, and not shown to anyone. They belong in **05.01 Descriptive Analysis**. Do not carry an exploratory plot into a report because it exists.
- **The purpose is decoration or density.** Where the brief is to make a page look substantial, or to include a visual because the client expects a certain number of them, say plainly what the visual would and would not show. A chart that answers no question is noise, and noise in a research document costs the reader's attention on the charts that matter.
- **The data is a full tabulation.** A data book, a cross-tab set or an appendix table exists to be looked up, not read. Table it. Charting a 40-row tabulation produces something neither readable nor lookup-able.

## 5. Required inputs

**Required.** Without these the skill cannot run. If absent, ask. If no answer is available and work must proceed, state the assumption at the point where it bites, per K5 §5.

- **The analytical question the visual answers**, in one sentence. This is the input the process actually runs on, and its absence is the reason most chart sets fail. Not "show Q14" but "is satisfaction different for customers who had a service issue".
- **The claim the visual must carry**, which becomes the chart title. If the claim cannot be written, the chart cannot be selected.
- **The data with its base description, base size and question reference**, per K2 §4.1. A chart without a base cannot be built to this standard.
- **Whether differences shown have been tested**, and at what threshold, per K4 §3.1. This determines whether the chart may imply a difference is real.
- **The output medium and viewing conditions**: printed, on screen, projected, on a phone, greyscale. Determines label sizes, how much a chart can carry, and whether colour can be relied on at all.

**Optional, and what each one adds.**

- **The narrative sequence** from **11.01**: lets a chart set be built as a sequence with consistent encodings, so the reader learns the visual grammar once rather than per page.
- **The full data rather than the summary table**: makes distribution and relationship visuals possible, which summary tables foreclose. Most reports show means because means are what was in the table, and the distribution frequently carries the finding.
- **Previous waves or previous reports**: fix the segment order, colour meanings and scales so a returning reader is not relearning the grammar, and allow genuine wave comparison where the questions match.
- **The client's accessibility requirements or public sector standards**: determine contrast ratios and whether colour-only encoding is permitted at all. Where none is stated, the defaults in step 8 apply anyway.
- **Statistical test results**: allow tested differences to be marked as tested and untested ones to be marked as observed, which is the difference between a chart that supports a claim and one that implies it.

## 6. Questions to ask before starting

1. **What question does this visual answer, and would the reader be able to state the answer after three seconds?** Determines the encoding and whether the chart should exist. *Default if unanswered:* do not build it. A chart with no question is deleted, per Section 4.
2. **Is the comparison the reader will make the comparison you intend?** Determines the encoding, because readers compare what is adjacent and aligned regardless of intent. *Default:* assume the reader compares whatever shares a baseline.
3. **Has the difference been tested?** Determines the annotation and the permitted language on the chart. *Default:* label as observed and untested, per K4 §3.1.
4. **What are the base sizes, and do they vary across the categories shown?** Determines whether percentages are permitted, what must be flagged, and whether categories are comparable at all. *Default:* show n for every category on the chart.
5. **How will this be viewed, and by whom?** Determines label size, colour reliance and information density. *Default:* build for greyscale print and for a reader with colour vision deficiency, which costs nothing and removes a category of failure.
6. **Is there an existing convention for segment order, colour meaning or scale in this report or previous waves?** Determines consistency, which is worth more than any individual chart improvement. *Default:* set the conventions once at the start of the chart set and apply them everywhere.
7. **How many charts is this section allowed?** Forces selection. *Default:* one chart per finding that needs one, and no chart for findings that read better as sentences.

## 7. Step-by-step methodology

**The founding rule.** Start from the analytical question, never from the data. A dataset does not imply a chart; a question does. The same table supports a dozen different charts making a dozen different claims, and choosing among them is the work.

**Step 1. Write the question and the claim.** In two sentences: the question the visual answers, and the finding it will show. The finding becomes the chart title, per step 9. Where the claim cannot be written, either the analysis is not finished or there is no finding, and in both cases the chart is premature. Where the claim is written and is dull, that is information: the chart may not be needed. *Correct result:* a question and a one-sentence claim, before any chart type is considered.

**Step 2. Decide whether a visual is the right form at all.** Three forms compete. A **sentence** wins for one or two numbers, or a comparison the reader grasps immediately. A **table** wins where the reader needs to look up values, where precision matters more than pattern, or where categories are numerous and unordered. A **chart** wins where the finding is a pattern: a shape, a difference, a trend, a distribution, a relationship, an outlier. The test: if the reader would have to read the values off the chart to get the point, the pattern is not the point and a table is more honest. *Correct result:* an explicit choice between sentence, table and chart, with the chart chosen because the finding is a pattern.

**Step 3. Map the question type to the encoding.** The mapping is not a menu of styles, it is a match between what the question asks and what visual variables human perception reads accurately. Position along a common scale is read most accurately, then length, then angle and area, then colour intensity. Encode the most important comparison in the most accurate channel available.

| Analytical question | Form | Why |
|---|---|---|
| How do these categories compare? | Bar chart, horizontal for long labels | Length on a common baseline, the most accurate comparison after position |
| How has this changed over time? | Line chart | Position encodes value and slope encodes rate, which is the finding |
| How is this distributed? | Histogram, or box plot for comparing distributions | Shows spread, skew and the cases a mean conceals |
| How do these two measures relate? | Scatter plot | Position on two scales; the pattern is the point |
| What is this made up of? | Stacked bar, or a single divided bar for one whole | Composition, with the caveat in step 3.1 |
| What is the order? | Ordered bar chart | Ranking is a comparison, and order is the finding |
| What is the sequence and where does it drop? | Flow or funnel | The finding is the transition, not the level |
| How do segments differ across many measures? | Small multiples, or a profile chart with an index line | Segment shape is read across panels, not within one crowded chart |
| How do two dimensions position options? | Matrix or two-by-two, with the axes defined and evidenced | Only where both axes are measured, not asserted |
| What themes came out of the qualitative? | Thematic framework: themes, prevalence, illustrative evidence | Prevalence as counts of participants, never percentages on small bases |
| How do these concepts relate? | Conceptual diagram | Where the relationship is the finding and no measurement encodes it |

**3.1 On stacked bars and composition.** A stacked bar is honest for one whole, and dishonest for comparing anything but the bottom segment across bars, because only the bottom segment shares a baseline. If the finding is about a middle segment across groups, break it out into its own bar chart. This is one of the four routine misleaders in step 10.

**3.2 On pie charts.** **Do not default to a pie chart.** Angle and area are read poorly, comparison between slices is unreliable, comparison between two pies is worse, and any pie with more than about four slices requires a legend, which defeats the purpose. The narrow cases where a pie is acceptable: two or three categories, a genuine part-to-whole relationship, where the finding is an approximate magnitude ("about half") rather than a comparison, and where the audience's convention expects it (budget or share-of-total contexts). Everything else is an ordered bar chart, which answers the same question better. Where the finding is a ranking, a pie is never correct.

*Correct result:* an encoding chosen from the question type, with the most important comparison in the most accurate perceptual channel.

**Step 4. Order the data by value, not by questionnaire.** Questionnaire order carries no information and actively hides the finding, because the reader has to reconstruct the ranking themselves. Order descending by value as the default. The exceptions are real and few: an inherent order (age bands, time, an ordered agreement or rating scale, stages of a journey), a fixed order carried from previous waves for comparability, or a grouping the reader navigates by (markets in a standard order). Where an inherent order applies, keep it, because reordering a scale destroys the shape that is the finding. State the ordering rule once and apply it across the chart set. *Correct result:* the reader can see the ranking without reading a single number.

**Step 5. Set the scale honestly.** The rule follows the encoding, not a blanket prohibition. **Bar charts start at zero, without exception**, because length encodes value and a truncated bar shows a false ratio. **Line charts may start above zero**, because position rather than length encodes value, and forcing a line to zero can flatten meaningful variation into nothing. A truncated line axis is defensible where the variation matters at the scale shown (an index moving between 78 and 82 is a real movement), the truncation is visible in the axis labels, and the full range is stated. It is not defensible where the truncation is what creates the appearance of change, which is the test to apply: **would the reader draw the same conclusion from the untruncated version?** If not, the truncation is doing the arguing. Hold scales constant across charts that will be compared, including small multiples, and never let a shared colour or axis change meaning between pages. *Correct result:* a scale decision that survives being shown next to its untruncated alternative.

**Step 6. Label directly and drop the legend.** A legend forces the reader to hold a colour-to-category mapping in memory and look back and forth. Put the label at the end of the line, next to the bar, on the segment. Where a legend is unavoidable, order it to match the data order rather than alphabetically. Values go on the chart where they are read (at the end of bars, at the ends and turning points of lines), not on every data point, which reintroduces the table. Axis labels state the unit and the measure, not the variable name from the dataset. *Correct result:* a chart readable without the reader's eye leaving the data.

**Step 7. Put the evidence on the chart.** Every chart carries its **question reference, base description and base size**, per K4 §4.3 and §7, in a note that stays with the chart when it is extracted into a deck or an email. Where bases vary across categories, show n per category, because uneven bases change what comparisons are legitimate. Flag any base below 100 on the chart and any base below 30 by not showing a percentage at all. Where a difference is tested, say what test and threshold; where it is not, mark it observed and untested. Where the data is weighted, say so and show the effective base. **A chart that has lost its base has lost its evidence status**, per K2 §7, and charts get extracted more often than any other element of a report. *Correct result:* a chart that could be forwarded on its own and still be checkable.

**Step 8. Use colour to encode meaning, and never colour alone.** Colour is a channel, not decoration, and it carries one of three meanings: **categorical** (different things, and the palette should not imply order), **sequential** (more of one thing, a single hue varying in lightness), or **highlight** (one series matters and the rest are context, in grey). Highlighting is the most under-used and most effective: a chart with one coloured bar and nine grey ones states the finding in the encoding. Fix colour meaning once across the report so that a colour means the same thing on every page. Then apply the accessibility rules, which are requirements rather than preferences: **do not rely on colour alone** to distinguish anything, so use direct labels, position or pattern as well; ensure adequate contrast between adjacent categories and against the background; check the chart in greyscale, which simulates the worst case and catches most failures; and avoid palettes that separate only in the red and green channels. Around one in twelve men has some form of colour vision deficiency, which in a board audience is a near certainty rather than an edge case. *Correct result:* a chart that survives being printed in black and white.

**Step 9. Annotate the finding on the chart.** The chart title states the finding, not the variable: "Satisfaction falls sharply after a customer's first service issue", not "Satisfaction by service issue experience". The title is a claim and it must be true of the chart and nothing more, so it carries the same discipline as a report headline: no causation the design does not license, no behaviour from stated preference, no change over time in a snapshot, per K4 §3. Add annotation where the reader would otherwise have to work: mark the point a change happens, note what is excluded, name the outlier. Where a chart needs three annotations to be understood, it is answering more than one question and should be split. **One visual, one idea.** *Correct result:* a reader who reads only the title and glances at the chart has the finding.

**Step 10. Strip to data-ink, then run the misleading-encoding check.** Remove everything that does not carry information: gridlines that are not needed for reading values, borders, backgrounds, shadows, redundant axis labels, tick marks. Every element removed makes the data more legible; the discipline is to remove until something is lost, then put that one thing back. Then check against the four routine misleaders, each of which is an encoding failure rather than a data failure:

- **Dual axes.** Two series on two independent scales, where the apparent relationship is created by the scaling choice. Any relationship can be manufactured or destroyed by moving either axis. Use two aligned charts, or index both series to a common base, and reserve the dual axis for the rare case where the two scales have a fixed, meaningful relationship.
- **Three-dimensional effects.** Perspective distorts the encoding, so the front of a 3D pie or the near end of a 3D bar reads larger. There is no research case for a 3D chart of two-dimensional data.
- **Area encoded by radius.** Where a circle or icon is sized to represent a value, doubling the radius quadruples the area, so the visual claims four times what the data says. Size by area, or use a bar.
- **Truncated stacked comparisons.** Comparing a middle or top segment across stacked bars, where only the bottom segment shares a baseline. Break the segment of interest out into its own chart.

*Correct result:* a chart set where every visual answers one question, carries its base, states its finding, and uses no encoding that claims more than the data.

## 8. Analytical framework

The selection chain, which is one decision worked through in order:

    Analytical question
      → Claim (which becomes the title)
        → Form (sentence, table or chart)
          → Encoding (matched to question type and perceptual accuracy)
            → Ordering, scale, labelling, base
              → Colour meaning and accessibility
                → Annotation
                  → Data-ink strip and misleading-encoding check

**Applying it.** The chain runs forward and is checked backward. Having built the chart, read it as a reader who does not know the answer and ask what question it appears to answer. If that is not the question at the top of the chain, the encoding has changed the claim, which is exactly what encodings do.

**The perceptual hierarchy, which is why the mapping is what it is.** Human judgement of visual quantity is most accurate for position along a common scale, then length, then angle and slope, then area, then colour intensity, then volume. Put the comparison that carries the finding as high in that list as the data allows. This single fact explains most of the guidance in step 3: it is why bars beat pies for comparison, why scatter beats bubble for relationship, why 3D always loses, and why radius-sized circles mislead.

**The three-second test.** A reader glances at the chart for three seconds and looks away. What do they now believe? If it is the finding, the chart works. If it is a different claim, the encoding is arguing with the title. If it is nothing, the chart is a table in disguise. Run this on every chart before it goes in, ideally with a colleague who does not know the answer.

## 9. Output format

**Every chart carries, without exception:**

- **A title stating the finding**, in a claim form that would survive being quoted alone.
- **The visual itself**, ordered by value unless an inherent order applies, on an honest scale, directly labelled.
- **A base note**: question reference, base description, base size, weighting status, and the test and threshold where a difference is claimed.
- **Flags where required by K4 §7**: small base, untested difference, non-probability sample, excluded cases, client-supplied data.
- **Annotation** where the reader would otherwise have to reconstruct the point.

**The chart plan**, produced before building a set:

| Ref | Analytical question | Claim (title) | Form | Encoding | Source and base | Order rule | Notes |
|---|---|---|---|---|---|---|---|

The plan's value is that it makes the deletions visible: a row with a question but no claim, or a claim that reads better as a sentence, gets removed before anyone spends time building it.

**Conventions to fix once across a chart set:** colour meanings, segment order, scale ranges for comparable charts, the base note format, the rounding rule, and how tested and untested differences are marked.

**When the evidence is thin.** The chart set contracts rather than filling. A finding on a base under 30 is reported as counts or verbatim, not charted as percentages, per K4 §7. A section with one finding gets one chart, and the space that would have held three is used for the finding's qualification. Where a comparison is untested and the difference is small, the chart says so on the page rather than letting adjacency imply a difference. Where a visual would be built only because the section looks empty, it is not built, per K4 §1. A report with six charts that each answer a question is stronger than one with twenty that display data, and it is also faster to QA.

## 10. Quality checks

Run on every chart before it is presented. Sits on top of K4 §8.

1. Does this chart answer a stated analytical question? If not, it is deleted.
2. Is the finding readable from the chart alone in about three seconds?
3. Does the title state a claim rather than name the variables, and is the claim true of this chart and nothing more?
4. Does the title claim causation, behaviour from stated preference, change over time, or reach beyond the sample?
5. Is the encoding matched to the question type, with the key comparison in the most accurate available channel?
6. Are the data ordered by value, or by a stated inherent order?
7. Do bar charts start at zero, and is any truncated axis visible, stated, and defensible against its untruncated version?
8. Do the base description, base size and question reference appear on the chart itself?
9. Are bases shown per category where they vary, and are small bases flagged per K4 §7?
10. Is any difference implied to be real without a test behind it, per K4 §3.1?
11. Is meaning encoded by anything other than colour alone, and does the chart survive greyscale?
12. Does each colour mean the same thing on every chart in the set?
13. Is the reader required to consult a legend to read the chart?
14. Does the chart use a dual axis, a 3D effect, area sized by radius, or a cross-bar comparison of a non-baseline stacked segment?
15. Would removing any remaining chart element lose information?
16. Are scales held constant across charts the reader will compare?

## 11. Common failure modes

| Failure | How to recognise it | How to prevent it |
|---|---|---|
| **Charting the table** | The chart shows what was in the analysis output, with no question behind it | Step 1. Write the question and the claim before choosing a form |
| **Questionnaire order** | Bars in the order the options appeared in the instrument | Order by value unless an inherent order applies. Step 4 |
| **The default pie** | A pie chart with five or more slices and a legend | Ordered bar chart. Pies only in the narrow cases in step 3.2 |
| **Truncated bars** | A bar axis starting at 40%, making a 6-point gap look like a chasm | Bars always start at zero. Length encodes value |
| **The manufactured relationship** | Two series on two axes, tracking each other suspiciously well | Dual axes create the relationship. Index both to a common base, or use two charts |
| **The vanished base** | A percentage chart with no n anywhere on it | Base note on every chart, per K4 §4.3. Charts get extracted |
| **False precision on a small base** | Percentages charted on a base of 22 | Counts or verbatim below 30, per K4 §7 |
| **Adjacency implies significance** | Two bars side by side, no test, and a title implying a real difference | Mark untested differences as observed, per K4 §3.1 |
| **Colour as decoration** | A palette applied per category with no meaning, or a rainbow on an ordered scale | Colour encodes categorical, sequential or highlight. Fix meaning across the set |
| **Colour-only encoding** | Two series distinguishable only by hue, indistinguishable in greyscale | Direct labels, position or pattern as well as colour. Test in greyscale |
| **Legend hunting** | The reader's eye moves between chart and legend to read anything | Direct labelling. Step 6 |
| **The stacked-segment comparison** | A claim about a middle segment across several stacked bars | Break that segment into its own chart |
| **The decorative visual** (the signature AI failure) | An attractive, novel or dense visual that answers no question | Every visual answers a question or is deleted, per K4 §1 |
| **Chart inflation** | More charts than findings; several charts per page | One chart per finding that needs one. Findings that read better as sentences get sentences |
| **The label title** | "Satisfaction by segment" above a chart whose finding is a 20-point gap | Titles are claims. Step 9 |

## 12. AI guardrails

Skill-specific only. Universal prohibitions are inherited from K4. Base disclosure follows K4 §7; traceability follows K2 §4.1.

1. **Never build a visual without a stated analytical question and a claim.** A chart generated from a table has no argument, and the encoding will supply one at random.
2. **Never produce a chart without its base description, base size and question reference on the chart itself.** Charts are the most extracted element of any report, and the base does not travel unless it is attached.
3. **Never truncate a bar axis.** Never truncate a line axis in a way that creates the appearance of change that the full range would not support, and always show the range where truncation is used.
4. **Never chart percentages on a base below 30**, per K4 §7, and never chart a comparison between categories whose bases differ materially without showing both.
5. **Never imply a difference is real through adjacency, annotation or title language where no test was run**, per K4 §3.1.
6. **Never use a dual axis, a 3D effect, or area sized by radius.** Each asserts a relationship or a magnitude the data does not contain.
7. **Never rely on colour alone to carry meaning**, and never let a colour change meaning between charts in the same document.
8. **Never write a chart title that claims more than the chart shows**, including causation, trend from a single wave, or a population the sample does not cover.
9. **Never build a visual to fill a page.** Where a section needs a visual and the evidence supports none, the section is shorter, per K4 §1.
10. **Never invent or interpolate a data point to complete a series or smooth a line.** A gap in a time series is shown as a gap, with the reason stated.
11. **Never present a conceptual diagram in a way that implies it is measured.** Diagrams of relationships are labelled as the analyst's model, per K2 §3.3, and never carry the visual grammar of data.

## 13. Best-practice principles

1. **Start from the question, never from the data.** A dataset does not imply a chart. The same table supports a dozen charts making a dozen different claims, and choosing between them is the entire job.
2. **Every visual answers a question, and one that answers none is deleted.** This single rule removes more bad charts than every design improvement combined.
3. **The title is the finding.** A chart whose title names the variables has left the interpretation to the reader, who will do it less well and less consistently than you would.
4. **One visual, one idea.** A chart carrying two findings has one that is being ignored, and the reader does not know which.
5. **Order by value.** Questionnaire order is a fact about your instrument. The reader is entitled to see the ranking without doing arithmetic.
6. **Encode in the most accurate channel the data allows.** Position, then length, then angle, then area, then colour intensity. Most chart errors are a decision to encode something important in a channel low on that list.
7. **The base belongs on the chart.** A chart travels further than any other part of a report, and it travels alone.
8. **Highlight rather than colour everything.** One coloured series against grey context states the finding in the encoding, and it is the most under-used technique in research visualisation.
9. **Design for greyscale and for colour vision deficiency by default.** It costs nothing, it improves the chart for everyone, and in any audience of size the alternative fails somebody.
10. **Remove until something is lost, then put that back.** Data-ink discipline is subtractive, and the instinct to add is almost always wrong.
11. **Consistency across the set beats optimisation of any single chart.** A reader who learns the grammar once reads the whole report faster; a reader relearning it per page reads none of it.
12. **A chart that needs explaining is a chart that failed.** If the caption has to teach the reader how to read the encoding, choose a different encoding.

## 14. Worked example

*Fictional scenario, used for illustration only. The organisation, findings, figures and quotes below are invented for the purpose of demonstrating method.*

**INPUT.** A software company has run a study into why users abandon its onboarding flow. The analysis is complete: a behavioural funnel from product logs, a survey of 890 users who started onboarding, and 18 usability sessions. The team has produced 22 charts for the report. The task is to build the chart set properly.

**PROCESS.**

*Step 1, and the cull.* Each of the 22 charts is given a question and a claim. Nine have neither: they exist because the analyst had a table. Four more have a claim that reads better as a sentence, including a two-category split that occupied a full pie chart. The set falls to nine before any chart is built, and the chart plan records why each was dropped, which turns out to matter when a stakeholder asks about one of them.

*Step 3, and the first judgement call.* The headline chart is the funnel: where users drop out across six onboarding steps. The team's version is a stacked bar of completion by step. It is wrong twice: the finding is the transition rather than the level, and the stacking invites comparison of segments that do not share a baseline. It becomes a flow chart of the six steps with the drop at each transition labelled directly, and the single step where 41% of the loss occurs is the only coloured element on an otherwise grey chart. The encoding now states the finding without the reader reading a number.

*Step 4.* A chart of self-reported reasons for abandoning sits in questionnaire order, so the largest reason appears fourth. Reordered descending, the ranking is visible instantly. The five-point confidence rating question keeps its inherent order, because reordering it would destroy the shape, which is the finding.

*Step 5, and the second judgement call.* A chart of weekly completion rate over twelve weeks shows movement between 61% and 68%. The team's version starts the axis at 60%, which makes a 7-point range fill the chart. The untruncated version shows a nearly flat line. The test is applied: would the reader draw the same conclusion from the untruncated version? No, and the movement is within the range that weekly base sizes (n between 55 and 90) would produce by chance in any case. The chart is rebuilt from zero, the flatness is the finding, and the title becomes "Completion has not moved since the redesign", with base sizes per week shown and the variation described as observed and untested.

*Step 7.* Three charts split by account type carry bases of 610, 198 and 82. The 82 is flagged on the chart and the associated title is softened from a comparative claim to a directional one. A fourth split, on a base of 24, is removed from the chart entirely and reported as counts in the text, per K4 §7.

*Step 8.* The palette uses red and green to distinguish completed from abandoned. In greyscale they are near-identical, and the report will be printed. The encoding changes to a single hue with direct labelling, and completion is marked by position rather than colour.

*Step 9.* Titles are rewritten from labels to claims. "Onboarding completion by step" becomes "Two in five users who abandon do so at the workspace setup step". The claim is checked against what the chart shows and nothing more: an earlier draft read "workspace setup is driving abandonment", which is causal language the funnel data does not license, per K4 §3.2.

**OUTPUT.** Nine charts, each with a question, a claim as its title, a base note carrying the reference and n, direct labelling, no legend, one highlight colour against grey, an honest scale, and greyscale-safe encoding. A chart plan recording the thirteen removals with reasons. Two findings converted to sentences. One finding on a base of 24 reported as counts. Total build time lower than the 22-chart set it replaced.

## 15. Advanced usage

**Small multiples.** Where a pattern must be compared across many segments, markets or time periods, a grid of small identical charts outperforms one crowded chart or a series of separate pages. The conditions: identical scales across every panel (otherwise the comparison is false), identical encoding, panels ordered by a meaningful variable rather than alphabetically, and one shared annotation rather than one per panel. The reader compares shapes, which is a perceptually cheap operation, and outliers appear without being pointed at.

**Visualising qualitative structure.** Themes, journeys and relationships have no numeric encoding, and the failure mode is borrowing the grammar of data charts, which implies measurement that does not exist. Use a thematic framework showing themes with participant counts (never percentages on qualitative bases), illustrative evidence attached, and prevalence stated as counts per K2 §4.2. Use a journey or process diagram where the sequence is the finding. Label any conceptual diagram as the analyst's model, per K2 §3.3, so a reader cannot mistake a synthesis for a measurement.

**Uncertainty on the chart.** Where a probability sample supports it, show confidence intervals rather than point estimates for comparisons that matter, because overlapping intervals communicate the uncertainty a reader would otherwise ignore. Do not attach intervals to non-probability samples, per K3 §6, where the number would be calculable and meaningless. For small bases, the more honest device is often to show counts as dots rather than a percentage bar, which makes the base visible in the encoding itself.

**Charts that will be extracted and reused.** Assume every chart will appear on a slide, in an email and in someone else's deck, without its surrounding text. Build each one to be self-sufficient: the finding in the title, the base in the note, the caveat on the chart rather than in the paragraph beside it, per K4 §4.3. The test is to look at the chart with everything else covered and ask what a reader would conclude.

**When the standard approach does not fit.** Some findings genuinely have no good chart: an interaction across three variables, a mechanism, a set of conditional relationships. Resist the novel visualisation, which requires the reader to learn a grammar for one use. Write the sentence, or use a small table, or split the finding into two charts that each answer one question. Novelty in research visualisation costs comprehension, and the cost is paid by every reader.

## 16. Skill chain

**Recommended previous skills:**
- **05.01 Descriptive Analysis** and **05.03 Cross-Tabulation.** Hand over the figures with their bases, conventions and test results attached, which is what a chart needs to be built to this standard.
- **08.01 Finding to Insight Development.** Hands over settled findings, so charts are built for claims rather than for tables.
- **11.01 Research Narrative Development.** Hands over the sequence, which lets a chart set be built as a consistent visual argument rather than page by page.
- **07.01 Thematic Analysis.** Hands over themes with prevalence, which is what a qualitative framework visual encodes.

**Recommended next skills:**
- **12.02 Research Report Design.** Takes correctly encoded charts and applies the visual system: typography, palette, grid and template. It does not change what is encoded.
- **11.05 Research Presentation Development.** Takes charts built for a report and adapts density and labelling for a projected or read-alone slide, which are different constraints.
- **12.03 Research Report Compilation.** Takes the chart set and holds it consistent with the report's conventions, bases and claims.
- **12.06 Research Report QA.** Checks the finished charts against their sources and against the misleading-encoding list.

**Runs well alongside:**
- **11.03 Research Report Writing**, because the decision between a sentence and a chart is made jointly, and a chart that fails is often a sentence that was never tried.
- **13.04 Bias Detection**, run against a chart set for framing introduced by encoding, ordering and scale choices.
- **K4 §7**, which sets the base and disclosure requirements every chart carries, and **K3 §4.4**, which governs the precision a chart may display.

---
A Yazi Supplied Skill and resource.
