---
title: "BK Blueprint Glossary — AI, Automation and Verification"
date-created: 2026-09-18
date-modified: 2026-10-03
note-type: literature
status: growing
tags: [type/literature, status/growing, domain/ai, domain/consulting, domain/brand, context/freelance]
aliases: ["public glossary", "bkbp glossary", "published glossary"]
related: ["[[bkfr_mattos_master_glossary]]", "[[bkbp_ai-context-dictionary]]"]
source: "GENERATED from _meta/bkfr_glossary.json — do not hand-edit"
---

> [!warning] Generated file
> Rendered from `_meta/bkfr_glossary.json` by the `glossary` command. Hand edits are overwritten. To change wording, edit `definition_public` in the source and regenerate.

# BK Blueprint Glossary

Plain-English definitions for the vocabulary of AI systems, automation, and operational verification. Written for leaders evaluating or deploying AI in their organization, not for engineers.

Every term leads with a one-or-two-sentence plain-language definition. Where a term has more depth worth having, it follows underneath. If a definition here needs you to already know another term to understand it, that is a defect: tell us.

71 terms. Last updated 2026-10-03.

## Git and version control

### Merge

Adding a reviewed set of changes into the shared version of a project.

Adding a proposed set of changes into the shared version of a project. Teams commonly review a pull request before merging it. A merge changes the project record; delivering that change to a running app or another device is a separate step.

### Pull request (PR)

A proposal to add changes to a project. People can review the differences and run checks before accepting them.

A proposal to add changes to a project. It shows what changed, keeps review discussion together, and can run automated checks. Accepting the proposal means merging it into the shared version; opening a pull request alone does not change that version.

## AI, agents and automation

### Agent

An AI given a goal and some tools and allowed to work out its own steps, instead of following a fixed script.

An AI system given a goal, a set of tools, and permission to decide its own steps, rather than following a fixed script. Agents are more capable than scripted automation and correspondingly harder to supervise, because the path they take is not known in advance. The practical requirement is not smarter agents but clearer boundaries: defined scope, defined tools, and a defined point where a human decides.

### AI-assisted software development

Building software with AI help while people still decide what it should do, review the changes, and check that they work.

An AI assistant can draft code, suggest fixes, or explain unfamiliar parts of a system. The team still sets the requirements, reviews the result, tests the behavior, and decides what to release. The quality of that review matters more than whether a person or an AI typed the first draft.

### Belt-and-suspenders (b-a-s)

Deliberately attaching a file to a project even though the AI could already find it by searching, to guarantee it loads every time rather than only when the AI thinks to look. Worth doing for a few genuinely important files, not for everything.

Manually attaching a document to an AI project or workspace even though the AI already has broader access that would let it find that same document on its own. The redundancy is deliberate: it guarantees a specific piece of context loads automatically at the start of every session, rather than depending on the AI to retrieve it when relevant. Worth reserving for a small set of genuinely load-bearing documents, since doing it for everything defeats the purpose of having broad access at all.

### C2PA / Content Credentials

A tamper-evident label attached to a file recording where it came from and what was done to it.

An open standard for attaching tamper-evident provenance to a media file, recording where it came from and what has been done to it since. Distinct from an invisible watermark: the watermark lives inside the content itself, the credential lives alongside it as signed metadata. Both are becoming default on generated media, and both are worth understanding before assuming a generated asset can pass as a photograph.

### Capture journal

A saved record of draft edits that can be recovered after a restart. Saving a draft does not approve it.

A saved record of draft edits that can be recovered after a restart. It helps an automated system preserve work before submitting it for review. The record shows what was captured; it does not establish that the draft was approved or delivered.

### Connector

An approved link between the AI and one of your accounts. Limited in what it can reach, revocable any time, and set per conversation, which is why one session can see something another cannot.

An authorized link between an AI system and an external account. Connectors are scoped and revocable, which is what makes them safe, and their scope is frequently narrower than people assume. A common source of confusion is an AI reporting that something does not exist when it simply falls outside what that particular connection was authorized to see.

### Context window

How much the AI can hold in its head at once. When it fills up, earlier details get fuzzy, which is why long conversations start to drift.

How much information an AI system can hold in working memory at once. When it fills, earlier detail degrades, which is why long sessions drift, repeat themselves, or lose established decisions. Practical implication: durable decisions belong in a document, not in a conversation.

### Cron job

A task set to run at a fixed time or on repeat, written in a compact five-part time code: minute, hour, day, month, weekday. Named after the decades-old scheduling tool that invented the format. Every scheduled task is one of these underneath.

A task set to run automatically at a fixed time or on a recurring schedule, defined by a compact five-part time pattern, minute, hour, day of month, month, day of week, known as a cron expression, for example one meaning "every day at 6:45 AM." The name comes from `cron`, the decades-old Unix scheduler this pattern originated in. It's now the standard way automated systems, including AI scheduled tasks, express recurring timing.

### Gem (Google Gemini)

Google's version of a saved AI setup, with a name, instructions, and reference files that can sync from Google Drive. The least independent of the major versions, closer to a saved prompt than to something that takes its own steps.

Google's version of a saved, custom-configured AI assistant inside Gemini, built from a name, instructions, and reference material. It serves the same basic purpose as OpenAI's custom GPTs: save a configuration once instead of re-explaining context every conversation. Of the major platforms' equivalents, it is generally the lightest-weight, closer to a saved prompt than to an autonomous agent that takes action on its own.

### Generally available vs preview

Whether a feature is finished and supported, or still being tested and liable to change without warning.

Whether a vendor considers a feature finished and supported, or still under test. The distinction is easy to skim past because both are available on the same account and often cost the same. Preview endpoints can change shape without notice, and a vendor's commercial protections, including copyright indemnity, usually do not extend to them. Anything load-bearing should be checked for which side of that line it sits on.

### Generative media

AI that produces pictures, video, music or speech rather than words.

AI that produces pictures, video, music or speech rather than text. It behaves differently to a chat model in three ways that matter commercially: it is priced per second, per image or per generation rather than per word, it often takes minutes rather than seconds to return, and its licensing terms are frequently stricter and less uniform than those covering text.

### GPT (OpenAI custom GPT)

OpenAI's version of a saved AI setup: custom instructions, uploaded reference files, and a chosen set of tools. Closer to a saved workspace than to something that acts on its own. OpenAI is actively renaming and reworking this, so check before relying on the term.

OpenAI's version of a saved, custom-configured AI assistant: built from custom instructions, uploaded reference material, and a chosen set of tools, layered on top of ChatGPT. It is closer to a saved project setup than to a fully autonomous agent, since it is mostly a persona and a knowledge scope rather than something that takes multi-step action on its own. OpenAI has begun evolving this concept toward more autonomous agent products, which is the direction the whole industry is heading.

### Grep/Glob

Two search tools that work together so an AI can have access to a huge folder without reading all of it. One finds files by name, the other searches inside them. Access and actual reading are separate, so nothing gets pulled in until a search says it is relevant.

The two-tool pattern behind targeted retrieval. One tool finds files by name or pattern; the other searches inside file contents for a match. Together they let an AI system have broad access to a folder of information without pulling all of it into working memory at once. This is the mechanism that makes wide access and a manageable context window compatible: an AI can reach everything in a connected folder while only the specific files relevant to the question at hand are actually read.

### Groups (Claude)

A way for a Claude org admin to bundle people into teams or departments, then set spending limits, permissions, and shared access for that whole bundle at once instead of person by person. It organizes who's in the org and what they can touch, not the work itself. Projects is still the workspace for that.

Groups is an organizational feature on Claude's Enterprise plan that lets an administrator bundle members by team or department, then manage spend limits, permissions, and access to shared resources for the whole group at once rather than person by person. It sits above Projects, which organizes a body of work — Groups is the layer that organizes people and what they're allowed to touch. Worth knowing before rolling AI access out past a handful of users.

### Hallucination

The AI stating something confidently that is not based on any real source. The reason you verify rather than trust.

Output stated confidently that is not grounded in any real source. The risk is not that AI systems are wrong sometimes, it is that wrong output arrives with the same fluency and confidence as correct output. This is why verification steps and source citation matter more than model quality in most business deployments.

### Human in the loop

A required human approval before something irreversible happens. Deliberately kept on anything that sends, spends, or deletes.

A required human approval before an action that is difficult or impossible to reverse. The practical rule is to place approval gates around anything that sends, spends, publishes, or deletes, and to let everything else run unattended. The goal is not to supervise AI constantly, it is to be deliberate about which decisions stay human.

### Idempotency

Being able to repeat an operation without doing its effect twice. For example, retrying a submission should return its first receipt rather than create a second submission.

Being able to repeat an operation without doing its effect twice. This matters when an automated system loses the response and cannot tell whether an action succeeded. A reliable retry recognizes the earlier action instead of creating a duplicate.

### Image-to-video

Giving a video generator a still picture to move, instead of describing the scene from scratch.

Giving a video generation model an existing still image to animate, rather than describing a scene from scratch and letting the model invent it. This matters more than it sounds in commercial work. Text-to-video produces something new every time, which means a client is approving a fresh subject with each attempt. Image-to-video anchors the motion to a frame already signed off, so what moves is the thing everyone agreed on.

### Instruction layer

Where your standing rules for the AI live. Some apply everywhere, some only inside one project, and some are only read when something asks for them. Only the everywhere ones are guaranteed to load.

Where standing rules for an AI system live. Most platforms have several layers: global instructions that apply everywhere, project or workspace instructions with narrower scope, and reference material loaded only on demand. Knowing which layer is guaranteed to load is essential, because a rule written into a layer that does not load is documented rather than active, and it will appear to be working right up until it matters.

### MCP (Model Context Protocol)

The common standard that lets an AI connect to outside systems like your email or your accounting software. Without it, every connection would have to be custom built.

An open standard that lets AI systems connect to external tools and data sources such as email, calendars, file storage, and business applications. It matters because it turns an AI from something that talks about your work into something that can act on it. It also means access, permissions, and scope become real operational concerns rather than theoretical ones.

### OCR (optical character recognition)

Turning a picture of text into text a computer can read. It is where character mistakes come from, like reading a zero as a letter O.

Optical character recognition turns a picture of text, such as a scan, into characters a computer can search and process. It saves retyping, but lookalike characters and poor scans can produce errors. Verify names, amounts, and exact reference numbers against the original page.

### Parallel session

A second AI conversation running at the same time. They cannot see each other, so both can change the same thing without either knowing.

Two or more AI sessions working at the same time on related material. Sessions do not share state, so each can act on the same resource without knowing the other exists, including undoing each other's work. As organizations deploy more agents, this becomes a foundational design question rather than an edge case: what is the shared source of truth, and how does one agent learn what another already changed?

### Parse (document parsing)

Converting a document into structured text a computer can work with. It is a translation, not a copy, which means something can always be lost in it.

Document parsing converts a file into organized text or data that software can work with. It is a translation of the original, so tables, labels, or whole passages can change or disappear. Compare important facts and completeness against the source before relying on the result.

### Parsing tier

The quality-versus-price setting on a document parsing service. Paying more is not automatically safer: a higher tier can be better at reading characters and worse at deciding what counts as content.

A parsing tier is a price and capability level offered by a document-processing service. Paying for a higher tier may improve one kind of extraction without improving every kind. Recheck both accuracy and completeness whenever you change methods or tiers.

### Plugin

A bundle that installs several skills, helpers, and connections as one package. The plugin is the box. A skill is one thing inside it.

A distributable bundle that packages skills, specialized agents, and connections to outside systems together so they can be installed as a single unit. A plugin is the package; a skill is one instruction set inside it. Plugins are how AI capability spreads in practice: rather than building a capability from scratch, an organization installs a plugin someone else built and gets its skills, agents, and connections all at once.

### Polling

Repeatedly asking a service whether a long job has finished, instead of waiting to be told. It is what sits behind most progress bars.

Polling means asking a service at intervals whether a job has finished. It is useful when a task takes longer than one request. Set a reasonable interval and a deadline so a stuck job does not wait forever or send unnecessary requests.

### Provider adapter

A thin translation layer so your tools ask for a capability rather than naming a specific AI company, making vendors swappable.

A design pattern for building on AI without being married to one supplier. Rather than writing code that calls a named vendor, you define the capability you want and put a thin translation layer behind it for each provider. Swapping or adding a supplier then means writing one small adapter instead of rewriting everything that depends on it. Given how fast models change and how often pricing and terms move with them, this is closer to basic hygiene than to over-engineering.

### Revision ledger

A record of which saved version came before another and which version has been confirmed at each step. It helps stop an older update from arriving out of order.

A record that tracks the order of saved versions and which one has been confirmed at each step. It helps an automated process reject an older update that arrives late. A ledger is evidence of what a process recorded; it still needs reliable checks of the actual files and approval decisions.

### Scheduled task

Instructions that run automatically on a clock, whether or not anyone is watching.

An instruction set that runs automatically on a clock, whether or not anyone is present. Scheduled tasks are where automation delivers compounding value and also where it fails most quietly, because nobody is watching at the moment they run. Any scheduled task worth having is worth monitoring for output rather than execution.

### Seed message

A short block of the essential facts from one conversation, pasted as the first message of the next one so it starts up to speed instead of from nothing. Deliberately short. Only what the new conversation cannot work without.

A short, paste-ready block of the essential facts from a finished session or project, written so a person can drop it into a brand-new AI conversation and have the assistant pick up right where things left off, without re-explaining everything from scratch. The same idea shows up across the AI industry under different names: some call it a 'seed,' others describe it as part of 'context engineering' or 'context rehydration.' Whatever the name, the practice is the same: rather than carrying a full conversation history forward (which eventually overflows or degrades), carry forward only the load-bearing facts a fresh conversation actually needs.

### Silent drift

An automation still following an old version of the rules while the written rules have moved on. Nothing breaks and nothing errors. The two just quietly stop matching.

An automated process still faithfully following instructions that no longer match the document meant to govern it. Nothing errors and nothing alerts. The procedure and the practice simply diverge over time, and the gap is usually discovered by accident, often long after it started causing damage. Preventing it requires periodically checking the live system against its documentation rather than assuming they match.

### Skill

A saved set of instructions the AI pulls up when a matching situation comes along. It defines how to do a type of task. It does not run on its own.

A reusable set of instructions an AI system loads when a matching situation appears. A skill defines how a particular kind of task should be done, encoding standards and procedure so the same work is performed consistently rather than improvised each time. Skills do not run on their own; they shape behavior when relevant work arrives.

### Subagent

A helper AI sent off to handle one specific piece of work and report back a single answer. It keeps its searching and dead ends to itself, so the main conversation does not fill up with the mess.

A specialized agent that a primary agent creates to handle one bounded piece of work, with its own working memory and a narrower set of tools, before reporting its result back and closing. Subagents exist mainly to protect the parent agent's own working memory: the exploratory searching, reading, and trial-and-error involved in a subtask stay contained inside the subagent rather than crowding out everything else the parent needs to remember. A useful mental model is a specialist a manager delegates a narrow task to, rather than doing the research personally.

### SynthID

An invisible marker baked into AI-generated images, audio and video that survives editing and identifies the file as machine-made.

An invisible watermark embedded into AI-generated images, audio and video at the moment they are created. It is designed to survive cropping, compression and re-encoding, so the file stays identifiable as machine-generated long after it leaves the tool that made it. Worth knowing before generated material goes into anything public: the marker travels with the file, and treating generated work as indistinguishable from captured work is a bet against detection improving.

### Token

The unit an AI reads and charges in. Roughly a word-fragment, about four characters. Images cost tokens too, which is why a scanned page costs far more to read than the same page as plain text.

A token is a small unit an AI system processes and often uses for usage limits or billing. It can be a word, part of a word, or another piece of input; the exact count depends on the model and format. Pictures and other media may also add to usage, so the same information can cost different amounts to process in different forms.

### Vibe coding

Telling AI what software you want and accepting the code mainly by whether it seems to work, without closely reviewing the code.

Vibe coding usually means describing a program to AI, trying what it produces, and continuing by feel without closely reading the code. It can be a fast way to explore an idea. For software people rely on, reviewing the code, testing its behavior, and controlling releases remain separate responsibilities.

## Client delivery and privacy

### Allowlist

A named list of the things allowed through a boundary. Anything not on the list stays out.

An allowlist starts with a closed door and names the items that may pass through it. For a document transfer, that might mean naming each approved file rather than copying an entire folder. The list still needs a trusted owner and a check that the transfer follows it.

### Model output indemnity

A vendor's written promise to defend you if someone claims the thing their AI generated for you infringes their copyright.

A vendor's contractual promise to defend a customer against third-party copyright claims arising from output their model generated. The detail that catches people out is that it is granted model by model rather than vendor by vendor, and usually only for generally available versions rather than preview ones. Two models from the same company, billed to the same account, can sit on opposite sides of that line. Anyone putting generated material into commercial work should check which specific models are covered rather than assuming the vendor covers all of them.

### Pre-download isolation

Keeping private files out of a cloud-connected collection before a cloud computer can download any of it.

If a cloud service should see only selected documents, make a separate collection containing only those approved documents before connecting it. Removing private documents after the service downloads a larger collection does not undo that exposure.

### Scratch asset

Something made to help decide, not to hand over. It exists to test an idea and then gets replaced.

Material made to help a decision rather than to be delivered. A rough music bed used to test whether a film wants something warm or something spare is a scratch asset; the licensed track that eventually sits under it is a delivery asset. The distinction is worth making explicit and deciding once, because generated material is convincing enough that the line blurs under deadline, and the cost of blurring it falls on whoever handed the work over.

### Sensitivity classification

Deciding how private a piece of information is and who is allowed to see it.

Sensitivity classification answers a practical question: who may see this exact information? A folder name can be a clue, but individual documents may need different treatment. The decision should come from an authorized reviewer and be checked before sharing.

## Verification and diagnostics

### Anchor check

A five-point spot check for converted documents: a key total, a value from the hardest table, a chart label, something from page one, and the very last line of real content. All five must match the original. The last one matters most, because it catches content that was quietly cut off.

After converting a document, compare a few known points with the original: an important number, a difficult table entry, a figure label, something near the beginning, and the last meaningful line. Checking the end helps reveal text that was quietly cut off. A spot check is useful, but it should be paired with a check for missing material between those points.

### Cache-buster

Adding a scrap of random text to the end of a web address to force a fresh copy, so you know you are not looking at a saved older version.

Adding a unique value to a request to force a fresh response rather than a stored copy. Used when diagnosing whether you are looking at current state or at something a cache is holding onto. A useful habit before concluding that a fix did not work.

### Content hash

A digital fingerprint calculated from a file. A changed fingerprint tells you the contents changed.

A digital fingerprint calculated from a file. Comparing fingerprints helps check whether two copies have the same contents or whether a file changed. A fingerprint alone does not tell you who created the file or whether its contents are trustworthy.

### Content type

What the server actually sent back: an image, a web page, a file. More useful than the status code, because a server can report success while sending you the wrong thing.

The declared format of what a server actually sent, for example an image or a web page. When you are verifying that something worked, content type is the more meaningful check, because it reveals cases where the request succeeded but the response was the wrong thing.

### Control test

Running a case you already know the answer to alongside the real one, so when something fails you can tell whether the thing is broken or your test is.

Running a case with a known outcome alongside the case you are actually investigating, so that a failure is distinguishable from a broken test. Without a control, an inconclusive result and a negative result look identical. It is a basic experimental discipline that is routinely skipped in technical troubleshooting, and skipping it is how teams confidently reach the wrong conclusion.

### Coverage vs accuracy

Two different questions that get treated as one. Accuracy asks whether the things you checked are correct. Coverage asks whether anything went missing between them. Spot checks answer the first and quietly assume the second.

Accuracy asks whether the parts you checked are correct. Coverage asks whether all the needed parts are present. A document conversion can pass several spot checks and still omit entire sections, so a good review checks both correctness and completeness.

### Endpoint over docs

When the documentation and the live service disagree about how to ask, believe the service and test it cheaply.

A working rule for integrating with fast-moving services: where published documentation and the live endpoint disagree, believe the endpoint. Documentation on rapidly shipping products is routinely behind the deployed reality. The cheapest way through is to send the smallest possible request and read the rejection carefully, because a well-built API usually lists its accepted values in the error. That makes a failed call a better specification than the guide, and it turns an afternoon of guessing into a few minutes of probing.

### False-positive 200

The server reports success and then sends the wrong thing entirely, like an error page where a picture should be. Green light, wrong result. Check what actually came back, not just whether something did.

A request that returns a success code while delivering the wrong thing entirely. In web systems, a server can answer 200 OK while actually serving an error page, a login redirect, or a placeholder. Anything checking only the status code sees success. This is the concrete, everyday version of a broader principle: a green signal is evidence that something responded, not evidence that it responded correctly. Always verify the content, not just the code.

### Free-baseline diff

Run the free version of a tool first, then compare simple counts against the paid version's output. It costs nothing and catches missing content that spot checks cannot see.

Run a lower-cost or free method as a comparison point before paying for a more advanced one. Compare simple signals such as page, word, or number counts to spot unexplained omissions. A difference is a reason to investigate, not automatic proof that either result is wrong.

### Ground truth

The authoritative original you compare against. Checking one copy against another copy proves nothing, since both can be wrong the same way.

Ground truth is the original evidence you use to judge a result. For a converted document, that usually means checking the visible source page rather than comparing two converted copies. Two copies can agree with each other and still share the same mistake.

### Homoglyph

Characters that look identical but are not, like a zero and a capital O. Harmless in a sentence, where the surrounding words fix it. Fatal in an invoice or account number, where nothing does.

A homoglyph is a character that looks like another character, such as the digit zero and the letter O. These mix-ups may be easy to infer in a sentence but dangerous in an account number, invoice ID, or other exact identifier. Check important identifiers against the original image or source.

### HTTP status code

The three-digit number a website gives back. 200 means it answered, 404 means the page does not exist, 403 means access refused.

The three digit code a web server returns with every response. 200 means it responded, 404 means the resource was not found, 403 means access was refused, 500 means the server itself failed. Useful as a first signal and insufficient as a final answer, because a server can return 200 while sending something entirely different from what was requested.

### Pre-registered prediction

Writing down what you expect, and how you will score it, before you run the test. It is the difference between an experiment and a demo, and it makes a wrong answer just as informative as a right one.

Write down what you expect a test to show, and how you will judge it, before running the test. That makes a surprising result useful evidence rather than something to explain away afterward. The practice helps a team learn from both success and failure.

### Regex gap

A search pattern that works on the format you built it for and silently finds nothing on a format you did not anticipate. No error appears. The field just comes back blank, and a blank looks like an answer.

A pattern used to find or extract something is written broadly enough to match the common cases in a dataset but not a format that shows up later or less often, and the mismatch produces no error, just a silently empty result for whatever it wasn't built to catch. A frequent failure mode in document and data extraction work (bank statements, invoices, structured PDFs, log parsing) where a field's format changes partway through a dataset, for example an account number that becomes partially masked after a certain date, and the extraction logic was written against the earlier format only. It looks exactly like a clean, successful run, because nothing errors and nothing gets flagged. The only way to catch it is to check the output itself for fields that are unexpectedly blank, rather than trusting a report that says zero issues were found.

### Silent success

Something that runs on schedule, reports success, and accomplishes nothing. Only caught by checking whether work came out, not whether the job ran.

An automated process that runs on schedule, reports success, and accomplishes nothing. The job starts, the log turns green, and no actual work happens. It is one of the most common and most expensive failures in automated systems, because every dashboard says everything is fine. The cause is almost always that the system measures execution rather than output: it checks whether the job ran, not whether the job did anything. The fix is to monitor results, not activity.

### Silent truncation

Output that quietly stops early and still reports success. There is no error, and the result reads as complete, because the missing part is exactly the part that would have told you.

Silent truncation happens when an output ends early but still looks complete and reports success. A quick skim may miss it because the missing ending is not there to raise an alarm. Compare length or counts with the original and check the last meaningful piece of content.

### Verification halo

Checking the part you touched and assuming the rest is fine. Confirming your change was made is not the same as confirming it did anything.

Confirming the part of a system you touched, then treating the entire system as verified. The check that was performed was real, which is what makes this so easy to miss. It just covered a narrower scope than the confidence it produced. Confirming that a change was saved is not the same as confirming the change had an effect, and the gap between those two is where a surprising number of production failures live.

## Identity, credentials and access

### API key

A single secret string that lets a program log into a service. There is no password prompt and usually no second factor, so anyone holding the key is you. Rotating it on the vendor's website does not update a program already running with the old one.

An API key is a secret credential that lets software use a service. Treat it like a password: anyone who obtains it may be able to act with its permissions. Keep it out of shared documents and chat, limit what it can access, and update the programs that use it when the key changes.

## File and content operations

### Atomic replacement

Switching one file from its old complete version to its new complete version in one visible step.

With atomic replacement, a reader sees either the complete old file or the complete new file. It avoids exposing a half-written file. It does not, by itself, protect against another editor racing to change the same file or make several file changes succeed as one unit.

### Canonical / Sole canonical

The one copy that counts. When several copies exist, this is the real one and everything else defers to it. It applies to whole accounts and systems too, not just to files.

The one version of something treated as the authoritative source when more than one copy exists. Calling something the "sole canonical copy" is a deliberate statement that it's the only version anyone should read from or edit going forward; everything else is a mirror, backup, export, or derivative that should defer to it, never the other way around. The term shows up constantly in any system with duplicated or synced data (files, folders, databases, documentation) because without naming one copy canonical, it's ambiguous which version is "true" when two versions disagree. Canonical applies to containers as much as to individual files: which account, workspace, or system is the one of record. That layer is the easier one to get wrong, because two accounts holding identically named folders both look correct until you check which one you are actually in. Naming the canonical container is worth doing before naming canonical copies inside it.

### Compare-and-write

Change a file only if it still matches the version you expected.

Before replacing a document, check that nobody has changed it since you last saw it. A true compare-and-write guarantee keeps that check and the update together. If they are separate steps, another editor can still slip in between them.

### Diff (verb: to diff / diffed)

Comparing two versions side by side and getting a list of exactly what differs, instead of eyeballing two long lists and hoping you spot the gap.

To compare two versions of something (two files, two folders, two document drafts) item-by-item and identify exactly what's only in one, only in the other, or changed between them. Same root as a software "diff" in version control. Useful any time two things are supposed to be identical, such as a backup, a synced folder, or a migrated copy, and you need to actually prove that rather than assume it.

### Page chrome

The furniture around a document's content: page numbers, headers, footers, navigation bars. Whether it counts as content depends entirely on the document, which is why automated tools get it wrong in both directions.

Page chrome is the material around the main content: page numbers, headers, footers, menus, and banners. Some of it is decoration; some carries context or makes citations possible. When a tool removes it, decide based on the document's purpose rather than assuming all surrounding material is disposable.

### Recovery copy

A saved copy of the earlier version so it can be examined or restored if an update goes wrong.

A recovery copy preserves what a file contained before a change. It gives you something concrete to inspect or restore after a mistake. Keeping the copy is one part of a recovery plan; the process for deciding when and how to restore it is another.

### Text layer

Real, invisible text stored behind the visual page of a PDF. Documents saved from a computer have one. Scans and screenshots usually do not. Without it, a PDF is just a picture and cannot be searched. To check, try selecting a sentence with your cursor.

A PDF text layer is computer-readable text stored behind the visible page. It lets you select, search, and copy words; a scan may contain only an image until text recognition is added. Check for a usable text layer before choosing how to process or verify a PDF.

## Vault and knowledge structure

### De-stuttering

A cleanup pass that removes the repeated part of a stuttering name, turning `bkfr_holdco-formation` into `bkfr_holdco-formation`. It is deliberately run as its own pass, never folded into the migration that created the stutter.

The cleanup pass that removes the repeated part of a stuttering name, so the prefix is left to carry that meaning by itself. The discipline worth keeping is that de-stuttering is run as its own pass, separately from the migration that created the stutter. A prefix migration is a single mechanical rule applied everywhere, which is exactly what makes it verifiable and reversible. Deciding which remaining words are redundant and which are load-bearing is a judgment call made file by file. Running both at once turns one auditable rule into two entangled ones, and leaves you unable to tell which change caused a given breakage.

### Sprawl

What happens when a folder structure grows without rules: two homes for the same kind of thing, folders nobody owns, names following no pattern. It is not about being deep. A deep structure can be perfectly clean if every level was deliberate.

The uncontrolled, inconsistent growth of a folder or file structure over time. It shows up as duplicate homes for the same kind of content, naming conventions that drift from folder to folder, and structure that no longer matches whatever documentation was meant to govern it. Sprawl is often confused with simple depth, but the two are different problems: a deeply nested structure can be perfectly clean if every level is deliberate and consistent, and a shallow structure can still be sprawling if it grew ad hoc. The tell for sprawl is not how many levels deep something is, it's whether new content keeps landing in the right place without anyone having to think about it.

### Stuttering

When a name says the same thing twice, because a prefix that already carries the meaning sits in front of words that repeat it. `bkfr_holdco-formation` stutters: `bkfr_` already means BK Forever, so `bk-forever` is dead weight.

When a name repeats itself because a prefix and the words after it carry the same meaning. A file named `acme_acme-vendor-list` stutters: the prefix already says Acme. It almost always appears after a naming convention is introduced or changed, because the older names spelled out what the new prefix now encodes, and a mechanical rename adds the prefix without removing the spelled-out version. On its own it looks cosmetic. At scale it is not: it makes every reference longer, it breaks autocomplete because typing the prefix stops narrowing anything, and it quietly teaches everyone that the prefix does not mean what it claims to mean.

---

## Publishing notes

Destination: bkblueprint.com glossary resource page. Each term should render as its own anchored heading so individual definitions are directly linkable and independently indexable.

Recommended structured data: `DefinedTermSet` with a `DefinedTerm` per entry. That is what makes a glossary eligible for rich results and citation by AI answer engines, which is the actual reason to publish one.
