Kelly Mears

Token

The sub-word unit a language model actually reads and writes; the unit of cost and of context.

Agents & Language Models1 min read199 words10 out · 12 in
also calledTokensTokenization

A token is the atomic unit a Large Language Model processes. Tokens are not words or characters but fragments produced by a learned segmentation — common words are usually a single token, rare words split into several, and whitespace and punctuation carry tokens of their own. English text averages roughly four characters per token.

Tokens matter for three practical reasons. They are the unit in which the Context Window is measured, so "how much can I show the model" is a token question. They are the unit of billing, so they are the unit of Token Budget. And they are the unit of latency, since generation is sequential.

Tokenisation also explains a family of otherwise-strange model behaviours: difficulty with character-level tasks like counting letters or reversing strings, uneven handling of unusual formatting, and the fact that the same content costs different amounts depending on how it is written. Dense structured formats and long identifier names are more expensive than their information content suggests.

Because tokens are consumed by everything in the request — instructions, tool definitions, prior turns, retrieved documents — reducing any of these frees capacity for the rest. That trade is the whole subject of context engineering.

See also7

Hand-picked in the note itself — the neighbours worth reading next.

Related2

Nearby in the graph rather than deliberately chosen. Looser, sometimes surprising.

Linked from12

Notes elsewhere in the wiki that reach for this one.