Token

LLM Foundations Beginner

A token is the unit an LLM actually reads and writes, a chunk of text, often a word piece or a few characters, not a whole word or letter.

In simple terms

Before any text reaches the model, it gets chopped into tokens by a fixed set of rules. "Tokenization" is common as one token; "unbelievable" often splits into two or three. Roughly, one token is three quarters of an English word.

Why it matters

Every cost, speed, and context-length number an AI provider publishes is measured in tokens, not words or characters. Reading a price sheet or an API response without knowing this leads to real math errors.

How it works

A tokenizer holds a fixed vocabulary, tens of thousands of common word pieces, learned in advance from a large text corpus. Text is greedily matched against that vocabulary and turned into a sequence of token IDs, integers the model actually operates on. The reverse happens on output: token IDs are decoded back into text.

Where it fits

Prompt text → Tokenizer → Token IDs → LLM → Token IDs → Detokenizer → Response text

Production impact

Input and output tokens are usually billed at different rates, and total token count drives both latency (more tokens to generate takes longer) and cost. A prompt that reads short in English can still be token-heavy if it's dense with rare words or code.

On howfastai.com

The token counts and per-million-token pricing shown on every result card come straight from the provider's own usage field on that recorded run, never estimated from word count.

Related terms

Learn next

← All terms · Knowledge map →