A token is the unit an LLM actually reads and writes, a chunk of text, often a word piece or a few characters, not a whole word or letter.
Before any text reaches the model, it gets chopped into tokens by a fixed set of rules. "Tokenization" is common as one token; "unbelievable" often splits into two or three. Roughly, one token is three quarters of an English word.
Every cost, speed, and context-length number an AI provider publishes is measured in tokens, not words or characters. Reading a price sheet or an API response without knowing this leads to real math errors.
A tokenizer holds a fixed vocabulary, tens of thousands of common word pieces, learned in advance from a large text corpus. Text is greedily matched against that vocabulary and turned into a sequence of token IDs, integers the model actually operates on. The reverse happens on output: token IDs are decoded back into text.
Input and output tokens are usually billed at different rates, and total token count drives both latency (more tokens to generate takes longer) and cost. A prompt that reads short in English can still be token-heavy if it's dense with rare words or code.
The token counts and per-million-token pricing shown on every result card come straight from the provider's own usage field on that recorded run, never estimated from word count.