← Back to all writing
Fundamentals13 Sept 2026 · 13 min read

What are tokens?

And why a model never reads your words.

Juan José Alonso
Juan José AlonsoWrites here, every other week

ChatGPT, Claude, Gemini: the products most of us use every day all run on large language models, or LLMs. Contrary to what you might assume, they do not understand language; at least not the way we humans do.

Part of the confusion comes from the name. Language model is an old term from computer science for a program that, given some text, predicts what comes next. It describes what the model does with language, not whether it understands any of it. And the way we use these models makes the mix-up even easier:

  • You: Give me a recipe for banana bread
  • LLM: Sure! Here's a recipe for Classic Banana Bread...

Sure: text goes in, text comes out. And text is language. So it's natural to assume that somewhere in between, the model is reading your words. It isn't.

Think of a pocket calculator. You type 7 + 5 and it shows 12, but there is no 7 and no 12 inside it, only electrical signals switching on and off. The digits are there for your benefit. The calculator is great at arithmetic without ever seeing a number the way you do.

LLMs are similar. You write words and read words back, but in between, the model works with neither. My point being: just because we talk to LLMs in language doesn't mean they process it as language. Whatever understanding happens in there (and I do think some does), it doesn't happen in words.

So if an LLM doesn't work with words, what does it work with? Tokens.

What a token is

A token is a chunk of text. Before your message reaches the model, a separate program called a tokenizer cuts it into chunks taken from a fixed list, its vocabulary, and replaces each chunk with its number on that list. Those numbers are all the model receives.

The sentence "Give me a recipe for banana bread" split into seven coloured tokens. Below it, inside a dark box, are the seven numbers the model receives: 50117, 668, 261, 14420, 395, 60252 and 19544.
What you type, and what the model receives. In this sentence, each space belongs to the word after it.

The cover of this article is another example. GPT-4o's tokenizer cuts incalculable into four chunks:

inc (2768) · al (280) · cul (2885) · able (562)

Common words usually get a token to themselves, while longer or rarer ones are split into pieces. GPT-4o's tokenizer has a vocabulary of about 200,000 chunks; GPT-4's had about 100,000.

Small changes produce different numbers. Written with a space in front, the way it appears in the middle of a sentence, incalculable starts with inc (4570) instead of inc (2768). Capitalised, it splits another way entirely: In · calcul · able. Nothing in those numbers says they are the same word. The model learns that from seeing them used the same way in its training text.

A token ID is only a label. Inside the model, each one is swapped again, this time for a long list of numbers learned during training. Tokens that are used in similar ways end up with similar lists, and those lists are what the model does its calculations with.

How the chunks are chosen

Nobody writes the vocabulary by hand. It is learned from a large pile of text before the model itself is trained, and OpenAI's tokenizers learn it with a method called byte pair encoding:

  1. Start with single characters (strictly speaking, bytes).
  2. Find the pair of neighbours that appears most often and glue it into a new token.
  3. Repeat until the vocabulary is full.

GPT-2's tokenizer, from 2019, numbers its tokens in the order it learned them. The first pair it glued together was a space followed by a t. Then came a space and an a, then he, in, re and on. By the seventh merge it had the.

Frequent sequences become single tokens and rare ones stay in pieces. The splits come from frequency alone, with no regard for syllables or meaning, which is how incalculable ends up cut between al and cul. Once trained, the tokenizer is frozen, and the model learns everything it knows from text that has already been through it.

Why not whole words, or single letters?

A vocabulary of whole words would never be complete. Every name, every typo and every verb form in every language would need an entry, and any word left out could be neither read nor written. With pieces, anything can be spelled out, down to single bytes if necessary, so tokenizers like OpenAI's never meet text they can't handle.

Single letters fail the other way. The English introduction to this article is 1,332 characters long but only 298 tokens, so reading it letter by letter would make it four and a half times longer. The work a model does grows with the length of what it reads, and so does the share of its context window, the limit on how much text it can take in at once, that the text uses up.

Tokens sit between the two, which keeps the vocabulary finite and the text short.

What tokens explain

Once you know a model reads numbered chunks, some of its best-known failures start to make sense.

The r's in strawberry

"How many r's are in strawberry?" comes to eight tokens. One of them, token 101830, is strawberry, space included. The model gets that single number. The letters aren't in it, so the model has to answer from whatever it picked up about the spelling of token 101830 during training.

Spell the word out as s-t-r-a-w-b-e-r-r-y and it becomes ten tokens, with the r's appearing as the same token, -r (6335), three times. Now there is something to count. Newer models usually get the question right, and asking any model to spell a word out before counting its letters helps.

Top: "How many r's are in strawberry?" split into eight tokens, with " strawberry" highlighted as a single token, number 101830. Bottom: "s-t-r-a-w-b-e-r-r-y" split into ten tokens, with the three "-r" tokens highlighted, each one number 6335.
Written normally, strawberry is a single token. Spelled out, each r becomes a token of its own.

Anything that depends on individual letters runs into the same wall: reversing a word, solving an anagram, counting syllables.

How many words are in your essay

Tokens don't line up with words. The English introduction to this article has 243 words and 298 tokens; the Spanish version has 270 words and 356 tokens. The ratio shifts with vocabulary, punctuation, numbers and language, so knowing one number doesn't give you the other.

Counting also needs a running tally, and keeping an exact count across thousands of tokens is hard for these models in general. Tokens add to the problem, because the model reads tokens and you asked about words. The same goes for "write exactly 500 words". If the app you're using can run code, ask it to count with code.

Why some languages cost more

A tokenizer only learns chunks that show up often in its training text. When most of that text is English, English words become whole tokens, while other languages keep more of theirs in pieces.

Here is the first article of the Universal Declaration of Human Rights in the UN's six official languages1, counted with GPT-4's tokenizer and with GPT-4o's:

Language GPT-4 GPT-4o
English 33 33
Spanish 44 (1.3×) 38 (1.2×)
French 50 (1.5×) 41 (1.2×)
Russian 74 (2.2×) 41 (1.2×)
Arabic 91 (2.8×) 44 (1.3×)
Chinese 50 (1.5×) 35 (1.1×)

The Arabic version says the same thing as the English one and takes almost three times as many tokens with GPT-4's tokenizer. More tokens means a higher bill for the same request, a slower answer and less room left in the context window. Across a much wider set of languages, researchers at the University of Oxford found gaps of up to 15 times2.

GPT-4o's tokenizer closes a lot of that gap. With twice the vocabulary, it has room for more chunks from more languages. English still gets the best deal, though. The Spanish introduction to this article takes 19% more tokens than the English one, and 36% more with GPT-4's tokenizer.

Long numbers

GPT-4o's tokenizer reads numbers in chunks of up to three digits, starting from the left. 48271 becomes 482 · 71 and 3591 becomes 359 · 1. To add them, a model has to line up units with units and tens with tens, and these chunks don't line up that way. The last chunk of the first number, 71, holds tens and units, while the last chunk of the second, 1, holds only units.

Commas fix it. Written as 48,271 and 3,591, the numbers split into 48 · , · 271 and 3 · , · 591, and both three-digit chunks hold hundreds, tens and units. In a 2024 study, writing numbers with commas raised GPT-3.5's accuracy at adding seven- to nine-digit numbers from 75.6% to 97.8%3. When a chat app can run code, it can hand the arithmetic to that instead.

Two sums written in columns. On the left, 48271 plus 3591, cut into 482 and 71, and 359 and 1, so the chunks sit over different digits. On the right, 48,271 plus 3,591, cut into 48 and 271, and 3 and 591, so the three-digit chunks line up.
Without commas, the chunks fall on different digits. With commas, they line up with the place values.

A word the model couldn't say

The tokenizer and the model are trained separately, on different text, and occasionally it shows. GPT-2 and GPT-3 shared a tokenizer that had SolidGoldMagikarp as a single token, number 43453. It was the username of one of the most active members of r/counting, a Reddit community that has spent years counting to infinity one comment at a time, so it appeared constantly in the text the tokenizer learned from. The model's own training text most likely left those threads out, which gave it a token it never learned anything about.

In early 2023, Jessica Rumbelow and Matthew Watkins found over a hundred tokens like it4. Asked to repeat "SolidGoldMagikarp", OpenAI's models of the time answered with something else entirely, such as "distribute". GPT-4o's tokenizer splits the name into five ordinary pieces.

Tokens and money

Prices are listed per million tokens, so it's easy to think of tokens as a unit of cost. They measure text. What you pay for a request depends on the model, on how many tokens go in and come out, and on tokens you never see.

Bigger models cost more per token, and in one company's lineup the largest model can cost ten or twenty times as much per token as the smallest. Output also costs five to six times as much as input at OpenAI, Google and Anthropic, at the time of writing5. Writing is the expensive part. A model takes in your whole prompt in one pass, but it has to run once for every token it writes.

Some of the tokens you pay for never appear on screen.

  • Thinking. Models that reason before answering write tokens for that reasoning. You usually see a summary of it or nothing at all, but OpenAI, Google and Anthropic all bill the full reasoning as output6.
  • The conversation so far. A model keeps no memory between messages, so every time you send one, the whole conversation goes back in as input. The fortieth message in a chat sends far more tokens than the first. Providers charge about a tenth of the usual price for input they have processed recently, a discount called caching, but that input still costs something.
  • Images and files. These become tokens too. Models cut an image into tiles or small square patches and count tokens for each one, so a bigger image costs more7.
Five rows of blocks, one for each message in a conversation. Every row repeats all the earlier messages and replies before the new message, so each row is longer than the one before.
What gets sent with each new message: every earlier message (pink) and reply (lilac), then the new message (dark).

The same text also comes to different counts on different tokenizers. GPT-4o's tokenizer turns the Spanish introduction to this article into 356 tokens, where GPT-4's needed 414. A new tokenizer can change a bill even when the price per token stays the same. When Anthropic released Claude Opus 4.7, it kept its predecessor's prices but switched to a tokenizer that turns the same input into up to 35% more tokens8. A price per million tokens only compares fairly between models that share a tokenizer.

Subscriptions hide all of this behind a monthly fee, but the model still reads the whole conversation with every message. Starting a new chat when you change topic keeps that reading short, and in apps whose usage limits depend on conversation length, it makes the limit last longer9.

Try it yourself

Tiktokenizer and OpenAI's Tokenizer show how a piece of text gets split. In Tiktokenizer, choose gpt-4o to match the numbers in this article. Paste in your name, a sentence in another language, or a long number with and without commas. Try an emoji as well. 😀 is a single token, and 🇨🇱 takes four.

Footnotes

  1. United Nations, Universal Declaration of Human Rights. The counts use the UN's official text of Article 1 in each language.

  2. Aleksandar Petrov, Emanuele La Malfa, Philip H.S. Torr and Adel Bibi, "Language Model Tokenizers Introduce Unfairness Between Languages", NeurIPS 2023.

  3. Aaditya K. Singh and DJ Strouse, "Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs", 2024. The figures are for gpt-3.5-turbo-0301 given eight worked examples.

  4. Jessica Rumbelow and Matthew Watkins, "SolidGoldMagikarp (plus, prompt generation)", LessWrong, February 2023.

  5. Pricing pages for OpenAI, Google and Anthropic, checked September 2026.

  6. Documentation from OpenAI, Google and Anthropic.

  7. OpenAI's newer models count 32×32-pixel patches, Google counts 768×768-pixel tiles and Anthropic counts 28×28-pixel patches. See the documentation from OpenAI, Google and Anthropic.

  8. Anthropic, Introducing Claude Opus 4.7, 16 April 2026.

  9. Claude's limits work this way, for example. See How do usage and length limits work?

A letter every other week.

A note from me about each new piece, in your language. Your address stays here.

Keep reading