The Decoder/What is a token in AI?

The Decoder • numbers checked 30 September 2026

What is a token in AI?

A token is the unit of text an AI model reads and writes: a chunk about four characters long in English, so a short word is one token and a long word is two or three.

A long strip of paper on a dark desk cut into uneven pieces with a pair of scissors, the pieces laid out in a row, two of them glowing

A token is the piece of text an AI model actually works with. It is not a word and not a letter. It is a chunk, about four characters long in English, cut by a fixed dictionary the model was trained with. Every price, every usage limit and every context window you see is counted in tokens, which is why the same paragraph costs a little more than its word count suggests.

Key facts

  • ~4characters per token in English on OpenAI models, the published rule of thumb
  • ~3.5characters per token on Claude, by Anthropic's own glossary
  • ~750words in 1,000 tokens (OpenAI); Anthropic puts 1 million tokens at about 555,000 words on its current tokenizer
  • 2 to 3tokens for a long or unusual word the dictionary has to split
  • OpenAI tokenizer help page (1 token is about 4 characters or three quarters of a word); Anthropic glossary and models overview; checked 30 September 2026

Is a token the same as a word?

No. A token is smaller than a word most of the time and occasionally bigger. Common short words such as the, cat or sat are one token each, and the space before a word usually rides along inside the same token. A rare or long word gets split into pieces the dictionary does know. Numbers, punctuation and emoji each take their own tokens too, which is why a table of figures costs more than a paragraph of prose with the same character count.

The dictionary is fixed. It was built once, before training, by scanning a huge pile of text and keeping the chunks that appeared most often. That is why English is cheap in tokens and other languages are often not. A sentence in Spanish or Hindi can take half again as many tokens as its English translation, because the dictionary saw less of it.

Sample texts with their word and token counts
Text Characters Words Tokens (approx.)
The cat sat on the mat. 23 6 7
Antidisestablishmentarianism 28 1 6
Call me at 203-555-0142 tomorrow. 33 5 12
A 1,500 word article 8,500 1,500 2,000 to 2,700
A 300 page book 540,000 90,000 120,000 to 160,000

How many tokens is 1,000 words?

About 1,333 on OpenAI’s models, and closer to 1,800 on Claude. The rule OpenAI publishes is that one token is roughly four characters or three quarters of a word, so 1,000 tokens is about 750 words. Anthropic’s glossary puts a Claude token at about 3.5 characters, and its models page says 1 million tokens is roughly 555,000 words on the current tokenizer. Same paragraph, different dictionary, different count. Plain prose lands close to those figures. Code, addresses, spreadsheets and anything with a lot of symbols land higher.

You can check any text yourself. OpenAI publishes a free tokenizer page where you paste text and watch it colour each token. The counts differ between vendors because each model has its own dictionary, but the ratio stays in the same neighbourhood for English.

Words versus tokens, same text WORDS VERSUS TOKENS, SAME TEXT 1,000 words of prose 1,000 words in OpenAI tokens ~1,333 in Claude tokens ~1,800
English prose runs at about 1.33 tokens per word on OpenAI's tokenizer and about 1.8 on Claude's current one. The gap grows with numbers, code and other languages.

Why do I pay per token?

Because the model does the same amount of work for every token, whether the text is brilliant or blank. Each token in your message has to be read, and each token in the reply has to be produced one at a time. Pricing follows the work. A vendor quotes a price per million tokens in and a separate, higher price per million tokens out, because producing text is more expensive than reading it.

The same rule sits behind the limits on the subscription plans. A cap on messages per five hours is really a cap on tokens, which is why a long conversation with big pasted documents runs into the wall sooner than a string of short questions.

What happens when I paste a whole document?

Every character in it becomes tokens, and they all count. A ten page PDF is roughly 5,000 words, which is around 6,700 tokens, before the model has written a single word back. If you then ask follow up questions, the document is re-read on every turn, because the model does not remember between messages. It is handed the whole conversation again each time. That is the mechanism behind the context window, and it is the reason long chats slow down and eventually forget the beginning.

Filed under Level: Start Here

Comments

No comments yet. Yours can be the first.

Comments are read before they appear.

No schedule, no filler

Get the next entry