Models & LLMsGuides & Tutorials

Tokenization Explained: How AI Reads Your Words (and Why It Costs Money)

See the exact pieces an AI sees when it reads your words. This guide shows how tokenization works, real token splits, why the strawberry mistake happens, how tokens are priced, and how to count your own tokens in one minute.

Toolbit AI - Team
16 min read
Tokenization Explained: How AI Reads Your Words (and Why It Costs Money)

"You miss 100% of the shots you don't take." Wayne Gretzky said that. It's a short, famous sentence. You can read it in about two seconds.

Now watch what an AI actually sees when it reads that same sentence. It gets cut into eleven pieces. The word "the" is one piece on its own. A word like "unbelievable" isn't one piece at all - it's three pieces: un + bel + ievable. The AI never saw the words. It saw the pieces.

A token is a puzzle piece of text. Sometimes it's a whole word. Sometimes it's a crumb of a word. AI reads in pieces, thinks in pieces, and charges money by the piece.

That's the whole secret, and once you see it, a lot of weird AI behavior suddenly makes sense. Why does AI sometimes miscount the letters in "strawberry"? Pieces. Why does a long chat slowly get more expensive? Pieces. Why does one message cost money at all? Pieces.

In short:

  • AI doesn't read words. It reads tokens - small puzzle pieces of text.
  • One token is about 4 characters, or about ¾ of an English word.
  • Everything is priced and counted in tokens: what you send, what it answers, even what it "thinks."
  • Weird AI mistakes (like counting r's in strawberry) happen because of how words get cut into pieces.
  • By the end of this post, you can count your own tokens with a free tool in one minute.

AI doesn't read words. It reads pieces

Imagine a big box of LEGO bricks. Some bricks are big and ready-made - press one down and you've built a whole wall in a second. Others are tiny, and you need five of them snapped together to make the same wall.

That's exactly how text arrives at an AI. Every sentence you type gets snapped apart into bricks first. Short, common words like "the" and "cat" are one big brick each. Long, rare words get built out of several small bricks.

So what is tokenization? It is the cutting step. Your text gets split into tokens, and those tokens are all an AI model ever sees. And tokens are not the same as words - a token can be a whole word or just a piece of one.

Now the proper definition, and it's friendlier than it sounds. OpenAI says tokens are the building blocks of text. A single token can be "as short as a single character or as long as a full word." Google says the same thing with different words: a token can be a single letter like z, or a whole word like cat, and long words get broken up into several tokens.

Here's the full round trip, in kid steps:

  1. You type some text.
  2. The text gets cut into pieces.
  3. Each piece gets a number, like a name tag on a brick.
  4. The AI works with those piece-numbers. This is all it ever sees.
  5. The answer comes back as a string of new pieces.
  6. The pieces get glued back together into normal words, and that's what you read.

Who does the cutting? A little helper program called a tokenizer. OpenAI's tokenizer is called tiktoken, and it's a fast BPE (byte pair encoding) tokenizer. In kid words: BPE is a recipe for cutting text the same way every single time, so nothing gets cut at random. You can read more about tokens and how they're counted in OpenAI's help article.

The 11-token sentence: a real example

A real sentence split into its 11 verified tokens

Think about eating a bowl of noodles. You don't swallow the whole bowl in one go. You also don't eat one noodle at a time, because that would take forever. You take bites - and one bite is about four noodles. That's the size of a piece.

Here's the short answer on how many tokens text turns into. A good rule of thumb: one token is about 4 characters, or about ¾ of an English word. So 100 words is roughly 130 to 150 tokens. The exact count depends on the words and the model.

Let's walk our star example, bite by bite. The Gretzky sentence - "You miss 100% of the shots you don't take" - is exactly 11 tokens. That's not a guess. It's the count printed in OpenAI's own help article, and you can paste the sentence into their free tool and count them yourself.

Now the tour of real splits. Every one of these was made with OpenAI's actual tokenizer, so you can reproduce any of them:

  • the, cat, hello, a - just 1 piece each. These are the big pre-made bricks. They're so common that the tokenizer knows them whole.
  • ChatGPT - 2 pieces: Chat + GPT.
  • tokenization - 2 pieces: token + ization.
  • unbelievable - 3 pieces: un + bel + ievable.
  • strawberry - 3 pieces: st + raw + berry. (Remember this one. It's about to become famous.)
  • antidisestablishmentarianism - 6 pieces: ant + idis + est + ablishment + arian + ism.

See the pattern? Common words ride alone. Rare, long words travel in crumbs.

So how many pieces is a normal piece of writing? The rule of thumb, from all three big AI companies: one token is about 4 characters, which is about ¾ of an English word. So 100 tokens is roughly 75 words. OpenAI says ~75, Anthropic agrees, and Google says 60 to 80 - so "about ¾ of a word, and it varies a bit" is the honest version.

For scale, here are two real reference points from OpenAI: the whole OpenAI Charter is about 476 tokens, and the US Declaration of Independence is about 1,695 tokens. A document you'd read in a few minutes is a few hundred to a couple thousand pieces.

One more thing, and it matters if you speak (or write in) more than English. A study presented at NeurIPS in 2023 measured the same meaning translated into different languages, and found the token count could be up to 15 times higher in some languages than in others. That study looked at 2023-era tokenizers; newer ones are better trained on more languages, but the gap hasn't fully gone away. Kid version: some languages get fewer big bricks in the box, so the same sentence needs more small bricks - which means more cost and more waiting.

Want to see your own words split into colored pieces? Try the free OpenAI Tokenizer tool - paste anything and watch it break.

Why pieces, not words

Back to the kitchen. Why not just swallow the whole sandwich (one piece per whole word)? And why not eat one crumb at a time (one piece per letter)? Because both extremes break down, and the middle - snack-sized bites - is just right.

The short answer for why AI uses tokens and not whole words: subword pieces are the sweet spot. Whole words would need a huge, brittle brick set. Single letters would be too many pieces. Subwords balance both, so every major AI uses them.

Here's the technical reason, in plain words. OpenAI describes tokens as "common sequences of characters found in a set of text." That means the tokenizer studied mountains of text and noticed which chunks show up a lot. Chunks that appear often became pieces. If "the" shows up millions of times, it earns its own single piece. A chunk like "ablishment" almost never appears, so it never gets its own piece - it only exists inside bigger cuts. Researchers who studied this design (you can read the paper on arXiv) describe it as: give pieces to the character-sequences that appear frequently.

Now the two failure modes, in kid words:

  • One piece per LETTER = way too many pieces per sentence. Every short word becomes four or five pieces, and the AI has to predict the next thing four or five times more often. Predictions get slow and wobbly.
  • One piece per WORD = way too many different pieces to memorize. Every language, every name, every typo, every weird compound word would need its own brick. The box of bricks becomes impossibly huge - and one typo would create a brick the AI had never seen.

Subword pieces are the middle ground, and that's what every major AI uses.

For a wow number: gpt-4o's tokenizer knows exactly 200,019 different pieces. That's the whole box of bricks it owns. The box is bigger than the word list of a big dictionary - and it has to be, because a full box of pieces needs every common word plus all the word-halves and crumbs to build rare words from.

And here's proof that pieces belong to the model, not to the text. The exact same sentence can come out as a different number of pieces on a different AI. Anthropic's docs say Claude 4.7 and later models use a newer tokenizer that produces about 30% more tokens for the same input text than earlier Claude models did. Kid version: different AI kitchens cut the sandwich differently. There is no one true way to slice text.

The strawberry mystery, solved

The strawberry split: st, raw, and berry hide the r letters inside pieces

Here's a domino trick. Imagine someone shows you dominoes only from across the room. You learn to recognize each one by its shape and its ID number - "that's #42, that's #7." But you've never walked up close and counted the printed dots. Now someone asks: "How many dots are on domino #42?" You might guess wrong. Not because you're dumb - because you were never close enough to count.

That is exactly what happens when you ask an AI "how many r's are in strawberry?" The short answer: the model never sees letters. It only sees token IDs, so it can't count letters inside a token unless it learned them by heart.

First, the split. With OpenAI's tokenizer, strawberry becomes 3 pieces: st + raw + berry. This isn't a rumor - the split is printed in a research paper on why language models struggle to count letters (arXiv:2412.18626), and it matches the actual tokenizer output exactly.

Now the mystery, in kid steps. The AI never sees letters - it sees piece-numbers. To answer the r-question, it would have to have learned, somewhere along the way, that st contains 0 r's, raw contains 1, and berry contains 2, and then add them up: 3. But the letters are hidden inside the piece-numbers, like dots printed on a domino you only know from afar. A 2025 paper called, yes, "The Strawberry Problem" (arXiv:2505.14172) found that the link between a token and its characters is weak - models only pick it up indirectly, if at all, during training.

Why do the errors pile up on words like this? The letter-counting paper found two things: models can often recognize a letter but not count it, and they struggle most when a letter appears more than twice. "Strawberry" has three r's spread across two pieces - about the worst-case puzzle you can build from a breakfast fruit.

And it's not just strawberry. The same root cause makes AI wobbly at reversing strings, swapping letters, or removing a letter from a word - all the things you'd find easy if you could see the letters, and hard if you only knew bricks by ID.

The fix is delightful: spell it out. If you ask "how many r's in s-t-r-a-w-b-e-r-r-y", each letter becomes its own token - now the model actually sees the letters, one piece each. That's why that trick works. You're not making the AI smarter. You're handing it the dominoes up close.

What tokens control: money, memory, and limits

Think about buying candy by the piece at one of those pick-a-mix shops. Every piece you hand over gets weighed and priced. Every piece the shopkeeper hands back does too. And some bins have fancier candy than others - same piece, different price.

That's the AI bill. You pay for the pieces you send (your question), the pieces you get back (the answer), and - at some shops - even the pieces the candy-maker "thinks with" behind the counter.

The short answer on cost: you pay for input tokens, output tokens, and thinking tokens, and output tokens cost the most. Same text can also cost different amounts on different models, because each model prices its pieces its own way.

Let's do the arithmetic with real prices, step by step.

Say you send an AI a 2,000-word article. At ¾ of a word per token, that's about 2,667 tokens of input.

  • Send it once to Claude Sonnet 5: 2,667 pieces at $2 per million input tokens ≈ $0.0053. About half a cent.
  • Send the same thing to gpt-5.6-luna: 2,667 at $0.20 per million ≈ $0.0005. Cheaper candy bin.
  • Now the model writes back 1,000 words (about 1,333 output tokens) on Sonnet 5: 1,333 at $10 per million ≈ $0.013. The answer costs more than the question did.

That last one surprises people: on every major provider, output tokens cost more than input tokens - usually several times more. Makes sense when you think about it: writing candy takes more effort than reading candy. You can check every number above on the official OpenAI pricing, Anthropic pricing, and Gemini pricing pages. Pricing and plan details are as published by the vendor around September 2026 and can change - confirm on the official site. Gemini even proves the point: its current Gemini 3.x Flash prices are listed as lasting "through December 31, 2026," after which they double. Prices really do change.

(Choosing between cheap and pricey candy bins is a whole skill of its own - if you're curious when a cheap model is enough, we wrote a guide on how to pick the right AI model.)

Next: memory. Every AI has a lunchbox. The lunchbox fits only so many pieces - OpenAI calls it a "maximum combined token limit" for input plus output together. Today's big models hold hundreds of thousands of pieces in the lunchbox (exact sizes vary by model). When the lunchbox is full, the oldest pieces have to fall out - that's why a very long chat "forgets" the beginning.

And here's the sneaky part: everything in the lunchbox counts. Your input, the output, any cached pieces, and even thinking tokens - the pieces a "reasoning" model makes while working out the answer before it replies. OpenAI's help article lists reasoning tokens as their own billed category, and Google's Gemini pricing explicitly says output is billed "including thinking tokens." Yes, you pay for the thinking too.

One more effect worth knowing: the conversation snowball. In a chat, every new turn re-sends the whole conversation so far as input. Turn 1 sends a little. Turn 20 sends everything from turns 1 through 19. Long chats get more expensive every turn - though providers discount the pieces they already have cached from earlier turns, which softens the snowball.

Count your own tokens in one minute

Ever weigh your suitcase before a flight so there are no surprises at the counter? The airlines give free scales. So do the AI companies - they want you to know your piece count before you send anything. Here's one per provider:

The quickest way to count tokens is a free token counter. OpenAI's Tokenizer web tool and Anthropic's count_tokens endpoint both give you an exact count before you spend anything.

  • OpenAI's Tokenizer web tool (at platform.openai.com/tokenizer): paste text, and it shows the total count and the colored split - every piece gets its own color so you can see exactly where the cuts land. You can even pick which model family's tokenizer to use, since (as we learned) each kitchen cuts differently.
  • The tiktoken package, for coders: if you write Python, tiktoken.encoding_for_model('gpt-4o') gives you the same tokenizer the models use, right in your code. Every split in this post was produced with it.
  • Anthropic's count_tokens endpoint: an API call that counts your input before you send the real request - it accepts the same message structure (system prompts, images, tools), and returns the token total. One caveat from their docs: it's an estimate, and system-added tokens aren't billed. Details in the Anthropic token-counting docs.
  • Google's count_tokens and usage meters: Gemini's API has a count_tokens call, and every response comes back with a usage breakdown showing input, output, thought, cached, and total tokens. Google's tokens documentation covers both.

Mini recipe, one minute from now:

  1. Open the OpenAI Tokenizer tool.
  2. Paste any sentence - "The quick brown fox jumps over the lazy dog." is a classic (it comes out as 10 pieces).
  3. Watch it split into colored pieces.
  4. Try "unbelievable" and see it break into three: un + bel + ievable.
  5. Try your own name. Now you know your personal token count.

And to turn pieces into money: take your token count, multiply by the price per million tokens, then divide by one million. Worked line: 2,667 tokens at $2 per million = 2,667 × 2 ÷ 1,000,000 ≈ $0.0053. That's it - that's the whole formula for estimating any AI bill.

Why this matters more every year

Tokens are the hidden currency of the AI world. New models arrive every few months, and every single one of them - no matter how fancy - still reads in pieces, still thinks in pieces, and still charges by the piece. Learn the currency once and it works forever.

Every large language model (LLM) uses tokens the same way. Tokens in AI are a shared currency across chat apps, coding helpers, and search assistants. So once you know how tokens work, you can read any LLM you will ever meet.

The three things worth keeping:

  • AI reads pieces, not words - about ¾ of an English word per piece.
  • Every piece has a price, and the lunchbox is finite - input, output, and thinking all count.
  • When AI trips over letters (strawberry, reversing, spelling), remember: the bricks hide the letters, and the model only sees brick numbers.

Want to pull the thread on where those pieces come from in the first place - how models learn from mountains of text during training? We have a walkthrough on how AI models actually get smarter with training.

Token FAQs: the five questions everyone asks

How many tokens are in 1,000 words? About 1,300 to 1,500 tokens in English, since one token ≈ ¾ of a word (or ~4 characters). It varies with the text and the model - code, unusual words, and non-English languages all change the count.

Do all AI companies use the same tokenizer? No. Each model family has its own tokenizer. The same text produces about 30% more tokens on Claude 4.7-and-later models than on earlier Claude models, and gpt-4o's tokenizer knows 200,019 pieces - each kitchen cuts the sandwich differently.

Why does AI fail at "how many r's in strawberry"? The model never sees letters - it sees the three token IDs for st, raw, and berry. It can only answer if it learned the letter counts inside each piece indirectly, which it often hasn't. Spelling it out letter by letter fixes it, because each letter then becomes its own token.

Do output tokens cost more than input tokens? Yes, on every major provider. For example, Claude Sonnet 5 was listed at $2 per million input tokens versus $10 per million output tokens as of September 2026. Writing candy takes more effort than reading candy.

Do "thinking" tokens cost money too? Yes. Google's Gemini pricing bills output "including thinking tokens," and OpenAI lists reasoning tokens as their own billed usage category. The pieces a model makes while reasoning are on your bill like any other.

Share this article

Related articles

Continue exploring similar guides and insights