A tokenizer is the first component your sentence passes through when you send it to a model, and its job is not understanding — it only splits and numbers. Before text enters a large model, the tokenizer cuts it into individual tokens and then replaces each one with an ID from its vocabulary. What the model actually reads is that string of numbers. Once this clicks, a lot of confusion clears up: a token is neither a character nor a word.
It does not split by characters, but by frequency
Most tokenizers today use an algorithm called BPE, short for Byte Pair Encoding. During training, the algorithm counts which character combinations appear together most often, and the more frequently a combination occurs, the more likely it is to be merged into a single token. As a result, common words and frequent subwords often survive as whole chunks — one English word can be a single token — while rare words, names, and newly coined terms get broken into smaller pieces, sometimes three or four fragments for one word.
Spaces, punctuation, numbers, and code are all part of the splitting, and the rules are not intuitive. In English, the space before a word is often bundled with the word into the same token, and line breaks, runs of digits, and indented code are all handled separately. So the same content can end up with a different token count just because you changed its formatting.
Why the same sentence does not match across models
Every model's vocabulary is trained separately. Vocabularies differ in size, in merging rules, and in which languages they favor. The larger the vocabulary, the more and longer fragments it can usually keep whole. With a vocabulary that covers Chinese well, one Chinese character may be exactly 1 token; with poor coverage, that same character may be split into several byte-level tokens.
That answers the question in the title. As a rule of thumb, 1 token in English corresponds to roughly 3 to 4 characters. Chinese varies far more: one character taking 1 token is very common, but 2 or more is hardly unusual either — it all depends on the model. Using model A's count to estimate model B's usage and getting a mismatch is normal; it does not mean either model counted wrong.
Token count also directly determines what things cost in practice. With usage-based billing, input and output are mostly charged per token. The context window limit and how much conversation history fits in one go are also measured in tokens, and the more tokens a reply has to generate, the slower it usually comes back. This is even more obvious when building a knowledge base: the document chunking discussed in What exactly is RAG? How it differs from fine-tuning and prompt engineering is also measured in tokens rather than characters — chunks that are too large waste the window, and chunks that are too small risk cutting meanings in half.
Three calculations people often get wrong
The first is estimating cost directly from character count. There is no fixed conversion between characters and tokens, and the gap grows when English, Chinese, and code are mixed together, so budgets estimated from character counts tend to come out too low.
The second is assuming Chinese is always cheaper, or always more expensive. The more accurate statement is this: with most mainstream vocabularies, saying the same thing in Chinese usually consumes more tokens than in English, because common English words are more likely to match as whole chunks while Chinese is more likely to be split apart. But this is not a law of nature. Switch to a vocabulary that is well optimized for Chinese and the gap shrinks noticeably — the conclusion always has to be tied to a specific model.
The third is treating token count as evidence of how strong a model is. If the same sentence is 80 tokens in one model and 110 in another, that only shows the two split text differently; it says nothing direct about which one is smarter. Vocabulary design is a trade-off: a large vocabulary splits coarsely but makes each token more expensive to represent, while a small vocabulary splits finely and produces longer sequences.
What it cannot do
A tokenizer is not the model itself. It only decides at what granularity text is fed in; it does not understand meaning and cannot judge whether a sentence is right or wrong. How sensibly text is split does affect how well a model performs — poor splitting makes arithmetic and rare words harder to handle — but understanding and reasoning ultimately come down to the model's parameters. Keeping the two separate stops you from expecting too much of the tokenizer, or blaming the wrong component.
There is a simple habit for everyday use: when estimating cost and length, do not count characters in your head. Run your actual text through the tokenizer tool that matches the model you are using, include the system prompt, the conversation history, and room for the output, and leave a buffer of about twenty percent. Once you budget in tokens instead of characters, most surprise bills and context overflows can be avoided in advance.