What is a token?
A token is the unit of text a language model reads and writes. Common English words are usually one token, while rare words, numbers, code symbols and non-Latin scripts are split into several. In English, one token averages about four characters, or roughly three quarters of a word.
How accurate is the token count?
OpenAI counts up to GPT-5.6 are exact because they use the same vocabularies as tiktoken. GPT-6, Claude, Gemini and Llama counts are estimates marked with ≈ and are usually within about 15% for English text and code. Unusual text, such as long runs of emoji or rare scripts, can drift further.
Why do Claude and GPT give different counts for the same text?
Each provider trains its own tokenizer with its own vocabulary, so the same sentence splits into different pieces. Claude's current tokenizer tends to produce more tokens than OpenAI's o200k_base for the same text, which affects both cost and how much fits in the context window.
How do I get an exact token count for Claude?
Use Anthropic's token counting endpoint, POST /v1/messages/count_tokens, which returns the exact input token count for a request and is free to call. Gemini offers a similar countTokens method in the Gemini API.
Does the count include chat formatting or system prompts?
No. The counter measures the text you paste. Chat APIs add a few tokens per message for role markers and formatting, and tools, images and system prompts count too. Paste your system prompt with the user message to include it.
Is my text uploaded anywhere?
No. Tokenizing runs in your browser. The page downloads the tokenizer data once, and after that nothing you type is sent to a server. Your last text is kept in this browser's local storage so it is still there when you come back.
How current are the prices?
Prices are standard API list prices per million tokens, checked against each provider's pricing page on the date shown below the estimate. Batch, regional and committed-use pricing differ, and providers change prices often, so confirm on the provider's page before you budget.
Why is JSON so expensive in tokens?
Pretty-printed JSON spends tokens on indentation, line breaks, quotes and keys repeated in every object. Minified JSON removes the whitespace, and CSV or YAML can remove repeated keys and most quotes. The JSON section on this page shows the exact saving for your data.