{ }jsonkit
LLM toolkit

Token counter for GPT, Claude and Gemini

Count LLM tokens as you type, see how text splits, and estimate API cost across 27 models.

  • OpenAI 18
  • Anthropic 4
  • Google 4
  • Meta 1
410 chars71 words5 lines410 bytesruns in your browser
Token streamo200k_base
Loading the tokenizer…
Hover a token to inspect its ID and bytes.

Compare models

Same text, 500 output tokens per request. Click a row to select it.
Model× 1K
est.
0% of 1.05M
—
est.
0% of 1.05M
—
est.
0% of 1.05M
—
est.
0% of 1.05M
—
0% of 1.05M
—
0% of 1.05M
—
0% of 1.05M
—
0% of 1.05M
—
0% of 1.05M
—
0% of 1.05M
—
0% of 400K
—
0% of 400K
—
0% of 1.05M
—
0% of 128K
—
0% of 128K
—
0% of 200K
—
0% of 200K
—
0% of 128K
——
est.
0% of 1M
—
est.
0% of 1M
—
est.
0% of 1M
—
est.
0% of 200K
—
est.
0% of 1.05M
—
est.
0% of 1.05M
—
est.
0% of 1.05M
—
est.
0% of 1.05M
—
est.
0% of 128K
——

About Token Counter

The token counter shows how many tokens a prompt, document or JSON payload uses in large language models such as GPT-6, GPT-5.6, GPT-4o, Claude, Gemini and Llama. Paste text or drop a file and the count updates as you type, together with the share of the model's context window and the API cost.

Counting happens in your browser. The tokenizer downloads once and then runs on your device, so private prompts, customer data and source code are never uploaded.

OpenAI models up to GPT-5.6 are counted exactly with their real tokenizers (o200k_base for GPT-4o through GPT-5.6, cl100k_base for GPT-4 and GPT-3.5). OpenAI hasn't published the GPT-6 tokenizer, and Anthropic and Google don't publish tokenizers that run in a browser, so GPT-6, Claude and Gemini counts are clearly labeled estimates.

How it works

Exact counts for OpenAI models

Text is split with byte pair encoding (BPE), the same algorithm and vocabulary OpenAI's tiktoken library uses. GPT-5.6, GPT-5.5, GPT-5.4, GPT-4.1, GPT-4o, o3 and o4-mini use o200k_base. Legacy GPT-4 and GPT-3.5 Turbo use cl100k_base. Special-token strings such as <|endoftext|> are counted as plain text, the way they arrive when pasted into a prompt.

Estimates for GPT-6, Claude, Gemini and Llama

For models without a browser tokenizer, the counter multiplies an OpenAI count by a calibration factor and marks the result with ≈. Claude's current tokenizer, introduced with Opus 4.7, uses roughly 1.0 to 1.35 times as many tokens as earlier Claude models, so current Claude models use 1.25 × o200k_base. GPT-6 and Gemini use 1.0 × o200k_base, and Llama 3 uses cl100k_base, which its vocabulary extends. Expect estimates to land within about 15% for English text and code. For billing-grade numbers, call the provider's token counting endpoint.

See every token

The token stream colors each token so you can see where the boundaries fall. Spaces, tabs and line breaks can be shown as ·, → and ↵. Hover a token to read its ID and byte length, switch to Token IDs to see the raw numbers, or copy every ID as a JSON array. An emoji or CJK character that is split across several tokens is shown as one piece.

Context window and cost

The gauge compares the prompt plus your expected output with the model's context window and warns when the total does not fit. The cost estimate uses each provider's standard list price per million input and output tokens. Set the output length, the number of requests and, where the model supports it, the share of the prompt served from cache. The comparison table runs the same numbers for all 27 models.

Spend fewer tokens on JSON

When the input is JSON, the counter also measures the same data as pretty-printed JSON, minified JSON, YAML and, for arrays of objects, CSV. Minifying removes whitespace tokens. CSV states each key once in a header row instead of in every record, which often saves the most. A version is only offered when it reads back as exactly the same data, so CSV is skipped for nested values, nulls or numbers stored as strings. Copy the cheapest version or load it into the editor.

Frequently asked questions

What is a token?

A token is the unit of text a language model reads and writes. Common English words are usually one token, while rare words, numbers, code symbols and non-Latin scripts are split into several. In English, one token averages about four characters, or roughly three quarters of a word.

How accurate is the token count?

OpenAI counts up to GPT-5.6 are exact because they use the same vocabularies as tiktoken. GPT-6, Claude, Gemini and Llama counts are estimates marked with ≈ and are usually within about 15% for English text and code. Unusual text, such as long runs of emoji or rare scripts, can drift further.

Why do Claude and GPT give different counts for the same text?

Each provider trains its own tokenizer with its own vocabulary, so the same sentence splits into different pieces. Claude's current tokenizer tends to produce more tokens than OpenAI's o200k_base for the same text, which affects both cost and how much fits in the context window.

How do I get an exact token count for Claude?

Use Anthropic's token counting endpoint, POST /v1/messages/count_tokens, which returns the exact input token count for a request and is free to call. Gemini offers a similar countTokens method in the Gemini API.

Does the count include chat formatting or system prompts?

No. The counter measures the text you paste. Chat APIs add a few tokens per message for role markers and formatting, and tools, images and system prompts count too. Paste your system prompt with the user message to include it.

Is my text uploaded anywhere?

No. Tokenizing runs in your browser. The page downloads the tokenizer data once, and after that nothing you type is sent to a server. Your last text is kept in this browser's local storage so it is still there when you come back.

How current are the prices?

Prices are standard API list prices per million tokens, checked against each provider's pricing page on the date shown below the estimate. Batch, regional and committed-use pricing differ, and providers change prices often, so confirm on the provider's page before you budget.

Why is JSON so expensive in tokens?

Pretty-printed JSON spends tokens on indentation, line breaks, quotes and keys repeated in every object. Minified JSON removes the whitespace, and CSV or YAML can remove repeated keys and most quotes. The JSON section on this page shows the exact saving for your data.

Sources & last verified

We checked this tool's behavior and the claims on this page against the specifications below on . The tool runs in your browser, so the output you see is exactly what the engine produced.

Found a wrong result or an outdated claim? Report an error and we will check it again.