token
A token is a unit of text processed by an AI model; it does not necessarily equal one word or character.
A token is a unit of text used by an AI model to process input and produce output. Depending on how text is split, a token may be a whole word, part of a word or punctuation; token counts therefore differ from word and character counts.
Token counts are used to describe context windows, usage and API pricing. In design systems, a token instead means a variable defining a value such as a color or spacing measurement.
This entry is based on AIPOST articles and widely known facts. If something is wrong, please send us a correction request.
Articles covering this entry
Price per million tokens Haiku 5.5, prompts up to 100K tokens Haiku 5.5, above 100K Haiku 4.5 Sonnet 5.5
The watermark works by adding a slight statistical bias to how the model picks tokens, the small chunks of text a language model reads and writes.
Total context used About 77,000 tokens About 8,700 tokens
In a nine-part test it found a hidden passkey in all 15 trials across a 256k-token context, passed 75 of 100 coding problems, and built working web apps and gam…
…an image along with a fixed list of possible answers, and it returns a probability for each option in about 150 milliseconds, with no output tokens to pay for.
It takes text in and predicts the most likely next token, the small chunk of text a model reads and writes.
Context tokens used About 84,000 About 129,000
Pretraining teaches a model to predict the next token, the small chunk of text a model processes, from the tokens before it.
Memories and Chronicle use more tokens, the units of text a model processes, but the team considers the extra cost worth it for an agent that remembers a projec…
…ace is a structured, nested record of one agent run: every conversational turn, tool call, prompt, model response, token count, latency figure, and dollar cost.
Time to first token 0.25 seconds
…hey run, Collison says nearly every modern software company is moving toward token credits, overage fees and tiered metering, much like ChatGPT and Claude, whic…
One model instance produces roughly 50 to 60 tokens per second, which is fine for a quick answer but too slow for multi-step engineering work.
…e copy such as "Welcome to our store, we are a family-owned business" adds irrelevant tokens, the small units of text a model processes, and confuses retrieval.
Muse is offered for free with generous token allowances (tokens are the units of text an AI model processes) to build a large user base.
The model's context window, the amount of information it can consider at once, is 128,000 tokens.
…anguage and reinforcement learning loop over several turns, using more than 2 million tokens, the units of text a model processes, without losing sandbox state.
…many models, open-weights models that anyone can download accounted for 55.8% of token volume from July 2024 to October 2025, ahead of 44.2% for closed models.
Ordinary use is free with an allowance of 100 million tokens a week; tokens are the small units of text an AI model processes.
Argon's introductory rate is $2 per million input tokens and $10 per million output tokens.
A 1-million-token output limit and introductory pricing
Because these models mainly predict the next token, the small unit of text a model processes, their reasoning cannot be fully checked against ground truth.
Then, on July 10, an agent announced that it had found valid Hugging Face write tokens exposed on the public internet.
After a credit-based phase, it moved to cost-plus pricing: tokens, compute, and third-party APIs are billed at cost, plus a separate orchestration fee.
Prompt caching, which stores content the model has already seen so it can be reused, cuts the cost of those input tokens by about 95 percent.
…5.5 and Meta's new Muse agent, the answer appears to lie less in leaderboard rankings and more in how organizations route work, track tokens and govern agents.
Most token waste comes from habits: buying credits without checking the meter, letting an agent click through websites like a person, running every task on the…
A typical development machine holds unlocked SSH keys, cloud credentials, personal access tokens and unrestricted internet access.
Faster responses come from a rewrite of the Responses API backend aimed at time to first token and time between tokens.
A token is a small unit of text a model processes or produces.
Usage limits and referral tokens
Its listed API prices are $2 per million input tokens and $10 per million output tokens, matching the prices given for Claude Sonnet 5.5.
The first raises the question of whether teams can keep their work in one place; the second, whether faster output is worth faster token consumption.
Alongside that change, it revised how plan usage and token consumption are calculated.
For developers, the most concrete published figure is $0.10 per million cached input tokens.
…multiple agents and continuously producing shared project information can also increase token use (tokens are the units of text a model processes) and latency.
Claude Sonnet 5.5 leads Claude Opus 5.5 on one coding benchmark, and its input and output token prices are half as high.
Its API prices are quoted per million tokens, the units used to measure text processed or generated by a model:
Tokens are units of text that an AI model processes, and token spending is one way to pay for that computing.
The repeated task illustrates the intended workflow, but it does not supply a measured time or token saving for that Reddit example.
On OpenRouter, Jev was listed at about $0.042 per million input tokens and $0 for output tokens.
The run used Gemini 3.5 Flash, processed 2.6 billion tokens and cost under $1,000 in API credits.
A token is a unit of text a model processes.
Higher effort can lengthen a turn and increase output tokens, the units of text that affect usage costs.
…used for that sync receives read access to the specific Skills database, and its token is stored as in GitHub Actions secrets rather than in a repository file.
The service grants five tokens on registration and charges one token per battle.
Across the evaluation, Codex processed 21.9 billion tokens—the units used to measure model input and output—against Claude Code’s 2.9 billion, a 7.5-fold differ…
Model token cost Approximately €300 per month
For comparison, the estimated cost for Claude Opus 4.8 at an identical token volume was $61.70.
In a presidential-debate test, it made 1,191 Jev calls across about 1.18 million input tokens.
Its profile area displays measures such as lifetime token use, longest task and daily activity in a layout that closely matches the Codex desktop application.
Language models calculate likely next tokens, while mnemonic methods link material to images and places.
Reviewing those records can reveal habits that waste tokens or repeatedly interrupt a task.
Yutori lists Navigator n2 API prices of $0.50 per million input tokens, $0.05 per million cached input tokens and $4.00 per million output tokens.
Jev’s listed OpenRouter price in September 2026 was $0.042 per million input tokens and $0.00 per million output tokens.
It did not always use fewer tokens, however, and the tests do not establish its cost per task.
The reported run involved 2.7 million messages and 130 billion output tokens.
…inexpensive NVIDIA RTX 3060 can run an 8-billion-parameter model at 40 to 50 tokens per second, while more expensive systems can fit substantially larger models…
A faster first token is not faster writing at the same rate
…clean policy document sat alongside a draft FAQ containing ordinary business text and a concealed instruction to reveal internal instructions and a test token.
The cost question is bigger than token prices