AI Agent Context Growth Calculator
See how an agent's prompt, previous responses and tool results can build up across model calls. Estimate peak context, total input tokens and the effect of summarizing older history.
15,200 tokens
12% of the 128,000 token context window
15,200 of 128,000 tokens at peak
Summarization reduces the estimated peak context from 15,200 to 10,000 tokens.
This planning estimate assumes retained history is sent again on every model call. It does not include summarization calls or overhead, provider prompt caching, retries, hidden framework prompts or tokenizer differences. Total input tokens are not the same as billed tokens, and context size alone does not predict cost.
Context grows when history is sent again
For each model invocation, the calculator adds the starting instructions and request to the conversation history carried into that call.
When a workflow includes several reasoning or tool steps, earlier messages may be included again in each later prompt. Peak context is the largest single prompt estimate. Total input tokens add the prompt sizes across all calls, so the total can exceed a model's context window even when each individual call fits.
Model summarization and pruning
Choose a summarization interval to replace accumulated variable history with the retained summary size after that many completed calls. This is an estimate of the next prompt after pruning. Real agent systems may summarize selectively, preserve important messages or use different context-management strategies.
Planning estimate, not provider billing
Providers can cache repeated prompt prefixes, count tokens with different tokenizers and add system prompts or tool schemas. This calculator estimates prompt input volume for a simple workflow. It does not reproduce provider billing, caching or a particular agent framework.
AI agent context calculator FAQ
Why can total input tokens exceed the context window?
The context window limits one model call. Total input tokens sum the inputs for every call in the workflow, including repeated history, so the cumulative total can be much larger.
Are tool calls counted as model calls?
No. Enter the number of model invocations. If an agent calls a tool and then invokes the model again with its result, count that next invocation and include the tool result tokens.
Does the estimate include summarization tokens?
No. The summary size represents history retained after pruning. The model call and input tokens needed to create a summary are not separately added.
Does total input tokens equal what an API bills me for?
Not necessarily. Prompt caching, provider-specific tokenization, hidden prompts, tool schemas, retries and pricing rules can change reported usage and charges.
How can context length affect GPU memory?
Longer context can increase KV cache memory during inference. Estimate model weights and runtime memory separately with the LLM VRAM Calculator.