AI TOOL

AI Agent Context Growth Calculator

See how an agent's prompt, previous responses and tool results can build up across model calls. Estimate peak context, total input tokens and the effect of summarizing older history.

Agent workflow assumptions

Count every model invocation, including calls made after tools return. Values describe a simplified repeated-history workflow.

Content sent on each model call before conversation history.

Tool calls count only when they trigger another model call.

Average response added to the history for the next call.

Average tool results produced after each call.

Use the content retained in the next model prompt.

Average new content added after each call.

Used to compare the estimated peak with a chosen limit.

Set to 0 to turn summarization off.

History size kept after each summarization point.

ESTIMATED PEAK CONTEXT

15,200 tokens

12% of the 128,000 token context window

15,200 of 128,000 tokens at peak

76,800total input tokens across calls
15,200peak without summarization
20,800estimated input tokens saved

Summarization reduces the estimated peak context from 15,200 to 10,000 tokens.

This planning estimate assumes retained history is sent again on every model call. It does not include summarization calls or overhead, provider prompt caching, retries, hidden framework prompts or tokenizer differences. Total input tokens are not the same as billed tokens, and context size alone does not predict cost.

HOW IT WORKS

Context grows when history is sent again

For each model invocation, the calculator adds the starting instructions and request to the conversation history carried into that call.

Input tokens per callstarting context + retained outputs + tool results + new user updates
Total input tokenssum of the input context for every model call

When a workflow includes several reasoning or tool steps, earlier messages may be included again in each later prompt. Peak context is the largest single prompt estimate. Total input tokens add the prompt sizes across all calls, so the total can exceed a model's context window even when each individual call fits.

Model summarization and pruning

Choose a summarization interval to replace accumulated variable history with the retained summary size after that many completed calls. This is an estimate of the next prompt after pruning. Real agent systems may summarize selectively, preserve important messages or use different context-management strategies.

Planning estimate, not provider billing

Providers can cache repeated prompt prefixes, count tokens with different tokenizers and add system prompts or tool schemas. This calculator estimates prompt input volume for a simple workflow. It does not reproduce provider billing, caching or a particular agent framework.

QUESTIONS

AI agent context calculator FAQ

Why can total input tokens exceed the context window?

The context window limits one model call. Total input tokens sum the inputs for every call in the workflow, including repeated history, so the cumulative total can be much larger.

Are tool calls counted as model calls?

No. Enter the number of model invocations. If an agent calls a tool and then invokes the model again with its result, count that next invocation and include the tool result tokens.

Does the estimate include summarization tokens?

No. The summary size represents history retained after pruning. The model call and input tokens needed to create a summary are not separately added.

Does total input tokens equal what an API bills me for?

Not necessarily. Prompt caching, provider-specific tokenization, hidden prompts, tool schemas, retries and pricing rules can change reported usage and charges.

How can context length affect GPU memory?

Longer context can increase KV cache memory during inference. Estimate model weights and runtime memory separately with the LLM VRAM Calculator.