AI TOOL

LLM VRAM Calculator

Estimate AI model weight memory, KV cache and total runtime VRAM. Adjust the model architecture and context assumptions to make a more useful GPU memory plan.

Model and runtime assumptions

Defaults describe one example 8B model. Replace architecture values with the model configuration when known.

For an 8B model, enter 8.

Real formats vary and include metadata.

Use the model config value.

Grouped-query models often have fewer KV than attention heads.

Per-head key/value width.

KV cache grows with context length.

Also called batch size.

Runtime support and cache format vary.

For buffers, workspace and other allocations.

Enter memory available to the workload.

Some runtimes support CPU offload.

APPROXIMATE RUNTIME ESTIMATE

6.5 GB

Estimated runtime VRAM including weights, KV cache and allowance

8 GB GPU VRAM compared with 6.5 GB estimated runtime VRAM

4.5 GBapproximate model weights
0.5 GBestimated KV cache
1.5 GB freeVRAM margin against estimate

The estimate adds model weights, KV cache and the runtime allowance you entered.

This is a planning estimate, not a guarantee that a model will run on a given GPU. Actual use depends on model details, runtime, context, batch size, buffers, other allocations and offload support.

HOW IT WORKS

Weights and runtime VRAM are different

The calculator shows the approximate weight memory separately from an estimated runtime total.

Approximate model weightsparameters in billions × bits per weight ÷ 8 ≈ decimal GB

For example, an 8B model at an estimated 4.5 bits per weight has about 4.5 GB of weights. Quantized files include scales, metadata and other values, so loaded weights can be larger than this simple estimate.

Estimated KV cache2 × layers × KV heads × head dimension × context tokens × sequences × cache bytes ÷ 1,000,000,000

The KV cache holds attention state while generating text. Longer context and more concurrent sequences increase its size. Defaults are illustrative; use the model configuration and runtime cache settings when available.

Runtime estimate is not a fit guarantee

The displayed runtime estimate adds approximate weights, the calculated KV cache and your other-memory allowance. Framework workspace, activations, quantization metadata, vision/audio components, fragmentation and other GPU workloads can change actual use. Some runtimes support CPU offload, but system RAM does not guarantee that offload is supported or fast.

QUESTIONS

LLM VRAM calculator FAQ

Does this tell me if a model will run on my GPU?

No. It estimates weights and selected runtime memory assumptions. Model architecture, runtime allocations, available memory and offload support determine whether inference works.

Why show weights separately from runtime VRAM?

Weights are only one part of memory use. The KV cache and runtime buffers add to the memory needed during inference, so weight size alone can understate GPU needs.

How accurate is the KV cache estimate?

It uses layer count, KV heads, head dimension, context length, sequence count and cache precision. Architectures and runtimes can use different cache layouts or optimizations, so treat the result as an estimate.

Why can real quantized weights be larger?

Quantized formats may store scales, metadata and some higher precision values. The selected bits per weight is an average planning assumption, not an exact file size.

Can system RAM make up for insufficient VRAM?

Some runtimes can offload layers or cache to system RAM. Support and speed depend on your hardware, model and software. This calculator cannot confirm compatibility.