Analysis
Size a local language model using a memory budget
Calculate weights, KV cache and headroom for a fictional local AI workload, then separate a memory estimate from evidence that a private pilot will work.
The short answer
Estimate raw weight storage, add explicit allowances for quantisation and runtime, calculate the KV cache for the maximum simultaneous sequences and total prompt-plus-output tokens, then reserve headroom in a consistent unit. A positive balance identifies a candidate for measurement; it does not prove loading success, useful speed, answer quality or security.
A private AI pilot needs a workload specification before it needs a hardware shortlist. “Seven billion parameters” describes neither the documents a team will submit nor how many requests will remain active. A small weight file can sit beside a much larger working cache. A successful load can still precede an unusably slow service.
This worksheet answers a narrower question: does a proposed inference workload remain inside a stated memory budget, under assumptions we can inspect? It uses a fictional dense decoder model, invented allowances and an assumed memory pool. No model was loaded, no device was measured, and no hardware purchase follows from the result. It extends the planning question in Unified’s own-hardware explainer with a calculation a reviewer can reproduce.
Specify the workload and the unit
Our fictional service allows two active sequences, each with up to 3,072 prompt tokens and 1,024 generated tokens: 4,096 total cached tokens per sequence. Count the complete tokenised input, including instructions, conversation history, retrieved material and formatting tokens. These are tokens, not words. Two active sequences means two resident request caches; it does not mean two registered users or merely a prefill batch size. We assume no beam expansion or shared-prefix saving.
The model has 7,000,000,000 parameters, 32 identical full-attention layers, 32 query heads, eight key-value heads and a head dimension of 128. Raw weights use four bits per parameter; cache values use FP16, two bytes each. The assumed device memory budget is 16 GiB, with every item below charged against that same pool. Host RAM and device memory are separate budgets; adding them together would conceal placement constraints.
Use bytes for arithmetic: 1 GB = 1,000,000,000 bytes; 1 GiB = 1,073,741,824 bytes. Thus 16 GB is about 14.901 GiB, not 16 GiB. Keep the original unit on each recorded specification. Our 16 GiB is an invented budget, not a claim about the free memory of a particular device.
Build the weight and cache ledger
Raw weight storage is parameters multiplied by bits divided by eight. Here that is 7,000,000,000 × 4 ÷ 8 = 3,500,000,000 bytes: 3.5 GB, or 3.259629 GiB. This idealised payload is not a download-size or resident-memory measurement. Real quantised formats need accounting for their representation; some tensors may retain another precision. Hugging Face’s bitsandbytes documentation, checked on 27 September 2026, describes separate dtypes for modules outside the quantised layers.
We therefore add an assumed 0.75 GiB delta above the raw payload for quantisation metadata and mixed-precision tensors. That number is deliberately visible and replaceable, not a universal percentage. Do not add the complete size of higher-precision tensors again if their low-bit baseline is already counted; add their extra cost, or replace the baseline with an exact tensor inventory.
For this conventional cache, the calculation is:
KV bytes = 2 × layers × KV heads × head dimension
× bytes per cached value × active sequences
× total cached tokens per sequence
= 2 × 32 × 8 × 128 × 2 × 2 × 4,096
= 1,073,741,824 bytes = 1 GiB
The factor of two represents keys and values. The dimensions follow the per-layer cache tensors in Hugging Face’s caching explanation; multiplying their dimensions and value width gives this worksheet’s formula. Use actual KV heads. The Llama configuration reference distinguishes them from query heads in grouped-query attention. Both references were checked on 27 September 2026; our configuration is fictional, not a named Llama model.
| Item | GiB | Evidence status |
|---|---|---|
| Raw weight payload | 3.259629 | Calculated from invented parameters and precision |
| Quantisation and mixed-tensor delta | 0.750000 | Assumed |
| KV cache | 1.000000 | Calculated under the stated cache model |
| Runtime working space | 2.000000 | Assumed; includes temporary buffers and allocator overhead |
| Estimated use, before reserve | 7.009629 | Sum of four items above |
| Uncommitted headroom | 2.000000 | Planning reserve, counted once |
| Total budget claim | 9.009629 | Estimated use plus reserve |
| Balance from 16 GiB | +6.990371 | Planning balance, not observed free memory |
The reserve is not another allocated tensor. Keep it separate so a reviewer can see how much margin the proposal intends to retain. Runtime space and headroom are different entries; neither has been measured. If a real budget already excludes background applications, do not subtract their memory a second time.
Make the fragile assumption fail on paper
Before reading the next table, predict which changes double the cache and which leave weight storage unchanged. For the long-context cases, reserve 24,576 prompt tokens plus 8,192 output tokens per sequence: 32,768 total. All comparisons keep the same invented weight, runtime and reserve allowances to isolate arithmetic sensitivity; actual runtime requirements may also change.
| Variation | KV GiB | Total GiB | Balance GiB |
|---|---|---|---|
| Four sequences; 4,096 tokens each | 2 | 10.009629 | +5.990371 |
| Two sequences; 32,768 tokens each | 8 | 16.009629 | −0.009629 |
| Four sequences; 32,768 tokens each | 16 | 24.009629 | −8.009629 |
| 32 KV heads instead of eight | 4 | 12.009629 | +3.990371 |
| One byte per KV value instead of two | 0.5 | 8.509629 | +7.490371 |
Doubling concurrency doubles this cache; multiplying per-sequence length by eight multiplies it by eight. Neither changes the raw weights. The second row exceeds the budget by about 9.860 MiB, where 1 MiB = 1,048,576 bytes. Rounding the total to “16.0” and accepting it would erase the failure. The stress case is short by more than eight GiB. Reducing the reserve until a column turns positive would change the policy, not validate the workload.
The 32-head row is an alternative architecture, not permission to reconfigure an existing model. The one-byte row is only an idealised sensitivity check: support, scaling data, residual higher-precision cache and quality effects need separate investigation. Four-bit weights do not automatically produce four-bit KV values. NVIDIA’s attention documentation, checked on 27 September 2026, describes cache precision choices and quantisation scaling separately.
Know what this formula leaves unresolved
Our formula assumes every layer retains every token with identical dimensions. Sliding-window, mixed-layer or compressed-cache architectures require their own accounting. Static caches may reserve their configured maximum immediately; dynamic caches grow; paged caches allocate blocks. Consult the actual runtime policy rather than treating current prompt length as allocated capacity. Hugging Face’s cache-strategy guide and NVIDIA’s paged-cache description, checked on 27 September 2026, explain these distinctions.
Loading is another phase. Temporary checkpoint storage, conversion and device placement can have a different peak from steady generation. Hugging Face’s large-model loading guide, checked on 27 September 2026, describes memory-aware loading and sharded checkpoints. It also documents moving weights between device, CPU and disk. Offloading changes where bytes reside and introduces transfers; it does not create a single interchangeable memory pool.
Prefill processes the prompt; decoding produces subsequent tokens. Their temporary allocations depend on the backend, as NVIDIA’s attention documentation illustrates. Our fixed 2 GiB runtime allowance proves nothing about either peak. Nor does a positive balance predict response time: a system can remain inside memory limits while transfers, computation or queuing make it unsuitable. This inference worksheet excludes training gradients and optimiser state, and does not cover every architecture or multi-device placement.
Turn the estimate into an acceptance record
Keep the worksheet with a separate measurement record. For this article, its fields are deliberately unfinished:
| Record field | What a later trial must capture | Current result |
|---|---|---|
| Model identity | Checkpoint revision, tensor inventory, tokenizer and chat-template revisions | NOT SELECTED |
| Execution configuration | Device, driver, runtime, attention backend, weight/cache formats, allocation and offload policies | NOT SELECTED |
| Memory by phase | Available bytes before loading, loading peak, resident use, prefill/decode peaks and host RAM peak; tool and metric definition | NOT MEASURED |
| Workload and lifecycle | Actual input/output lengths, active sequences, cancellations, repeated-run memory and repetition count | NOT MEASURED |
| Service behaviour | First-token time, subsequent-token latency, completion time and failures under the specified load | NOT MEASURED |
First, choose a concrete configuration and verify that its architecture supports the requested context. Set workload-specific response-time limits, a measurement method and a headroom requirement before a trial. Record whether a reading is device-wide usage, process allocation or reserved allocator space; those quantities answer different questions.
Next, replace allowances with documented estimates, then measure controlled loading, maximum input and output, and simultaneous requests using approved test data. Include release of caches after completion or cancellation and repeat the workload. Keep raw units, observed peaks, repetition counts and failures. If limits are exceeded or evidence is missing, retain HOLD and revise the workload or implementation. Memory acceptance remains separate from answer-quality, access-control and data-handling acceptance for a private deployment.
Compare an observed total peak against the estimated total use and the reserve requirement. Do not add that entire peak as an “overhead” on top of weights and KV: it already contains overlapping costs. Replacing one allowance requires an isolated incremental measurement or a justified component breakdown for the same phase. Our assumed 2 GiB runtime entry was not derived from a framework’s peak reading.
For further practice, Trust Agent’s Level 3 GPU memory maths course expects multiplication, division and familiarity with parameters and precision; its calculator requires a terminal, with Which AI model? listed as prior study. This article supplies a paper calculation without requiring a download. Reusing the worksheet should produce a defensible test proposal, not an unsupported assurance that a model will run well.
Contribution and ownership: This AI-assisted piece is credited to Mickarle Wagstaff-Irons - Micky Irons, full name Mickarle Sean Junior Wagstaff-Irons. Unified, Trust Agent and Mickai share ownership. Mickai’s AI readiness programme is an optional commercial route for discussing a pilot, not independent evidence for this worksheet or a guarantee of deployment suitability. Confirm its current scope separately.
Questions readers ask
- Do 4-bit model weights imply a 4-bit KV cache?
- No. Weight storage and cache precision are separate choices. This fictional worksheet uses 4-bit raw weights and FP16 cache values at two bytes each; a different cache format needs explicit backend support and its own overhead accounting.
- Does context mean prompt tokens or prompt plus generated tokens?
- Here it means the total reserved cache length per active sequence: tokenised prompt plus maximum generated output. Two simultaneous sequences each reserve their own cache. Static or paged allocation can require a different accounting of the actual reserved capacity.
- Does a positive memory balance mean a model is ready to deploy?
- No. It is a planning result under stated assumptions. Loading and runtime peaks, response times, workload quality, data handling and operational controls still need separate evidence.
Published by Mickai LTD.
Unified covers the field broadly and treats Mickai as one example within it. About the journal and the team.
Mickarle Wagstaff-Irons - Micky Irons, full name Mickarle Sean Junior Wagstaff-Irons. Article author. Biography and related work.