Open table of contents
Conclusion
Separate weights, KV cache and runtime memory under explicit concurrency conditions.
Context
Weights fitting on a GPU do not establish capacity for long contexts or concurrent requests. Budget peak inference memory by component.
Design and verification scope
Assess the following responsibilities and boundaries when designing and verifying a configuration.
- Parameter count
- Precision
- Quantization
- Context length
- KV cache
- Batch
- Runtime
Weight memory is approximately bytes, where is the parameter count and the bits per weight. A single-GPU estimate including the KV cache separates these terms:
is concurrent sequences, layers, sequence length, KV heads, head dimension and bits per KV element. The factor of two accounts for keys and values. Quantization metadata, working buffers and fragmentation require additional capacity; measure peak memory with the intended runtime.
Decision rationale
Use the relationship between Parameter count and Runtime to compare the responsibilities of the selected approach and alternatives. Separate retained constraints from what the new boundary can change.
Trade-offs
Compare the implementation, maintenance and review work introduced by Precision with the control it provides. Include failure paths, operator effort and conditions in which the approach should not be adopted.
Limitations
Record versions, runtime, quantization, sequence length and peak memory by batch. Estimates are not measured performance.
Related case context
These cases provide attributed design context. They do not establish that the proposed experiments or configurations were delivered in those engagements.