Open table of contents
Conclusion
Keeping data local leaves capacity, updates, monitoring and recovery as operational responsibilities.
Context
For sensitive workloads, deployment choices depend on ownership of inference, logs, model updates and incident response as well as data location.
Design and verification scope
Assess the following responsibilities and boundaries when designing and verifying a configuration.
- Data sovereignty
- GPU capacity and concurrency
- Inference
- Model lifecycle
- Monitoring and security
- Cloud boundaries and cost
Decision rationale
Use the relationship between Data sovereignty and Cloud boundaries and cost to compare the responsibilities of the selected approach and alternatives. Separate retained constraints from what the new boundary can change.
Trade-offs
Compare the implementation, maintenance and review work introduced by GPU capacity and concurrency with the control it provides. Include failure paths, operator effort and conditions in which the approach should not be adopted.
Limitations
Record GPU, quantization, context length, concurrency, latency and cost conditions. Local hosting alone does not guarantee compliance or prevent data leakage.
Related case context
These cases provide attributed design context. They do not establish that the proposed experiments or configurations were delivered in those engagements.