Open table of contents
Conclusion
Control how failures propagate into external actions, beyond model response quality.
Context
Tool-using agents turn decisions into data changes and notifications. The commitment boundary and post-failure state need design attention alongside generated responses.
Design and verification scope
Assess the following responsibilities and boundaries when designing and verifying a configuration.
- Identity and authorization
- State and replay
- Tool execution and human approval
- Observability
- Evaluation
- Failure recovery
Decision rationale
Place authority, state, execution, observation, evaluation and recovery along the path that changes external state. Identity supplies the acting principal; human approval controls commitment. Design state and idempotency together to prevent duplicate actions on replay.
Trade-offs
Approvals and intermediate checks add latency and review work. Explain their auditability and containment benefits against reduced automation. Human approval can still rely on bad context, so review the evidence presented to approvers and the evaluation data.
Limitations
Review permission matrices, state transitions, approval points, failure injection and evaluations. The SaaS case supports integration context, not AI delivery or improvement claims.
Related case context
These cases provide attributed design context. They do not establish that the proposed experiments or configurations were delivered in those engagements.