Open table of contents
When an AI agent operates a CRM, mailbox, calendar or ticket system, automation expands from generating answers to performing work. A shared API token leaves important questions unresolved: whose authority is used, what was approved, and when did the external state actually change?
“Send this customer a follow-up” involves separate decisions about reading the customer, using the sender’s mailbox, drafting the message and approving this particular send. If the request times out, the system must also distinguish a failed send from a lost response.
This article separates authority, capability, approval, execution, finality and audit. It develops Six Designs for Production AI Agents and System Boundaries Before Adding AI into concrete SaaS execution decisions. Examples are design and review materials; SaaS writes, authentication, SDK or SQL execution and fault injection have not been performed.
1. Separate requester, executing actor and approver
Consider an agent using a shared administrator credential to search customers on a user’s behalf. If the user can see 20 customers but the agent can read 20,000, the agent becomes a privilege-escalation path. Filtering the final answer to 20 records is insufficient if unauthorized information has already entered model context, caches or traces.
Microsoft Graph distinguishes delegated and app-only access. Delegated applications act on behalf of a signed-in user and cannot access information that user could not access. App-only applications use their own permissions without a signed-in user. [1]
| Principal | Meaning to record | Trusted source |
|---|---|---|
| Requester / subject | Whose request and delegation authorize the work? | Validated session and delegation |
| Agent / workload actor | Which runtime and deployment propose or execute it? | Workload authentication and deployment management |
| Approver | Who approved this concrete operation? | Authentication and authorization at approval time |
| SaaS principal | Which user, application and account does the provider recognize? | Actual credential and provider evidence |
Do not accept user ID, tenant ID or connector account directly from model-generated tool arguments. Bind them from authenticated context and verify their relationship to the resource. An agent name in a log is not actor authentication either. Separating actor and subject is a useful internal model; do not assume every token contains an identically structured act claim.
Scroll horizontally to view the complete diagram.
Read the diagram as text
The agent proposes capabilities and arguments. The gateway checks authenticated context, policy and concrete approval. The executor uses the broker and journal to call SaaS, which retains authorization; outcomes require evidence.
2. Use separate workflows for delegated and app-only access
Adding an appointment to a user’s calendar calls for user delegation. A nightly organizational ticket report may need a dedicated workload identity. Even then, constrain tenants, projects, mailboxes and operations. App-only does not inherently mean all data: the provider’s application permissions and resource-level restrictions determine its effective reach. [1]
Microsoft’s On-Behalf-Of flow lets a middle-tier API obtain a downstream delegated token representing a user. It works with user principals; it does not turn an app-only token into user delegation. The incoming token’s audience must also match the middle-tier API. [2]
RFC 8693 Token Exchange and RFC 8707 Resource Indicators help express delegation and intended resources. Not every SaaS supports them, and they are not interchangeable with Microsoft’s OBO protocol. Asking a broker for narrower scopes does not establish that the issued token has the expected limits; inspect the actual issuance and authorization contract. [3][4]
Mail and calendar operations exposed through Microsoft Graph do not necessarily have distinct audiences just because the products have different names. Combine resource audience, OAuth scopes, mailbox permissions and gateway capability restrictions. Internal capability names such as customer.read are not necessarily real OAuth scopes.
A delegated 403 must not trigger fallback to an administrator app token. Treat it as denial of that workflow. App-only automation has a separate identity, policy and audit path.
3. Expose business capabilities and constrain reads too
Prefer read_customer, prepare_followup and send_customer_email over an unrestricted generic_http_request. An adapter translates the capability into a provider API; the gateway validates schema, tenant, resource, connector account and policy.
# Proposed contract; provider behavior must be reviewed per endpoint.
send_customer_email:
authority: delegated
model_may_choose_principal: false
risk: external_side_effect
approval: exact_prepared_action
credential_owner: token_broker
retry: reconcile_before_retry
provider_idempotency: NOT_VERIFIED_DO_NOT_ASSUME
acceptance: graph_sendMail_202_without_response_body
completion_evidence: DEFINE_FOR_BUSINESS_OUTCOME
compensation: no_general_unsend_guarantee
execution_status: NOT_RUN
This is a tool-contract example. The backend derives risk and approval requirements from the target, amount, visibility and environment; the model does not declare them. Do not permanently classify every CRM note as reversible either. If creating it triggers notifications or webhooks, deleting the note cannot retract those effects.
Reads need delegated authorization, resource scope, query restrictions, result limits and field projection. Minimize the DTO sent to the model instead of retrieving unauthorized data and hiding it later. Apply the same tenant and user boundaries to caches, vector retrieval and attachments.
SaaS documents and ticket bodies are untrusted data. Text asking the agent to export a token or “continue as already approved” cannot modify authority. Address prompt injection through constrained capabilities, destinations, authorization and approval enforcement, alongside model-level mitigations. [13]
4. Freeze execution content during preparation
Prepare stores a concrete action; Commit executes that stored action. These are application workflow phases, not a distributed two-phase commit protocol involving the SaaS provider.
Freeze the sender mailbox, To/Cc/Bcc, subject, body content and format, Reply-To, immutable attachment versions and hashes, tenant, connector account, operation ID, target, expected revision and expiry. Financial operations also bind amount, currency and beneficiary.
{
"schema_version": "agent-action/v1",
"operation_id": "00000000-0000-4000-8000-000000000001",
"tenant_id": "tenant-example",
"subject_id": "user-example",
"actor_id": "followup-worker-example",
"authority_mode": "delegated",
"connector_account_id": "mailbox-connection-example",
"capability": "send_customer_email",
"target_id": "customer-example",
"expected_revision": "crm-revision-example",
"policy_version": "policy-example-v1",
"expires_at": "2026-09-24T12:10:00Z",
"payload": {
"from_mailbox": "sales@example.com",
"to": ["customer@example.net"],
"cc": [],
"bcc": [],
"reply_to": [],
"subject": "Follow-up",
"body_type": "Text",
"body": "Thank you for the meeting. May we discuss the next steps?",
"attachments": []
}
}
Compute an action fingerprint from the validated execution envelope after resolving defaults. One option is RFC 8785 JSON Canonicalization Scheme, followed by SHA-256 with a domain separator and schema version. Hashing arbitrary serializer output can introduce differences in property ordering or number representation. Reject ambiguous inputs such as duplicate keys. [7]
# Algorithm contract, not an implemented hashing function.
canonical_bytes = JCS(validated_execution_envelope)
fingerprint = SHA256(UTF8("agent-action:v1") || 0x00 || canonical_bytes)
# Store separately; the fingerprint is not a field in its own input.
approval = (tenant_id, operation_id, fingerprint, approver, expires_at)
A hash is not approver authentication or a signature. Bind the approved digest, approver, approval ID and expiry in a protected server-side record. Comparing hashes is meaningless if the client can replace both the action and its hash.
Show the concrete recipients and content, not just the hash. Refetching an attachment from a mutable URL or expanding a recipient group at Commit can change what the person approved. Payload changes require new approval.
Scroll horizontally to view the complete diagram.
Read the diagram as text
Prepare freezes tenant, subject, actor, recipients, content, attachment versions and revision. An authorized reviewer approves the content; the server binds approval to its fingerprint and operation. Commit checks current authority and the same saved action.
5. Manage approval as durable state
Official OpenAI documentation describes returning an approval interruption, retaining state and resuming the same run. This supplies a pause/resume primitive. The application still implements who may approve, what they approve and when that approval expires. [6]
Send the browser only the display data and an opaque approval ID. Keep complete run state, tokens and internal context in server-controlled storage. This storage arrangement is the design proposed here; serialized state and an unpredictable ID are not substitutes for authentication and authorization.
The approval endpoint checks session, tenant, required role, operation and protections such as CSRF handling. Define whether self-approval is permitted or requester/approver separation is required. Rejection, expiry, cancellation, retries and concurrent clicks are state transitions. Approval must not be reusable for a different operation.
High-impact actions can require step-up authentication, but stronger authentication is distinct from approving an action. Bind step-up evidence to the approval session, and approval to the concrete envelope. A framework example that automatically approves pending interruptions is not a production human-review implementation.
6. Reauthorize before execution and enforce concurrency at the API
Between approval and execution, the user may be disabled, customer ownership may change, policy may be updated or mailbox permission may disappear. Immediately before Commit, recheck current authorization, approval validity, the saved fingerprint and target state. Approval does not replace permission.
However, GET → compare revision → PATCH leaves a race between the comparison and update. When supported, send the approved revision with If-Match or the provider’s equivalent so the SaaS evaluates the condition with the write. HTTP defines conditional requests and 412 Precondition Failed; verify actual endpoint support. [8]
# Conceptual CRM endpoint; verify support before adopting an adapter.
PATCH /opportunities/example
If-Match: "approved-revision"
Content-Type: application/json
{"stage":"closed_won"}
# Precondition mismatch -> reload and re-plan; no blind PATCH retry.
# This is not an If-Match example for Microsoft Graph sendMail.
A revision mismatch returns to retrieval and planning, not blind retry. Changed state or action content requires renewed approval. Comparing updated_at without provider-enforced conditional writes does not provide equivalent protection. If the residual race is unacceptable, exclude that operation from automatic Commit or provide a stronger domain API.
Authorization revalidation and the external commit are not one distributed transaction either. Combine downstream SaaS authorization, a short execution window and supported revocation controls, and document residual propagation delays. A local policy check alone cannot guarantee that every request stops immediately when a user is disabled.
7. Journal execution ownership and crash outcomes
Assign a server-generated operation ID to each side effect and persist it with the tenant. Retries of the same execution request return to the same operation. A new UUID on every attempt defeats deduplication. Conversely, do not use a payload hash to suppress all intentionally repeated actions, such as sending the same content on a later day.
Atomically claim an approved operation locally. The following PostgreSQL fragment illustrates that claim. It still needs current authorization, approval-record verification, immutable prepared content and an outbox; it is not a complete execution engine.
-- PostgreSQL design fragment: $1 tenant, $2 operation, $3 digest (bytea).
-- Only a trusted executor may call this after current policy/approval checks.
UPDATE agent_operations
SET state = 'executing',
attempt = attempt + 1,
claimed_at = CURRENT_TIMESTAMP
WHERE tenant_id = $1
AND operation_id = $2
AND state = 'approved'
AND action_digest = $3
AND approved_by IS NOT NULL
AND approval_expires_at > CURRENT_TIMESTAMP
AND action_expires_at > CURRENT_TIMESTAMP
RETURNING operation_id, canonical_action, attempt;
-- Zero rows: do not dispatch. Commit this local claim before calling SaaS.
-- Unresolved executing attempts are NOT automatically returned to approved.
If the update returns zero rows, that worker does not send. Commit the executing transition before initiating the external request. This limits concurrent local consumption of one approval, but the SaaS side effect and the database result cannot be committed in one transaction.
The provider may finish immediately before the worker crashes while recording the response. A journal still showing executing does not prove nothing was sent. Reassigning work on lease expiry can also duplicate a request from the original worker that is still alive. Route unresolved attempts to reconciliation unless the provider’s deduplication contract safely permits retry.
Scroll horizontally to view the complete diagram.
Read the diagram as text
Atomically claim an approved operation before calling SaaS. Acceptance is distinct from business completion. Lost responses or worker crashes lead to unknown; do not resend without reconciliation under the provider contract.
Reserve failed for failures known, under the endpoint contract, to have produced no effect. A timeout after possible dispatch is unknown. Do not turn uncertainty into failure or tell the customer the message was definitely not sent.
8. Verify idempotency for each provider operation
A local operation ID does not prevent external duplicates unless the SaaS uses a corresponding identifier under a defined deduplication contract. Design local atomic claims, provider idempotency and reconciliation separately.
| Example operation | Documented contract | Design consequence |
|---|---|---|
Microsoft Graph sendMail | 202 Accepted, no response body; processing may be incomplete | Do not assume a returned message ID or a generic idempotency key |
| Microsoft Graph event creation | transactionId is intended to avoid redundant POSTs on retries | Reuse it for the same creation; verify unspecified retention and related behavior separately |
| Supported Stripe POSTs | Store and reuse the result for an idempotency key | Observe parameter matching, key retention and endpoint conditions |
| Local adapter journal | Identifies operations and stores local state | Does not establish external exactly-once execution |
A Graph sendMail success response does not necessarily give you a SaaS message ID. Client and provider request IDs are not message IDs or delivery evidence. A draft-first design still needs separate scope, draft-mutation and send-reconciliation decisions. [9]
Calendar transactionId supports create retries, not a universal guarantee for every update and deletion. Stripe allows keys to be removed after they are at least 24 hours old; reuse after removal creates a new request. Different parameters with the same key cause an error, and saved 500 responses are replayed too. A provider key is not an indefinite deduplication record. [10][11]
Record the account, endpoint, identical payload, key lifetime and conditions under which a response is stored. Do not fill unknowns with provider_supported: true. The email example in this article assumes no verified provider-side idempotency contract.
For 429 responses, follow provider instructions such as Retry-After and enforce backoff, jitter, concurrency and retry budgets in the runtime. First classify whether the failed attempt could have caused an effect. Apply the same operation management when a model loop proposes another tool call as a retry. [12]
9. Separate acceptance, completion, delivery and compensation
An API success is not necessarily the final business result. Email acceptance, provider processing, delivery and recipient reading are distinct states. Record Graph 202 as acceptance evidence, not a delivered message. [9]
| State | What is known | Next step |
|---|---|---|
prepared / approved | Content is prepared or approved | Verify validity and claim execution |
executing | Local execution ownership was claimed | Await the external outcome |
accepted | Provider accepted the request | Confirm the required business outcome separately |
completed | Evidence establishes the defined outcome | Preserve the outcome definition and evidence |
failed | A no-effect failure is established | Correct or retry under policy |
unknown | An external effect remains possible | Reconcile; do not automatically resend while unresolved |
Finality is the boundary where a difficult-to-retract external effect occurs. Define cancellable stages and the evidence required to report completion for each capability. Cancellation and refund do not erase history: they are additional effects requiring their own authorization, approval, operation ID and audit. Calendar cancellation can send another external notification.
For multi-step workflows, perform validation and reversible preparation earlier and difficult-to-retract actions later. Reordering does not make the whole workflow atomic. Record partial success and the compensation decision.
10. Keep credentials and browser actions behind the execution boundary
Do not put credentials in prompts, tool results or traces. The model sees capability names and validated DTOs; an authenticated executor calls the token broker. The broker must not issue a token merely because a caller supplied a user_id. It validates delegation, tenant, connector, required scope and current policy.
Keep refresh tokens in the broker’s secret system. Give workers only necessary short-lived credentials, or perform provider calls inside the broker/adapter boundary. Workload federation may be an option where supported. Credential lifetime does not need to equal agent-run lifetime.
DPoP binds tokens to a client key and requires authorization-server and resource-server support. It can reduce misuse of a stolen token alone, but does not solve compromise of an executor that can use the key. Audience and destination validation and secure credential custody remain necessary. [5]
Browser cookies and profiles can carry broad authority. A prompt saying “ask before clicking Send” is insufficient when the model can freely click, execute JavaScript or issue network requests within the same session. Those routes may bypass an approval tool.
Prefer APIs or domain capabilities with enforceable backend authorization for high-impact operations. When only a UI is available, evaluate browser isolation, fixed accounts, action restrictions, concrete-content review and whether Commit can actually be controlled. If the boundary cannot be enforced, have a human perform the final action rather than claiming a prompt implements it.
11. Connect audit and evaluation to execution sequences
Audit tenant, requester, actor, capability, target, authority mode, connector account, policy version, approval, fingerprint, execution time, attempt, provider request ID and outcome evidence. Exclude access and refresh tokens. A digest of predictable content does not guarantee confidentiality either; control access and retention.
Recording a local state transition and an audit outbox event in the same database transaction reduces the chance of losing the event before delivery to the audit system. It still does not make the SaaS operation atomic. Do not place the business journal solely in sampled traces; audit delivery also needs retry and deduplication.
Use the operation ID to connect traces, the journal and provider records. If no external resource ID exists, retain available request IDs, time, account and reconciliation evidence with its confidence. Do not invent a correspondence.
| Negative or failure case | Expected boundary | Current status |
|---|---|---|
| Cross-tenant or unauthorized customer request | Deny before exposing data to the model | NOT_RUN |
| A document instructs policy bypass | Treat it as untrusted data, without modifying policy | NOT_RUN |
| Client changes recipient or approval state | Reject against the server-controlled action | NOT_RUN |
| Another user’s approval ID, expiry or replay | Reject using reviewer, tenant, operation and state | NOT_RUN |
| Permission revoked after approval | Reauthorize at Commit and document propagation limits | NOT_RUN |
| Resource changed after approval | Detect conflict through provider conditional update | NOT_RUN |
| Two workers claim one operation | Only one obtains local execution ownership | NOT_RUN |
| Provider processes, then response or worker is lost | Reconcile unknown; no automatic resend | NOT_RUN |
| Provider idempotency retention elapsed | No blind retry after the guarantee has lapsed | NOT_RUN |
| Delegated denial triggers app-only fallback | Deny without elevating authority | NOT_RUN |
Evaluate tool choice, actual principal, approved envelope, state transitions and external outcome alongside final-answer quality. A fake adapter exercising the intended sequence and a live provider exercise establishing idempotency, revocation or delivery are different evidence.
12. Start with one capability’s operating contract
Download the capability contract, action schema, SQL fragments and acceptance materials. These are review artifacts, not a complete agent framework or SaaS connector. Fingerprinting implementation, authentication/authorization, approval UI, provider adapters and outbox workers are not included. All execution cases remain NOT_RUN; no SDK, SQL or SaaS validation is claimed.
Scroll horizontally to view the complete diagram.
Read the diagram as text
Use the operation ID to connect requester/actor, approved content, provider attempts and outcome evidence. Exercise revocation, concurrent updates, replay and timeouts as well as positive cases. These plans remain NOT_RUN.
Begin with one capability, such as sending one follow-up to a customer the requester owns. Define authority and readable fields, freeze the prepared envelope, show what approval covers, implement Commit-time reauthorization and specify timeout reconciliation. Assign an owner for unresolved outcomes and provide capability/account suspension controls. Stopping future execution does not cancel a request already sent.
The model interprets intent, prepares candidates and proposes capabilities. The backend determines whose authority is used, what was approved, whether execution is still allowed and whether the operation has already been attempted. The SaaS enforces authorization for the actual principal.
Separate autonomy to propose work from authority to commit external effects. Track approved content, dispatched requests and evidenced outcomes as one operation. That is the operating foundation for an agent that manipulates SaaS.
References
Endpoint and SDK features are not treated as universal guarantees across SaaS providers.
- Microsoft Graph permissions overview
- Microsoft identity platform On-Behalf-Of flow
- RFC 8693: OAuth 2.0 Token Exchange
- RFC 8707: Resource Indicators
- RFC 9449: DPoP
- OpenAI Agents: Guardrails and approvals
- RFC 8785: JSON Canonicalization Scheme
- RFC 9110: HTTP conditional requests / If-Match
- Microsoft Graph sendMail
- Microsoft Graph event / transactionId
- Stripe idempotent requests
- Microsoft Graph throttling
- OpenAI safety in building agents