Choose the maximum exposure first
Begin with the maximum provider cost the owner is willing to approve for the month. AI Gateway HQ enforces that boundary before an eligible request can leave the gateway.
Protect the whole workspace
Create an organization-wide cost limit first. Workload, key, route, model, and principal limits can narrow exposure further, but they should not replace the workspace-wide stop.
Use strict mode for money
Select strict enforcement when the limit must stop spend. The gateway reserves a conservative request amount before dispatch, then reconciles supported provider usage and releases unused exposure afterward.
Warn before the stop
Set a warning percentage below the hard ceiling so an owner can review unusual growth. The warning adds visibility; it does not raise or bypass the approved limit.
Prepay managed provider exposure
Managed Bedrock requires settled purchased credit and a reusable payment authorization. Promotional balances cannot fund provider calls, and zero eligible credit stops the request before Amazon Bedrock is invoked.
Bound every automatic reload
Automatic reload begins disabled. If the owner enables it, the threshold, reload amount, and monthly cap must all be set, and a pending charge never becomes spendable credit.
See the cost basis
Review each request's provider spend, platform charge, token usage source, reservation outcome, and cost basis. Prompt and response bodies are not required to explain the financial decision.
Expand from observed evidence
Start with one bounded workload, review its usage and operational result, and raise the approved limit only when the owner understands why the additional exposure is justified.