One cost vocabulary
Connect request, token, model, provider account, configured rate, workload, and execution context in an explainable metadata ledger, then inspect the latest retained window in the customer overview.
Reserve provider exposure before dispatch and attribute recent gateway-observed usage across operational owners.
AI Gateway HQ attributes modeled and supported provider-reported usage to the workload, environment, client mode, data class, provider account, route, and model that caused it. Strict limits reserve budget before forwarding, while cost-ranked routes consider only models that already satisfy technical and policy requirements.
Organization and workload scopes separate centrally funded use from product, team, environment, or customer-specific consumption.
Name the organization, workload, administrator, environment, and provider account behind usage.
Combine monthly cost, request, RPM, TPM, concurrency, prepaid, and output-cap controls.
Break down the latest retained request window across seven operational dimensions, with exact usage and conservative settlement visibly separated.
Estimate the maximum gateway and provider exposure from the request and configured output cap.
Atomically reserve against the applicable organization, workload, and prepaid limits before forwarding.
Choose the least-cost eligible target only after capability, policy, and health requirements pass.
Reconcile the reservation against reported usage and preserve a reason-coded ledger entry.
Organization and workload-key limits are reserved before provider traffic instead of reported only after the bill arrives.
Organization and workload-key limits are reserved before provider traffic instead of reported only after the bill arrives.
Each control has an operating path, an owner, and evidence that can be reviewed without collecting prompt bodies by default.
Connect request, token, model, provider account, configured rate, workload, and execution context in an explainable metadata ledger, then inspect the latest retained window in the customer overview.
Free-tier and internal quotas can be enforced without AI Gateway HQ funding customer inference.
Provider invoice import and variance detection are not currently included; current evidence reports gateway-observed usage.
Current evidence provides a bounded, tenant-private operational allocation—not an invoice or forecast. Durable periods, provider invoice import, discount schedules, and automated variance detection are not currently included.
Review security boundaries and current statusClear answers for buyers, administrators, and security reviewers.
Current organization and workload-key scopes can represent a team, product, environment, or customer. Deeper hierarchy and cost-center synchronization are planned.
Administrators configure the rates used by request-cost routing. Automated provider invoice and contract-rate ingestion are not yet shipped.
Strict prepaid mode stops eligible paid requests before forwarding. If the account owner has enabled capped auto-reload, the configured Stripe charge can restore credit; otherwise traffic remains stopped. Managed Bedrock also requires positive purchased, managed-eligible cash; promotional credit cannot fund AWS inference.
Connect a provider credential, create a workload key, and begin in Observe mode. Move a tested rule to Enforce when your team is ready.