Set the ceiling before the request.
Your estimated total is the AI Gateway HQ fee plus model-provider usage. Try the Test Lab without a card or provider key. For live traffic, gateway credit is reserved before forwarding. Bring a provider account or prepay managed Bedrock; either path can stop before an approved ceiling is crossed.
- No-card Test Lab
- No BYOK inference markup
- Managed balance stops at zero
One control point between your AI tools and model providers.
Your application, coding agent, or automation sends model requests through AI Gateway HQ. The gateway applies your keys, routes, budgets, and security rules before an approved provider can charge.
Use it when
- Several people or tools use AI and you need one spending ceiling.
- You want to switch models or provider accounts without rewriting every client.
- You need routing, prompt-attack controls, workload keys, and decision evidence.
- You want provider costs billed at your own BYOK rates with no inference markup.
It does not
- Design or rebuild a WordPress website by itself.
- Replace the ChatGPT or Claude consumer chat application.
- Turn a consumer subscription login into a transferable API key.
- Promise that every cheaper model is suitable for every task.
Set the limit and its stop behavior in the same workflow.
The illustration above explains the boundary. This sanitized capture shows the API-backed budget console used to configure organization, request-rate, token, and concurrency limits.
Free
Explore real policy logic before connecting a provider.
- No-card, no-provider Test Lab
- $1 gateway-request credit · not model credit
- 1 administrator
- 2 provider credentials
- Observe and Enforce
- 7-day metadata
Flex
Prepaid, hard-capped gateway usage with no subscription.
- $0.10 / 1,000 successful requests
- Buy credits before use
- Optional threshold auto-reload
- Customer-set monthly reload ceiling
- No charge for failed upstream calls
Company
Self-service controls for several teams, tools, and environments.
- 2,000,000 successful requests per UTC calendar month
- OIDC, SAML, SCIM, and built-in roles
- Policy modes and route controls
- Insights with approval-bound changes
- No percentage fee on BYOK inference
Portfolio
Consented cost, control, and value reporting across participating companies.
- 10,000,000 successful requests per UTC calendar month
- One sponsor-owned portfolio workspace
- Company-owned gateway boundaries
- Up to 100 approved companies per bounded view
- Separate operations, finance, and outcome consent
- Board PDF, scheduled executive reports, and action queue
Gateway charge + provider inference + any contracted services.
Free, Flex, and Company customers choose one gateway plan or usage meter. They then pay model inference either directly to a BYOK provider or from prepaid managed Bedrock credit. Implementation, support, SLA, DPA, and private-deployment work apply only when stated in an order.
Published Company and Portfolio request allowances are UTC calendar-month hard ceilings. They do not roll over and do not create an automatic overage charge. At the ceiling, the gateway stops before another provider call until the next calendar month or an approved plan change; a tighter customer budget can stop traffic sooner.
The published $1,500 monthly base includes the sponsor-owned portfolio workspace, a 10,000,000-successful-request UTC calendar-month ceiling, approved company aggregation and drill-down, up to 100 approved companies per bounded synchronous view, the executive action queue, Board PDF, and scheduled audience-specific reports.
Each participating company keeps its own gateway and provider boundary. Its provider inference, managed Bedrock credit, applicable gateway plan or usage, implementation work, contracted support, SLA, DPA, and security obligations are separate unless the written order says otherwise.
Define companies, responsibilities, and pilot termsThe BYOK gateway fee does not rise with the model bill.
Flex charges for successful governed requests, not a percentage of the inference purchased from your provider. That keeps negotiated provider discounts and high-value workloads from increasing the gateway's percentage take.
- 1,000 successful requests
- $0.10 Gateway usage. Failed upstream calls release the reservation.
- 1,000,000 successful requests
- $100.00 The same meter whether the eligible model is inexpensive or premium.
- BYOK provider inference
- 0% markup Your provider bills its own rates directly to your account.
Request-heavy workloads on very inexpensive models can have different economics from long or premium-model requests. Use the estimator below with your expected request and token shape; it is an illustration, not a savings claim. BYOK cost enforcement reserves against the rate catalog you configure and can differ from a provider invoice; set a provider-side quota or spend limit when you require an external billing boundary.
See how requests, tokens, and model choice change the bill.
A short question may use one request. An agentic coding job may loop through dozens. This estimator separates provider inference from the gateway fee so you can see where the money goes.
Estimate a prompt, workflow, or agent run.
A “task” is not a billing unit. It may send one model request or hundreds. Start with a familiar example, then change every assumption.
Tokens are pieces of text, not words. As a rough English-language guide, 1,000 words is often about 1,300 tokens; code, tables, tools, and retained conversation history can change that substantially.
per month for 100 tasks
- Provider inference
- $1.40
- AI Gateway HQ
- $0.01
- Total requests
- 100
Illustration, not a quote. BYOK provider inference is billed directly by the provider; contract rates, caching, batch discounts, taxes, context tiers, retries, and model behavior can change the result. Welcome credit is not deducted here.
View OpenAI pricing sourceCatalog 2026-08-31.1 · reviewed Aug 31, 2026 · review due Sep 30, 2026Hosted enterprise from $36,000/year. Private-deployment design from $60,000/year.
Self-service OIDC, SAML, and SCIM controls are included in the Company product. Assisted identity architecture, evidence review, and contracted support can be scoped now. Native SIEM delivery, custom retention enforcement, custom roles, and customer-VPC deployment are not generally available and are not sold as completed features. The private-deployment price is for a separately defined design engagement, not a self-service deployment license.
BYOK gateway usage is $0.10 per 1,000 successful requests and has no inference markup; customers pay their provider directly. Optional managed Bedrock is self-service for funded workspaces. It requires a settled purchased-credit balance and a reusable card or Stripe Link payment method, and hard-stops at zero. The managed catalog contains reviewed conversational models routed in AWS US Regions, non-streaming, at the rates displayed in the application. Failed BYOK upstream calls release gateway credit; an ambiguous managed call may capture its conservative provider reserve. Published prices exclude taxes, provider inference, a formal SLA, and certifications unless stated in an executed agreement.
USD pricing, digital delivery, and a visible stop switch.
Authenticated workspaces use Stripe-hosted checkout. It identifies the product, total, renewal terms, cancellation path, and statement descriptor before confirmation. Credit or plan access begins only after the signed paid event is verified.