OpenAI API

OpenAI API pricing charges per million tokens, separately for input and output, with a different rate for every model.

Pricing Model:

Pricing Model:

Pure usage

Pure usage

Pure usage

Packaging Model:

Packaging Model:

Good / Better / Best (GBB)

Good / Better / Best (GBB)

Good / Better / Best (GBB)

Credit Model:

Credit Model:

Prepaid credits. Terms not stated on the pricing page

Prepaid credits. Terms not stated on the pricing page

Prepaid credits. Terms not stated on the pricing page

Updated on:

OpenAI API pricing: the rate card and the discounts

OpenAI API pricing charges per million tokens, separately for input and output, with a different rate for every model. There is no subscription, no seat and no included allowance. You pay for what the model reads and what it writes. The rate card is the entire product, which makes the discounts the only real lever: cached input runs 90% below standard, and the Batch API runs at half price.

Key takeaways

  • Rates span roughly 67x, from $0.15 per million input tokens on GPT-4o-mini to $10 on GPT-6 Astra.

  • Output costs about 5 to 8 times input on current models, so response length drives the bill more than prompt length.

  • GPT-5.6 Sol's headline rate is promotional and guaranteed only through 21 November 2026.

  • Cached input is discounted 90%, which makes prompt structure a genuine cost decision.

  • Crossing the long-context threshold doubles input rates and adds 50% to output, so prompt length is a second price lever.

  • The Batch API halves the rate for work that tolerates asynchronous delivery.

OpenAI API pricing in 2026

Model

Input /1M

Cached input

Cache writes

Output /1M

 

GPT-6 Astra

$10.00

$1.00

$12.50

$50.00

GPT-5.6-Sol

$4.00*

$0.40*

$5.00*

$20.00*

GPT-5.6-Terra

$2.00

$0.20

$2.50

$12.00

GPT-5.6-Luna

$0.20

$0.02

$0.25

$1.20

Source: developers.openai.com/api/docs/pricing, read 22 September 2026.

What OpenAI actually meters

OpenAI meters tokens, and it splits them three ways rather than two. Input, cached input and output each carry their own rate, and the gaps between them are large enough to change how you build.

Output is the expensive half. GPT-6 Astra reads at $10 per million and writes at $50, a 5x spread. GPT-5.6-Terra reads at $2 and writes at $12, a 6x spread. A verbose system prompt is cheap by comparison with a verbose answer, so instructing a model to be concise is a pricing decision as much as a product one.

Cached input is where the real discount sits. Most current models price cached input at 90% below standard input, so GPT-5.6-Sol drops from $4 to $0.40 per million. That only pays off if the stable part of your prompt sits at the front and stays byte-identical between calls, which makes prompt ordering a billing concern.

GPT-4o is the exception worth noting. Its cached input is $1.25 against $2.50 standard, a 50% discount rather than 90%, so the older model does not benefit nearly as much from caching.

What happens when you hit the limit

Nothing to hit. There is no included allowance on the API, so there is no overage concept. Every token bills at list rate from the first call.

Spend control therefore lives entirely in account-level configuration rather than in the pricing model: prepaid credit balances, usage limits and rate limits. The pricing page does not document credit expiry, auto top-up behaviour or what happens when a prepaid balance reaches zero, so those mechanics need verifying against account settings before publish.

How OpenAI API pricing has changed

OpenAI dates every pricing change in its API changelog, and the direction is consistently downward.

Date

Milestone

Source

 

10 Sep 2026

GPT-Live voice sessions priced at $0.05 per minute, billed per second

developers.openai.com/api/docs/changelog

21 Aug 2026

GPT-5.6 Sol drops to $4 input and $20 output, 20% and 33% lower, as promotional pricing through at least 21 November 2026

developers.openai.com/api/docs/changelog

20 Aug 2026

Prompt Caching dashboard released, exposing cache hit rate and cache-read versus cache-write token splits

developers.openai.com/api/docs/changelog

5 Aug 2026

Fast mode extended to long-context prompts above 272K tokens

developers.openai.com/api/docs/changelog

30 Jul 2026

GPT-5.6 Luna cut 80%, GPT-5.6 Terra cut 20%. Fast mode replaces Priority Processing, running 2.5x faster at twice the price

developers.openai.com/api/docs/changelog

2 Jun 2026

Container sessions move to per-minute billing with a 5-minute minimum, replacing a flat 20-minute session rate

developers.openai.com/api/docs/changelog

21 Apr 2026

GPT Image 2 ships with token-based image pricing and Batch API support at a 50% discount

developers.openai.com/api/docs/changelog

Source: developers.openai.com/api/docs/changelog, read 22 September 2026.

Flexprice’s Take

OpenAI publishes the cleanest rate card in AI and the least about how you actually get billed.

OpenAI's rate card is easy to model and easy to compare across providers. Every model gets a row, and cached input gets its own column at 90% below standard, which tells developers prompt architecture has a price and roughly what it is.

What the table won't tell you is when a rate expires. GPT-5.6 Sol's $4 input is promotional and guaranteed only through 21 November 2026, so a 2027 budget built on it is a guess.

The billing mechanics are missing entirely. The pricing page never says whether prepaid credits expire, how auto-recharge behaves, or what happens at zero balance. On a product with no allowance to fall back on, those matter as much as the rates.

Best For

Teams that can measure token flow and optimise against a published rate.

Watch Out For

Output-heavy workloads, where the 5x spread compounds fast.

Manish Choudhary

CEO & Co-founder, Flexprice

Reselling model access and need per-customer, per-model cost tracking?

Flexprice meters it and reports margin by account.

Flexprice’s Take

OpenAI publishes the cleanest rate card in AI and the least about how you actually get billed.

OpenAI's rate card is easy to model and easy to compare across providers. Every model gets a row, and cached input gets its own column at 90% below standard, which tells developers prompt architecture has a price and roughly what it is.

What the table won't tell you is when a rate expires. GPT-5.6 Sol's $4 input is promotional and guaranteed only through 21 November 2026, so a 2027 budget built on it is a guess.

The billing mechanics are missing entirely. The pricing page never says whether prepaid credits expire, how auto-recharge behaves, or what happens at zero balance. On a product with no allowance to fall back on, those matter as much as the rates.

Best For

Teams that can measure token flow and optimise against a published rate.

Watch Out For

Output-heavy workloads, where the 5x spread compounds fast.

Manish Choudhary

CEO & Co-founder, Flexprice

Reselling model access and need per-customer, per-model cost tracking?

Flexprice meters it and reports margin by account.

Customer
Sentiment Highlights

~$500-1500/month token spend at API OpenAI/Anthropic pricing seems pretty realistic for full-time engineers

Engineering lead describing company LLM budgets, Hacker News, August 2026

Frequently Asked Questions

Frequently Asked Questions

How much does the OpenAI API cost?

Is there a free tier for the OpenAI API?

How much does OpenAI prompt caching save?

What is the OpenAI Batch API discount?

Launch usage-based billing this week, not next quarter

Launch usage-based billing this week, not next quarter

Get Instant Feedback on Your Pricing | Join the Flexprice Community with 400+ Builders on Slack

Join the Flexprice Community on Slack