Skip to content
TokenCostLLM cost calculator

LLM API pricing table

Every model in the calculator, with SKU, context window, and standard USD per million input/output tokens. Snapshot date . This page exists so search engines and humans can read the rate card without running the form.

Public list prices per million tokens
ModelSKUProviderInput $/MOutput $/MContext
GPT-6 AstraStandard short-context. Input >272K reprices the entire request to $20 / $75 (cached input $2, cache writes $25). Source: developers.openai.com/api/docs/pricing.gpt-6-astraOpenAI$10.00$50.00See vendor
GPT-5.6 SolOfficial Standard short-context $4 / $20 (cached $0.40, cache writes $5). Promotional at least through 2026-11-21. Input >272K → $8 / $30 for the full request.gpt-5.6-solOpenAI$4.00$20.00See vendor
GPT-5.6 TerraGPT-5.6 mid-tier. Input >272K → $4 / $18 for the full request.gpt-5.6-terraOpenAI$2.00$12.00See vendor
GPT-5.6 LunaGPT-5.6 high-volume SKU. Input >272K → $0.40 / $1.80 for the full request.gpt-5.6-lunaOpenAI$0.20$1.20See vendor
GPT-5.5Previous flagship; still on the public card (<272K). Long context $10 / $45.gpt-5.5OpenAI$5.00$30.00See vendor
Claude Fable 5.1Current Fable SKU. Cache hits $0.25 / MTok (0.025× base input).claude-fable-5-1Anthropic$10.00$50.00See vendor
Claude Fable 5Prior Fable SKU; cache hits $1 / MTok (0.1×). Same $10 / $50 base.claude-fable-5Anthropic$10.00$50.00See vendor
Claude Opus 5Current public Opus SKU. Official first-party list, uncached text.claude-opus-5Anthropic$5.00$25.00See vendor
Claude Sonnet 5Standard rate. The scheduled 2026-09-01 raise to $3/$15 will not occur.claude-sonnet-5Anthropic$2.00$10.00See vendor
Claude Haiku 4.5Current public Haiku SKU; high-volume / low-latency work.claude-haiku-4-5Anthropic$1.00$5.00See vendor
Gemini 3.8 FlashIntroductory Developer API pricing through 2026-12-31. Planned step-up after that date: $1.50 / $7.50.gemini-3.8-flashGoogle$0.75$3.75See vendor
Gemini 3.7 FlashGoogle Cloud Agent Platform global intro through 2026-12-31: $0.75 / $3.75, then $1.50 / $7.50 (same intro tier as 3.8 Flash). Non-global is 10% higher. Developer API pricing page fetched 2026-09-08 did not list 3.7. Source: cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricing.gemini-3.7-flashGoogle$0.75$3.751M
Gemini 3.1 ProStandard text tier for prompts ≤ 200k tokens. Longer prompts are billed higher; this row will understate that case.gemini-3.1-proGoogle$2.00$12.00See vendor
Gemini 2.5 ProStandard paid tier for prompts ≤ 200k tokens. Prompts > 200k are $2.50 / $15. Output includes thinking tokens.gemini-2.5-proGoogle$1.25$10.00See vendor
Gemini 2.5 FlashText/image/video input. Audio input is billed higher ($1.00 / 1M).gemini-2.5-flashGoogle$0.30$2.50See vendor
Grok 4.6Prompts <200k tokens. If the prompt is ≥200k, the whole request bills at $4 / $12. Cached input $0.50 (<200k). Source: docs.x.ai/developers/pricing.grok-4.6xAI$2.00$6.00500k
Mistral Large 3Official USD Standard $0.50 / $1.50 (cached input $0.05). API id mistral-large-2512; alias mistral-large-3. Source: docs.mistral.ai/inference/pricing and docs.mistral.ai/models/mistral-large-3-25-12. FAQ on mistral.ai/pricing cites Large $0.5/$1.5.mistral-large-2512Mistral$0.50$1.50256k
Mistral Medium 3.5Official USD Standard $1.50 / $7.50 (cached input $0.15). Higher than Large 3 on the first-party card. Source: docs.mistral.ai/inference/pricing and docs.mistral.ai/models/mistral-medium-3-5-26-04.mistral-medium-3-5Mistral$1.50$7.50256k
Amazon Nova MicroBedrock OnDemand us-east-1 text tokens. AWS Price List AmazonBedrock 2026-09-01: $0.000035 / $0.00014 per 1K = $0.035 / $0.14 per 1M. Host: Amazon Bedrock (in-region).amazon.nova-micro-v1:0Amazon$0.035$0.14128k
Amazon Nova LiteBedrock OnDemand us-east-1. AWS Price List 2026-09-01: $0.00006 / $0.00024 per 1K = $0.06 / $0.24 per 1M. Host: Amazon Bedrock (in-region).amazon.nova-lite-v1:0Amazon$0.06$0.24300k
Amazon Nova ProBedrock OnDemand us-east-1. AWS Price List 2026-09-01: $0.0008 / $0.0032 per 1K = $0.80 / $3.20 per 1M. Host: Amazon Bedrock (in-region).amazon.nova-pro-v1:0Amazon$0.80$3.20300k
Amazon Nova PremierBedrock OnDemand us-east-1. AWS Price List 2026-09-01: $0.0025 / $0.0125 per 1K = $2.50 / $12.50 per 1M. Host: Amazon Bedrock (in-region).amazon.nova-premier-v1:0Amazon$2.50$12.501M
Amazon Nova 2 LiteBedrock OnDemand global cross-region (USE1-Nova2.0Lite-*-cross-region-global): $0.30 / $2.50 per 1M. In-region us-east-1 is $0.33 / $2.75 — not used here. One host/tier only. AWS Price List 2026-09-01. ID amazon.nova-2-lite-v1:0; global profile global.amazon.nova-2-lite-v1:0.amazon.nova-2-lite-v1:0Amazon$0.30$2.501M
Llama 4 ScoutAmazon Bedrock OnDemand us-east-1 (not Meta Model API). AWS Price List 2026-09-01: $0.00017 / $0.00066 per 1K = $0.17 / $0.66 per 1M. Meta first-party Model API fetched this date only listed Muse Spark ($1.25/$4.25), not Llama 4. One host; not averaged.meta.llama4-scout-17b-instruct-v1:0Meta (Bedrock)$0.17$0.663.5M on Bedrock
Llama 4 MaverickAmazon Bedrock OnDemand us-east-1. AWS Price List 2026-09-01: $0.00024 / $0.00097 per 1K = $0.24 / $0.97 per 1M. Same host as Scout. Meta first-party Llama 4 rates were not on the Model API card.meta.llama4-maverick-17b-instruct-v1:0Meta (Bedrock)$0.24$0.971M
Kimi K3Cache-miss input. Cache hit $0.30 / 1M (not applied). Official: platform.kimi.ai/docs/pricing/chat-k3.kimi-k3Moonshot$3.00$15.001M
MiniMax M3Standard ≤512k input, permanent 50% off ($0.60/$2.40 list → $0.30/$1.20 billed). >512k is $0.60 / $2.40. Cache read $0.06. Source: platform.minimax.io/docs/guides/pricing-paygo.MiniMax-M3MiniMax$0.30$1.20See vendor
GLM-5.3Z.ai international USD card $1.4 / $4.4 (docs.z.ai/guides/overview/pricing). BigModel.cn China card is still ¥8 / ¥28 — not used here. Cached input $0.26 (not applied).glm-5.3Z.ai$1.40$4.401M
GLM-5.3-FlashList $0.15 / $0.50 on docs.z.ai/guides/overview/pricing. A 50% launch promo ($0.075 / $0.25) ends 2026-09-09 24:00 UTC+8 — calculator uses list for stability. Cached input list $0.03 (not applied).glm-5.3-flashZ.ai$0.15$0.501M
DeepSeek V4 FlashPeak cache-miss / output (01:00–04:00 and 06:00–10:00 UTC, Mon–Fri). Off-peak is half ($0.22 / $0.66). Source: api-docs.deepseek.com/quick_start/pricing.deepseek-v4-flashDeepSeek$0.44$1.321M
DeepSeek V4 ProPeak cache-miss / output. Off-peak is half ($0.66 / $1.98). Same peak-hour window as Flash.deepseek-v4-proDeepSeek$1.32$3.961M
Qwen3.8-MaxAlibaba Model Studio China (Beijing) / Global: CNY 12 / 36 per 1M (help.aliyun.com/en/model-studio/qwen3-8-max). USD ≈ $1.67 / $5.00 at CNY 7.2 per USD. Singapore lists CNY 14.988 / 44.965 (not used here).qwen3.8-maxAlibaba$1.67¥12$5.00¥361M
Tencent Hunyuan Hy3Tencent Cloud TokenHub Intl. Input $0.132 / output $0.528 / cache hit $0.033 per 1M (campaign card tencentcloud.com/act/pro/tokenhub). API parameter hy3 (tencentcloud.com/document/product/1300/80632). Host: TokenHub, not mainland Hunyuan Open.hy3Tencent$0.132$0.528256k
Tencent Hunyuan Hy4TokenHub Intl Preview SKU hy4-preview: $0.834 / $2.501 (cache hit $0.042). Same TokenHub campaign card as Hy3. Preview — rates and availability can move.hy4-previewTencent$0.834$2.5011M
Xiaomi MiMo-V2.5Overseas pay-as-you-go cache-miss $0.14 / $0.28 (cache hit $0.0028). Source: mimo.mi.com/docs/en-US/price/pay-as-you-go, updated 2026-08-06. China mainland is billed in CNY (¥1 / ¥2) — not used here. Pro $0.435/$0.87 omitted from this packed table.mimo-v2.5Xiaomi$0.14$0.28See vendor

Prices last updated September 8, 2026 (). Rates change without notice. Figures are estimates from public list prices for standard, uncached, real-time text tokens — not an invoice. Cached tokens, batch discounts, thinking tokens you did not enter, and image/audio/video usage are not modeled in v1. GPT-6 Astra (and GPT-5.6) prompts with more than 272K input tokens reprice the entire request at the long-context column — Astra becomes $20 / $75. Grok 4.6 prompts at or above 200k bill the whole request at $4 / $12. Confirm on the provider’s official pricing page before you budget.

How to read this table

These are standard, uncached, real-time text rates. They are not OpenRouter blended prices. GPT-6 Astra and GPT-5.6 use the short-context column; input above 272K reprices the whole request. Grok 4.6 uses the <200k band. DeepSeek rows are official peak cache-miss rates; off-peak is half. Qwen3.8-Max is converted from official CNY at 7.2 CNY per USD (original yuan is shown under the dollar figure). GLM-5.3 uses the Z.ai international USD card ($1.4 / $4.4), not a yuan conversion. GLM-5.3-Flash uses list $0.15 / $0.50 (a 50% promo ends 9 September 2026 UTC+8). Gemini Pro rows use the ≤200k prompt tier. Gemini 3.8 and 3.7 Flash are introductory through 31 December 2026. Claude Sonnet 5 uses the $2 / $10 standard rate. Mistral Medium 3.5 is listed above Large 3 on the first-party card. Llama 4 and Nova are Amazon Bedrock us-east-1 OnDemand (Nova 2 Lite uses the global cross-region SKU). Hunyuan Hy3/Hy4 are Tencent Cloud TokenHub Intl. Use the calculator filters — All / OpenAI / Anthropic / Google / Mistral / Amazon / xAI / China — rather than scanning this dump unfiltered. When a provider publishes a new card, change src/data/pricing.ts and bump PRICES_LAST_UPDATED.