GPT-5.6 on Bedrock costs exactly 10% more than OpenAI direct: the August 2026 cost math

Bedrock lists every GPT-5.6 tier at exactly 1.10x OpenAI's direct rate. The reason is documented on both sides.

Read time
17 min
Word count
2.7K
Sections
11
FAQs
8
Share
Balance scale weighing two server units, comparing Bedrock and OpenAI direct inference cost
Bedrock lists every GPT-5.6 tier at 1.10x OpenAI direct rates as of 8 August 2026.
On this page · 11 sections
  1. The two price lists, side by side
  2. Why the 10% exists
  3. Five monthly bills
  4. What you cannot do on the Bedrock path
  5. The plumbing is genuinely different
  6. The retention setting that breaks background mode
  7. A decision rule you can apply in an afternoon
  8. India-specific considerations
  9. FAQ
  10. How eCorpIT can help
  11. References

Summary. On 30 July 2026 OpenAI cut GPT-5.6 Luna by 80% and Terra by 20%, and Amazon Bedrock applied the same reductions the same day. That looked like the end of the Bedrock-versus-direct pricing question. It is not. Read the two published tables side by side and Bedrock's in-region rate is exactly 1.10 times OpenAI's list rate on every GPT-5.6 tier: Sol at $5.50 per 1M input tokens against $5.00, Terra at $2.20 against $2.00, Luna at $0.22 against $0.20, and the identical multiple on outputs, on long-context rates and on cache reads. Amazon's newsroom says "Pricing matches OpenAI first-party rates." Amazon's pricing page adds the qualifier that resolves the arithmetic: "In-region inference is priced at parity with OpenAI data residency tier." OpenAI's pricing page supplies the number, stating that regional data-residency endpoints carry a 10% uplift for models released on or after 5 March 2026.

So both vendors document it, neither headlines it, and the gap is wider than 10% for any workload that could have used a discount tier. OpenAI sells Luna through Batch and Flex at $0.10 input and $0.60 output, half of standard. Bedrock publishes no Batch, Flex, Priority or Reserved price for GPT-5.6 at all, and OpenAI's own Bedrock guide records the position plainly: service tiers are "On-demand inference only." For a batchable classification pipeline the real difference is not 10%. It is 120%.

This is a procurement question rather than an engineering one, and it has a defensible answer in both directions. Below are the published numbers, five worked monthly bills, the capabilities that do not exist on the Bedrock path, and the retention setting that will quietly break your background jobs.

The two price lists, side by side

Both tables below carry the published rates as of 8 August 2026. OpenAI's figures come from its API pricing page; Bedrock's come from the Amazon Bedrock pricing page, whose regional heading reads "Regions: US East (N. Virginia) & US East (Ohio)". Both vendors call the cheaper band short context, which Bedrock labels "Short Context Window (272K)".

Model OpenAI input Bedrock input OpenAI output Bedrock output Multiple
GPT-5.6 Sol $5.00 $5.50 $30.00 $33.00 1.10x
GPT-5.6 Terra $2.00 $2.20 $12.00 $13.20 1.10x
GPT-5.6 Luna $0.20 $0.22 $1.20 $1.32 1.10x
GPT-5.5 $5.00 $5.50 $30.00 $33.00 1.10x
GPT-5.4 $2.50 $2.75 $15.00 $16.50 1.10x

All figures are USD per 1M tokens at short context. The multiple holds at long context too, and it holds on cached tokens, which is the detail that rules out a rounding coincidence.

Rate line, GPT-5.6 Luna OpenAI Bedrock Multiple
Long-context input $0.40 $0.44 1.10x
Long-context output $1.80 $1.98 1.10x
Cached input read $0.02 $0.022 1.10x
Cache write $0.25 $0.275 1.10x
Standard output $1.20 $1.32 1.10x

Bedrock's cache-write column is labelled "Price per 1M input tokens (30m cache write)", so the write rate buys a 30-minute window. Cache reads land at a tenth of the input rate on both platforms, matching Amazon's newsroom description of "prompt caching with a 90% cached-input discount."

Two regional asymmetries are worth knowing before you pick a Region. GPT-5.6 Sol does not appear in the US West (Oregon) table at all, which lists only Terra, Luna and GPT-5.4. And AWS GovCloud (US-West) lists a single OpenAI model, GPT-5.4, at $3.30 input and $19.80 output, a 20% premium over the same model in commercial US Regions.

Why the 10% exists

The mechanism is stated on OpenAI's pricing page, repeated under each of its flagship tables: "Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency." Bedrock's OpenAI section then benchmarks itself to exactly that tier: "Note: In-region inference is priced at parity with OpenAI data residency tier. Global cross-region inference pricing coming soon."

Put the two together and the claim "Bedrock matches OpenAI pricing" is true against the data-residency tier and false against the standard list price. Which one you should compare against depends on a question about your own workload: would you have used OpenAI's data-residency endpoint anyway? If your compliance posture already requires regional processing, Bedrock is at parity and the argument is over. If you were happily calling the standard endpoint at $5.00, you are being quoted the residency rate whether or not you need residency.

OpenAI's own Bedrock guide says as much without the number: "AWS bills Amazon Bedrock usage. Bedrock-specific pricing can differ from direct OpenAI API pricing, including regional processing premiums or other AWS-specific commercial terms." Its closing advice is one line long and worth following: "Compare Bedrock pricing and direct API pricing before launch."

The second half of that comparison is the part teams skip. AWS's own pricing page states the tier economics for other providers in a footnote: "Priority tier pricing is at 75% premium to Standard tier pricing. Flex tier and Batch pricing is at 50% discount to Standard tier pricing." That footnote does not appear under the OpenAI frontier-model tables, and no Priority, Flex, Reserved or Batch rate is published for GPT-5.6 anywhere on the page. The GPT-5.6 Luna model card confirms it from the other direction: Standard is supported, Priority, Flex and Reserved are not.

Five monthly bills

The workloads below use the published rates above, short context unless stated. Every Bedrock figure is in-region standard, because that is the only tier available for GPT-5.6.

Monthly workload OpenAI standard Bedrock in-region OpenAI Batch or Flex Cheapest Bedrock against cheapest OpenAI
Luna: 100M in, 20M out $44.00 $48.40 $22.00 2.20x
Terra: 20M in, 5M out $100.00 $110.00 $50.00 2.20x
Sol: 5M in, 1M out $55.00 $60.50 $27.50 2.20x
Luna at 90% cache read: 100M in, 20M out $27.80 $30.58 tier not applicable 1.10x
Luna long context: 50M in, 10M out $38.00 $41.80 $19.00 2.20x

The pattern is blunt. Against OpenAI's standard tier the Bedrock premium is a flat 10% and rounds to nothing on most bills. Against OpenAI's Batch or Flex tier it is 2.2x, because Bedrock removes the discount tier rather than pricing it higher. If your traffic is interactive and latency-bound, the discount tiers were never available to you and the 10% is the whole story. If your traffic is a nightly enrichment job, a backfill, a bulk classification pass or an evaluation sweep, Bedrock more than doubles it.

That is also why the cache-heavy row behaves differently. Prompt caching is available on both platforms at the same 90% discount, so heavy cache reuse compresses the absolute gap without changing the ratio. Caching and batching are the two levers that move an inference bill by more than a version bump does, and the budget LLM tier cost comparison covers how the cheap tiers compare once both are applied.

One accounting point pushes the other way, and it is real money rather than a rounding argument. Amazon states that OpenAI usage on Bedrock "counts toward existing AWS commitments." If your organisation holds a large AWS spend commitment it is not on track to consume, a 10% list premium paid in committed dollars can beat a standard-price invoice paid in new cash. That is a finance question, not an engineering one, and it belongs to whoever owns the commitment. The mechanics of that trade sit alongside the rest of the cloud FinOps playbook for Indian teams.

What you cannot do on the Bedrock path

Price is the easy half. OpenAI publishes a capability table comparing its API against Bedrock, current "as of July 13, 2026", and the gaps decide more migrations than the rates do. Listed as available on the OpenAI API and not available on Amazon Bedrock:

Capability OpenAI API Amazon Bedrock
Audio input Available Not available
WebSocket connections Available Not available
Pro mode Available on supported models Not available
Programmatic Tool Calling Available on supported models Not available
Multi-agent Beta on supported models Not available
Hosted file search Available Not available
Computer use Available Not available
Shell tool Available Not available
Remote MCP servers Available Not available

Two of those are load-bearing for anyone building agents in 2026. Remote MCP servers are unavailable, so an agent that reaches tools through hosted MCP has to be rewritten to client-side tool calling before it can run on Bedrock. The image generation tool is also listed as unavailable, alongside the shell tool and computer use. Hosted web search does work. OpenAI states the split explicitly: "Hosted web search is available on Amazon Bedrock, but hosted file search and remote MCP servers are unavailable."

What does carry over is substantial: text generation, image input, file input for supported types, structured outputs, function calling, streaming, reasoning effort "including max on supported models", persisted reasoning, custom tools, client-side tool_search, and both implicit and explicit prompt caching. OpenAI's own summary sentence is the right test to apply: "Use the OpenAI API directly when you need the broadest feature coverage, the latest first-party platform capabilities, or functionality unavailable in Bedrock."

The plumbing is genuinely different

Bedrock does not serve GPT-5.6 through the familiar bedrock-runtime endpoint, and it does not expose it through Converse or InvokeModel. The GPT-5.6 Luna model card marks bedrock-runtime, Chat Completions, Invoke and Converse all as unsupported, and marks bedrock-mantle and the Responses API as the only supported pair. AWS adds a note that the path differs from every other model on the responses endpoint.

The base URL is regional and contains the word mantle. OpenAI's guide gives the working client for it:


            from openai import BedrockOpenAI

client = BedrockOpenAI(aws_region="us-east-2")

response = client.responses.create(
    model="openai.gpt-5.6-sol",
    input="Write a haiku about cloud infrastructure.",
)

print(response.output_text)
          

Or with a plain HTTP call, using a Bedrock API key from the environment:


            curl "https://bedrock-mantle.us-east-2.api.aws/openai/v1/responses" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $AWS_BEARER_TOKEN_BEDROCK" \
  -d '{
    "model": "openai.gpt-5.6-sol",
    "input": "Write a haiku about cloud infrastructure."
  }'
          

Three details there are easy to get wrong. Model IDs take an openai. prefix, such as openai.gpt-5.6-sol and openai.gpt-5.6-luna. The path is openai/v1/responses, which AWS notes "is different from the v1/responses path used by other models on the responses endpoint." And for long-running services OpenAI recommends a token provider over a static key, because the AWS token generators return a cached short-term key and mint a new one when it expires, using the normal AWS credential chain.

Context limits differ from the direct API too. OpenAI's guide states 272,000 tokens for GPT-5.4 and GPT-5.5 on Bedrock, and 1,050,000 tokens for GPT-5.6 Sol, Terra and Luna, and warns that "Amazon Bedrock rejects requests that exceed the applicable model limit." The pricing band and the hard limit are not the same line: 272K is where the price doubles on GPT-5.6, not where the request fails.

The retention setting that breaks background mode

This is the trap most likely to reach production, and it lives in one paragraph that nobody reads twice.

Bedrock separates two controls. Zero operator access means, in AWS's words, that "AWS operators have no technical mechanism to sign in to Mantle's underlying compute systems or access customer data, including inference prompts and completions." Zero data retention means "AWS does not write model inputs or outputs to durable storage when the effective retention mode is none."

Under the default retention mode, Responses API requests use store: true by default and AWS retains the response, including its input and output, for 30 days. AWS also retains classifier-flagged traffic for up to 30 days for automated offline abuse detection on specific OpenAI GPT models. Turning that off is not a request-level flag: OpenAI's guide states flatly that "Setting store: false does not guarantee ZDR." Full zero data retention requires contacting your AWS account manager, per-account and per-model approval, confirming none appears in the model's allowed_modes, then setting the account or project retention mode to none.

Here is the consequence nobody plans for: "When the effective retention mode is none, AWS rejects store: true, and background mode is unavailable." The compliance win costs you a feature. If your architecture submits long jobs in background mode and polls for them, achieving zero data retention on Bedrock breaks that pattern, and you find out after a compliance sign-off has already been given. Decide the retention mode before you design the job runner.

For teams whose reason to be on Bedrock is governance rather than price, the offsetting controls are the actual product: IAM-based access management, AWS PrivateLink connectivity, guardrails, encryption at rest and in transit, and logging through AWS CloudTrail. Ben Kus, CTO at Box, framed the appeal of the AWS path as letting developers "build optimized, production-scale AI applications that bring together the strengths and capabilities of OpenAI's latest models with the scale, security, and infrastructure of AWS." If those controls are what your auditor asks about, 10% is a cheap answer.

A decision rule you can apply in an afternoon

Work through it in this order, because the first two questions settle most cases without any arithmetic.

  1. Does the workload need a capability on the unavailable list, especially remote MCP servers, hosted file search, computer use, the shell tool or audio input? If yes, go direct. No price makes an absent feature present.
  1. Is the workload batchable? If a nightly or asynchronous pass is acceptable, OpenAI's Batch and Flex tiers are half of standard and have no Bedrock equivalent for GPT-5.6, so the gap is 2.2x rather than 10%.
  1. Do you have an unconsumed AWS spend commitment? If yes, price the 10% premium against committed dollars rather than new cash, and get finance to sign that comparison.
  1. Do you actually need regional processing? If your compliance posture already requires OpenAI's data-residency endpoint, Bedrock is at parity and the premium disappears.
  1. Is IAM, PrivateLink, CloudTrail and guardrail integration worth 10%? For a regulated workload, usually yes. For a side service, usually no.
  1. Can you get zero data retention approved before launch, and can your job design survive losing background mode? Test that early.
  1. Whatever you choose, put the base URL, model ID and tier behind one resolvable configuration layer rather than in each service. That is the argument for an LLM hybrid routing decision framework instead of two hard-coded clients.

The honest summary is that the 80% Luna cut did not change this decision at all. It moved both sides down together and left the ratio untouched. What moves the decision is whether you can use a discount tier, and whether you need a feature Bedrock does not carry.

India-specific considerations

There is no Indian Region for GPT-5.6 on Bedrock. Published availability is US East (N. Virginia), US East (Ohio) and US West (Oregon), all In-Region only, with Geo and Global cross-region inference both marked unsupported on the Luna model card, and Bedrock's own note confirming that "Global cross-region inference pricing coming soon." An Indian team running GPT-5.6 through Bedrock is therefore making a trans-Pacific call on every request, paying a 10% premium benchmarked to a data-residency tier, and getting no Indian data residency for it.

That combination deserves a hard look under the Digital Personal Data Protection Act 2023. A data fiduciary stays accountable for personal data processed on its behalf, so prompts containing customer data crossing to a US Region is an architectural decision with a compliance owner, not a latency footnote. OpenAI's guide is careful about this and worth quoting to whoever signs off: "AWS Regions are physical deployment locations, which differ from OpenAI data residency jurisdictions. Teams with residency requirements should evaluate the Bedrock Region itself and the corresponding AWS terms."

One useful India data point does exist on the same pricing page. The gpt-oss-safeguard open-weight models are priced for Asia Pacific (Mumbai) at $0.08 input and $0.24 output per 1M tokens for the 20B, and $0.18 and $0.71 for the 120B, against $0.07 and $0.20, and $0.15 and $0.60, in the US Regions. Mumbai capacity for OpenAI open-weight models is therefore roughly 14% to 20% dearer than US East, but it exists, which the frontier models do not. For a moderation or guardrail layer that must stay in India while the reasoning model sits elsewhere, that split is a workable pattern. Building it is the same exercise as any other data residency and DPDP cloud architecture decision: classify the data, then place each hop.

The wider cost picture for Indian teams across all three hyperscalers sits in FinOps for AI cloud cost across AWS, Azure and GCP, and the price cut itself is covered in the OpenAI GPT-5.6 price cut and Fast mode migration.

FAQ

How eCorpIT can help

eCorpIT sizes inference channels against the workload rather than the headline rate: which calls can move to a discount tier, where prompt caching actually pays, what an AWS spend commitment changes about the comparison, and which capabilities on the unavailable list your architecture quietly depends on. As a CMMI Level 5 and ISO 27001:2022 certified organisation, we design these deployments aligned with DPDP Act 2023 requirements, including the Region placement and retention-mode decisions that have to be made before launch rather than during an audit. If you are choosing between Bedrock and a direct API and want the arithmetic run against your own token volumes, talk to our engineering team.

References

  1. Amazon Bedrock announces up to 80% lower prices for OpenAI GPT-5.6 models, AWS What's New, 30 July 2026
  1. Amazon Bedrock pricing, Amazon Web Services
  1. OpenAI API pricing, OpenAI Developers
  1. OpenAI models in Amazon Bedrock, OpenAI Developers, feature table as of 13 July 2026
  1. GPT-5.6 Luna model card, Amazon Bedrock User Guide
  1. OpenAI models available in Amazon Bedrock, Amazon Bedrock User Guide
  1. OpenAI GPT-5.6 models now generally available on Amazon Bedrock, Amazon News, updated 9 July 2026
  1. Advancing the price-performance frontier with GPT-5.6, OpenAI
  1. AWS Weekly Roundup, 3 August 2026, AWS News Blog
  1. OpenAI GPT-5.6 Sol, Terra, and Luna are now generally available on Amazon Bedrock, AWS Machine Learning Blog
  1. Amazon Bedrock service tiers, Amazon Bedrock User Guide
  1. Model support by AWS Region, Amazon Bedrock User Guide
  1. Amazon Bedrock API keys, Amazon Bedrock User Guide

Last updated: 8 August 2026.

Frequently asked

Quick answers.

01 Is Bedrock really more expensive than OpenAI direct for GPT-5.6?
Yes, at list price. Bedrock's published in-region rates are exactly 1.10 times OpenAI's standard rates on every GPT-5.6 tier: Sol at $5.50 against $5.00 input, Terra at $2.20 against $2.00, Luna at $0.22 against $0.20, with the same multiple on outputs, on long-context rates and on cached tokens.
02 Why is the premium exactly 10%?
Because Bedrock benchmarks itself to a different OpenAI tier. Bedrock's pricing page states that in-region inference is priced at parity with the OpenAI data residency tier, and OpenAI's pricing page states that regional data-residency endpoints carry a 10% uplift for models released on or after 5 March 2026. Those two statements reconcile the arithmetic exactly.
03 Amazon says pricing matches OpenAI first-party rates. Is that wrong?
It is true against one tier and not the other. Amazon's newsroom says pricing matches OpenAI first-party rates; its pricing page adds that the match is to the data residency tier. OpenAI's Bedrock guide states that Bedrock pricing "can differ from direct OpenAI API pricing, including regional processing premiums."
04 Can I get the Batch discount on Bedrock?
Not for GPT-5.6. Bedrock publishes no Batch, Flex, Priority or Reserved rate for GPT-5.6, the Luna model card lists Standard as the only supported tier, and OpenAI's comparison table records Bedrock service tiers as on-demand inference only. Direct OpenAI Batch and Flex are half of standard, so batchable work costs 2.2 times more on Bedrock.
05 Which regions serve GPT-5.6 on Bedrock?
US East (N. Virginia), US East (Ohio) and US West (Oregon), all In-Region only, with Geo and Global cross-region inference unsupported. GPT-5.6 Sol is absent from the Oregon pricing table. AWS GovCloud (US-West) lists only GPT-5.4, at $3.30 input and $19.80 output per 1M tokens.
06 What features do I lose by going through Bedrock?
OpenAI's comparison table, current as of 13 July 2026, lists audio input, WebSocket connections, Pro mode, Programmatic Tool Calling, multi-agent, hosted file search, computer use, the shell tool, the image generation tool and remote MCP servers as unavailable on Bedrock. Hosted web search, function calling, structured outputs and prompt caching do work.
07 How do I call GPT-5.6 on Bedrock?
Through the Responses API on the bedrock-mantle endpoint, not bedrock-runtime. The regional base URL looks like https://bedrock-mantle.us-east-2.api.aws/openai/v1, the path is openai/v1/responses rather than v1/responses, and model IDs take an openai. prefix such as openai.gpt-5.6-sol. Converse, InvokeModel and Chat Completions are all unsupported, so an existing Bedrock integration needs rewriting.
08 Does zero data retention have a downside on Bedrock?
Yes. AWS documents that when the effective retention mode is none, it rejects store: true and background mode is unavailable. Setting store: false alone does not deliver zero data retention; that needs per-account and per-model approval from AWS. Decide the retention mode before designing any long-running job runner.

About the author

Manu Shukla

Founder & Director

Founder of eCorpIT. Hands-on engineer leading senior-only delivery for AI apps, custom software, and cloud systems for global clients.

Subscribe

One engineering note a week. No fluff, no spam.

Senior-architect playbooks on AI agents, mobile apps, cloud, security, data, and marketing — delivered every Wednesday.

Past the reading

Read enough. Let's build something.

A senior architect responds in 24 working hours with scope, indicative cost, and a timeline. NDA before any technical conversation.