Together AI's $0.24 DeepSeek rollout is billed at DeepSeek's peak rate in 2026

Together AI charges DeepSeek's peak token rate around the clock, and its 18 August routing study priced GPT-5.6 Sol three days before OpenAI cut it.

Read time
12 min
Word count
1.9K
Sections
12
FAQs
8
Share
Price table: Together AI lists DeepSeek V4 Pro at $1.32 and $3.96 per million tokens, matching DeepSeek's peak rate
On this page · 12 sections
  1. What Together published, and when
  2. Together's price for DeepSeek V4 Pro is DeepSeek's peak rate
  3. The $0.24 figure does not reconcile with Together's own list price
  4. The Sol side of the comparison expired three days later
  5. The 272,000-token cliff the changelog does not mention
  6. Two DeepSeek V4 Pro entries, and one of them is not the current model
  7. What to do
  8. India-specific considerations
  9. What is still unknown
  10. FAQ
  11. How eCorpIT can help
  12. References

Summary. Together AI published a 904-rollout DeepSWE study on 18 August 2026 recommending a cascade: run DeepSeek V4 Pro 0813 first, escalate to GPT-5.6 Sol on test failure, 83.0% solved at $3.35 a task against Sol alone at 72.7% and $8.37. Two of the three prices in that recommendation have a problem. Together's own serverless table lists DeepSeek-V4-Pro-0813 at $1.32 per million input tokens and $3.96 per million output tokens, which are DeepSeek's peak-hour rates to the cent; DeepSeek's direct off-peak rates are $0.66 and $1.98, and off-peak covers 133 of the 168 hours in a week. And on 21 August 2026, three days after the study, OpenAI cut GPT-5.6 Sol to $4 per million input and $20 per million output, a 20% and 33% reduction from the $5 and $30 the $8.37 figure was computed against.

None of that makes the benchmark wrong. It makes the routing decision it recommends a decision you have to re-price yourself, on your own traffic pattern, before you wire it into a gateway.

What Together published, and when

The post, "DeepSeek V4 Pro 0813 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing", is dated 18 August 2026 and credited to Zain Hasan and Shobhit Dixit. It reports 904 rollouts: all 113 DeepSWE tasks, four trials each, 452 rollouts per model, both at maximum reasoning effort.

The measured result is a clean split. Sol takes single-shot quality at 72.7% pass@1 against 62.8%, holds pass@2 at 81.0% to 78.5%, and then loses pass@4, 85.8% to 88.5%. Sol is faster and tighter: a median 17 minutes and 53 steps per rollout against 35 minutes and 146 steps, and 59,000 output tokens against 101,000. Sol also breaks tests that already passed in 20% of its failures, against 11% for Pro. DeepSeek's own 13 August changelog puts V4 Pro at 62.7 on DeepSWE, within a tenth of a point of Together's 62.8, so the accuracy half of the study corroborates against the model vendor.

The cost half is the part to check.

Together's price for DeepSeek V4 Pro is DeepSeek's peak rate

Set Together's published serverless price beside DeepSeek's own published API price for the same model version.

Line item, per 1M tokens Together AI serverless DeepSeek direct, off-peak DeepSeek direct, peak
Input, cache miss $1.32 $0.66 $1.32
Input, cache hit $0.13 $0.022 $0.044
Output $3.96 $1.98 $3.96
Context window 1,048,576 1,000,000 1,000,000
Concurrency limit not published per model 500 500

Together's uncached input and output prices match DeepSeek's peak-hour rates exactly. DeepSeek states that off-peak rates are half the peak rates, and defines peak as 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday. That is 35 hours out of 168, or 20.8% of the week. For 79.2% of the week, the same model on DeepSeek's own API costs half what Together charges for it.

The cache line is wider still. Together's cached-input rate of $0.13 per million is 2.95 times DeepSeek's peak cache-hit rate of $0.044 and 5.91 times the off-peak rate of $0.022. That matters more than the headline rate for agent work, because DeepSeek enables disk-based context caching by default for every user, and a 146-step rollout re-sends a growing prefix on every step.

The $0.24 figure does not reconcile with Together's own list price

The study reports $109 total across 452 DeepSeek V4 Pro rollouts, an average of $0.24 each, and a median of 101,000 output tokens per rollout.

Take Together's published output price of $3.96 per million and apply it to that median: 101,000 tokens is $0.40 in output charges alone, 67% more than the $0.24 average total, before a single input token is billed. A median is not a mean, and output distributions in agent runs usually skew so that the mean sits above the median, not below it. Run the same check on the Sol side and it holds together: 59,000 output tokens at the then-current $30 per million is $1.77, leaving about $6.60 of input inside the reported $8.37, which is what roughly 1.3 million billed input tokens across 53 steps would look like.

So the Sol cost reconciles against public pricing and the DeepSeek cost does not. The likely explanation is benign, that the run was not billed at serverless list price. It is still a number you cannot reproduce from any published rate card, which is exactly the number a routing decision hangs on.

The Sol side of the comparison expired three days later

On 21 August 2026 the OpenAI API changelog recorded that GPT-5.6 Sol now costs $4 per million input tokens and $20 per million output tokens, "representing 20% lower input pricing and 33% lower output pricing", and that the promotional pricing is available at least through 21 November 2026. The implied previous rates are $5 and $30, which is what the study's $8.37 was computed against.

By our arithmetic, the same rollout now costs between $5.61 and $6.70, depending on the input and output split: $6.70 if the whole bill were input, $5.61 if it were all output. The cascade's $3.35 falls too, but by less, because only its escalation stage repriced and the DeepSeek first stage did not. The study's own margin, ten points of accuracy for less than half of Sol's price, is therefore narrower now than when it was published, and it narrows again if you move the first stage to DeepSeek's off-peak API.

The 272,000-token cliff the changelog does not mention

The changelog gives one price for Sol. The model card gives two. GPT-5.6 Sol has a 1,050,000-token context window and a 922,000-token maximum input, and the card states that prompts above 272,000 input tokens are priced at 2x input and 1.5x output "for the full request". The OpenAI pricing table carries that as separate long-context columns, $8 input and $30 output against $4 and $20.

GPT-5.6 Sol tier, per 1M tokens Input Cached input Output
Standard, up to 272K input $4.00 $0.40 $20.00
Standard, above 272K input $8.00 $0.80 $30.00
Batch, up to 272K input $2.00 $0.20 $10.00
Batch, above 272K input $4.00 $0.40 $15.00
Fast mode, up to 272K input $8.00 $0.80 $40.00

Two things are worth knowing before you size an agent on this. The multiplier applies to the entire request, not to the tokens above the threshold, so crossing 272,001 input tokens reprices every token in that call. And the threshold is printed in the pricing table's row labels for gpt-5.5, gpt-5.5-pro, gpt-5.4 and gpt-5.4-pro, all marked "(<272K context length)", but for no row in the GPT-5.6 family. The string "long context" appears zero times in the changelog. A reader who priced Sol from the 21 August entry alone would have no way to know a second tier exists.

Sol's median peak context in the Together run was 177,000 tokens, comfortably under the cliff. The 922,000-token input ceiling means 70.5% of the model's addressable input range sits above it.

Two DeepSeek V4 Pro entries, and one of them is not the current model

Together's serverless table lists DeepSeek V4 Pro twice.

Together model ID Context Input Cached input Output
deepseek-ai/DeepSeek-V4-Pro-0813 1,048,576 $1.32 $0.13 $3.96
deepseek-ai/DeepSeek-V4-Pro 512,000 $1.74 $0.20 $3.48
deepseek-ai/DeepSeek-V4-Flash-0731 1,000,000 $0.14 $0.03 $0.28

DeepSeek's own pricing page documents exactly one V4 Pro, and gives its model version as DeepSeek-V4-Pro-0813 with a 1M context length and a 384K maximum output. The unsuffixed deepseek-ai/DeepSeek-V4-Pro entry, at 512,000 context, does not correspond to a version DeepSeek currently documents. It is also the name a developer reaches for first, and picking it costs 31.8% more per input token for half the context window. The output price is the one line that goes the other way, $3.48 against $3.96.

What to do

Re-run the arithmetic on your own traffic before you adopt the cascade. Three checks, in order.

First, plot your request volume against DeepSeek's peak windows. If most of your agent traffic falls outside 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, going direct to DeepSeek halves the first stage's token cost against Together's flat rate, and Together's convenience has to earn that gap back in throughput or reliability.

Second, measure your cache-hit ratio before you choose a host. At a 5.91x spread on cache-hit input, the hosting decision for a long multi-step agent is decided by the cache line, not the headline rate.

Third, instrument the 272,000-token boundary as a hard alert, not a dashboard metric. A request that crosses it reprices in full at 2x input and 1.5x output, and nothing in your token counter will flag it.

The general lesson is one we keep re-learning on gateway work: a published benchmark is a measurement of a moment, and inference price lists now move faster than the studies that cite them. If you are choosing between hosts, compare the gateway layer itself rather than only the model, and treat any per-task cost as a variable with a date attached. The wider decision framework sits in our note on hybrid model routing and API spend, and the peak-window mechanics are worked through in detail in our piece on DeepSeek V4 peak and off-peak pricing. For the OpenAI side of this month's repricing, see our coverage of the GPT-5.6 Sol price cut across AI gateways and the broader frontier model comparison.

India-specific considerations

DeepSeek's peak windows are defined in UTC, and India is UTC+5:30, so the two windows land at 06:30 to 09:30 IST and 11:30 to 15:30 IST. The second one sits squarely inside the Indian working day. A Gurugram or Bengaluru team running agent workloads at their own desks pays peak rates for four hours of every working afternoon, while a US-hours team runs almost entirely off-peak on the same account.

That is a scheduling problem with a real number on it. At $0.66 against $1.32 per million uncached input tokens, an Indian team that can move batch and evaluation runs out of the 11:30 to 15:30 IST window halves the token cost of that slice with no code change. Interactive traffic cannot move, but regression suites, nightly evaluations and bulk re-indexing can. Under the Digital Personal Data Protection Act 2023, note separately that routing prompts through a third-party inference host is a processor relationship you have to document, whichever rate you pay.

What is still unknown

Together has not published how the study's DeepSeek costs were computed, so the gap between the $0.24 average and the $0.40 that list price implies for the median rollout stays unexplained. Together's serverless table does not publish a per-model concurrency limit against DeepSeek's stated 500. And OpenAI's promotional pricing for Sol is guaranteed only "at least through November 21, 2026", with no published rate after that date, so any cascade economics you fix today have a 21 November review point built in.

FAQ

How eCorpIT can help

Inference price lists now move faster than the benchmarks that cite them, and a routing rule written against last week's rate card quietly overspends. Our senior engineering teams build and operate the gateway layer that makes model routing measurable, including per-model cost attribution, cache-hit instrumentation and alerting on repricing thresholds like the 272,000-token boundary. That work is described on our AI gateway and model routing FinOps service page. To have your current model spend re-priced against today's published rates, book a model routing cost review.

References

  1. Together AI, DeepSeek V4 Pro 0813 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing, published 18 August 2026.
  1. Together AI documentation, Serverless models and pricing.
  1. DeepSeek API documentation, Models and pricing.
  1. DeepSeek API documentation, Change log, entries dated 13 and 21 August 2026.
  1. DeepSeek API documentation, Context caching.
  1. DeepSeek API documentation, Rate limit and isolation.
  1. OpenAI, API changelog, entry dated 21 August 2026.
  1. OpenAI, API pricing.
  1. OpenAI, GPT-5.6 Sol model card.
  1. OpenAI, Your data guide.
  1. Together AI, Rate limits.
  1. Together AI, Dedicated inference.

Last updated: 24 August 2026.

Frequently asked

Quick answers.

01 Is Together AI more expensive than DeepSeek's own API?
For DeepSeek V4 Pro 0813, yes, outside DeepSeek's peak hours. Together lists $1.32 per million input and $3.96 output. DeepSeek charges the same at peak but half that off-peak, at $0.66 and $1.98. Off-peak is 133 of the 168 hours in a week.
02 What are DeepSeek's peak hours?
DeepSeek defines peak as 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday through Friday, and states that off-peak rates are half the peak rates. That is 35 hours out of 168, or 20.8% of the week. In Indian Standard Time those windows fall at 06:30 to 09:30 and 11:30 to 15:30.
03 Did GPT-5.6 Sol get cheaper after the Together study?
Yes. On 21 August 2026, three days after the study was published, OpenAI's API changelog recorded Sol at $4 per million input tokens and $20 per million output, a 20% and 33% cut from $5 and $30. The promotional pricing runs at least through 21 November 2026.
04 What is the 272,000-token pricing cliff on GPT-5.6 Sol?
The GPT-5.6 Sol model card states that prompts above 272,000 input tokens are priced at 2x input and 1.5x output for the full request, moving the rate to $8 and $30. The threshold is not printed against any GPT-5.6 row in the pricing table.
05 Is the Together AI cascade recommendation still valid?
The accuracy result is unaffected: 83.0% for the cascade against 72.7% for Sol alone, measured over 904 rollouts. The cost gap has narrowed, because Sol repriced downward on 21 August 2026 while the DeepSeek first stage did not. Re-run the arithmetic on current rates.
06 Which DeepSeek V4 Pro model ID should I use on Together AI?
Together lists two. The 0813 entry gives 1,048,576 context at $1.32 input and $3.96 output. The unsuffixed entry gives 512,000 context at $1.74 and $3.48. DeepSeek documents only version 0813, so the unsuffixed ID costs 31.8% more per input token for half the context.
07 Does DeepSeek caching change the comparison?
Substantially, for long agent runs. DeepSeek enables disk context caching by default and charges $0.022 per million cache-hit input tokens off-peak. Together charges $0.13 for the same line, 5.91 times more. Because a 146-step rollout re-sends a growing prefix, that line often decides the bill.
08 How accurate is DeepSeek V4 Pro against GPT-5.6 Sol on DeepSWE?
Together measured 62.8% pass@1 for DeepSeek V4 Pro 0813 and 72.7% for GPT-5.6 Sol, over four trials on 113 tasks. At four attempts the order inverts, 88.5% against 85.8%. DeepSeek's own changelog reports 62.7 for V4 Pro, matching Together's figure.

About the author

Manu Shukla

Founder & Director

Founder of eCorpIT. Hands-on engineer leading senior-only delivery for AI apps, custom software, and cloud systems for global clients.

Subscribe

One engineering note a week. No fluff, no spam.

Senior-architect playbooks on AI agents, mobile apps, cloud, security, data, and marketing — delivered every Wednesday.

Past the reading

Read enough. Let's build something.

A senior architect responds in 24 working hours with scope, indicative cost, and a timeline. NDA before any technical conversation.