On this page · 12 sections
- What Together published, and when
- Together's price for DeepSeek V4 Pro is DeepSeek's peak rate
- The $0.24 figure does not reconcile with Together's own list price
- The Sol side of the comparison expired three days later
- The 272,000-token cliff the changelog does not mention
- Two DeepSeek V4 Pro entries, and one of them is not the current model
- What to do
- India-specific considerations
- What is still unknown
- FAQ
- How eCorpIT can help
- References
Summary. Together AI published a 904-rollout DeepSWE study on 18 August 2026 recommending a cascade: run DeepSeek V4 Pro 0813 first, escalate to GPT-5.6 Sol on test failure, 83.0% solved at $3.35 a task against Sol alone at 72.7% and $8.37. Two of the three prices in that recommendation have a problem. Together's own serverless table lists DeepSeek-V4-Pro-0813 at $1.32 per million input tokens and $3.96 per million output tokens, which are DeepSeek's peak-hour rates to the cent; DeepSeek's direct off-peak rates are $0.66 and $1.98, and off-peak covers 133 of the 168 hours in a week. And on 21 August 2026, three days after the study, OpenAI cut GPT-5.6 Sol to $4 per million input and $20 per million output, a 20% and 33% reduction from the $5 and $30 the $8.37 figure was computed against.
None of that makes the benchmark wrong. It makes the routing decision it recommends a decision you have to re-price yourself, on your own traffic pattern, before you wire it into a gateway.
What Together published, and when
The post, "DeepSeek V4 Pro 0813 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing", is dated 18 August 2026 and credited to Zain Hasan and Shobhit Dixit. It reports 904 rollouts: all 113 DeepSWE tasks, four trials each, 452 rollouts per model, both at maximum reasoning effort.
The measured result is a clean split. Sol takes single-shot quality at 72.7% pass@1 against 62.8%, holds pass@2 at 81.0% to 78.5%, and then loses pass@4, 85.8% to 88.5%. Sol is faster and tighter: a median 17 minutes and 53 steps per rollout against 35 minutes and 146 steps, and 59,000 output tokens against 101,000. Sol also breaks tests that already passed in 20% of its failures, against 11% for Pro. DeepSeek's own 13 August changelog puts V4 Pro at 62.7 on DeepSWE, within a tenth of a point of Together's 62.8, so the accuracy half of the study corroborates against the model vendor.
The cost half is the part to check.
Together's price for DeepSeek V4 Pro is DeepSeek's peak rate
Set Together's published serverless price beside DeepSeek's own published API price for the same model version.
| Line item, per 1M tokens | Together AI serverless | DeepSeek direct, off-peak | DeepSeek direct, peak |
|---|---|---|---|
| Input, cache miss | $1.32 | $0.66 | $1.32 |
| Input, cache hit | $0.13 | $0.022 | $0.044 |
| Output | $3.96 | $1.98 | $3.96 |
| Context window | 1,048,576 | 1,000,000 | 1,000,000 |
| Concurrency limit | not published per model | 500 | 500 |
Together's uncached input and output prices match DeepSeek's peak-hour rates exactly. DeepSeek states that off-peak rates are half the peak rates, and defines peak as 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday. That is 35 hours out of 168, or 20.8% of the week. For 79.2% of the week, the same model on DeepSeek's own API costs half what Together charges for it.
The cache line is wider still. Together's cached-input rate of $0.13 per million is 2.95 times DeepSeek's peak cache-hit rate of $0.044 and 5.91 times the off-peak rate of $0.022. That matters more than the headline rate for agent work, because DeepSeek enables disk-based context caching by default for every user, and a 146-step rollout re-sends a growing prefix on every step.
The $0.24 figure does not reconcile with Together's own list price
The study reports $109 total across 452 DeepSeek V4 Pro rollouts, an average of $0.24 each, and a median of 101,000 output tokens per rollout.
Take Together's published output price of $3.96 per million and apply it to that median: 101,000 tokens is $0.40 in output charges alone, 67% more than the $0.24 average total, before a single input token is billed. A median is not a mean, and output distributions in agent runs usually skew so that the mean sits above the median, not below it. Run the same check on the Sol side and it holds together: 59,000 output tokens at the then-current $30 per million is $1.77, leaving about $6.60 of input inside the reported $8.37, which is what roughly 1.3 million billed input tokens across 53 steps would look like.
So the Sol cost reconciles against public pricing and the DeepSeek cost does not. The likely explanation is benign, that the run was not billed at serverless list price. It is still a number you cannot reproduce from any published rate card, which is exactly the number a routing decision hangs on.
The Sol side of the comparison expired three days later
On 21 August 2026 the OpenAI API changelog recorded that GPT-5.6 Sol now costs $4 per million input tokens and $20 per million output tokens, "representing 20% lower input pricing and 33% lower output pricing", and that the promotional pricing is available at least through 21 November 2026. The implied previous rates are $5 and $30, which is what the study's $8.37 was computed against.
By our arithmetic, the same rollout now costs between $5.61 and $6.70, depending on the input and output split: $6.70 if the whole bill were input, $5.61 if it were all output. The cascade's $3.35 falls too, but by less, because only its escalation stage repriced and the DeepSeek first stage did not. The study's own margin, ten points of accuracy for less than half of Sol's price, is therefore narrower now than when it was published, and it narrows again if you move the first stage to DeepSeek's off-peak API.
The 272,000-token cliff the changelog does not mention
The changelog gives one price for Sol. The model card gives two. GPT-5.6 Sol has a 1,050,000-token context window and a 922,000-token maximum input, and the card states that prompts above 272,000 input tokens are priced at 2x input and 1.5x output "for the full request". The OpenAI pricing table carries that as separate long-context columns, $8 input and $30 output against $4 and $20.
| GPT-5.6 Sol tier, per 1M tokens | Input | Cached input | Output |
|---|---|---|---|
| Standard, up to 272K input | $4.00 | $0.40 | $20.00 |
| Standard, above 272K input | $8.00 | $0.80 | $30.00 |
| Batch, up to 272K input | $2.00 | $0.20 | $10.00 |
| Batch, above 272K input | $4.00 | $0.40 | $15.00 |
| Fast mode, up to 272K input | $8.00 | $0.80 | $40.00 |
Two things are worth knowing before you size an agent on this. The multiplier applies to the entire request, not to the tokens above the threshold, so crossing 272,001 input tokens reprices every token in that call. And the threshold is printed in the pricing table's row labels for gpt-5.5, gpt-5.5-pro, gpt-5.4 and gpt-5.4-pro, all marked "(<272K context length)", but for no row in the GPT-5.6 family. The string "long context" appears zero times in the changelog. A reader who priced Sol from the 21 August entry alone would have no way to know a second tier exists.
Sol's median peak context in the Together run was 177,000 tokens, comfortably under the cliff. The 922,000-token input ceiling means 70.5% of the model's addressable input range sits above it.
Two DeepSeek V4 Pro entries, and one of them is not the current model
Together's serverless table lists DeepSeek V4 Pro twice.
| Together model ID | Context | Input | Cached input | Output |
|---|---|---|---|---|
| deepseek-ai/DeepSeek-V4-Pro-0813 | 1,048,576 | $1.32 | $0.13 | $3.96 |
| deepseek-ai/DeepSeek-V4-Pro | 512,000 | $1.74 | $0.20 | $3.48 |
| deepseek-ai/DeepSeek-V4-Flash-0731 | 1,000,000 | $0.14 | $0.03 | $0.28 |
DeepSeek's own pricing page documents exactly one V4 Pro, and gives its model version as DeepSeek-V4-Pro-0813 with a 1M context length and a 384K maximum output. The unsuffixed deepseek-ai/DeepSeek-V4-Pro entry, at 512,000 context, does not correspond to a version DeepSeek currently documents. It is also the name a developer reaches for first, and picking it costs 31.8% more per input token for half the context window. The output price is the one line that goes the other way, $3.48 against $3.96.
What to do
Re-run the arithmetic on your own traffic before you adopt the cascade. Three checks, in order.
First, plot your request volume against DeepSeek's peak windows. If most of your agent traffic falls outside 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, going direct to DeepSeek halves the first stage's token cost against Together's flat rate, and Together's convenience has to earn that gap back in throughput or reliability.
Second, measure your cache-hit ratio before you choose a host. At a 5.91x spread on cache-hit input, the hosting decision for a long multi-step agent is decided by the cache line, not the headline rate.
Third, instrument the 272,000-token boundary as a hard alert, not a dashboard metric. A request that crosses it reprices in full at 2x input and 1.5x output, and nothing in your token counter will flag it.
The general lesson is one we keep re-learning on gateway work: a published benchmark is a measurement of a moment, and inference price lists now move faster than the studies that cite them. If you are choosing between hosts, compare the gateway layer itself rather than only the model, and treat any per-task cost as a variable with a date attached. The wider decision framework sits in our note on hybrid model routing and API spend, and the peak-window mechanics are worked through in detail in our piece on DeepSeek V4 peak and off-peak pricing. For the OpenAI side of this month's repricing, see our coverage of the GPT-5.6 Sol price cut across AI gateways and the broader frontier model comparison.
India-specific considerations
DeepSeek's peak windows are defined in UTC, and India is UTC+5:30, so the two windows land at 06:30 to 09:30 IST and 11:30 to 15:30 IST. The second one sits squarely inside the Indian working day. A Gurugram or Bengaluru team running agent workloads at their own desks pays peak rates for four hours of every working afternoon, while a US-hours team runs almost entirely off-peak on the same account.
That is a scheduling problem with a real number on it. At $0.66 against $1.32 per million uncached input tokens, an Indian team that can move batch and evaluation runs out of the 11:30 to 15:30 IST window halves the token cost of that slice with no code change. Interactive traffic cannot move, but regression suites, nightly evaluations and bulk re-indexing can. Under the Digital Personal Data Protection Act 2023, note separately that routing prompts through a third-party inference host is a processor relationship you have to document, whichever rate you pay.
What is still unknown
Together has not published how the study's DeepSeek costs were computed, so the gap between the $0.24 average and the $0.40 that list price implies for the median rollout stays unexplained. Together's serverless table does not publish a per-model concurrency limit against DeepSeek's stated 500. And OpenAI's promotional pricing for Sol is guaranteed only "at least through November 21, 2026", with no published rate after that date, so any cascade economics you fix today have a 21 November review point built in.
FAQ
How eCorpIT can help
Inference price lists now move faster than the benchmarks that cite them, and a routing rule written against last week's rate card quietly overspends. Our senior engineering teams build and operate the gateway layer that makes model routing measurable, including per-model cost attribution, cache-hit instrumentation and alerting on repricing thresholds like the 272,000-token boundary. That work is described on our AI gateway and model routing FinOps service page. To have your current model spend re-priced against today's published rates, book a model routing cost review.
References
- Together AI, DeepSeek V4 Pro 0813 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing, published 18 August 2026.
- Together AI documentation, Serverless models and pricing.
- DeepSeek API documentation, Models and pricing.
- DeepSeek API documentation, Change log, entries dated 13 and 21 August 2026.
- DeepSeek API documentation, Context caching.
- DeepSeek API documentation, Rate limit and isolation.
- OpenAI, API changelog, entry dated 21 August 2026.
- OpenAI, API pricing.
- OpenAI, GPT-5.6 Sol model card.
- OpenAI, Your data guide.
- Together AI, Rate limits.
- Together AI, Dedicated inference.
Last updated: 24 August 2026.