On this page · 10 sections
Summary. NVIDIA announced at Hot Chips on 24 August 2026 that Groq 3 LPX, its interactive inference accelerator for the Vera Rubin NVL72 platform, is in full production. The press release states plainly that "Nebius is the first AI cloud to adopt NVIDIA Groq 3 LPX", and that "following Nebius", Groq "plans to be among the platform's earliest adopters". Groq published its own post the same day headlined "Groq Among the First to Bring NVIDIA Groq 3 LPX and Vera Rubin NVL72 to Market". The single benchmark behind every performance claim is one model at one context length: 3,400 output tokens per second on Gemma 4 31B with a 100,000-token context in Artificial Analysis testing, and 4x faster responsiveness than an unnamed "nearest alternative platform". No party published a general-availability date, a price per token, or a Region. Groq raised $350 million on 17 August 2026, eight days before the announcement.
What NVIDIA actually announced
Groq 3 LPX is an accelerator that extends Vera Rubin NVL72 by raising token generation rate, which NVIDIA calls interactivity: the rate at which tokens are produced for an individual user, and therefore how quickly an agent completes each step of its work. The framing is agentic workloads, which run hundreds or thousands of inference steps and generate large token volumes per task.
"Full production" is a manufacturing status. It says the part is being built at volume. It does not say a cloud is serving it, and NVIDIA's own wording is careful about that distinction throughout.
Three companies, one day, three tenses
| Source, 24 August 2026 | What it says about availability | Tense |
|---|---|---|
| NVIDIA press release | Groq 3 LPX "is now in full production" | Present, about manufacturing |
| NVIDIA press release, on Nebius | Nebius "plans to bring NVIDIA Groq 3 LPX to Nebius Token Factory" | Future |
| NVIDIA press release, on Groq | "Following Nebius", Groq "plans to be among the platform's earliest adopters" | Future, and second in order |
| Groq blog post headline | "Groq Among the First to Bring NVIDIA Groq 3 LPX and Vera Rubin NVL72 to Market" | Present-tense framing |
| Groq blog post body | "When Groq brings NVIDIA Groq 3 LPX capacity online" | Future, undated |
| Dell Technologies quote in Groq's post | "turning industry-leading performance into deployable infrastructure that customers can use today" | Present |
Read the six rows together and the picture is clear enough. The silicon is in volume production. No cloud is serving it yet. One quote in one press release says customers can use it today, and the two companies that would actually serve it both use the word "plans".
Nebius is named first by NVIDIA, in two separate NVIDIA documents on the same day: the press release and the accompanying NVIDIA blog post on the Vera Rubin platform. Groq's post does not mention Nebius.
None of this is a false statement by anyone. "Among the first" is true of a company NVIDIA places second. The point for a buyer is narrower: if you are planning capacity around this part, the queue position matters, and NVIDIA published it while Groq did not.
What the benchmark covers, and what it does not
Every performance number in circulation traces to one Artificial Analysis run. NVIDIA describes it as "a record 3,400 output tokens per second in Artificial Analysis benchmarking running Gemma 4 31B, an open source agentic model, with a 100,000-token context critical for agentic systems", and calls it "the fastest performance ever recorded for the model". Groq's post restates the same figure. The two match, which is worth saying, because vendor restatements of a partner's benchmark often do not.
| Claim | What is specified | What is not |
|---|---|---|
| 3,400 output tokens per second | Model Gemma 4 31B, 100,000-token context, Artificial Analysis | Batch size, concurrency, precision, rack configuration |
| "Fastest performance ever recorded for the model" | Scoped to Gemma 4 31B | Any other model, any other context length |
| 4x faster responsiveness | Compared against "the nearest alternative platform" | Which platform, which configuration, which model |
| "Coding in minutes versus hours" | Agentic coding as the workload class | Task set, baseline, measurement method |
| Cost per token | Nothing published | Price, Region, instance shape, commitment terms |
One model at one context length is a legitimate benchmark and a narrow one. Gemma 4 31B is a 31-billion-parameter open model; a 31B model at 100,000 tokens of context is not the same workload as a large mixture-of-experts model at 8,000 tokens, and the ratio between platforms will not carry across unchanged. Treat 3,400 tokens per second as a ceiling observed on one configuration, not a rate you will see on your model.
The unnamed comparator is the weaker part. "4x faster responsiveness than the nearest alternative platform" cannot be reproduced by a reader, because the reader does not know what was on the other side.
The naming is not a coincidence
The product is called NVIDIA Groq 3 LPX, and the Groq name in it is not marketing overlap. Groq and NVIDIA entered a non-exclusive inference technology licensing agreement, announced on Groq's blog on 24 December 2025. Groq became an NVIDIA Cloud Partner on 12 August 2026, which its own post describes as designing, deploying and operating accelerated computing to NVIDIA's reference architecture and operational standards. Twelve days later, Groq announced it will deploy an NVIDIA-branded product carrying its own name, built into Dell Technologies infrastructure.
That is an unusual position for an accelerator company: the LPU architecture Groq built its cloud on is now also an NVIDIA product line, and Groq is a customer of it. Groq's advantage in the announcement is stated as operational rather than architectural. Its post says Groq is "the only team with hands-on experience operating LPUs in production at scale", with trillions of tokens generated weekly across data centres in North America, Europe, the Middle East and APAC.
Capital timing is worth noting alongside it. Groq's blog lists a $650 million raise on 22 June 2026 and a $350 million Series A closing on 17 August 2026, eight days before this announcement. Deploying rack-scale Vera Rubin systems is a capital-intensive commitment, and the funding sequence is the part of the story that has a date attached.
What a buyer should do with this
Nothing yet, and that is the honest answer. There is no way to buy Groq 3 LPX capacity today from either named cloud, no published price per million tokens, and no Region list.
Three things are worth doing in the meantime.
Benchmark your own model, not Gemma 4 31B. If your agent runs a 70B-class model at 32,000 tokens of context, the published figure tells you very little. Record your current tokens per second per user and your p95 time to first token now, so that when capacity appears you have a baseline to compare against rather than a vendor ratio to trust.
Separate interactivity from throughput in your requirements. NVIDIA is explicit that Groq 3 LPX targets the generation phase and the per-user token rate. If your cost problem is batch throughput on long prompts rather than latency in an interactive loop, this part is aimed at a different bottleneck and a Vera Rubin NVL72 configuration without it may serve you better.
Ask both vendors for the queue position in writing. NVIDIA published an order. If your capacity plan depends on being early, the order is the number that matters, and it is the one that changes quietly.
One more thing to keep in view: NVIDIA's press release carries standard forward-looking-statement language covering expectations about availability of its products and its third-party arrangements. That safe-harbour paragraph is boilerplate in every such release, and it is also an accurate description of what a partner adoption statement is.
India-specific considerations
Groq lists data centres in North America, Europe, the Middle East and APAC without naming Indian sites, and Nebius Token Factory advertises multi-region routing with 99.9% uptime and sub-second targets without naming an Indian Region either. For an Indian team with data-residency obligations under the Digital Personal Data Protection Act 2023, that means the interactivity gain is not yet available on terms that keep inference in-country.
The practical read for a team in Gurugram or Bengaluru planning agent workloads through 2027: treat this as a signal about where per-user token rates are heading, not as capacity you can schedule. If residency is a hard requirement, the near-term options remain in-country GPU capacity or an in-Region endpoint from a provider that names the Region, and the decision belongs in the same review as the rest of your inference spend.
What is still unknown
No general-availability date from Nebius or Groq. No price per million tokens, no instance shape and no commitment terms. No Region list for either provider. No benchmark on a second model or a second context length. No named comparator behind the 4x interactivity claim. No public detail on how a Groq 3 LPX rack is configured relative to a standard Vera Rubin NVL72 rack, or what the power draw difference is. Until those exist, the only firm facts are the production status and the adoption order.
FAQ
How eCorpIT can help
The decision this announcement creates is a planning decision, not a purchase: whether to hold an inference workload on current capacity or reserve budget for an interactivity gain with no date and no price attached. eCorpIT sizes and runs inference platforms for teams shipping agent workloads, and our engineering organisation is CMMI Level 5 and ISO 27001:2022 certified. If you are building a 2027 inference budget, book an inference capacity review and we will baseline your current tokens per second and latency against the configurations that are actually purchasable now. Teams staffing this internally can hire AI engineers to own the measurement.
References
- NVIDIA Groq 3 LPX Now in Full Production With World-Class Speed for Agentic AI, NVIDIA Newsroom, 24 August 2026.
- With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agents, NVIDIA Blog, 24 August 2026.
- Groq Among the First to Bring NVIDIA Groq 3 LPX and Vera Rubin NVL72 to Market, Groq, 24 August 2026.
- Groq blog index, Groq, listing post dates for the 22 June 2026 and 17 August 2026 funding announcements and the 24 December 2025 NVIDIA licensing agreement.
- SpaceXAI Adopts NVIDIA Vera CPU to Accelerate Agentic AI at Massive Scale, NVIDIA Newsroom, 24 August 2026.
- NVIDIA Newsroom, latest news index, retrieved 25 August 2026.
- Nebius Token Factory, Nebius, retrieved 25 August 2026.
- Groq documentation, supported models, Groq, retrieved 25 August 2026.
- Batch processing with GroqCloud for AI inference workloads, Groq, on the throughput path distinct from interactivity.
- The Groq LPU explained, Groq, on the architecture now licensed to NVIDIA.
Related reading: NVIDIA Vera Rubin cloud instances for AI workloads, Rubin GPU buildout and cloud budgets, Groq's Llama 3.3 70B shutdown and its replacement, and AI compute capacity planning.
Last updated: 25 August 2026.