Groq 3 LPX hit full production on 24 August 2026: NVIDIA names Nebius first, Groq second

Groq 3 LPX reached full production 24 August 2026. NVIDIA names Nebius first and Groq second; no GA date.

Read time
11 min
Word count
1.6K
Sections
10
FAQs
8
Share
Groq 3 LPX reached full production on 24 August 2026 with Nebius named the first AI cloud adopter
NVIDIA Groq 3 LPX, announced in full production on 24 August 2026.
On this page · 10 sections
  1. What NVIDIA actually announced
  2. Three companies, one day, three tenses
  3. What the benchmark covers, and what it does not
  4. The naming is not a coincidence
  5. What a buyer should do with this
  6. India-specific considerations
  7. What is still unknown
  8. FAQ
  9. How eCorpIT can help
  10. References

Summary. NVIDIA announced at Hot Chips on 24 August 2026 that Groq 3 LPX, its interactive inference accelerator for the Vera Rubin NVL72 platform, is in full production. The press release states plainly that "Nebius is the first AI cloud to adopt NVIDIA Groq 3 LPX", and that "following Nebius", Groq "plans to be among the platform's earliest adopters". Groq published its own post the same day headlined "Groq Among the First to Bring NVIDIA Groq 3 LPX and Vera Rubin NVL72 to Market". The single benchmark behind every performance claim is one model at one context length: 3,400 output tokens per second on Gemma 4 31B with a 100,000-token context in Artificial Analysis testing, and 4x faster responsiveness than an unnamed "nearest alternative platform". No party published a general-availability date, a price per token, or a Region. Groq raised $350 million on 17 August 2026, eight days before the announcement.

What NVIDIA actually announced

Groq 3 LPX is an accelerator that extends Vera Rubin NVL72 by raising token generation rate, which NVIDIA calls interactivity: the rate at which tokens are produced for an individual user, and therefore how quickly an agent completes each step of its work. The framing is agentic workloads, which run hundreds or thousands of inference steps and generate large token volumes per task.

"Full production" is a manufacturing status. It says the part is being built at volume. It does not say a cloud is serving it, and NVIDIA's own wording is careful about that distinction throughout.

Three companies, one day, three tenses

Source, 24 August 2026 What it says about availability Tense
NVIDIA press release Groq 3 LPX "is now in full production" Present, about manufacturing
NVIDIA press release, on Nebius Nebius "plans to bring NVIDIA Groq 3 LPX to Nebius Token Factory" Future
NVIDIA press release, on Groq "Following Nebius", Groq "plans to be among the platform's earliest adopters" Future, and second in order
Groq blog post headline "Groq Among the First to Bring NVIDIA Groq 3 LPX and Vera Rubin NVL72 to Market" Present-tense framing
Groq blog post body "When Groq brings NVIDIA Groq 3 LPX capacity online" Future, undated
Dell Technologies quote in Groq's post "turning industry-leading performance into deployable infrastructure that customers can use today" Present

Read the six rows together and the picture is clear enough. The silicon is in volume production. No cloud is serving it yet. One quote in one press release says customers can use it today, and the two companies that would actually serve it both use the word "plans".

Nebius is named first by NVIDIA, in two separate NVIDIA documents on the same day: the press release and the accompanying NVIDIA blog post on the Vera Rubin platform. Groq's post does not mention Nebius.

None of this is a false statement by anyone. "Among the first" is true of a company NVIDIA places second. The point for a buyer is narrower: if you are planning capacity around this part, the queue position matters, and NVIDIA published it while Groq did not.

What the benchmark covers, and what it does not

Every performance number in circulation traces to one Artificial Analysis run. NVIDIA describes it as "a record 3,400 output tokens per second in Artificial Analysis benchmarking running Gemma 4 31B, an open source agentic model, with a 100,000-token context critical for agentic systems", and calls it "the fastest performance ever recorded for the model". Groq's post restates the same figure. The two match, which is worth saying, because vendor restatements of a partner's benchmark often do not.

Claim What is specified What is not
3,400 output tokens per second Model Gemma 4 31B, 100,000-token context, Artificial Analysis Batch size, concurrency, precision, rack configuration
"Fastest performance ever recorded for the model" Scoped to Gemma 4 31B Any other model, any other context length
4x faster responsiveness Compared against "the nearest alternative platform" Which platform, which configuration, which model
"Coding in minutes versus hours" Agentic coding as the workload class Task set, baseline, measurement method
Cost per token Nothing published Price, Region, instance shape, commitment terms

One model at one context length is a legitimate benchmark and a narrow one. Gemma 4 31B is a 31-billion-parameter open model; a 31B model at 100,000 tokens of context is not the same workload as a large mixture-of-experts model at 8,000 tokens, and the ratio between platforms will not carry across unchanged. Treat 3,400 tokens per second as a ceiling observed on one configuration, not a rate you will see on your model.

The unnamed comparator is the weaker part. "4x faster responsiveness than the nearest alternative platform" cannot be reproduced by a reader, because the reader does not know what was on the other side.

The naming is not a coincidence

The product is called NVIDIA Groq 3 LPX, and the Groq name in it is not marketing overlap. Groq and NVIDIA entered a non-exclusive inference technology licensing agreement, announced on Groq's blog on 24 December 2025. Groq became an NVIDIA Cloud Partner on 12 August 2026, which its own post describes as designing, deploying and operating accelerated computing to NVIDIA's reference architecture and operational standards. Twelve days later, Groq announced it will deploy an NVIDIA-branded product carrying its own name, built into Dell Technologies infrastructure.

That is an unusual position for an accelerator company: the LPU architecture Groq built its cloud on is now also an NVIDIA product line, and Groq is a customer of it. Groq's advantage in the announcement is stated as operational rather than architectural. Its post says Groq is "the only team with hands-on experience operating LPUs in production at scale", with trillions of tokens generated weekly across data centres in North America, Europe, the Middle East and APAC.

Capital timing is worth noting alongside it. Groq's blog lists a $650 million raise on 22 June 2026 and a $350 million Series A closing on 17 August 2026, eight days before this announcement. Deploying rack-scale Vera Rubin systems is a capital-intensive commitment, and the funding sequence is the part of the story that has a date attached.

What a buyer should do with this

Nothing yet, and that is the honest answer. There is no way to buy Groq 3 LPX capacity today from either named cloud, no published price per million tokens, and no Region list.

Three things are worth doing in the meantime.

Benchmark your own model, not Gemma 4 31B. If your agent runs a 70B-class model at 32,000 tokens of context, the published figure tells you very little. Record your current tokens per second per user and your p95 time to first token now, so that when capacity appears you have a baseline to compare against rather than a vendor ratio to trust.

Separate interactivity from throughput in your requirements. NVIDIA is explicit that Groq 3 LPX targets the generation phase and the per-user token rate. If your cost problem is batch throughput on long prompts rather than latency in an interactive loop, this part is aimed at a different bottleneck and a Vera Rubin NVL72 configuration without it may serve you better.

Ask both vendors for the queue position in writing. NVIDIA published an order. If your capacity plan depends on being early, the order is the number that matters, and it is the one that changes quietly.

One more thing to keep in view: NVIDIA's press release carries standard forward-looking-statement language covering expectations about availability of its products and its third-party arrangements. That safe-harbour paragraph is boilerplate in every such release, and it is also an accurate description of what a partner adoption statement is.

India-specific considerations

Groq lists data centres in North America, Europe, the Middle East and APAC without naming Indian sites, and Nebius Token Factory advertises multi-region routing with 99.9% uptime and sub-second targets without naming an Indian Region either. For an Indian team with data-residency obligations under the Digital Personal Data Protection Act 2023, that means the interactivity gain is not yet available on terms that keep inference in-country.

The practical read for a team in Gurugram or Bengaluru planning agent workloads through 2027: treat this as a signal about where per-user token rates are heading, not as capacity you can schedule. If residency is a hard requirement, the near-term options remain in-country GPU capacity or an in-Region endpoint from a provider that names the Region, and the decision belongs in the same review as the rest of your inference spend.

What is still unknown

No general-availability date from Nebius or Groq. No price per million tokens, no instance shape and no commitment terms. No Region list for either provider. No benchmark on a second model or a second context length. No named comparator behind the 4x interactivity claim. No public detail on how a Groq 3 LPX rack is configured relative to a standard Vera Rubin NVL72 rack, or what the power draw difference is. Until those exist, the only firm facts are the production status and the adoption order.

FAQ

How eCorpIT can help

The decision this announcement creates is a planning decision, not a purchase: whether to hold an inference workload on current capacity or reserve budget for an interactivity gain with no date and no price attached. eCorpIT sizes and runs inference platforms for teams shipping agent workloads, and our engineering organisation is CMMI Level 5 and ISO 27001:2022 certified. If you are building a 2027 inference budget, book an inference capacity review and we will baseline your current tokens per second and latency against the configurations that are actually purchasable now. Teams staffing this internally can hire AI engineers to own the measurement.

References

  1. NVIDIA Groq 3 LPX Now in Full Production With World-Class Speed for Agentic AI, NVIDIA Newsroom, 24 August 2026.
  1. With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agents, NVIDIA Blog, 24 August 2026.
  1. Groq Among the First to Bring NVIDIA Groq 3 LPX and Vera Rubin NVL72 to Market, Groq, 24 August 2026.
  1. Groq blog index, Groq, listing post dates for the 22 June 2026 and 17 August 2026 funding announcements and the 24 December 2025 NVIDIA licensing agreement.
  1. SpaceXAI Adopts NVIDIA Vera CPU to Accelerate Agentic AI at Massive Scale, NVIDIA Newsroom, 24 August 2026.
  1. NVIDIA Newsroom, latest news index, retrieved 25 August 2026.
  1. Nebius Token Factory, Nebius, retrieved 25 August 2026.
  1. Groq documentation, supported models, Groq, retrieved 25 August 2026.
  1. Thank you, 1M developers building with GroqCloud, Groq.
  1. Batch processing with GroqCloud for AI inference workloads, Groq, on the throughput path distinct from interactivity.
  1. From speed to scale: how Groq is optimized for MoE and other large models, Groq.
  1. The Groq LPU explained, Groq, on the architecture now licensed to NVIDIA.

Related reading: NVIDIA Vera Rubin cloud instances for AI workloads, Rubin GPU buildout and cloud budgets, Groq's Llama 3.3 70B shutdown and its replacement, and AI compute capacity planning.

Last updated: 25 August 2026.

Frequently asked

Quick answers.

01 Which cloud gets Groq 3 LPX first?
NVIDIA states that Nebius is the first AI cloud to adopt NVIDIA Groq 3 LPX, in both its press release and its Vera Rubin blog post of 24 August 2026. The same release says that following Nebius, Groq plans to be among the platform's earliest adopters. Groq's own post does not mention Nebius.
02 Is Groq 3 LPX available to buy today?
No. NVIDIA says the accelerator is in full production, which describes manufacturing rather than cloud availability. Nebius plans to bring it to Nebius Token Factory and Groq plans to be an early adopter. Neither company published a general-availability date, a price per token, or a Region list.
03 What does the 3,400 tokens per second figure actually measure?
One Artificial Analysis benchmark run on Gemma 4 31B, an open source agentic model, with a 100,000-token context. NVIDIA calls it the fastest performance ever recorded for that model. Batch size, concurrency, precision and rack configuration are not published, so the figure is not directly portable to another model.
04 What is the 4x claim compared against?
An unnamed platform. NVIDIA writes that Groq 3 LPX provides 4x faster responsiveness for agents and latency-sensitive workloads than the nearest alternative platform, without identifying it. Groq restates the same claim as 4x higher interactivity. Without a named baseline and configuration, a reader cannot reproduce the comparison.
05 Why does an NVIDIA product carry the Groq name?
Groq and NVIDIA announced a non-exclusive inference technology licensing agreement on 24 December 2025, and Groq became an NVIDIA Cloud Partner on 12 August 2026. The result is that the LPU architecture Groq built its inference cloud on now also ships as an NVIDIA product line, which Groq itself plans to deploy through Dell Technologies infrastructure.
06 Does this change anything for a team running inference today?
Not yet. The useful preparation is measuring your own workload now: tokens per second per user and p95 time to first token on your actual model and context length. That gives you a baseline to test against when capacity appears, instead of relying on a ratio measured on a 31B model at 100,000 tokens.
07 Who else is adopting the Vera Rubin platform?
NVIDIA's blog names SpaceXAI, which announced that NVIDIA Vera CPUs will power its next generation of agentic AI, and CoreWeave, which has deployed Spectrum-X Multiplane into production to connect Vera Rubin racks. Nebius is named as the first AI cloud specifically for Groq 3 LPX, with Groq following.
08 What did Groq raise recently?
Groq's blog lists a $650 million raise on 22 June 2026 and a $350 million Series A closing on 17 August 2026, eight days before the Groq 3 LPX announcement. Deploying rack-scale Vera Rubin systems with Dell Technologies is capital-intensive, so the funding sequence is relevant context for the deployment timeline.

About the author

Manu Shukla

Founder & Director

Founder of eCorpIT. Hands-on engineer leading senior-only delivery for AI apps, custom software, and cloud systems for global clients.

Subscribe

One engineering note a week. No fluff, no spam.

Senior-architect playbooks on AI agents, mobile apps, cloud, security, data, and marketing — delivered every Wednesday.

Past the reading

Read enough. Let's build something.

A senior architect responds in 24 working hours with scope, indicative cost, and a timeline. NDA before any technical conversation.