Technical positioning
Nvidia occupies the general-purpose end of the architecture spectrum. A Blackwell GPU will run a mixture-of-experts model, a dense transformer, a diffusion model, a recommendation model and a classical HPC workload, at good-to-excellent efficiency on all of them. An inference ASIC will beat it decisively on the workload it was designed for and may not run the others at all.
The strategic question is therefore about the stability of model architecture. If transformer-shaped decoding remains the dominant production workload for many years, the economic argument for specialisation strengthens over time. If architectures keep shifting — and the move toward reasoning models with long generated outputs and heavy KV cache pressure is a recent example of exactly that — flexibility retains real option value.
The Groq licensing deal is best read in this context. It gives Nvidia access to a deterministic, SRAM-heavy inference design that is architecturally unlike a GPU, at the point where the company's main competitive exposure is precisely the claim that GPUs are the wrong shape for decode.
That reading is no longer speculative. At GTC 2026 Nvidia unveiled the Groq 3 LPU, the first product of the licensing agreement, and positioned it inside the Vera Rubin platform as a dedicated decode co-processor: Rubin GPUs process long input contexts during prefill, and the SRAM-based LPU generates output tokens. It is a significant concession in principle — Nvidia is shipping a heterogeneous system in which the GPU is explicitly not the right device for half the workload. It is also a strong competitive position in practice, because Nvidia is now selling both halves rather than defending the GPU against a specialist selling one of them.