

Groq 3 LPX
#1 en Matériel d'inférence IAGroq · v3 · lpx · depuis 24. August 2026 (Vollproduktion angekündigt); Auslieferung/Verfügbarkeit ursprünglich für Q3 2026 geplant · 18× · vu le 29 août 2026
NVIDIA Groq 3 LPX is a rack-scale AI inference accelerator that extends the NVIDIA Vera Rubin platform as a dedicated decode-phase co-processor. An LPX rack consists of 256 interconnected Groq 3 LPU chips (the next generation of the Groq LPU architecture NVIDIA licensed for about $20 billion in December 2025), working alongside Vera Rubin GPUs to handle long-context, low-latency agentic AI workloads. The product entered full production on August 24, 2026, is manufactured on a 4nm process at Samsung Foundry, and Nebius is the first announced customer. In Artificial Analysis benchmarking it achieved 3,400 output tokens/second on the Gemma 4 31B model.
Fonctionnalités
| Fertigungsprozess (nm) | 4-nm-Prozess (Samsung Foundry SF4X) |
| Plattform | Erweiterung der NVIDIA Vera-Rubin-Plattform (Vera-Rubin NVL72), Rubin-GPUs übernehmen Prefill, LPX übernimmt Decode |
| Preis | Kein öffentlicher Stückpreis; NVIDIA nennt Zielpreis von 45 USD pro Million Tokens im Systembetrieb |
| Rechenleistung (FLOPS/TOPS) | 3.400 Output-Token/Sek. im Artificial-Analysis-Benchmark (Gemma 4 31B, 100.000-Token-Kontext); 4x schnellere Reaktionszeit als nächstbeste Alternativplattform |
| Release-Datum | Vollproduktion angekündigt am 24. August 2026 (Hot Chips 2026); Vorstellung auf GTC 2026 im März 2026 |
| Speicher | 500 MB SRAM pro LPU-Chip, 150 TB/s SRAM-Bandbreite, 2,5 TB/s Scale-up-Bandbreite; volles Rack (256 Chips) = 128 GB SRAM, 40 PB/s Bandbreite |
| Verfügbarkeit | Seit 24. August 2026 in Vollproduktion; erster Kunde Nebius, Rack-Deployment soll noch 2026 online gehen |