

Groq · seit 2024-02-19 (Soft-Launch GroqCloud Developer Platform) · 16× · zuletzt 24. Juli 2026
Groq ist ein US-amerikanisches KI-Hardware-Unternehmen (gegründet 2016), das mit der Language Processing Unit (LPU) einen speziell für KI-Inferenz entwickelten Chip anbietet. Die LPU nutzt statt externem HBM-Speicher ausschließlich On-Chip-SRAM (230 MB pro Chip) und erreicht dadurch sehr niedrige, deterministische Latenzen bei Textgenerierung. Über die Cloud-Plattform GroqCloud wird der Chip nutzungsbasiert (Pay-per-Token) für offene Modelle wie Llama, Mixtral oder DeepSeek angeboten; zusätzlich ist die Hardware als GroqCard, GroqNode und GroqRack für On-Premise-Einsatz verfügbar. Ende 2025 lizenzierte NVIDIA die LPU-Architektur für rund 20 Mrd. USD, GroqCloud wird jedoch weiterhin unabhängig betrieben.
Features
| Fertigungsprozess (nm) | 14 nm (erste LPU-Generation, GlobalFoundries) |
| Lizenz | Proprietäre Hardware/Architektur (LPU™, GroqChip™, GroqCard™ sind eingetragene Marken von Groq, Inc.) |
| Plattform | GroqCloud (Cloud-API, OpenAI-kompatibel), GroqRack/GroqNode für On-Premise-Einsatz |
| Preis | GroqCard-Beschleuniger ca. $19.948; GroqCloud API ab $0,05 pro 1 Mio. Input-Token (Llama 3.1 8B), kostenloser Tarif ohne Kreditkarte verfügbar |
| Rechenleistung (FLOPS/TOPS) | Bis zu 750 TOPS (INT8) bzw. 188 TFLOPS (FP16 @ 900 MHz) pro GroqChip; GroqRack bis zu 12 PetaFLOPS |
| Release-Datum | Erste LPU-Generation 2019 eingeführt; GroqCloud Developer-Plattform ab 19. Februar 2024 öffentlich verfügbar |
| Speicher | 230 MB On-Chip-SRAM pro Chip, 80 TB/s On-Die-Speicherbandbreite; GroqRack bis zu 14 GB globales SRAM |
| Verfügbarkeit | GroqCloud API weltweit online (vier Regionen), Free-Tier ohne Kreditkarte; GroqCard/GroqRack auf Anfrage/On-Premise |