

Groq Llama
#14 v LLM inference a servingGroq · 2× · naposledy 30. 6. 2026
1
Momentum
Groq Llama refers to the deployment of Meta's Llama models (including Llama 3.1, 3.3, 4 Scout/Maverick) on GroqCloud, Groq's cloud Inference platform. The models run on Groq's proprietary LPU chips (Language Processing Unit) instead of GPUs, which according to the provider enables significantly higher throughput and lower, predictable latency. Access is provided via a Token-based API in OpenAI-compatible format, available either through GroqCloud (cloud) or GroqRack (on-premise, upon request). Pricing varies considerably depending on model size, ranging from very affordable smaller models to more expensive 70B-class models.
Vývoj momenta
19.05.17.08.