Synthszr Charts — die großen AI-Marken im Wettkampf ums Podium
synthszr charts
sglang

Sglang · 3× · vu le 24 août 2026

43
Momentum

SGLang is an Open Source framework for programming and serving Large Language Models and multimodal models, maintained by the nonprofit organization LMSYS. Its core is a backend runtime featuring RadixAttention for automatic KV-Cache reuse, along with capabilities such as continuous batching, speculative decoding, disaggregated Prefill/Decode, and quantization. The system is designed for low latency and high throughput in production workloads, supporting a wide range of open models (including Llama, Qwen, and DeepSeek) via an OpenAI-compatible API. According to its creators, SGLang runs in production on over 400,000 GPUs worldwide and processes trillions of Tokens daily.

Historique du momentum
26.05.24.08.

Fonctionnalités

Deployment (Self-Hosted/Cloud)Self-hosted from single GPU to distributed clusters; used in production across more than 400,000 GPUs worldwide
Throughput/LatencyDesigned for low latency and high throughput via RadixAttention, prefix caching, and multi-GPU parallelism
LicenseApache License 2.0
PlatformNVIDIA, AMD, Intel Xeon, Google TPU, and Ascend NPU hardware; install via pip, source, or Docker
PriceFree, open source (no license fees)
Protocol CompatibilityCompatible with OpenAI API and Hugging Face models
Release DateInitial release January 17, 2024; current versions ship continuously (e.g., v0.4 in December 2024)
Supported Models/ProvidersBroad support including Llama, Qwen, DeepSeek, Gemma, Mistral, GLM, LLaVA, plus embedding and reward models

Plus de produits dans cette catégorie: Inférence & serving LLM

Preuves (3)

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.