Synthszr Charts — die großen AI-Marken im Wettkampf ums Podium
synthszr charts
lmsys

Lmsys · depuis Januar 2024 · 11× · vu le 08 juil. 2026

38
Momentum

SGLang is an Open Source serving framework for large language models (LLMs) and multimodal models, developed by researchers from UC Berkeley/Stanford and hosted by the nonprofit organization LMSYS. At its core is RadixAttention for efficient KV-Cache reuse, complemented by zero-overhead scheduling, continuous batching, speculative decoding (including the new DSpark method), and broad quantization support. The framework is deployed in production on over 400,000 GPUs worldwide and supports numerous models such as Llama, Qwen, DeepSeek, GLM, Gemma, and Mistral. SGLang offers an OpenAI-compatible API and runs on NVIDIA, AMD, Intel Xeon, Google TPU, and Ascend NPU hardware.

Historique du momentum
19.05.17.08.

Fonctionnalités

Throughput/LatencyDeepSeek-V4-Pro: 383.7 tok/s at B=1 on B300 with DSpark; up to ~20% higher throughput under high concurrency vs. fixed budget
LicenseApache License 2.0
PlatformNVIDIA, AMD, Intel Xeon, Google TPU, Ascend NPU; Linux, Docker, Kubernetes
PriceFree, open source (no license fees)
Protocol CompatibilityOpenAI-compatible API; compatible with Hugging Face APIs
Release DateInitial release: January 17, 2024
Supported Models/ProvidersLlama, Qwen, DeepSeek, Kimi, GLM, GPT, Gemma, Mistral and more; also embedding, reward and diffusion models (WAN, Qwen-Image)

Plus de produits dans cette catégorie: Inférence & serving LLM

Preuves (11)

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.