Synthszr Charts — die großen AI-Marken im Wettkampf ums Podium
synthszr charts
vllm

Vllm · since Juni 2023 (offizieller Launch-Blogpost); zugehöriges Paper September 2023 · 59× · last seen Aug 13, 2026

100
Momentum

vLLM is an open-source inference and serving engine for large language models, originally developed at UC Berkeley's Sky Computing Lab. Its core innovation is PagedAttention, a memory-management technique for the KV cache, combined with continuous batching, enabling high throughput and memory-efficient model serving. The project supports over 200 HuggingFace model architectures, numerous quantization formats, and various hardware backends (NVIDIA, AMD, Intel, TPU, Trainium). vLLM is distributed as free open-source software under the Apache 2.0 license and can be self-hosted or run via an OpenAI-compatible HTTP server.

Momentum trend
19.05.17.08.

Features

Deployment (Self-host/Cloud)Self-hosted via pip/uv-Installation oder Docker-Image (vllm/vllm-openai); K8s-natives Deployment über production-stack
Durchsatz/LatenzBis zu 24x höherer Durchsatz als HuggingFace Transformers (offizieller Launch-Benchmark 2023)
LizenzApache License 2.0
PlattformPython-Bibliothek/Server; unterstützt NVIDIA-, AMD-GPUs, Intel-CPUs, TPU, AWS Trainium/Inferentia, Apple Silicon (via vLLM-Metal); Python 3.10–3.13
PreisKostenlos, Open Source (keine Lizenzgebühren; Kosten entstehen nur durch selbst betriebene Infrastruktur)
Protokoll-KompatibilitätOpenAI-kompatibler HTTP-Server (Chat Completions, Responses API, SageMaker-kompatibler Endpunkt)
Release-DatumJuni 2023 (Launch-Blogpost 'Easy, Fast, and Cheap LLM Serving with PagedAttention')
Unterstützte Modelle/Provider200+ Modellarchitekturen auf HuggingFace (u.a. Llama, DeepSeek, Qwen); Quantisierung: FP8, MXFP8/MXFP4, NVFP4, INT8, INT4, GPTQ/AWQ, GGUF

More products in this category: LLM Inference & Serving

Sources (59)

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.