Synthszr Charts — die großen AI-Marken im Wettkampf ums Podium
synthszr charts
vllm

Vllm · v0.20.0 · 4× · tolest 30. Juni 2026

1
Momentum

vLLM 0.20.0 is an Open Source release of the quelloffenen Inference and Serving Engine vLLM (Apache-2.0 license), designed for high throughput and efficient memory management (PagedAttention) in LLM Inference. The release encompasses 752 commits from 320 contributors and introduces initial DeepSeek-V4 support, TurboQuant 2-bit KV-Cache quantization for quadrupled KV-Cache capacity, FA4 as the standard for MLA-Prefill, and a DeepSeek-specific MegaMoE path on Blackwell GPUs. Additionally, the standard toolchain switches to CUDA 13.0, PyTorch 2.11, and HuggingFace Transformers v5.

Momentum-Verloop
19.05.17.08.

Features

Deployment (Self-host/Cloud)Self-Hosting via pip/Docker (vllm/vllm-openai Image), Kubernetes, Multi-GPU/Multi-Node; keine eigene Cloud-Hosting-Option
Durchsatz/LatenzTurboQuant 2-Bit-KV-Cache: 4x KV-Kapazität; fused RMSNorm: ca. 2,1% End-to-End-Latenzverbesserung
LizenzApache 2.0 (permissiv, kommerzielle Nutzung erlaubt)
PlattformLinux; NVIDIA CUDA 13.0 (Standard), AMD ROCm, Intel XPU, TPU, CPU (inkl. ARM/RISC-V/PowerPC); PyTorch 2.11, Python bis 3.14
PreisKostenlos, Open Source (keine Lizenzgebühren)
Protokoll-KompatibilitätOpenAI-kompatibler HTTP-Server (Completions, Chat, Embeddings, Rerank u.a. Endpunkte)
Release-Datum28. April 2026 (v0.20.0)
Unterstützte Modelle/ProviderDeepSeek V4 (initial), Hunyuan v3 (Preview), plus breite HF-Transformers-v5-Modellbasis (Llama, Mistral, Qwen u.a.)

Mehr Produkten in disse Kategorie: LLM-Inferenz & Serving

Belege (4)

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.