Synthszr Charts — die großen AI-Marken im Wettkampf ums Podium
synthszr charts
vllm-project

Vllm Project · 2× · naposledy 25. 8. 2026

35
Momentum

vLLM is an Open Source library for efficient Inference and serving of Large Language Models, originally developed at UC Berkeley's Sky Computing Lab. Its core is the PagedAttention algorithm, which organizes KV-Cache memory in blocks analogous to virtual memory management, drastically reducing memory waste. vLLM supports hundreds of model architectures from Hugging Face, offers an OpenAI-compatible HTTP server, and runs on diverse hardware (NVIDIA, AMD, Intel GPUs, TPUs, AWS Trainium/Inferentia, and others). The project is licensed under Apache 2.0 and is being further developed by a broad community from academia and industry.

Vývoj momenta
27.05.25.08.

Vlastnosti

Deployment (Self-Hosted/Cloud)Self-hosted via Docker/Kubernetes (Helm charts, 'Production Stack'); deployable on cloud platforms such as Google Cloud (GKE, Compute Engine, Vertex AI)
Throughput/LatencyUp to 24x higher throughput than HuggingFace Transformers and 2.2x-3.5x higher than HuggingFace TGI (depending on scenario)
LicenseApache License 2.0
PlatformPython library/server; runs on NVIDIA, AMD, Intel GPUs/CPUs, TPU, PowerPC, AWS Trainium/Inferentia
PriceFree, open source (self-hosted); costs only for own infrastructure/cloud compute
Protocol CompatibilityHTTP server compatible with OpenAI Completions, Chat, and Embeddings API; endpoints include /v1/chat/completions, /v1/completions, /v1/embeddings
Release DateInitial release June 2023; ongoing active development with frequent version releases (e.g. v0.14.1)
Supported Models/Providers200+ model architectures via Hugging Face, incl. Llama, Qwen, GPT-OSS, LLaVA (multimodal)

Další produkty v této kategorii: LLM inference a serving

Zdroje (2)

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.