Synthszr Charts — die großen AI-Marken im Wettkampf ums Podium
synthszr charts
vllm

Vllm · 3× · vu le 03 sept. 2026

51
Momentum

vLLM-Omni is an Open Source extension of the vLLM serving framework that supports omni-modal and non-autoregressive models (image, audio, video, robotics actions) alongside classical text-LLM Inference. It extends the efficiency techniques known from vLLM (e.g., KV-Cache management) with pipeline-like multi-stage execution for Diffusion Transformers and TTS architectures. The software is available as an open-source project under the Apache-2.0 license and is primarily self-hosted, though it is also used in production by third-party providers such as Baseten. Initially released in November 2025, it has since followed a continuous release cycle in parallel with vLLM major versions.

Historique du momentum
05.06.03.09.

Fonctionnalités

Deployment (Self-Hosted/Cloud)Primarily self-hosted on own GPU infrastructure; managed deployment also available via third parties such as Baseten
Throughput/LatencyTTS benchmarks measure RTF (realtime factor), TTFP (time to first packet), and Tput (generated audio seconds per wall-clock second); values are model-dependent
LicenseApache License 2.0
PlatformSupports CUDA (NVIDIA), ROCm (AMD), MUSA, NPU, and XPU as hardware backends
PriceFree, open source (Apache 2.0); costs only arise from self-managed GPU infrastructure or third-party managed hosting (e.g. Baseten, RunPod, Modal)
Protocol CompatibilityOpenAI-compatible API, including for text-to-speech (TTS) endpoints
Release DateFirst official release on November 30, 2025 (v0.11.0rc, built on vLLM v0.11.0)
Supported Models/ProvidersOmni models (Qwen3-Omni, MiniCPM-o 4.5, Cosmos3, HunyuanImage, BAGEL), TTS (Qwen3-TTS, IndexTTS 2.5, CosyVoice3), diffusion (MiniMax H3, Wan2.2, LTX), robotics (π0, GR00T-N1.7)

Plus de produits dans cette catégorie: Inférence & serving LLM

Preuves (3)

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.