Synthszr Charts — die großen AI-Marken im Wettkampf ums Podium
synthszr charts
llama-cpp

Llama Cpp · seit März 2023 (Projektstart durch Georgi Gerganov) · 28× · zuletzt 30. Juni 2026

16
Momentum

llama.cpp ist eine quelloffene C/C++-Bibliothek für die lokale und cloud-basierte Inferenz großer Sprachmodelle, ursprünglich von Georgi Gerganov entwickelt. Sie läuft ohne externe Abhängigkeiten wie Python oder PyTorch, unterstützt zahlreiche Hardware-Backends (CPU, CUDA, Metal, Vulkan, SYCL u.a.) und nutzt das quantisierte GGUF-Modellformat, um LLMs auch auf Consumer-Hardware effizient auszuführen. Der integrierte llama-server bietet eine OpenAI-kompatible HTTP-API und wird von Projekten wie Ollama als Backend genutzt. Das Projekt wird von der ggml-org-Community weiterentwickelt und veröffentlicht statt klassischer Versionsnummern fortlaufend build-getaggte Releases.

Momentum-Verlauf
19.05.17.08.

Features

Deployment (Self-host/Cloud)Self-hosted (lokal, Server) sowie Cloud-Deployment möglich, u.a. via Hugging Face Inference Endpoints
Durchsatz/Latenz3–8x schnellere Inferenz als Python-Frameworks, besonders auf CPU (laut Anbieter, hardwareabhängig)
LizenzMIT-Lizenz
PlattformLinux, macOS, Windows, Android, Raspberry Pi, Browser (WebGPU); x86 (AVX/AVX2/AVX512/AMX), ARM, Apple Silicon (Metal)
PreisKostenlos, Open Source (keine Lizenzgebühren)
Protokoll-KompatibilitätOpenAI-kompatible API (/v1/completions, /v1/chat/completions, /v1/embeddings) über llama-server
Release-DatumMärz 2023 (Projektstart); fortlaufende Build-Releases, aktuell z.B. b9838 (Ende Juni 2026)
Unterstützte Modelle/ProviderAlle GGUF-kompatiblen Modelle: Llama 1/2/3, Mistral, Phi, Gemma, Qwen, DeepSeek, Yi, Solar, StableLM u.a.

Weitere Produkte in dieser Kategorie: Lokale LLM-Runtimes

Belege (28)

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.