

llama.cpp
#3 in Lokale LLM-RuntimesGeorgi Gerganov · seit March 10, 2023 · 3× · zuletzt 09. Okt. 2026
75
Momentum
llama.cpp ist eine quelloffene C/C++-Bibliothek für LLM-Inferenz, ursprünglich von Georgi Gerganov entwickelt und heute im ggml-org-Projekt gepflegt. Sie läuft lokal auf CPU und GPU und bildet die Grundlage vieler lokaler LLM-Tools wie Ollama und LM Studio. Die offizielle Website llama.app stellt einen Installer und die Einstiegsdokumentation bereit.
Momentum-Verlauf
11.07.09.10.
Features
| Lizenz | MIT |
| Plattform | Windows, macOS, Linux (Docker, Homebrew) |
| Preis | Kostenlos (Open Source) |
| Protokoll-Kompatibilität | OpenAI-kompatible API (/v1/chat/completions, /v1/embeddings), Anthropic-Messages-kompatibel, Ollama-Shim (/api/tags) |
| Release-Datum | 10. März 2023 |
| Unterstützte Modelle/Provider | GGUF-Modelle (quantisiert), z.B. Qwen, Llama, Mistral, Gemma |