

Ollama Cloud
#2 in Local LLM RuntimesOllama · 7× · last seen Jul 29, 2026
82
Momentum
Ollama Cloud is a hosted Inference service by the Ollama team that runs large Open-Weight models (e.g., gpt-oss:120b, Kimi K3, DeepSeek, Qwen3-Coder) on Ollama's own data center GPUs without requiring local hardware. The service was introduced in September 2025 (with Ollama v0.12) as "Cloud models" in Preview and uses the same local CLI, REST API, and OpenAI-compatible API as the local Ollama Runtime. Billing is based on GPU time rather than Tokens through tiered subscription levels (Free, Pro, Max); the local Ollama software itself remains MIT-licensed and free to use without limitations.
Momentum trend
01.05.30.07.