

Ollaya
#4 v Lokální LLM runtimeUnknown · 2× · naposledy 28. 9. 2026
Ollaya is an Open Source local runtime environment that executes specialized "Decision Models" (including Laya, Decider, NLI, GLiClass, Kev, Von, Qwen3Guard) on your own machine instead of generating text. It delivers typed, calibrated probabilities for classification, routing, and scoring tasks in milliseconds through an interface compatible with TypeSafe's System-One API (/v1/systemone). The software runs entirely locally (127.0.0.1), uses ONNX Runtime or llama.cpp for Inference on CPU/GPU, and is available as a CLI, desktop app (macOS/Windows/Linux), and Docker image. Ollaya is an independent community project and is not affiliated with Ollama or TypeSafe AI.
Vlastnosti
| Throughput/Latency | Laya (ModernBERT-large, 421M): 8–10 ms for 5 questions on RTX 4090; decider:2b ~178 ms; JevK5 (4B) ~288 ms for 5 questions on RTX 4090 |
| License | Apache-2.0 (runtime); individual models keep their own license (e.g. nli:deberta-v3-large is MIT) |
| Platform | Linux (x86_64/arm64), macOS (Apple silicon), Windows x64 (CPU); desktop app for macOS/Windows/Linux; Docker images incl. CUDA variant |
| Price | Free, open source; no metering, no API bill |
| Protocol Compatibility | TypeSafe-compatible: POST /v1/systemone, /v1/decisions, GET /v1/models wire-identical; official TypeSafe SDK works unchanged via TYPESAFE_BASE_URL |
| Release Date | Circa September 2026, shortly after TypeSafe AI's Jev launch (September 15, 2026) |
| Supported Models/Providers | Laya, Decider, NLI, GLiClass, Kev, Von, Qwen3Guard, JevK5 and other decision models, pullable and runnable by name |