Synthszr Charts — die großen AI-Marken im Wettkampf ums Podium
synthszr charts
superlinked

SIE (Superlinked Inference Engine)

#4 v LLM inference a serving

Superlinked · 2× · naposledy 10. 9. 2026

41
Momentum

SIE (Superlinked Inference Engine) is an Open Source Inference server by Superlinked that provides over 100 small AI models (Embeddings, Reranking, Sparse-Retrieval, Extraction, OCR, and Text Generation) through a unified API on shared GPU resources. Instead of operating a separate server for each model type, SIE loads models on-demand and manages them via LRU-Eviction in a single cluster. The core product is released under Apache-2.0 license for Self-Hosting (from laptop to Kubernetes), and Superlinked additionally offers a managed Cloud version of the same Engine. The API is partially OpenAI-compatible, enabling simple migration.

Vývoj momenta
15.06.13.09.

Vlastnosti

Throughput/LatencySelf-hosted embedding returns results 3–7x faster than hosted frontier APIs; p99 latency close to p50; batch-then-route reaches ~89% GPU efficiency vs ~51% for route-then-batch (~1.8x throughput per GPU)
LicenseOpen source, Apache 2.0
PlatformRuns on own hardware: macOS (Apple Silicon), Linux (CPU/NVIDIA GPU), Kubernetes (AWS EKS, GCP GKE, Azure AKS)
PriceCore engine free (open source, self-hosted); managed SIE Cloud on request/contact; free hosted capacity for selected projects via inference grant
Protocol CompatibilityOpenAI-compatible API endpoints: /v1/embeddings, /v1/chat/completions, /v1/completions, /v1/responses
Release DateLaunch blog post dated February 2026
Supported Models/Providers100+ pre-configured models (incl. Stella, SPLADE, Qwen3, GLiNER, SigLIP, BGE-M3, ColQwen2.5, BGE-reranker, Florence-2) covering embeddings, reranking, extraction, OCR and generation

Další produkty v této kategorii: LLM inference a serving

Zdroje (2)

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.