Synthszr Charts — die großen AI-Marken im Wettkampf ums Podium
synthszr charts

Unknown · seit 17. August 2026 (arXiv-Preprint und Open-Source-Release) · 2× · zuletzt 26. Aug. 2026

43
Momentum

FreeToken ist eine Open-Source-Inferenz-Engine von FlashML (mit Forschern der UC Berkeley, UT Austin und MIT) für das lokale Ausführen großer Mixture-of-Experts-Modelle auf Consumer-Hardware. Die Software verteilt Modellgewichte dynamisch zwischen GPU-VRAM, System-RAM und CPU, sodass Modelle wie Qwen3.6-35B, DeepSeek-V4-Flash 284B oder GLM-5.2 753B auch auf einer einzelnen Consumer- oder Workstation-GPU mit begrenztem VRAM lauffähig sind. Sie bietet OpenAI- und Anthropic-kompatible API-Endpunkte für die Integration mit Coding-Agents wie Claude Code, Codex oder OpenCode. Der Quellcode ist unter Apache-2.0-Lizenz auf GitHub verfügbar, zusätzlich existiert eine Desktop-App für Windows und Linux sowie ein PyPI-Paket.

Momentum-Verlauf
01.06.30.08.

Features

Durchsatz/LatenzQwen3.6-35B: 39,3 Tok/s auf 8GB RTX 4060 Laptop; DeepSeek-V4-Flash 284B: 22-25 Tok/s auf RTX 5090; GLM-5.2 753B: 14,9 Tok/s auf RTX PRO 6000; 2-4x schneller als Ollama
LizenzApache License 2.0
PlattformLinux x86_64 (CLI) und Windows/Linux Desktop-App; NVIDIA GPU (RTX 30/40/50-Serie), CUDA 13, Treiber r580+
PreisKostenlos/Open Source (nur eigene Hardwarekosten)
Protokoll-KompatibilitätOpenAI-kompatible API (/v1/chat/completions, /v1/responses, /v1/models) und Anthropic-kompatible API (/v1/messages)
Release-Datum17. August 2026 (arXiv-Preprint, Open-Source-Release)
Unterstützte Modelle/ProviderÜber 20 MoE-Modelle, u.a. DeepSeek-V4-Flash, Qwen3.6-35B-A3B, GLM-5.2; Formate MXFP4, NVFP4, FP8, BF16

Weitere Produkte in dieser Kategorie: Lokale LLM-Runtimes

Belege (2)

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.