

DeepSeek · 6× · tolest 01. Juli 2026
31
Momentum
DFlash 2 is a "drafter" for speculative decoding of large language models, released by Inco AI (building on the original DFlash technique developed at Z Lab). It is not physical hardware, but an Open Source model/software method that predicts the output of a target LLM (e.g., Qwen3.8-27B, Meta Muse Glimmer) in parallel blocks, thereby increasing Inference speed without changing output quality. The models are available open source under Apache 2.0 license on Hugging Face and run on common Inference engines such as SGLang, vLLM, llama.cpp, and oMLX/MLX on Apple Silicon.
Momentum-Verloop
26.05.24.08.