Synthszr Charts — die großen AI-Marken im Wettkampf ums Podium
synthszr charts
baai

Baai · siet Januar/Februar 2024 (GitHub-Release 30.01.2024; arXiv-Paper veröffentlicht 5. Februar 2024) · 2× · tolest 30. Juni 2026

3
Momentum

BGE-M3 is an open-source embedding model by the Beijing Academy of Artificial Intelligence (BAAI), based on a fine-tuned XLM-RoBERTa architecture that encodes text into 1024-dimensional vectors. It is distinguished by three properties: support for more than 100 languages (Multi-Linguality), simultaneous dense, sparse, and multi-vector retrieval (Multi-Functionality), and processing of inputs from short sentences up to 8192 tokens (Multi-Granularity). The model is freely available via Hugging Face, GitHub (FlagEmbedding) and various cloud platforms such as NVIDIA NIM or Ollama, and is distributed under the MIT license.

Momentum-Verloop
19.05.17.08.

Features

Deployment (Self-host/Cloud)Self-hosting via Python-Bibliothek (FlagEmbedding, Transformers, ONNX) oder Cloud-Inferenz via NVIDIA NIM/Ollama
Durchsatz/LatenzKeine offiziellen Benchmark-Zahlen zu Durchsatz/Latenz in den geprüften Quellen gefunden
LizenzMIT-Lizenz, frei für akademische und kommerzielle Nutzung
PlattformHugging Face, GitHub (FlagEmbedding), NVIDIA NIM, Ollama, ONNX
PreisKostenlos, Open-Source-Gewichte (Download via Hugging Face/GitHub)
Protokoll-KompatibilitätMax. 8192 Tokens Kontextlänge; unterstützt Dense-, Sparse- (BM25-ähnlich) und Multi-Vektor-Retrieval (ColBERT-Stil); Integration mit Milvus und Vespa für Hybrid-Retrieval
Release-Datum30. Januar 2024 (GitHub-Ankündigung); Paper veröffentlicht 5. Februar 2024
Unterstützte Modelle/ProviderBasis: XLM-RoBERTa (fine-tuned), 568 Mio. Parameter, Embedding-Dimension 1024; nutzbar via FlagEmbedding, sentence-transformers, HuggingFace transformers, LangChain, Milvus, Vespa

Mehr Produkten in disse Kategorie: Embeddings & Vektor-DBs

Belege (2)

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.