Synthszr Charts — die großen AI-Marken im Wettkampf ums Podium
synthszr charts
baai

Baai · od Januar/Februar 2024 (GitHub-Release 30.01.2024; arXiv-Paper veröffentlicht 5. Februar 2024) · 2× · naposledy 30. 6. 2026

3
Momentum

BGE-M3 is an open-source embedding model by the Beijing Academy of Artificial Intelligence (BAAI), based on a fine-tuned XLM-RoBERTa architecture that encodes text into 1024-dimensional vectors. It is distinguished by three properties: support for more than 100 languages (Multi-Linguality), simultaneous dense, sparse, and multi-vector retrieval (Multi-Functionality), and processing of inputs from short sentences up to 8192 tokens (Multi-Granularity). The model is freely available via Hugging Face, GitHub (FlagEmbedding) and various cloud platforms such as NVIDIA NIM or Ollama, and is distributed under the MIT license.

Vývoj momenta
19.05.17.08.

Vlastnosti

Deployment (Self-host/Cloud)Self-hosting via Python-Bibliothek (FlagEmbedding, Transformers, ONNX) oder Cloud-Inferenz via NVIDIA NIM/Ollama
Durchsatz/LatenzKeine offiziellen Benchmark-Zahlen zu Durchsatz/Latenz in den geprüften Quellen gefunden
LizenzMIT-Lizenz, frei für akademische und kommerzielle Nutzung
PlattformHugging Face, GitHub (FlagEmbedding), NVIDIA NIM, Ollama, ONNX
PreisKostenlos, Open-Source-Gewichte (Download via Hugging Face/GitHub)
Protokoll-KompatibilitätMax. 8192 Tokens Kontextlänge; unterstützt Dense-, Sparse- (BM25-ähnlich) und Multi-Vektor-Retrieval (ColBERT-Stil); Integration mit Milvus und Vespa für Hybrid-Retrieval
Release-Datum30. Januar 2024 (GitHub-Ankündigung); Paper veröffentlicht 5. Februar 2024
Unterstützte Modelle/ProviderBasis: XLM-RoBERTa (fine-tuned), 568 Mio. Parameter, Embedding-Dimension 1024; nutzbar via FlagEmbedding, sentence-transformers, HuggingFace transformers, LangChain, Milvus, Vespa

Další produkty v této kategorii: Embeddings a vektorové DB

Zdroje (2)

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.