

Mistral OCR
#7 en Transcription (STT)Mistral AI · depuis 6. März 2025 (Original Mistral OCR); aktuelle Version OCR 4: 23. Juni 2026 · 2× · vu le 14 août 2026
Mistral OCR is a specialized document processing model by Mistral AI that extracts text, tables, formulas, and embedded images from PDFs and images, converting them into structured Markdown/HTML. Originally launched on March 6, 2025, it has since evolved through several versions (OCR 3 in December 2025, OCR 4 in June 2026), with the current OCR 4 version offering bounding boxes, block classification, confidence scores, and support for 170 languages. The model is available via the Mistral API, Le Chat/Vibe, Document AI, and cloud partners such as Microsoft Foundry and AWS SageMaker; it is proprietary (closed-weight) but can also be self-hosted for enterprise customers.
Fonctionnalités
| License | Proprietary, closed-weight; API access via la Plateforme; self-hosting available for enterprise customers on request |
| Platform | Available via Mistral API (la Plateforme), Le Chat/Vibe, Document AI/Mistral Studio, Microsoft Foundry, Amazon SageMaker; Snowflake Parse Document coming soon |
| Price | OCR 4: $4 per 1,000 pages (API), $2 with Batch API discount (50%), Document AI $5 per 1,000 pages |
| Release Date | Original: March 6, 2025; OCR 3: December 2025; OCR 4 (current): June 23, 2026 |
| Languages | 170 languages across 10 language groups (OCR 4); original 2025 version supported 11 languages |