

DeepEval
#3 in Observability & EvalsDeepeval · siet Erste PyPI-Veröffentlichung: 15. August 2023 (als "deepeval") · 3× · tolest 05. Okt. 2026
DeepEval is an open-source framework by Confident AI for evaluating and testing LLM applications (RAG, chatbots, AI agents, voice). It ships 50+ ready-to-use, research-backed metrics (e.g. hallucination, faithfulness, answer relevancy, G-Eval) and integrates Pytest-style into CI/CD pipelines. DeepEval runs locally/self-hosted and can optionally connect to the commercial Confident AI cloud platform for dashboards, tracing, and team collaboration. As an LLM judge it supports virtually any model provider, including OpenAI, Anthropic, Gemini, Azure OpenAI, Ollama, and 100+ more models via LiteLLM.
Features
| License | Open source, Apache License 2.0 |
| Platform | Python library (pip install deepeval); TypeScript SDK recently in beta |
| Price | DeepEval (open source): free. Confident AI cloud: free tier, Starter from $9.99–$19.99/user/month |
| Protocol Compatibility | Native OpenTelemetry/OTLP support for tracing; metrics for AI agent MCP interactions |
| Release Date | First released on PyPI on August 15, 2023 |
| Supported Models/Providers | Any LLM judge: OpenAI (default), Anthropic, Gemini, Azure OpenAI, Ollama, Amazon Bedrock, Vertex AI, Grok, plus 100+ models via LiteLLM |