Observability & Evaluation
23 agents tracked · Open-source-first, with limited public GitHub comparators · Ranked by trust score · Updated 2026-08-29 03:32 UTC
HVTracker independently evaluates 23 observability & evaluation using daily signals from GitHub, package registries, and security databases. The category is open-source-first, with a limited set of public GitHub-hosted proprietary or source-available comparators when they are important ecosystem reference points. Each agent is scored on activity, adoption, transparency, safety, and identity. The top-ranked observability & evaluation is MLflow with a trust score of 89.8/100 (Grade A). Other leading projects include Promptfoo and Weights & Biases Weave.
| # | Agent | Trust | Stars | Language |
|---|---|---|---|---|
| 1 | MLflow A Listed The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, eva | 89.8 | 27.7k | Python |
| 2 | Promptfoo A Listed Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, C | 88.6 | 24.7k | TypeScript |
| 3 | Weights & Biases Weave A Listed Weave is a toolkit for developing AI-powered applications, built by Weights & Biases. | 86.4 | 1.1k | Python |
| 4 | OpenLLMetry A Listed Open-source observability for your GenAI or LLM application, based on OpenTelemetry | 84.5 | 7.4k | Python |
| 5 | Langfuse B Listed πͺ’ Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integ | 78.2 | 33.9k | TypeScript |
| 6 | Arize Phoenix B Listed AI Observability & Evaluation | 76.1 | 11.2k | Python |
| 7 | LangWatch B Listed The platform for LLM evaluations and AI agent testing | 75.6 | 3.5k | TypeScript |
| 8 | Evidently B Listed Evidently is ββan open-source ML and LLM observability framework. Evaluate, test, and monitor any AI-powered system or d | 75.5 | 7.9k | Jupyter Notebook |
| 9 | Agenta B Listed Agenta is a workspace where you and your team build agents and automations. | 71.0 | 4.6k | TypeScript |
| 10 | DeepEval B Listed The LLM Evaluation Framework | 70.9 | 17.9k | Python |
| 11 | Giskard B Listed π’ Open-Source Evaluation & Testing library for LLM Agents | 70.3 | 5.8k | Python |
| 12 | Holmesgpt B Listed SRE Agent - CNCF Sandbox Project | 65.0 | 3.2k | Python |
| 13 | Helicone C Listed π§ Open source LLM observability platform. One line of code to monitor, evaluate, and experiment. YC W23 π | 63.5 | 6.1k | TypeScript |
| 14 | Traceroot C Listed TraceRoot - open-source observability and self-improving layer for AI agents. YC S25 | 61.4 | 748 | TypeScript |
| 15 | OpenLIT C Listed Open source platform for AI Engineering: OpenTelemetry-native LLM Observability, GPU Monitoring, Guardrails, Evaluations | 59.8 | 2.7k | TypeScript |
| 16 | agentacct C Listed See what your coding agents did and what it cost. Breaks each task down into work steps β tools used, files changed, tes | 58.0 | 667 | Python |
| 17 | AgentOps C Listed Python SDK for AI agent monitoring, LLM cost tracking, benchmarking, and more. Integrates with most LLMs and agent frame | 56.6 | 5.8k | Python |
| 18 | Opik D Listed Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, autom | 48.5 | 21.7k | Python |
| 19 | Ragas D Listed Supercharge Your LLM Application Evaluations π | 44.7 | 15.5k | Python |
| 20 | TruLens D Listed Evaluation and Tracking for LLM Experiments and AI Agents | 44.3 | 3.5k | Python |
| 21 | AgentBench D Listed A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24) | 37.3 | 3.7k | Python |
| 22 | Webarena D Listed Code repo for "WebArena: A Realistic Web Environment for Building Autonomous Agents" | 36.3 | 1.6k | Python |
| 23 | SimWorld D Listed SimWorld: An Open-ended Realistic Simulator for Autonomous Agents in Physical and Social Worlds | 17.3 | 759 | Python |
Grade = trust band: A β₯ 80, B β₯ 65, C β₯ 50, D < 50 β methodology