Observability & Evaluation
22 agents tracked · Open-source-first, with limited public GitHub comparators · Ranked by trust score · Updated 2026-08-09 20:01 UTC
HVTracker independently evaluates 22 observability & evaluation using daily signals from GitHub, package registries, and security databases. The category is open-source-first, with a limited set of public GitHub-hosted proprietary or source-available comparators when they are important ecosystem reference points. Each agent is scored on activity, adoption, transparency, safety, and identity. The top-ranked observability & evaluation is MLflow with a trust score of 89.8/100 (Grade A). Other leading projects include Weights & Biases Weave and OpenLLMetry.
| # | Agent | Trust | Stars | Language |
|---|---|---|---|---|
| 1 | MLflow A Listed The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, eva | 89.8 | 27.4k | Python |
| 2 | Weights & Biases Weave A Listed Weave is a toolkit for developing AI-powered applications, built by Weights & Biases. | 86.4 | 1.1k | Python |
| 3 | OpenLLMetry A Listed Open-source observability for your GenAI or LLM application, based on OpenTelemetry | 84.7 | 7.4k | Python |
| 4 | Langfuse B Listed πͺ’ Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integ | 78.2 | 32.8k | TypeScript |
| 5 | Evidently B Listed Evidently is ββan open-source ML and LLM observability framework. Evaluate, test, and monitor any AI-powered system or d | 76.8 | 7.8k | Jupyter Notebook |
| 6 | LangWatch B Listed The platform for LLM evaluations and AI agent testing | 75.7 | 3.5k | TypeScript |
| 7 | Promptfoo B Listed Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, C | 75.3 | 24.1k | TypeScript |
| 8 | Arize Phoenix B Listed AI Observability & Evaluation | 72.0 | 11.0k | Python |
| 9 | DeepEval B Listed The LLM Evaluation Framework | 71.3 | 17.5k | Python |
| 10 | Agenta B Listed Agenta is a workspace where you and your team build agents and automations. | 70.8 | 4.5k | TypeScript |
| 11 | Giskard B Listed π’ Open-Source Evaluation & Testing library for LLM Agents | 69.6 | 5.7k | Python |
| 12 | Holmesgpt C Listed SRE Agent - CNCF Sandbox Project | 62.6 | 3.0k | Python |
| 13 | Helicone C Listed π§ Open source LLM observability platform. One line of code to monitor, evaluate, and experiment. YC W23 π | 62.3 | 6.0k | TypeScript |
| 14 | Traceroot C Listed TraceRoot - open-source observability and self-improving layer for AI agents. YC S25 | 61.4 | 702 | TypeScript |
| 15 | OpenLIT C Listed Open source platform for AI Engineering: OpenTelemetry-native LLM Observability, GPU Monitoring, Guardrails, Evaluations | 60.9 | 2.7k | TypeScript |
| 16 | AgentOps C Listed Python SDK for AI agent monitoring, LLM cost tracking, benchmarking, and more. Integrates with most LLMs and agent frame | 57.9 | 5.8k | Python |
| 17 | Opik D Listed Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, autom | 48.5 | 21.3k | Python |
| 18 | Ragas D Listed Supercharge Your LLM Application Evaluations π | 45.6 | 15.2k | Python |
| 19 | TruLens D Listed Evaluation and Tracking for LLM Experiments and AI Agents | 44.3 | 3.5k | Python |
| 20 | AgentBench D Listed A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24) | 37.3 | 3.7k | Python |
| 21 | Webarena D Listed Code repo for "WebArena: A Realistic Web Environment for Building Autonomous Agents" | 36.3 | 1.6k | Python |
| 22 | agentacct D Listed See what your coding agents did and what it cost. Breaks each task down into work steps β tools used, files changed, tes | 25.2 | 569 | Python |
Grade = trust band: A β₯ 80, B β₯ 65, C β₯ 50, D < 50 β methodology