Observability & Evaluation

22 agents tracked · Open-source-first, with limited public GitHub comparators · Ranked by trust score · Updated 2026-08-09 20:01 UTC

HVTracker independently evaluates 22 observability & evaluation using daily signals from GitHub, package registries, and security databases. The category is open-source-first, with a limited set of public GitHub-hosted proprietary or source-available comparators when they are important ecosystem reference points. Each agent is scored on activity, adoption, transparency, safety, and identity. The top-ranked observability & evaluation is MLflow with a trust score of 89.8/100 (Grade A). Other leading projects include Weights & Biases Weave and OpenLLMetry.

22
Agents
63
Avg Trust
206.7k
Total Stars
3
Grade A
# Agent Trust Stars Language
1 MLflow A Listed The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, eva 89.8 27.4k Python
2 Weights & Biases Weave A Listed Weave is a toolkit for developing AI-powered applications, built by Weights & Biases. 86.4 1.1k Python
3 OpenLLMetry A Listed Open-source observability for your GenAI or LLM application, based on OpenTelemetry 84.7 7.4k Python
4 Langfuse B Listed πŸͺ’ Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integ 78.2 32.8k TypeScript
5 Evidently B Listed Evidently is ​​an open-source ML and LLM observability framework. Evaluate, test, and monitor any AI-powered system or d 76.8 7.8k Jupyter Notebook
6 LangWatch B Listed The platform for LLM evaluations and AI agent testing 75.7 3.5k TypeScript
7 Promptfoo B Listed Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, C 75.3 24.1k TypeScript
8 Arize Phoenix B Listed AI Observability & Evaluation 72.0 11.0k Python
9 DeepEval B Listed The LLM Evaluation Framework 71.3 17.5k Python
10 Agenta B Listed Agenta is a workspace where you and your team build agents and automations. 70.8 4.5k TypeScript
11 Giskard B Listed 🐒 Open-Source Evaluation & Testing library for LLM Agents 69.6 5.7k Python
12 Holmesgpt C Listed SRE Agent - CNCF Sandbox Project 62.6 3.0k Python
13 Helicone C Listed 🧊 Open source LLM observability platform. One line of code to monitor, evaluate, and experiment. YC W23 πŸ“ 62.3 6.0k TypeScript
14 Traceroot C Listed TraceRoot - open-source observability and self-improving layer for AI agents. YC S25 61.4 702 TypeScript
15 OpenLIT C Listed Open source platform for AI Engineering: OpenTelemetry-native LLM Observability, GPU Monitoring, Guardrails, Evaluations 60.9 2.7k TypeScript
16 AgentOps C Listed Python SDK for AI agent monitoring, LLM cost tracking, benchmarking, and more. Integrates with most LLMs and agent frame 57.9 5.8k Python
17 Opik D Listed Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, autom 48.5 21.3k Python
18 Ragas D Listed Supercharge Your LLM Application Evaluations πŸš€ 45.6 15.2k Python
19 TruLens D Listed Evaluation and Tracking for LLM Experiments and AI Agents 44.3 3.5k Python
20 AgentBench D Listed A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24) 37.3 3.7k Python
21 Webarena D Listed Code repo for "WebArena: A Realistic Web Environment for Building Autonomous Agents" 36.3 1.6k Python
22 agentacct D Listed See what your coding agents did and what it cost. Breaks each task down into work steps β€” tools used, files changed, tes 25.2 569 Python

Grade = trust band: A β‰₯ 80, B β‰₯ 65, C β‰₯ 50, D < 50 β€” methodology