Observability & Evaluation

23 agents tracked · Open-source-first, with limited public GitHub comparators · Ranked by trust score · Updated 2026-08-29 03:32 UTC

HVTracker independently evaluates 23 observability & evaluation using daily signals from GitHub, package registries, and security databases. The category is open-source-first, with a limited set of public GitHub-hosted proprietary or source-available comparators when they are important ecosystem reference points. Each agent is scored on activity, adoption, transparency, safety, and identity. The top-ranked observability & evaluation is MLflow with a trust score of 89.8/100 (Grade A). Other leading projects include Promptfoo and Weights & Biases Weave.

23
Agents
63
Avg Trust
211.7k
Total Stars
4
Grade A
# Agent Trust Stars Language
1 MLflow A Listed The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, eva 89.8 27.7k Python
2 Promptfoo A Listed Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, C 88.6 24.7k TypeScript
3 Weights & Biases Weave A Listed Weave is a toolkit for developing AI-powered applications, built by Weights & Biases. 86.4 1.1k Python
4 OpenLLMetry A Listed Open-source observability for your GenAI or LLM application, based on OpenTelemetry 84.5 7.4k Python
5 Langfuse B Listed πŸͺ’ Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integ 78.2 33.9k TypeScript
6 Arize Phoenix B Listed AI Observability & Evaluation 76.1 11.2k Python
7 LangWatch B Listed The platform for LLM evaluations and AI agent testing 75.6 3.5k TypeScript
8 Evidently B Listed Evidently is ​​an open-source ML and LLM observability framework. Evaluate, test, and monitor any AI-powered system or d 75.5 7.9k Jupyter Notebook
9 Agenta B Listed Agenta is a workspace where you and your team build agents and automations. 71.0 4.6k TypeScript
10 DeepEval B Listed The LLM Evaluation Framework 70.9 17.9k Python
11 Giskard B Listed 🐒 Open-Source Evaluation & Testing library for LLM Agents 70.3 5.8k Python
12 Holmesgpt B Listed SRE Agent - CNCF Sandbox Project 65.0 3.2k Python
13 Helicone C Listed 🧊 Open source LLM observability platform. One line of code to monitor, evaluate, and experiment. YC W23 πŸ“ 63.5 6.1k TypeScript
14 Traceroot C Listed TraceRoot - open-source observability and self-improving layer for AI agents. YC S25 61.4 748 TypeScript
15 OpenLIT C Listed Open source platform for AI Engineering: OpenTelemetry-native LLM Observability, GPU Monitoring, Guardrails, Evaluations 59.8 2.7k TypeScript
16 agentacct C Listed See what your coding agents did and what it cost. Breaks each task down into work steps β€” tools used, files changed, tes 58.0 667 Python
17 AgentOps C Listed Python SDK for AI agent monitoring, LLM cost tracking, benchmarking, and more. Integrates with most LLMs and agent frame 56.6 5.8k Python
18 Opik D Listed Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, autom 48.5 21.7k Python
19 Ragas D Listed Supercharge Your LLM Application Evaluations πŸš€ 44.7 15.5k Python
20 TruLens D Listed Evaluation and Tracking for LLM Experiments and AI Agents 44.3 3.5k Python
21 AgentBench D Listed A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24) 37.3 3.7k Python
22 Webarena D Listed Code repo for "WebArena: A Realistic Web Environment for Building Autonomous Agents" 36.3 1.6k Python
23 SimWorld D Listed SimWorld: An Open-ended Realistic Simulator for Autonomous Agents in Physical and Social Worlds 17.3 759 Python

Grade = trust band: A β‰₯ 80, B β‰₯ 65, C β‰₯ 50, D < 50 β€” methodology