Registry › Compare › Evidently vs Langfuse

Evidently vs Langfuse

Both are Grade B and 0.1 points apart, so choose on what you weigh most. Evidently leads on provenance and supply-chain integrity; Langfuse leads on maintenance and adoption.

B Evidently 78.4

Evidently is ​​an open-source ML and LLM observability framework. Evaluate, test, and monitor any AI-powered system or data pipeline. From tabular data to Gen AI. 100+ metrics.

evidentlyai/evidently · #167 overall · #8 Observability & Evaluation · coverage A (4/5)

Choose Evidently if provenance and supply-chain integrity matter most.

  • +7.2Identity / Provenance: package provenance attested, against none
  • +3.0Safety / Integrity: package provenance attested, against none
B Langfuse 78.5

🪢 Open source agent evals & observability: Trace, evaluate, and improve LLM applications with one open platform.

langfuse/langfuse · #165 overall · #7 Observability & Evaluation · coverage A (4/5)

Choose Langfuse if maintenance and adoption matter most.

  • +6.6Maintenance: last push today, against 15d ago
  • +3.5Adoption: 5.5M weekly downloads against 199.6k
  • +3.0Transparency: OSSF Scorecard 6.9 against 3.4

Where they differ

16.2
Safety / IntegrityEvidently +3.0
13.2
18.0
Identity / ProvenanceEvidently +7.2
10.8
11.4
TransparencyLangfuse +3.0
14.4
13.4
MaintenanceLangfuse +6.6
20.0
16.4
AdoptionLangfuse +3.5
19.9
+3.0
Runtime calibrationEvidently +2.8
+0.2

Full evidence table

An independent, evidence-based trust comparison of Evidently and Langfuse, two Observability & Evaluation projects in the HVTracker registry. Scores come from public, checkable signals — supply-chain provenance, OSSF Scorecard, maintenance, and adoption — not popularity.

Full evidence

Signal Evidentlyevidentlyai/evidently Langfuselangfuse/langfuse
HVTrust score 78.4 78.5
Evidence grade B B
Coverage grade A A
Overall rank #167 #165
Rank in Observability & Evaluation #8 #7
GitHub stars 7.9k 35.1k
Last updated 15d ago today
Build provenance Yes No
OSSF Scorecard 3.4 / 10 6.9 / 10
License Apache-2.0 NOASSERTION
Downloads 200k/wk 5.5M/wk
Trust dimensions (points earned)
Safety / integrity / 25 16.2 13.2
Identity & provenance / 18 18.0 10.8
Transparency / 17 11.4 14.4
Maintenance / 20 13.4 20.0
Adoption / 20 16.4 19.9
Runtime capability surface (full matrix)
MCP server — Implemented
External providers 3 — Multi-provider (LiteLLM), OpenAI, Postgres 2 — Anthropic, OpenAI
Requires API keys No Yes
Plugin surface — plugins
Provenance drift Match Unknown
Open in the live compare tool → Evidently profile Langfuse profile More Observability & Evaluation →

How to read this: HVTrust (0–100) weighs supply-chain signals (provenance, OSSF Scorecard, signed commits, open license) alongside real-world adoption, scaled by an evidence-confidence factor. Grade bands: A ≥ 80, B ≥ 65, C ≥ 50, D < 50. Signals refresh daily. Full methodology v4.3 →