RegistryCompare › Arize Phoenix vs Langfuse

Arize Phoenix vs Langfuse

An independent, evidence-based trust comparison of Arize Phoenix and Langfuse, two Observability & Evaluation projects in the HVTracker registry. Scores come from public, checkable signals — supply-chain provenance, OSSF Scorecard, maintenance, and adoption — not popularity.

Langfuse leads on trust — 78.8/100 (Grade B) vs 74.7/100 (Grade B), a 4.1-point gap. Full breakdown below.
Signal Arize PhoenixArize-AI/phoenix Langfuselangfuse/langfuse
HVTrust score 74.7 78.8
Evidence grade B B
Coverage grade A A
Overall rank #91 #75
Rank in Observability & Evaluation #9 #7
GitHub stars 11.6k 34.9k
Last updated today today
Build provenance Yes No
OSSF Scorecard 6.9 / 10
License NOASSERTION NOASSERTION
Downloads 155k/wk 5.3M/wk
Trust dimensions (points earned)
Safety / integrity / 25 10.3 13.5
Identity & provenance / 18 18.0 10.8
Transparency / 17 8.5 14.4
Maintenance / 20 20.0 20.0
Adoption / 20 16.7 19.9
Runtime capability surface (full matrix)
MCP server Implemented Implemented
External providers 6 — Amazon Bedrock, Anthropic, E2B, … 2 — Anthropic, OpenAI
Requires API keys No Yes
Plugin surface plugins plugins
Provenance drift Partial Unknown
Open in the live compare tool → Arize Phoenix profile Langfuse profile More Observability & Evaluation →

How to read this: HVTrust (0–100) weighs supply-chain signals (provenance, OSSF Scorecard, signed commits, open license) alongside real-world adoption, scaled by an evidence-confidence factor. Grade bands: A ≥ 80, B ≥ 65, C ≥ 50, D < 50. Signals refresh daily. Full methodology v4.3 →