Arize Phoenix vs LangWatch
An independent, evidence-based trust comparison of Arize Phoenix and LangWatch, two Observability & Evaluation projects in the HVTracker registry. Scores come from public, checkable signals — supply-chain provenance, OSSF Scorecard, maintenance, and adoption — not popularity.
Arize Phoenix leads on trust — 76.2/100 (Grade B) vs 72.7/100 (Grade B), a 3.5-point gap. Full breakdown below.
| Signal | Arize PhoenixArize-AI/phoenix | LangWatchlangwatch/langwatch |
|---|---|---|
| HVTrust score | 76.2 | 72.7 |
| Evidence grade | B | B |
| Coverage grade | A | A |
| Overall rank | #70 | #91 |
| Rank in Observability & Evaluation | #4 | #6 |
| GitHub stars | 10.9k | 3.5k |
| Last updated | today | today |
| Build provenance | Yes | Yes |
| OSSF Scorecard | — | — |
| License | NOASSERTION | Apache-2.0 |
| Downloads | 481k/wk | 88k/wk |
| Trust dimensions (points earned) | ||
| Safety / integrity / 25 | 12.3 | 11.9 |
| Identity & provenance / 18 | 18.0 | 18.0 |
| Transparency / 17 | 8.5 | 8.5 |
| Maintenance / 20 | 20.0 | 20.0 |
| Adoption / 20 | 17.3 | 15.1 |
| Runtime capability surface (full matrix) | ||
| MCP server | Implemented | Implemented |
| External providers | 8 — Amazon Bedrock, Anthropic, E2B, … | 5 — Amazon Bedrock, Anthropic, OpenAI, … |
| Requires API keys | No | Yes |
| Plugin surface | plugins | extensions |
| Provenance drift | Partial | Partial |
Open in the live compare tool →
Arize Phoenix profile
LangWatch profile
More Observability & Evaluation →
How to read this: HVTrust (0–100) weighs supply-chain signals (provenance, OSSF Scorecard, signed commits, open license) alongside real-world adoption, scaled by an evidence-confidence factor. Grade bands: A ≥ 80, B ≥ 65, C ≥ 50, D < 50. Signals refresh daily. Full methodology v4.2 →