LangWatch vs Latitude
LangWatch leads on trust: 83.0/100 (Grade A) against 76.9/100 (Grade B), a 6.1-point gap. LangWatch leads on supply-chain integrity and transparency, and rests on broader evidence.
Open-Source Agent Observability & Evals for AI Platform teams. Trace, test, route and govern every agent & coding assistant in the company.
Choose LangWatch if supply-chain integrity and transparency matter most.
- +5.8Safety / Integrity
- +4.2Transparency
- +1.7Adoption: 88.7k weekly downloads against 5.4k
- A vs BEvidence coverage: 4 of 5 independent signal types, against 3
Open-source observability for AI agents. Find where your agents fail, dispatch your coding agent to fix it, and verify the fix against real traces.
Latitude doesn't lead on any scored dimension in this pair.
Where they differ
2 dimensions identical: Identity 18.0 · Maintenance 19.9 · Full evidence table
An independent, evidence-based trust comparison of LangWatch and Latitude, two Observability & Evaluation projects in the HVTracker registry. Scores come from public, checkable signals — supply-chain provenance, OSSF Scorecard, maintenance, and adoption — not popularity.
Full evidence
| Signal | LangWatchlangwatch/langwatch | Latitudelatitude-dev/latitude-llm |
|---|---|---|
| HVTrust score | 83.0 | 76.9 |
| Evidence grade | A | B |
| Coverage grade | A | B |
| Overall rank | #111 | #210 |
| Rank in Observability & Evaluation | #5 | #8 |
| GitHub stars | 4.9k | 4.7k |
| Last updated | 1d ago | 2d ago |
| Build provenance | Yes | Yes |
| OSSF Scorecard | 4.9 / 10 | — |
| License | Apache-2.0 | MIT |
| Downloads | 89k/wk | 5k/wk |
| Trust dimensions (points earned) | ||
| Safety / integrity / 25 | 17.8 | 12.0 |
| Identity & provenance / 18 | 18.0 | 18.0 |
| Transparency / 17 | 12.7 | 8.5 |
| Maintenance / 20 | 19.9 | 19.9 |
| Adoption / 20 | 15.5 | 13.8 |
| Runtime capability surface (full matrix) | ||
| MCP server | Implemented | Implemented |
| External providers | 6 — Amazon Bedrock, Anthropic, ElevenLabs, … | — |
| Requires API keys | Yes | Yes |
| Plugin surface | plugins | plugins |
| Provenance drift | Partial | Match |
How to read this: HVTrust (0–100) weighs supply-chain signals (provenance, OSSF Scorecard, signed commits, open license) alongside real-world adoption, scaled by an evidence-confidence factor. Grade bands: A ≥ 80, B ≥ 65, C ≥ 50, D < 50. Signals refresh daily. Full methodology v4.4 →