Langfuse vs Weights & Biases Weave
Weights & Biases Weave leads on trust: 84.8/100 (Grade A) against 78.4/100 (Grade B), a 6.4-point gap. Langfuse leads on adoption and transparency; Weights & Biases Weave leads on provenance and supply-chain integrity.
🪢 Open source agent evals & observability: Trace, evaluate, and improve LLM applications with one open platform.
Choose Langfuse if adoption and transparency matter most.
- +5.5Adoption: 5.6M weekly downloads against 184.4k
- +0.6Transparency: OSSF Scorecard 6.9 against 6.2
- +0.6Maintenance
Weave is a toolkit for developing AI-powered applications, built by Weights & Biases.
Choose Weights & Biases Weave if provenance and supply-chain integrity matter most.
- +7.2Identity / Provenance: package provenance attested, against none
- +6.4Safety / Integrity: package provenance attested, against none
Where they differ
An independent, evidence-based trust comparison of Langfuse and Weights & Biases Weave, two Observability & Evaluation projects in the HVTracker registry. Scores come from public, checkable signals — supply-chain provenance, OSSF Scorecard, maintenance, and adoption — not popularity.
Full evidence
| Signal | Langfuselangfuse/langfuse | Weights & Biases Weavewandb/weave |
|---|---|---|
| HVTrust score | 78.4 | 84.8 |
| Evidence grade | B | A |
| Coverage grade | A | A |
| Overall rank | #167 | #88 |
| Rank in Observability & Evaluation | #6 | #3 |
| GitHub stars | 35.2k | 1.1k |
| Last updated | 1d ago | 1d ago |
| Build provenance | No | Yes |
| OSSF Scorecard | 6.9 / 10 | 6.2 / 10 |
| License | NOASSERTION | Apache-2.0 |
| Downloads | 5.6M/wk | 184k/wk |
| Trust dimensions (points earned) | ||
| Safety / integrity / 25 | 13.2 | 19.6 |
| Identity & provenance / 18 | 10.8 | 18.0 |
| Transparency / 17 | 14.4 | 13.8 |
| Maintenance / 20 | 19.9 | 19.3 |
| Adoption / 20 | 19.9 | 14.4 |
| Runtime capability surface (full matrix) | ||
| MCP server | Implemented | Implemented |
| External providers | 2 — Anthropic, OpenAI | 9 — Amazon Bedrock, Anthropic, Cohere, … |
| Requires API keys | Yes | No |
| Plugin surface | plugins | — |
| Provenance drift | Unknown | Partial |
How to read this: HVTrust (0–100) weighs supply-chain signals (provenance, OSSF Scorecard, signed commits, open license) alongside real-world adoption, scaled by an evidence-confidence factor. Grade bands: A ≥ 80, B ≥ 65, C ≥ 50, D < 50. Signals refresh daily. Full methodology v4.3 →