Registry › Compare › Langfuse vs Weights & Biases Weave

Langfuse vs Weights & Biases Weave

Weights & Biases Weave leads on trust: 84.8/100 (Grade A) against 78.4/100 (Grade B), a 6.4-point gap. Langfuse leads on adoption and transparency; Weights & Biases Weave leads on provenance and supply-chain integrity.

B Langfuse 78.4

🪢 Open source agent evals & observability: Trace, evaluate, and improve LLM applications with one open platform.

langfuse/langfuse · #167 overall · #6 Observability & Evaluation · coverage A (4/5)

Choose Langfuse if adoption and transparency matter most.

  • +5.5Adoption: 5.6M weekly downloads against 184.4k
  • +0.6Transparency: OSSF Scorecard 6.9 against 6.2
  • +0.6Maintenance

Weave is a toolkit for developing AI-powered applications, built by Weights & Biases.

wandb/weave · #88 overall · #3 Observability & Evaluation · coverage A (4/5)

Choose Weights & Biases Weave if provenance and supply-chain integrity matter most.

  • +7.2Identity / Provenance: package provenance attested, against none
  • +6.4Safety / Integrity: package provenance attested, against none

Where they differ

13.2
Safety / IntegrityWeights & Biases Weave +6.4
19.6
10.8
Identity / ProvenanceWeights & Biases Weave +7.2
18.0
14.4
TransparencyLangfuse +0.6
13.8
19.9
MaintenanceLangfuse +0.6
19.3
19.9
AdoptionLangfuse +5.5
14.4
+0.2
Runtime calibrationLangfuse +0.5
-0.3

Full evidence table

An independent, evidence-based trust comparison of Langfuse and Weights & Biases Weave, two Observability & Evaluation projects in the HVTracker registry. Scores come from public, checkable signals — supply-chain provenance, OSSF Scorecard, maintenance, and adoption — not popularity.

Full evidence

Signal Langfuselangfuse/langfuse Weights & Biases Weavewandb/weave
HVTrust score 78.4 84.8
Evidence grade B A
Coverage grade A A
Overall rank #167 #88
Rank in Observability & Evaluation #6 #3
GitHub stars 35.2k 1.1k
Last updated 1d ago 1d ago
Build provenance No Yes
OSSF Scorecard 6.9 / 10 6.2 / 10
License NOASSERTION Apache-2.0
Downloads 5.6M/wk 184k/wk
Trust dimensions (points earned)
Safety / integrity / 25 13.2 19.6
Identity & provenance / 18 10.8 18.0
Transparency / 17 14.4 13.8
Maintenance / 20 19.9 19.3
Adoption / 20 19.9 14.4
Runtime capability surface (full matrix)
MCP server Implemented Implemented
External providers 2 — Anthropic, OpenAI 9 — Amazon Bedrock, Anthropic, Cohere, …
Requires API keys Yes No
Plugin surface plugins —
Provenance drift Unknown Partial
Open in the live compare tool → Langfuse profile Weights & Biases Weave profile More Observability & Evaluation →

How to read this: HVTrust (0–100) weighs supply-chain signals (provenance, OSSF Scorecard, signed commits, open license) alongside real-world adoption, scaled by an evidence-confidence factor. Grade bands: A ≥ 80, B ≥ 65, C ≥ 50, D < 50. Signals refresh daily. Full methodology v4.3 →