Registry › Compare › Langfuse vs MLflow

Langfuse vs MLflow

MLflow leads on trust: 88.6/100 (Grade A) against 78.5/100 (Grade B), a 10.1-point gap. Langfuse leads on transparency; MLflow leads on provenance and supply-chain integrity.

B Langfuse 78.5

🪢 Open source agent evals & observability: Trace, evaluate, and improve LLM applications with one open platform.

langfuse/langfuse · #167 overall · #6 Observability & Evaluation · coverage A (4/5)

Choose Langfuse if transparency matters most.

  • +1.2Transparency: OSSF Scorecard 6.9 against 5.5
A MLflow 88.6

The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models and data.

mlflow/mlflow · #33 overall · #2 Observability & Evaluation · coverage A (4/5)

Choose MLflow if provenance and supply-chain integrity matter most.

  • +7.2Identity / Provenance: package provenance attested, against none
  • +5.1Safety / Integrity: package provenance attested, against none

Where they differ

13.2
Safety / IntegrityMLflow +5.1
18.3
10.8
Identity / ProvenanceMLflow +7.2
18.0
14.4
TransparencyLangfuse +1.2
13.2
19.9
AdoptionLangfuse +0.3
19.6
+0.2
Runtime calibrationLangfuse +0.7
-0.5

1 dimension identical: Maintenance 20.0 · Full evidence table

An independent, evidence-based trust comparison of Langfuse and MLflow, two Observability & Evaluation projects in the HVTracker registry. Scores come from public, checkable signals — supply-chain provenance, OSSF Scorecard, maintenance, and adoption — not popularity.

Full evidence

Signal Langfuselangfuse/langfuse MLflowmlflow/mlflow
HVTrust score 78.5 88.6
Evidence grade B A
Coverage grade A A
Overall rank #167 #33
Rank in Observability & Evaluation #6 #2
GitHub stars 35.2k 28.2k
Last updated today today
Build provenance No Yes
OSSF Scorecard 6.9 / 10 5.5 / 10
License NOASSERTION Apache-2.0
Downloads 5.7M/wk 4.5M/wk
Trust dimensions (points earned)
Safety / integrity / 25 13.2 18.3
Identity & provenance / 18 10.8 18.0
Transparency / 17 14.4 13.2
Maintenance / 20 20.0 20.0
Adoption / 20 19.9 19.6
Runtime capability surface (full matrix)
MCP server Implemented Implemented
External providers 2 — Anthropic, OpenAI 3 — Amazon Bedrock, Anthropic, Postgres
Requires API keys Yes No
Plugin surface plugins plugins
Provenance drift Unknown Unknown
Open in the live compare tool → Langfuse profile MLflow profile More Observability & Evaluation →

How to read this: HVTrust (0–100) weighs supply-chain signals (provenance, OSSF Scorecard, signed commits, open license) alongside real-world adoption, scaled by an evidence-confidence factor. Grade bands: A ≥ 80, B ≥ 65, C ≥ 50, D < 50. Signals refresh daily. Full methodology v4.3 →