RegistryCompare › Opik vs Promptfoo

Opik vs Promptfoo

An independent, evidence-based trust comparison of Opik and Promptfoo, two Observability & Evaluation projects in the HVTracker registry. Scores come from public, checkable signals — supply-chain provenance, OSSF Scorecard, maintenance, and adoption — not popularity.

Promptfoo leads on trust — 88.6/100 (Grade A) vs 78.4/100 (Grade B), a 10.2-point gap. Full breakdown below.
Signal Opikcomet-ml/opik Promptfoopromptfoo/promptfoo
HVTrust score 78.4 88.6
Evidence grade B A
Coverage grade A A
Overall rank #103 #20
Rank in Observability & Evaluation #6 #2
GitHub stars 21.8k 24.8k
Last updated today today
Build provenance Yes Yes
OSSF Scorecard 7.4 / 10
License Apache-2.0 MIT
Downloads 428k/wk 384k/wk
Trust dimensions (points earned)
Safety / integrity / 25 12.1 16.8
Identity & provenance / 18 18.0 18.0
Transparency / 17 8.5 14.8
Maintenance / 20 20.0 20.0
Adoption / 20 17.9 18.0
Runtime capability surface (full matrix)
MCP server Implemented Implemented
External providers 4 — Amazon Bedrock, Anthropic, Multi-provider (LiteLLM), … 2 — Anthropic, OpenAI
Requires API keys No Yes
Plugin surface extensions plugins
Provenance drift Partial Match
Open in the live compare tool → Opik profile Promptfoo profile More Observability & Evaluation →

How to read this: HVTrust (0–100) weighs supply-chain signals (provenance, OSSF Scorecard, signed commits, open license) alongside real-world adoption, scaled by an evidence-confidence factor. Grade bands: A ≥ 80, B ≥ 65, C ≥ 50, D < 50. Signals refresh daily. Full methodology v4.3 →