Registry › Compare › Promptfoo vs Weights & Biases Weave

Promptfoo vs Weights & Biases Weave

Promptfoo leads on trust: 92.2/100 (Grade A) against 84.9/100 (Grade A), a 7.3-point gap. Promptfoo leads on adoption and supply-chain integrity.

A Promptfoo 92.2

Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.

promptfoo/promptfoo · #8 overall · #1 Observability & Evaluation · coverage A (4/5)

Choose Promptfoo if adoption and supply-chain integrity matter most.

  • +4.1Adoption: 905.4k weekly downloads against 212.8k
  • +2.1Safety / Integrity: 100% of recent commits signed, against 87%
  • +1.0Transparency: OSSF Scorecard 7.4 against 6.2
  • +0.5Maintenance

Weave is a toolkit for developing AI-powered applications, built by Weights & Biases.

wandb/weave · #87 overall · #3 Observability & Evaluation · coverage A (4/5)

Weights & Biases Weave doesn't lead on any scored dimension in this pair.

Where they differ

21.7
Safety / IntegrityPromptfoo +2.1
19.6
14.8
TransparencyPromptfoo +1.0
13.8
20.0
MaintenancePromptfoo +0.5
19.5
18.5
AdoptionPromptfoo +4.1
14.4
-0.8
Runtime calibrationWeights & Biases Weave +0.4
-0.4

1 dimension identical: Identity 18.0 · Full evidence table

An independent, evidence-based trust comparison of Promptfoo and Weights & Biases Weave, two Observability & Evaluation projects in the HVTracker registry. Scores come from public, checkable signals — supply-chain provenance, OSSF Scorecard, maintenance, and adoption — not popularity.

Full evidence

Signal Promptfoopromptfoo/promptfoo Weights & Biases Weavewandb/weave
HVTrust score 92.2 84.9
Evidence grade A A
Coverage grade A A
Overall rank #8 #87
Rank in Observability & Evaluation #1 #3
GitHub stars 25.7k 1.1k
Last updated today today
Build provenance Yes Yes
OSSF Scorecard 7.4 / 10 6.2 / 10
License MIT Apache-2.0
Downloads 905k/wk 213k/wk
Trust dimensions (points earned)
Safety / integrity / 25 21.7 19.6
Identity & provenance / 18 18.0 18.0
Transparency / 17 14.8 13.8
Maintenance / 20 20.0 19.5
Adoption / 20 18.5 14.4
Runtime capability surface (full matrix)
MCP server Implemented Implemented
External providers 3 — Amazon Bedrock, Anthropic, OpenAI 9 — Amazon Bedrock, Anthropic, Cohere, …
Requires API keys Yes No
Plugin surface plugins —
Provenance drift Match Partial
Open in the live compare tool → Promptfoo profile Weights & Biases Weave profile More Observability & Evaluation →

How to read this: HVTrust (0–100) weighs supply-chain signals (provenance, OSSF Scorecard, signed commits, open license) alongside real-world adoption, scaled by an evidence-confidence factor. Grade bands: A ≥ 80, B ≥ 65, C ≥ 50, D < 50. Signals refresh daily. Full methodology v4.4 →