Registry › Compare › Latitude vs Promptfoo

Latitude vs Promptfoo

Promptfoo leads on trust: 92.5/100 (Grade A) against 76.9/100 (Grade B), a 15.6-point gap. Promptfoo leads on supply-chain integrity and transparency, and rests on broader evidence.

B Latitude 76.9

Open-source observability for AI agents. Find where your agents fail, dispatch your coding agent to fix it, and verify the fix against real traces.

latitude-dev/latitude-llm · #210 overall · #8 Observability & Evaluation · coverage B (3/5)

Latitude doesn't lead on any scored dimension in this pair.

A Promptfoo 92.5

Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.

promptfoo/promptfoo · #7 overall · #1 Observability & Evaluation · coverage A (4/5)

Choose Promptfoo if supply-chain integrity and transparency matter most.

  • +10.0Safety / Integrity: 100% of recent commits signed, against 91%
  • +6.5Transparency
  • +4.7Adoption: 904.5k weekly downloads against 5.4k
  • A vs BEvidence coverage: 4 of 5 independent signal types, against 3

Where they differ

12.0
Safety / IntegrityPromptfoo +10.0
22.0
8.5
TransparencyPromptfoo +6.5
15.0
13.8
AdoptionPromptfoo +4.7
18.5
+4.7
Runtime calibrationLatitude +5.6
-0.9

2 dimensions identical: Identity 18.0 · Maintenance 19.9 · Full evidence table

An independent, evidence-based trust comparison of Latitude and Promptfoo, two Observability & Evaluation projects in the HVTracker registry. Scores come from public, checkable signals — supply-chain provenance, OSSF Scorecard, maintenance, and adoption — not popularity.

Full evidence

Signal Latitudelatitude-dev/latitude-llm Promptfoopromptfoo/promptfoo
HVTrust score 76.9 92.5
Evidence grade B A
Coverage grade B A
Overall rank #210 #7
Rank in Observability & Evaluation #8 #1
GitHub stars 4.7k 25.9k
Last updated 2d ago 1d ago
Build provenance Yes Yes
OSSF Scorecard — 7.6 / 10
License MIT MIT
Downloads 5k/wk 905k/wk
Trust dimensions (points earned)
Safety / integrity / 25 12.0 22.0
Identity & provenance / 18 18.0 18.0
Transparency / 17 8.5 15.0
Maintenance / 20 19.9 19.9
Adoption / 20 13.8 18.5
Runtime capability surface (full matrix)
MCP server Implemented Implemented
External providers — 3 — Amazon Bedrock, Anthropic, OpenAI
Requires API keys Yes Yes
Plugin surface plugins plugins
Provenance drift Match Match
Open in the live compare tool → Latitude profile Promptfoo profile More Observability & Evaluation →

How to read this: HVTrust (0–100) weighs supply-chain signals (provenance, OSSF Scorecard, signed commits, open license) alongside real-world adoption, scaled by an evidence-confidence factor. Grade bands: A ≥ 80, B ≥ 65, C ≥ 50, D < 50. Signals refresh daily. Full methodology v4.4 →