DeepEval
The LLM Evaluation Framework
Compare DeepEval
How does it stack up against its Observability & Evaluation neighbours?
Pick any agent to compare →Promising trust profile, but some evidence still deserves review.
Is DeepEval safe? DeepEval scores 70.6/100 (Grade B), ranked #138 of 1308 tracked open-source AI agent projects, on evidence coverage A (4 of 5 independent signal types). The public evidence: no package-provenance attestation found; OSSF Scorecard rates its supply-chain practices 3.8/10; 22% of recent commits are signed; last pushed 2026-08-17. Every point is earned from checkable signals — never paid placement. How scoring works →
Ranked neighbours in Observability & Evaluation
Quick Trust Read
How to read this: HVTrust (0–100) weighs supply-chain signals (provenance, OSSF Scorecard, signed commits, open license) alongside real-world adoption. Grade B reflects the trust score band: A ≥ 80, B ≥ 65, C ≥ 50, D < 50. Evidence coverage A is separate — it grades how many independent signal types back the score (4 of 5), so a high score on thin evidence stays visible. Full methodology →
Rank Trend
Activity & Reach
Analysis
HVTrust Dimensions vs Observability & Evaluation
70.6 / 100 · 100.0% confidenceDeepEval Observability & Evaluation average (23 agents)
Activity Inputs
90.6 / 100Supply Chain Trust
Is DeepEval safe?
Does DeepEval publish package provenance?
Does DeepEval have an OpenSSF Scorecard?
Is DeepEval actively maintained?
What license does DeepEval use?
Are DeepEval's commits signed?
Not a safety endorsement. HVTracker describes what public signals show, not whether a project is safe for your use case. Run your own security review before adopting in production.
AI agent surface
MCP, providers, tool surface
These runtime-trust fields — detected from public repo docs and manifests — contribute a bounded adjustment to this project's HVTrust score alongside supply-chain evidence. The exact values each field can add or subtract are documented in the methodology → Compare this surface across every listed agent in the capability matrix →
- MCP signal live
- External deps live
- Tool / plugin surface live
- Package provenance drift live
Detected changes to DeepEval's runtime surface and supply-chain posture, from daily public-signal snapshots. A change here means our detectors see something different — a genuinely changed capability, or better evidence of an existing one.
Maintain DeepEval?
For maintainers
HVTrust scores DeepEval from public signals only — we never contact maintainers first. If a signal is wrong, stale, or missing (provenance you publish, a Scorecard you run, signed releases), tell us and we'll review it. Corrections are public and tracked on GitHub.
Reputation Timeline
Signal history
Embed Badge Badge guide for maintainers →
For maintainers
[](https://hvtracker.net/agents/deepeval)
<a href="https://hvtracker.net/agents/deepeval"><img src="https://hvtracker.net/badge/deepeval.svg" alt="HVTrust"></a>
Other agents in Observability & Evaluation
GitHub REST API (repo, commits, stars, forks, license) · PyPI / pypistats (downloads, provenance) · OpenSSF Scorecard CLI · Algolia HN Search API
Each agent's signals refresh once daily across 6 staggered batches. Methodology v4.3 · Raw JSON