DeepEval

The LLM Evaluation Framework

Observability & Evaluation Python Grade B Listed Apache-2.0
70.6/100
Rank #138 of 1308
Compare DeepEval

How does it stack up against its Observability & Evaluation neighbours?

Pick any agent to compare →

Promising trust profile, but some evidence still deserves review.

Open compare tool Suggest correction
Listing state
Listed
Evidence coverage
Grade A · 4/5 signals
Last push
2026-08-17 · 2d ago
Recent change
Rank +1

Is DeepEval safe? DeepEval scores 70.6/100 (Grade B), ranked #138 of 1308 tracked open-source AI agent projects, on evidence coverage A (4 of 5 independent signal types). The public evidence: no package-provenance attestation found; OSSF Scorecard rates its supply-chain practices 3.8/10; 22% of recent commits are signed; last pushed 2026-08-17. Every point is earned from checkable signals — never paid placement. How scoring works →

Ranked neighbours in Observability & Evaluation

Quick Trust Read

What Would Improve It
Publish package provenance or release attestations for stronger supply-chain evidence.
Recent Changes
2026-08-17
Rank Moved
Rank rose 12 spots (#142 → #130)
2026-08-12
Rank Moved
Rank rose 10 spots (#136 → #126)
2026-08-10
Rank Moved
Rank dropped 14 spots (#123 → #137)
Maintainer Checklist
Raise Scorecard signals Current OSSF Scorecard is 3.8/10. Tighten the weakest checks to improve public safety evidence.
Publish provenance Add package provenance or release attestations so users can verify where shipped artifacts came from.
Increase signed commits Raise the share of verified-signed commits to make maintainer identity and release history easier to trust.
90.6
Activity sub-score · out of 100
#9

How to read this: HVTrust (0–100) weighs supply-chain signals (provenance, OSSF Scorecard, signed commits, open license) alongside real-world adoption. Grade B reflects the trust score band: A ≥ 80, B ≥ 65, C ≥ 50, D < 50. Evidence coverage A is separate — it grades how many independent signal types back the score (4 of 5), so a high score on thin evidence stays visible. Full methodology →

Signals refreshed 2026-08-19 05:02 UTC · Repo last pushed 2 days ago

Rank Trend

2026-08-11 2026-08-19

Activity & Reach

Stars
17.7k
Forks
1.8k
Last Push
2026-08-17
2 days ago
Commits (4 wk)
191
Downloads (7d)
2,166,901
pypi
HN mentions (30d)
1
Open Issues
247
Rank Change
▼8
was #130

Analysis

HVTrust Dimensions vs Observability & Evaluation

70.6 / 100 · 100.0% confidence

DeepEval Observability & Evaluation average (23 agents)

Safety / Integrity50% OSSF Scorecard · 30% provenance · 20% signed commits
5.9 / 25
5.1 below avg 11.0
Identity / Provenance60% listing status · 40% build provenance
10.8 / 18
2.2 below avg 13.0
Transparency50% declared license · 50% OSSF Scorecard
11.7 / 17
in line with avg 11.9
Maintenance60% last-push freshness · 40% commit activity
19.9 / 20
4.8 above avg 15.1
AdoptionLog-scaled stars · package downloads
18.6 / 20
4.7 above avg 13.9

Activity Inputs

90.6 / 100
StarsRepository reach
25.5 / 30
FreshnessLast push recency
24.7 / 25
ActivityRecent commits
25 / 25
CommunityFork signal
15.2 / 20

Supply Chain Trust

Package Provenance
None
No package attestations found
OSSF Scorecard
3.8 / 10
OpenSSF Scorecard · scanned Aug 18, 2026
Signed Commits
22%
of last 100 commits verified
Binary-Artifacts 10
Branch-Protection 0
CI-Tests 10
CII-Best-Practices 0
Code-Review 2
Contributors 10
Dangerous-Workflow 10
Dependency-Update-Tool 0
Fuzzing 0
License 10
Maintained 10
Packaging -1
Pinned-Dependencies 2
SAST 0
Security-Policy 0
Signed-Releases -1
Token-Permissions 0
Vulnerabilities 0

Is DeepEval safe?

DeepEval has a mixed signal profile. Some trust indicators are present, others are missing. Whether it is safe for your use case depends on which gaps matter to you — review the breakdown below before adopting in production.
Does DeepEval publish package provenance?
No published build provenance is currently detected for DeepEval. This is common for open-source projects but means consumers cannot independently verify that the package on the registry matches the GitHub source.
Does DeepEval have an OpenSSF Scorecard?
DeepEval has an OpenSSF Scorecard score of 3.8/10. The Scorecard checks for branch protection, signed releases, dependency updates, fuzzing, code review, and other supply-chain hygiene items. See the full check breakdown on this page.
Is DeepEval actively maintained?
Actively maintained. The repository was pushed to within the last 2 day(s).
What license does DeepEval use?
DeepEval ships under Apache-2.0. A declared, OSI-approved license is one of the transparency signals HVTrust scores.
Are DeepEval's commits signed?
22% of the last 100 commits to DeepEval are verified-signed (GPG, SSH, S/MIME, or GitHub's signing flow). Signed commits help confirm that code was authored by who the commit claims.

Not a safety endorsement. HVTracker describes what public signals show, not whether a project is safe for your use case. Run your own security review before adopting in production.

AI agent surface

MCP, providers, tool surface
Scored in HVTrust

These runtime-trust fields — detected from public repo docs and manifests — contribute a bounded adjustment to this project's HVTrust score alongside supply-chain evidence. The exact values each field can add or subtract are documented in the methodology → Compare this surface across every listed agent in the capability matrix →

MCP Server Support
medium confidence
Implemented
DeepEval appears to expose MCP server capabilities.
Detailed evidence is not shown in the public view.
External Service Dependencies
high confidence
3 detected
Public provider/service dependencies detected.
Credential signal: API keys or service config markers documented.
Tool / Plugin Surface
high confidence
Declared
Declared plugin/integration surface detected.
Detailed evidence is not shown in the public view.
Package Provenance Drift
high confidence
Match
Published package metadata matches the tracked repo
Detailed evidence is not shown in the public view.
  • MCP signal live
  • External deps live
  • Tool / plugin surface live
  • Package provenance drift live
How this surface has changed

Detected changes to DeepEval's runtime surface and supply-chain posture, from daily public-signal snapshots. A change here means our detectors see something different — a genuinely changed capability, or better evidence of an existing one.

2026-06-16
Provider Added
Runtime surface grew — new detected provider dependencies: Amazon Bedrock, Google Gemini
2026-06-15
Provider Removed
Runtime surface shrank — no longer detected: Amazon Bedrock, Google Gemini
2026-06-13
Provider Added
Runtime surface grew — new detected provider dependencies: Amazon Bedrock, Google Gemini
2026-06-11
Provider Removed
Runtime surface shrank — no longer detected: Amazon Bedrock, Google Gemini
2026-06-10
Provider Added
Runtime surface grew — new detected provider dependencies: Amazon Bedrock, Google Gemini

Maintain DeepEval?

For maintainers

HVTrust scores DeepEval from public signals only — we never contact maintainers first. If a signal is wrong, stale, or missing (provenance you publish, a Scorecard you run, signed releases), tell us and we'll review it. Corrections are public and tracked on GitHub.

Reputation Timeline

Signal history
Rank 7Surface 3HVTrust 2Surface 2Listed 1Grade 1Scorecard 1Score 1
2026-08-17
Rank Moved
Rank rose 12 spots (#142 → #130)
2026-08-12
Rank Moved
Rank rose 10 spots (#136 → #126)
2026-08-10
Rank Moved
Rank dropped 14 spots (#123 → #137)
2026-08-08
Rank Moved
Rank dropped 12 spots (#110 → #122)
2026-07-13
Rank Moved
Rank dropped 13 spots (#87 → #100)
2026-06-30
Rank Moved
Rank rose 11 spots (#117 → #106)
2026-06-24
Rank Moved
Rank dropped 11 spots (#100 → #111)
2026-06-16
Provider Added
Runtime surface grew — new detected provider dependencies: Amazon Bedrock, Google Gemini
2026-06-15
Provider Removed
Runtime surface shrank — no longer detected: Amazon Bedrock, Google Gemini
2026-06-13
Provider Added
Runtime surface grew — new detected provider dependencies: Amazon Bedrock, Google Gemini
2026-06-11
Provider Removed
Runtime surface shrank — no longer detected: Amazon Bedrock, Google Gemini
2026-06-10
Provider Added
Runtime surface grew — new detected provider dependencies: Amazon Bedrock, Google Gemini
2026-05-29
HVTrust Changed
HVTrust up 9.9pts (52.0 → 61.9)
2026-05-28
Activity Score Changed
Activity score up 25pts (65 → 90)
2026-05-27
Scorecard Added
OSSF Scorecard: 3.8/10
2026-05-27
Grade Changed
Trust grade C → B
2026-05-27
HVTrust Changed
HVTrust up 22.5pts (36.8 → 59.3)
2026-05-25
Newly Listed
First tracked at rank #76

Embed Badge Badge guide for maintainers →

For maintainers
HVTrust 70.6 Grade B
Markdown:
[![HVTrust](https://hvtracker.net/badge/deepeval.svg)](https://hvtracker.net/agents/deepeval)
HTML:
<a href="https://hvtracker.net/agents/deepeval"><img src="https://hvtracker.net/badge/deepeval.svg" alt="HVTrust"></a>

Other agents in Observability & Evaluation

Data sources
GitHub REST API (repo, commits, stars, forks, license) · PyPI / pypistats (downloads, provenance) · OpenSSF Scorecard CLI · Algolia HN Search API
Each agent's signals refresh once daily across 6 staggered batches. Methodology v4.3 · Raw JSON