AgentBench

A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)

Observability & Evaluation Python Grade D Listed Apache-2.0
37.3/100
Rank #1201 of 1328
Compare AgentBench

How does it stack up against its Observability & Evaluation neighbours?

Pick any agent to compare →

Thin or incomplete trust evidence. Review carefully before production use.

Open compare tool Suggest correction
Listing state
Listed
Evidence coverage
Grade C · 2/5 signals
Last push
2026-02-08 · 224d ago
Recent change
No recent signal change

Is AgentBench safe? AgentBench scores 37.3/100 (Grade D), ranked #1201 of 1328 tracked open-source AI agent projects, on evidence coverage C (2 of 5 independent signal types). The public evidence: no package-provenance attestation found; OSSF Scorecard rates its supply-chain practices 3.4/10; 51% of recent commits are signed; last pushed 2026-02-08. Every point is earned from checkable signals — never paid placement. How scoring works →

Ranked neighbours in Observability & Evaluation

Quick Trust Read

What Would Improve It
Publish package provenance or release attestations for stronger supply-chain evidence.
Recent Changes
2026-09-09
Rank Moved
Rank dropped 12 spots (#1182 → #1194)
2026-09-08
Rank Moved
Rank dropped 12 spots (#1170 → #1182)
2026-09-07
Rank Moved
Rank dropped 13 spots (#1157 → #1170)
Maintainer Checklist
Raise Scorecard signals Current OSSF Scorecard is 3.4/10. Tighten the weakest checks to improve public safety evidence.
Publish provenance Add package provenance or release attestations so users can verify where shipped artifacts came from.
Refresh maintenance signals The repo was last pushed 224 days ago. Fresh activity helps separate stable projects from stale ones.
32.8
Activity sub-score · out of 100
#22

How to read this: HVTrust (0–100) weighs supply-chain signals (provenance, OSSF Scorecard, signed commits, open license) alongside real-world adoption. Grade D reflects the trust score band: A ≥ 80, B ≥ 65, C ≥ 50, D < 50. Evidence coverage C is separate — it grades how many independent signal types back the score (2 of 5), so a high score on thin evidence stays visible. Full methodology →

Signals refreshed 2026-09-20 15:32 UTC · Repo last pushed 224 days ago — may be stale

Rank Trend

2026-08-11 2026-09-20

Activity & Reach

Stars
3.7k
Forks
280
Last Push
2026-02-08
224 days ago
Commits (4 wk)
0
Downloads (7d)
HN mentions (30d)
Open Issues
65
Rank Change
▲5
was #1206

Analysis

HVTrust Dimensions vs Observability & Evaluation

37.3 / 100 · 100.0% confidence

AgentBench Observability & Evaluation average (23 agents)

Safety / Integrity50% OSSF Scorecard · 30% provenance · 20% signed commits
6.8 / 25
5.1 below avg 11.9
Identity / Provenance60% listing status · 40% build provenance
10.8 / 18
2.5 below avg 13.3
Transparency50% declared license · 50% OSSF Scorecard
11.4 / 17
1.0 below avg 12.4
Maintenance60% last-push freshness · 40% commit activity
0 / 20
15.0 below avg 15.0
AdoptionLog-scaled stars · package downloads
8.6 / 20
5.2 below avg 13.8

Activity Inputs

32.8 / 100
StarsRepository reach
21.4 / 30
FreshnessLast push recency
0.0 / 25
ActivityRecent commits
0.0 / 25
CommunityFork signal
11.4 / 20

Supply Chain Trust

Package Provenance
None
No package attestations found
OSSF Scorecard
3.4 / 10
OpenSSF Scorecard · scanned Sep 13, 2026
Signed Commits
51%
of last 100 commits verified
Binary-Artifacts 10
Branch-Protection 6
CI-Tests 0
CII-Best-Practices 0
Code-Review 6
Contributors 10
Dangerous-Workflow 10
Dependency-Update-Tool 0
Fuzzing 0
License 10
Maintained 0
Packaging -1
Pinned-Dependencies 0
SAST 0
Security-Policy 0
Signed-Releases -1
Token-Permissions 0
Vulnerabilities 0

Is AgentBench safe?

Public trust evidence for AgentBench is thin: several supply-chain signals are missing or weak. This does not mean the project is unsafe — it means an outside observer cannot easily verify the usual integrity checks. Treat with extra scrutiny.
Does AgentBench publish package provenance?
No published build provenance is currently detected for AgentBench. This is common for open-source projects but means consumers cannot independently verify that the package on the registry matches the GitHub source.
Does AgentBench have an OpenSSF Scorecard?
AgentBench has an OpenSSF Scorecard score of 3.4/10. The Scorecard checks for branch protection, signed releases, dependency updates, fuzzing, code review, and other supply-chain hygiene items. See the full check breakdown on this page.
Is AgentBench actively maintained?
Stale. The repository has not been pushed to in 224 days. Consider whether the project is still being maintained.
What license does AgentBench use?
AgentBench ships under Apache-2.0. A declared, OSI-approved license is one of the transparency signals HVTrust scores.
Are AgentBench's commits signed?
50% of the last 100 commits to AgentBench are verified-signed (GPG, SSH, S/MIME, or GitHub's signing flow). Signed commits help confirm that code was authored by who the commit claims.

Not a safety endorsement. HVTracker describes what public signals show, not whether a project is safe for your use case. Run your own security review before adopting in production.

AI agent surface

MCP, providers, tool surface
Scored in HVTrust

These runtime-trust fields — detected from public repo docs and manifests — contribute a bounded adjustment to this project's HVTrust score alongside supply-chain evidence. The exact values each field can add or subtract are documented in the methodology → Compare this surface across every listed agent in the capability matrix →

MCP Server Support
None detected
No MCP server signal detected.
Detailed evidence is not shown in the public view.
External Service Dependencies
medium confidence
1 detected
Public provider/service dependencies detected.
Credential signal: No explicit API-key/config marker detected.
Tool / Plugin Surface
medium confidence
1 tags
Broad capability areas detected.
  • browser
Detailed evidence is not shown in the public view.
Package Provenance Drift
N/A
No package source configured
Detailed evidence is not shown in the public view.
  • MCP signal live
  • External deps live
  • Tool / plugin surface live
  • Package provenance drift live

Maintain AgentBench?

For maintainers

HVTrust scores AgentBench from public signals only — we never contact maintainers first. If a signal is wrong, stale, or missing (provenance you publish, a Scorecard you run, signed releases), tell us and we'll review it. Corrections are public and tracked on GitHub.

Reputation Timeline

Signal history
Rank 24HVTrust 1Scorecard 1
2026-09-09
Rank Moved
Rank dropped 12 spots (#1182 → #1194)
2026-09-08
Rank Moved
Rank dropped 12 spots (#1170 → #1182)
2026-09-07
Rank Moved
Rank dropped 13 spots (#1157 → #1170)
2026-09-06
Rank Moved
Rank dropped 15 spots (#1142 → #1157)
2026-09-05
Rank Moved
Rank dropped 35 spots (#1107 → #1142)
2026-09-04
Rank Moved
Rank dropped 20 spots (#1087 → #1107)
2026-09-03
Rank Moved
Rank dropped 51 spots (#1036 → #1087)
2026-09-02
Rank Moved
Rank dropped 42 spots (#994 → #1036)
2026-09-01
Rank Moved
Rank dropped 45 spots (#949 → #994)
2026-08-31
Rank Moved
Rank dropped 31 spots (#918 → #949)
2026-08-30
Rank Moved
Rank dropped 44 spots (#874 → #918)
2026-08-29
Rank Moved
Rank dropped 47 spots (#827 → #874)
2026-08-28
Rank Moved
Rank dropped 23 spots (#804 → #827)
2026-08-27
Rank Moved
Rank dropped 22 spots (#782 → #804)
2026-08-26
Rank Moved
Rank dropped 101 spots (#681 → #782)
2026-08-16
Rank Moved
Rank dropped 12 spots (#663 → #675)
2026-08-10
Rank Moved
Rank dropped 29 spots (#635 → #664)
2026-08-08
Rank Moved
Rank dropped 234 spots (#401 → #635)
2026-07-29
Rank Moved
Rank dropped 11 spots (#389 → #400)
2026-07-15
Rank Moved
Rank dropped 27 spots (#354 → #381)
2026-07-14
Rank Moved
Rank dropped 20 spots (#334 → #354)
2026-07-13
Rank Moved
Rank dropped 36 spots (#298 → #334)
2026-06-24
Rank Moved
Rank dropped 35 spots (#250 → #285)
2026-06-23
Scorecard Added
OSSF Scorecard: 3.4/10
2026-06-23
Rank Moved
Rank rose 48 spots (#298 → #250)
2026-06-23
HVTrust Changed
HVTrust up 23.8pts (16.7 → 40.5)

Embed Badge Badge guide for maintainers →

For maintainers
HVTrust 37.3 Grade D
Markdown:
[![HVTrust](https://hvtracker.net/badge/agentbench.svg)](https://hvtracker.net/agents/agentbench)
HTML:
<a href="https://hvtracker.net/agents/agentbench"><img src="https://hvtracker.net/badge/agentbench.svg" alt="HVTrust"></a>

Other agents in Observability & Evaluation

Data sources
GitHub REST API (repo, commits, stars, forks, license) · OpenSSF Scorecard CLI
Each agent's signals refresh once daily across 6 staggered batches. Methodology v4.3 · Raw JSON