Eval Skills

Skills that guide AI coding agents to help you build product-specific AI evals.

Agent Skills Listed Apache-2.0

Is Eval Skills safe? Thin or incomplete trust evidence. Review carefully before production use.

OSSF Scorecard 4.1 / 10
Provenance None
Signed commits 29%
Last push 16d ago
Compare Eval Skills
Suggest a correction

In detail: Eval Skills scores 51.4/100 (Grade C), ranked #221 of 371 tracked open-source AI agent projects, on evidence coverage C (2 of 5 independent signal types). The public evidence: no package-provenance attestation found; OSSF Scorecard rates its supply-chain practices 4.1/10; 29% of recent commits are signed; last pushed 2026-09-24. Every point is earned from checkable signals — never paid placement. How scoring works →

How Eval Skills could raise its score

Each line changes one public signal and recomputes with the live scoring function. The gains don't add up exactly, because the score is capped near the top.

Chain of custody

Scanners check what the code says. This traces who ships Eval Skills and whether that has changed: from the source repository, through how changes are reviewed and released, to the code you run. These are the checks for OWASP ASI04 Agentic Supply Chain Vulnerabilities and, for skills, AST02 Supply Chain Compromise.

  1. Source Partly signed

    github.com/ai-evals-course/evals-skills, Apache-2.0, last pushed 2026-09-24. 29% of the last 100 commits carry a verified signature; the rest can't be tied to a verified identity.

  2. Review and release Mixed

    OpenSSF Scorecard rates the repository's practices 4.1/10 (scanned Oct 10, 2026). The checks that decide who can get a change released:

    • Code Review0
    • Branch Protection0
    • Signed Releasesn/a
    • Dangerous Workflown/a
    • Token Permissionsn/a
    • Pinned Dependenciesn/a
    All 18 Scorecard checks
    Binary-Artifacts 10
    Branch-Protection 0
    CI-Tests 4
    CII-Best-Practices 0
    Code-Review 0
    Contributors 10
    Dangerous-Workflow -1
    Dependency-Update-Tool 0
    Fuzzing 0
    License 10
    Maintained 10
    Packaging -1
    Pinned-Dependencies -1
    SAST 0
    Security-Policy 0
    Signed-Releases -1
    Token-Permissions -1
    Vulnerabilities 10
  3. Published packages None

    No registry package. Eval Skills is used from its repository, so what you run is whatever you check out: pin a release tag or commit.

Changes to this chain

  • 2026-10-10 OSSF Scorecard: 4.1/10

Evidence behind this score

Coverage C: 2 of 5 independent evidence types found.

  • GitHub repository data
  • Package downloads (missing)
  • Supply-chain checks
  • Public actions (missing)
  • Community mentions (missing)

Verify this score yourself

Every build re-issues Eval Skills’s score as an Ed25519-signed credential, valid for 7 days. You can check it offline against HVTracker’s published key; if anyone changes a number after signing, verification fails.

What this score doesn’t check

It covers who ships the code, not what the code does. It doesn’t read the skill’s instructions for injected instructions, watch runtime behaviour, or find bugs nobody has disclosed yet. For that, run a content scanner before you install it, such as Cisco Skill Scanner.

How Eval Skills compares in Agent Skills

  1. #219 Digital Marketing Pro 51.7 +0.3
  2. #220 Oxylabs Agent Skills 51.7 +0.3
  3. #221 Eval Skills 51.4 this agent
  4. #222 SimpleEnglish 51.4 ±0
  5. #223 Claude Skills Governance Risk and Compliance 51.3 −0.1

Bars show each HVTrust score; the tick marks Eval Skills’s 51.4.

Where the 51.4 comes from

HVTrust dimensions vs the Agent Skills average

51.4 / 100 · 100.0% confidence

Eval Skills Agent Skills average (371 agents)

Safety / Integrity50% OSSF Scorecard · 30% provenance · 20% signed commits
6.6 / 25
1.2 below avg 7.8
Identity / Provenance60% listing status · 40% build provenance
10.8 / 18
0.6 below avg 11.4
Transparency50% declared license · 50% OSSF Scorecard
12.0 / 17
in line with avg 12.0
Maintenance60% last-push freshness · 40% commit activity
14.7 / 20
0.6 above avg 14.1
AdoptionLog-scaled stars · package downloads
7.6 / 20
1.4 below avg 9.0

Quick Trust Read

What Would Improve It
Publish package provenance or release attestations for stronger supply-chain evidence.
Recent Changes
2026-10-10
Scorecard Added
OSSF Scorecard: 4.1/10
2026-10-10
Grade Changed
Trust grade D → C
2026-10-10
Rank Moved
Rank rose 145 spots (#366 → #221)
Maintainer Checklist
Raise Scorecard signals Current OSSF Scorecard is 4.1/10. Tighten the weakest checks to improve public safety evidence.
Publish provenance Add package provenance or release attestations so users can verify where shipped artifacts came from.
Increase signed commits Raise the share of verified-signed commits to make maintainer identity and release history easier to trust.
63.0
Activity sub-score · out of 100
#221

How to read this: HVTrust (0–100) weighs supply-chain signals (provenance, OSSF Scorecard, signed commits, open license) alongside real-world adoption. Grade C reflects the trust score band: A ≥ 80, B ≥ 65, C ≥ 50, D < 50. Evidence coverage C is separate — it grades how many independent signal types back the score (2 of 5), so a high score on thin evidence stays visible. Full methodology →

Signals refreshed 2026-10-10 20:07 UTC · Repo last pushed 16 days ago

Rank Trend

2026-10-09 2026-10-10

Activity & Reach

Stars
1.5k
Forks
101
Last Push
2026-09-24
16 days ago
Commits (4 wk)
8
Downloads (7d)
—
HN mentions (30d)
—
Open Issues
0
Rank Change
▲145
was #366

Analysis

Activity Inputs

63.0 / 100
StarsRepository reach
19.0 / 30
FreshnessLast push recency
22.8 / 25
ActivityRecent commits
11.9 / 25
CommunityFork signal
9.3 / 20

Common questions about Eval Skills

Public trust evidence for Eval Skills is thin: several supply-chain signals are missing or weak. This does not mean the project is unsafe — it means an outside observer cannot easily verify the usual integrity checks. Treat with extra scrutiny.
Does Eval Skills publish package provenance?
No published build provenance is currently detected for Eval Skills. This is common for open-source projects but means consumers cannot independently verify that the package on the registry matches the GitHub source.
Does Eval Skills have an OpenSSF Scorecard?
Eval Skills has an OpenSSF Scorecard score of 4.1/10. The Scorecard checks for branch protection, signed releases, dependency updates, fuzzing, code review, and other supply-chain hygiene items. See the full check breakdown on this page.
Is Eval Skills actively maintained?
Maintained. Last push was 16 days ago.
What license does Eval Skills use?
Eval Skills ships under Apache-2.0. A declared, OSI-approved license is one of the transparency signals HVTrust scores.
Are Eval Skills's commits signed?
28% of the last 100 commits to Eval Skills are verified-signed (GPG, SSH, S/MIME, or GitHub's signing flow). Signed commits help confirm that code was authored by who the commit claims.

Not a safety endorsement. HVTracker describes what public signals show, not whether a project is safe for your use case. Run your own security review before adopting in production.

AI agent surface

MCP, providers, tool surface
Scored in HVTrust

These runtime-trust fields — detected from public repo docs and manifests — contribute a bounded adjustment to this project's HVTrust score alongside supply-chain evidence. The exact values each field can add or subtract are documented in the methodology → Compare this surface across every listed agent in the capability matrix →

MCP Server Support
None detected
No MCP server signal detected.
Detailed evidence is not shown in the public view.
External Service Dependencies
medium confidence
1 detected
Public provider/service dependencies detected.
Credential signal: No explicit API-key/config marker detected.
Tool / Plugin Surface
high confidence
Declared
Declared plugin/integration surface detected.
Detailed evidence is not shown in the public view.
Package Provenance Drift
N/A
No package source configured
Detailed evidence is not shown in the public view.
  • MCP signal live
  • External deps live
  • Tool / plugin surface live
  • Package provenance drift live

Maintain Eval Skills?

For maintainers

HVTrust scores Eval Skills from public signals only — we never contact maintainers first. If a signal is wrong, stale, or missing (provenance you publish, a Scorecard you run, signed releases), tell us and we'll review it. Corrections are public and tracked on GitHub.

Reputation Timeline

Signal history
Listed 1HVTrust 1Rank 1Grade 1Scorecard 1
2026-10-10
Scorecard Added
OSSF Scorecard: 4.1/10
2026-10-10
Grade Changed
Trust grade D → C
2026-10-10
Rank Moved
Rank rose 145 spots (#366 → #221)
2026-10-10
HVTrust Changed
HVTrust up 30.1pts (21.3 → 51.4)
2026-10-09
Newly Listed
First tracked at rank #366

Embed Badge Badge guide for maintainers →

For maintainers
HVTrust 51.4 Grade C
Markdown:
[![HVTrust](https://hvtracker.net/badge/eval-skills.svg)](https://hvtracker.net/agents/eval-skills)
HTML:
<a href="https://hvtracker.net/agents/eval-skills"><img src="https://hvtracker.net/badge/eval-skills.svg" alt="HVTrust"></a>

Other agents in Agent Skills

Data sources
GitHub REST API (repo, commits, stars, forks, license) · OpenSSF Scorecard CLI
Each agent's signals refresh once daily across 6 staggered batches. Methodology v4.4 · Raw JSON