LangWatch
Open-Source Agent Observability & Evals for AI Platform teams. Trace, test, route and govern every agent & coding assistant in the company.
Is LangWatch safe? Strong public trust posture, backed by multiple independent signals.
Compare LangWatch
How does it stack up against its Observability & Evaluation neighbours?
Pick any agent to compare →In detail: LangWatch scores 83.1/100 (Grade A), ranked #113 of 1371 tracked open-source AI agent projects, on evidence coverage A (4 of 5 independent signal types). The public evidence: its packages ship with cryptographic provenance; OSSF Scorecard rates its supply-chain practices 4.9/10; 83% of recent commits are signed; no published advisory affects its latest release (@langwatch/server); last pushed 2026-10-09. Every point is earned from checkable signals — never paid placement. How scoring works →
How LangWatch could raise its score
Each line changes one public signal and recomputes with the live scoring function. The gains don't add up exactly, because the score is capped near the top.
- Raise the OSSF Scorecard from 4.9 to 9.0lowest checks: CII-Best-Practices 0, Dangerous-Workflow 0, Fuzzing 0 +6.8 → 89.9 A · #2 of 26
- Sign every commit83% signed today +0.6 → 83.7 A · #5 of 26
Chain of custody
Scanners check what the code says. This traces who ships LangWatch and whether that has changed: from the source repository, through how changes are reviewed and released, to the package you install. These are the checks for OWASP ASI04 Agentic Supply Chain Vulnerabilities.
-
Source Signed
github.com/langwatch/langwatch, Apache-2.0, last pushed 2026-10-09. 83% of the last 100 commits carry a verified signature, so changes trace to a verified account.
-
Review and release Mixed
OpenSSF Scorecard rates the repository's practices 4.9/10 (scanned Aug 30, 2026, refresh pending). The checks that decide who can get a change released:
- Code Review2
- Branch Protection4
- Signed Releases0
- Dangerous Workflow0
- Token Permissions0
- Pinned Dependencies7
All 18 Scorecard checks
Binary-Artifacts 10Branch-Protection 4CI-Tests 10CII-Best-Practices 0Code-Review 2Contributors 10Dangerous-Workflow 0Dependency-Update-Tool 10Fuzzing 0License 10Maintained 10Packaging 10Pinned-Dependencies 7SAST 7Security-Policy 10Signed-Releases 0Token-Permissions 0Vulnerabilities 0 -
Published packages Partly attested
- npm @langwatch/server Source link points back to this repo Build provenance attested
- PyPI langwatch Source link isn't a GitHub repo No build attestation
-
What you install today No known advisories
Latest release checked: @langwatch/server 3.20.1. No published advisory affects it (OSV, checked 2026-10-09).
Changes to this chain
- 2026-08-30 OSSF Scorecard: 4.9/10
Evidence behind this score
Coverage A: 4 of 5 independent evidence types found.
- GitHub repository data
- Package downloads
- Supply-chain checks
- Public actions (missing)
- Community mentions
Verify this score yourself
Every build re-issues LangWatch’s score as an Ed25519-signed credential, valid for 7 days. You can check it offline against HVTracker’s published key; if anyone changes a number after signing, verification fails.
What this score doesn’t check
It covers who ships the code, not what the code does. It doesn’t read tool descriptions or prompts for injected instructions, watch runtime behaviour, or find bugs nobody has disclosed yet. For that, run a content scanner before you connect it, such as Cisco MCP Scanner or Snyk Agent Scan.
How LangWatch compares in Observability & Evaluation
- #3 OpenLLMetry 86.1 +3.0
- #4 Weights & Biases Weave 84.3 +1.2
- #5 LangWatch 83.1 this agent
- #6 Langfuse 79.8 −3.3
- #7 Evidently 78.9 −4.2
Bars show each HVTrust score; the tick marks LangWatch’s 83.1.
Where the 83.1 comes from
HVTrust dimensions vs the Observability & Evaluation average
83.1 / 100 · 100.0% confidenceLangWatch Observability & Evaluation average (26 agents)
Quick Trust Read
How to read this: HVTrust (0–100) weighs supply-chain signals (provenance, OSSF Scorecard, signed commits, open license) alongside real-world adoption. Grade A reflects the trust score band: A ≥ 80, B ≥ 65, C ≥ 50, D < 50. Evidence coverage A is separate — it grades how many independent signal types back the score (4 of 5), so a high score on thin evidence stays visible. Full methodology →
Rank Trend
Activity & Reach
Analysis
Activity Inputs
84.3 / 100Common questions about LangWatch
Does LangWatch publish package provenance?
Does LangWatch have an OpenSSF Scorecard?
Is LangWatch actively maintained?
What license does LangWatch use?
Are LangWatch's commits signed?
Not a safety endorsement. HVTracker describes what public signals show, not whether a project is safe for your use case. Run your own security review before adopting in production.
Compare LangWatch head-to-head
AI agent surface
MCP, providers, tool surface
These runtime-trust fields — detected from public repo docs and manifests — contribute a bounded adjustment to this project's HVTrust score alongside supply-chain evidence. The exact values each field can add or subtract are documented in the methodology → Compare this surface across every listed agent in the capability matrix →
- database
- MCP signal live
- External deps live
- Tool / plugin surface live
- Package provenance drift live
Detected changes to LangWatch's runtime surface and supply-chain posture, from daily public-signal snapshots. A change here means our detectors see something different — a genuinely changed capability, or better evidence of an existing one.
Maintain LangWatch?
For maintainers
HVTrust scores LangWatch from public signals only — we never contact maintainers first. If a signal is wrong, stale, or missing (provenance you publish, a Scorecard you run, signed releases), tell us and we'll review it. Corrections are public and tracked on GitHub.
Reputation Timeline
Signal history
Embed Badge Badge guide for maintainers →
For maintainers
[](https://hvtracker.net/agents/langwatch)
<a href="https://hvtracker.net/agents/langwatch"><img src="https://hvtracker.net/badge/langwatch.svg" alt="HVTrust"></a>
Other agents in Observability & Evaluation
GitHub REST API (repo, commits, stars, forks, license) · npm Registry (downloads, provenance) · PyPI / pypistats (downloads, provenance) · OpenSSF Scorecard CLI · Algolia HN Search API
Each agent's signals refresh once daily across 6 staggered batches. Methodology v4.4 · Raw JSON