Home /Benchmarks, evals and calibration /jev-agent-failure-benchmark
jev-agent-failure-benchmark
Jev against a strong LLM on the Who and When agent-failure-attribution benchmark.
From the readme
jev-agent-failure-benchmark Can a fast, cheap decision model find what broke an AI agent as well as a frontier LLM? This benchmarks Jev (Typesafe.ai) on the text subset of Who&When Pro, an agent-failure-attribution benchmark: given a failed multi-agent run, predict the responsible agent, the decisive step, and the error type. Result On all 6,257 text …
Details
- Section
- Benchmarks, evals and calibration
- owner
- TokenTrim
- stars
- 1
- forks
- 0
- pushed
- 2026-09-17
- Language
- Python
- License
- Apache-2.0
What it is
- Form
- Dataset or benchmark
- Host agent
- Standalone
- Audience
- Researchers
Maturity
Docs
reports measured numbers