Home /Benchmarks, evals and calibration /jev-agent-failure-benchmark

jev-agent-failure-benchmark

Jev against a strong LLM on the Who and When agent-failure-attribution benchmark.

From the readme

jev-agent-failure-benchmark Can a fast, cheap decision model find what broke an AI agent as well as a frontier LLM? This benchmarks Jev (Typesafe.ai) on the text subset of Who&When Pro, an agent-failure-attribution benchmark: given a failed multi-agent run, predict the responsible agent, the decisive step, and the error type. Result On all 6,257 text …

Details

Section
Benchmarks, evals and calibration
owner
TokenTrim
stars
1
forks
0
pushed
2026-09-17
Language
Python
License
Apache-2.0

What it is

Form
Dataset or benchmark
Host agent
Standalone
Audience
Researchers
Maturity
Docs
reports measured numbers

Useful for

Best intent matches