首页 /基准、评测与校准 /jev-agent-failure-benchmark

jev-agent-failure-benchmark

在 Who and When 智能体失败归因基准上把 Jev 与一个强力 LLM 作对比。

本页由英文清单自动生成。

来自 readme

jev-agent-failure-benchmark Can a fast, cheap decision model find what broke an AI agent as well as a frontier LLM? This benchmarks Jev (Typesafe.ai) on the text subset of Who&When Pro, an agent-failure-attribution benchmark: given a failed multi-agent run, predict the responsible agent, the decisive step, and the error type. Result On all 6,257 text …

详情

分区
基准、评测与校准
owner
TokenTrim
星标
1
复刻
0
最近提交
2026-09-17
语言
Python
许可证
Apache-2.0

这是什么

形态
数据集或基准
宿主智能体
独立运行
面向人群
研究人员
成熟度
文档
报告了实测数据

适合用来

最匹配的意图