ホーム /ベンチマーク・評価・較正 /jev-agent-failure-benchmark

jev-agent-failure-benchmark

Who and When というエージェントの失敗要因特定ベンチマークで、Jev と強力な LLM を比較します。

このページは英語版リストから生成しています。

readme より

jev-agent-failure-benchmark Can a fast, cheap decision model find what broke an AI agent as well as a frontier LLM? This benchmarks Jev (Typesafe.ai) on the text subset of Who&When Pro, an agent-failure-attribution benchmark: given a failed multi-agent run, predict the responsible agent, the decisive step, and the error type. Result On all 6,257 text …

詳細

セクション
ベンチマーク・評価・較正
owner
TokenTrim
スター
1
フォーク
0
最終更新
2026-09-17
言語
Python
ライセンス
Apache-2.0

これは何か

形態
データセットやベンチマーク
ホストエージェント
単体で動作
対象
研究者
成熟度
ドキュメント
実測値を報告しています

こんなときに役立ちます

近い用途