ホーム /ベンチマーク・評価・較正 /jev-agent-failure-benchmark
jev-agent-failure-benchmark
Who and When というエージェントの失敗要因特定ベンチマークで、Jev と強力な LLM を比較します。
このページは英語版リストから生成しています。
readme より
jev-agent-failure-benchmark Can a fast, cheap decision model find what broke an AI agent as well as a frontier LLM? This benchmarks Jev (Typesafe.ai) on the text subset of Who&When Pro, an agent-failure-attribution benchmark: given a failed multi-agent run, predict the responsible agent, the decisive step, and the error type. Result On all 6,257 text …
詳細
- セクション
- ベンチマーク・評価・較正
- owner
- TokenTrim
- スター
- 1
- フォーク
- 0
- 最終更新
- 2026-09-17
- 言語
- Python
- ライセンス
- Apache-2.0
これは何か
- 形態
- データセットやベンチマーク
- ホストエージェント
- 単体で動作
- 対象
- 研究者
成熟度
ドキュメント
実測値を報告しています