/벤치마크, 평가, 보정 /jev-agent-failure-benchmark

jev-agent-failure-benchmark

Who and When 에이전트 실패 귀인 벤치마크에서 Jev와 강력한 LLM을 비교합니다.

이 페이지는 영문 목록에서 생성했습니다.

readme에서

jev-agent-failure-benchmark Can a fast, cheap decision model find what broke an AI agent as well as a frontier LLM? This benchmarks Jev (Typesafe.ai) on the text subset of Who&When Pro, an agent-failure-attribution benchmark: given a failed multi-agent run, predict the responsible agent, the decisive step, and the error type. Result On all 6,257 text …

상세

섹션
벤치마크, 평가, 보정
owner
TokenTrim
스타
1
포크
0
최근 커밋
2026-09-17
언어
Python
라이선스
Apache-2.0

어떤 프로젝트인가

형태
데이터셋 또는 벤치마크
호스트 에이전트
단독 실행
대상
연구자
성숙도
문서
실측치를 보고합니다

이럴 때 유용합니다

가장 가까운 용도