홈 /벤치마크, 평가, 보정 /jev-benchmark by themsquared
jev-benchmark by themsquared
실행 간 분산을 함께 보고하는 도구 호출 위험 분류이며, 오답은 모두 유보적인 신뢰도를 동반했습니다.
이 페이지는 영문 목록에서 생성했습니다.
readme에서
jev-benchmark 📖 Read the write-up: I Benchmarked Jev on Agent Tool-Call Risk. Calibration Held. A reproducible benchmark for TypeSafe AI's Jev on a real agent-infrastructure task: classifying an agent tool call as readonly, destructive, privileged, or exfiltration. Everything here is runnable. The task set is published, the harness is 200 lines, and the …
상세
- 섹션
- 벤치마크, 평가, 보정
- owner
- themsquared
- 스타
- 0
- 포크
- 0
- 최근 커밋
- 2026-09-19
- 언어
- Python
- 라이선스
- Apache-2.0
어떤 프로젝트인가
- 형태
- 데이터셋 또는 벤치마크
- 호스트 에이전트
- 단독 실행
- 대상
- 연구자
성숙도
문서
실측치를 보고합니다