/벤치마크, 평가, 보정 /jev-benchmark by themsquared

jev-benchmark by themsquared

실행 간 분산을 함께 보고하는 도구 호출 위험 분류이며, 오답은 모두 유보적인 신뢰도를 동반했습니다.

이 페이지는 영문 목록에서 생성했습니다.

readme에서

jev-benchmark 📖 Read the write-up: I Benchmarked Jev on Agent Tool-Call Risk. Calibration Held. A reproducible benchmark for TypeSafe AI's Jev on a real agent-infrastructure task: classifying an agent tool call as readonly, destructive, privileged, or exfiltration. Everything here is runnable. The task set is published, the harness is 200 lines, and the …

상세

섹션
벤치마크, 평가, 보정
owner
themsquared
스타
0
포크
0
최근 커밋
2026-09-19
언어
Python
라이선스
Apache-2.0

어떤 프로젝트인가

형태
데이터셋 또는 벤치마크
호스트 에이전트
단독 실행
대상
연구자
성숙도
문서
실측치를 보고합니다

이럴 때 유용합니다

가장 가까운 용도