홈 /벤치마크, 평가, 보정 /jev-phishing-bench
jev-phishing-bench
피싱 이메일 2,000건에서 Jev와 Claude Haiku를 비교합니다: 정확도, 보정, 지연 시간, 비용을 다룹니다.
이 페이지는 영문 목록에서 생성했습니다.
readme에서
Jev vs LLM: a phishing decision benchmark with a calibration audit Public, reproducible comparison of Jev (TypeSafe AI's System One model, launched 15 September 2026) against a classic LLM on one security decision: should an email agent click the link in this email? The question this repo answers with numbers: is Jev accurate enough, are its probabilities …
상세
- 섹션
- 벤치마크, 평가, 보정
- owner
- anisselbd
- 스타
- 2
- 포크
- 0
- 최근 커밋
- 2026-09-19
- 언어
- Python
어떤 프로젝트인가
- 형태
- 데이터셋 또는 벤치마크
- 호스트 에이전트
- 단독 실행
- 대상
- 연구자
성숙도
문서
실측치를 보고합니다