/벤치마크, 평가, 보정 /jev-phishing-bench

jev-phishing-bench

피싱 이메일 2,000건에서 Jev와 Claude Haiku를 비교합니다: 정확도, 보정, 지연 시간, 비용을 다룹니다.

이 페이지는 영문 목록에서 생성했습니다.

readme에서

Jev vs LLM: a phishing decision benchmark with a calibration audit Public, reproducible comparison of Jev (TypeSafe AI's System One model, launched 15 September 2026) against a classic LLM on one security decision: should an email agent click the link in this email? The question this repo answers with numbers: is Jev accurate enough, are its probabilities …

상세

섹션
벤치마크, 평가, 보정
owner
anisselbd
스타
2
포크
0
최근 커밋
2026-09-19
언어
Python

어떤 프로젝트인가

형태
데이터셋 또는 벤치마크
호스트 에이전트
단독 실행
대상
연구자
성숙도
문서
실측치를 보고합니다

이럴 때 유용합니다

가장 가까운 용도