首页 /基准、评测与校准 /jev-phishing-bench
jev-phishing-bench
在 2,000 封钓鱼邮件上把 Jev 与 Claude Haiku 作对比:准确率、校准、延迟、成本。
本页由英文清单自动生成。
来自 readme
Jev vs LLM: a phishing decision benchmark with a calibration audit Public, reproducible comparison of Jev (TypeSafe AI's System One model, launched 15 September 2026) against a classic LLM on one security decision: should an email agent click the link in this email? The question this repo answers with numbers: is Jev accurate enough, are its probabilities …
这是什么
- 形态
- 数据集或基准
- 宿主智能体
- 独立运行
- 面向人群
- 研究人员
成熟度
文档
报告了实测数据