首页 /基准、评测与校准 /jev-phishing-bench

jev-phishing-bench

在 2,000 封钓鱼邮件上把 Jev 与 Claude Haiku 作对比:准确率、校准、延迟、成本。

本页由英文清单自动生成。

来自 readme

Jev vs LLM: a phishing decision benchmark with a calibration audit Public, reproducible comparison of Jev (TypeSafe AI's System One model, launched 15 September 2026) against a classic LLM on one security decision: should an email agent click the link in this email? The question this repo answers with numbers: is Jev accurate enough, are its probabilities …

详情

分区
基准、评测与校准
owner
anisselbd
星标
2
复刻
0
最近提交
2026-09-19
语言
Python

这是什么

形态
数据集或基准
宿主智能体
独立运行
面向人群
研究人员
成熟度
文档
报告了实测数据

适合用来

最匹配的意图