ホーム /ベンチマーク・評価・較正 /jev-phishing-bench
jev-phishing-bench
2,000 通のフィッシングメールで Jev と Claude Haiku を比較し、精度、較正、レイテンシ、コストを測ります。
このページは英語版リストから生成しています。
readme より
Jev vs LLM: a phishing decision benchmark with a calibration audit Public, reproducible comparison of Jev (TypeSafe AI's System One model, launched 15 September 2026) against a classic LLM on one security decision: should an email agent click the link in this email? The question this repo answers with numbers: is Jev accurate enough, are its probabilities …
詳細
- セクション
- ベンチマーク・評価・較正
- owner
- anisselbd
- スター
- 2
- フォーク
- 0
- 最終更新
- 2026-09-19
- 言語
- Python
これは何か
- 形態
- データセットやベンチマーク
- ホストエージェント
- 単体で動作
- 対象
- 研究者
成熟度
ドキュメント
実測値を報告しています