ホーム /ベンチマーク・評価・較正 /jev-phishing-bench

jev-phishing-bench

2,000 通のフィッシングメールで Jev と Claude Haiku を比較し、精度、較正、レイテンシ、コストを測ります。

このページは英語版リストから生成しています。

readme より

Jev vs LLM: a phishing decision benchmark with a calibration audit Public, reproducible comparison of Jev (TypeSafe AI's System One model, launched 15 September 2026) against a classic LLM on one security decision: should an email agent click the link in this email? The question this repo answers with numbers: is Jev accurate enough, are its probabilities …

詳細

セクション
ベンチマーク・評価・較正
owner
anisselbd
スター
2
フォーク
0
最終更新
2026-09-19
言語
Python

これは何か

形態
データセットやベンチマーク
ホストエージェント
単体で動作
対象
研究者
成熟度
ドキュメント
実測値を報告しています

こんなときに役立ちます

近い用途