Home /Benchmarks, evals and calibration /jev-phishing-bench

jev-phishing-bench

Jev against Claude Haiku on 2,000 phishing emails: accuracy, calibration, latency, cost.

From the readme

Jev vs LLM: a phishing decision benchmark with a calibration audit Public, reproducible comparison of Jev (TypeSafe AI's System One model, launched 15 September 2026) against a classic LLM on one security decision: should an email agent click the link in this email? The question this repo answers with numbers: is Jev accurate enough, are its probabilities …

Details

Section
Benchmarks, evals and calibration
owner
anisselbd
stars
2
forks
0
pushed
2026-09-19
Language
Python

What it is

Form
Dataset or benchmark
Host agent
Standalone
Audience
Researchers
Maturity
Docs
reports measured numbers

Useful for

Best intent matches