Home /Benchmarks, evals and calibration /jev-phishing-bench
jev-phishing-bench
Jev against Claude Haiku on 2,000 phishing emails: accuracy, calibration, latency, cost.
From the readme
Jev vs LLM: a phishing decision benchmark with a calibration audit Public, reproducible comparison of Jev (TypeSafe AI's System One model, launched 15 September 2026) against a classic LLM on one security decision: should an email agent click the link in this email? The question this repo answers with numbers: is Jev accurate enough, are its probabilities …
Details
- Section
- Benchmarks, evals and calibration
- owner
- anisselbd
- stars
- 2
- forks
- 0
- pushed
- 2026-09-19
- Language
- Python
What it is
- Form
- Dataset or benchmark
- Host agent
- Standalone
- Audience
- Researchers
Maturity
Docs
reports measured numbers