Home /Benchmarks, evals and calibration /jev-benchmark by themsquared
jev-benchmark by themsquared
Tool-call risk classification with the run-to-run variance reported; every wrong answer came with hedged confidence.
From the readme
jev-benchmark 📖 Read the write-up: I Benchmarked Jev on Agent Tool-Call Risk. Calibration Held. A reproducible benchmark for TypeSafe AI's Jev on a real agent-infrastructure task: classifying an agent tool call as readonly, destructive, privileged, or exfiltration. Everything here is runnable. The task set is published, the harness is 200 lines, and the …
Details
- Section
- Benchmarks, evals and calibration
- owner
- themsquared
- stars
- 0
- forks
- 0
- pushed
- 2026-09-19
- Language
- Python
- License
- Apache-2.0
What it is
- Form
- Dataset or benchmark
- Host agent
- Standalone
- Audience
- Researchers
Maturity
Docs
reports measured numbers