Home /Benchmarks, evals and calibration /jev-benchmark by themsquared

jev-benchmark by themsquared

Tool-call risk classification with the run-to-run variance reported; every wrong answer came with hedged confidence.

From the readme

jev-benchmark 📖 Read the write-up: I Benchmarked Jev on Agent Tool-Call Risk. Calibration Held. A reproducible benchmark for TypeSafe AI's Jev on a real agent-infrastructure task: classifying an agent tool call as readonly, destructive, privileged, or exfiltration. Everything here is runnable. The task set is published, the harness is 200 lines, and the …

Details

Section
Benchmarks, evals and calibration
owner
themsquared
stars
0
forks
0
pushed
2026-09-19
Language
Python
License
Apache-2.0

What it is

Form
Dataset or benchmark
Host agent
Standalone
Audience
Researchers
Maturity
Docs
reports measured numbers

Useful for

Best intent matches