Home /Benchmarks, evals and calibration /jev-benchmarks

jev-benchmarks

Probability-aware evaluation for typed decision models: calibration, selective risk, latency, reproducible.

From the readme

jev-benchmarks Probability-aware evaluation for typed decision models. jev-benchmarks measures more than whether a model selects the right label. It evaluates whether the reported probabilities are calibrated enough to support automation, how much work can be accepted at a fixed error budget, what resources each decision uses, and how long it takes end to …

Details

Section
Benchmarks, evals and calibration
owner
AbdelStark
stars
12
forks
0
pushed
2026-09-17
Language
Python
License
Apache-2.0

What it is

Form
Dataset or benchmark
Host agent
Standalone
Audience
Researchers
Maturity
Docs
reports measured numbers

Useful for

Best intent matches