Home /Benchmarks, evals and calibration /jev-eval by 4esv

jev-eval by 4esv

Jev against GPT-5.6 Terra on three labeled tasks: equal on the easy ones, 6.7 points lower on 77-way routing, 5 times faster, 41 to 50 times cheaper.

From the readme

jev-eval Benchmark TypeSafe Jev against any OpenRouter model on labelled classification data: accuracy, calibration, latency, cost, determinism. Runs on your own data or the three public tasks included; the results below are those three against GPT-5.6 Terra. Results 300 items per task, run 2026-09-17 with jev-1.13.0 and openai/gpt-5.6-terra via OpenRouter. …

Details

Section
Benchmarks, evals and calibration
owner
4esv
stars
1
forks
0
pushed
2026-09-19
Language
Python

What it is

Form
Dataset or benchmark
Host agent
Standalone
Audience
Researchers
Maturity
Docs
reports measured numbers

Useful for

Best intent matches