Home /Benchmarks, evals and calibration /jev-eval by 4esv
jev-eval by 4esv
Jev against GPT-5.6 Terra on three labeled tasks: equal on the easy ones, 6.7 points lower on 77-way routing, 5 times faster, 41 to 50 times cheaper.
From the readme
jev-eval Benchmark TypeSafe Jev against any OpenRouter model on labelled classification data: accuracy, calibration, latency, cost, determinism. Runs on your own data or the three public tasks included; the results below are those three against GPT-5.6 Terra. Results 300 items per task, run 2026-09-17 with jev-1.13.0 and openai/gpt-5.6-terra via OpenRouter. …
Details
- Section
- Benchmarks, evals and calibration
- owner
- 4esv
- stars
- 1
- forks
- 0
- pushed
- 2026-09-19
- Language
- Python
What it is
- Form
- Dataset or benchmark
- Host agent
- Standalone
- Audience
- Researchers
Maturity
Docs
reports measured numbers