Home /Benchmarks, evals and calibration /jev-benchmark

jev-benchmark

Two benchmarks, chess and predator identification, one inside Jev's lane and one outside, both with results.

From the readme

Jev Benchmark & Playground Two hands-on benchmarks of Jev, the "System One" model from TypeSafe. Jev doesn't generate text or reason step by step. You hand it state (a JSON blob) and typed questions (yes/no, pick-one, or rate-on-a-scale) and it returns calibrated probabilities in about 200 ms. The pitch is "programmable common sense": code owns the …

Details

Section
Benchmarks, evals and calibration
owner
wondertwins
stars
2
forks
1
pushed
2026-09-16
Language
Python
License
MIT

What it is

Form
Dataset or benchmark
Host agent
Standalone
Audience
Researchers
Maturity
Docs
reports measured numbers

Useful for

Best intent matches