Home /Benchmarks, evals and calibration /jev-playground by hegargarcia

jev-playground by hegargarcia

Jev against other models in games with explicit states, legal actions, and a measurable outcome.

From the readme

Jev Playground A playground for benchmarking TypeSafe AI’s Jev against other evaluation models in games with explicit states, legal actions, and measurable outcomes. The core question: how well does each model choose the next action when the rules and available choices are clearly defined? Games provide a small, inspectable environment for exploring …

Details

Section
Benchmarks, evals and calibration
owner
hegargarcia
stars
0
forks
0
pushed
2026-09-17
Language
TypeScript

What it is

Form
Dataset or benchmark
Host agent
Standalone
Audience
Researchers
Maturity
Docs

Useful for

Best intent matches