Home /Benchmarks, evals and calibration /jev-playground by hegargarcia
jev-playground by hegargarcia
Jev against other models in games with explicit states, legal actions, and a measurable outcome.
From the readme
Jev Playground A playground for benchmarking TypeSafe AI’s Jev against other evaluation models in games with explicit states, legal actions, and measurable outcomes. The core question: how well does each model choose the next action when the rules and available choices are clearly defined? Games provide a small, inspectable environment for exploring …
Details
- Section
- Benchmarks, evals and calibration
- owner
- hegargarcia
- stars
- 0
- forks
- 0
- pushed
- 2026-09-17
- Language
- TypeScript
What it is
- Form
- Dataset or benchmark
- Host agent
- Standalone
- Audience
- Researchers
Maturity
Docs