Home /Benchmarks, evals and calibration /jev-behavior-study

jev-behavior-study

Controlled prompt experiments on jev-1.13.0, raw results and offline verification.

From the readme

How does Jev behave? A browsable field guide to Jev 1.13.0: where explicit questions work, where small changes alter answers, and where harder tasks expose failures. 11,621 text-study requests · 3 Snake studies · 3D City lab · 7 detailed reports Live API observations from September 16–17, 2026 (UTC). Independent, AI-assisted research. Explore findings · How …

Details

Section
Benchmarks, evals and calibration
owner
RINNECODER
stars
3
forks
0
pushed
2026-09-17
Language
Python
License
MIT

What it is

Form
Dataset or benchmark
Host agent
Standalone
Audience
Researchers
Maturity
Docs
reports measured numbers

Useful for

Best intent matches