Home /Benchmarks, evals and calibration /jev-orderby-bench
jev-orderby-bench
Measures whether ORDER BY over a Jev probability is defensible: pairwise inversion, Score ordinality against a human grade, calibration, and wording invariants under a pre-registered gate; passes on 20 Newsgroups topics, fails four of six conditions on Amazon ESCI product relevance, and shows that a 40-row batched state through a DuckDB extension fails the ranking gate one row per request passes.
From the readme
jev-orderby-bench Does ORDER BY over a Jev probability put rows in a defensible order? An independent measurement of TypeSafe AI's Jev (jev-1.13.0) on the properties a semantic sort actually depends on: pairwise inversion rate, Score ordinality against a graded target, and whether the probabilities move with evidence or with wording. Calibration (ECE, …
Details
- Section
- Benchmarks, evals and calibration
- owner
- yodablocks
- stars
- 0
- forks
- 0
- pushed
- 2026-09-20
- Language
- Python
- License
- MIT
- claim
- pre-registered
What it is
- Form
- Dataset or benchmark
- Host agent
- Standalone
- Audience
- Researchers