Home /Benchmarks, evals and calibration /jev-orderby-bench

jev-orderby-bench

Measures whether ORDER BY over a Jev probability is defensible: pairwise inversion, Score ordinality against a human grade, calibration, and wording invariants under a pre-registered gate; passes on 20 Newsgroups topics, fails four of six conditions on Amazon ESCI product relevance, and shows that a 40-row batched state through a DuckDB extension fails the ranking gate one row per request passes.

From the readme

jev-orderby-bench Does ORDER BY over a Jev probability put rows in a defensible order? An independent measurement of TypeSafe AI's Jev (jev-1.13.0) on the properties a semantic sort actually depends on: pairwise inversion rate, Score ordinality against a graded target, and whether the probabilities move with evidence or with wording. Calibration (ECE, …

Details

Section
Benchmarks, evals and calibration
owner
yodablocks
stars
0
forks
0
pushed
2026-09-20
Language
Python
License
MIT
claim
pre-registered

What it is

Form
Dataset or benchmark
Host agent
Standalone
Audience
Researchers
Maturity
Docs
reports measured numbers

Useful for

Best intent matches