Home /Benchmarks, evals and calibration /jev-rerank-bench

jev-rerank-bench

Jev against Cohere Rerank, ZeroEntropy, and a chat baseline on 14 datasets, raw responses included.

From the readme

jev-rerank-bench I gave TypeSafe's Jev thirty search results and asked it which ones were useful. Then I gave Cohere and ZeroEntropy the same passages. This repository contains the experiments, saved responses and scoring code. The ranking average put Jev's rubric at 0.692 and Cohere Pro at 0.691, without establishing a winner. Giving every query equal …

Details

Section
Benchmarks, evals and calibration
owner
anessbelbati
stars
2
forks
0
pushed
2026-09-17
Language
Python
License
MIT

What it is

Form
Dataset or benchmark
Host agent
Standalone
Audience
Researchers
Maturity
Docs
reports measured numbers

Useful for

Best intent matches