Home /Benchmarks, evals and calibration /jev-rerank-bench
jev-rerank-bench
Jev against Cohere Rerank, ZeroEntropy, and a chat baseline on 14 datasets, raw responses included.
From the readme
jev-rerank-bench I gave TypeSafe's Jev thirty search results and asked it which ones were useful. Then I gave Cohere and ZeroEntropy the same passages. This repository contains the experiments, saved responses and scoring code. The ranking average put Jev's rubric at 0.692 and Cohere Pro at 0.691, without establishing a winner. Giving every query equal …
Details
- Section
- Benchmarks, evals and calibration
- owner
- anessbelbati
- stars
- 2
- forks
- 0
- pushed
- 2026-09-17
- Language
- Python
- License
- MIT
What it is
- Form
- Dataset or benchmark
- Host agent
- Standalone
- Audience
- Researchers
Maturity
Docs
reports measured numbers