Home /Benchmarks, evals and calibration /jev-search-rerank-eval

jev-search-rerank-eval

Does a Jev rerank beat embedding search? 9,831 graded pairs, with the judge-circularity bias measured.

From the readme

jev-search-rerank-eval Does a TypeSafe Jev rerank beat embedding search? Measured, with the judge-bias removed. A graded relevance evaluation over the Agent Skills Hub catalog (33,047 skills, MCP servers and coding-agent tools; snapshot 2026-09-18) with 164 real Chinese / English / mixed queries, 9,831 labelled (query, skill) pairs, and a bake-off of: - …

Details

Section
Benchmarks, evals and calibration
owner
zhuyansen
stars
4
forks
0
pushed
2026-09-18
Language
Python
License
MIT

What it is

Form
Dataset or benchmark
Host agent
Standalone
Audience
Researchers
Maturity
Docs
reports measured numbers

Useful for

Best intent matches