Home /Benchmarks, evals and calibration /jev-search-rerank-eval
jev-search-rerank-eval
Does a Jev rerank beat embedding search? 9,831 graded pairs, with the judge-circularity bias measured.
From the readme
jev-search-rerank-eval Does a TypeSafe Jev rerank beat embedding search? Measured, with the judge-bias removed. A graded relevance evaluation over the Agent Skills Hub catalog (33,047 skills, MCP servers and coding-agent tools; snapshot 2026-09-18) with 164 real Chinese / English / mixed queries, 9,831 labelled (query, skill) pairs, and a bake-off of: - …
Details
- Section
- Benchmarks, evals and calibration
- owner
- zhuyansen
- stars
- 4
- forks
- 0
- pushed
- 2026-09-18
- Language
- Python
- License
- MIT
What it is
- Form
- Dataset or benchmark
- Host agent
- Standalone
- Audience
- Researchers
Maturity
Docs
reports measured numbers