/벤치마크, 평가, 보정 /jev-rerank-bench

jev-rerank-bench

데이터셋 14개에서 Jev를 Cohere Rerank, ZeroEntropy, 챗 기준선과 비교하며, 원본 응답도 포함합니다.

이 페이지는 영문 목록에서 생성했습니다.

readme에서

jev-rerank-bench I gave TypeSafe's Jev thirty search results and asked it which ones were useful. Then I gave Cohere and ZeroEntropy the same passages. This repository contains the experiments, saved responses and scoring code. The ranking average put Jev's rubric at 0.692 and Cohere Pro at 0.691, without establishing a winner. Giving every query equal …

상세

섹션
벤치마크, 평가, 보정
owner
anessbelbati
스타
2
포크
0
최근 커밋
2026-09-17
언어
Python
라이선스
MIT

어떤 프로젝트인가

형태
데이터셋 또는 벤치마크
호스트 에이전트
단독 실행
대상
연구자
성숙도
문서
실측치를 보고합니다

이럴 때 유용합니다

가장 가까운 용도