홈 /벤치마크, 평가, 보정 /jev-eval by 4esv
jev-eval by 4esv
라벨된 세 가지 과제에서 Jev와 GPT-5.6 Terra를 비교합니다: 쉬운 과제에서는 동등하고, 77지 라우팅에서는 6.7포인트 낮으며, 5배 빠르고 41~50배 저렴합니다.
이 페이지는 영문 목록에서 생성했습니다.
readme에서
jev-eval Benchmark TypeSafe Jev against any OpenRouter model on labelled classification data: accuracy, calibration, latency, cost, determinism. Runs on your own data or the three public tasks included; the results below are those three against GPT-5.6 Terra. Results 300 items per task, run 2026-09-17 with jev-1.13.0 and openai/gpt-5.6-terra via OpenRouter. …
상세
- 섹션
- 벤치마크, 평가, 보정
- owner
- 4esv
- 스타
- 1
- 포크
- 0
- 최근 커밋
- 2026-09-19
- 언어
- Python
어떤 프로젝트인가
- 형태
- 데이터셋 또는 벤치마크
- 호스트 에이전트
- 단독 실행
- 대상
- 연구자
성숙도
문서
실측치를 보고합니다