/벤치마크, 평가, 보정 /jev-eval by 4esv

jev-eval by 4esv

라벨된 세 가지 과제에서 Jev와 GPT-5.6 Terra를 비교합니다: 쉬운 과제에서는 동등하고, 77지 라우팅에서는 6.7포인트 낮으며, 5배 빠르고 41~50배 저렴합니다.

이 페이지는 영문 목록에서 생성했습니다.

readme에서

jev-eval Benchmark TypeSafe Jev against any OpenRouter model on labelled classification data: accuracy, calibration, latency, cost, determinism. Runs on your own data or the three public tasks included; the results below are those three against GPT-5.6 Terra. Results 300 items per task, run 2026-09-17 with jev-1.13.0 and openai/gpt-5.6-terra via OpenRouter. …

상세

섹션
벤치마크, 평가, 보정
owner
4esv
스타
1
포크
0
최근 커밋
2026-09-19
언어
Python

어떤 프로젝트인가

형태
데이터셋 또는 벤치마크
호스트 에이전트
단독 실행
대상
연구자
성숙도
문서
실측치를 보고합니다

이럴 때 유용합니다

가장 가까운 용도