홈 /벤치마크, 평가, 보정 /jev-benchmarks
jev-benchmarks
타입 지정 결정 모델을 위한 확률 인지 평가입니다: 보정, 선택적 위험, 지연 시간을 재현 가능하게 다룹니다.
이 페이지는 영문 목록에서 생성했습니다.
readme에서
jev-benchmarks Probability-aware evaluation for typed decision models. jev-benchmarks measures more than whether a model selects the right label. It evaluates whether the reported probabilities are calibrated enough to support automation, how much work can be accepted at a fixed error budget, what resources each decision uses, and how long it takes end to …
상세
- 섹션
- 벤치마크, 평가, 보정
- owner
- AbdelStark
- 스타
- 12
- 포크
- 0
- 최근 커밋
- 2026-09-17
- 언어
- Python
- 라이선스
- Apache-2.0
어떤 프로젝트인가
- 형태
- 데이터셋 또는 벤치마크
- 호스트 에이전트
- 단독 실행
- 대상
- 연구자
성숙도
문서
실측치를 보고합니다