ホーム /ベンチマーク・評価・較正 /jev-benchmarks
jev-benchmarks
型付き判断モデル向けの確率を考慮した評価で、較正、選択的リスク、レイテンシを再現可能な形で測ります。
このページは英語版リストから生成しています。
readme より
jev-benchmarks Probability-aware evaluation for typed decision models. jev-benchmarks measures more than whether a model selects the right label. It evaluates whether the reported probabilities are calibrated enough to support automation, how much work can be accepted at a fixed error budget, what resources each decision uses, and how long it takes end to …
詳細
- セクション
- ベンチマーク・評価・較正
- owner
- AbdelStark
- スター
- 12
- フォーク
- 0
- 最終更新
- 2026-09-17
- 言語
- Python
- ライセンス
- Apache-2.0
これは何か
- 形態
- データセットやベンチマーク
- ホストエージェント
- 単体で動作
- 対象
- 研究者
成熟度
ドキュメント
実測値を報告しています