ホーム /ベンチマーク・評価・較正 /jev-benchmarks

jev-benchmarks

型付き判断モデル向けの確率を考慮した評価で、較正、選択的リスク、レイテンシを再現可能な形で測ります。

このページは英語版リストから生成しています。

readme より

jev-benchmarks Probability-aware evaluation for typed decision models. jev-benchmarks measures more than whether a model selects the right label. It evaluates whether the reported probabilities are calibrated enough to support automation, how much work can be accepted at a fixed error budget, what resources each decision uses, and how long it takes end to …

詳細

セクション
ベンチマーク・評価・較正
owner
AbdelStark
スター
12
フォーク
0
最終更新
2026-09-17
言語
Python
ライセンス
Apache-2.0

これは何か

形態
データセットやベンチマーク
ホストエージェント
単体で動作
対象
研究者
成熟度
ドキュメント
実測値を報告しています

こんなときに役立ちます

近い用途