首页 /基准、评测与校准 /jev-benchmarks

jev-benchmarks

面向带类型决策模型的概率感知评估:校准、选择性风险、延迟,可复现。

本页由英文清单自动生成。

来自 readme

jev-benchmarks Probability-aware evaluation for typed decision models. jev-benchmarks measures more than whether a model selects the right label. It evaluates whether the reported probabilities are calibrated enough to support automation, how much work can be accepted at a fixed error budget, what resources each decision uses, and how long it takes end to …

详情

分区
基准、评测与校准
owner
AbdelStark
星标
12
复刻
0
最近提交
2026-09-17
语言
Python
许可证
Apache-2.0

这是什么

形态
数据集或基准
宿主智能体
独立运行
面向人群
研究人员
成熟度
文档
报告了实测数据

适合用来

最匹配的意图