jev-benchmarks
面向带类型决策模型的概率感知评估:校准、选择性风险、延迟,可复现。
本页由英文清单自动生成。
来自 readme
jev-benchmarks Probability-aware evaluation for typed decision models. jev-benchmarks measures more than whether a model selects the right label. It evaluates whether the reported probabilities are calibrated enough to support automation, how much work can be accepted at a fixed error budget, what resources each decision uses, and how long it takes end to …
详情
- 分区
- 基准、评测与校准
- owner
- AbdelStark
- 星标
- 12
- 复刻
- 0
- 最近提交
- 2026-09-17
- 语言
- Python
- 许可证
- Apache-2.0
这是什么
- 形态
- 数据集或基准
- 宿主智能体
- 独立运行
- 面向人群
- 研究人员
成熟度
文档
报告了实测数据