首页 /基准、评测与校准 /jev-benchmark by themsquared
jev-benchmark by themsquared
工具调用风险分类,附多次运行的方差;每一个错误答案都伴随着含糊的置信度。
本页由英文清单自动生成。
来自 readme
jev-benchmark 📖 Read the write-up: I Benchmarked Jev on Agent Tool-Call Risk. Calibration Held. A reproducible benchmark for TypeSafe AI's Jev on a real agent-infrastructure task: classifying an agent tool call as readonly, destructive, privileged, or exfiltration. Everything here is runnable. The task set is published, the harness is 200 lines, and the …
详情
- 分区
- 基准、评测与校准
- owner
- themsquared
- 星标
- 0
- 复刻
- 0
- 最近提交
- 2026-09-19
- 语言
- Python
- 许可证
- Apache-2.0
这是什么
- 形态
- 数据集或基准
- 宿主智能体
- 独立运行
- 面向人群
- 研究人员
成熟度
文档
报告了实测数据