首页 /基准、评测与校准 /jev-benchmark by themsquared

jev-benchmark by themsquared

工具调用风险分类,附多次运行的方差;每一个错误答案都伴随着含糊的置信度。

本页由英文清单自动生成。

来自 readme

jev-benchmark 📖 Read the write-up: I Benchmarked Jev on Agent Tool-Call Risk. Calibration Held. A reproducible benchmark for TypeSafe AI's Jev on a real agent-infrastructure task: classifying an agent tool call as readonly, destructive, privileged, or exfiltration. Everything here is runnable. The task set is published, the harness is 200 lines, and the …

详情

分区
基准、评测与校准
owner
themsquared
星标
0
复刻
0
最近提交
2026-09-19
语言
Python
许可证
Apache-2.0

这是什么

形态
数据集或基准
宿主智能体
独立运行
面向人群
研究人员
成熟度
文档
报告了实测数据

适合用来

最匹配的意图