ホーム /ベンチマーク・評価・較正 /jev-benchmark by themsquared
jev-benchmark by themsquared
ツール呼び出しのリスク分類を実行ごとの分散付きで報告し、誤答はいずれも confidence が低めに出ていました。
このページは英語版リストから生成しています。
readme より
jev-benchmark 📖 Read the write-up: I Benchmarked Jev on Agent Tool-Call Risk. Calibration Held. A reproducible benchmark for TypeSafe AI's Jev on a real agent-infrastructure task: classifying an agent tool call as readonly, destructive, privileged, or exfiltration. Everything here is runnable. The task set is published, the harness is 200 lines, and the …
詳細
- セクション
- ベンチマーク・評価・較正
- owner
- themsquared
- スター
- 0
- フォーク
- 0
- 最終更新
- 2026-09-19
- 言語
- Python
- ライセンス
- Apache-2.0
これは何か
- 形態
- データセットやベンチマーク
- ホストエージェント
- 単体で動作
- 対象
- 研究者
成熟度
ドキュメント
実測値を報告しています