ホーム /ベンチマーク・評価・較正 /jev-benchmark by themsquared

jev-benchmark by themsquared

ツール呼び出しのリスク分類を実行ごとの分散付きで報告し、誤答はいずれも confidence が低めに出ていました。

このページは英語版リストから生成しています。

readme より

jev-benchmark 📖 Read the write-up: I Benchmarked Jev on Agent Tool-Call Risk. Calibration Held. A reproducible benchmark for TypeSafe AI's Jev on a real agent-infrastructure task: classifying an agent tool call as readonly, destructive, privileged, or exfiltration. Everything here is runnable. The task set is published, the harness is 200 lines, and the …

詳細

セクション
ベンチマーク・評価・較正
owner
themsquared
スター
0
フォーク
0
最終更新
2026-09-19
言語
Python
ライセンス
Apache-2.0

これは何か

形態
データセットやベンチマーク
ホストエージェント
単体で動作
対象
研究者
成熟度
ドキュメント
実測値を報告しています

こんなときに役立ちます

近い用途