ホーム /ベンチマーク・評価・較正 /jev-eval-agent
jev-eval-agent
100 個のモックツールを持つパーソナルアシスタントのエージェントで、Jev でゲートしたエージェントが何ステップを要するかを測定します。
このページは英語版リストから生成しています。
readme より
jev-eval-agent 🇺🇸 English · 🇧🇷 Leia em português A personal-assistant agent built with eve (Vercel), with 100 mocked tools, served through OpenRouter. The repository exists to answer one question: how many steps does the agent need to finish the same task when the LLM picks the tool itself vs. when Jev (TypeSafe's classifier) picks it? AGENTMODE Who …
詳細
- セクション
- ベンチマーク・評価・較正
- owner
- vinilana
- スター
- 95
- フォーク
- 10
- 最終更新
- 2026-09-17
- 言語
- HTML
これは何か
- 形態
- データセットやベンチマーク
- ホストエージェント
- 単体で動作
- 対象
- 研究者
成熟度
ドキュメント