Home /Benchmarks, evals and calibration /jev-eval-agent
jev-eval-agent
Personal-assistant agent with 100 mocked tools, measuring how many steps a Jev-gated agent needs.
From the readme
jev-eval-agent 🇺🇸 English · 🇧🇷 Leia em português A personal-assistant agent built with eve (Vercel), with 100 mocked tools, served through OpenRouter. The repository exists to answer one question: how many steps does the agent need to finish the same task when the LLM picks the tool itself vs. when Jev (TypeSafe's classifier) picks it? AGENTMODE Who …
Details
- Section
- Benchmarks, evals and calibration
- owner
- vinilana
- stars
- 95
- forks
- 10
- pushed
- 2026-09-17
- Language
- HTML
What it is
- Form
- Dataset or benchmark
- Host agent
- Standalone
- Audience
- Researchers
Maturity
Docs