Home /Benchmarks, evals and calibration /jev-eval-agent

jev-eval-agent

Personal-assistant agent with 100 mocked tools, measuring how many steps a Jev-gated agent needs.

From the readme

jev-eval-agent 🇺🇸 English · 🇧🇷 Leia em português A personal-assistant agent built with eve (Vercel), with 100 mocked tools, served through OpenRouter. The repository exists to answer one question: how many steps does the agent need to finish the same task when the LLM picks the tool itself vs. when Jev (TypeSafe's classifier) picks it? AGENTMODE Who …

Details

Section
Benchmarks, evals and calibration
owner
vinilana
stars
95
forks
10
pushed
2026-09-17
Language
HTML

What it is

Form
Dataset or benchmark
Host agent
Standalone
Audience
Researchers
Maturity
Docs

Useful for

Best intent matches