Home /Benchmarks, evals and calibration /jev-research-eval

jev-research-eval

Reproducible harness over a pinned jev-ultrafast commit, with baseline and stress suites.

From the readme

jev-research-eval Reproducible evaluation harness for Jared’s Jev Ultrafast research-browser session (17 Sep 2026): 11 baseline cases (R1–R11), human + quant stress suites (S1–S10, QS1–QS8+QS7b), CoS-locked QC grades, the v4 HTML field note, and research notebooks (v1 baseline / v2 baseline+stress) with per-step Trace reports. This is not a fork of Jev. It …

Details

Section
Benchmarks, evals and calibration
owner
jgridifier
stars
2
forks
0
pushed
2026-09-17
Language
HTML
License
NOASSERTION

What it is

Form
Dataset or benchmark
Host agent
Standalone
Audience
Researchers
Maturity
Docs
reports measured numbersmocked or unclear Jev call

Useful for

Best intent matches