Home / Benchmarks, evals and calibration

Benchmarks, evals and calibration

22 entries

Useful for: Benchmarks and evals

20

jev-benchmarks

Probability-aware evaluation for typed decision models: calibration, selective risk, latency, reproducible.

★ 12PythonDataset or benchmark
Benchmarks and evalsOutput verification

jevcal

Stop guessing thresholds: calibrate, threshold, and drift-check against an LLM teacher.

★ 8PythonCLI
Benchmarks and evalsOutput verification

jev-rerank-bench

Jev against Cohere Rerank, ZeroEntropy, and a chat baseline on 14 datasets, raw responses included.

★ 2PythonDataset or benchmark
Benchmarks and evalsOutput verification

jev-search-rerank-eval

Does a Jev rerank beat embedding search? 9,831 graded pairs, with the judge-circularity bias measured.

★ 4PythonDataset or benchmark
Benchmarks and evalsSearch reranking

jev-sec-bench

Blind benchmarks for prompt injection and vulnerable-code detection.

★ 2GoDataset or benchmark
Benchmarks and evalsOutput verification

jev-phishing-bench

Jev against Claude Haiku on 2,000 phishing emails: accuracy, calibration, latency, cost.

★ 2PythonDataset or benchmark
Benchmarks and evalsOutput verification

jev-spam-eval

Zero-shot spam filtering with Noul questions against TF-IDF baselines.

★ 0Jupyter NotebookDataset or benchmark
Benchmarks and evalsPlaygrounds and demos

jev-agent-failure-benchmark

Jev against a strong LLM on the Who and When agent-failure-attribution benchmark.

★ 1PythonDataset or benchmark
Benchmarks and evalsOutput verification

jev-korean-benchmark

Korean understanding and medical text, with runtime and cost evidence.

★ 5PythonDataset or benchmark
Benchmarks and evalsOutput verification

jev-report

Independent Chinese research report: 52 pages, 50 reproducible tests, 143 traceable data rows.

★ 0PythonDataset or benchmark
Benchmarks and evalsOutput verification

jev-benchmark

Two benchmarks, chess and predator identification, one inside Jev's lane and one outside, both with results.

★ 2PythonDataset or benchmark
Benchmarks and evalsPlaygrounds and demos

jev-dspy-lab

Reproducible calibration and selective-risk benchmarks for Jev decisions in DSPy.

★ 0PythonDataset or benchmark
Benchmarks and evalsOutput verification

jev-eval-agent

Personal-assistant agent with 100 mocked tools, measuring how many steps a Jev-gated agent needs.

★ 95HTMLDataset or benchmark
Benchmarks and evalsTool gating and guardrails

jev-synergy-screening

Choice and Noul questions scored against ASReview SYNERGY gold labels for abstract screening.

★ 1PythonDataset or benchmark
Benchmarks and evalsOutput verification

jev-orderby-bench

Measures whether ORDER BY over a Jev probability is defensible: pairwise inversion, Score ordinality against a human grade, calibration, and wording invariants under a pre-registered gate; passes on 20 Newsgroups topics, fails four of six conditions on Amazon ESCI product relevance, and shows that a 40-row batched state through a DuckDB extension fails the ranking gate one row per request passes.

★ 0PythonDataset or benchmark
Benchmarks and evalsOutput verification

cultivar

Pinecone's skill-testing CLI, with a Jev grading backend it reports at about 30 times cheaper than the LLM grader.

★ 40PythonCLI
Benchmarks and evalsOutput verification

jev-eval by 4esv

Jev against GPT-5.6 Terra on three labeled tasks: equal on the easy ones, 6.7 points lower on 77-way routing, 5 times faster, 41 to 50 times cheaper.

★ 1PythonDataset or benchmark
Benchmarks and evalsOutput verification

jev-benchmark by themsquared

Tool-call risk classification with the run-to-run variance reported; every wrong answer came with hedged confidence.

★ 0PythonDataset or benchmark
Benchmarks and evalsOutput verification

jev-research-eval

Reproducible harness over a pinned jev-ultrafast commit, with baseline and stress suites.

★ 2HTMLDataset or benchmark
Benchmarks and evalsOutput verification

jev-playground by hegargarcia

Jev against other models in games with explicit states, legal actions, and a measurable outcome.

★ 0TypeScriptDataset or benchmark
Benchmarks and evalsOutput verification

Useful for: Output verification

1

jev-harness

Confidence gates, shadow mode, recipes, and evals; reports Claude CLI at 48.9 s against Jev at 1.3 s on the same row-filter job.

★ 5TypeScriptLibrary or SDK
Output verificationBenchmarks and evals

Useful for: Playgrounds and demos

1

jev-behavior-study

Controlled prompt experiments on jev-1.13.0, raw results and offline verification.

★ 3PythonDataset or benchmark
Playgrounds and demosBenchmarks and evals