Home /Benchmarks, evals and calibration /jev-dspy-lab

jev-dspy-lab

Reproducible calibration and selective-risk benchmarks for Jev decisions in DSPy.

From the readme

Jev DSPy Lab Reproducible calibration, confidence-gating, latency, and modeled-cost benchmarks for Jev / TypeSafe System One decisions used in DSPy workflows. This is not another DSPy fork. It is a companion measurement lab for typesafeainate/dspy-typesafeify and other callers that use TypeSafe decisions inside DSPy-style pipelines. Why this exists The …

Details

Section
Benchmarks, evals and calibration
owner
jmanhype
stars
0
forks
0
pushed
2026-09-17
Language
Python
License
MIT

What it is

Form
Dataset or benchmark
Host agent
Standalone
Audience
Researchers
Maturity
Docs
reports measured numbers

Useful for

Best intent matches