首页 /基准、评测与校准 /jev-eval by 4esv

jev-eval by 4esv

在三项带标注的任务上把 Jev 与 GPT-5.6 Terra 作对比:简单任务打平,77 路路由低 6.7 分,速度快 5 倍,成本低 41 到 50 倍。

本页由英文清单自动生成。

来自 readme

jev-eval Benchmark TypeSafe Jev against any OpenRouter model on labelled classification data: accuracy, calibration, latency, cost, determinism. Runs on your own data or the three public tasks included; the results below are those three against GPT-5.6 Terra. Results 300 items per task, run 2026-09-17 with jev-1.13.0 and openai/gpt-5.6-terra via OpenRouter. …

详情

分区
基准、评测与校准
owner
4esv
星标
1
复刻
0
最近提交
2026-09-19
语言
Python

这是什么

形态
数据集或基准
宿主智能体
独立运行
面向人群
研究人员
成熟度
文档
报告了实测数据

适合用来

最匹配的意图