jev-benchmark
两项基准测试,国际象棋和猛兽辨识,一项在 Jev 的能力区内、一项在区外,都附结果。
本页由英文清单自动生成。
来自 readme
Jev Benchmark & Playground Two hands-on benchmarks of Jev, the "System One" model from TypeSafe. Jev doesn't generate text or reason step by step. You hand it state (a JSON blob) and typed questions (yes/no, pick-one, or rate-on-a-scale) and it returns calibrated probabilities in about 200 ms. The pitch is "programmable common sense": code owns the …
详情
- 分区
- 基准、评测与校准
- owner
- wondertwins
- 星标
- 2
- 复刻
- 1
- 最近提交
- 2026-09-16
- 语言
- Python
- 许可证
- MIT
这是什么
- 形态
- 数据集或基准
- 宿主智能体
- 独立运行
- 面向人群
- 研究人员
成熟度
文档
报告了实测数据