Introducing System One Models and Jev
The launch post: architecture, RLCD, benchmarks, pricing, and the FAQ.
Home / Articles and talks
33 entries
The launch post: architecture, RLCD, benchmarks, pricing, and the FAQ.
The founder's case that RLCD-trained decision models are a shorter path to value than chat.
Where the sceptical reading of the benchmarks lives.
Launch coverage, the Doom demo, and the $40M seed round.
Launch-day roundup.
The funding announcement.
Short launch summary.
Third-party explainer of the primitives, pricing, and vendor evals.
What is public about the training method.
Release-week technical roundup: API, evals, adapter, and skill.
Two days in: demand briefly took the API down, and Almeida on not wanting to be a frontier lab.
The business angle, with the reported valuation and Every's speed and cost numbers.
16,000 calls against two GPT models: where it wins, where it breaks, and a threshold procedure.
Three classification tasks, one direct question against a dozen scored dimensions with fitted weights.
Head to head on local event listings, with cost and latency.
Every runs 1,709 judgments over a writing archive for under a cent.
Legal-move Choices land it next to reasoning models.
Which launch claims survive a reading of the primary sources.
Separates the published claims from what public evidence establishes.
Japanese; reproduces the logit shortcut on Gemma and compares against LLMs on the Mario harness.
Pre-registered test of about 9,750 calls: calibration error by question type, Jev ahead on commit classification and behind on knowledge-base filing, abstention closing the gap.
What is known, what is guessed, and what it is good for, with the independent numbers pulled together.
LangChain on model routing and gating dangerous tool calls behind a typed decision.
Triage, RAG filtering, citation checks, and confidence gates, with code.
An essay on what a calibrated, non-generating model is for.
Guillermo Rauch on Jev reviewing every fx command.
Steve Krouse's sixteen-judgment demo and video.
Walkthrough of the Vercel AI SDK integration.
Builders using Jev as a tool-use safety layer.
Japanese; Jev against Jev at gomoku, with source and timing logs.
A 10,000-call probe that reconstructs a shared-state, parallel-branch architecture; kev above is built from it.
Whether reading logits directly is new at all, argued at length.
A prior-art claim for non-autoregressive typed decisions, and the counter that zero-shot generality is the actual difference.