The clone wars have begun.
It’s a former OpenAI veteran who says so. He now runs TypeSafe AI. This startup has garnered quite a buzz since it emerged from stealth in mid-September. The product that earned it attention is called Jev. A nod to British economist William Stanley Jevons (1835-1882). He theorized the idea that the more technological progress enables us to use a resource efficiently, the more consumption of that resource tends to rise.
There is indeed in Jev the promise of more efficient use… of tokens, specifically. Through it, TypeSafe AI pushes a new typology of models, perched between LLMs and classifiers. Called System One, it does not produce text, but “calibrated decisions.” More precisely, structured answers to three types of framed questions:
- Select options from a finite list
- Distribute probability scores over a defined sample
- Assess whether an item is true or false
Each question has its primitive (respectively Choice, Score and Noul). They can be combined in a single call. They are processed in parallel, but isolated to avoid context contamination.
More reliable, faster and cheaper than LLMs… within a limited operating scope
The name TypeSafe AI alludes to type safety, a property that guarantees that a program does not misuse data relative to their type (for example, trying to multiply a price by a date). It reflects the promise of the System One approach: to obtain answers that machines can directly use… and handle more simply than LLM outputs, with greater reliability (thanks to embedded confidence scores), lower latency and a lower cost (for Jev, “under 100 ms for most queries,” at $0.042 per million input tokens).
The finite-list selection scenario, for example, can help route a customer request to the right team. The scoring of probability measures their frustration level. The true/false evaluation, namely whether the request involves a refund. In an agentic system, the whole set can serve both request routing and tool selection, argument verification or the application of guardrails.
System One models are, in any case, suited to rapid responses — the kind of quick answers a seasoned human would instantly provide, TypeSafe AI explains. They will eventually be paired with deterministic checks in code.
AWS responds with a dedicated model; OpenAI with GPT Luna
The approach inspired AWS, which published its own incarnation: Strands Decider. It largely advances the same arguments as TypeSafe AI, but adds one particular point: training requires less expertise than for a traditional classifier.
This remark is not trivial: Strands Decider is open source. Designed for local use, it builds on Qwen3.5-2B. AWS has adopted the trunk (adapted with LoRA), but removed the head that models text. It replaced it with a pointer-type head that evaluates each option.
AWS reports performance indicators on the public JevBench benchmark. On an RTX 3090 in WSL2, the latest Strands Decider (v19) records a median response time of just over 100 ms, with an overall accuracy score of 0.7229. On a 36 GB MacBook Pro with an M3 Pro, it rises to 234 ms when hot (consecutive requests).
OpenAI has also responded, with the Decisions API, which is currently access-limited. Unlike AWS and TypeSafe AI, it did not develop a dedicated model. Under the hood lies GPT Luna, whose intelligence is “concentrated on a defined set of user-defined questions, with a finite number of predefined answers.” This choice has the advantage of accepting images as input in addition to text.