Skip to content
3rdLoopSolutions
  • AI
  • Engineering

Jev and System One models: fast, calibrated decisions, and what they mean for LLMs

TypeSafe AI’s Jev doesn’t write text. It makes typed decisions in milliseconds and reports how sure it is. Here is what that is good for, how we plan to use it in our products, and why it changes the job of the large language model.

3rdLoop Solutions

· 8 min read

On 27 September 2026, TypeSafe AI introduced Jev, the first of what it calls System One models. Jev does not write text. Instead, it makes a decision from a set of answers you define in advance, and it tells you how confident it is.

That sounds like a small change. We think it matters a lot for software that has to make many small decisions quickly, and for anyone deciding where a large language model (LLM) belongs in a product.

This article explains what Jev is, where it fits, how we plan to use it at 3rdLoop, and what it means for LLMs.

What Jev is

The name System One comes from Daniel Kahneman’s Thinking, Fast and Slow. System 1 is fast and intuitive. System 2 is slow and deliberate. Today’s LLMs mostly work like System 2: they reason and generate one token at a time, and that makes them slow and expensive to use for quick decisions. Jev is built for System 1 work.

TypeSafe describes Jev as “a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out.” In practice:

  • You define the possible answers first. The output is a typed value, such as a category, a yes or no, or a score, and it always matches the structure you defined. TypeSafe says the model never makes type errors.
  • Every answer comes with a probability. Jev is trained with what TypeSafe calls Reinforcement Learning for Calibrated Decisions (RLCD). The goal is calibration: when the model reports higher confidence, it should be right more often.
  • It samples in parallel. An LLM produces its answer token by token. Jev scores the possible answers in a single pass.
  • It reads program state, not only chat. Jev accepts unstructured input such as text, but it is designed for the structured state your software already holds: records, fields, and events.

TypeSafe’s published numbers:

  • Latency: 70 to 500 milliseconds end to end, which it puts at 40 to 200 times faster than frontier LLMs on comparable tasks.
  • Price: $0.042 per million input tokens. Output is free.
  • Workflow evaluation: up to 193.6 times faster and 444.6 times cheaper than an average of two frontier LLMs.

TypeSafe itself calls the workflow figures likely upper bounds. Its own team built the test workflows, and the demo inputs were short, dense, and easy to read. We treat these numbers as a direction, not a guarantee, until we have measured Jev on our own workloads.

The name refers to William Stanley Jevons, the economist who observed that more efficient steam engines increased demand for coal. In TypeSafe’s words: “Every order of magnitude drop in the cost of intelligence unlocks orders of magnitude more use cases.”

What Jev is not

Jev’s limits matter as much as its strengths:

  • It can’t generate text. It won’t draft a reply, summarize a contract, or explain its answer in prose.
  • It chooses from a fixed set. A single decision supports up to 255 choices. For larger sets, TypeSafe scores the options separately and then makes a final choice.
  • It doesn’t read images yet. Scans and photos need another step first.
  • It is in early access. Developers are being admitted from a waitlist.

Also be clear about what “can’t hallucinate” means. Jev can’t invent an answer outside the structure you defined, such as a category that doesn’t exist, a malformed value, or a made-up citation. It can still choose the wrong option. The difference is that it tells you how likely it is to be right, so your software can decide what to do with a low-confidence answer.

Where it fits

Jev is suited to decisions that are frequent, bounded, and time-sensitive:

  • Classification and routing. Which queue, which team, what urgency.
  • Scoring and extraction. How risky this is, which clause type this is, which field a value belongs in.
  • Smart conditional logic in workflows. Replacing brittle if rules with a judgment that reads context.
  • Processing large datasets. Labeling or triaging millions of records, where LLM costs add up quickly.
  • Real-time features. Decisions that must finish in under 100 milliseconds, such as live form checks.
  • Checking LLM output. Verifying, scoring, and guarding what a language model produced before anyone sees it.

How we plan to use it

Our software is built on one rule: routine work runs within the limits you set, what matters gets escalated, and consequential decisions stay with a person. Every action has a risk level that determines whether a human is out of, on, or in the loop.

To enforce that rule, the system has to make many small judgments. Is this routine? How sure are we? Who should see it? Today we answer many of those questions with LLM calls or fixed rules. LLMs are slow and costly at that volume, and fixed rules are brittle. A fast model that reports its confidence is a better fit.

These are the uses we are evaluating first.

Confidence-based escalation

Our handoff rules already say the system escalates when confidence is low, facts conflict, or a case is sensitive. With calibrated probabilities, “confidence is low” becomes a threshold we can set, test, and show to an auditor, instead of a judgment buried in a prompt. A decision above the threshold proceeds within its limits. A decision below it goes to a named person, along with the probability that triggered the escalation.

Risk tiering before an action runs

Before an action runs, the system decides where it sits on our risk model: low, medium, high, critical, or prohibited. That classification runs on every action, so it has to be fast and cheap. Jev returns a typed tier, so the result can’t fall outside the model. When a case falls between two tiers, the probabilities show it, and the system treats it as the higher tier.

Guarding LLM output

Where we still use an LLM to generate text, such as a draft reply, a summary, or an answer, Jev can check the result before anyone sees it. Does the answer cite a source? Does the draft stay within policy? Should it go to a reviewer? This adds a fast second check without doubling the cost of every request.

In our products

  • Batayan. Route each legal question by area of law and urgency, flag documents that need licensed-lawyer review, and check that every answer links to the law or ruling it relies on.
  • ariarian.ai. Classify listings and documents as they arrive, score how complete the evidence behind a value is, and mark where it runs out, so a person can decide how far to rely on it.
  • Habi. Sort leave requests, attendance exceptions, and payroll anomalies into routine items and items for HR. Decisions about people still stay with a qualified person. Jev decides who needs to look, not what happens to an employee.
  • 3C. Screen accelerator applications for completeness, and tag startups by sector and stage so the right people see them. People still choose who they build with and who joins a batch.

How we will adopt it

We will adopt Jev the way we adopt any new component. We start with one workflow, one metric, and one accountable owner. We run Jev alongside what is already in production and compare accuracy, latency, cost, and calibration on our own data before it makes any decision on its own. We also treat sensitive data the same way we do today: nothing from a customer workspace goes to a new provider until the data processing terms are in place.

What it means for LLMs

Jev doesn’t replace large language models. It changes what they should be used for.

LLMs return to the work they are good at. Many production systems today use an LLM as an expensive if statement. They ask it to classify, route, or score, then parse its text and hope it follows the format. Those are System 1 tasks run on System 2 hardware. Moving them to a model like Jev leaves LLMs with the work that actually needs language and reasoning: drafting, explaining, summarizing, and multi-step analysis.

The typical architecture becomes two-speed. A fast model handles the many small decisions, and a slow model handles the few that need depth. The fast model can also decide when to call the slow one. This matches how people work: most decisions are quick, and a few deserve careful thought.

Confidence becomes something you can build on. LLMs are often confidently wrong, and the confidence they state is unreliable. A model that reports calibrated probabilities makes the phrase “escalate when unsure” something you can engineer and audit. For regulated industries, that may matter more than raw accuracy.

Structured output becomes the default. When the model can only return values that fit a type, a whole class of failures disappears: parsing errors, retries, and invented options. Your code can use the result directly.

Cheap intelligence means more of it. If Jevons is right, lower prices won’t reduce spending on AI. They will spread it into places that were never worth an LLM call: every form field, every record, every event. That makes oversight more important, not less. When a system makes millions of decisions a day, you need to know which of them a person should see.

The bottom line

Jev is a new kind of model for a specific job: fast, typed decisions that report their confidence. It won’t write your emails, and it can still be wrong. It can make the routine judgments inside software faster and cheaper, and, most important for us, measurable.

That fits how we build. Software handles routine work within limits you set, escalates what matters, and leaves consequential decisions with a person. A model that knows when it is unsure makes the escalation step more reliable.

We will share what we learn as we test it. If you want to explore where fast, calibrated decisions could fit in one of your workflows, contact us.

Source: TypeSafe AI, “Introducing System One models and Jev,” 27 September 2026. Figures are TypeSafe’s own and have not yet been independently verified.

Your people stop doing the routine work. They still make the calls that matter, and they can always step in.

Tell us about one workflow that eats time or carries risk. We’ll reply with what software could handle, what it should hand to a person, and who stays accountable.