Most AI inside enterprise software does not write anything. It decides. Is this ticket urgent? Does this transaction need a second look? Is this agent action safe to run? Today those decisions run through chat models built for writing, which makes every one of them slow, expensive and hard to govern.
Jev by TypeSafe AI is the first model built for the decision alone. TypeSafe calls it a System One model. It launched on September 15, 2026, returns typed answers with calibrated probabilities instead of text, and answers in a few hundred milliseconds. Within ten days, Vercel, Cloudflare, OpenRouter and Databricks all added access, and open-source clones appeared.
System One models add a fourth option between rules, trained machine learning and large language models, and the enterprise architecture has to make room for them in every layer. This post explains the category, the trade-offs, where a decision model fits in data integration, process intelligence and agentic AI, what it means for trust and sovereignty, and when not to use one. It is the first post of a three-part series on System One models like Jev. The second goes deep on data streaming with Apache Kafka, Flink and Fluss, the third on workflow orchestration and process intelligence.
![]()
What Is a System One Model?
A System One model is an AI model that evaluates a piece of state and returns typed answers with calibrated probabilities instead of generated text. The term is TypeSafe’s. The generic name this post also uses is decision model.
A System One model is not a rules engine either. Business rules and decision tables (the DMN standard in the BPM world) spell out every condition and outcome in advance, and a System One model can replace them where the rules stopped being writable.
The difference is what “deterministic” means. A decision table does exactly what someone wrote down. It can be read line by line, shown to a regulator, and it changes only when a person edits it. A System One model’s logic is learned from data. Nobody can read the rule it applied, it can be confidently wrong on inputs unlike its training data, and its behavior shifts with every model version. It is probably reproducible, since TypeSafe describes the outputs as consistent for similar inputs, but it is not specified. It gives a statistical answer, computed the same way each time.
How Jev Answers a Question
Jev is the reference example. You send a state (text or JSON) plus a set of questions. Each question is a Choice among options, a Score on a scale, or a yes/no probability. All questions are answered in parallel in one call, and the answer is always inside the schema you defined. TypeSafe’s own summary of the use case is “smart if-statements“: classify, route, score, extract or branch where hand-written logic is too brittle, with the surrounding code constraining what the model can do.
The name comes from Daniel Kahneman’s book “Thinking, Fast and Slow” (2011) and its split between fast, automatic judgment (System 1) and slow, deliberate reasoning (System 2). A System One model makes the snap judgment. The LLM plays the System 2 role: slow, deliberate reasoning on the cases that need it. The model itself is named after William Stanley Jevons, whose paradox says that making a resource cheaper increases its total consumption. The bet is that decisions cheap enough to make ten times a second will be made ten times a second.
The launch numbers support the bet. Pricing is $0.042 per million input tokens with no charge for output, and TypeSafe reports 70 to 500 milliseconds end to end. Independent measurements land around 0.35 seconds. TypeSafe says the model is trained with a method it calls Reinforcement Learning for Calibrated Decisions (RLCD) and has not published a paper. All of this is as of September 2026 and will move.
What “Cannot Hallucinate” Actually Means
TypeSafe markets Jev as a model that cannot hallucinate. The claim is narrower than it sounds. The model cannot return a value outside your schema. It can still pick the wrong one.
TypeSafe publishes the known failure modes per version, which is unusual and useful. The list for version 1.13 includes literal reading (the model answers the question you wrote, not the one you meant), weak counting and arithmetic, dates read as text rather than as ordered values, and accuracy that drops as the state fills with content unrelated to the decision. The state is treated as data, not as hostile input, so text written to steer the model can move the answer. Each of these weaknesses is a design rule for the surrounding architecture: compute the numbers in code, keep the state small, write the criteria explicitly, and never let untrusted text into the state unfiltered.
Decision Models vs Rules, Machine Learning, and LLMs
Enterprises already have three ways to make a decision in software. A decision model like Jev is the fourth. The differences are in what each needs, what it returns and how fast it answers.

What Each Option Needs and Returns
Rules are fixed logic in code. They need nothing but a developer, answer in microseconds, and are the right tool for thresholds, validation and anything deterministic. They cannot judge a free-text ticket.
Self-trained machine learning means a model you build on your own labeled data, for one task. It needs labeled data, a data scientist and a retraining cycle whenever the categories or the data drift. Jev is also a trained model, but TypeSafe trained it once for all tasks, and you only supply the questions. Performance varies a lot. An optimized model artifact from a platform like H2O.ai runs in microseconds inside the application. A typical Python model behind a REST endpoint answers in milliseconds to tens of milliseconds. On a stable, well-labeled task, a trained model is usually the most accurate and the cheapest option per decision.
Large language models reason and write. They need a prompt, answer in seconds, and cost per output token. They are the right tool for explanation, summarization and open-ended reasoning, and the wrong tool for a decision that has to be made a million times a day.
Decision models are pre-trained and general-purpose. The vendor trained the model once, for every task, and you supply only the questions. There is no training step on your side. The questions and options are defined in natural language at request time, so changing a criterion is a text edit. They handle unstructured input out of the box and answer in a few hundred milliseconds as a remote call. On a stable, labeled task they are usually less accurate than a trained model. One independent intent-classification test with 77 categories put Jev at 79.0% accuracy against 83.9% and 86.2% for two OpenAI models, and the difference was statistically significant.
The Four Options Side by Side
| Rules | Trained ML | Decision Model | LLM | |
|---|---|---|---|---|
| Needs | Code | Labeled data, retraining | Questions defined at runtime | Prompts |
| Returns | Boolean or value | Class or score | Typed answer with probability | Text, optionally structured |
| Latency | Microseconds | Microseconds to tens of ms | Roughly 100 to 500 ms | Seconds |
| Cost driver | Developer time | Infrastructure and people | Input tokens | Output tokens |
| Handles free text | No | With feature engineering | Yes | Yes |
| Best fit | Deterministic logic | Stable, labeled, high volume | Changing criteria, no labels, many questions per item | Reasoning, generation, explanation |
From a Decision Model to Your Own Trained Model
The 0.87 in the figures of this article is an example: the calibrated probability Jev returns for one question, here 87% confidence that the chosen option is right. The four options are stages of a lifecycle rather than competitors.
A new decision starts with a decision model, because there are no labels yet and the criteria are still moving. Every answer is logged together with the state. After a few weeks, that log is a labeled dataset. The stable, high-volume share of the decisions then moves into a trained model, and the decision model keeps the share where the criteria still change. The LLM handles the cases both of them are unsure about. Rules cover whatever turned out to be deterministic after all.
Where System One Models Fit in the Enterprise Architecture
I describe the modern enterprise stack as three layers: event-driven integration, process intelligence, and trusted agentic AI. A decision model touches all three, because a cheap, typed decision with a probability is the kind of primitive that connects layers. The integration layer produces it, the process layer branches on it, and the agentic layer is gated by it.

Data Integration: APIs, Batch, Messaging, and Streaming
The integration layer covers four patterns, and a decision model can be called from each of them. The call shape is the same. What differs is who asks, how often, and what the answer becomes.
- API. An API gateway or a service asks one question per request and gets a typed answer back in the response. This is the simplest form: a fraud check inside a checkout call, a routing decision inside a support API. Latency adds directly to the user-facing request, so the state has to be small. Typical platforms are MuleSoft, Kong and Workato.
- Batch. A scheduled job scores a whole table in a platform such as Databricks or Snowflake. Databricks showed within days of the launch how to call an open decision model from SQL through a function, with the selected option and the probabilities returned as columns. This is the cheapest way to label history and the natural way to build the dataset for your own trained model later.
- Messaging. A consumer scores each message on a queue in IBM MQ, RabbitMQ or Solace and publishes a decision message. Order and delivery guarantees come from the broker, and a slow model call does not block the producer.
- Streaming. A stream processor scores every event with a condensed state and writes a decision event, with its confidence, to a topic. This is where the per-event economics and the latency budget matter most, and where the model’s weaknesses with numbers and dates are best handled, because the stream processor computes the aggregates before the question is asked.

The second post in this series covers why System One models like Jev fit data streaming with Kafka, Flink and Fluss, and where they do not.
Process Intelligence: Confidence as the Decision Gate
Process intelligence combines process mining, workflow orchestration and the decision gate that sits between them. The calibrated probability is that gate. High confidence lets the process proceed automatically. Medium confidence routes the case to an LLM that reasons and explains. Low confidence pauses for a person.
The thresholds belong to the process, not to the model. They are tuned per use case and per model version, and early adopters warn that they do not transfer between versions or between question types.
Workflow orchestration runs everything after the gate: the escalation flow, the human review, the retry after a timeout, the rollback. The orchestrator does not make the decision. It makes sure the right thing happens once the decision exists, across data pipelines, infrastructure, applications and business processes alike. This is also where the determinism lives. A System One model like Jev gives a statistical answer. The workflow decides, by rule, what happens at each confidence level, keeps the audit trail, and sends the cases that must be explainable by rule to a decision table or a person.
Camunda shipped a Jev connector within days of the launch, which shows the BPM vendors see the same split. Every decision the gate makes lands in a log with its probability. That log is what process mining needs to show where thresholds are too tight, where reviewers overrule the model, and where a step could be automated.
The third post in this series covers confidence gates in workflow orchestration in detail.
Trusted Agentic AI: A Gate Before the Agent Acts
The most direct use of a decision model like Jev is in front of an agent. Before an agent runs a command, a decision model answers a few typed questions about the pending action: Is it irreversible? Is it off-task? Is it outside the agreed scope? Low confidence on any of them pauses the agent and opens a review. The LLM inside the agent stays the reasoner. The decision model is the check.
Vercel is the first public production example. Its engineering team reports that its safety classifiers ran 18 times faster after swapping the model behind them for Jev. Calibrated confidence turns “trusted” from a slogan into a measurable property, because every action an agent takes carries the probability that allowed it, and that number can be audited later.
Trust, Data Sovereignty, and Open Decision Models
Jev is hosted only. TypeSafe publishes docs, prices and rate limits, but no weights, no container and no on-prem option. Access goes through TypeSafe’s own API, through Vercel AI Gateway (which offers a zero-data-retention option), or through OpenRouter, which lists the model since September 18. For any enterprise with regulated data, in banking, healthcare or the public sector, that means the state you send leaves your perimeter, and the state is exactly the sensitive part. In Europe, data sovereignty requirements make this even stricter.
The open alternatives formed within days. The request shape is public, so open models that copy it work with the official SDKs. Laya is a purpose-built decision model based on ModernBERT, with 322M to 421M parameters and support for more than 100 languages. Kev is a family of LoRA adapters on Qwen3.5 bases that serves TypeSafe’s own API contract, so switching means changing a base URL. SemIf-OpenJev is the model Databricks used for its walkthrough on serverless GPUs. All three run inside the perimeter, and all three are different models with their own accuracy. Calibrated confidence is the part nobody has matched in public, so an open model’s probabilities need your own evaluation before they drive a gate.
The architecture answer for regulated workloads is the same one that works for LLMs: put the decision model behind one internal interface, run the open model inside for data that cannot leave, use the hosted model outside for the rest, and keep the schema identical so the gate logic does not care which one answered.
Decision Models in a Multi-Model AI Strategy
Multi-model used to mean several LLMs from different vendors. Decision models extend it to different kinds of models for different kinds of work: a System One model for the decision, an LLM for the reasoning.
I made that case in an earlier post on multi-model AI strategy with five drivers: availability and model churn, the cost of running agents, regulation and compliance, sovereignty and jurisdiction, and picking the right model for the task. A decision model hits all five.
The right model for the task and the cost are the obvious ones. Availability is the driver people underestimate. Jev has one published model version, paused new signups a week after launch under demand, and publishes rate limits without an SLA. An application that hard-depends on it has a single point of failure. Plan for a second decision model from day one.
The follow-up post on multi-model AI orchestration introduced the principle “decide high, route low.” The governed decision gate defines which models are allowed for a unit of work. The technical router executes within that set. A decision model fits the principle without changes: the gate allows a hosted decision model for triage, an open one for regulated data, and an approved LLM for escalations. The router calls the allowed model, falls back to the open one when the hosted API is rate-limited or down, and sends low-confidence cases to the LLM.
When a System One Model Like Jev Is the Right Choice
A decision model like Jev is a good fit when several of the following are true.
- The input is unstructured or semi-structured. Tickets, logs, alerts, messages and agent actions need judgment rather than a threshold.
- The criteria change often. New categories, new policies or new teams should not mean a retraining project.
- There are no labels yet. The decision model produces them from day one.
- Several questions apply to the same item. One call answers all of them.
- A latency budget of a few hundred milliseconds is acceptable. Most triage, routing and gating decisions meet this.
- Escalation by confidence adds value. A person or an LLM should see the uncertain cases.
When Not to Use a System One Model Like Jev
- Ultra-low latency. Payment authorization, trading and OT control need answers in well under 50 milliseconds. Use rules or an embedded trained model.
- Exact, rule-defined logic. Thresholds, validation and joins have one correct answer that someone can write down. Code computes it the same way every time, and anyone can read the rule. A System One model would return a probability for something that has no uncertainty, and nobody could read the rule it applied.
- Stable, labeled, high-volume tasks. A trained model is cheaper per decision, faster and often more accurate once the task stops changing. The distillation pattern above exists for this reason.
- Tasks that need text. An explanation, a summary or a customer reply requires an LLM.
- Numbers and dates as the core judgment. Compute them in code and pass the result as a statement.
- Data that cannot leave your infrastructure. A System One model that is only available as a hosted API, as Jev is today, is not an option here. Use a decision model you can run yourself, such as one of the open models inside the perimeter, with your own calibration check.
- Very large option sets. Jev handles up to 255 options per choice and falls back to a two-stage score-then-choose process above that. Routing across a big category tree needs a hierarchy of questions, not one question.
Jev Production Readiness as of September 2026
Treat Jev as a promising early-access service. The published limits at launch were 250K tokens per second and 1,200 requests per minute, changeable without notice, with no SLA. There is one model version, and thresholds calibrated on it do not carry over to the next. Pin the versioned model ID, run the model in shadow mode next to the current logic, and keep a fallback path.
The fast platform adoption needs the same care. Vercel, Cloudflare, OpenRouter and Databricks added access within ten days, and Vercel reported more teams trying Jev on day one than any previous model. That shows distribution. It does not show validation. A gateway can add an endpoint in an afternoon.
The category is proven when several vendors ship competing decision models with their own weights and enterprises publish production numbers. Neither has happened yet. The same goes for the numbers. The 194x faster and 445x cheaper figures on TypeSafe’s site come from its own workflow evals, built by its own team and scored against the average of two frontier LLMs, and TypeSafe says in the launch post that these sit at the high end of real-world gains. JevBench is a vendor benchmark and should be read as one.
What System One Models Like Jev Mean for Enterprise Architects
The model will change. The category will stay, because the underlying observation is correct: most AI in enterprise software needs a decision, not a paragraph, and paying LLM prices and LLM latency for every decision was never going to scale. Add decision models to the reference architecture now, behind an interface that lets you swap the model, and start in shadow mode where the criteria keep changing and nobody has labels. The integration layer produces the decision, the process layer branches on its confidence, and the agentic layer is gated by it. Which models are allowed in each place stays a governed decision.
The next post in this series shows the streaming pattern on Apache Kafka and Flink, where the per-event economics make the case sharpest. The third covers confidence gates in workflow orchestration across data, infrastructure, applications and business processes.
To follow this work across data integration, workflow orchestration, process intelligence and trusted agentic AI, subscribe to the newsletter and connect on LinkedIn.