Why Jev by TypeSafe AI Fits Apache Kafka and Flink Better Than LLMs

Every event in a data stream carries a small decision. Is this ticket urgent? Is this alert noise? Should this agent action run without a human looking at it? The economics of that decision decide what can answer it. Rules are rigid, your own machine learning model needs labels, and an LLM is far too slow and expensive to ask a million times a day.

Jev by TypeSafe AI is built for exactly this gap. TypeSafe calls it a System One model, after Kahneman’s fast, intuitive thinking. The LLM plays the System 2 role: slow, deliberate reasoning on the cases that need it. Jev returns typed answers with calibrated probabilities in a few hundred milliseconds instead of generating text. The overview of System One models in enterprise architecture explains the category, the trade-offs against rules, machine learning and LLMs, and where it fits across integration, process intelligence and agentic AI. This post goes into the event-driven data streaming pattern.

For per-event decisions, a decision model fits stream processing better than a generative LLM. Apache Flink is the example throughout, but the pattern works the same way with Kafka Streams or Spark Structured Streaming. I compared Kafka Streams and Flink in an earlier post, and the trade-offs there apply here unchanged.

Thumbnail showing Kafka events flowing into Apache Flink for a condensed state, then to the Jev System One model returning a 0.87 confidence, then to a decisions topic, with high confidence acting, medium escalating to an LLM and low going to a person, and the tagline Confidence At Stream Speed

Why LLMs Struggle as a Per-Event Decision Engine

LLMs are excellent at open-ended reasoning and generation. They are a poor fit for a decision that has to be made on every event of a Kafka topic.

  • Latency. A generative call takes seconds, because the model produces the answer token by token. In a Flink job, a remote call of that length turns into backpressure.
  • Cost. Output tokens are the expensive part. Even a short JSON answer costs more than the decision is worth at high volume.
  • Parsing. Free text has to be parsed, validated and retried. Structured-output modes help, but the model still generates.
  • Confidence. Most LLM APIs return no calibrated probability. The application cannot tell a confident answer from a guess.

LLMs keep their place in the architecture as System 2. The expensive model should see the few events that need reasoning, not every event that needs a label.

Jev via Request-Response API vs Inside the Stream

A decision model can be called from any integration pattern. The enterprise architecture post walks through APIs, batch, messaging and streaming. For Jev, the two patterns that compete in practice are request-response and streaming, and they trade off against each other.

Behind an API, the caller waits. The model’s latency adds directly to the user-facing request, retries and timeouts live in every client, and throughput scales with the number of callers, which makes rate limits hard to manage. The state has to be small because the caller assembles it on the spot. In return, the pattern is simple, the answer is available immediately, and no new infrastructure is needed.

Inside the stream, the call is decoupled. The stream processor assembles the state from windows and joins, calls the model asynchronously, and writes the decision to a topic that any consumer can read later. Backpressure and rate limits are handled in one place. Replay makes re-evaluation cheap. The price is that the answer arrives after the event, not inside a request, so the pattern fits decisions that drive downstream processes rather than an interactive response.

Use the API pattern when a person is waiting for the answer. Use the stream when the decision feeds other systems. The decisions that run at high volume are almost always the second kind, because nobody sits in front of a screen a million times a day.

Why Event-Driven Architecture With Kafka and Flink Fits System One Decisions

A decision is an event. It happens at a point in time, it refers to a business fact, and other systems need to react to it. So a System One model fits an event-driven architecture better than a chain of synchronous calls. Producers publish events without knowing who consumes them. The stream processor prepares the question and calls the model without blocking. The decision goes back into the event log, where any number of consumers act on it, now or later. Apache Kafka is the de facto standard for the event log and Apache Flink for stateful stream processing, so they are the reference implementation in this post.

Kafka Carries the State and the Typed Decision

Jev consumes a state and returns a typed answer. Both fit Kafka’s model exactly. The state is an event, or an aggregation of events, on a topic. The decision is another event, bound to a schema, on another topic. Consumers subscribe to decisions without knowing how they were made. Replay turns re-evaluation into a normal operation. When the criteria change, the Flink job reprocesses the topic from an earlier offset and produces a new decision stream. The original decisions stay in the log as an audit trail.

Flink Does the Exact, Rule-Based Work Before the Question

Flink and Jev complement each other exactly where it matters. Jev’s documented weaknesses are arithmetic, dates and large states full of irrelevant detail. TypeSafe publishes these failure modes per model version. Flink’s strengths are those things. Windows, aggregations, joins, time handling and filtering all run in Flink before the question is asked.

A fraud triage job does not send a raw transaction to Jev. It computes the count of transactions in the last ten minutes, the deviation from the customer’s usual amount, and the country mismatch. Then it sends a condensed state of a few hundred tokens with those facts already spelled out. Flink turns the numbers into statements, and Jev judges the statements. The model never has to count or compare dates, and the state stays small enough to avoid what TypeSafe calls context rot: accuracy that drops as the state fills with content unrelated to the decision.

This split between Flink and Jev also keeps the exact logic testable in code, where it belongs. Flink computes the same result from the same events every time, by a rule someone wrote down. Jev is probably reproducible too, since TypeSafe describes its outputs as consistent for similar inputs, but its logic is learned, not written. Nobody can read the rule it applied. Everything that must be exact and explainable stays in Flink, and only the judgment goes to the model.

Calling Jev From Apache Flink

Flink has two ways to call a remote model. The first is async I/O in the DataStream API, which issues the request without blocking the operator and handles timeouts and retries. The second is model inference in Flink SQL. Flink has supported the ML_PREDICT function since version 2.1, and Flink 2.2 added model inference to the Table API. Jev’s endpoint follows its own request shape rather than the OpenAI protocol, so calling it from Flink SQL needs a custom model provider or a user-defined function. OpenRouter lists Jev behind an OpenAI-compatible interface, which is the fastest way to try the SQL route before writing a provider. Either way, group several questions on one state into a single call. Jev evaluates every question in a request in parallel, so adding questions barely changes the response time.

Rate limits matter more than latency. The throughput of a model inference operator is bound by the rate limits of the model provider. When the limit is reached, the Flink job backpressures, and that can turn into timeouts or job restarts. The throughput section below covers what that means for Jev.

Confidence Routing as a Stream Topology

The calibrated probability is what makes the streaming design work. The Flink job compares the confidence against thresholds and routes each event to a topic per confidence band. Three bands are a common starting point. High confidence goes to an action topic, and the downstream system acts automatically. Medium confidence goes to an escalation topic, where an LLM reasons about the case and produces an explanation. Low confidence goes to a review topic, where a workflow routes the case to a person.

The thresholds live in code, not in the model. Teams tune them per use case and per model version. One caution from early adopters: thresholds do not transfer between versions or between question types, so pin the versioned model ID and never share one threshold across types. The 0.87 in the figures is an example of that probability: 87% confidence that the chosen answer is right.

Architecture diagram showing events flowing from Kafka into a stream processor, the processor condensing the state and calling the Jev decision model as TypeSafe's System One model, and the typed decision routed by confidence into an action topic for an automated system, an escalation topic for an LLM, and a review topic for a person

Where Apache Fluss Fits in a Jev Pipeline

Apache Fluss is streaming storage built for real-time analytics. It offers sub-second streaming reads and writes, cheap partial updates without expensive joins, and fast lookups by key, so a Flink job can fetch the current record for one customer or order at high frequency. Those properties make it useful in a Jev pipeline.

  • Decision cache. A primary-key table stores the last decision per entity. Unchanged states do not trigger another paid call.
  • Enrichment. A lookup join assembles the full state from reference data before the question is asked.
  • Audit and evaluation. Decisions tiered into the lakehouse become the dataset for evaluating Jev and for training your own ML model later.

Fluss pays off most when the same entity produces many events and most of them should not cost a model call. A customer, a machine or an order changes state constantly, but the decision only needs to be asked again when something relevant changed. Fluss holds the current state and the last decision per entity, Flink compares the two, and Jev is called only for the real changes. That cuts the model cost and the rate-limit pressure at the same time.

Real-Time Use Cases for Jev on Kafka and Flink

The same pattern works wherever a stream carries text or semi-structured events that need judgment rather than a rule: Flink condenses the state, Jev returns a typed decision with a confidence, and the confidence routes the event.

Four examples show the range:

  • Support ticket triage. Each incoming ticket gets an urgency score, a routing choice, and a yes/no on whether it mentions a legal or safety issue. The team changes the categories without retraining anything.
  • Fraud alert triage. The authorization decision stays with the existing rules and models, in milliseconds. Jev triages the alerts those systems raise, so analysts see the plausible cases first.
  • IT operations. Alerts from monitoring systems are scored for noise, grouped by likely cause, and routed to the right on-call team. The criteria evolve with every postmortem.
  • AI agent guardrails. Before an agent runs a command, a Flink job asks whether the action is irreversible, off-task, or outside the agreed scope. Low confidence pauses the agent and opens a review.

In every example, the confidence decides whether a machine, an LLM, or a person takes the next step.

From Jev Decisions to Your Own Trained Model

A decision model and your own machine learning model are stages of one lifecycle. Jev is trained too, by TypeSafe, once and for all tasks. Your own ML model is the one you build on your own labeled data for one task. A new decision starts with Jev, because there are no labels yet and the criteria are still moving. Every answer lands in the decision topic together with the state and the confidence. After a few weeks, that topic is a labeled dataset. Nobody had to build a labeling project.

The stable, high-volume share of the decisions then moves into your own ML model, trained on that log and embedded in the Flink job, and Jev keeps handling the share where the criteria still change. The embedded model answers in microseconds and costs nothing per call. The stream topology stays the same. Only the model behind one operator changes, and the confidence routing downstream does not notice. An LLM can sit behind the same operator, and for a low-volume topic with complex judgments it sometimes should. The trade-offs are the ones from the start of the post: seconds per call, output-token cost, and no confidence value to route on.

Diagram of a Flink job with three possible models behind one operator, the pre-trained Jev decision model for changing criteria, a model you trained on your own labels for the stable share, and an LLM marked as possible with trade-offs, writing to a decision topic, with the decision log becoming the labeled dataset that trains your own model

Kafka replay closes the loop. When TypeSafe ships a new Jev version, the Flink job re-reads a slice of old events from the topic, scores them again with the new version, and compares the answers with the logged ones before anything goes live.

When to Use Jev in a Data Streaming Pipeline

Jev is a good fit in a data streaming pipeline when several of the following are true:

  • The events are unstructured or semi-structured. Tickets, logs, alerts, chat messages, and agent actions need judgment rather than a threshold.
  • The criteria change often. New categories, new policies, or new teams should not mean a retraining project.
  • There are no labels yet. The decision topic produces them.
  • Several questions apply to the same event. One call answers all of them.
  • A latency budget of a few hundred milliseconds is acceptable. Most triage and routing use cases meet this.
  • Escalation by confidence adds value. A person or an LLM should see the uncertain cases.

When Not to Use Jev in a Data Streaming Pipeline

In the following cases a System One model like Jev is the wrong tool. Each bullet names what to use instead:

  • Hard real-time. Robot control, safety loops and other OT use cases need guaranteed deadlines with no latency spikes. Neither Jev nor Kafka and Flink provide that. Data streaming is soft real-time, as I explained in a separate post on hard versus soft real-time. Those loops stay in PLCs and embedded systems.
  • Latency budgets under about 50 milliseconds. Payment authorization and trading run on Kafka and Flink, but a remote call of a few hundred milliseconds breaks the budget. Use rules or your own ML model embedded in the Flink job.
  • Exact, rule-defined logic. Thresholds, schema validation and joins have one correct answer that someone can write down. Flink computes it the same way every time, and anyone can read the rule. Jev would return a probability for something that has no uncertainty, and nobody could read the rule it applied.
  • Stable, labeled, high-volume tasks. Once the categories stop changing and the decision log has labels, your own ML model embedded in the job is cheaper per event, answers in microseconds instead of a remote call, and is often more accurate on that one task. The lifecycle above exists for this reason.
  • Numbers and dates as the core judgment. TypeSafe lists weak arithmetic and dates read as text among Jev’s known failure modes. A model that cannot reliably tell which of two timestamps is later should not decide a late-payment rule. Compute the numbers in Flink and send the result as a statement, such as “three payments late in the last 30 days”.
  • Data that cannot leave your infrastructure. Jev is a hosted API, so every state you send leaves your perimeter, and the state is the sensitive part. Run an open decision model such as Laya or Kev next to the Flink cluster instead, with your own calibration check. The enterprise architecture post covers data sovereignty and the open models in detail.

Throughput, Cost, and Backpressure

The cost is easy to estimate. At the launch price of September 2026, a condensed state of about 500 tokens costs around two thousandths of a cent per decision. TypeSafe expects the price to go down, not up, so the ratios matter more than the absolute numbers.

At a million events per day that is about twenty dollars. At ten thousand events per second it is closer to twenty thousand dollars a day. TypeSafe’s own Doom demo gives a vendor-published anchor: ten queries a second cost about seven dollars an hour. The economics work for triage and routing volumes, not for raw sensor streams.

Rate limits arrive before the cost does. The limits TypeSafe published in September 2026 were 250K tokens per second and 1,200 requests per minute, changeable without notice, with no SLA. Twenty requests per second is far below what a Flink job with a dozen parallel operators can generate, so the job will hit the limit long before the budget. Grouping questions per state, caching decisions per entity, and sampling low-value events are the mitigations.

As long as Jev runs as an early-access service, run it in shadow mode next to the current logic first, and keep a fallback lane in the topology that routes to an open decision model when the hosted API is rate-limited or down.

What Jev Means for Real-Time AI Architecture

Whichever decision model wins this category, the pattern it makes affordable will stay. System One decisions run inside the stream, and System Two reasoning runs on the escalation path. In an event-driven architecture, Kafka carries the state and the typed decision. Flink does the exact, rule-based work and prepares a condensed state. The decision model returns a probability. The confidence decides whether a machine, an LLM, or a person acts, and the decision log becomes the training set for your own model.

Start with one use case where the criteria keep changing and nobody has labels. Run Jev in shadow mode, compare its answers with the current logic, and tune the thresholds on real traffic. Then decide which decisions stay with the hosted model, which move to an open model next to the cluster, and which get distilled into your own ML model in Flink.

The first post in this series covers System One models in enterprise architecture. The third covers confidence gates in process intelligence and workflow orchestration, across data, infrastructure, applications and business processes.

To follow this work across data integration, workflow orchestration, process intelligence and trusted agentic AI, subscribe to the newsletter and connect on LinkedIn.

Don't miss my next post. Subscribe!

We don’t spam! Read more in our privacy policy

Share this post :