The data streaming market changed shape in less than a year. IBM closed its acquisition of Confluent in March 2026 and retired its own Kafka products in Confluent’s favor. In August, founder Jay Kreps announced he is stepping back. CoreWeave absorbed Bufstream into its internal AI platform. Decodable disappeared into Redis. StreamNative, the company that spent years arguing against Kafka, now ships its own Kafka offering. The Data Streaming Landscape Q3 2026 maps this market on two axes that did not exist in earlier editions: the workload a platform serves, and who controls where it runs.
Any one of these events would justify an update. The reason for a new model is bigger. Who controls where your data runs, and under whose jurisdiction, moved from a compliance footnote to a board-level architecture decision, and the sovereignty debate that started in AI has arrived in data infrastructure. That question reshaped how the landscape is organized, and it returns in almost every chapter of the report.
![]()
What is data streaming?
Data streaming moves and processes data continuously as events happen, instead of in scheduled batches or on-demand requests. Apache Kafka is the de facto standard and the protocol most vendors and frameworks now implement, used by more than 150,000 organizations. It is not the only implementation, and the landscape covers the alternatives as well: stream processing frameworks, cloud-native services that never adopted the protocol, and analytics platforms that now speak it.
Two criteria decide whether a technology gets a bubble on the landscape map. Its core purpose has to be streaming: moving, processing, or storing events as a continuous, replayable log. And it needs production adoption that customer references or revenue can back up, not marketing claims. Message brokers such as IBM MQ, RabbitMQ, or Solace move data too, but they solve messaging rather than streaming, and the Data Integration Landscape covers them.
Everything else is covered in the text of the report instead: table formats such as Apache Iceberg and Apache Paimon, consuming-only OLAP engines such as Apache Pinot, ClickHouse, and Apache Doris, Kafka gateways, and the long tail of managed Kafka services from smaller clouds.
Streaming itself is not a goal. Plenty of use cases are served better by batch, a message queue, or a simple API call, which I covered in When NOT to Use Stream Processing. The same holds at the edge, where MQTT, NATS, and event meshes carry most of the traffic and Kafka owns only site-to-site replication, as described in Edge to Cloud and Back.
Why did the landscape model change this year?
Earlier editions organized vendors by deployment model: self-managed, PaaS, and SaaS in 2023 and 2024, with BYOC added as a row in 2025 after WarpStream pioneered it. Two things broke that model. First, deployment stopped being a vendor property. The same product family now ships as fully managed, BYOC, and self-managed at once, so a row no longer describes anything. Second, sovereignty stopped being a footnote. It is now a structural dimension of the buying decision, and the old model had nowhere to put it.

How to read the landscape
So the operating model became an axis. The vertical axis is workload focus: operational at the top (applications, transactions, low latency), analytical at the bottom (lakehouse, real-time analytics, reporting). The horizontal axis is the primary operating model, from fully managed and vendor-operated on the left to self-managed and air-gapped on the right. Each vendor appears once, at the model that leads its business today, and the report names every deployment option each one offers. Bubble size shows market footprint, and bigger is not better. Color marks vendor origin: US, EU, China, or vendor-neutral open source. A grey ring marks ownership that spans jurisdictions.
The second change is the analytical half. The map now has an analytical side for platforms such as Databricks and Snowflake and for open-source projects such as Apache Spark and Apache Fluss, because the boundary between the streaming layer and the lakehouse is where much of the market movement happens. The operational half remains the core of the landscape.

The full vendor-by-vendor analysis is available as a free PDF without registration.
The Four Quadrants: Vendors by Workload and Operating Model
Vendors appear alphabetically within each quadrant, and the order carries no ranking.
Operational and fully managed
Aiven, Alibaba Cloud, Amazon, Azure Event Hubs, Confluent (IBM), Google Cloud, Oracle, and WarpStream (IBM). The hyperscalers sell managed Kafka right beside their proprietary services, which is the clearest evidence that the protocol won. Confluent sits at the midline because Confluent Cloud leads the business while Platform and Private Cloud carry many of the most critical deployments. WarpStream sits there for the opposite reason: strong isolation in your VPC, but it cannot run without the vendor-operated control plane.
Operational and self-managed
Apache Flink, Apache Kafka, AutoMQ, Axual, Cloudera, Red Hat (IBM) with its Strimzi-based streams for Apache Kafka, Redpanda, StreamNative, Strimzi, and Ververica. Open-source projects follow the same placement rule as vendors: the bubble sits where the project reaches production today, so Kafka and Flink, with broad managed ecosystems, sit closer to the midline than Strimzi, which is run by hand. Axual shows sovereignty winning deals on its own in the EU public sector. Redpanda’s most substantial new capability in 2026 is Shadowing, in-broker replication that migrates topics, schemas, offsets, and ACLs from other Kafka services.
Analytical and fully managed
Databricks and Snowflake, both new on the map. Each now offers two ways in: the established connector path, and a native Kafka-protocol path. Databricks Zerobus accepts Kafka producers directly. Snowflake Datastream implements the Kafka wire protocol for producers and consumers without Kafka underneath and lands topics as governed tables, still in private preview. Both are analytics platforms first, with one cloud perimeter and no self-managed path. They shorten the path into the lakehouse. The operational event log belongs outside the perimeter.
Analytical and self-managed
Apache Fluss, Apache Spark, Materialize, and RisingWave. This is where the analytical side of streaming is being rebuilt in the open, on several engines rather than one. Fluss is the one to watch: streaming storage built for analytics, columnar where Kafka is row-oriented, with Apache Paimon as the table format where the data settles. Materialize and RisingWave both dropped the streaming database label this year.
The full report has an entry per vendor with deployment options, plus a guidance paragraph per quadrant on when that operating model is the right decision. No registration required: Download the Data Streaming Landscape Q3 2026
Confluent Inside IBM: What Changes for the Market
Confluent remains the reference for a complete data streaming platform. What changes is the story the market has heard for a decade. Confluent grew by evangelizing Kafka as the central nervous system that displaces legacy messaging and batch middleware. Inside IBM, that pitch shares a portfolio with IBM MQ, webMethods, App Connect, and DataStage.
The likely evolution is from displacement to coexistence: IBM will position Confluent for data streaming alone, next to MQ for messaging, webMethods and App Connect for integration, and DataStage for batch. The portfolio logic is sound. The message is quieter.
The portfolio also has more Kafka in it than the announcement suggested. IBM now owns Confluent Cloud, Platform, and Private Cloud, WarpStream for diskless BYOC, and through Red Hat the Strimzi-based streams for Apache Kafka, which kept shipping releases in 2026. Two supported Kafka distributions under one owner is an open roadmap question, next to the pace of a founder-led organization after the founder’s departure and the future of Flink inside IBM. None of this changes what the platform does today. All of it belongs in a multi-year platform decision.
Data Sovereignty and the Return of Self-Managed Deployments
For a decade the direction of travel was one way: out of the data center and into the cloud. Self-managed and private deployments are recovering ground, and the reasons are jurisdiction, control, and continuity rather than cost. Confluent is the clearest public evidence. It built its growth story around Confluent Cloud, then the hybrid side reasserted itself: Confluent Platform grew 18 percent year over year in Q1 2025, its strongest first quarter in three years, and in October 2025 the company launched Confluent Private Cloud for regulated industries.

The regulation is loudest in Europe, with the Cloud and AI Development Act and DORA treating concentration on a single hyperscaler as a risk to manage. But this is not a European problem. Middle East enterprises building national capability, APAC data-localization rules, US public sector and defense running air-gapped, and China requiring in-country operation all ask the same question, and only the weighting differs.
A Vendor Switched Off by Its Own Government
The strongest argument for deployment control came from AI, not from data infrastructure. In June 2026 the US government suspended access to Anthropic’s most capable models for foreign nationals, and the vendor disabled them globally until access was restored almost three weeks later. A leading vendor was switched off by a government its customers had not chosen. Every dependency in the stack carries the same question, including the platform carrying your operational data. The full argument is in the Trusted Agentic AI Landscape Q3 2026.
Does the Kafka topic become the lakehouse table?
The direction is the same at every vendor: Confluent Tableflow, StreamNative Ursa, Snowflake Datastream, Databricks Zerobus, and Fluss all turn the topic into a table. The reality in most enterprises is different. Zero-copy is rare, in the same way zero-ETL turned out to be rare. Most teams still run a connector from Kafka into object storage, and the data exists in both places. Storing it twice is often the better pattern: Kafka holds the raw event log for replay and the operational workloads that need low latency, while Iceberg holds the governed tables for analytics. What no copy count solves is the engineering in between: schematization, type conversion, schema evolution, and data quality before the write. The report covers this in its own chapter, together with why streaming databases never became a category.
Five Questions Before Choosing a Data Streaming Platform
Start with the workload, not the vendor. Five questions place a vendor on this landscape more accurately than any marketing claim.
- Workload: Which parts of your architecture are operational, with polyglot consumers and delivery guarantees, and which are analytical? This places a platform on the vertical axis.
- Control: What does your regulator require about where data runs, what stops working if the vendor becomes unreachable, and can you run the software yourself if the relationship ends? This places it on the horizontal axis.
- Jurisdiction: Whose law binds the vendor, and the parent behind the vendor? Shown on the map by color and the grey ring.
- Economics: What does the platform cost under object-storage math rather than last year’s broker math, at your actual throughput and retention?
- Exit: What breaks if you leave? Protocol compatibility makes the technical migration plausible, while governance, connectors, and operational tooling are where the real switching cost sits.
Four of the five have concrete answers before you ever read a feature list.
Data Streaming in the Trinity of Modern Data Architecture
Data streaming is the foundation of what I call the Trinity of modern data architecture. It is how most enterprises implement event-driven architecture, because Kafka brings the guarantees that layer needs: ordering, replay, and decoupling. Process intelligence and workflow orchestration turn that data into an understanding of how the organization runs and coordinate what happens next. Trusted agentic AI acts on both. AI agents make the dependency concrete: an agent approving a payment needs the current state of the business, not yesterday’s snapshot. The Data Integration Landscape 2026 maps how streaming connects to APIs and batch, and the Process Intelligence Landscape 2026 covers the layer above it.
Download the Data Streaming Landscape Q3 2026
The full report works through every vendor and open-source project, the Kafka protocol and the architecture shifts inside Kafka, the IBM chapter, sovereignty, the lakehouse boundary, governance and gateways, ten trends through 2027, and the five questions. It is a free PDF with no registration: Download the Data Streaming Landscape Q3 2026
If you are making a streaming platform decision this year, start with the two axes, not the vendor list. Apache Kafka is the dominant implementation and the protocol most vendors now speak, but it is not the only one, and the analytical side runs on different engines entirely. Every enterprise has to decide which workloads belong in the stream, and who controls where that stream runs.
To follow this work across data integration, workflow orchestration, process intelligence, and trusted agentic AI, subscribe to the newsletter (kai-waehner.de/news) and connect on LinkedIn (linkedin.com/in/kaiwaehner).
About this landscape: This is an independent analyst perspective. No vendor paid for inclusion or placement. No vendor reviewed its section before publication. I run an advisory practice through Kai Waehner GmbH and have held roles at Talend, TIBCO, and Confluent. I also serve as Global Field CTO at Kestra, a workflow orchestration vendor outside this landscape’s scope. The full methodology note is in the PDF.