No industry has thought harder about orchestration than telecom. Operators have standards for it, maturity levels for it, and an entire vendor category selling it. A tier-one operator can retune tens of thousands of antennas during a winter storm with almost nobody in the loop.
The same operator still runs its billing cycle on a scheduler bought in the 1990s. It provisions servers through tickets. It coordinates a fibre migration campaign across millions of lines with spreadsheets, email, and a handful of people who know how the sequence works.
Telecom orchestration maturity is deep and narrow. It is excellent inside the network domain and absent everywhere else. Autonomous network frameworks govern the radio and the core. They say nothing about the invoice run, the partner settlement, the data platform, or the GPU cluster the operator now sells to enterprise customers.
This post maps unified orchestration in telecom. It starts with what the industry already calls orchestration and the four silos underneath that vocabulary. It covers the tools inside each silo and the flows that cross all four: the billing cycle, copper switch-off, and governed AI agents in network operations.
![]()
What Telecom Already Calls Orchestration
Before mapping anything new onto telecom, it helps to be precise about the words already in use. Orchestration in telecom is not one discipline. It is at least half a dozen, each with its own standards body, tooling, and team.
OSS, BSS, and OTT: What Each Layer Owns
Business Support Systems (BSS) own the commercial side: product catalog, pricing, eligibility, customer orders, charging, and invoicing. Operations Support Systems (OSS) own the network side: activating those services, assuring their performance, and managing inventory and resources.
Over-the-Top (OTT) services are third-party digital offerings such as video, messaging, and cloud applications that an operator sells and bundles alongside its own connectivity. They add a third coordination problem. A customer expects a streaming bundle to activate in the same minute as the SIM, which means BSS, OSS, and a partner platform all have to agree.
I covered this split in detail in Telecom OSS Modernization with Data Streaming and, years earlier, in Telco-OTT services with OSS and BSS integration.
Service Orchestration, Resource Orchestration, MANO, and the Closed Loop
Inside the OSS, the word orchestration carries several meanings at once:
- Service orchestration decomposes a customer-facing service into resource-facing services and drives activation across domains.
- Resource orchestration and NFV MANO manage the lifecycle of network functions and the cloud resources under them. ONAP and OSM are the open-source entries here.
- The O-RAN Service Management and Orchestration framework (SMO) does the equivalent job for a disaggregated radio access network, with rApps on top.
- Closed-loop assurance detects a degradation, decides, and acts, usually in seconds and usually inside one domain.
- Intent-based networking expresses a declarative goal and leaves the path to the system.

Each of these is a real capability, and none of them was designed to coordinate work outside the network.
Where Autonomous Network Levels Stop
The TM Forum measures progress on a scale from Level 0, fully manual, to Level 5, fully autonomous. Most operators sit between Level 2 and Level 3 as of 2026. Validated Level 4 results are arriving in specific domains such as fault management and energy optimization, and roughly four in five operators target Level 4 or above by 2030. I wrote about the direction of travel after TM Forum Innovate Americas.
Read the scope of those levels carefully. They describe the network, not the operator. A validated Level 4 fault-management domain says nothing about whether the monthly invoice run recovers cleanly from a failed step, or whether a partner settlement dispute has an audit trail. That work is orchestration too, and it lives in a different part of the company.
The Four Orchestration Silos Inside an Operator
Strip away the vocabulary and the same structure appears in a telco that appears everywhere else. Unified orchestration means one control plane for data, infrastructure, applications, and business processes. Agentic AI is governed across all four by the same control plane.

Data Orchestration: Mediation, Telemetry, and AI Feature Pipelines
Mediation feeds usage records into rating and billing. Radio and core telemetry lands in the data platform for performance and capacity work. Customer data feeds churn models, next-best-offer engines, and the care desk. Regulatory extracts leave on a schedule because the regulator expects them on a schedule. Data teams coordinate all of it with pipeline schedulers built for that job.
Infrastructure Orchestration: Telco Cloud, Change Windows, and AI Factories
Cloud-native network functions run on Kubernetes across central sites and edge locations. Firmware and configuration campaigns roll across thousands of devices in maintenance windows. Landing zones, certificates, and capacity are provisioned by one team and consumed by everyone.
Operators are also becoming infrastructure providers in their own right. Sovereign cloud offerings and GPU capacity are products now, sold to enterprise customers. Deutsche Telekom and NVIDIA opened an industrial AI cloud in Munich in 2026, operated by T-Systems. Around ten thousand accelerators are available to enterprise customers. An AI factory needs orchestration nobody in the OSS world ever scoped: tenant onboarding, job admission, quota enforcement, and teardown.
Application Orchestration: Order to Activate, Charging, and Network APIs
Order to activate is the classic multi-step operation. Validate the order, reserve inventory, call the activation system, update charging, confirm to the customer, and compensate cleanly when a step fails halfway.
Network APIs added a new one. Open Gateway and CAMARA moved from pilots to commercial volume in 2026. Identity and anti-fraud APIs lead the revenue, and aggregators distribute them globally. Exposing a network capability as a product means partner onboarding, entitlement, consent handling, usage capture, and settlement. Every one of those is a coordinated multi-system operation.
Business Process Orchestration: Fulfilment, Field Dispatch, and Number Portability
Some flows have people in them by design. An engineer visits a site. A porting request waits on another operator. A complaint escalates toward an SLA credit. A wholesale dispute goes to a commercial review. These processes have deadlines, escalation paths, and audit requirements, and they historically live in BPM suites or in the order management module of a BSS vendor.
The Flows That Cross All Four Categories
Telecom already has names for the flows that matter, and all of them cross every category. Lead to cash starts in a shop or an app and ends in the ledger. Trouble to resolve starts with an alarm and ends with a customer notified and a ticket closed. Concept to market ends with a product live in every channel.
No silo owner has ever been accountable for one of those flows end to end. The seams between the silos are where the coordination code lives, and that code is usually somebody’s homegrown framework.
One Orchestrator Per Silo: Airflow, Cisco NSO, Temporal, and Camunda
The orchestration market served each silo very well, which is how the fragmentation happened. Each category has a strong incumbent, and each is a reasonable choice for a real set of problems.
Data Pipelines: Airflow and the Analytics Schedulers
Apache Airflow dominates data pipeline orchestration, with Dagster, Prefect, and the cloud providers’ pipeline services alongside it. In a telco it typically runs the analytical work: telemetry aggregation, model feature pipelines, reporting loads. Its limits appear at the edges of that scope. Python-defined DAGs create a bottleneck when engineers outside the data team need to contribute, and coordinating infrastructure tasks or human steps works against the grain of the tool.
Network and Infrastructure: Cisco NSO, Nokia NSP, Itential, Ansible, and ONAP
Model-driven network orchestration is its own mature market. Cisco NSO and Crosswork, Nokia NSP, Itential, Red Hat Ansible Automation Platform, Terraform, and the open-source MANO projects each solve a well-defined part of network and infrastructure automation. They are strong where the target is a network element or a cloud resource. They were never meant to wait for a credit check, call a partner API, or route a task to a field engineer.
Enterprise Scheduling: The Billing Batch Backbone
BMC Control-M, Broadcom Automic, IBM Workload Automation, Redwood, and Stonebranch run the overnight core of the operator’s IT: the billing cycle, the general ledger, the extracts, the file handoffs. They are reliable, well understood, and expensive, and they were designed for a world with a quiet window at night.
Durable Execution Inside Services: Temporal and Restate
Temporal, Restate, AWS Step Functions, and Orkes Conductor give developers durable execution: long-running stateful code that survives crashes and retries safely. An activation service or a charging integration benefits from exactly that, as I described in the rise of the durable execution engine. These engines orchestrate code inside an application boundary. Stretch one into a company-wide control plane and the team writes the glue itself.
Business Process Suites and BSS Order Management
Camunda, Pega, Appian, ServiceNow, SAP, and open-source engines such as Flowable own the human-centric processes. The BSS vendors ship their own order management and fulfilment engines inside Amdocs, Netcracker, Comarch, Blue Planet, and Salesforce Communications Cloud. They are strong where the process is the product and business analysts co-own the model. None of them is a data pipeline tool or a network automation tool.
Starting Unified Inside One Category
The four-silo argument does not require an operator to start with all four, and in practice many do not. Teams typically adopt a unified orchestration platform for one category and stay there for a while: a data team replacing cron jobs and a Python scheduler for telemetry pipelines, a platform team sequencing change windows, a billing team moving job chains off a legacy scheduler. The platform earns its place in that one category on its own terms. The question is only whether it can grow when the second category shows up, without paying for that breadth up front.
Four properties decide whether a platform can grow from one category to the next without that up-front cost:
- Flow definitions any team can read and change, not one team’s programming language.
- Schedules and event triggers in the same engine from the first flow, so a nightly job can go event-driven later without a migration.
- A footprint small enough to run self-managed without a dedicated operations team.
- Independence from the cloud provider, the network vendor, and the BSS suite, so the platform outlives all of them.
A platform with those properties costs a single team nothing extra on day one. When the second category shows up, and in a telco it always does, it plugs into a control plane that is already running instead of into new glue code.
Unified Orchestration: One Control Plane Above OSS, BSS, and IT
Airflow, Cisco NSO, Temporal, and Camunda each earn their place, and none of them is the problem. The problem sits in the seams between them, where a flow leaves one tool and arrives in another through code somebody wrote once and now maintains forever.
Unified orchestration platforms address this gap. The category is young and the market is still small. Kestra is the platform I use to make the concept concrete here. I wrote separately about why this category is worth the move. Adoption in telcos follows the pattern above: one team, one category, one scheduler or pile of scripts replaced, and the control plane grows from there. What matters for a telco is the shape of the project rather than the company behind it. The core is open source under the Apache 2.0 license, the same license model that put Kafka, Kubernetes, and ONAP into operator stacks. Flows are declared in YAML that any team can read, schedules and event triggers run in one engine, and the workers run wherever the operator needs them.
What makes it unified is the task reach on top of that. One flow can run a pipeline step, provision infrastructure, call an activation API, wait for a partner confirmation, and hold for a human approval, in one definition and one audit trail.

Unified orchestration does not re-model every BPMN diagram and does not take work away from the network orchestrator. It adds one layer above, where nothing was accountable before. The next section draws the boundary precisely.
Where Workflow Orchestration Fits in the Telecom Latency Spectrum
With the control plane on the table, the first question every network architect asks is which paths it may enter. In telecom the answer is unusually clear, because the network has hard timing requirements that nobody argues about.
Latency Layers in a Telco Network
- Critical real-time, microseconds to milliseconds. The radio scheduler, the user plane, IMS call setup, and online charging at session start. Purpose-built network technology, often specialized silicon. Not a workflow engine, and not a general-purpose orchestrator of any kind.
- Streaming and APIs, milliseconds. Telemetry distribution, mediation, assurance, fraud scoring, and request-response calls including CAMARA network APIs. Apache Kafka and Apache Flink do this work, as covered in Open RAN and Data Streaming.
- Orchestration, seconds to days. Migration campaigns, billing cycles, settlement, change windows, approvals, dispatch, and calendars.

These are three different jobs, not three speeds of the same job. I made the same boundary argument for data streaming years ago in Apache Kafka is NOT real real-time data streaming, and for the shop floor in unified orchestration in manufacturing.
What a Unified Orchestrator Should Not Touch
Being specific here builds more credibility with network architects than any feature list.
A unified orchestrator does not replace NFV MANO or the O-RAN SMO, does not run a sub-second closed loop, and does not take over the catalog-driven decomposition logic inside an OSS service orchestrator. Nor does it move high-volume data: Kafka, change data capture, and the integration middleware around them stay exactly where they are. I mapped that tooling space in the Data Integration Landscape.
What it does own is the coordination around all of those systems, on a timescale of seconds to days, across teams that do not share a tool today.
What Unified Orchestration Looks Like Inside a Telecom Operator
The architecture argument only counts if it survives contact with what an operator actually runs. Three places decide that: the billing cycle, legacy network retirement, and regulated deployment.
The Billing Cycle and the Batch That Never Went Away
Conference talks cover 5G, AI, and autonomous networks. Nobody presents their bill run, yet it is where a large share of telecom revenue is actually produced. Mediation, rating, invoicing, dunning, revenue assurance, interconnect and roaming settlement, and wholesale reconciliation all run on cycles with hard dependencies and hard dates.
Charging itself moved to real time years ago. The cycle around it did not. A convergent charging system rates a session in milliseconds and still feeds a monthly process that takes hours and must not be rerun casually. The dependency graph of that process, the calendars, the reruns, and the operator interventions are an orchestration problem, not a data movement problem.
Modernization here follows the same route as anywhere else. The Strangler Fig pattern applies to scheduler migration exactly as it applies to applications. Move job chains to a modern orchestrator incrementally while the legacy scheduler keeps running what it still owns.
Copper Switch-Off and Legacy Retirement as an Orchestration Problem
Every operator in the world is retiring something large right now. Copper and PSTN lines, 2G and 3G radio, legacy platforms with a decade of accumulated dependencies.
The dates are real. In the UK, the PSTN is scheduled to shut down permanently on 31 January 2027, after one delay. Exchange-by-exchange stop-sell already covers millions of premises. The UK is one regional example of a global pattern. The proposed EU Digital Networks Act would add conditional copper switch-off deadlines across the bloc. Operators would have to file switch-off plans, get regulator approval, and report progress.

A Line Migration, Step by Step
- Eligibility and coverage check (data): fibre availability, line characteristics, and installed devices are joined from inventory and the data platform. A line that is not yet eligible is parked and re-checked automatically when the next fibre release lands.
- Customer notice, slot, and consent (business process): the customer is contacted, a slot is agreed, and the clock on the regulatory notice period starts.
- Order and service activation (application), in parallel with port and line configuration (infrastructure): the orchestrator creates the order, calls provisioning, updates charging and the catalog, and at the same time configures the new port and validates the change. Both branches have to complete before the cutover starts.
- Cutover and verification (infrastructure and data): the old service is decommissioned, the new line is tested, and the result is checked against the order.
- Engineer visit for exceptions (people): telecare alarms, lift lines, payment terminals, failed cutovers, and no-shows route to a field visit. The workflow waits, escalates on deadline, records who decided what, and retries the cutover once the visit closes.
- Case close and regulator report (data): confirmation flows back, the case closes, and the progress report to the regulator is generated from the same audit trail.
No single-silo orchestrator covers that sequence. Multiply it by millions of lines over several years, with parallel branches, exception paths, and a regulator watching the progress reports, and the coordination layer stops being a detail.
Sovereign and Hybrid Deployment: NIS2, Multi-Tenancy, and the Edge
In telecom, deployment constraints shape the platform decision before any feature comparison begins.
- Workers inside the operator’s own data centers. The control plane can run self-managed or as a managed service where policy allows. The workers that touch subscriber data and network elements stay inside the operator, and in some markets fully air-gapped. Lawful intercept and data retention obligations make a SaaS-only platform a non-starter for that part of the workload.
- Logical tenant isolation. Country units, wholesale and MVNO customers, shared service centers, and business units must be compartmentalized. Without logical multi-tenancy, operators deploy dozens of physical instances out of necessity and pay the operational cost forever.
- Fine-grained access control and compute isolation. Role-based permissions down to who may execute which flow, and isolated worker pools so one team’s heavy job does not starve another team’s critical one.
- Workflows as code, with tests. Version control, CI/CD, and tests before deployment. A configuration change that reaches a live network untested is an outage waiting for a date.
- Secrets from an external vault. Credentials for hundreds of network elements and partner systems cannot live inside the orchestrator and be rotated by hand.
- Lineage and a live asset inventory. NIS2 and national security rules put evidence requirements on operators. An orchestrator that knows every resource its workflows touch turns audit evidence into a query. The inventory of the orchestrator complements the CMDB rather than replacing it.

Hybrid execution belongs here too. Core systems stay in the operator’s own data centers for years while analytics and AI workloads run in the public cloud. Both sides need to be coordinated by the same control plane and land in the same audit trail.
Agentic AI in Telecom: Governed Across All Four Categories
Telecom is further along with operational AI agents than most industries, because it has the telemetry, the cost pressure, and a standards community pushing in the same direction.
The agent use cases with traction are operational. A fault triage agent correlates alarms and proposes a root cause before an engineer opens the ticket. An assurance agent enriches an incident with topology and change history. An energy agent proposes which cells to put to sleep and when. A care agent drafts the response and the credit calculation.
An Agent Step Is a Workflow Step
The governance requirement follows directly. An agent step needs the same operational rigor as any other step: inspectable runs, the ability to replay from a specific input to understand a decision, alerting when behavior drifts, and a human approval as a first-class step before anything consequential happens.
On a live network, consequential means irreversible. A wrong configuration change pushed at scale is an outage with a regulatory reporting obligation attached. The approval gate is not ceremony.

Governance extends to the platform’s own AI features. Operators routinely require the ability to restrict or disable functions that call external model APIs, down to specific roles, because subscriber data going to an external model provider is a compliance event regardless of which feature sent it.
Flows as Tools Instead of an MCP Server per Use Case
A pattern I see repeatedly in operators and telco IT service providers: every team that wants an agent starts by hand-building its own tool server. One for incident management, one for provisioning, one for reporting. Each one re-implements access control badly or skips it.
Exposing governed flows as tools removes that work. An agent calls a flow, the flow runs under the platform’s existing permissions, and every invocation lands in the same audit trail as a scheduled run. The engineering effort moves from building agent infrastructure to defining which flows an agent may call.
Why Agentic AI Is Not a Fifth Silo
The industry arrived at the same conclusion independently. The TM Forum’s AI-native extensions to its Open Digital Architecture, launched in 2026 as part of the Race to 2030, add a governed execution layer for AI agents inside the existing architecture. Policy control points and Model Context Protocol support are part of it. The industry chose a governed layer inside the operations platforms over a parallel agent stack beside them.
An agent that reads network data, calls application APIs, and hands work to an engineer crosses all four categories like every other flow. Bolting a separate agent platform next to the existing orchestrators creates the next generation of glue code, this time with a non-deterministic component inside it.
The Trinity of Modern Data Architecture in Telecom
Unified orchestration is not a pillar of its own. It sits inside process intelligence. The Trinity of Modern Data Architecture has three layers: process intelligence, event-driven integration, and trusted agentic AI.
Process intelligence covers three capabilities. Process mining shows how order to activate or trouble to resolve actually runs, including the rework nobody documented. Orchestration executes and governs what happens next. A decision gate keeps automation and agents inside auditable boundaries. The Process Intelligence Landscape maps the vendors across both halves.
Event-driven integration moves network and customer events in real time. Data streaming with Apache Kafka and Flink, event brokers, and the messaging middleware around them turn telemetry, mediation records, and OSS events into a shared backbone that OSS, BSS, and OTT partners can all consume. Process intelligence coordinates the work around those events, from the bill run to the migration campaign. Trusted agentic AI acts inside the boundaries that layer sets. A fault triage agent is only as good as the telemetry beneath it and the gate above it.
To follow this work across data integration, workflow orchestration, process intelligence, and trusted agentic AI, subscribe to the newsletter and connect on LinkedIn.