Skip to main content

AI Monetization in Telecom: What's Real, What's Not

Get our latest reports straight to your inbox. Subscribe
Share this article

At the Telco AI Forum this week, I moderated a panel on AI monetization with Shahed Mazumder, Senior Director for AI-RAN and Product Strategy at SoftBank, and Rick Lievano, CTO Telecom at Microsoft. The question we set out to answer: which parts of the AI monetization story have actually moved from strategy deck to P&L?

 

Cost savings are happening. New revenue is not — yet.

Both speakers started from the same baseline: in 2026, the value operators are capturing from AI sits almost entirely in the cost savings bucket. Predictive maintenance, energy efficiency, AI-optimized RAN operations — these are real, measurable, and increasingly embedded in how operators run their networks. Shahed, who co-authored the AI-RAN Alliance’s new commercialization white paper presented in Phoenix two weeks ago, was direct about this: the TCO savings story is in motion; the new revenue story is nascent (but without a doubt, it’s making solid progress - see below for the 5 revenue buckets).

Rick added an important qualifier: monetization isn’t just cost saving. That distinction matters more than it sounds, because most of the industry’s AI narrative has jumped straight to new revenue without establishing the baseline. Operators who haven’t yet extracted value from efficiency gains are not positioned to sell AI services externally.

 

Five revenue buckets — two are near-term

The AI-RAN Alliance framework Shahed presented identifies five new revenue streams unlocked by AI-RAN: inference as a service, physical AI and robotics, network APIs for AI services, personalized AI services, and sovereign AI hosting. Of these five, two are further along commercially than the rest.

Inference as a service is the most concrete. SoftBank has already deployed a large GPU fleet in Japan — as the top 10 data center providers in the country, the first customer on NVIDIA’s GB200 GPU infrastructure — starting with training workloads and now expanding to inference, with general availability planned for October 2026.

Sovereign AI hosting is the other near-term story, and SoftBank’s position in Japan illustrates why vertical integration matters here. Data centers, RAN, GPU fleet, and enterprise relationships all under one roof — plus  Infrinia, their sovereign cloud product launched in April 2026 — making the story even more compelling is Japan’s first local Large Language Model (LLM) “SaraShina” developed by SoftBank. Across a number of markets such as Indonesia, local LLMs are applied to government services, where data residency and privacy requirements make a locally-governed LLM a commercial differentiator, not just a compliance checkbox. That vertical integration is what turns “sovereignty” from a regulatory story into a commercial one.

Physical AI and robotics, although trailing the above two commercially, is dictating conversations everywhere with the proof of concept already here. SoftBank has run a demo with Yaskawa Denki, one of the top four robotics manufacturers globally, showing what differentiated connectivity — the network actively managing latency and reliability for robotic workloads — can do that best-effort connectivity cannot.

 

Voice AI: the network that understands the call

 

Rick brought a use case that doesn’t fit neatly into the five-bucket framework but may be the most commercially immediate: running AI directly in the live voice path.

Working with partners including  Norwood Systems, Microsoft is using Azure transcription to apply AI to live calls in real time. The use cases include fraud detection — flagging in the moment when a caller is being asked for their PIN — live translation, and customer service coaching.  T-Mobile’s live translation service is a live example of this category in the market today.

These services require being embedded into network infrastructure — IMS, gateways — not bolted on top. For operators, that means they need to be exposed as a consumable capability: APIs, premium features, packages with a clear link to business value. Real-time translation only becomes a monetizable product when it drives a measurable outcome — reduced churn, higher ARPU — and that requires commercial discipline as much as technical execution.

 

Sovereignty: trust is the asset, not the real estate

The sovereignty conversation predictably took up airtime. But the more useful framing that emerged wasn’t about data residency regulation — it was about trust.

Microsoft’s  Azure Local Disconnected Operation (ALDO for short), announced at MWC 2026, lets operators run AI infrastructure entirely on-premises, addressing regulatory requirements without ceding control. The operator stays in charge; Microsoft provides the underlying software model. The operator brings what Microsoft cannot easily replicate: edge locations, regulatory standing, and the trust of being the incumbent connectivity provider in a given country.

That trust — being the guaranteed, walled-garden provider for data that cannot leave the country — is the asset. Privacy is a major factor in how this plays out commercially, particularly in markets where government and enterprise customers have hard requirements around data sovereignty.

 

The agentic layer: distribution, not infrastructure

 

Microsoft also demoed an agentic marketplace framework — a blueprint for operators to build AI-agent-led customer experiences. The demo had a customer guided through a device purchase by an AI agent that asks lifestyle questions and makes recommendations accordingly. VEON already operates a marketplace in this vein.

This sits at the intersection of personalized AI services and network APIs, and it gives operators the ability to move from static service offerings to dynamic, outcome-based models. But it’s worth being clear about what it is: a retail distribution model, not a network monetization model. Operators who pursue this are competing on customer relationship and billing trust, not on network differentiation.

Rick was direct about the risk: telcos face a real window of commoditization if they don’t act. The asset they’re bringing to this layer — existing customer base, billing relationships, regulatory standing, and investment-grade borrowing capacity to deploy GPU fleets — is genuinely valuable. But it’s only valuable for a period. How the value chain is structured, and whether hyperscalers capture the underlying model economics while telcos capture the application and marketplace layer, remains an open question.

The AI-RAN Alliance, now 140 members in two years, exists in part to prevent fragmentation across this emerging stack. At 27 billion tokens processed per day, the tokenomics of AI inference — and tools like Microsoft’s Fabric and Foundry IQ that help optimize token consumption — are becoming a meaningful part of how operators think about the economics of running AI workloads at scale.

 

What the data will tell us

At Opensignal, we track how network performance translates into market share and subscriber behavior. The AI monetization thesis ultimately rests on a chain: better network operations → better user experience → measurable commercial outcomes. The cost savings part of that chain is already visible in how operators are managing their networks.

Whether the new revenue side materializes — and shows up in the metrics that matter to subscribers — is what the next 12 to 18 months will answer. The operators with clear proof points, and the discipline to build on them rather than leapfrog them, are the ones to watch.

You can watch the full recording now by following this link: https://attendees.bizzabo.com/804288?source=magic_link