Your AI comms bill is about to look very different

Your AI Comms Bill Is About to Look Very Different: What UK Businesses Need to Ask Their Vendors

For most businesses, AI spend used to be theoretical. Pilot projects, proof-of-concept budgets, a line in the innovation roadmap. By 2026 that has changed. AI features are live inside the tools staff use every day — phone systems, contact centre platforms, meeting software — and the costs have moved from experimental to operational. The bills are arriving, and a lot of IT Directors are finding they weren’t budgeted for what AI actually costs at scale.

The Training/Inference Problem Nobody Explained

When organisations talk about AI costs, they tend to think about the upfront work: building models, training them, buying into a platform. Those are real costs, but they’re largely fixed. The cost that compounds is inference — the compute that runs every time an AI feature does something. Every call transcribed, every meeting summarised, every customer interaction routed through a conversational AI workflow: that’s inference. It runs continuously, at volume, and it scales directly with usage.

Aragon Research estimates that by 2026, inference accounts for two-thirds of all AI computing power consumed by enterprises. Lenovo’s 2026 Total Cost of Ownership analysis found that for sustained AI inference workloads, on-premises infrastructure can reach cost breakeven against cloud providers in as little as four months. Those aren’t abstract figures. They describe what happens when a platform that looked affordable in a pilot runs at production volume across a business.

Most organisations didn’t choose the wrong platform. The cost model was never properly explained to them — because vendors have little incentive to explain it.

What Happens When Usage Doubles

Here’s a practical test. Ask your communications vendor a single question: what happens to my costs when usage doubles?

A platform built with AI inference economics in mind will have a clear answer. The pricing model accounts for scale. The architecture doesn’t rely on passing every request through an external cloud API at per-call rates that compound with adoption. The vendor can tell you, plainly, what your bill looks like at twice the current volume.

A platform with AI bolted on — an existing product that added AI features as the market demanded them — will give you a murkier answer. The inference costs sit upstream, in a dependency on a third-party model provider, and the vendor is passing those costs through with a margin on top.

AudioCodes make this point clearly in their own analysis of inference economics, drawing on Aragon Research. Their Meeting Insights On-Prem product is a useful illustration: meeting intelligence software designed from the ground up to run within a customer’s own infrastructure, specifically to avoid the runaway inference costs that come with cloud-dependent AI at scale. The architecture reflects a deliberate choice about where inference runs and who bears the cost. That’s what the vendor question is designed to reveal.

Why Regulated UK Businesses Face a Harder Version of This Problem

For firms in financial services, legal, and professional services more broadly, inference economics isn’t only a cost question. It’s a compliance question.

Under UK GDPR and the Data Protection Act 2018, organisations are data controllers. When AI inference runs on an external cloud platform — processing call recordings, meeting transcripts, customer interaction data — the firm may not have adequate visibility or control over where that data is processed and by whom. The Data (Use and Access) Act 2025, which came into force in February 2026, introduced further clarifications on automated processing that firms in regulated sectors need to account for.

FCA-regulated businesses and those operating under legal professional privilege obligations face a sharper version still. The cheapest inference option on the market may simply not be available to them. That makes the architecture question — where does the AI actually run, and on whose infrastructure — a material factor in platform selection, not a technical footnote.

This intersects directly with a risk many professional services firms are already carrying. Shadow AI — staff using unsanctioned tools without IT oversight — is already present in most firms, and communications data is among the most sensitive it touches. Read our related post on the compliance implications in more detail: Shadow AI in Professional Services: Why Meeting Data Is Your Biggest Compliance Blind Spot.

The Portfolio Question, Not the Product Question

The right frame for this isn’t cloud versus on-premises. Most organisations will end up with both, and that’s probably correct. Cloud infrastructure suits variable, experimental, and frontier-capability workloads. Predictable, high-volume, data-sensitive workloads — the kind that run continuously inside a communications platform — are increasingly better served by architecture that doesn’t meter every inference call back to a hyperscaler.

A useful rule of thumb: when cloud AI costs approach 60–70% of what equivalent on-premises infrastructure would cost over the same period, an on-premises evaluation is worth running. For communications workloads running all day, every working day, that threshold arrives sooner than most budget forecasts assume.

When a platform renewal comes up — whether that’s your voice infrastructure, your contact centre, or the AI layer sitting across both — the inference cost question belongs on the agenda alongside the feature comparison. Ask what the architecture looks like. Ask where the AI runs. And when the vendor talks about scale, ask for the numbers. The answers will tell you more about long-term cost than any headline price.

If you’d like to work through those questions against your own communications setup, our team of Solutions Consultants be happy to help.

About Marlin Communications

Marlin Communications is committed to providing expert guidance on Zero Trust solutions.

Marlin Communications is an independent, single-source provider of business communications & collaboration solutions including voice, data, mobile, video, network security and contact centre technology for businesses of 50 – 5,000 staff.

We operate throughout the UK – with global reach – and our own, on-premises, 1,000 ft² Technology Suite at our Bath office, where we host regular events and showcase technology solutions for our clients. Contact us for your free comms audit or product demo.

Marlin Communications is ISO 27001 certified by BSI under certificate number IS795313.

Get the latest tech news & reviews – straight to your inbox

Sign up to receive exclusive business communications, tech content, new tech launches, tips, articles and more.

SUBSCRIBE NOW

Click here to follow our LinkedIn company  page and stay up-to-date with our LinkedIn newsletter