For years, voice occupied an uncomfortable position in communications compliance. Everyone knew the risk was there. Regulated firms were generating thousands of hours of calls, across trading floors, back offices, and client-facing teams. But the tools to manage that data properly were either too expensive, too labour-intensive, or too unreliable to make systematic supervision realistic.
So firms did what they could. They sampled. They stored. And they hoped.
That approach is no longer acceptable – and for compliance leaders who haven’t yet addressed voice properly, the window to act quietly is closing.
The Blind Spot Nobody Talked About
The reasons voice stayed unaddressed for so long were largely practical. As Daniel Yates, Voice SME at Global Relay, explains, the options available to firms were limited and costly.
“For a long time, the only options were to have people on desks doing what was known as random sampling of voice calls, which was pretty ineffective. It was very difficult for humans to random sample more than a small percentage of the overall number of calls.”
The alternative – investing in AI transcription – was theoretically available, but the reality rarely matched the promise. Early language models required significant training, heavy investment, and still delivered limiting results. The cost-benefit calculation rarely worked in voice’s favour, particularly when regulators weren’t yet pushing hard on enforcement.
That calculus has now changed entirely. The arrival of large language models capable of accurate transcription straight out of the box – with no training required – has removed the primary technical barrier that kept voice in the compliance blind spot.
“Since we’ve had the ability to use large language models for transcription, that has been the biggest game changer,” Yates says.
“These LLMs offer very high levels of accuracy that we’ve never seen before and there’s no need for training. You switch them on and off you go.”
Why Regulators Are Now Enforcing
The shift in technology has had a direct impact on regulatory expectations. For a long time, a store-and-forget approach to voice recording was broadly acceptable. Firms captured calls, stored them somewhere, and retrieved them if needed. The bar was low, and regulators largely accepted it.
That is no longer the case.
“The store and forget record keeping approach was acceptable to most regulators for a long time,” Yates explains. “But now, due to the advances in technology, regulators expect a proactively tracked, structured, and monitored approach. It’s no longer acceptable to just store some recordings on a server somewhere and hope that it’s working.”
The expectation now is that voice data is securely stored, reconciled, and actively monitored – and that firms can evidence those procedures to regulators on a regular basis. The shift from passive retention to active supervision is significant, and it applies broadly across regulated industries.
The Misconceptions Holding Firms Back
Despite the changing landscape, a number of persistent misconceptions continue to leave firms exposed.
The first is the assumption that voice is simply harder to manage than text-based communications – that messaging is a solved problem but voice remains out of reach. That gap has narrowed considerably. Advances in capture and transcription mean that voice monitoring and risk detection can now sit within the same compliance technology ecosystem as text-based communications. The technical argument for treating voice differently no longer holds.
The second misconception is more dangerous. Many firms assume their existing eComms archiving solution is providing voice coverage. It is not.
“Voice is its own data type with its own capture, retention, and retrieval requirements,” Yates notes.
“Treating it as an afterthought, or assuming it’s handled, is where firms tend to find themselves exposed.”
There are further assumptions worth challenging. The idea that only frontline traders need to be captured is no longer accepted by regulators. Nor is the assumption that only external calls fall within scope – internal communications related to transactions carry the same obligations. And the belief that the telecoms provider bears responsibility for compliance controls fundamentally misunderstands where accountability sits.
“Overall responsibility falls to the firm and its senior management,” Yates says plainly.
The Risks Voice Monitoring Actually Surfaces
Getting voice supervision right matters not just for regulatory box-ticking, but for the quality of risk detection it enables.
Voice channels have long been associated with financial misconduct risk – insider trading, market manipulation, improper client communication. The prominence of voice in roles where those risks are highest has always made it a priority channel, even when the tools to supervise it properly were lacking.
But the risk picture has broadened. Non-financial misconduct – bullying, harassment, toxic behaviour – is an increasingly explicit regulatory priority, particularly for the FCA, which has been clear that toxic culture breeds financial risk. Voice is a critical channel for detecting that kind of behaviour, in part because people often use it precisely because they believe it leaves no record.
The argument that a conversation “doesn’t leave a paper trail” has always been part of how misconduct on voice goes undetected. Modern surveillance tools, equipped with LLM-based detection that understands context and intent rather than relying on keyword lists, are changing that assumption fundamentally.
The Window Is Closing
The compliance case for voice supervision has never been clearer. The technology is mature, the regulatory expectations are explicit, and the misconceptions that allowed firms to defer action are running out of road.
Firms that have treated voice as an afterthought – relying on random sampling, assuming their eComms solution covers it, or simply storing recordings and hoping – are carrying a visible gap that regulators are increasingly equipped to identify.
As Yates puts it: “It’s no longer acceptable to just store some recordings on a server somewhere and hope that it’s working. The expectation is now that the data is securely stored and reconciled, but proactively monitored.”