BUILT BY YAHOO

What It Takes To Build AI Voice That Sounds Natural

Talking to an AI assistant feels simple: You ask a question. It answers.

Building one that can carry on a natural conversation is anything but. A voice experience has to hear accurately, respond fast enough to feel conversational, protect sensitive data and—quite literally—decide what the product should sound like.

Those were some of the challenges Yahoo's product teams confronted when they set out to bring voice to Yahoo Scout, Yahoo’s AI answer engine. Working with the AI audio platform ElevenLabs, the team had to make a series of choices that increasingly confront any company building for the voice era: where to rely on specialized AI partners rather than build from scratch, how to design user data protection into an experience meant to feel effortless, and how to give an AI voice a personality that feels native to the product and brand.

“Voice is a very different interaction pattern than text,” says Matt Linford, a senior engineering director at Yahoo. “The way people use it is different. You might be pulling out your phone to ask a quick question, or you’re hands-free. We need to think differently about how the experience works and what makes a conversation feel natural.”

Voice is a very different interaction pattern than text.

Know what not to build

In a fast-moving area like AI voice, building everything yourself isn’t necessarily an advantage. The fastest path to a differentiated product is identifying which capabilities are unique to your experience and which are better handled by a specialist. That was the calculation Yahoo made with voice.

Creating a natural conversation requires more than converting speech to text and generating an audio response. The experience has to account for background noise, interruptions, and the ability to jump in while the AI is still talking. Rather than spend months recreating those foundational capabilities, Yahoo partnered with ElevenLabs, whose conversational AI technology already addressed many of them.

“ElevenLabs had already invested deeply in solving many of the foundational challenges of conversational voice,” Linford says. “Rather than recreate those capabilities ourselves, it made sense to build on what they had already developed and focus our efforts on the Yahoo Scout experience.”

The partnership went much deeper than plugging in an API. Yahoo integrated ElevenLabs’ React Native SDK directly into the Yahoo Search mobile app, using AI coding tools to accelerate integration. The primary technical challenge was connecting ElevenLabs’ public cloud infrastructure with Yahoo’s strictly secured internal network. To avoid compromising on security or redesigning core architecture, Yahoo engineers collaborated with ElevenLabs to build dedicated gateway endpoints through the corporate firewall. This allowed real-time audio streams to pass back and forth safely without adding latency.

Treat voice as data, not just an interface

Voice interactions raise design questions that text search does not. Audio carries more than a typed query does, and the decisions about what gets processed, what gets kept, and in what form have to be settled before a feature ships rather than retrofitted afterward. 

For Yahoo Scout, those decisions were settled during design. Voice sessions are handled under Yahoo's published privacy practices for Yahoo Scout, and the third-party providers that process them are contractually restricted from using that data for their own purposes, including model training. How Yahoo handles data from the feature is described in the Yahoo Privacy Policy. 

Entering a voice conversation is an explicit choice. A user starts the conversation, selects a voice, and grants microphone access on the device. For product teams generally, that's the shift voice requires: treating retention and processing as architecture decisions, made once and documented, rather than as disclosures added to the interface later.

Make voice part of the product identity

Once the backend architecture and privacy parameters were established, the team faced a creative challenge: What should the product sound like?

‍

They started with two different voice and tone options, but the team also saw an opportunity to make voice feel more distinctly Yahoo. Working with ElevenLabs’ voice-cloning technology, they created a custom voice based on Yahoo co-founder Jerry Yang—with his permission and strict prompt guardrails baked in. It captures not just the sound of his voice but elements such as his pacing, pauses, and speech patterns.

The choice was partly an experiment. Rather than immediately pursuing a wide range of recognizable voices, the team wanted to test the technology with someone closely associated with Yahoo and learn how users respond to an AI experience that sounds familiar.

“It was really a proving ground for us,” Linford says. “How do people react to it? How do they engage with it?”

The experimentation extended beyond the assistant itself. The teams also explored recreating Yahoo’s signature yodel with generative voice technology, another way of asking how recognizable elements of a decades-old brand might translate into an AI-native experience.

Linford believes familiar voices could change the relationship people have with conversational products. Users may be more comfortable talking—and potentially more willing to engage—with a voice they already know and trust than with a generic AI voice.

As voice becomes an interface, how an AI sounds can become as much a product decision as what it says.

Key takeaways

Outsource baseline tech, build your differentiator. Don't waste cycles reinventing complex AI infrastructure. Partnering directly with specialized engine vendors allows your engineering teams to bypass foundational challenges and focus strictly on creating unique, native product experiences.

Make data handling an architecture decision. Decide what a voice feature processes, what it stores, and in what form before you build it. Those choices are far harder to change once the experience ships.

Treat AI voice as a core brand asset. Audio interfaces require a distinct identity. Cloning familiar voices or legacy sonic brand markers creates immediate trust, transforming generic AI interactions into memorable, differentiated user experiences.

‍