Table of Contents
Today Deepgram welcomes Ed Anuff as Chief Product Officer. Ed brings more than 30 years of experience building and shaping products across data, APIs and enterprise platforms. Most recently, he served as Chief Product Officer for IBM watsonx.data, leading IBM’s product strategy for data after its 2025 acquisition of DataStax, where he previously held the same role. Earlier in his career, Ed helped launch one of the first internet search engines at Wired, founded enterprise portal pioneer Epicentric, and guided Apigee through its IPO and acquisition by Google.
Ed has closely followed Deepgram’s evolution from a speech-to-text provider into a voice AI platform built for production-scale agentic applications. We sat down with him to discuss what drew him to the company, his priorities as CPO, and where he sees voice AI heading next.
You’ve been close to Deepgram for several years. Why did you decide to join now?
I first met Scott back in 2021 and have been proud to be an advisor to Scott and the company since then. Back then, Deepgram was one of the first AI companies developing great speech-to-text, but that was fairly straightforward compared to what we see today.
The timing to join felt right for a few reasons. First, voice is becoming the primary interaction layer for AI, and Deepgram is entering a pivotal stage of growth and operationalization. Secondly, the technical foundation is here, and now we have an opportunity to turn it into products that can define this next phase of the market. I’ve spent a lot of time helping pre-IPO companies get to the next level of scale, but even more important to me is the great alignment with where I’ve focused from a product and technology standpoint.
A consistent thread throughout my career has been working on new interaction layers. The web changed how people accessed information. Mobile changed where and when people could interact with software. APIs changed how applications connected and how digital experiences were built. I see voice as the next major shift because it changes how people provide context to AI and allows those interactions to happen anywhere.
You’ve said that “voice is the new keyboard.” What do you mean by that?
I strongly believe that voice will become the dominant way people interact with AI within the next two years.
The keyboard has been the default interface for decades, but it keeps you tied to a screen and requires you to stop what you’re doing. Voice removes that constraint. Like mobile and cloud, it makes technology more accessible wherever you are and whatever you’re doing.
Voice also gives AI more of the tacit context it needs to be useful. Most people don’t naturally write long, perfectly constructed prompts. We’re much better at explaining what we need by talking through it, adding context, and clarifying as we go.
Voice gives AI far more context than the words alone. It can capture whether someone is speaking quickly or slowly, where they pause, how certain they sound and how their meaning changes as they explain or correct themselves. Deepgram can capture those signals with low latency and, through text-to-speech, close the loop by responding in voice.
That’s why I think voice AI is significantly underestimated right now. It’s not just another way to give AI a command. It creates a more natural and contextual way to interact with it, and I believe that will ultimately make it the primary interface for AI.
What has shifted in the industry to make this moment possible?
Voice agents are starting to shift from interesting demos into products people actually use.
For a long time, the industry was focused on individual pieces of the stack: better transcription, a better language model, or more realistic speech. Those pieces still matter, but customers now need complete systems that work reliably in real-world environments.
At the same time, agentic AI has made agents significantly more capable of actually doing things. Voice becomes much more valuable when a conversation can lead directly to an action or transaction.
What separates a compelling demo from a voice product people can actually rely on?
A demo happens in a controlled environment. A production system has to work over and over again for different people, on different devices, and in messy environments you can’t control.
It needs to understand different accents and languages, to hear in a noisy room, to know when someone has stopped speaking and when they are just pausing to think, and to be able to recover from interruptions without losing the thread of the conversation.
Latency is also incredibly important. Even small delays compound across the stack and make the interaction feel unnatural. You have to balance speed, accuracy, and conversational quality across the entire system.
That’s where testing, observability, and infrastructure become essential: when something goes wrong, developers need to know where and why it happened.
How is Deepgram approaching the broader voice AI lifecycle?
Deepgram started with a strong speech-to-text foundation, which gave the company a deep understanding of the speed, accuracy, and reliability that real-time voice applications require.
Now we’re building across the broader lifecycle, including text-to-speech and conversational voice models. The goal is to help developers build systems that don’t just listen and respond, but can understand context and react as a conversation unfolds, moving voice AI closer to a truly real-time, continuous experience.
As open-source frameworks and models become a bigger part of how developers build AI products, we want to be thoughtful about where Deepgram can add real value. The goal is to give developers more flexibility while building on what Deepgram already does well: production-grade voice infrastructure that is fast, accurate, and reliable. We don’t want to force everyone into a rigid, one-size-fits-all stack. We want to give teams the pieces they need and make the path from experimentation to deployment much easier.
What will you be most focused on as CPO?
What sets Deepgram apart from other voice AI is how the team is overcoming the robotic voice responses, and that is focusing on emotionally resonant conversational experiences. My first priority will be to continue building out our text-to-speech capabilities so that these conversations feel more and more natural to the user. The second will be to connect voice much more tightly to agents that can complete real tasks and transactions. The third, which is still taking shape, is defining our open-source strategy.
Across all three, my focus is to make it easier for developers and enterprises to build, deploy, and scale voice agents. Deepgram has real technical advantages, but those advantages have to translate into clear value for our customers.
That means staying close to the people building with our products, while also looking ahead to where voice AI is going next. As the market evolves, we want to anticipate the capabilities developers will need – from more continuous, context-aware interactions to tighter connections between voice and reasoning – and make sure our customers can take advantage of those advances as quickly and reliably as possible.
What do you think voice AI will look like two years from now?
Voice will no longer be treated as a feature added to an existing product but instead be the default interface for how we use AI both professionally and in our own lives.
People will talk to AI while they are working, driving, walking, or moving between tasks. And those interactions will be richer because voice provides more context. They’ll also be more useful because the agents on the other side will be able to take on more and more actions on our behalf.
Deepgram has spent years building the technical foundation for that future. I joined because I believe the next opportunity is to push beyond today’s voice stack – toward systems that can follow context as it unfolds, respond more naturally, and become a more capable interface between people and AI.









