How Realtime Voice AI Is Giving Developers More Control Over Delivery
How Realtime Voice AI Is Giving Developers More Control Over Delivery
Natural-language direction lets developers shape delivery during live interactions and choose the model that fits each application.
Artificial intelligence (AI) can generate dialogue in seconds, leaving voice technology with the job of turning those words into a believable conversation. How a line is spoken becomes especially important when no one can predict the exact sentence ahead of time. Inworld has built its latest text-to-speech technology around giving developers more control at that moment.
Take a simple line like “I knew you’d come back.” A developer could direct the voice to deliver it softly, with the relief of someone who has been waiting for hours. The exact same words could carry nervous excitement in another scene, complete with a breath before the sentence or a laugh afterward.
TTS-2, a next-generation conversational text-to-speech voice model launched by Inworld, lets developers describe that delivery in natural language, the company says. They can also add pauses and nonverbal sounds including a laugh or a sigh. Voices can be designed from a written description or cloned from five to 15 seconds of authorized audio, then carried across different languages while retaining the same identity.
Choosing what each experience needs
Speed becomes especially noticeable when someone is talking back and forth with an AI. Inworld offers a second model, Realtime TTS-2 Flash, for products handling high volumes of interactions or putting a greater emphasis on fast responses and lower serving costs.
Flash keeps the steering, cloning, and language coverage available in TTS-2. Developers can send an identical line with the same delivery directions to each model, choosing TTS-2 for its stronger emphasis on voice quality or Flash for applications where milliseconds have a larger effect on the experience.
Realtime TTS-2 Flash has a reported 25-millisecond time to first byte. TTS-2 comes in under 100 milliseconds at P99 according to Inworld’s launch figures. Coval, an AI-native simulation, evaluation, and production monitoring platform designed to test and validate voice and chat AI agents, also measured Flash at 25 milliseconds under its published benchmark conditions and methodology.
How realtime voice AI gives developers more control over delivery
The difference between speech models can reach further than how much someone likes a voice sample. Talkpal, a language learning platform with more than 5 million learners, tested Inworld’s realtime speech technology for 4 weeks. Its A/B test found feature usage reportedly rose 7% and retention increased 4%, alongside a 40% reduction in TTS costs.
Interactive storytelling puts a different kind of pressure on generated speech. ARX Media CEO Louis Muk has spent years pursuing voices that can cross the line from technically impressive to believable for ISEKAI ZERO. He describes TTS-2 as bringing generated speech closer to the moment when a character talks and the technology behind the voice fades from attention.
LiveKit CTO David Zhao points to responsiveness as another part of that experience. Live conversation moves through several processes before a spoken answer comes back, so expressive speech has to work within the speed of the interaction itself. LiveKit makes Inworld’s realtime speech products available through its Agents platform.
Changing the model without rebuilding the app
Developers also have to think about what happens after a model ships. Inworld says a speech model update won a neutral A/B test for one customer even though the endpoint stayed unchanged. The customer could receive the improved model through its existing integration.
That setup also gives Inworld room to route requests between models according to what an application needs. Speech is one part of a business that also includes language-model serving, routing, realtime APIs, and compute for consumer applications.
Live AI leaves plenty of room for surprise. A conversation can change direction in seconds, and the voice has to change with it. By giving developers greater control over delivery in real time, TTS-2 brings more of that spontaneity within reach.
Entrepreneur Media was not involved in the creation of this content.