Voice agents in practice
Phone scheduling, customer outreach, surveys, market research, shopping, and other deployed voice interfaces—along with whether they deliver useful outcomes.
40%
Best tweets about AI Voice
Browse the best tweets about AI voice, covering speech synthesis, cloning, voice agents, dubbing, audio models, creator workflows, and safety.
AI voice models, speech generation, cloning, agents, dubbing, audio quality, workflows, consent, safety, and practical applications.
Original Xholic analysis
AI voice discussion centers on practical agents, open/local tools and more expressive speech. Enthusiastic cloning and dubbing demonstrations sit alongside challenges about real-world outcomes, conversational timing and consent. The supplied analytics give tutorials a high median score, but only two posts carry that format.
62% of posts
All-time engagement
52% of posts
Published in 90 days
Conversation map
Phone scheduling, customer outreach, surveys, market research, shopping, and other deployed voice interfaces—along with whether they deliver useful outcomes.
40%
Frameworks and workflows that connect transcription, language models, speech generation, streaming, tools, deployment, and evaluation.
40%
Realistic text-to-speech, emotion and pacing controls, pronunciation accuracy, and low-latency generation.
36%
Open-weight speech models, self-hosting, licensing, privacy, and competition with subscription voice services.
30%
Speech-to-speech reasoning, full-duplex dialogue, interruptions, turn-taking, and the trade-off between intelligence and response speed.
30%
Short-sample cloning, preserving a speaker’s style, and using personal voice replicas in content and agents.
24%
Translation, localized narration, cross-language voice continuity, and expressive multi-speaker video dubbing.
16%
Speech recognition, diarization, speaker isolation, and reliable agent input in crowded or noisy environments.
14%
Tone and stance
Performance benchmark
Posts with media make up 90% of this collection. Their median all-time score is 15.3, compared with 2.41 for text-only posts.
Format mix
Consensus and debate
Shared view
Builders emphasize interruption handling, precise pronunciation and clean transcription as requirements for usable agents. These posts locate failures across the audio pipeline, rather than treating speech realism alone as sufficient.
Shared view
Posts describe replacing a paid transcription subscription, self-hosting a voice platform and switching models over licensing. The shared appeal is control over deployment, spending and permitted use—not a demonstrated universal advantage over hosted tools.
Shared view
Cloning and dubbing posts emphasize preserving pacing, emotion and speaker identity across content and languages. A YouTube viewer's account adds a reception-side example of dubbing that sounded natural to them.
Open debate
A home-services operator celebrates an agent scheduling an appointment; a satirical critique argues bookings can still become no-shows. Olivia Moore reframes the benefit around performance and business expansion rather than assuming cost reduction.
Open debate
Cloning demonstrations celebrate personal rhythm and convincing delivery. Other posts challenge broader human-like claims, pointing to broken turn-taking and an architectural gap between sequential agents and people who think while listening.
Open debate
Short-sample cloning is promoted as accessible and locally deployable. Cautionary posts raise impersonation risks, while the Ryza announcement describes purpose-recorded actor audio and royalties. These are contrasting deployment approaches, not evidence that safeguards are universal.
What performs
The supplied analytics assign TUTORIAL a median all-time score of 462.48 across 2 posts. Those examples combine a full-stack build with a voice-cloning comparison; the small category does not establish that tutorials reliably outperform other formats.
Open and local voice AI has a supplied theme median all-time score of 36.047. The local-transcription post is the highest listed outlier at 879.9, and the open-source platform tutorial scores 707.56. These standout posts do not establish why audiences engaged.
Statistical standouts
Creator landscape
The five most represented creators account for 20% of the selected posts.
1. Artificial Analysis
@ArtificialAnlys
2 posts
2. divyansh tiwari
@DivyanshT91162
2 posts
3. kwindla
@kwindla
2 posts
4. LiveKit
@livekit
2 posts
5. Olivia Moore
@omooretweets
2 posts
6. Shushant Lakhyani
@shushant_l
2 posts
Artificial Analysis contributes 2 posts, with a supplied median all-time score of 43.37. Its coverage distinguishes speech reasoning from conversational dynamics and reports the reasoning–latency trade-off, offering a counterweight to single-number quality claims.
LiveKit contributes 2 posts, with a supplied median all-time score of 135.38. Its examples focus on distinguishing interruptions from incidental sounds and fixing domain-specific pronunciation, connecting infrastructure releases to identifiable user-facing problems.
Since the previous snapshot
Themes, sentiment, stance, and post format are classified per tweet. All counts, shares, medians, creator concentration, freshness, and performance comparisons are then calculated directly from the published snapshot.
Xholic's all-time score compares engagement while accounting for reach, post age, and creator consistency. It is used for relative comparisons within this collection.
This report analyzes the exact 50-post snapshot shown below. AI identifies editorial categories and drafts explanations; all statistics are calculated from the snapshot, and every narrative claim is checked against cited posts before publication.
Best AI Voice tweets
Ranked 01–50
@thepatwalls ·
I'm bullish on open source AI. Was paying $15/month for a popular AI voice to text tool. But switched to an open source one where you download the model to your computer, and everything is done locally. It's better, faster, more private, and it's free! I wonder what other subscriptions I can get rid of?

@codewithantonio ·
🔊 Introducing Resonance, a full-stack AI voice platform like ElevenLabs No, this isn't my SaaS announcement, it's a completely free, open source tutorial 😎 🎤 Clone voices from audio samples or record your own 🎙️ Self-host an AI voice model with @modal 📦 Store audio with @Cloudflare R2 🔐 Authentication & multi-tenant orgs with @clerk 🗄️ PostgreSQL database with @prisma ORM 💳 Pay-as-you-go billing with @polar_sh 🪲 Each pull request reviewed with @coderabbitai 🚀 Learn to deploy on @Railway 📊 Monitor errors and logs with @sentry 🎨 @shadcn UI + @tailwindcss v4
@garrytan ·
GBrain just shipped v0.40.0 gives your OpenClaw/Hermes Agent + GBrain a voice agent. It's based on Gemini Live. (Thanks @demishassabis it's amazing) Large context, great tool use, full brain access. Mars is a friend, Venus is your EA. My open source gift to you.

@livekit ·
How can a voice agent tell when you’re actually interrupting it? VAD is too sensitive—laughs, “mm-hmm,” or a sneeze shouldn’t stop the agent. We trained an audio model for adaptive interruption handling so agents can distinguish real interruptions from noise.
@heyshrutimishra ·
I cloned my voice in 10 seconds and I'm still not over it. Not a generic AI voice. Not a "slightly sounds like me" voice. My timing. My pauses. My quirks. I ran it against the most popular TTS tool out there. Wasn't expecting much of a difference. There was a huge one. The reason most AI voices feel slightly off isn't audio quality. It's that they never sound like they're thinking. Real speech has micro-pauses, irregular pacing, emphasis shifts… because a brain is working in real time. Somehow Lightning V3.1 from @smallest_ai captured that. The clone doesn't just copy my timbre. It copies the irregularity. The small imperfections that make a voice feel like a person. Watch the comparison at the end. You'll hear exactly what I mean. For creators: one piece of writing, your voice, podcast clips, reels, multilingual versions. At scale.
@sentient_agency ·
R.I.P ElevenLabs Open-source voice AI is getting way too good now. OmniVoice just dropped with: - 600+ languages - zero-shot voice cloning - voice design from text - pronunciation control - non-verbal sounds like [laughter] - inference up to 40x faster than real time The wild part is not that it can clone voices. We already knew that was coming. The wild part is that it makes one voice travel across hundreds of languages. A creator can record once and localize everywhere. A teacher can turn one lesson into 20 languages. A founder can make product demos for every market. A YouTuber can dub their entire channel without hiring a studio. This is where paid voice tools start looking fragile. Because once open source gets “good enough,” the business stops being magic. It becomes packaging. And packaging is a very dangerous moat.

@tannerdripjobs ·
Are you kidding me?! 🤯 Last night at 9:40pm my house painting business got a phone call. Our @routemize AI Voice agent handled the entire call, end to end and scheduled a perfect appointment 7.1 miles away from an existing appointment scheduled that same day. Here’s what would’ve normally happened: 1. Call goes to voicemail 2. Admin calls in morning 3. Admin fumbles trying to find a spot for the customer. Doesn’t have time to see which spot is closest while on the phone 4. Picks a slot that “feels right” 5. *IF the customer even answers in the morning Routemize will be the favorite app of anyone who runs a home service business. API friendly and can plug into any software stack. Preferably @dripjobs 😉

@heygurisingh ·
Alibaba just open-sourced the world's FIRST AI dubbing model that handles multi-speaker scenes. It's called Fun-CineForge and it nails what every other dubbing model fails at: Lip-sync. Emotion. Voice stability. Timing. Across multiple characters. Not a single open-source model could do this before. Here's how it works: ↓ @Ali_TongyiLab @AlibabaGroup
@wildmindai ·
You know why AI voice assistants feel robotic? Not the synth voice- it's that polite walkie-talkie turn-taking. Real humans interrupt & talk over each other. Sommelier. Scalable open multi-turn audio pre-processing. - preserves overlapping speech and backchanneling for full-duplex SLMs; - cuts WER by 37% in noisy environments - ensemble ASR via Whisper-v3, Canary, and Parakeet; - metadata via Qwen3-Omni-Captioner. https://t.co/1MYxrlgQ0j

@shedntcare_ ·
Most AI voice agents still assume everyone speaks English. I wanted to test what happens if a voice agent can understand and respond across dozens of languages without rebuilding the entire stack. So I connected Krisp's new Voice Translation API. Here's what happened 👇

@ArtificialAnlys ·
NVIDIA has released Nemotron 3 VoiceChat! A ~12B parameter Speech to Speech model that leads our open weights Conversational Dynamics vs. Speech Reasoning pareto frontier Understanding Speech to Speech model performance is multidimensional - two key and distinct dimensions are raw intelligence and conversational dynamics: how well a model handles the natural rhythms of human conversation such as turn-taking, interruptions. Amongst full duplex open weights models, NVIDIA’s new Nemotron 3 VoiceChat, V1, leads in balancing these dimensions, setting itself apart from other models on the Conversational Dynamics vs. Speech Reasoning pareto frontier. Key benchmarking results: ➤ Conversational Dynamics (Full Duplex Bench): Nemotron 3 VoiceChat (V1) scores 77.8%, second among open weights speech to speech models behind NVIDIA's own PersonaPlex (91.0%) and ahead of FLM-Audio (62.0%), Moshi (61.0%) and Freeze-Omni (58.7%) ➤ Speech Reasoning (Big Bench Audio): Nemotron 3 VoiceChat (V1) scores 29.2%, second among open weights speech to speech models behind Freeze-Omni (33.9%) and well ahead of PersonaPlex (12.6%), FLM-Audio (5.3%) and Moshi (1.7%) ➤ Pareto leader: While Freeze-Omni leads on speech reasoning and PersonaPlex leads on conversational dynamics, Nemotron 3 VoiceChat (V1) is the only open weights model that performs amongst the top 3 on both - making it the clear leader on the pareto frontier between these two critical dimensions ➤ Larger than other open weights models but still relatively small compared to LLMs: Nemotron 3 VoiceChat (V1) has 12B parameters, making it one of the larger open weights speech to speech models, while NVIDIA's PersonaPlex is ~7B. While larger compared to other larger open weights speech to speech models the model still is relatively small compared to leading LLMs ➤ Context vs. proprietary models: While this release materially advances open weights performance, open weights speech to speech models still significantly underperform leading proprietary offerings. For comparison, proprietary models on our Big Bench Audio benchmark score substantially higher - Step-Audio R1.1 at 96%, Grok Voice Agent at 92%, Gemini 2.5 Flash (Thinking) at 92%, and Nova 2.0 Sonic at 87%. The gap between open weights and proprietary remains large in this modality. As the capability and adoption of Speech to Speech models increases, we expect to expand our set of benchmarks to include elements such as tool-calling and multi-turn instruction following. See more details below ⬇️

@kwindla ·
NVIDIA Nemotron 3 Super launches today! We've been building voice agents with Super's pre-release checkpoints and running all our various tests and benchmarks. Nemotron 3 Super matches both GPT-5.4 and GPT-4.1 in tool calling and instruction following performance on our realtime conversation, long context, real-world benchmarks. GPT-4.1 is the most widely used LLM today for production voice agents. So an open model that performs as well as GPT-4.1 on hard, voice-specific benchmarks is a big deal. (Side note: we don't think a benchmark "tells the story" about a model's voice agent performance unless it tests model correctness across at least 20 human/agent conversation turns.) The Nemotron models are *fully* open: weights, data sets, training code, inference code. Nemotron 3 Super is 120B params, with a hybrid Mamba-Transformer MoE architecture for efficient inference. You can run it on NVIDIA data center hardware or on a DGX Spark mini-desktop machine. 1M token context. Blog post with full benchmarks, thinking budget notes, inference setup on @Modal, and where we think this goes next. 👇

@livekit ·
Pronunciation is one of the fastest ways to break trust in a voice agent, especially in healthcare, legal, and finance where terminology matters. Rime's Mist v3 introduces phonetic brackets that let you define the exact pronunciation for any word and reproduce it deterministically. We built a demo nurse agent that stumbled on words like "levothyroxine" and "gastroesophageal," then fixed every one with a few lines of config. It's also fast.. as low as 100ms TTFB. Try it on LiveKit Inference today.
@aiwithjainam ·
10 GITHUB REPOS THAT GIVE YOU A STUDIO-QUALITY AI VOICE FOR FREE Bookmark every single one. Each one generates natural speech on your own machine, the thing ElevenLabs charges by the character for, running locally for $0. 1. https://t.co/xjtiT0Jnqi The open model a huge slice of the AI voice world is built on. XTTS-v2 turns text into natural speech in 17 languages and can carry your own recorded voice across all of them from a short sample. Best-in-class for multilingual narration. Note the model license is non-commercial, so it's for personal and research use. 2. https://t.co/teYZIcca5a Fast, lightweight text-to-speech that runs on almost anything, even a Raspberry Pi, fully offline. The voice engine behind countless accessibility tools and home assistants. When you need clean narration that just works on any hardware, this is the one. MIT. 3. https://t.co/gmtno0grIG One of the most natural-sounding open voice models out right now, the closest open rival to the paid services on realism. Generates speech so smooth it's hard to tell it apart from a real recording. The frontier of free, local narration. 4. https://t.co/PXW5BSICA4 A tiny, blazing-fast English voice model with preset voices you can ship in real products, since it's openly licensed for commercial use. No cloning, just clean, professional narration that runs almost instantly. The practical workhorse for apps and tools. 5. https://t.co/z8eLeeD9xI A model that doesn't just speak, it laughs, sighs, hums, and adds the little human sounds flat narration is missing. Great for expressive characters and creative audio. The one you reach for when you want personality, not just words. 35K+ stars. 6. https://t.co/oUNjf3JgSv Generates speech with fine control over tone, emotion, rhythm, and accent, in multiple languages. Built so creators can craft exactly the delivery they want. The control panel for how your synthetic voice actually performs. 7. https://t.co/lCIHlPRb4p The other half of the pipeline. It transcribes audio with precise word-level timing, so you can subtitle, edit, and re-voice content cleanly. Pair it with a voice model and you have a full audio production line on your laptop. 8. https://t.co/NYT3uQcqvs Translate and re-voice a video into 100+ languages automatically, using Whisper to transcribe, a model to translate, and a voice engine to narrate. Localize your whole channel without a studio. The repo creators use to go global overnight. 9. https://t.co/fShdI8whks Convert one singing or speaking take into a different voice you've trained, the tool the music and dubbing communities live in. Built for creators working on their own material and original characters. One of the most-starred audio repos on GitHub. 10. https://t.co/jLK5yk41vj The reading list behind all of it. A curated map of the research papers that the modern voice models grew out of, so you understand how the thing actually works instead of just running it. Where the builders in this space start. This used to need a studio and a sound engineer. Now it needs a GPU.




@ArtificialAnlys ·
Google has released Gemini 3.1 Flash Live Preview, achieving #2 in our Big Bench Audio Speech to Speech model benchmark, and now features configurable thinking levels With thinking level set to high, it scores 95.9% on Big Bench Audio, making it the second-highest scoring speech reasoning model behind Step-Audio R1.1 Realtime (97.0%) and ahead of Grok Voice Agent (92.9%). Switching to minimal thinking brings the score down to 70.5%, but opens up a faster option for latency-sensitive applications. The flexibility in thinking levels also provides a range of latency profiles. On high, average Time to First Audio (TTFA) is 2.98 seconds, slower than Step-Audio R1.1 Realtime (1.51s) and Grok Voice Agent (0.78s). On minimal, TTFA drops to 0.96 seconds, closer to the pack but still behind Google's own Gemini 2.5 Flash Native Audio Dialog (0.63s), which trades ~5 points of intelligence for the fastest response time on our leaderboard. Key takeaways: ➤ Model introduces configurable thinking levels (minimal, low, medium, high) that let developers dial reasoning depth up or down ➤ "High" thinking level: 95.9% Big Bench Audio score (2nd overall, behind only Step-Audio R1.1 Realtime ), 2.98s TTFA ➤ "Minimal" thinking level: 70.5% score, 0.96s TTFA ➤ Pricing remains stable at $0.35 per hour of audio input, and $1.38 per hour audio output, matching Gemini 2.5 Flash Native Audio Dialog See below for more details 🔽

@DivyanshT91162 ·
THIS OPEN-SOURCE VOICE AI IS GETTING SCARY GOOD. 20,000+ GitHub stars. #1 on Trending. VoxCPM2 can generate voices that are almost impossible to distinguish from real humans. → Type: "calm woman in her 30s" and it creates the voice → Clone tone, speaking style, and pacing from just a short audio sample → Studio-quality 48kHz output → Apache 2.0 licensed, including commercial use The gap between AI voices and real voices just got a lot smaller. Open source is moving insanely fast. Repo👇

@SahilPanhotra ·
Meet Mati Staniszewski → Grew up in Poland → Studied Mathematics and Computer Science → Started his career at Palantir → Worked on large scale software systems → Loved movies, games, and audiobooks → Kept noticing how terrible AI dubbing sounded → Wondered why AI voices still felt robotic → Believed AI could sound indistinguishable from humans → Teamed up with former Google ML engineer Piotr Dąbkowski → Founded ElevenLabs in 2022 → Set out to build the world's most realistic AI voice platform → Launched the first public beta in January 2023 → The internet was blown away → People started generating incredibly realistic voices within days → The product spread rapidly across creators, developers, and businesses → Then came the backlash → Users began cloning celebrity voices without permission → Deepfake clips started circulating online → The company was criticized for making voice cloning too accessible → Many believed the technology was too dangerous → Some thought ElevenLabs had launched too early → Instead of shutting the product down → Mati focused on making it safer → Added voice verification → Built AI speech detection tools → Introduced stronger moderation and security controls → Kept improving the models while earning enterprise trust → Expanded beyond text to speech → Built AI dubbing → Speech-to-text → Conversational AI → Voice agents → Developer APIs → Thousands of companies integrated ElevenLabs into their products → Reached about $100M ARR in less than 2 years → Crossed $200M ARR less than a year later → Raised a $500M Series D in 2026 → Reached an $11B valuation → Today ElevenLabs is one of the fastest growing AI companies in the world Imagine if Mati had backed down after the deepfake controversy ElevenLabs might never have become the voice behind the AI revolution

@DivyanshT91162 ·
🤯 Someone just open-sourced what feels like an ElevenLabs competitor. A few months ago, voice cloning this good would've cost you a monthly subscription. Now it's sitting on GitHub. Meet VoxCPM2. A 2B parameter voice model trained on 2 million hours of audio that can clone a voice from as little as 3 seconds of reference audio. And the scary part? Most people won't be able to tell the difference. • Clone almost any voice with 3–10 seconds of audio • Supports 30 languages, including Spanish, without language tagging • Generates 48kHz studio-quality speech • Create entirely new voices from simple text descriptions • Real-time streaming with RTF 0.3 on an RTX 4090 • Compatible with ComfyUI, vLLM, and OpenAI-style APIs • Includes LoRA fine-tuning for training custom voices • Apache 2.0 licensed for commercial use The biggest flex isn't the quality. It's that everything runs locally. No subscriptions. No cloud dependency. No sending voice samples to a third party. Just download the model and start cloning. Open source is moving way faster than most people realize. Repo👇
@omooretweets ·
One of the biggest misconceptions about voice AI is that the core benefit is cutting costs In many cases, human call center agents are cheaper than AI 👇 Most cos scale voice AI because it improves performance or helps expand their business

@shushant_l ·
I'm shocked most people still don't know how powerful Voice AI has become. Here's how Voice AI agents can listen, think, talk, and take actions automatically. --- 1. Voice AI combines speech recognition, LLM reasoning, voice generation, tools, and automation. --- 2. Speech-to-Text converts your voice into text that AI systems can understand. --- 3. Large Language Models like GPT, Gemini, and Claude power Voice AI reasoning. --- 4. Text-to-Speech tools turn AI responses into realistic human-like voices. --- 5. Tools like ElevenLabs, Deepgram, Vapi, and Retell AI help build voice agents. --- 6. Voice AI can automate customer support, sales calls, meetings, and content creation. --- 7. Modern AI voice agents can clone voices, remember context, and have real-time conversations. --- 8. Developers can combine STT, LLMs, TTS, APIs, and workflows to create custom agents. --- 9. Voice AI is evolving towards autonomous agents, multilingual support, and emotional intelligence. --- 10. The future of AI interaction is moving from typing prompts to natural voice conversations. --- To learn more, check the infographic.

@DataChaz ·
THE FOUNDERS OF @ELEVENLABS HAVE TO BE SWEATING RIGHT NOW For years, they absolutely dominated TTS because no open-source options could actually sound human. That era is officially over. @FishAudio just dropped their S2.1 Pro model, stepping up as one of the rare voice AI companies offering open-weight models. Devs now get now natural-language control over emotion, pacing, and delivery, alongside serious real-time performance: The specs speak for themselves: → 2x faster than Cartesia → 1/6th the cost of ElevenLabs → 56.3 chars/s throughput, ahead of GPT-Realtime-2 → Native support for 83+ languages → Around 90ms latency for natural conversations Need to keep everything in-house? S2.1 Pro also supports on-prem deployment with zero data retention. If you're building production-grade voice agents, this is THE most interesting open-source releases to test right now 👀

@FutureStacked ·
The first part of this audio was recorded. The second part wasn’t. Listen closely. Same voice. Same tone. Same pacing. But the second half was generated from a short voice sample — not recorded. This is voice cloning with Lightning V3.1 by @smallest_ai. With just 5–10 seconds of audio, it can recreate a voice and generate speech that still holds the natural rhythm, tone, and emotional delivery. It runs at sub-100ms latency (real-time factor of 0.01), so it’s not just realistic — it responds almost instantly. The voice quality holds up too, with one of the highest MOS scores, so it doesn’t fall apart mid-sentence. We’ve seen voice cloning in other tools before, but it often comes with trade-offs — longer sample requirements, slower generation, or a drop in quality. This is one of the few cases where speed, quality, and responsiveness all hold up at the same time. If you’re building anything involving voice — assistants, content, or automation — this opens up a lot of possibilities across different languages and use cases. Try it here: https://t.co/qVbbUqmNYD
@agenticgirl ·
Microsoft just dropped VibeVoice An open source voice AI stack that’s quietly pushing the limits of what speech models can do. Here’s why this is actually a big deal: → It can process 60 minutes of audio in a single pass No chunking. No broken context. No stitching errors. → It doesn’t just transcribe it understands structure You get: • Who spoke • When they spoke • What they said → Built-in speaker tracking + timestamps Basically ASR + diarization + formatting all in one model. → Supports 50+ languages natively Not an afterthought. Designed multilingual from the start. → You can guide it with custom hotwords Perfect for domain-heavy use cases (meetings, tech, healthcare, etc.) And that’s just the ASR side. There’s also: • Long-form TTS (up to 90 minutes, multi-speaker) • Real-time streaming voice (~300ms latency) • Lightweight deployment options (0.5B model) Under the hood, the interesting part: → Continuous speech tokenization at ultra-low frame rates → LLM + diffusion hybrid for better audio quality → Designed for long-context understanding (not just short clips) But here’s the part people shouldn’t ignore: High-quality voice generation = high risk of misuse. Deepfakes, impersonation, misinformation: all very real concerns. Even Microsoft explicitly warns against production use without safeguards. Still If you’re building in: • AI meeting assistants • Podcast tools • Voice agents • Transcription pipelines This repo is worth studying. Here's the GitHub: https://t.co/S1EpT40Ry0

@0xIngresso ·
One of the most interesting AI use cases I have seen lately started from an extremely valid frustration: overpaying for a pint of Guinness. A guy got charged too much in Dublin, realized there was no longer an official up to date tracker for Guinness prices across Ireland, and decided to do what almost nobody else would do. He built an AI voice agent, called more than 3,000 pubs, and created his own pricing map. The interesting part is that this stopped being just a funny story almost immediately. Once voice agents start collecting real world data at scale, the outcome is bigger than automation. It becomes market intelligence. It becomes price discovery. It creates competitive pressure. And that is exactly what happened, with pubs starting to lower prices to stay competitive. So yes, the story sounds funny on the surface, but it points to something very real. AI is not only entering digital workflows or internal ops. It is starting to operate inside fragmented, opaque markets full of inefficiency, exactly where manual research used to be too slow and too annoying to do properly. Today it is Guinness. Tomorrow it could be rent, insurance, hotel rates, freight, local services, any market where bad information still protects margins. If a voice agent can already reshape competition between pubs in Ireland, imagine what happens when the same model starts being used across much bigger sectors.

@TheGeorgePu ·
We used an AI voice LLM in our open-source voice tool. Checked the license. Non-commercial. Swapping to Qwen3-TTS. Apache 2.0. No restrictions. The switching cost with AI tools is approaching zero. If your license is restrictive, I just leave. This is what open source made possible.
@mark_k ·
I'm super impressed by the latest iteration of the YouTube auto-dubbing feature. I just watched a video which seemed 100% normal to me, super expressive voice, American accent. Later I noticed that the original audio was in Russian language! AI voice is getting crazy good!
@neilpatel ·
Are AI calls really saving a lot of time? Well, at least for SMBs it is. Check out the data from High Level. When they reviewed 562,000 AI voice calls through their system, they found 8,789 hours saved. Some people believe AI voice calls don't work, but the usage shows otherwise. The issue isn't the technology, it's the training you provide the AI. The more you put in, the better the result. And one of the biggest missed opportunities is how fast you get back to your leads. We've found that you are around 40% more likely to reach a prospect if you call within the first 5 minutes after a lead comes in, and although that might be hard sometimes with humans, it isn't with AI.

@TheAIColony ·
The whole TTS industry has been optimizing for how well a voice reads text, while voice agents live or die on how well a voice talks in real time. Smallest AI just launched Lightning v3.1 to solve that problem, speaking naturally when the model is still figuring out what it wants to say. We recorded a short voice sample earlier, just to see how it would hold up. Then we let it handle this: “Hey, I saw your message come in earlier, so I wanted to respond here.” What came back wasn’t just a response, it was the same voice, carried through the entire interaction, like the person never left. - No shift in tone, no break in delivery. - It didn’t sound like a system picking up halfway. - It sounded like a continuation. This is voice cloning with Lightning V3.1 by @smallest_ai. You’re not just generating speech you’re preserving how someone sounds, and letting that carry across conversations, even when they’re not there. We’ve tried a few tools in this space, but this is one of the first times it actually holds together from start to finish without losing that identity. If you’re building anything where voice matters, this changes how present someone can feel. Try it here: https://t.co/OVbyM5qnqZ
@Shruti_0810 ·
Another billion-dollar AI category is becoming open source. Voice AI is following the same path as image generation and coding assistants. Pipecat is an open-source framework for building real-time voice AI agents. It handles the entire voice stack: → Speech-to-Text → LLM reasoning → Text-to-Speech → Real-time streaming → Conversation orchestration Everything you need to build: • AI phone agents • Customer support bots • Voice copilots • AI companions • Interactive assistants • Multimodal experiences What makes it different: • Swap any AI provider without rewriting your app • Plug in OpenAI, Claude, Gemini, Groq, Ollama, and more • Choose Deepgram, Whisper, AssemblyAI, or other STT providers • Use ElevenLabs, OpenAI, Cartesia, Azure, and more for TTS • WebRTC + WebSockets built in for ultra-low latency The biggest shift isn't a new model. It's that the entire infrastructure layer is becoming a commodity. The winners won't be the companies selling voice APIs. They'll be the ones building products users actually want. GitHub: https://t.co/fkj7JO2ul7
@kwindla ·
I'm super impressed with GPT-5.4 for general use and for coding. I'm also a tiny bit disappointed (though not surprised) that it's not a standout model for voice agent use cases. - reasoning_effort = none | performs slightly worse then GPT-4o - reasoning_effort = low | performs slightly worse then GPT-5.1 ("medium" and "high" reasoning_effort are too slow for most voice agent use cases.) Every token the model generates adds to latency. And for voice agents, we have pretty hard latency caps. We need a TTFT of less than 700ms. (The actual content TTFT; the first post-thinking token!) I've had a similar conversation with several teams training models recently: I totally understand the focus on RL for reasoning. The models are getting really good at some very hard things. But ... I think we could also keep improving some capabilities in low-reasoning configurations. In particular, most new models are not very good at tool calling with reasoning turned off or set very low. My intuition from doing just enough ML work to be over-confident about my knowledge is that we should be able to have our cake and eat it too, and that this is just a data sets and engineering focus issue. Today's model's could and should be better at low-thinking budget tool calling than last year's models, while still having all the higher thinking budget gains that are so impressive.

@trikcode ·
this guy built a voice agent from their terminal and it took only 10 minutes from blank file to a deployed agent that handles a real call. It can run evals on it, find where it broke, and ship the fix with the same loop you’d use for any other service.
@alex_prompter ·
Fish Audio just launched S2.1 Pro. But the release I was more interested in was S2, its newly open-sourced speech model trained on 10 million hours of audio across 50+ languages. I went through the research paper to see what those 10 million hours of training actually produced. S2 is a text-to-speech model that lets you control how an AI voice speaks using plain-text instructions rather than preset emotions. You can add [laugh], [whisper nervously], or [professional broadcast tone] anywhere in your script, and the model follows it. The benchmarks are strong. On the Audio Turing Test, S2 scored 0.515, outperforming Seed-TTS by 24% and MiniMax-Speech by 33%. On EmergentTTS-Eval, it achieved an 81.88% win rate against GPT-4o-mini-tts. Its strongest result was emotional expression and non-verbal cues, where it won 91.61% of comparisons. In other words, when people compared the voices side by side, S2 was picked as more natural more often than Seed-TTS, MiniMax-Speech, and even GPT-4o-mini-tts. It also recorded the lowest word error rate of any model tested, including closed-source systems: 0.54% in Chinese and 0.99% in English. The architecture is smart too. Instead of using one huge model for everything, Fish split the job between two models. One generates the speech. The other adds the tiny acoustic details that make it sound natural. That keeps inference fast without sacrificing quality. They also used the same scoring system to clean the training data and fine-tune the model. Most TTS systems treat those as separate steps, which can make the final model less consistent. Fish released the model weights, fine-tuning code, and a production-ready inference engine built on SGLang. On an H200 GPU, it reaches roughly 100ms time-to-first-audio, fast enough for real-time conversations. If you’re building anything with voice, this paper is worth your time. Alongside S2, Fish Audio also launched S2.1 Pro, its newest hosted production model with 83+ languages, sub-90ms time-to-first-audio, and a managed API for developers who don’t want to self-host.

@RoundtableSpace ·
VOXCPM2 IS A FREE OPEN SOURCE VOICE MODEL THAT CAN CLONE A VOICE FROM A SHORT SAMPLE AND GENERATE NEW SPEECH IN THAT SAME STYLE. The big deal is that it’s being framed as studio quality, supports 30 languages, runs locally, and doesn’t need the usual subscription or per character pricing. WHAT MAKES IT STAND OUT IS THE FULL STACK OF FEATURES. Voice design from text, voice cloning, emotion control, longer context, real time streaming, and even fine tuning your own voice model are all packed into one open source system. That’s why this matters. A lot of paid voice AI products charge heavily for this exact workflow, while this is positioning itself as a local open alternative that can do realistic voice generation without locking you into monthly costs.

@alexabelonix ·
AI voices are getting dangerously close to “wait, is this a real person?” Microsoft just unveiled MAI-Voice-2 and MAI-Voice-2-Flash, new text-to-speech models built for more natural, expressive, multilingual speech. One is focused on high-quality narration for audiobooks, podcasts, assistants, and long-form content. The Flash version is built for real-time agents and voice assistants, where latency actually matters. This is the part of AI that will feel very personal very fast. Text AI changed how we write. Voice AI will change how we talk to software. And once agents have memory, personality, and a voice that doesn’t sound like a customer support robot from 2014, people will absolutely start treating them like real companions.
@KettlebellDan ·
I had Grok Build (v4.6) create this animated video message - it used the Voice Agent Builder and generated this message using my trained clone - it wrote the script for the message itself based on my posts - it pulled my pfp and created the 3D voxels and animated them to the waveform Sounds EXACTLY like me and like things I would actually say The future is gonna get pretty wild!
@SEBI_India ·
Launch of AI-driven calling campaign by Chairman, SEBI leveraging @SarvamAI's multilingual AI-enabled technology to make investor awareness calls. An AI conversational voice agent will proactively reach out to investors from the SEBI-authorised number 1600-313-384 to spread awareness about the SEBI Check Tool and Validated UPI handles. Know more: https://t.co/siYi2G6Yw2 #SEBI #SEBICheck


@aaliya_va ·
Voice AI has come a long way, but getting it right in a real product is still tricky. That’s why the latest launch from @soniox_ai caught my attention. Soniox just announced the launch of TTS v2 - with 60+ languages, ultra-low latency and much more accurate pronunciation for things like names, numbers, emails and verification codes. That last part is huge for voice agents - even a small pronunciation mistake can completely change the meaning of what’s being said. I also like the control over emotion and delivery with audio tags. Feel like this gives developers way more control over how natural and conversational the voice sounds. At $0.70 per generated hour, TTS v2 is looking like one of the best option for teams building voice AI at scale.

@smallest_AI ·
Last weekend was our first conference appearance in SF at the AI+ Renaissance Conference as the Title sponsor. @kamath_sutra took the stage at the Voice AI panel, and we launched Hydra – our Async Thinking Multimodal LLM – live in front of the room. This is the statement we opened with: “we are not close to passing the Turing test in voice. Not even for a single speaker, in a single language, in a single use case. And that's exactly the problem we're here to solve” The gap between AI voice agents and human conversation isn't subtle. Today's agents listen, then think, then respond. Humans do something fundamentally different – they think while listening, act while listening, and respond with contextual emotion. That's not a feature gap. That's an architectural gap. And offline LLMs can't be retrofitted to close it. That's the conviction behind everything we build at https://t.co/PIY2H0Nzcz. Small, real-time models – built from the ground up for async inference, partial context, and sub-500ms multimodal response – are the path to human-level voice intelligence. Not bigger models. Faster ones. Hydra is our step in that direction: an async thinking Speech-to-Speech model that listens and reasons in parallel, with ~50ms latency. Paired with our Lightning TTS, Lightning ASR, and Electron SLM (which outperforms GPT-4.1 on realtime conversational tasks) – the full stack is finally coming together. A massive thank you to Joshua and @lynn_aisv for building @Aiplus__ into the kind of event where everyone can have meaningful conversations, and learn from those around them. And to @Sky9Capital and @Topify_AI for co-organizing the afterparty with us – 300+ signups speaks for itself. That kind of momentum doesn't happen without people who care about the ecosystem as much as the technology. We're just getting started. The question we left the room with: Attention is all you need -but attention on what?
@bayareawriter ·
Another interesting startups whose fundraise I covered this week: @MiravoiceAI a startup using AI voice agents to conduct long-form phone surveys, raised $6.3M in a seed funding round. Miravoice has developed an AI interviewer that it says can conduct phone surveys and voice interviews for “precision data collection” without human interviewers. The surveys are long-form and quantitative, with some including more than 120 questions and lasting over 40 minutes. They span open-ended responses, numerical inputs, multiple choice questions, Likert scales and matrix questions. “Imagine talking to 100,000 people and instantly capturing what they know,” said CEO and co-founder @najain. “We make that as simple as creating a Google Form.” @Unusual_VC led the financing, which included participation from Neo, @25m_official and angel investors from companies such as Ramp, PubMatic, Atlassian and Google.

@omooretweets ·
Fun to chat with @EricNewcomer, @graceisford, and @jakesaper on the state of voice AI in 2026 Over the past two years, voice agents (and the startups building them) have exploded Now, it’s time to transition from making calls -> delivering outcomes -> building platforms


@moneyfetishist ·
Oh you built an AI voice agent? For dentists? It books appointments? Oh wow Does it also handle the part where the patient doesn’t show up? Oh it doesn’t So you automated the process of creating empty calendar slots? The dentist had empty calendar slots before But now he has AUTOMATED empty calendar slots? For $297 a month? Previously his receptionist Karen created the empty slots for free because at least when Karen booked someone she’d guilt them into showing up with a passive aggressive reminder call and a tone of voice that implied cancelling would be a moral failing Your bot doesn’t do that Your bot has no guilt Your bot has no Karen energy Your bot is a polite machine that confirms appointments for people who have already decided they’re not coming You have automated the generation of no-shows and you are charging for it monthly The dentist will figure this out in 60 days when his chair is empty and his Stripe bill isn’t But by then you’ll be at a mastermind event explaining your “retention strategy” Which is finding new dentists faster than the old ones cancel Incredible business model Really disruptive stuff
@BhosalePratim ·
One shotted a fun voice agent with @pipecat_ai and @GradiumAI to figure out with sessions to attend based on my profile at @aiDotEngineer . My agent recommended @danielhanchen ( ofcourse! ) and few more. Should I ship this so that everyone coming to conf can find their way around? Wdyt @swyx ?
@RaulJuncoV ·
A friend called me to fix their voice agent. For some reason, it was getting rogue from time to time; they were not able to understand the way. We were 100% focused on the agent and the models (tried many), and still, they didn't realize what the problem was. We had to listen a full conversation but only one was enough to understand the problem. Turns out we were looking in the wrong direction the whole time. A customer called from a busy airport to change a reservation. People were talking nearby. An announcement played in the background. The customer gave the agent a new date. And of course, the speech-to-text engine mixed the customer’s voice with the people around them, and the agent just got the wrong transcript. The mistake happened before the agent even had a chance. It was a deadly mix of bad input with confident action. My recommendation was Krisp. I personally use it for all my meetings, and their latest release takes voice isolation to the next level. It runs before speech-to-text and separates the main speaker from competing voices, background chatter, and echo. With background voices specifically, the error rate dropped from 36% to 11%. There is a moral here: it is not enough if the audio sounds better to us; we need to make it easier for machines to understand. BTW, Krisp is launching VIVA 2.5. You can learn more about the latest launch here. https://t.co/0wPXFcc5QS And here is a full demo you can play with. https://t.co/CG2viBzCbv Thanks to Krisp for partnering with me in this post. It's pretty cool to talk about a tool I use daily.

@BITKRAFTVC ·
Portfolio Spotlight: Inworld AI The most natural human interface is not a keyboard. It is voice. As AI agents become the primary layer between humans and information, voice infrastructure becomes foundational. That is why we backed @inworld_ai. Realtime TTS-2 — now the top-ranked realtime TTS model on Artificial Analysis, ahead of ElevenLabs and OpenAI. Talkpal: 5M language learners. Bible Chat: 800K DAUs, 90% TTS cost reduction. Little Umbrella: 3M hours of gameplay across 20M users. When the voice feels right, everything built on top of it grows. 🎙️

@shushant_l ·
Everyone is hyping “human-like” AI voice agents. Meanwhile most of them still interrupt you mid-sentence like a broken customer support bot. ChatGPT and Gemini Live sound smooth. But real-time conversation is still broken. Because true full-duplex AI is not about sounding smart. It is about knowing when to speak… and when to shut up. That’s the exact problem the 2026 FinVolution Global Data Science Challenge is tackling: → low-latency speech event prediction → natural turn-taking → multi-dialect Chinese conversations → zero awkward pauses $43K prize pool. Direct entry to NLPCC 2026. This is the real benchmark for voice AI. Not demo videos. Join here: https://t.co/YrX3LfNvbp #AI #LLM #SpeechAI #ConversationalAI #FullDuplex #MachineLearning #NLP #ChatGPT

@ActivateSignal ·
Ravindra Yadav, Senior Director of Data Science at Meesho, sits down with Varun Mayya and Pratyush Choudhury at Mumbai Tech Week to break down how one of India's most distinctly Bharat-first companies is rebuilding e-commerce around how people actually shop. With 264 million annual transacting users who browse rather than search, Meesho's challenge isn't adding AI for its own sake, it's making shopping feel as natural as walking into a kirana store. Ravindra unpacks the insight behind Vani, Meesho's multimodal voice assistant, and the on-ground research that revealed a striking truth: nearly 80% of new e-commerce users make their first purchase on someone else's account because the interfaces were never built for them. In this conversation, they go deep on: 0:00 How Meesho gets users: organic vs influencer channels 1:08 Why Meesho users are browsing-first, not search-first 1:32 Where AI sits in the Meesho stack 2:04 The vision behind Vani: a virtual kirana you can talk to 2:42 DICE and the on-ground insights that shaped Vani 4:04 Why 80% of new users buy on someone else's account 5:00 Making offline shopping behavior internet-native 5:31 How Vani's multimodal voice and visual understanding works 6:39 Building trust in a brand-new shopping modality If you're a founder, builder or data scientist thinking about AI, voice interfaces, or building for the next few hundred million Indian shoppers, this one's for you. @aakrit @waitin4agi_ @177pc @Meesho_Official
@DanKornas ·
Building a live voice agent requires coordinating audio streaming, turn detection, interruptions, model calls, and media routing in the same session. VideoSDK AI Agents is an open-source Python framework for building production-ready real-time voice and multimodal AI agents that join VideoSDK rooms as participants. You configure a unified Pipeline with the components you need, and the framework connects the agent lifecycle, live media processing, and the appropriate cascade, realtime, or hybrid execution mode. Key features: • Real-time room participation: agents can listen, speak, and interact live in meetings while the framework manages joining, live audio processing, and clean teardown. • Unified Pipeline configuration accepts STT, LLM, TTS, VAD, turn detection, and avatar components, then selects the optimal execution mode automatically. • Cascade mode composes STT → LLM → TTS providers for custom transcription, model, and voice choices. • Realtime and hybrid modes support a single unified realtime model or combinations such as external STT with a realtime LLM and custom TTS. • Decorator-based pipeline hooks can intercept and transform STT, LLM, TTS, vision frames, and user or agent turn events without subclassing. The README describes VideoSDK AI Agents as an open-source Python framework. Link in the reply 👇

@BenjaminBadejo ·
People keep saying this or that or another thing is a voice agent killer. Wrong. Voice functionality is a non-voice adjunct. And no single app or product that uses voice can have a moat, any more than one can corner the market for…talking. Few understand this. The market is huge. There will be many players. AI is not a market to be captured. Voice AI is not a market to be captured. They are means of interaction. The possibilities and permutations are endless.
@Kaperskyguru ·
Most voice agents fall apart the moment the room gets noisy. The fix isn't better code. It's a model that hears the speaker instead of everything happening around them. AssemblyAI's new Universal-3.5 Pro Realtime model does exactly that. It locks onto the primary speaker and tunes out the background, so a noisy room doesn't turn into phantom words. Its turn detection also reads tonality, pacing, and rhythm rather than just silence, so the agent stops talking over people. The result is cleaner transcripts and a clear read on what the speaker actually said. “Get your API key free at https://t.co/0tvnhO3FBI” Link: https://t.co/3GJPKPBmxA
@animeupdates ·
'Atelier Ryza' Officially Announces an AI chat RPG developed by SpiralAI in collaboration with KOEI TECMO's Gust The developers say the game uses no AI-generated artwork or videos, with all illustrations supervised by original Ryza illustrator Toridamono SpiralAI also said Ryza's AI voice uses recordings made specifically for the game by original voice actress Yuri Noguchi, while the dialogue model was trained on authorized 'Atelier Ryza' scenarios and reference materials The company also says revenue will be shared with illustrators, voice actors, and IP holders through royalties tied to the assets used

Best AI Voice tweets
Xholic studies what works in your niche, drafts posts in your voice and schedules them for the hours your audience is online.
$0 today · Cancel anytime
Browse all tweet collectionsKeep exploring