Speech generation, cloning, and voice quality
Text-to-speech models, voice cloning, expressive delivery, human likeness, prosody, multilingual speech, and model benchmarks.
34%
Best tweets about AI Voice
Browse the best tweets about AI voice, covering speech synthesis, cloning, voice agents, dubbing, audio models, creator workflows, and safety.
AI voice models, speech generation, cloning, agents, dubbing, audio quality, workflows, consent, safety, and practical applications.
Original Xholic analysis
The 50-post AI-voice set is predominantly supportive (76%), covering speech generation and cloning, conversational agents, open/local tooling, operational deployments, dubbing, and dictation. Posts also raise recurring concerns about fraud, disclosure, authorized voice use, latency, interruption handling, and the limits of automation in some customer-facing workflows.
76% of posts
All-time engagement
50% of posts
Published in 90 days
Conversation map
Text-to-speech models, voice cloning, expressive delivery, human likeness, prosody, multilingual speech, and model benchmarks.
34%
Full-duplex interaction, interruption handling, backchannels, VAD, speech reasoning, time-to-first-audio, and responsiveness.
32%
Customer service, sales, scheduling, lead response, calling campaigns, surveys, market research, and vertical enterprise adoption.
28%
Frameworks, code stacks, pipelines, tools, evaluations, deployment, and self-hosted platforms for creating real-time voice agents.
24%
Open-weight models, local/offline dictation and synthesis, permissive licensing, self-hosting, and low-cost voice stacks.
20%
AI dubbing, lip sync, multilingual re-voicing, video translation, voiceovers, UGC remixing, and content production.
16%
Voice-clone scams, SIM-swap impersonation, deepfakes, watermarking, detection, disclosure, and authorized voice use.
12%
Voice input for writing, task dispatch, mobile work, accessibility, and replacing keyboard-centric workflows.
10%
Tone and stance
Performance benchmark
Posts with media make up 80% of this collection. Their median all-time score is 39.3, compared with 16.9 for text-only posts.
Format mix
Consensus and debate
Shared view
Posts describe pauses, pacing, backchannels, interruption handling, and full-duplex turn-taking as important components of natural voice interaction, alongside speech quality itself.
Shared view
The dataset includes self-hosted voice-agent platforms, an open-source voice-platform tutorial, local dictation, and discussion of permissive model licensing.
Shared view
Examples include scheduling, outbound investor-awareness calls, long-form phone surveys, and voice-to-task dispatch.
Shared view
Posts present cloning, dubbing, and AI voiceovers as tools for multilingual versions, UGC remixes, and producing multiple concepts from existing assets.
Open debate
Some posts characterize full-duplex voice interaction as highly human-like, while others argue that latency, interruption handling, and turn-taking remain difficult production requirements.
Open debate
One operational case study reports an end-to-end automated booking, while another post argues that high-value buyers still want a human salesperson. A separate opinion questions whether booking automation alone addresses appointment no-shows.
Open debate
Creator-oriented posts promote cloning and re-voicing, while other posts raise voice-cloning fraud, watermarking/disclosure, and authorized-recording practices.
What performs
Media appeared in 40 of 50 tweets (80%). The supplied median all-time score is 39.34 for media posts, compared with 16.95 for text-only posts.
The five supplied score outliers are an open-source voice-platform tutorial (937.44), a ChatGPT Voice workflow post (827.79), a voice-cloning/SIM-swap safety post (428.31), an interruption-handling announcement (298.11), and a cloning comparison post (297.49).
Announcements account for 25 tweets (50%) and have a median all-time score of 48.63, above the supplied medians for case studies (16.95), tutorials (11.32), and opinions (2.69).
The dictation and voice-first productivity theme has the highest supplied theme median all-time score, 122.245, and contains five tweets.
Statistical standouts
Creator landscape
The five most represented creators account for 20% of the selected posts.
1. Artificial Analysis
@ArtificialAnlys
2 posts
2. Igor
@IgorWoorts
2 posts
3. kwindla
@kwindla
2 posts
4. LiveKit
@livekit
2 posts
5. Tanner Mullen | Biz Ops
@tannerdripjobs
2 posts
6. X Freeze
@XFreeze
2 posts
Artificial Analysis is represented by benchmark posts, while LiveKit is represented by posts on interruption handling and voice-agent development. These are among the creators listed with two tweets in the supplied top-voices table.
The evidence includes route-aware appointment scheduling, a personal ChatGPT Voice work routine, and local dictation and translation tooling.
The dataset contains 44 creators across 50 tweets. The supplied analytics report a top-five placement share of 20%, indicating that the posts are not concentrated solely among the five listed top voices.
Since the previous snapshot
Themes, sentiment, stance, and post format are classified per tweet. All counts, shares, medians, creator concentration, freshness, and performance comparisons are then calculated directly from the published snapshot.
Xholic's all-time score compares engagement while accounting for reach, post age, and creator consistency. It is used for relative comparisons within this collection.
This report analyzes the exact 50-post snapshot shown below. AI identifies editorial categories and drafts explanations; all statistics are calculated from the snapshot, and every narrative claim is checked against cited posts before publication.
Best AI Voice tweets
Ranked 01–50
@codewithantonio ·
🔊 Introducing Resonance, a full-stack AI voice platform like ElevenLabs No, this isn't my SaaS announcement, it's a completely free, open source tutorial 😎 🎤 Clone voices from audio samples or record your own 🎙️ Self-host an AI voice model with @modal 📦 Store audio with @Cloudflare R2 🔐 Authentication & multi-tenant orgs with @clerk 🗄️ PostgreSQL database with @prisma ORM 💳 Pay-as-you-go billing with @polar_sh 🪲 Each pull request reviewed with @coderabbitai 🚀 Learn to deploy on @Railway 📊 Monitor errors and logs with @sentry 🎨 @shadcn UI + @tailwindcss v4
@AlexFinn ·
ChatGPT Voice has 100% transformed how I work Instead of spending 12+ hours a day at my desk, I now spend at most 2 The rest is outside in nature. Picture below is me hiking this morning, getting WAY more done then I ever have at my desk You need to be using it right tho Here are my best tips for getting the most out of Voice: 1. Use it to delegate, not actually do the work. Every command you give to your Voice agent, ask it to spin up a new thread and have another agent do the work. Voice is powered by a lower intelligence model. By delegating tasks, it gives the task to 5.6 Sol and allows your Voice agent to free up time to keep working with you 2. Frequently ask for status updates on all the work it delegates. I have found a higher silent failure rate than I'd like with delegate agents. By forcing your Voice/chief of staff agent to constantly check in on delegate agents, you can assure they are on top of their work 3. 'Spruce' is the best voice. Most pleasant to talk to 4. When on the go, frequently ask your Voice agent to create HTML sites for research/tasks it does. This way when you get back to your computer you have well designed HTML sites explaining work your agent got done. 5. My favorite new routine is waking up at 6:00am, chugging water, putting on my weighted vest, grabbing my phone and airpods, getting outside, booting up the Voice Agent, then brain dumping everything on my mind about what I need to get done that day. The Voice agent then proceeds to spin up 10-15 new threads/agents to start tackling all of that work. By the time the clock hits 7:00am, I already have more work done than I was getting done in a full 8 hour work day before AI. Steal this routine I'm probably the biggest power user of ChatGPT Voice outside of OpenAI employees. Truly blown away by this tech If you take these tips and get the most out of Voice, I promise your productivity will 100x
@FutureStacked ·
🚨 BREAKING: Hackers no longer need to impersonate you. AI does it for them. In 3 seconds of audio. It is called SIM swapping. And in 2026 it has become something different entirely. Your phone number receives your password reset codes. Your banking OTPs. Your two-factor authentication. Whoever controls your number controls your accounts. Attackers collect your name, date of birth, and last four digits of your SSN. All available from data breaches for a few dollars on dark web markets. Then they call your carrier pretending to be you. AI voice cloning tools replicate your voice from 3 seconds of audio. Your voicemail greeting, a social media video, or a podcast appearance is enough. The carrier hears you. It is not you. And with eSIM, there is no waiting for a physical card. The number transfers digitally in under 5 minutes. Your phone shows no service. The attacker’s device is now you. They reset your email. Your bank. Your crypto wallet. Your iCloud. Everything tied to that number. One arrest last year: a 19-year-old Canadian. $13 million in Bitcoin. How to protect yourself: Call your carrier and add a SIM lock or port freeze. It is free. Switch from SMS-based 2FA to an authenticator app or a hardware key. Your phone number should not be the recovery method for your email or bank. Ever.
@livekit ·
How can a voice agent tell when you’re actually interrupting it? VAD is too sensitive—laughs, “mm-hmm,” or a sneeze shouldn’t stop the agent. We trained an audio model for adaptive interruption handling so agents can distinguish real interruptions from noise.
@heyshrutimishra ·
I cloned my voice in 10 seconds and I'm still not over it. Not a generic AI voice. Not a "slightly sounds like me" voice. My timing. My pauses. My quirks. I ran it against the most popular TTS tool out there. Wasn't expecting much of a difference. There was a huge one. The reason most AI voices feel slightly off isn't audio quality. It's that they never sound like they're thinking. Real speech has micro-pauses, irregular pacing, emphasis shifts… because a brain is working in real time. Somehow Lightning V3.1 from @smallest_ai captured that. The clone doesn't just copy my timbre. It copies the irregularity. The small imperfections that make a voice feel like a person. Watch the comparison at the end. You'll hear exactly what I mean. For creators: one piece of writing, your voice, podcast clips, reels, multilingual versions. At scale.
@svpino ·
A professional voice agent in a hundred lines of Python code. Listen to this agent speak: This is one of the best, human-like voices and intonation you'll find. Stack: LiveKit + Rime + gpt-4o-mini.
@kritarthmittal ·
Tanay Kothari lore is insane > spawn in Delhi, India > watch Iron Man, get obsessed with Jarvis > build one of the first voice assistants at age 11 > hit 2.5M users > Google shuts it down > learn 20+ programming languages by age 13 > represent India at a global hackathon (won) 2014 > Delhi Public School RK Puram mafia > youngest Microsoft Certified Developer > lead NASA space settlement team > research in cancer detection > co-found Proximity, a social travel app 2016-2020 > leave Delhi to study at Stanford > work at Stanford AI Lab > teach Deep Learning alongside Andrew Ng > join Microsoft as a Program Manager > launch Convert, a music discovery platform > product goes viral > Google shuts it down too > co-found FeatherX > AI personalization for e-commerce sites > turned down YCombinator at age 21 > got competing bids for acquiring FeatherX > get acquired by Cerebra Technologies > joins as Head of Product at Cerebra 2021 > quit Cerebra to follow the childhood dream > co-found Wispr, a wearable brain-computer interface that turns thoughts into text > raise $12m, scale to 40 employees 2023 > spend years building hardware > get featured in Forbes 30 Under 30 > realize voice software sucks > quietly build a second company to fix it Early 2024 > launch Flow, a software dictation layer > personally onboard first 500 users on Gmeet > realize the software layer is the real game > kill the hardware company > go from 40 employees to 5 overnight > pivot entirely to Wispr Flow > an AI voice dictation across every app October 2024 > launch Flow publicly with 100+ languages > goes viral on X, LinkedIn, ProductHunt > rank #1 on PH for the day and the week > millions of views Early 2025 > paid conversion rate jumps to ~20% > ~90% MOM organic growth > revenue grows 50% month-over-month > raise $30M Series A led by Menlo Ventures > Pinterest founder joins the round > total funding: $56M November 2025 > raise $25M Series A-II led by Notable Capital > total raised: $81M > acquire Yapify AI > team grows back to ~33 and climbing > $700 million valuation DPS RK Puram mafia is so real — Karan Goel (Cartesia), Kunal Bahl (Snapdeal), Aman Gupta (boAt), Vineeta Singh (Sugar), and now him. Tanay Kothari for you, ladies and gentlemen. While most kids were playing games, he was getting cease & desists from Google. By 27, he had built $700M voice OS. What's your excuse?
@XFreeze ·
Grok TTS is already sounding insanely human In Vapi’s blind voting Humaneness Index, Grok TTS ranked as the top AI voice model in the chart with a humaneness score of 96.....just 4 points below the real human benchmark • Top AI voice model shown • 96/100 humaneness score • Only 4 points behind the human benchmark What makes this even more impressive is that Grok TTS is combining natural-sounding speech with low latency and aggressive pricing The gap between AI-generated speech and real human voices is disappearing faster than most people realize Grok is starting to speak like a real person
@XFreeze ·
Brad Smith is Neuralink’s 3rd human recipient overall and the first with ALS (also the first non-verbal patient) He received the N1 brain implant on November 8, 2024. ALS had left him fully paralyzed and unable to speak… until now. The implant is his only way to communicate Using only his thoughts, Brad can: • Move a cursor on a MacBook Pro at blazing speed • Edit entire videos in iMovie (he created & posted the first video ever edited purely by brain control) • Type messages and speak using an AI voice cloned from his own pre-ALS recordings • Control a motorized AI webcam His family now hears his real, warm voice again - not robotic text-to-speech ❤️ All powered by Neuralink’s brain-computer interface. As of April 2026, Brad is doing incredibly well. He’s active on X (@ALScyborg), giving interviews, and still showing impressive progress Neuralink is not just advancing technology......it’s restoring humanity and giving people their lives back
@ai_for_success ·
If you have a Mac or iPhone, try Google AI Edge Eloquent. It is a free local AI voice dictation app from Google powered by Gemma 12B, and it's seriously impressive. You can transcribe, translate, dictate, and polish text locally.
@AlphaSignalAI ·
You can now run unlimited voice agents for $0. Building a voice agent today usually means renting one. You pay per minute. You hand over call data. You hope the closed platform you depend on does not change terms. A team of YC alumni just shipped the open source way out. Dograh is a self-hosted voice agent platform. One Docker command runs everything. You drag and drop a workflow, name your bot, describe the use case, and have a working agent in two minutes. The stack is fully swappable: > Any LLM provider > Any speech to text engine > Any text to speech voice > Inbound and outbound calls > WebRTC and phone numbers It runs on Pipecat and FastAPI with a Next.js frontend. BSD-2 licensed. 1,516 stars on GitHub. Closed alternatives cost a sales team $14,400 a year. This one costs zero. What gets built next?
@tannerdripjobs ·
Are you kidding me?! 🤯 Last night at 9:40pm my house painting business got a phone call. Our @routemize AI Voice agent handled the entire call, end to end and scheduled a perfect appointment 7.1 miles away from an existing appointment scheduled that same day. Here’s what would’ve normally happened: 1. Call goes to voicemail 2. Admin calls in morning 3. Admin fumbles trying to find a spot for the customer. Doesn’t have time to see which spot is closest while on the phone 4. Picks a slot that “feels right” 5. *IF the customer even answers in the morning Routemize will be the favorite app of anyone who runs a home service business. API friendly and can plug into any software stack. Preferably @dripjobs 😉
@heygurisingh ·
Alibaba just open-sourced the world's FIRST AI dubbing model that handles multi-speaker scenes. It's called Fun-CineForge and it nails what every other dubbing model fails at: Lip-sync. Emotion. Voice stability. Timing. Across multiple characters. Not a single open-source model could do this before. Here's how it works: ↓ @Ali_TongyiLab @AlibabaGroup
@wildmindai ·
You know why AI voice assistants feel robotic? Not the synth voice- it's that polite walkie-talkie turn-taking. Real humans interrupt & talk over each other. Sommelier. Scalable open multi-turn audio pre-processing. - preserves overlapping speech and backchanneling for full-duplex SLMs; - cuts WER by 37% in noisy environments - ensemble ASR via Whisper-v3, Canary, and Parakeet; - metadata via Qwen3-Omni-Captioner. https://t.co/1MYxrlgQ0j
@ArtificialAnlys ·
NVIDIA has released Nemotron 3 VoiceChat! A ~12B parameter Speech to Speech model that leads our open weights Conversational Dynamics vs. Speech Reasoning pareto frontier Understanding Speech to Speech model performance is multidimensional - two key and distinct dimensions are raw intelligence and conversational dynamics: how well a model handles the natural rhythms of human conversation such as turn-taking, interruptions. Amongst full duplex open weights models, NVIDIA’s new Nemotron 3 VoiceChat, V1, leads in balancing these dimensions, setting itself apart from other models on the Conversational Dynamics vs. Speech Reasoning pareto frontier. Key benchmarking results: ➤ Conversational Dynamics (Full Duplex Bench): Nemotron 3 VoiceChat (V1) scores 77.8%, second among open weights speech to speech models behind NVIDIA's own PersonaPlex (91.0%) and ahead of FLM-Audio (62.0%), Moshi (61.0%) and Freeze-Omni (58.7%) ➤ Speech Reasoning (Big Bench Audio): Nemotron 3 VoiceChat (V1) scores 29.2%, second among open weights speech to speech models behind Freeze-Omni (33.9%) and well ahead of PersonaPlex (12.6%), FLM-Audio (5.3%) and Moshi (1.7%) ➤ Pareto leader: While Freeze-Omni leads on speech reasoning and PersonaPlex leads on conversational dynamics, Nemotron 3 VoiceChat (V1) is the only open weights model that performs amongst the top 3 on both - making it the clear leader on the pareto frontier between these two critical dimensions ➤ Larger than other open weights models but still relatively small compared to LLMs: Nemotron 3 VoiceChat (V1) has 12B parameters, making it one of the larger open weights speech to speech models, while NVIDIA's PersonaPlex is ~7B. While larger compared to other larger open weights speech to speech models the model still is relatively small compared to leading LLMs ➤ Context vs. proprietary models: While this release materially advances open weights performance, open weights speech to speech models still significantly underperform leading proprietary offerings. For comparison, proprietary models on our Big Bench Audio benchmark score substantially higher - Step-Audio R1.1 at 96%, Grok Voice Agent at 92%, Gemini 2.5 Flash (Thinking) at 92%, and Nova 2.0 Sonic at 87%. The gap between open weights and proprietary remains large in this modality. As the capability and adoption of Speech to Speech models increases, we expect to expand our set of benchmarks to include elements such as tool-calling and multi-turn instruction following. See more details below ⬇️
@kwindla ·
NVIDIA Nemotron 3 Super launches today! We've been building voice agents with Super's pre-release checkpoints and running all our various tests and benchmarks. Nemotron 3 Super matches both GPT-5.4 and GPT-4.1 in tool calling and instruction following performance on our realtime conversation, long context, real-world benchmarks. GPT-4.1 is the most widely used LLM today for production voice agents. So an open model that performs as well as GPT-4.1 on hard, voice-specific benchmarks is a big deal. (Side note: we don't think a benchmark "tells the story" about a model's voice agent performance unless it tests model correctness across at least 20 human/agent conversation turns.) The Nemotron models are *fully* open: weights, data sets, training code, inference code. Nemotron 3 Super is 120B params, with a hybrid Mamba-Transformer MoE architecture for efficient inference. You can run it on NVIDIA data center hardware or on a DGX Spark mini-desktop machine. 1M token context. Blog post with full benchmarks, thinking budget notes, inference setup on @Modal, and where we think this goes next. 👇
@ArtificialAnlys ·
Google has released Gemini 3.1 Flash Live Preview, achieving #2 in our Big Bench Audio Speech to Speech model benchmark, and now features configurable thinking levels With thinking level set to high, it scores 95.9% on Big Bench Audio, making it the second-highest scoring speech reasoning model behind Step-Audio R1.1 Realtime (97.0%) and ahead of Grok Voice Agent (92.9%). Switching to minimal thinking brings the score down to 70.5%, but opens up a faster option for latency-sensitive applications. The flexibility in thinking levels also provides a range of latency profiles. On high, average Time to First Audio (TTFA) is 2.98 seconds, slower than Step-Audio R1.1 Realtime (1.51s) and Grok Voice Agent (0.78s). On minimal, TTFA drops to 0.96 seconds, closer to the pack but still behind Google's own Gemini 2.5 Flash Native Audio Dialog (0.63s), which trades ~5 points of intelligence for the fastest response time on our leaderboard. Key takeaways: ➤ Model introduces configurable thinking levels (minimal, low, medium, high) that let developers dial reasoning depth up or down ➤ "High" thinking level: 95.9% Big Bench Audio score (2nd overall, behind only Step-Audio R1.1 Realtime ), 2.98s TTFA ➤ "Minimal" thinking level: 70.5% score, 0.96s TTFA ➤ Pricing remains stable at $0.35 per hour of audio input, and $1.38 per hour audio output, matching Gemini 2.5 Flash Native Audio Dialog See below for more details 🔽
@aiwithjainam ·
10 GITHUB REPOS THAT GIVE YOU A STUDIO-QUALITY AI VOICE FOR FREE Bookmark every single one. Each one generates natural speech on your own machine, the thing ElevenLabs charges by the character for, running locally for $0. 1. https://t.co/xjtiT0Jnqi The open model a huge slice of the AI voice world is built on. XTTS-v2 turns text into natural speech in 17 languages and can carry your own recorded voice across all of them from a short sample. Best-in-class for multilingual narration. Note the model license is non-commercial, so it's for personal and research use. 2. https://t.co/teYZIcca5a Fast, lightweight text-to-speech that runs on almost anything, even a Raspberry Pi, fully offline. The voice engine behind countless accessibility tools and home assistants. When you need clean narration that just works on any hardware, this is the one. MIT. 3. https://t.co/gmtno0grIG One of the most natural-sounding open voice models out right now, the closest open rival to the paid services on realism. Generates speech so smooth it's hard to tell it apart from a real recording. The frontier of free, local narration. 4. https://t.co/PXW5BSICA4 A tiny, blazing-fast English voice model with preset voices you can ship in real products, since it's openly licensed for commercial use. No cloning, just clean, professional narration that runs almost instantly. The practical workhorse for apps and tools. 5. https://t.co/z8eLeeD9xI A model that doesn't just speak, it laughs, sighs, hums, and adds the little human sounds flat narration is missing. Great for expressive characters and creative audio. The one you reach for when you want personality, not just words. 35K+ stars. 6. https://t.co/oUNjf3JgSv Generates speech with fine control over tone, emotion, rhythm, and accent, in multiple languages. Built so creators can craft exactly the delivery they want. The control panel for how your synthetic voice actually performs. 7. https://t.co/lCIHlPRb4p The other half of the pipeline. It transcribes audio with precise word-level timing, so you can subtitle, edit, and re-voice content cleanly. Pair it with a voice model and you have a full audio production line on your laptop. 8. https://t.co/NYT3uQcqvs Translate and re-voice a video into 100+ languages automatically, using Whisper to transcribe, a model to translate, and a voice engine to narrate. Localize your whole channel without a studio. The repo creators use to go global overnight. 9. https://t.co/fShdI8whks Convert one singing or speaking take into a different voice you've trained, the tool the music and dubbing communities live in. Built for creators working on their own material and original characters. One of the most-starred audio repos on GitHub. 10. https://t.co/jLK5yk41vj The reading list behind all of it. A curated map of the research papers that the modern voice models grew out of, so you understand how the thing actually works instead of just running it. Where the builders in this space start. This used to need a studio and a sound engineer. Now it needs a GPU.
@kimmonismus ·
1/ Nvidia Introduces PersonaPlex: An Open-Source, Real-Time Conversational AI Voice It has the lowest latency ive ever seen, absolutely amazing. It is a full-duplex conversational AI that combines real-time, human-like dialogue (interruptions, backchannels, natural timing) with fully customizable roles and voices, eliminating the trade-off between naturalness and personalization. Love it!
@hasantoxr ·
"AI dubbing? That's easy, right? Just match the voice to the lips." That's what most people think. But here's the thing: multi-speaker dubbing is actually ONE OF THE HARDEST problems in AI video. And Tongyi Lab just solved it. World's FIRST OPEN-SOURCE AI dubbing model that handles multi-speaker conversations. Watch this magic @Ali_TongyiLab @AlibabaGroup
@0xIngresso ·
One of the most interesting AI use cases I have seen lately started from an extremely valid frustration: overpaying for a pint of Guinness. A guy got charged too much in Dublin, realized there was no longer an official up to date tracker for Guinness prices across Ireland, and decided to do what almost nobody else would do. He built an AI voice agent, called more than 3,000 pubs, and created his own pricing map. The interesting part is that this stopped being just a funny story almost immediately. Once voice agents start collecting real world data at scale, the outcome is bigger than automation. It becomes market intelligence. It becomes price discovery. It creates competitive pressure. And that is exactly what happened, with pubs starting to lower prices to stay competitive. So yes, the story sounds funny on the surface, but it points to something very real. AI is not only entering digital workflows or internal ops. It is starting to operate inside fragmented, opaque markets full of inefficiency, exactly where manual research used to be too slow and too annoying to do properly. Today it is Guinness. Tomorrow it could be rent, insurance, hotel rates, freight, local services, any market where bad information still protects margins. If a voice agent can already reshape competition between pubs in Ireland, imagine what happens when the same model starts being used across much bigger sectors.
@aiDotEngineer ·
🆕 Build a Real-Time AI Sales Agent https://t.co/Ai0g0ZR5uV One of the great highest speed AI needs is in realtime voice agents, who need to keep humans engaged. @SarahChieng and @zhennydez show you how LiveKit + Cartesia + Cerebras is the SOTA stack for building a grounded voice agent with clear rules, VAD, and a full STT - LLM - TTS orchestration pipeline, with tool calling and routing! Thanks also to @Cerebras for working with us on the official AIE CODE Afterparty!
@TechCabal ·
In a university hostel in Nairobi, a 19-year-old founder has assembled what he describes as a full-stack artificial intelligence company, featuring a large language model (LLM) trained on Kenyan dialects, a voice agent already handling customer queries at a savings and credit cooperative (SACCO), and a prompt tool aimed at everyday users. https://t.co/AvyyAacTVU The company, Map Maven GMB, founded in 2025, claims it is worth millions, based on a formal valuation that leans on projected revenue growth in an expanding AI market.
@heynavtoor ·
A mother in Scottsdale, Arizona answered her phone in January 2023. Her 15-year-old daughter Brianna was sobbing on the line. "Mom, I messed up." Then a man took the phone. "Listen, I have your daughter. If you contact the police, I'll inject her with drugs and leave her in Mexico." He demanded $1 million. Brianna was on a ski trip, completely safe. The US Senate later heard the case as an example of AI voice cloning fraud. McAfee Labs found 3 seconds of audio produces an 85% voice match. Less than one TikTok clip. Here's the one-word system every family needs tonight 👇
@neilpatel ·
Are AI calls really saving a lot of time? Well, at least for SMBs it is. Check out the data from High Level. When they reviewed 562,000 AI voice calls through their system, they found 8,789 hours saved. Some people believe AI voice calls don't work, but the usage shows otherwise. The issue isn't the technology, it's the training you provide the AI. The more you put in, the better the result. And one of the biggest missed opportunities is how fast you get back to your leads. We've found that you are around 40% more likely to reach a prospect if you call within the first 5 minutes after a lead comes in, and although that might be hard sometimes with humans, it isn't with AI.
@kwindla ·
I'm super impressed with GPT-5.4 for general use and for coding. I'm also a tiny bit disappointed (though not surprised) that it's not a standout model for voice agent use cases. - reasoning_effort = none | performs slightly worse then GPT-4o - reasoning_effort = low | performs slightly worse then GPT-5.1 ("medium" and "high" reasoning_effort are too slow for most voice agent use cases.) Every token the model generates adds to latency. And for voice agents, we have pretty hard latency caps. We need a TTFT of less than 700ms. (The actual content TTFT; the first post-thinking token!) I've had a similar conversation with several teams training models recently: I totally understand the focus on RL for reasoning. The models are getting really good at some very hard things. But ... I think we could also keep improving some capabilities in low-reasoning configurations. In particular, most new models are not very good at tool calling with reasoning turned off or set very low. My intuition from doing just enough ML work to be over-confident about my knowledge is that we should be able to have our cake and eat it too, and that this is just a data sets and engineering focus issue. Today's model's could and should be better at low-thinking budget tool calling than last year's models, while still having all the higher thinking budget gains that are so impressive.
@VraserX ·
So what does everyone think of OpenAI’s new BiDi voice model? The bidirectional, full duplex stuff is honestly uncanny. It doesn’t feel like talking to a chatbot anymore. It literally feels like talking to a real human who can listen, interrupt, react, and keep up in real time. This is the first AI voice mode that actually feels alive.
@trikcode ·
this guy built a voice agent from their terminal and it took only 10 minutes from blank file to a deployed agent that handles a real call. It can run evals on it, find where it broke, and ship the fix with the same loop you’d use for any other service.
@alex_prompter ·
Fish Audio just launched S2.1 Pro. But the release I was more interested in was S2, its newly open-sourced speech model trained on 10 million hours of audio across 50+ languages. I went through the research paper to see what those 10 million hours of training actually produced. S2 is a text-to-speech model that lets you control how an AI voice speaks using plain-text instructions rather than preset emotions. You can add [laugh], [whisper nervously], or [professional broadcast tone] anywhere in your script, and the model follows it. The benchmarks are strong. On the Audio Turing Test, S2 scored 0.515, outperforming Seed-TTS by 24% and MiniMax-Speech by 33%. On EmergentTTS-Eval, it achieved an 81.88% win rate against GPT-4o-mini-tts. Its strongest result was emotional expression and non-verbal cues, where it won 91.61% of comparisons. In other words, when people compared the voices side by side, S2 was picked as more natural more often than Seed-TTS, MiniMax-Speech, and even GPT-4o-mini-tts. It also recorded the lowest word error rate of any model tested, including closed-source systems: 0.54% in Chinese and 0.99% in English. The architecture is smart too. Instead of using one huge model for everything, Fish split the job between two models. One generates the speech. The other adds the tiny acoustic details that make it sound natural. That keeps inference fast without sacrificing quality. They also used the same scoring system to clean the training data and fine-tune the model. Most TTS systems treat those as separate steps, which can make the final model less consistent. Fish released the model weights, fine-tuning code, and a production-ready inference engine built on SGLang. On an H200 GPU, it reaches roughly 100ms time-to-first-audio, fast enough for real-time conversations. If you’re building anything with voice, this paper is worth your time. Alongside S2, Fish Audio also launched S2.1 Pro, its newest hosted production model with 83+ languages, sub-90ms time-to-first-audio, and a managed API for developers who don’t want to self-host.
@SEBI_India ·
Launch of AI-driven calling campaign by Chairman, SEBI leveraging @SarvamAI's multilingual AI-enabled technology to make investor awareness calls. An AI conversational voice agent will proactively reach out to investors from the SEBI-authorised number 1600-313-384 to spread awareness about the SEBI Check Tool and Validated UPI handles. Know more: https://t.co/siYi2G6Yw2 #SEBI #SEBICheck
@bayareawriter ·
Another interesting startups whose fundraise I covered this week: @MiravoiceAI a startup using AI voice agents to conduct long-form phone surveys, raised $6.3M in a seed funding round. Miravoice has developed an AI interviewer that it says can conduct phone surveys and voice interviews for “precision data collection” without human interviewers. The surveys are long-form and quantitative, with some including more than 120 questions and lasting over 40 minutes. They span open-ended responses, numerical inputs, multiple choice questions, Likert scales and matrix questions. “Imagine talking to 100,000 people and instantly capturing what they know,” said CEO and co-founder @najain. “We make that as simple as creating a Google Form.” @Unusual_VC led the financing, which included participation from Neo, @25m_official and angel investors from companies such as Ramp, PubMatic, Atlassian and Google.
@dappOS_com ·
DAPPOS BUILD | Dev Update February 2026 Key Development Upgrades ✅ Path Routing Engine: Shipped dual-layer Path distribution where every request auto-routes between Standard Universal Paths and Vertically Optimized Paths based on confidence thresholds. Initial vertical coverage targets e-commerce ad creation, material sourcing, and product explainer workflows, with zero manual path selection required. ✅ AI Voice Dictation (Typeless Input): Native voice-to-task interface integrated with the PI routing layer. Users speak intent directly and the system handles transcription, intent parsing, and task dispatch in a single pipeline, collapsing the gap between thought and execution. ✅ Train Bubble (Community Intelligence Loop): One-tap feedback pipeline turns underperforming interactions into structured optimization tasks for Bubble Engine. Every subpar output becomes a training signal that makes the next one better.
we replaced a departing customer service team with a bunch of if-else statements and AI handles 80% of the tickets. here's everything from this week: the custom software era is fully here. tom requests CRM features by text, "hey jacky I want this," and it ships same day. jacky's agency fulfillment is basically automated now: article generation for clients, audits, monthly reporting, all one click from an internal dashboard hooked to kimi K3. the take: stop renting one-size-fits-all SaaS. in 2026 you should be vibe coding your own tools and making small custom improvements daily. over time you end up with software built exactly for how you work. tony's customer service lead quit and jacky's immediate reaction was "isn't this the best time to one-shot a replacement?" text-based support is literally a rules engine: bought within 30 days, issue refund. AI handles 80% of tickets minimum, flags the ones that can be upsold, humans take the rest. if someone quits in 2026 the first question should always be "does this role even need a human anymore?" where AI still loses: whales. rank surge started as pure self-checkout, but the big spenders want a human on the phone before they buy. tom tried AI voice agents and the verdict is brutal: smart people can hear the AI instantly, and smart people are the whales. so tom un-retired himself from sales and takes every call personally. the stack right now is AI for fulfillment, humans for closing. that's the split that works. the AI webinar hack from tom: don't write your webinar from scratch. find someone who crushes webinars, transcribe their entire presentation, and have AI rewrite it for your niche. jacky's funnel needed it: 120 signups, 30 showed, 5 stayed to the pitch, 1 converted. the structure is the fix, and the structure is now a copy-paste job. the @nicksaraev lunch was a masterclass in AI-era content. he grew using 1of10, a tool that surfaces viral trending youtube topics to emulate. views exploded, revenue didn't. so he switched to documentation-style videos, "I used AI in this exact way in my business," and the money followed because it attracts buyers instead of viewers. he runs a $200K/month skool community completely alone, no ads, no upsells, answers comments in the morning and records videos. then he launched a dialer SaaS into that audience just by mentioning it in videos. ARR ripped to nearly eight figures in enterprise value with zero paid spend. the next SaaS meta: partner with creators who own mind share but haven't monetized. distribution is the moat now that AI killed product moats, so launching with a creator means banger after banger, guaranteed. the group chat plan for a mutual friend building her first product: bring a laptop to dinner, whisper flow her spec while she talks, hit build when she goes to the washroom, and show her the finished product before dessert. her timeline was six months. the real timeline is one dinner. NGMI with @itstonyyu and @tomwang24 — watch/listen ↓
@CryptoSnooper_ ·
AI deepfakes just became the newest rug tool. That "proof of life" video? AI-generated. Six fingers gave it away. 👀 Now picture this in crypto: - Fake founder AMA videos - Deepfake "partnership" announcements - AI voice clones for Spaces rugs How to spot them in 10 seconds: 🔍 Count the fingers 🔍 Check ear shape/symmetry 🔍 Watch for lip sync lag on hard consonants (b, p, t) 🔍 Reverse image search the background 🔍 Run the clip through AI detectors 🔍 Look for weird background artifacts Scammers leveled up. Your snooping needs to level up too. 🕵️♂️
@moneyfetishist ·
Oh you built an AI voice agent? For dentists? It books appointments? Oh wow Does it also handle the part where the patient doesn’t show up? Oh it doesn’t So you automated the process of creating empty calendar slots? The dentist had empty calendar slots before But now he has AUTOMATED empty calendar slots? For $297 a month? Previously his receptionist Karen created the empty slots for free because at least when Karen booked someone she’d guilt them into showing up with a passive aggressive reminder call and a tone of voice that implied cancelling would be a moral failing Your bot doesn’t do that Your bot has no guilt Your bot has no Karen energy Your bot is a polite machine that confirms appointments for people who have already decided they’re not coming You have automated the generation of no-shows and you are charging for it monthly The dentist will figure this out in 60 days when his chair is empty and his Stripe bill isn’t But by then you’ll be at a mastermind event explaining your “retention strategy” Which is finding new dentists faster than the old ones cancel Incredible business model Really disruptive stuff
@EvanKirstel ·
March 26, 1876 doesn’t get nearly enough credit. Two weeks earlier, Alexander Graham Bell and Thomas Watson had already done the impossible in private. Voice traveling over a wire. That “Mr. Watson, come here” moment is the one history remembers, and fair enough. But March 26 is where the story gets more interesting. Bell started demonstrating the telephone to outside observers, scientists, investors, curious skeptics, and the reactions were all over the map. Some were genuinely stunned hearing a disembodied voice come through a wire. Others shrugged it off as a novelty. The telegraph was already a massive, reliable industry. Why would anyone need this? Here’s the detail that sticks with me: early listeners didn’t just hear sound. They heard something slightly distorted, unfamiliar, almost eerie. Think about the first time you heard a synthetic AI voice. That same uncanny sensation likely hit people in 1876. The technology worked. It just didn’t feel normal yet. The telephone had crossed from possible to demonstrable. But it hadn’t yet crossed into trusted or inevitable. Every major breakthrough shows up the same way. The first reaction isn’t adoption. It’s doubt. Telegraph operators didn’t see the need. Early telephone audio wasn’t perfect. The use case wasn’t obvious at scale. The people closest to the old technology are almost always the last to believe in the new one. We’re in that moment right now with AI voice agents. The tech works. The demos are genuinely impressive. But there’s still a wide gap between “interesting” and “essential.” March 26 is a useful reminder. Innovation doesn’t just need a breakthrough. It needs a belief moment. And that part almost always takes longer than the technology itself. #TechHistory #Innovation #Telecom #AI #DigitalTransformation
@shushant_l ·
Everyone is hyping “human-like” AI voice agents. Meanwhile most of them still interrupt you mid-sentence like a broken customer support bot. ChatGPT and Gemini Live sound smooth. But real-time conversation is still broken. Because true full-duplex AI is not about sounding smart. It is about knowing when to speak… and when to shut up. That’s the exact problem the 2026 FinVolution Global Data Science Challenge is tackling: → low-latency speech event prediction → natural turn-taking → multi-dialect Chinese conversations → zero awkward pauses $43K prize pool. Direct entry to NLPCC 2026. This is the real benchmark for voice AI. Not demo videos. Join here: https://t.co/YrX3LfNvbp #AI #LLM #SpeechAI #ConversationalAI #FullDuplex #MachineLearning #NLP #ChatGPT
@IgorWoorts ·
How to 5x your creative volume overnight: - collect all your ugc assets - split them into b-rolls - do research + make scripts - add AI voice over + background music - now the magic: editing Now you have what I call the ‘remix’ system. Instead of relying on 1 output of a creator, you have now full control over the end result and can make endless combinations with hooks, b rolls, angles etc. If you implement this system the right way with a top creative strategist, script writer + editor, you can easily 5x your creative output. But remember: the right script + edit is where the magic happens.
@ActivateSignal ·
Ravindra Yadav, Senior Director of Data Science at Meesho, sits down with Varun Mayya and Pratyush Choudhury at Mumbai Tech Week to break down how one of India's most distinctly Bharat-first companies is rebuilding e-commerce around how people actually shop. With 264 million annual transacting users who browse rather than search, Meesho's challenge isn't adding AI for its own sake, it's making shopping feel as natural as walking into a kirana store. Ravindra unpacks the insight behind Vani, Meesho's multimodal voice assistant, and the on-ground research that revealed a striking truth: nearly 80% of new e-commerce users make their first purchase on someone else's account because the interfaces were never built for them. In this conversation, they go deep on: 0:00 How Meesho gets users: organic vs influencer channels 1:08 Why Meesho users are browsing-first, not search-first 1:32 Where AI sits in the Meesho stack 2:04 The vision behind Vani: a virtual kirana you can talk to 2:42 DICE and the on-ground insights that shaped Vani 4:04 Why 80% of new users buy on someone else's account 5:00 Making offline shopping behavior internet-native 5:31 How Vani's multimodal voice and visual understanding works 6:39 Building trust in a brand-new shopping modality If you're a founder, builder or data scientist thinking about AI, voice interfaces, or building for the next few hundred million Indian shoppers, this one's for you. @aakrit @waitin4agi_ @177pc @Meesho_Official
@DanKornas ·
Building a live voice agent requires coordinating audio streaming, turn detection, interruptions, model calls, and media routing in the same session. VideoSDK AI Agents is an open-source Python framework for building production-ready real-time voice and multimodal AI agents that join VideoSDK rooms as participants. You configure a unified Pipeline with the components you need, and the framework connects the agent lifecycle, live media processing, and the appropriate cascade, realtime, or hybrid execution mode. Key features: • Real-time room participation: agents can listen, speak, and interact live in meetings while the framework manages joining, live audio processing, and clean teardown. • Unified Pipeline configuration accepts STT, LLM, TTS, VAD, turn detection, and avatar components, then selects the optimal execution mode automatically. • Cascade mode composes STT → LLM → TTS providers for custom transcription, model, and voice choices. • Realtime and hybrid modes support a single unified realtime model or combinations such as external STT with a realtime LLM and custom TTS. • Decorator-based pipeline hooks can intercept and transform STT, LLM, TTS, vision frames, and user or agent turn events without subclassing. The README describes VideoSDK AI Agents as an open-source Python framework. Link in the reply 👇
@tannerdripjobs ·
Most people won't believe this, but the future is here. This is our Routemize AI agent handling a booking end-to-end. Sure, AI Voice agents are here, but without Route Optimization, they're worthless. It handled the objection and found a new slot that is perfectly aligned with our estimators schedule. (These are transcription logs from the conversation today) https://t.co/cqTrujj1Iq
@animeupdates ·
'Atelier Ryza' Officially Announces an AI chat RPG developed by SpiralAI in collaboration with KOEI TECMO's Gust The developers say the game uses no AI-generated artwork or videos, with all illustrations supervised by original Ryza illustrator Toridamono SpiralAI also said Ryza's AI voice uses recordings made specifically for the game by original voice actress Yuri Noguchi, while the dialogue model was trained on authorized 'Atelier Ryza' scenarios and reference materials The company also says revenue will be shared with illustrators, voice actors, and IP holders through royalties tied to the assets used
Best Tweets by Topic