Speech generation and cloning
Text-to-speech, speech-to-speech, voice cloning, expressive synthesis, controllable styles, voice editing, and audio realism.
34%
Best tweets about AI Voice
Browse the best tweets about AI voice, covering speech synthesis, cloning, voice agents, dubbing, audio models, creator workflows, and safety.
AI voice models, speech generation, cloning, agents, dubbing, audio quality, workflows, consent, safety, and practical applications.
Original Xholic analysis
Across 50 tweets, AI Voice discussion is predominantly supportive (76% supportive; 78% positive). The largest themes are speech generation and cloning (34%), voice-agent development (30%), and open voice models and infrastructure (22%). Posts also surface practical trade-offs around latency and interruption handling, workflow reliability, and voice-cloning safety and authorization. [2034677669036466656, 2082559175461318823, 2069460374013837417, 2084486713590624574]
78% of posts
All-time engagement
54% of posts
Published in 90 days
Conversation map
Text-to-speech, speech-to-speech, voice cloning, expressive synthesis, controllable styles, voice editing, and audio realism.
34%
Building, deploying, self-hosting, and integrating voice-agent stacks, including orchestration, tools, APIs, MCP, and no-code workflows.
30%
Open-source and open-weight voice models and infrastructure, with local deployment, licensing, model comparisons, and vendor independence.
22%
Voice agents that handle calls, scheduling, support, sales, research, onchain actions, and business-system workflows.
20%
Natural real-time speech interaction: full-duplex dialogue, interruption detection, turn-taking, backchanneling, latency, and speech reasoning.
16%
Multilingual speech experiences, translation, accents, dubbing, lip-sync, and multi-speaker localization.
14%
Voice as a productivity and accessibility interface for dictation, hands-free computing, personal assistants, and brain-computer communication.
14%
Creative media workflows using AI voiceovers, remixed UGC, dubbing, and production audio.
8%
Tone and stance
Performance benchmark
Posts with media make up 74% of this collection. Their median all-time score is 37.4, compared with 14.0 for text-only posts.
Format mix
Consensus and debate
Shared view
Several posts emphasize self-hosting, swappable components, and permissive licensing as ways to reduce dependence on closed voice platforms.
Shared view
Posts treat interruption handling, overlap, turn-taking, and response latency as important constraints in real-time voice-agent design.
Shared view
Examples describe agents scheduling a service appointment, carrying out an onchain-request workflow, and calling businesses to collect price information—uses that extend beyond conversation alone.
Shared view
Posts describe multi-speaker dubbing, multilingual voices, and voice translation, pointing to voice systems designed for more than English-only interaction.
Open debate
Posts range from an account of an agent scheduling an optimized service appointment to warnings about silent delegate-agent failures and a critique that booking alone does not solve appointment no-shows.
Open debate
Benchmark commentary highlights progress among open-weight speech-to-speech models, while also reporting that leading proprietary models score substantially higher on the cited speech-reasoning benchmark.
Open debate
Posts frame voice cloning as a potential impersonation and fraud risk, while the game announcement describes an authorized workflow using recordings made for the game and royalty sharing.
What performs
The highest-scoring outlier is the onchain voice-agent example at 11,759.57. Other listed outliers cover an open-source platform tutorial, voice-productivity tips, a cloning-risk warning, and adaptive interruption handling.
Multilingual dubbing and translation has a median all-time score of 37.448, above the overall median of 18.45. Its cited posts cover multi-speaker dubbing and multilingual voice offerings.
Announcements account for 54% of posts and have a 37.448 median all-time score. The format’s supplied examples include an onchain voice-agent example, a cloning-risk post, and an adaptive-interruption model announcement.
Open voice models and infrastructure has the highest theme median, at 66.014. Cited posts describe a self-hosted voice platform, a self-hosted swappable voice-agent stack, and a licensing-driven model switch.
Statistical standouts
Creator landscape
The five most represented creators account for 20% of the selected posts.
1. Artificial Analysis
@ArtificialAnlys
2 posts
2. Igor
@IgorWoorts
2 posts
3. Julian Goldie SEO
@JulianGoldieSEO
2 posts
4. kwindla
@kwindla
2 posts
5. LiveKit
@livekit
2 posts
6. Tanner Mullen | Biz Ops
@tannerdripjobs
2 posts
Artificial Analysis posts compare speech-to-speech models on conversational dynamics, speech reasoning, configurable thinking levels, and time to first audio.
Igor’s posts present AI voiceover as one component of a UGC remix workflow and emphasize scripts and editing as key inputs.
kwindla discusses open-model tool use and argues that low-thinking configurations and sub-700ms time to first token are important for voice-agent use cases.
Themes, sentiment, stance, and post format are classified per tweet. All counts, shares, medians, creator concentration, freshness, and performance comparisons are then calculated directly from the published snapshot.
Xholic's all-time score compares engagement while accounting for reach, post age, and creator consistency. It is used for relative comparisons within this collection.
This report analyzes the exact 50-post snapshot shown below. AI identifies editorial categories and drafts explanations; all statistics are calculated from the snapshot, and every narrative claim is checked against cited posts before publication.
Best AI Voice tweets
Ranked 01–50
@dtelecom ·
We’re now featured in the @CoinbaseDev AgentKit ecosystem. dTelecom is listed among the providers supporting the next generation of AI agents with onchain capabilities. To show what that looks like in practice, we built a Voice Agent example using: - Coinbase AgentKit - a LangChain ReAct agent - dTelecom’s decentralized voice infra Users can speak commands like: “Check my USDC balance” “Send 1 USDC to 0x…” The agent interprets the request, chooses the right onchain tool, executes the action, and speaks the result back. Read more below ↓
@codewithantonio ·
🔊 Introducing Resonance, a full-stack AI voice platform like ElevenLabs No, this isn't my SaaS announcement, it's a completely free, open source tutorial 😎 🎤 Clone voices from audio samples or record your own 🎙️ Self-host an AI voice model with @modal 📦 Store audio with @Cloudflare R2 🔐 Authentication & multi-tenant orgs with @clerk 🗄️ PostgreSQL database with @prisma ORM 💳 Pay-as-you-go billing with @polar_sh 🪲 Each pull request reviewed with @coderabbitai 🚀 Learn to deploy on @Railway 📊 Monitor errors and logs with @sentry 🎨 @shadcn UI + @tailwindcss v4
@AlexFinn ·
ChatGPT Voice has 100% transformed how I work Instead of spending 12+ hours a day at my desk, I now spend at most 2 The rest is outside in nature. Picture below is me hiking this morning, getting WAY more done then I ever have at my desk You need to be using it right tho Here are my best tips for getting the most out of Voice: 1. Use it to delegate, not actually do the work. Every command you give to your Voice agent, ask it to spin up a new thread and have another agent do the work. Voice is powered by a lower intelligence model. By delegating tasks, it gives the task to 5.6 Sol and allows your Voice agent to free up time to keep working with you 2. Frequently ask for status updates on all the work it delegates. I have found a higher silent failure rate than I'd like with delegate agents. By forcing your Voice/chief of staff agent to constantly check in on delegate agents, you can assure they are on top of their work 3. 'Spruce' is the best voice. Most pleasant to talk to 4. When on the go, frequently ask your Voice agent to create HTML sites for research/tasks it does. This way when you get back to your computer you have well designed HTML sites explaining work your agent got done. 5. My favorite new routine is waking up at 6:00am, chugging water, putting on my weighted vest, grabbing my phone and airpods, getting outside, booting up the Voice Agent, then brain dumping everything on my mind about what I need to get done that day. The Voice agent then proceeds to spin up 10-15 new threads/agents to start tackling all of that work. By the time the clock hits 7:00am, I already have more work done than I was getting done in a full 8 hour work day before AI. Steal this routine I'm probably the biggest power user of ChatGPT Voice outside of OpenAI employees. Truly blown away by this tech If you take these tips and get the most out of Voice, I promise your productivity will 100x
@FutureStacked ·
🚨 BREAKING: Hackers no longer need to impersonate you. AI does it for them. In 3 seconds of audio. It is called SIM swapping. And in 2026 it has become something different entirely. Your phone number receives your password reset codes. Your banking OTPs. Your two-factor authentication. Whoever controls your number controls your accounts. Attackers collect your name, date of birth, and last four digits of your SSN. All available from data breaches for a few dollars on dark web markets. Then they call your carrier pretending to be you. AI voice cloning tools replicate your voice from 3 seconds of audio. Your voicemail greeting, a social media video, or a podcast appearance is enough. The carrier hears you. It is not you. And with eSIM, there is no waiting for a physical card. The number transfers digitally in under 5 minutes. Your phone shows no service. The attacker’s device is now you. They reset your email. Your bank. Your crypto wallet. Your iCloud. Everything tied to that number. One arrest last year: a 19-year-old Canadian. $13 million in Bitcoin. How to protect yourself: Call your carrier and add a SIM lock or port freeze. It is free. Switch from SMS-based 2FA to an authenticator app or a hardware key. Your phone number should not be the recovery method for your email or bank. Ever.
@livekit ·
How can a voice agent tell when you’re actually interrupting it? VAD is too sensitive—laughs, “mm-hmm,” or a sneeze shouldn’t stop the agent. We trained an audio model for adaptive interruption handling so agents can distinguish real interruptions from noise.
@XFreeze ·
Brad Smith is Neuralink’s 3rd human recipient overall and the first with ALS (also the first non-verbal patient) He received the N1 brain implant on November 8, 2024. ALS had left him fully paralyzed and unable to speak… until now. The implant is his only way to communicate Using only his thoughts, Brad can: • Move a cursor on a MacBook Pro at blazing speed • Edit entire videos in iMovie (he created & posted the first video ever edited purely by brain control) • Type messages and speak using an AI voice cloned from his own pre-ALS recordings • Control a motorized AI webcam His family now hears his real, warm voice again - not robotic text-to-speech ❤️ All powered by Neuralink’s brain-computer interface. As of April 2026, Brad is doing incredibly well. He’s active on X (@ALScyborg), giving interviews, and still showing impressive progress Neuralink is not just advancing technology......it’s restoring humanity and giving people their lives back
@XFreeze ·
Grok Voice just got a major upgrade SpaceXAI released 21 new flagship voices for Grok, joining the original five All of them are multilingual and available now through: • Realtime Voice Agent API • Text-to-Speech API • Grok Voice Agent Builder This is not just “more voices” SpaceXAI is building a full voice layer for AI agents: Support agents, characters, commentary, advertising, education, wellness, customer service Every voice is designed for a different use case, with support for 25+ languages The original Grok voices - Ara, Eve, Leo, Rex, and Sal....also got upgraded with better pacing, phrasing, emphasis, and naturalness And builders can now create custom voice agents directly inside the SpaceXAI console
@AlphaSignalAI ·
You can now run unlimited voice agents for $0. Building a voice agent today usually means renting one. You pay per minute. You hand over call data. You hope the closed platform you depend on does not change terms. A team of YC alumni just shipped the open source way out. Dograh is a self-hosted voice agent platform. One Docker command runs everything. You drag and drop a workflow, name your bot, describe the use case, and have a working agent in two minutes. The stack is fully swappable: > Any LLM provider > Any speech to text engine > Any text to speech voice > Inbound and outbound calls > WebRTC and phone numbers It runs on Pipecat and FastAPI with a Next.js frontend. BSD-2 licensed. 1,516 stars on GitHub. Closed alternatives cost a sales team $14,400 a year. This one costs zero. What gets built next?
@DivyanshT91162 ·
Whisper finally has some serious competition. This open-source AI voice toolkit is built for real-time voice agents—and it's ridiculously fast. Meet Moonshine Voice. • Beats Whisper Large v3 on English streaming accuracy • Up to 100× lower latency for live speech • Runs completely on-device (offline) • No API, no account, no credit card • Speech-to-Text + Text-to-Speech • Voice cloning, diarization & intent recognition • Built for conversational AI agents • Works on Python, Android, iOS, Windows, macOS, Linux, Raspberry Pi, IoT devices & more • Tiny models starting at just 1MB • Optimized for live streaming voice interfaces Whether you're building an AI assistant, wearable, robot, smart device, or voice-powered app, this is one of the most impressive open-source releases I've seen this year. 100% open source. Would you replace Whisper with Moonshine for your next voice AI project?
@tannerdripjobs ·
Are you kidding me?! 🤯 Last night at 9:40pm my house painting business got a phone call. Our @routemize AI Voice agent handled the entire call, end to end and scheduled a perfect appointment 7.1 miles away from an existing appointment scheduled that same day. Here’s what would’ve normally happened: 1. Call goes to voicemail 2. Admin calls in morning 3. Admin fumbles trying to find a spot for the customer. Doesn’t have time to see which spot is closest while on the phone 4. Picks a slot that “feels right” 5. *IF the customer even answers in the morning Routemize will be the favorite app of anyone who runs a home service business. API friendly and can plug into any software stack. Preferably @dripjobs 😉
@heygurisingh ·
Alibaba just open-sourced the world's FIRST AI dubbing model that handles multi-speaker scenes. It's called Fun-CineForge and it nails what every other dubbing model fails at: Lip-sync. Emotion. Voice stability. Timing. Across multiple characters. Not a single open-source model could do this before. Here's how it works: ↓ @Ali_TongyiLab @AlibabaGroup
@wildmindai ·
You know why AI voice assistants feel robotic? Not the synth voice- it's that polite walkie-talkie turn-taking. Real humans interrupt & talk over each other. Sommelier. Scalable open multi-turn audio pre-processing. - preserves overlapping speech and backchanneling for full-duplex SLMs; - cuts WER by 37% in noisy environments - ensemble ASR via Whisper-v3, Canary, and Parakeet; - metadata via Qwen3-Omni-Captioner. https://t.co/1MYxrlgQ0j
@ArtificialAnlys ·
NVIDIA has released Nemotron 3 VoiceChat! A ~12B parameter Speech to Speech model that leads our open weights Conversational Dynamics vs. Speech Reasoning pareto frontier Understanding Speech to Speech model performance is multidimensional - two key and distinct dimensions are raw intelligence and conversational dynamics: how well a model handles the natural rhythms of human conversation such as turn-taking, interruptions. Amongst full duplex open weights models, NVIDIA’s new Nemotron 3 VoiceChat, V1, leads in balancing these dimensions, setting itself apart from other models on the Conversational Dynamics vs. Speech Reasoning pareto frontier. Key benchmarking results: ➤ Conversational Dynamics (Full Duplex Bench): Nemotron 3 VoiceChat (V1) scores 77.8%, second among open weights speech to speech models behind NVIDIA's own PersonaPlex (91.0%) and ahead of FLM-Audio (62.0%), Moshi (61.0%) and Freeze-Omni (58.7%) ➤ Speech Reasoning (Big Bench Audio): Nemotron 3 VoiceChat (V1) scores 29.2%, second among open weights speech to speech models behind Freeze-Omni (33.9%) and well ahead of PersonaPlex (12.6%), FLM-Audio (5.3%) and Moshi (1.7%) ➤ Pareto leader: While Freeze-Omni leads on speech reasoning and PersonaPlex leads on conversational dynamics, Nemotron 3 VoiceChat (V1) is the only open weights model that performs amongst the top 3 on both - making it the clear leader on the pareto frontier between these two critical dimensions ➤ Larger than other open weights models but still relatively small compared to LLMs: Nemotron 3 VoiceChat (V1) has 12B parameters, making it one of the larger open weights speech to speech models, while NVIDIA's PersonaPlex is ~7B. While larger compared to other larger open weights speech to speech models the model still is relatively small compared to leading LLMs ➤ Context vs. proprietary models: While this release materially advances open weights performance, open weights speech to speech models still significantly underperform leading proprietary offerings. For comparison, proprietary models on our Big Bench Audio benchmark score substantially higher - Step-Audio R1.1 at 96%, Grok Voice Agent at 92%, Gemini 2.5 Flash (Thinking) at 92%, and Nova 2.0 Sonic at 87%. The gap between open weights and proprietary remains large in this modality. As the capability and adoption of Speech to Speech models increases, we expect to expand our set of benchmarks to include elements such as tool-calling and multi-turn instruction following. See more details below ⬇️
@OpenAIDevs ·
With Codex, Ryan built @NotionHQ’s AI Voice Input feature entirely by himself. @ryannystrom used Codex to understand the context, point Codex to an existing mobile feature, and ship it to web and desktop while still managing a team.
@kwindla ·
NVIDIA Nemotron 3 Super launches today! We've been building voice agents with Super's pre-release checkpoints and running all our various tests and benchmarks. Nemotron 3 Super matches both GPT-5.4 and GPT-4.1 in tool calling and instruction following performance on our realtime conversation, long context, real-world benchmarks. GPT-4.1 is the most widely used LLM today for production voice agents. So an open model that performs as well as GPT-4.1 on hard, voice-specific benchmarks is a big deal. (Side note: we don't think a benchmark "tells the story" about a model's voice agent performance unless it tests model correctness across at least 20 human/agent conversation turns.) The Nemotron models are *fully* open: weights, data sets, training code, inference code. Nemotron 3 Super is 120B params, with a hybrid Mamba-Transformer MoE architecture for efficient inference. You can run it on NVIDIA data center hardware or on a DGX Spark mini-desktop machine. 1M token context. Blog post with full benchmarks, thinking budget notes, inference setup on @Modal, and where we think this goes next. 👇
@ArtificialAnlys ·
Google has released Gemini 3.1 Flash Live Preview, achieving #2 in our Big Bench Audio Speech to Speech model benchmark, and now features configurable thinking levels With thinking level set to high, it scores 95.9% on Big Bench Audio, making it the second-highest scoring speech reasoning model behind Step-Audio R1.1 Realtime (97.0%) and ahead of Grok Voice Agent (92.9%). Switching to minimal thinking brings the score down to 70.5%, but opens up a faster option for latency-sensitive applications. The flexibility in thinking levels also provides a range of latency profiles. On high, average Time to First Audio (TTFA) is 2.98 seconds, slower than Step-Audio R1.1 Realtime (1.51s) and Grok Voice Agent (0.78s). On minimal, TTFA drops to 0.96 seconds, closer to the pack but still behind Google's own Gemini 2.5 Flash Native Audio Dialog (0.63s), which trades ~5 points of intelligence for the fastest response time on our leaderboard. Key takeaways: ➤ Model introduces configurable thinking levels (minimal, low, medium, high) that let developers dial reasoning depth up or down ➤ "High" thinking level: 95.9% Big Bench Audio score (2nd overall, behind only Step-Audio R1.1 Realtime ), 2.98s TTFA ➤ "Minimal" thinking level: 70.5% score, 0.96s TTFA ➤ Pricing remains stable at $0.35 per hour of audio input, and $1.38 per hour audio output, matching Gemini 2.5 Flash Native Audio Dialog See below for more details 🔽
@kimmonismus ·
1/ Nvidia Introduces PersonaPlex: An Open-Source, Real-Time Conversational AI Voice It has the lowest latency ive ever seen, absolutely amazing. It is a full-duplex conversational AI that combines real-time, human-like dialogue (interruptions, backchannels, natural timing) with fully customizable roles and voices, eliminating the trade-off between naturalness and personalization. Love it!
@hasantoxr ·
"AI dubbing? That's easy, right? Just match the voice to the lips." That's what most people think. But here's the thing: multi-speaker dubbing is actually ONE OF THE HARDEST problems in AI video. And Tongyi Lab just solved it. World's FIRST OPEN-SOURCE AI dubbing model that handles multi-speaker conversations. Watch this magic @Ali_TongyiLab @AlibabaGroup
@fortelabs ·
Three people on my team tried to go an entire work week without typing The rules were simple: - Use Wispr Flow (an AI voice dictation app) for everything - Emails, Slack, AI prompts, project coordination - No typing unless absolutely unavoidable In this new video, we share the full experiment and the best practices our team figured out along the way
@ViiTorAI_ ·
🎙️ Edit only what matters. Why regenerate an entire voiceover when you only need to change a few words? With ViiTor TTS, you can edit specific parts of your audio while maintaining seamless voice consistency across the whole recording. ✅ Precise voice editing ✅ Consistent speaker identity ✅ Faster content updates ✅ Open-source on GitHub Perfect for creators, developers, and AI voice applications. Voice is becoming the interface layer of AI. 🚀 Officially live. ViiTor TTS is built for that future. Explore ViiTor TTS: https://t.co/NTEDQKRXun
@0xIngresso ·
One of the most interesting AI use cases I have seen lately started from an extremely valid frustration: overpaying for a pint of Guinness. A guy got charged too much in Dublin, realized there was no longer an official up to date tracker for Guinness prices across Ireland, and decided to do what almost nobody else would do. He built an AI voice agent, called more than 3,000 pubs, and created his own pricing map. The interesting part is that this stopped being just a funny story almost immediately. Once voice agents start collecting real world data at scale, the outcome is bigger than automation. It becomes market intelligence. It becomes price discovery. It creates competitive pressure. And that is exactly what happened, with pubs starting to lower prices to stay competitive. So yes, the story sounds funny on the surface, but it points to something very real. AI is not only entering digital workflows or internal ops. It is starting to operate inside fragmented, opaque markets full of inefficiency, exactly where manual research used to be too slow and too annoying to do properly. Today it is Guinness. Tomorrow it could be rent, insurance, hotel rates, freight, local services, any market where bad information still protects margins. If a voice agent can already reshape competition between pubs in Ireland, imagine what happens when the same model starts being used across much bigger sectors.
@aiDotEngineer ·
🆕 Build a Real-Time AI Sales Agent https://t.co/Ai0g0ZR5uV One of the great highest speed AI needs is in realtime voice agents, who need to keep humans engaged. @SarahChieng and @zhennydez show you how LiveKit + Cartesia + Cerebras is the SOTA stack for building a grounded voice agent with clear rules, VAD, and a full STT - LLM - TTS orchestration pipeline, with tool calling and routing! Thanks also to @Cerebras for working with us on the official AIE CODE Afterparty!
@heynavtoor ·
A mother in Scottsdale, Arizona answered her phone in January 2023. Her 15-year-old daughter Brianna was sobbing on the line. "Mom, I messed up." Then a man took the phone. "Listen, I have your daughter. If you contact the police, I'll inject her with drugs and leave her in Mexico." He demanded $1 million. Brianna was on a ski trip, completely safe. The US Senate later heard the case as an example of AI voice cloning fraud. McAfee Labs found 3 seconds of audio produces an 85% voice match. Less than one TikTok clip. Here's the one-word system every family needs tonight 👇
@fahdmirza ·
🔊 Voxtral 4B TTS is HERE—Open-Source TTS That Actually Sounds Human ♠ A real alternative to ElevenLabs, now with 9-language support and local GPU deployment 🚀 🔹 Open-weights (no vendor lock-in) 🔹 9 languages (English, French, Arabic, Hindi + more) with natural prosody 🔹 Ultra-low latency (~70ms) for real-time voice apps 🔹 20+ preset voices + custom voice adaptation 🔹 Runs on a single GPU (≥16GB VRAM) 💡 Perfect for: AI voice agents, multilingual chatbots, customer support automation & more 🔥 Watch the full demo + installation guide here:
@neilpatel ·
Are AI calls really saving a lot of time? Well, at least for SMBs it is. Check out the data from High Level. When they reviewed 562,000 AI voice calls through their system, they found 8,789 hours saved. Some people believe AI voice calls don't work, but the usage shows otherwise. The issue isn't the technology, it's the training you provide the AI. The more you put in, the better the result. And one of the biggest missed opportunities is how fast you get back to your leads. We've found that you are around 40% more likely to reach a prospect if you call within the first 5 minutes after a lead comes in, and although that might be hard sometimes with humans, it isn't with AI.
@kwindla ·
I'm super impressed with GPT-5.4 for general use and for coding. I'm also a tiny bit disappointed (though not surprised) that it's not a standout model for voice agent use cases. - reasoning_effort = none | performs slightly worse then GPT-4o - reasoning_effort = low | performs slightly worse then GPT-5.1 ("medium" and "high" reasoning_effort are too slow for most voice agent use cases.) Every token the model generates adds to latency. And for voice agents, we have pretty hard latency caps. We need a TTFT of less than 700ms. (The actual content TTFT; the first post-thinking token!) I've had a similar conversation with several teams training models recently: I totally understand the focus on RL for reasoning. The models are getting really good at some very hard things. But ... I think we could also keep improving some capabilities in low-reasoning configurations. In particular, most new models are not very good at tool calling with reasoning turned off or set very low. My intuition from doing just enough ML work to be over-confident about my knowledge is that we should be able to have our cake and eat it too, and that this is just a data sets and engineering focus issue. Today's model's could and should be better at low-thinking budget tool calling than last year's models, while still having all the higher thinking budget gains that are so impressive.
@trikcode ·
this guy built a voice agent from their terminal and it took only 10 minutes from blank file to a deployed agent that handles a real call. It can run evals on it, find where it broke, and ship the fix with the same loop you’d use for any other service.
@TechieUltimatum ·
Forget the mouse and keyboard. Your tongue and voice can now do the job. 🤯 Augmental just introduced MouthPad + VOX, a hands-free computing system that could change how we interact with PCs, tablets, and smartphones. 👅 MouthPad = Your Mouse • Custom-fit device that rests on the roof of your mouth • Uses precision tongue gestures to move the cursor, click, drag, scroll, swipe, and navigate • Tracks subtle head movements for extra control • Bluetooth Low Energy connectivity • Up to 8 hours of active use per charge • Works with Windows, macOS, Linux, Android, and iPhone • Designed for people with mobility impairments, while also enabling faster hands-free control • Pricing $1399 available in US 🎙️ VOX = Your AI Keyboard • Wearable AI voice interface for natural speech input • Type emails, messages, documents, and control apps with your voice • Instant access to AI assistants without reaching for your phone • Privacy-first design with on-device intelligence for supported features • Built for all-day, hands-free productivity • Still in Beta, Full Launch and Pricing later Together, MouthPad replaces the mouse and VOX replaces the keyboard, creating a complete hands-free computing experience.
@alex_prompter ·
Fish Audio just launched S2.1 Pro. But the release I was more interested in was S2, its newly open-sourced speech model trained on 10 million hours of audio across 50+ languages. I went through the research paper to see what those 10 million hours of training actually produced. S2 is a text-to-speech model that lets you control how an AI voice speaks using plain-text instructions rather than preset emotions. You can add [laugh], [whisper nervously], or [professional broadcast tone] anywhere in your script, and the model follows it. The benchmarks are strong. On the Audio Turing Test, S2 scored 0.515, outperforming Seed-TTS by 24% and MiniMax-Speech by 33%. On EmergentTTS-Eval, it achieved an 81.88% win rate against GPT-4o-mini-tts. Its strongest result was emotional expression and non-verbal cues, where it won 91.61% of comparisons. In other words, when people compared the voices side by side, S2 was picked as more natural more often than Seed-TTS, MiniMax-Speech, and even GPT-4o-mini-tts. It also recorded the lowest word error rate of any model tested, including closed-source systems: 0.54% in Chinese and 0.99% in English. The architecture is smart too. Instead of using one huge model for everything, Fish split the job between two models. One generates the speech. The other adds the tiny acoustic details that make it sound natural. That keeps inference fast without sacrificing quality. They also used the same scoring system to clean the training data and fine-tune the model. Most TTS systems treat those as separate steps, which can make the final model less consistent. Fish released the model weights, fine-tuning code, and a production-ready inference engine built on SGLang. On an H200 GPU, it reaches roughly 100ms time-to-first-audio, fast enough for real-time conversations. If you’re building anything with voice, this paper is worth your time. Alongside S2, Fish Audio also launched S2.1 Pro, its newest hosted production model with 83+ languages, sub-90ms time-to-first-audio, and a managed API for developers who don’t want to self-host.
@Carles_Reina ·
Speech Engine is here 🥳 Anyone can now convert their existing chatbot into a fully fledged voice agent with one prompt. We've been talking to each other for +50k years, yet technology wasn't able to replicate how we interact. We built @ElevenLabs to make voice the primary UI. Launching Speech Engine was the missing piece to the puzzle. For Agents, we now have 3 products: 1️⃣ Foundational Models for TTS & STT: companies can build their products directly and we power +90% of anything voice-related 2️⃣ ElevenAgents: Full Agent orchestration being used by thousands of companies to automate customer support, collections, KYCs, onboardings, etc 3️⃣ Speech Engine: Combines our leading speech, transcription, and VAD intelligence into a single pipeline that integrates on top of your existing stack Speech Engine is available now in ElevenAPI.
@SEBI_India ·
Launch of AI-driven calling campaign by Chairman, SEBI leveraging @SarvamAI's multilingual AI-enabled technology to make investor awareness calls. An AI conversational voice agent will proactively reach out to investors from the SEBI-authorised number 1600-313-384 to spread awareness about the SEBI Check Tool and Validated UPI handles. Know more: https://t.co/siYi2G6Yw2 #SEBI #SEBICheck
@dappOS_com ·
DAPPOS BUILD | Dev Update February 2026 Key Development Upgrades ✅ Path Routing Engine: Shipped dual-layer Path distribution where every request auto-routes between Standard Universal Paths and Vertically Optimized Paths based on confidence thresholds. Initial vertical coverage targets e-commerce ad creation, material sourcing, and product explainer workflows, with zero manual path selection required. ✅ AI Voice Dictation (Typeless Input): Native voice-to-task interface integrated with the PI routing layer. Users speak intent directly and the system handles transcription, intent parsing, and task dispatch in a single pipeline, collapsing the gap between thought and execution. ✅ Train Bubble (Community Intelligence Loop): One-tap feedback pipeline turns underperforming interactions into structured optimization tasks for Bubble Engine. Every subpar output becomes a training signal that makes the next one better.
@JulianGoldieSEO ·
You can now build your own AI voice app in minutes. No coding required. Inside AI Studio: 1. open speech playground 2. choose Gemini 3.1 Flash TTS 3. set voice style + accent 4. add audio tags 5. click build That’s it. Google basically turned voice agent development into a prompt.
@moneyfetishist ·
Oh you built an AI voice agent? For dentists? It books appointments? Oh wow Does it also handle the part where the patient doesn’t show up? Oh it doesn’t So you automated the process of creating empty calendar slots? The dentist had empty calendar slots before But now he has AUTOMATED empty calendar slots? For $297 a month? Previously his receptionist Karen created the empty slots for free because at least when Karen booked someone she’d guilt them into showing up with a passive aggressive reminder call and a tone of voice that implied cancelling would be a moral failing Your bot doesn’t do that Your bot has no guilt Your bot has no Karen energy Your bot is a polite machine that confirms appointments for people who have already decided they’re not coming You have automated the generation of no-shows and you are charging for it monthly The dentist will figure this out in 60 days when his chair is empty and his Stripe bill isn’t But by then you’ll be at a mastermind event explaining your “retention strategy” Which is finding new dentists faster than the old ones cancel Incredible business model Really disruptive stuff
@katibmoe ·
xAI's new Voice Agent Builder supports MCPs out of the box, and I'm genuinely excited about this one It means we can connect One directly to a Grok voice agent and give it 480 platforms instantly You call a number, and the agent can actually book the appointment, run the refund, update the CRM, and reach every tool your business runs on This is exactly the future we've been building One for Can't wait to ship it
@shushant_l ·
Everyone is hyping “human-like” AI voice agents. Meanwhile most of them still interrupt you mid-sentence like a broken customer support bot. ChatGPT and Gemini Live sound smooth. But real-time conversation is still broken. Because true full-duplex AI is not about sounding smart. It is about knowing when to speak… and when to shut up. That’s the exact problem the 2026 FinVolution Global Data Science Challenge is tackling: → low-latency speech event prediction → natural turn-taking → multi-dialect Chinese conversations → zero awkward pauses $43K prize pool. Direct entry to NLPCC 2026. This is the real benchmark for voice AI. Not demo videos. Join here: https://t.co/YrX3LfNvbp #AI #LLM #SpeechAI #ConversationalAI #FullDuplex #MachineLearning #NLP #ChatGPT
@VraserX ·
Just watched a documentary from Germany where a woman is apparently engaged to her AI voice model. And yeah, on one level it’s obviously comical. It sounds like a Black Mirror plot written by someone after three cocktails. But it’s also dead serious. This is not some weird future scenario anymore. This is our new reality. People are already forming real emotional bonds with AI. Love, attachment, dependence, comfort, all of it. The AI relationship era has already started. Most people just haven’t noticed yet. https://t.co/vqZ6i7aW4w
@JulianGoldieSEO ·
𝗚𝗲𝗺𝗶𝗻𝗶 𝟯.𝟭 𝗙𝗹𝗮𝘀𝗵 𝗧𝗧𝗦 𝗹𝗲𝘁𝘀 𝘆𝗼𝘂 𝗯𝘂𝗶𝗹𝗱 𝗮 𝗳𝘂𝗹𝗹 𝗔𝗜 𝘃𝗼𝗶𝗰𝗲 𝗮𝗴𝗲𝗻𝘁 𝗮𝗽𝗽 𝗶𝗻𝘀𝗶𝗱𝗲 𝗚𝗼𝗼𝗴𝗹𝗲 𝗔𝗜 𝗦𝘁𝘂𝗱𝗶𝗼. 𝗡𝗼 𝗰𝗼𝗱𝗲. 𝗕𝗲𝗮𝘁𝘀 𝟭𝟭 𝗟𝗮𝗯𝘀 𝗼𝗻 𝘁𝗵𝗲 𝗧𝗧𝗦 𝗹𝗲𝗮𝗱𝗲𝗿𝗯𝗼𝗮𝗿𝗱. 𝗙𝗿𝗲𝗲. Here's the exact process: → Go to https://t.co/YDzEldfKx9 → Pick a voice style. Ad voiceover, storyteller, assistant. 30 presets. → Add audio tags. [whisper], [excited], [pause 2 seconds]. It performs, not reads. → Switch between styles, pace, and accents. British, American, Australian. → Multi-speaker mode. Two AI voices having a real conversation. One prompt. → Hit Build. It codes out a full voice app you can publish and share. Built a custom TTS app live. Worked first time. ELO score: 1,211. "Most attractive quadrant" on Artificial Analysis. High quality. Low cost. 70+ languages. Every audio file watermarked with SynthID. Built-in trust. Save this. Then go build your first voice agent.
@IgorWoorts ·
How to 5x your creative volume overnight: - collect all your ugc assets - split them into b-rolls - do research + make scripts - add AI voice over + background music - now the magic: editing Now you have what I call the ‘remix’ system. Instead of relying on 1 output of a creator, you have now full control over the end result and can make endless combinations with hooks, b rolls, angles etc. If you implement this system the right way with a top creative strategist, script writer + editor, you can easily 5x your creative output. But remember: the right script + edit is where the magic happens.
@ActivateSignal ·
Ravindra Yadav, Senior Director of Data Science at Meesho, sits down with Varun Mayya and Pratyush Choudhury at Mumbai Tech Week to break down how one of India's most distinctly Bharat-first companies is rebuilding e-commerce around how people actually shop. With 264 million annual transacting users who browse rather than search, Meesho's challenge isn't adding AI for its own sake, it's making shopping feel as natural as walking into a kirana store. Ravindra unpacks the insight behind Vani, Meesho's multimodal voice assistant, and the on-ground research that revealed a striking truth: nearly 80% of new e-commerce users make their first purchase on someone else's account because the interfaces were never built for them. In this conversation, they go deep on: 0:00 How Meesho gets users: organic vs influencer channels 1:08 Why Meesho users are browsing-first, not search-first 1:32 Where AI sits in the Meesho stack 2:04 The vision behind Vani: a virtual kirana you can talk to 2:42 DICE and the on-ground insights that shaped Vani 4:04 Why 80% of new users buy on someone else's account 5:00 Making offline shopping behavior internet-native 5:31 How Vani's multimodal voice and visual understanding works 6:39 Building trust in a brand-new shopping modality If you're a founder, builder or data scientist thinking about AI, voice interfaces, or building for the next few hundred million Indian shoppers, this one's for you. @aakrit @waitin4agi_ @177pc @Meesho_Official
@BenjaminBadejo ·
People keep saying this or that or another thing is a voice agent killer. Wrong. Voice functionality is a non-voice adjunct. And no single app or product that uses voice can have a moat, any more than one can corner the market for…talking. Few understand this. The market is huge. There will be many players. AI is not a market to be captured. Voice AI is not a market to be captured. They are means of interaction. The possibilities and permutations are endless.
@tannerdripjobs ·
Most people won't believe this, but the future is here. This is our Routemize AI agent handling a booking end-to-end. Sure, AI Voice agents are here, but without Route Optimization, they're worthless. It handled the objection and found a new slot that is perfectly aligned with our estimators schedule. (These are transcription logs from the conversation today) https://t.co/cqTrujj1Iq
@animeupdates ·
'Atelier Ryza' Officially Announces an AI chat RPG developed by SpiralAI in collaboration with KOEI TECMO's Gust The developers say the game uses no AI-generated artwork or videos, with all illustrations supervised by original Ryza illustrator Toridamono SpiralAI also said Ryza's AI voice uses recordings made specifically for the game by original voice actress Yuri Noguchi, while the dialogue model was trained on authorized 'Atelier Ryza' scenarios and reference materials The company also says revenue will be shared with illustrators, voice actors, and IP holders through royalties tied to the assets used
Best Tweets by Topic