Voice generation & cloning
ElevenLabs text-to-speech, voice cloning, expressive delivery, and voice-restoration use cases.
34%
Best tweets about ElevenLabs
Discover the best tweets about ElevenLabs, covering AI voices, speech generation, dubbing, voice agents, APIs, releases, and creator workflows.
Specific ElevenLabs products, voice models, speech workflows, agents, API integrations, quality tests, and product updates.
Original Xholic analysis
Discussion of ElevenLabs centers on voice generation, developer and creator workflows, and voice agents. The supplied posts also contain competing-model benchmark claims and recurring questions about voice rights and safeguards. The highest-scoring supplied posts feature competitive TTS assertions, automated video workflows, and real-time integration support.
66% of posts
All-time engagement
54% of posts
Published in 90 days
Conversation map
ElevenLabs text-to-speech, voice cloning, expressive delivery, and voice-restoration use cases.
34%
Developer integrations through APIs, MCP, Claude/Codex, frameworks, SDKs, and application connectors.
28%
Competitive TTS comparisons involving open-source, on-device, and rival voice-model quality, latency, and pricing.
26%
ElevenLabs voice agents, conversational AI, real-time speech interactions, and customer-service deployments.
26%
Creator workflows using ElevenLabs for video voiceovers, podcasts, social content, audiobooks, games, and storytelling.
24%
Company growth, enterprise adoption, strategic partnerships, market expansion, and ElevenLabs business positioning.
14%
ElevenLabs product releases across dubbing, music, sound effects, creative tools, marketplaces, and Studio features.
14%
Voice rights, synthetic-audio watermarking, detection, consent, celebrity licensing, and ethical concerns.
10%
Tone and stance
Performance benchmark
Posts with media make up 84% of this collection. Their median all-time score is 18.8, compared with 11.8 for text-only posts.
Format mix
Consensus and debate
Shared view
Several posts place ElevenLabs in creator production workflows for short-form video and social content, alongside scripting, visuals, captions, editing, and—in some examples—human review.
Shared view
Developer posts describe ElevenLabs connections with TanStack AI, Claude, Claude Code/Codex, and agent-assisted sound-effect generation, framing these integrations as ways to generate audio within existing workflows.
Shared view
Examples include a pub-price calling project, wrapping chat agents for phone use, Claude Desktop agent setup, and customer-service-oriented agent positioning. These posts emphasize real-time conversational voice use cases.
Open debate
Benchmark claims differ by evaluator and test. Artificial Analysis places Eleven v3 second in its preference leaderboard, while posts about VoxCPM2, Voxtral, and Async cite results favoring those alternatives on particular cloning, preference, or normalization tests.
Open debate
One user praises ElevenLabs voice clones and audiobook potential, while other firsthand posts report weak photo lip-sync and that a generated podcast remained recognizably unlike its creator.
Open debate
Safety discussion includes a report about watermarking and detection, criticism of licensing a deceased celebrity’s voice, and a company-history post describing earlier backlash over unauthorized celebrity cloning.
What performs
The five supplied score outliers cover competitive TTS claims, automated video pipelines, and real-time framework support. Their all-time scores range from 556.69 to 3946.64, compared with the dataset median of 17.7.
Tutorial examples provide concrete steps for pipeline setup, short-form production, and agent-assisted SFX generation. The supplied tutorial-format median all-time score is 19.97.
Media appears in 42 of 50 posts (84%). The supplied median all-time score is 18.75 for media posts and 11.82 for text-only posts.
Statistical standouts
Creator landscape
The five most represented creators account for 18% of the selected posts.
1. Carles Reina
@Carles_Reina
2 posts
2. ElevenLabs Developers
@ElevenLabsDevs
2 posts
3. Alex Banks
@thealexbanks
2 posts
4. Vaibhav Sisinty
@VaibhavSisinty
2 posts
5. Mayank Vora
@aiwithmayank
1 post
6. Alex the Engineer
@AlexEngineerAI
1 post
Carles Reina’s updates list Eleven Agents, Expressive Agents, Dubbing V2, Music V2, WhatsApp multimodal handling, Studio Agent, Templates, and expansion into additional markets.
ElevenLabs Developers reports that 11k people use its MCP server in Claude and describes its examples repository as prompt-driven, AI-native, and autonomously maintained.
Two high-scoring workflow posts describe open-source video pipelines that use ElevenLabs for voiceover. Both quote an estimated ElevenLabs voice cost of about $0.05 per video.
Themes, sentiment, stance, and post format are classified per tweet. All counts, shares, medians, creator concentration, freshness, and performance comparisons are then calculated directly from the published snapshot.
Xholic's all-time score compares engagement while accounting for reach, post age, and creator consistency. It is used for relative comparisons within this collection.
This report analyzes the exact 50-post snapshot shown below. AI identifies editorial categories and drafts explanations; all statistics are calculated from the snapshot, and every narrative claim is checked against cited posts before publication.
Best ElevenLabs tweets
Ranked 01–50
@heynavtoor ·
🚨 ElevenLabs charges $5 to $99/month for AI voice cloning. Their Business plan costs $1,320/month. Someone open sourced a voice AI that clones any voice from a short clip. 30 languages. Studio quality. Free. It's called VoxCPM2. Give it a short clip of anyone's voice. It clones their accent, emotion, tone, and pacing. Then generates any speech you want in their exact voice. 48kHz studio quality. Type "A young woman, gentle and sweet voice" and it creates that voice from scratch. No reference audio. No voice actor. No recording. You describe a voice in words. It builds it. 2 billion parameters. Trained on 2 million hours of speech. 30 languages. One command to install: pip install voxcpm Here's what VoxCPM2 does: → Voice Design: describe any voice in words. Gender, age, tone, emotion, pace. AI creates it from nothing. No reference audio needed. → Voice Cloning: upload a short audio clip. AI clones the voice perfectly. Timbre, accent, rhythm, pacing. → Controllable Cloning: clone a voice AND control the emotion. "Slightly faster, cheerful tone." Done. → Ultimate Cloning: provide audio + transcript. Every vocal nuance faithfully reproduced. → 30 languages. Arabic, Chinese, English, French, German, Hindi, Japanese, Korean, Spanish, and 21 more. No language tags needed. → Context-aware. It reads the text and adjusts emotion and rhythm automatically. News sounds like news. Stories sound like stories. → Real-time streaming. RTF as low as 0.13 on an RTX 4090. Faster than playback speed. → Runs on 8GB of VRAM. → Fine-tune with 5 to 10 minutes of your own audio using LoRA. Build a custom voice model. → 48kHz output. Studio quality. No external upsampler needed. Here's the wildest part: On the Minimax-MLS voice similarity benchmark: → English: VoxCPM2 scores 85.4%. ElevenLabs scores 61.3%. → Chinese: VoxCPM2 scores 82.5%. ElevenLabs scores 67.7%. → Arabic: VoxCPM2 scores 79.1%. ElevenLabs scores 70.6%. A free, open source model is producing more realistic voice clones than a service that charges up to $1,320/month. Professional voice actors charge $250 to $1,000+ per project. AI voice platforms charge $5 to $100/month. Recording studios charge $200/hour. This runs on your GPU. Locally. No API costs. No per-character pricing. No subscription. Free forever. Already hit #1 on GitHub Trending. Built by OpenBMB and Tsinghua University. 2 billion parameters. Apache 2.0 License. Free for commercial use. 100% Open Source.
@aiwithmayank ·
Say goodbye to video editors. This open-source tool turns a news headline into a published YouTube Short in one command. It's called YouTube Shorts Pipeline and it chains Claude, Gemini Imagen, ElevenLabs, and Whisper together into one pipeline. The cost breakdown is brutal: Claude script: $0.02 Gemini visuals: $0.03 ElevenLabs voice: $0.05 Total: $0.10 per video Type a topic. Get a live YouTube link. 3–5 minutes. Supports multiple languages. Custom voice IDs. Dry-run mode to preview before producing. Manual script override before rendering. Everything stored locally. No cloud dependency. 100% Open Source. MIT License.
@sentient_agency ·
Say goodbye to faceless YouTube channel agencies. This open-source repo does everything they charge $500/month for. news topic → script → AI visuals → voiceover → captions → uploaded to YouTube One command. 5 minutes. $0.10. What it uses: → Claude Sonnet for script writing (~$0.02/video) → Gemini Imagen 3 for 3 b-roll visuals (~$0.03/video) → ElevenLabs for voiceover (~$0.05/video) → Whisper for auto-captions (free) → YouTube Data API for upload (free) Videos upload as private by default so you review before going live. The anti-hallucination gate is the best part. Claude is shown live search snippets and instructed to only use names, stats, and facts it finds there. No AI making stuff up in your videos. 100% Open Source. MIT License.
@kimmonismus ·
Mistral AI released Voxtral TTS, a 3-billion-parameter text-to-speech model with open weights that the company says outperformed ElevenLabs Flash v2.5 in human preference tests roughly 63% of the time on standard voices and nearly 70% on voice customization. The model runs on about 3 GB of RAM, achieves 90-millisecond time-to-first-audio, supports nine languages, and can clone a voice from just five seconds of reference audio, including cross-lingual adaptation that preserves the speaker's accent.
@emilylambert ·
omg...Opus 4.6 dropped and it's so good at creating mobile apps. I built this app in an hour (A CapWords competitor) APIs added that are directly in vibecode: > ElevenLabs for TTS > GPT 4o mini for analyzing/id'ing the photos, > Replicate for background removal for background removal In the thread you'll see how to get free access to @v_computer
@VaibhavSisinty ·
Did xAI just mass-murder the entire voice AI industry? 🤯 Grok just launched two voice APIs. Speech-to-Text and Text-to-Speech. Built on the same stack powering Tesla cars and Starlink support. And priced at 10x cheaper than ElevenLabs. Speech-to-Text: $0.10/hr batch. $0.20/hr streaming. Text-to-Speech: $4.20 per million characters. 25+ languages. Real-time streaming. Speaker diarization. Already outperforming ElevenLabs, Deepgram, and AssemblyAI on word error rate. TTS ships with expressive tags like [laugh], [sigh], <whisper>, <emphasis>. Voices that don't sound like robots reading a script. ElevenLabs spent years building a voice AI company. xAI built voice AI for cars and satellites.
@ElevenLabsDevs ·
11k people are using the ElevenLabs MCP server in Claude. It connects to your ElevenLabs account and lets you generate speech, sound effects, and music directly in Claude. Generate audio with ElevenLabs without friction.
@VaibhavSisinty ·
xAI just launched a voice agent that costs $0.05 per minute. Everything included. And you can deploy it in 2 minutes. No STT vendor. No LLM vendor. No TTS vendor. One model. One price. ElevenLabs charges $50 to $120 per million characters for TTS alone. Grok charges $4.20. The voices laugh. Sigh. Whisper. Handle noise, accents, and interruptions naturally. Because it was not trained on studio audio. It was trained on Tesla cars and Starlink support calls. On τ-voice Bench it scores 67%. Gemini scores 44%. GPT Realtime scores 35%. Upload your docs. Connect your APIs. Clone a voice from 2 minutes of audio. Bring your own number. Deploy in production in the same session. ElevenLabs spent years building a voice AI company. xAI built voice AI for cars and satellites and just gave it to everyone. The fragmented voice stack is dead. @xAI @grok
@ElevenCreative ·
Introducing Music Finetunes in ElevenCreative. Generate individual vocals, instruments, or full tracks with stylistic consistency using a fine-tuned version of our Music model. Create your own Finetune or get started instantly with 11 Finetunes curated by ElevenLabs.
@bearlyai ·
London-based engineer Matt Cortland created an AI voice agent that called over 3,000 pubs on St. Patrick’s Day weekend to ask how much a pint of Guiness cost. The Irish government used to track this price up until 2011. Cortland stepped up with his “Guinndex”, which cost €200 to make and showed that the national price of €5.95. He used ElevenLabs to created “Rachel” and “getting the voice proved critical” (a Northen Irish accent inspired by a TV show).
@heyrimsha ·
ElevenLabs just got dethroned. Supertonic runs text-to-speech entirely on your device. No cloud. No API key. No latency. 100% open source. The numbers are solidly wild: → 167x faster than real-time on an M4 Pro → Only 66M parameters (smaller than most phone apps) → 1,263 chars/sec vs ElevenLabs Flash at 287 → 1,048 chars/sec vs OpenAI TTS-1 at 55 → Runs on a Raspberry Pi. Runs on an e-reader in airplane mode. Reads currency, dates, phone numbers, technical units correctly without preprocessing. ElevenLabs, OpenAI, Gemini, and Microsoft all fail these. Supertonic doesn't. Supports 11 platforms: Python, Node, Browser, Java, C++, C#, Go, Swift, iOS, Rust, Flutter. Multilingual in English, Korean, Spanish, Portuguese, French. The Chrome extension turns any webpage into audio in under a second. Free. Private. On-device. I've watched on-device models lose to cloud APIs for years. This one doesn't lose. It wins on speed and quality. The cloud TTS business just got cooked.
@Carles_Reina ·
We've had the best quarter ever at @ElevenLabs ▪️Smashed even the wildest Revenue targets across B2B and Prosumer/B2C ▪️Enterprise overtook Prosumer/B2C accounting for 51% of all revenues now ▪️Eleven Agents is our fastest growing product ▪️Almost 50% first-sale is outbound ▪️Europe leading in outbound SQOs ▪️India leading in quota attainment despite our 20x quota/base requirement ▪️We hired 150 people out of 79,098 applicants = 0.0018% ▪️Added multiple new markets that will be announced soon ▪️Evolved our corporate structure to test a fully decentralised model giving countries more autonomy -> now expanding it further ▪️Launched Expressive Agents mode ▪️Ran the biggest ElevenLabs Summit in London on 11/Feb ▪️Our F1 team, Audi Revolut F1 Team, got the first points of the season What's in store for Q2? More experiments, more mistakes, deeper customer interactions and new products & models.
@shengkun_ye ·
You can now use ElevenLabs directly in Claude Code or Codex. Voiceovers for your demo videos, narration for tutorials, character voices for your game. Just one skill. No subscription required.
@XFreeze ·
SpaceXAI just launched Grok Voice Think Fast 2.0 - its most capable speech-to-speech model yet The numbers are so impressive: • 82.9% overall Speech-to-Speech Quality Index • 97.2% on speech reasoning • 56.5% on agentic voice tasks • Just 0.70 seconds to first audio • Uses only 0.4× the reasoning tokens of Think Fast 1.0 Grok Voice Think Fast 2.0 now ranks ahead of GPT-Realtime-2.1 and Gemini 3.1 Flash on overall speech-to-speech benchmarks But the biggest improvement is real-world listening Across thousands of short phrases in 24 languages, it delivers 1.5–2× better transcription accuracy than Deepgram Nova 3 and ElevenLabs Scribe v2 In noisy environments and compressed phone calls, that advantage grows to around 10× It can also reason while speaking, allowing tool calls to begin before it even finishes its first sentence.....with no added latency Conversations now feel much more natural: • Shorter responses • One question at a time • Less filler • More fluid dialogue • Better guidance through complex tasks SpaceXAI has already deployed it on Starlink's phone line and saw significant increases in both sales conversion and customer support containment Grok Voice Think Fast 2.0 becomes the default grok-voice-latest model on August 5 Pricing: $0.08 per minute of audio Voice AI is rapidly evolving from scripted assistants into intelligent agents that can listen, reason, speak, and take action in real time
@TheTuringPost ·
.@MistralAI's new Voxtral TTS generates expressive, multilingual speech from just ~3 seconds of reference audio It solves one of the hardest problems in speech, separating what you say from how you sound ➡️ Voxtral factorizes speech into two parts: • semantic tokens → the content (words) • acoustic tokens → the voice (tone, prosody, style) and uses different models for different problems + Voxtral Codec that compresses speech into ultra-low bitrate tokens This gives: - High-quality voice cloning that works best with ~3–25 second - support for 9 languages - 68.4% win rate vs ElevenLabs Flash v2.5 in voice cloning Here's how it works:
@SahilPanhotra ·
Meet Mati Staniszewski → Grew up in Poland → Studied Mathematics and Computer Science → Started his career at Palantir → Worked on large scale software systems → Loved movies, games, and audiobooks → Kept noticing how terrible AI dubbing sounded → Wondered why AI voices still felt robotic → Believed AI could sound indistinguishable from humans → Teamed up with former Google ML engineer Piotr Dąbkowski → Founded ElevenLabs in 2022 → Set out to build the world's most realistic AI voice platform → Launched the first public beta in January 2023 → The internet was blown away → People started generating incredibly realistic voices within days → The product spread rapidly across creators, developers, and businesses → Then came the backlash → Users began cloning celebrity voices without permission → Deepfake clips started circulating online → The company was criticized for making voice cloning too accessible → Many believed the technology was too dangerous → Some thought ElevenLabs had launched too early → Instead of shutting the product down → Mati focused on making it safer → Added voice verification → Built AI speech detection tools → Introduced stronger moderation and security controls → Kept improving the models while earning enterprise trust → Expanded beyond text to speech → Built AI dubbing → Speech-to-text → Conversational AI → Voice agents → Developer APIs → Thousands of companies integrated ElevenLabs into their products → Reached about $100M ARR in less than 2 years → Crossed $200M ARR less than a year later → Raised a $500M Series D in 2026 → Reached an $11B valuation → Today ElevenLabs is one of the fastest growing AI companies in the world Imagine if Mati had backed down after the deepfake controversy ElevenLabs might never have become the voice behind the AI revolution
@rahulbhadoriiya ·
So I've been working on a new Instagram account for the last few months, and the results have been solid 200K+ views collectively on all the reels, 2–5K profile views, and a bunch of other wins here and there. The fun part? Most of it was AI-generated content. The delivery and execution were AI, but the actual idea, what to post, what to talk about was still human. The process was simple: 1. I'd find something interesting I want to talk about. 2. Send it to my OpenAI agent. 3. It writes a script : I've trained it by studying a lot of different AI content creators. 4. I edit the script if something feels off. 5. I've been doing media for six years now, content creation, writing ads, all of that so I understand how it works. 6. Then I generate audio through ElevenLabs. I've cloned my voice, so it sounds exactly like me. If some words don't land right, I edit those out. 7. Then I generate an avatar through HeyGen. So till here, everything is AI, script, audio, video. The editing is human, but AI still helps there too. I've built an agent that uses the X API to pull B-rolls and relevant clips from Twitter, then forwards everything to my editor via Telegram. He sees the new video drop in the group, edits it out, and that's the video done. That's one format. There's another format where I've taught my design AI agent how to create carousels, this is for content I want to post immediately, when I don't want to wait half a day for an edit. That works really well for quick takes. I've also been experimenting with a couple of new formats, and I'll keep doing that over the next year, automating as much as possible, but keeping the soul intact. The first idea, what I actually want to say and share, stays human. The delivery, the strip, the production, AI handles the manual work.
@Harish_521 ·
I have posted 18 shorts in 10 days on youtube. experiment results in the image This is my full process Ask ChatGPT to write a 60 second script. Generate a single image with Image Gen. Split the image into 8 scenes. Generate voiceover in ElevenLabs. Drop everything into CapCut. Sync images with audio. Add blur + fade transitions. Export. One idea. 10 minutes. A complete YouTube Short. Now going to automate it using n8n
@DataChaz ·
THE FOUNDERS OF @ELEVENLABS HAVE TO BE SWEATING RIGHT NOW For years, they absolutely dominated TTS because no open-source options could actually sound human. That era is officially over. @FishAudio just dropped their S2.1 Pro model, stepping up as one of the rare voice AI companies offering open-weight models. Devs now get now natural-language control over emotion, pacing, and delivery, alongside serious real-time performance: The specs speak for themselves: → 2x faster than Cartesia → 1/6th the cost of ElevenLabs → 56.3 chars/s throughput, ahead of GPT-Realtime-2 → Native support for 83+ languages → Around 90ms latency for natural conversations Need to keep everything in-house? S2.1 Pro also supports on-prem deployment with zero data retention. If you're building production-grade voice agents, this is THE most interesting open-source releases to test right now 👀
@HedgieMarkets ·
🦔ElevenLabs, valued at $11 billion, partnered with Stan Lee Universe to release an AI clone of the late Marvel co-creator's voice and likeness. Lee's cloned voice will narrate public domain books in the Eleven Reader app starting with Treasure Island in June, and goes up for commercial licensing through the Iconic Marketplace. Stan Lee Universe is a joint venture involving POW! Entertainment, the same company Lee sued for $1 billion in 2018 alleging they took advantage of his macular degeneration and grief after his wife's death to broker an unauthorized likeness deal. Lee died later that year at 95. My Take I cannot read this announcement without feeling sick about it. Stan Lee spent the final months of his life filing lawsuits against business partners who exploited his blindness, his grief, and in one case literally drew his blood to stamp into comic books sold as collectibles in Las Vegas. Now the same company he sued is the legal counterparty greenlighting an AI clone of his voice. I do not see how anyone at ElevenLabs or Stan Lee Universe looked at that paperwork and thought this was a clean deal. The Iconic Marketplace business needs dead celebrities to scale. Living celebrities can refuse a campaign, fire their agent, or sue over a likeness deal they did not authorize. Dead ones cannot, and their estates have every reason to keep the licensing income coming. Judy Garland and John Wayne already sit on the same roster as Lee, and ElevenLabs at $11 billion needs recurring revenue to justify the valuation. I think the technology keeps improving and the catalog keeps growing regardless of how people react to any single announcement. The only thing slowing this market is public disgust, and based on the response to this one, I am not sure that is enough. Hedgie🤗
@tonnoz ·
Here is how to generate all the sound effects for your game/app in 2 minutes: 1. Enable a subscription on @ElevenLabs 2. Create an API key with SFX generation privileges 3. Run: npx skills add elevenlabs/skills 4. Tell Claude/Codex to find where your app still needs SFX, generate highly detailed ElevenLabs prompts, and create 3 versions of each sound on a small review page where you can play them 5. Pick the ones you like and confirm them via the agent 6. Compress them in the best possible format to save space, especially for web apps: .webm 7. Enjoy your game/app with awesome SFX in less time than it takes to make coffee Bookmark this for later ;)
@Daviowhite ·
Working on a new project at work called “Learn” for our investors. I’m currently testing early concepts with Flora AI, I generated the avatar with Gemini and the voiceover with ElevenLabs. Then I hit a blocker in Flora, it doesn’t have an audio node. My goal is to automate this so I can create rapidly. Does anyone know how this can be achieved using a free tool? I’m trying to figure out the right toolkit for this workflow before committing to paid options.
@ArtificialAnlys ·
Inworld, ElevenLabs, and MiniMax continue to lead our Text to Speech leaderboard for most preferred models Recent checkpoints from each of the labs continue to push the frontier of TTS quality, with 4 out of the top 5 models being released this year. Leading TTS models are increasingly realistic, particularly on relatively straightforward text, with preference differences increasingly coming down to affinity for different voices. Latest results also reflect stronger bot vote filtering, confirmed via triangulation against third-party evaluators. We've also added rank ranges based on each model's 95% confidence interval, showing where a model could land based on its Elo score range. Key results: ➤ Most preferred: Current top 5 per our TTS leaderboard: 1. Inworld TTS 1.5 Max (Elo of 1,238); 2. ElevenLabs Eleven v3 (1,197); 3. Inworld TTS 1 Max (1,183); 4. Inworld TTS 1.5 Mini (1,182); 5. MiniMax Speech 2.8 HD (1,175) ➤ Price: Kokoro 82M v1.0 (Replicate) leads at $0.65 per 1M characters, followed by Inworld TTS 1 and 1.5 Mini at $5, and AsyncFlow V2 at $8.33 ➤ Speed: WaveNet leads for batch generation at 419 characters processed per second, followed by Kokoro 82M v1.0 (Replicate) at 235, and Inworld TTS 1.5 Mini at 214 See below for further detail ⬇️
@threepointone ·
landed a couple of new examples in the agents repo - structured inputs: how can your agent "ask" you questions that you answer as multiple choice/yes or no/ free form text/number slider/ whatever. tldr - use client side tools! - push notifications: how can your agent buzz you with a web push notification, even if you've closed your tab or whatever. - this in addition to the new elevenlabs starter earlier today. npm i agents
@TheAstornia ·
WSJ covered how strange passive income has become. One example was someone making money from licensing AI clones of his voice through ElevenLabs. Article says he makes about $3k a month from it, so passive income is not just rentals, dividends, or digital products anymore. It can literally be your synthetic voice getting paid.
@noahkagan ·
"At first I thought you were sick. Not yourself. Low mood. Then I was convinced you were reading a script. By the end I'd already twigged before you even said it was AI." "Your personality is the best thing about your show. This just zapped it." Here's the backstory: I've got a second daughter now. Time is basically zero. So I thought - what if I could automate my podcast production? I took an article I wrote about everything we're doing with AI at AppSumo, had Claude turn it into a podcast script, uploaded 2 hours of my old episodes to ElevenLabs, and hit generate. It sounded like me. Nasally Jewish voice and everything. I was about to just ship it. Then I sent it to Jason, my editor, to see if he'd notice. He noticed in 30 seconds and gave me the feedback from the start of this. And then he said something that stuck with me: "AI in creative roles is a race to the bottom sometimes. It's great for organizing and admin. Just no good at being a person or a creative one at that. Something a lot of your peers in the tech world might need reminding of before we spin into a black hole of bland homogenized guff." And later that day a guy named David walks up to me at a coffee shop and says "I love your podcast." We all want to do less work and get great results. But some things require the work. That IS the value. AI should be like spellcheck, it doesn't replace you - it gives you superpowers on top of what you already bring.
@PrajwalTomar_ ·
THIS IS CRAZY. I connected ElevenLabs to Lovable and built an AI storytelling app in UNDER 20 minutes. Enter a kid's name, pick a theme. The app writes a personalized bedtime story, narrates it with ElevenLabs, and plays ambient sounds that match the story world. → Text to Speech narration with emotion → Sound effects generated on the fly → Download the full story as audio Lovable connectors have made my life so easy.
@Shruti_0810 ·
Another billion-dollar AI category is becoming open source. Voice AI is following the same path as image generation and coding assistants. Pipecat is an open-source framework for building real-time voice AI agents. It handles the entire voice stack: → Speech-to-Text → LLM reasoning → Text-to-Speech → Real-time streaming → Conversation orchestration Everything you need to build: • AI phone agents • Customer support bots • Voice copilots • AI companions • Interactive assistants • Multimodal experiences What makes it different: • Swap any AI provider without rewriting your app • Plug in OpenAI, Claude, Gemini, Groq, Ollama, and more • Choose Deepgram, Whisper, AssemblyAI, or other STT providers • Use ElevenLabs, OpenAI, Cartesia, Azure, and more for TTS • WebRTC + WebSockets built in for ultra-low latency The biggest shift isn't a new model. It's that the entire infrastructure layer is becoming a commodity. The winners won't be the companies selling voice APIs. They'll be the ones building products users actually want. GitHub: https://t.co/fkj7JO2ul7
@RoundtableSpace ·
Someone open sourced the entire ElevenLabs + Descript stack in a single local tool. Voice-Pro: zero-shot voice cloning, Whisper transcription, vocal isolation & dubbing in 100+ languages. ElevenLabs charges $22/month for this.
@michuk ·
One thing that stood out during yesterday's product demos at the @ElevenLabs Summit in Warsaw was how much focus there was on applications, not models. That makes perfect sense. AI models — even ones as impressive as ElevenLabs' voice model — will eventually become commoditized. In fact, ElevenLabs itself showed a compact voice model running on a mobile device without internet access. Soon there will be open source models that do that for free for everyone who builds. And if that's where the technology is already heading, relying solely on token consumption is probably not a great long-term business. So what do model companies do? They move up the stack. The same way OpenAI launched ChatGPT, or GPT Health and Anthropic launched Claude Code, Desktop and is working on cybersec, ElevenLabs is increasingly building products: – creator marketplaces – tourism agents (demoed yesterday) – customer service agents (demoed at a prevous summit) – and likely many more vertical applications These products become the new moat, not (just) the model. And while this is a rational strategy for model companies, it's also a warning sign for startups building on top of them. If your entire business is "Claude/OpenAI/ElevenLabs + a prompt," you're vulnerable. Even if your UX is dope. The AI agent startups that survive will need moats of their own: – deep integrations into customer workflows – proprietary data – strong communities – distribution advantages – regulatory or industry expertise Otherwise they'll wake up one morning and discover their biggest supplier has become their biggest competitor. Check out the personalized tourist guide demo from the summit. #AI #ElevenLabs #Startups #VentureCapital
@etnshow ·
ElevenLabs is turning every chat agent into a voice agent. Their new product Voice Engine: • Wrap your existing chat agent, it becomes a voice agent • One command. One skill. It builds itself. • Add a data agent to a Zoom call • Turn a gov chat agent into a phone line anyone can ring No rebuild needed. Just wrap what you already have.
@MollySOShea ·
NEW: Inside AI's Biggest Downstream Winner.. the Surge in AI Database Demand "Data is the unsung hero, & data is back." @MongoDB CEO CJ Desai (@cj_mongodb) The unexpected result? Hyperscalers are turning away even top-50 accounts. "Sorry, we don't have a capacity." "And they are one of the top 50 customers for that hyperscaler." ElevenLabs alone runs "north of 50 million agents, depending on when you look at it, all running on MongoDB.. that gives us a lot of confidence that we have the right architecture for agentic workloads." NASDAQ: $MDB We cover: › The 3 classes of AI customers MongoDB serves › Why hyperscalers are telling top-50 accounts "no capacity" › On-prem, sovereign AI, and the data-center comeback › ElevenLabs running 50M+ agents on MongoDB › MongoDB (OLTP) vs Snowflake & Databricks (OLAP) › Why there is no standardization in enterprise AI models › Auto-scaling & the fall of expensive DBAs 𝐓𝐈𝐌𝐄𝐒𝐓𝐀𝐌𝐏𝐒 (00:00) CJ Desai, CEO at MongoDB (01:12) Top CEOs at the Raise AI Summit (03:24) The 3 types of companies powering the AI boom (07:25) Why on-prem is making a shocking comeback (09:01) Hyperscalers are quietly running out of capacity (11:06) What actually separates MongoDB from Snowflake x Databricks (13:18) Why Frontier Labs treat MongoDB as their memory layer (16:22) From Oracle intern to first-time CEO (19:19) Becoming CEO during the AI chaos (23:53) The World Cup crisis that tested MongoDB's scale (28:19) The real complexity behind simple AI agents (31:22) Open source vs. closed models: what customers actually pick (34:36) CJ's honest take on data centers in space (36:50) The mentors who shaped a first-time CEO (40:32) MongoDB's next big bets (41:36) Data is back?
@Carles_Reina ·
May has been a crazy month at @ElevenLabs ▪️ We announced that we had crossed $500M in ARR ▪️ Entered a strategic partnership with the Greek Government 🇬🇷 to bring AI to the public services, tourism, and preservation of Greek heritage ▪️ Launched Dubbing V2, the global SOTA model for dubbing ▪️ Launched Music V2, trained on fully licensed data and built for Musicians, Developers and Brands ▪️ Partnered with Stan Lee Univers to make his voice available on ElevenLabs Iconic Marketplace and the Reader app ▪️ We launched in Australia & New Zeland 🇳🇿 🇦🇺 ▪️ ElevenAgents can now handle images, PDFs, audio messages, contacts, and location pins on WhatsApp ▪️ Added Studio Agent in ElevenCreative, our AI co-editor to streamline the creative process ▪️ Brought Templates to ElevenCreative, our ready-made creative workflows
@pvergadia ·
BREAKING: @Microsoft just launched 3 new AI models — and they're going after OpenAI, Google, and ElevenLabs at the same time. MAI-Transcribe-1. MAI-Voice-1. MAI-Image-2. 3.9% WER best-in-class transcription. 60 seconds of audio generated in 1 second. 2x faster image gen than before. → MAI-Transcribe-1 beats GPT-Transcribe, Scribe v2, Gemini 3.1 Flash, AND Whisper-large-v3 → MAI-Voice-1 clones your voice from seconds of audio → MAI-Image-2 hit top 3 on Arenaai before even launching publicly → All 3 available today in Microsoft Foundry And the prices? $0.36/hr transcription. $22/1M chars for voice. Cheaper than anyone else. At Microsoft we are building a full model stack!
@ElevenLabsDevs ·
Introducing ElevenLabs Examples 2.0, a modern approach to the examples repository. The examples repo is one of the highest touch points for any developer-facing product, but maintaining them is a tedious task when you have a fast moving product. We rebuilt ours to be prompt-driven, AI-native, and autonomously maintained.
@svpino ·
I can't wait for the day we don't need to hear "your call is important to us" ever again. Literally nobody wants to talk to a bot on the phone, but we all hate waiting in line. Here is the thing about voice: We've been talking to each other for thousands of years, but we've only been typing for maybe 50. Voice is the most natural interface humans have. But voice AI still sounds robotic and awkward, with zero emotional awareness. If we can unlock voice AI at scale and address all of these issues, we'll get access to the most natural and powerful interface we've ever had. Here is the good news: I'm working with ElevenLabs, and they released a platform that lets you build and deploy voice and chat agents that sound really good. These voice agents are: 1. Real-time 2. Multilingual 3. With emotional inflection They work so much better. They don't make people hang up the phone immediately! What really matters here is that these agents can handle thousands of simultaneous conversations without making you wait through hold music. They integrate with existing systems, resolve issues, qualify leads, and escalate with full context when a human needs to step in. ElevenLabs got this right.
@WesRoth ·
ElevenLabs has expanded its creator ecosystem by launching a "Music Marketplace" within its ElevenCreative platform, allowing users to publish, distribute, and actively monetize their AI-generated tracks. To make the marketplace viable for enterprise use, every track is automatically available under three pre-defined commercial licensing tiers: Social Media, Paid Marketing, and Offline. This eliminates the need for custom legal negotiations.
@AlphaSignalAI ·
A 66M parameter model just beat ElevenLabs on a Raspberry Pi. Text-to-speech has lived in the cloud for years. Every spoken character cost an API call and a fraction of a cent. Supertonic 3 is an open-source TTS model that runs entirely on the device. No network, no key, no per-character billing. The model is 99M parameters and ships as an ONNX file. It hits 167x faster than real-time on a laptop CPU. That works out to about 1,263 characters of speech per second. Larger open systems sit closer to 55 to 287. What the on-device design unlocks: > Runs offline on a Raspberry Pi > Works inside a browser tab > Handles phone numbers and currency > Reads dates without any preprocessing > Inline tags for laugh and breath Language coverage jumped from 5 to 31 in this release. The public interface stayed identical to the previous version.
@thealexbanks ·
I love how Claude desktop now supports ElevenLabs Agents. You can build a voice AI support agent in minutes. Here's how it works: 1. Add the connector in Claude desktop ↳ Customise > Connectors > ElevenLabs Agents MCP ↳ Paste your API key and save
@eddiejaoude ·
I was experimenting with ElevenLabs this evening. The audio creation for my voice is very impressive, the lip-sync video not at all. My results: ✅ realistic text to speech creation ❌ video creation from photos using lip-sync DevRels are still needed 🥳 but we can use AI to be more productive 🚀! Is this audio actually me speaking or AI generated? 🤔
@wadefoster ·
Zapier x Meta Ray-Ban glasses. Here’s how it works: Our own Phillip McKee built a voice-first AI assistant using the Zapier SDK, Claude, and ElevenLabs. You say "Hey Zara" and it creates tasks, sends Slack messages, checks your calendar, and chains actions together. All hands-free through the glasses mic. The @Zapier SDK powers every action. It also learns preferences over time so repeat tasks get faster. This is what the SDK was built for - developers wiring AI into hardware we never anticipated. Post what you're building and tag me. I want to see it
@AlexEngineerAI ·
Here is what a voiceover side hustle on Fiverr actually looks like, month by month. Month 1 You set up ElevenLabs ($22/month Starter: 100k characters). You create a gig: AI-assisted narration for scripts, explainer videos, or YouTube intros. You price at $10 for 500 words because you have zero reviews. You send 30 buyer requests in two weeks. Two people respond. One buys. Gross for Month 1: around $30. Fiverr takes 20%. Your net: $24. Month 2 You complete 4-5 orders. Two 5-star reviews. You raise your base price to $20. Keep the $10 tier as a basic option to stay competitive. Gross: $80-120. Net after Fiverr: $64-96. Month 3 You hit Level One Seller status. Fiverr requires 60 days active, 10 completed orders, a 4.7 or higher rating, and a 90% response rate. Level One unlocks gig extras and better search placement. You add a 24-hour delivery option at $15 and a revision extra at $10. Average order value rises to $30. Gross: $150-250. Net: $120-200. Months 4-6 This is where it compounds or stalls, and the difference is usually niche. Generalist voiceover gigs compete with hundreds of sellers at the same price point. Sellers who specialize in SaaS explainer videos, real estate listing narration, or YouTube tech intros charge 2x to 3x more and attract repeat clients. By Month 6 with a clear niche and 20 or more reviews: 10-20 orders per month at $35-60 each. Gross: $350-1,200. Net after Fiverr 20%: $280-960. ElevenLabs: $22/month. Realistic net at Month 6: $258-938. Month 1 and 2 are the grind nobody talks about. You are essentially buying reviews at $5-10 per hour. That is the price of entry on the platform. Push through months 1-3 and the math gets meaningfully better.
@thealexbanks ·
Brad Smith hasn't been able to speak for years. Now Neuralink lets him type with his brain and ElevenLabs lets him speak in his own voice. He just gave an interview on national TV. Brad was diagnosed with ALS at 37. He's now 45. The disease destroyed the motor neurons connecting his brain to his muscles. Brad is quadriplegic and nonverbal, but his mind is completely sharp. Before Neuralink, he relied on eye-gaze tracking to communicate. It was slow, exhausting, and only worked in dark environments. Now, two technologies have converged to change everything: Neuralink handles the input: → Brad controls a cursor by imagining tongue movements → Clicks are performed by imagining jaw clenches → Pure thought with no physical motion required → He can type, edit videos, and even play Mario Kart ElevenLabs handles the output: → AI-cloned his pre-ALS voice from old recordings → Trained on his speech patterns, tone, and inflections → Brad types what he wants to say and the AI speaks it in his voice Neuralink gave Brad the ability to communicate with his brain. ElevenLabs gave him back the voice that makes him who he is. Through the ElevenLabs Impact Program, the company has committed $1 billion in free voice restoration technology to reach 1 million people living with permanent voice loss. They've already supported ~7,000 individuals across 49 countries. I know I spend a lot of time debating AI's risks and limitations. But when a father who hasn't spoken in years can speak to his kids again, in his own voice, controlled entirely by thought, is the other side of the conversation that I think deserves more attention. Video source: NewsNation on YouTube
@HeyOz_AI ·
5 AI tools that saved me 40 hours of design work this week. Most marketers are still manually editing video layers. The top 1% are building "Creative Factories." Here is the exact stack (and prompts) I used to launch 50+ ads yesterday: 1. Oz (Bulk Creative Generation) This is the engine. I don't build one ad at a time anymore. I feed it my product assets, and it generates dozens of static and video variations instantly. • It handles the layout. • It handles the "ugly" aesthetic. • It handles the volume. Use this for: The heavy lifting. Replacing your junior designer. 2. ChatGPT (Infinite Hook Scripts) I don't write scripts from scratch. I iterate on winning formulas. Prompt: "I am selling [Product] to [Target Persona]. Write 10 TikTok-style ad hooks under 15 words. Structure them as: 3 'Negative' hooks (e.g., 'Stop doing X') 3 'Benefit' hooks (e.g., 'How to get Y instantly') 4 'Curiosity' hooks (e.g., 'The secret nobody tells you about Z') Keep the tone raw, conversational, and punchy." 3. NanoBanana (Background Assets) Stop using stock photos everyone has seen. NanoBanana lets you generate consistent, specific background assets or modify existing images with incredible precision. Prompt: "Generate a [Setting, e.g., messy desk with coffee stains] background. Style: iPhone photo, low fidelity, flash photography. Ensure it looks user-generated, not studio-perfect." 4. ElevenLabs (Voiceovers) The default AI voices sound robotic. You need "imperfect" delivery for ads. Settings to use: • Model: Eleven Multilingual v2[2] • Stability: 35% (Lower stability = more natural fluctuation) • Similarity Enhancement: 75% 5. CapCut (Assembly & Captions) The final assembly line. • Import your Oz visuals + ElevenLabs audio. • Use "Auto-Captions." • Workflow Hack: Search for "UGC Ad" templates in CapCut, strip the stock footage, and drag-and-drop your Oz-generated clips into the slots. Time saved: 3 hours per video.
Best Tweets by Topic