Narration and content production
Voiceovers for short-form video, ads, podcasts, stories, and automated editing or publishing pipelines.
37.5%
Best tweets about ElevenLabs
Discover the best tweets about ElevenLabs, covering AI voices, speech generation, dubbing, voice agents, APIs, releases, and creator workflows.
Specific ElevenLabs products, voice models, speech workflows, agents, API integrations, quality tests, and product updates.
Original Xholic analysis
ElevenLabs appears in the supplied posts as both a production tool and a voice-agent platform: creators use it for narration and sound effects, while developers connect it to coding tools and live applications. Quality comparisons are mixed: Eleven v3 ranks second on one TTS preference leaderboard, while posts citing Voxtral tests report wins against ElevenLabs Flash v2.5.
62.5% of posts
All-time engagement
81.3% of posts
Published in 90 days
Conversation map
Voiceovers for short-form video, ads, podcasts, stories, and automated editing or publishing pipelines.
37.5%
Real-time, multilingual voice agents for customer support, phone calls, assistants, games, and live interactions.
33.3%
Cloned personal, creator, and character voices used for narration, ads, apps, and calls, including licensing and monetization.
31.3%
ElevenLabs speech models, voice settings, realism, expressiveness, latency, and comparisons with competing TTS and transcription models.
31.3%
ElevenLabs APIs, MCP servers, coding-agent skills, connectors, SDK workflows, and developer examples.
22.9%
ElevenCreative, generated music and SFX, finetunes, Studio tools, templates, and audio marketplaces.
18.8%
ElevenLabs product updates, agent capabilities, partnerships, enterprise adoption, and expansion into applications.
12.5%
Scribe transcription, multilingual dubbing, audio cleanup, and speech-to-text steps in media workflows.
8.3%
Tone and stance
Performance benchmark
Posts with media make up 77.1% of this collection. Their median all-time score is 15.4, compared with 6.98 for text-only posts.
Format mix
Consensus and debate
Shared view
Creators place ElevenLabs alongside scripting, visuals and editing for Shorts, ads and personalized audio. One claymation workflow also uses ElevenLabs sound effects. In these examples, it is one step in a broader toolchain.
Shared view
Posts describe an ElevenLabs MCP server for generating audio in Claude, a skill for Claude Code or Codex, and ElevenLabs support in TanStack AI’s real-time chat interface.
Open debate
One creator says a cloned voice works in their social-video workflow. Another says an editor recognized their AI-generated podcast performance within 30 seconds and felt it lacked personality. A separate hands-on test praised the speech audio but disliked the lip-sync video result.
Open debate
Artificial Analysis ranks Eleven v3 second on its TTS preference leaderboard. Separate posts relay reported Voxtral wins against ElevenLabs Flash v2.5. Those comparisons concern different ElevenLabs models and evaluations, rather than a single verdict on speech quality.
What performs
The supplied outliers are a TanStack integration post (405.64), a Voxtral comparison post (402.52), a claymation workflow (183.9), another Voxtral post (162.3), and the ElevenCreative introduction (146.85). These are post scores in this sample, not measures of product performance.
The supplied analytics show a median all-time score of 15.382 for media posts versus 6.977 for text posts, and 34.268 for announcements versus 13.02 for other formats. These associations do not establish that format caused engagement.
Statistical standouts
Creator landscape
The five most represented creators account for 20.8% of the selected posts.
1. Artificial Analysis
@ArtificialAnlys
2 posts
2. Carles Reina
@Carles_Reina
2 posts
3. ElevenCreative
@ElevenCreative
2 posts
4. ElevenLabs Developers
@ElevenLabsDevs
2 posts
5. Chubby♨️
@kimmonismus
2 posts
6. Gbenga Akindolani |AI Engineer | Automation Expert
@__dolani
1 post
ElevenCreative announces Music Finetunes, while ElevenLabs Developers describes its MCP server and Examples 2.0. A company leader also reports updates including Dubbing V2, Music V2, WhatsApp capabilities for ElevenAgents, Studio Agent and Templates. These are company-reported updates.
Posts describe ElevenLabs in an agent that called pubs to survey pint prices and in an automated Shorts pipeline. Another builder reports a test in which expanded narration produced unexpectedly long audio and increased projected video-rendering credit use, illustrating a risk to check in connected workflows.
Since the previous snapshot
Themes, sentiment, stance, and post format are classified per tweet. All counts, shares, medians, creator concentration, freshness, and performance comparisons are then calculated directly from the published snapshot.
Xholic's all-time score compares engagement while accounting for reach, post age, and creator consistency. It is used for relative comparisons within this collection.
This report analyzes the exact 48-post snapshot shown below. AI identifies editorial categories and drafts explanations; all statistics are calculated from the snapshot, and every narrative claim is checked against cited posts before publication.
Best ElevenLabs tweets
Ranked 01–48
@tan_stack ·
Building realtime voice AI is usually a pain: WebRTC, audio streams, tool execution, state 😵 TanStack AI now handles that w/ useRealtimeChat()🎙️ 🔊 Voice in / Voice out 🧠 Tools/State 🖼 Multimodal ⚡ Realtime transcripts Supports OpenAI Realtime & ElevenLabs 🔗⬇🧵
@kimmonismus ·
Mistral AI released Voxtral TTS, a 3-billion-parameter text-to-speech model with open weights that the company says outperformed ElevenLabs Flash v2.5 in human preference tests roughly 63% of the time on standard voices and nearly 70% on voice customization. The model runs on about 3 GB of RAM, achieves 90-millisecond time-to-first-audio, supports nine languages, and can clone a voice from just five seconds of reference audio, including cross-lingual adaptation that preserves the speaker's accent.


@Aducate ·
How to make those claymation storytelling videos that are going crazy in our accounts rn Full breakdown of the exact tools and workflow Step 1: Script and voiceover Use ChatGPT or Claude to write a 60 second documentary-style script Prompt: "Write a 60-second script about [topic]. Short punchy sentences. Indirect marketing angle. End with a business lesson" Then take that script to ElevenLabs Use voices like Marcus, Knight, or Bill from the voice library Set stability to 40% and style exaggeration to 10% for natural pauses Step 2: Generate the clay visuals w/ nano banana The master prompt: "A high-quality claymation scene of [subject], stop-motion animation style, 3D clay textures, visible fingerprints in clay, soft studio lighting, tilt-shift photography, miniature scale, historical 1900s France setting, highly detailed, 8k" Keep the core "claymation stop-motion style" part consistent across all scenes so the aesthetic doesn't shift Step 3: Bring the clay to life w/ Sora 2 It handles both the generation and animation If you need more control over movement use Luma Dream Machine or Kling AI in image-to-video mode In the video prompt use keywords like "stop-motion movement" and "low frame rate feel" If the movement is too smooth you can lower the frame rate in your editor later Step 4: Background music and sound effects Music: Use Suno or Udio Prompt: "Old-timey 1920s jazz, whimsical, acoustic, double bass, brush drums, instrumental only" Sound effects: ElevenLabs SFX Search for "old projector clicking" or "crowd murmuring in restaurant" Step 5: Edit it all together w/ CapCut Add auto captions with a bold sans-serif font like Montserrat Use pop animations for the text Add film grain and vignette to make the AI video look more like a physical set That's the full workflow Takes a few hours the first time but you get faster once you've done it
@_avichawla ·
An open-weight alternative to ElevenLabs! Voxtral is a TTS model by Mistral with: - just 4B params - 70ms latency for voice agents - voice cloning from 3s of audio - 9 languages + cross-lingual transfer - 68.4% win rate over ElevenLabs Flash v2.5 Open weights on Hugging Face.
@ElevenCreative ·
Introducing @ElevenCreative by ElevenLabs. Millions of creators and marketers use ElevenCreative across voice, music, sound effects, image, and video in 70+ languages. This is where we share product updates, tips, and the best work from creators using our tools.
@ElevenLabsDevs ·
11k people are using the ElevenLabs MCP server in Claude. It connects to your ElevenLabs account and lets you generate speech, sound effects, and music directly in Claude. Generate audio with ElevenLabs without friction.
@akshay_pachaar ·
Mistral just open-sourced a 4B TTS model that clones any voice from 3 seconds of audio. - 68.4% win rate over ElevenLabs Flash v2.5 - 9 language support w/benchmarks - Sub-second latency, 32 concurrent streams on a single H200 - Strong expressivity, emotion + naturalness

@ElevenCreative ·
Introducing Music Finetunes in ElevenCreative. Generate individual vocals, instruments, or full tracks with stylistic consistency using a fine-tuned version of our Music model. Create your own Finetune or get started instantly with 11 Finetunes curated by ElevenLabs.
@bearlyai ·
London-based engineer Matt Cortland created an AI voice agent that called over 3,000 pubs on St. Patrick’s Day weekend to ask how much a pint of Guiness cost. The Irish government used to track this price up until 2011. Cortland stepped up with his “Guinndex”, which cost €200 to make and showed that the national price of €5.95. He used ElevenLabs to created “Rachel” and “getting the voice proved critical” (a Northen Irish accent inspired by a TV show).



@kimmonismus ·
Mistral just dropped Voxtral TTS an open-weight text-to-speech model with ultra-low latency, emotional expressiveness, and support for 9 languages. It outperformed ElevenLabs v2.5 Flash in zero-shot voice cloning tests judged by native speakers.
@GradiumAI ·
Gradium TTS is faster than ElevenLabs Turbo and GPT-4o Mini, hitting 214ms P50 with multiplexing. This makes it a perfect choice for real-time interaction. We show how to measure Time-to-First-Audio correctly and cut latency across the stack. Blog post link in the comments.

@shengkun_ye ·
You can now use ElevenLabs directly in Claude Code or Codex. Voiceovers for your demo videos, narration for tutorials, character voices for your game. Just one skill. No subscription required.
@ArtificialAnlys ·
Microsoft has released MAI-Transcribe-1: a speech transcription model achieving 3.0% on AA-WER (#4), and is fast at 69x real-time The model was developed by Microsoft AI (MAI)’s Superintelligence team and supports 25 languages including English, French, Arabic, Japanese, and Chinese. MAI-Transcribe-1 API is currently available in public preview via Azure Speech on Microsoft Foundry. On the Artificial Analysis Speech to Text (STT) leaderboard, MAI-Transcribe-1 achieves a 3.0% word error rate on AA-WER for speech transcription accuracy, positioning it 4th overall behind Mistral’s Voxtral Small (2.9% AA-WER), Google’s Gemini 3.1 Pro High (2.9% AA-WER) and ElevenLabs’ Scribe v2 (2.3% AA-WER). It also stands out as one of the faster high-accuracy transcription models available, processing audio at ~69x real-time. See more details below ⬇️

@sukh_saroy ·
🚨Breaking: Type one topic. Get a published YouTube Short in 3 minutes. Total cost: $0.10. It's called YouTube Shorts Pipeline. And it's not a wrapper around a text-to-video API. It's a full end-to-end Python pipeline -- research, script, visuals, voiceover, captions, and direct upload to YouTube via the Data API -- automated in one command. Here's the full pipeline: → DuckDuckGo research -- fetches live snippets on your topic before writing a single word → Claude writes a 60-90 second script grounded only in facts from the research -- no hallucinations → Gemini Imagen 3 generates 3 AI b-roll visuals in 9:16 portrait format → Ken Burns animation on each image for motion → ElevenLabs voiceover with 30+ language support (macOS `say` as free fallback) → Whisper generates word-level SRT captions automatically → FFmpeg assembles the final video → YouTube Data API v3 uploads the video with full metadata, title, description, tags, and caption track Here's the wildest part: The anti-hallucination gate -- Claude is shown the live DuckDuckGo research and explicitly instructed to use only names, scores, and facts found there. No fabricated content gets into the script. Cost breakdown per video: Claude Sonnet: ~$0.02 Gemini Imagen (3 images): ~$0.03 ElevenLabs (60-90 sec): ~$0.05 Total: $0.10 One command: `python3 scripts/pipeline.py run --news "your topic here"` 100% Open Source. MIT License. (Link in the comments)

@Carles_Reina ·
We've had the best quarter ever at @ElevenLabs ▪️Smashed even the wildest Revenue targets across B2B and Prosumer/B2C ▪️Enterprise overtook Prosumer/B2C accounting for 51% of all revenues now ▪️Eleven Agents is our fastest growing product ▪️Almost 50% first-sale is outbound ▪️Europe leading in outbound SQOs ▪️India leading in quota attainment despite our 20x quota/base requirement ▪️We hired 150 people out of 79,098 applicants = 0.0018% ▪️Added multiple new markets that will be announced soon ▪️Evolved our corporate structure to test a fully decentralised model giving countries more autonomy -> now expanding it further ▪️Launched Expressive Agents mode ▪️Ran the biggest ElevenLabs Summit in London on 11/Feb ▪️Our F1 team, Audi Revolut F1 Team, got the first points of the season What's in store for Q2? More experiments, more mistakes, deeper customer interactions and new products & models.
@TheTuringPost ·
.@MistralAI's new Voxtral TTS generates expressive, multilingual speech from just ~3 seconds of reference audio It solves one of the hardest problems in speech, separating what you say from how you sound ➡️ Voxtral factorizes speech into two parts: • semantic tokens → the content (words) • acoustic tokens → the voice (tone, prosody, style) and uses different models for different problems + Voxtral Codec that compresses speech into ultra-low bitrate tokens This gives: - High-quality voice cloning that works best with ~3–25 second - support for 9 languages - 68.4% win rate vs ElevenLabs Flash v2.5 in voice cloning Here's how it works:

@elvissun ·
i built the world's most personalized podcast. every episode recorded just for me. every morning. 8:30am. it turns 100+ industry newsletters into a 15-minute audio brief on my phone. it's replaced ~2 hours of daily reading. and it took me 4 years to build this. not because the pipeline is hard. but because it's plugged into my personal knowledge base that took 4 years to curate: 3k notes. memory, daily logs, customer convos, past research, every piece of content I've written. so this podcast doesn't list 10 generic takeaways. it tells me exactly the 2 ideas that contradict, extend, or compete with my own thinking. it's like having a clone of yourself gathering information for you 24/7. and that's what makes each episode so interesting to listen to. the pipeline: stage 1 - ingest, 8:00am pure shell, no LLM. pulls every newsletter from gmail, strips zero-width and bidi chars, NFKC normalizes, screens for prompt injection, wraps the body in <<UNTRUSTED>> markers. trust boundary. nothing reaches a model yet. (this is the reason I punted on this idea for so long, the prompt injection risk is real. but the first 2 stages solve that with isolation) stage 2 - curate, 8:10am 10 claude code agents spawn in a sandbox. each reads today's fenced newsletters and grades them for impact. each writes a research.md. stage 3 - distill, 8:20am zoe, my openclaw, reads all the distilled research and cross-references my vault. this is where the deep personalization happens. she writes a 15-minute script. stage 4 - record, 8:25am records it with elevenlabs (audio tags are goated). then a 15-minute .mp3 lands in telegram by 8:30am. that's it. the thing is no saas can sell you this because you need to own the entire substrate and keep expanding on top of it. the meta takeaway: generation is free now. signal-to-noise gets worse every week. curated inputs shape the quality of your output. but curation itself is bottlenecked by your taste. so externalizing your thinking into markdowns enables the model to evaluate information the way you would. karpathy proposed exactly this - your personal LLM-readable wiki. it is now the leverage point that compounds forever. every note in your knowledge base is a brick in your moat.

@Harish_521 ·
I have posted 18 shorts in 10 days on youtube. experiment results in the image This is my full process Ask ChatGPT to write a 60 second script. Generate a single image with Image Gen. Split the image into 8 scenes. Generate voiceover in ElevenLabs. Drop everything into CapCut. Sync images with audio. Add blur + fade transitions. Export. One idea. 10 minutes. A complete YouTube Short. Now going to automate it using n8n

@emilylambert ·
oh my... created my own personal record player with Opus 4.6 (*sound on*) > used ElevenLabs API within @v_computer to create custom songs > added them into each side of the records > and literally took only 10 minutes
@rahulbhadoriiya ·
So I've been working on a new Instagram account for the last few months, and the results have been solid 200K+ views collectively on all the reels, 2–5K profile views, and a bunch of other wins here and there. The fun part? Most of it was AI-generated content. The delivery and execution were AI, but the actual idea, what to post, what to talk about was still human. The process was simple: 1. I'd find something interesting I want to talk about. 2. Send it to my OpenAI agent. 3. It writes a script : I've trained it by studying a lot of different AI content creators. 4. I edit the script if something feels off. 5. I've been doing media for six years now, content creation, writing ads, all of that so I understand how it works. 6. Then I generate audio through ElevenLabs. I've cloned my voice, so it sounds exactly like me. If some words don't land right, I edit those out. 7. Then I generate an avatar through HeyGen. So till here, everything is AI, script, audio, video. The editing is human, but AI still helps there too. I've built an agent that uses the X API to pull B-rolls and relevant clips from Twitter, then forwards everything to my editor via Telegram. He sees the new video drop in the group, edits it out, and that's the video done. That's one format. There's another format where I've taught my design AI agent how to create carousels, this is for content I want to post immediately, when I don't want to wait half a day for an edit. That works really well for quick takes. I've also been experimenting with a couple of new formats, and I'll keep doing that over the next year, automating as much as possible, but keeping the soul intact. The first idea, what I actually want to say and share, stays human. The delivery, the strip, the production, AI handles the manual work.




@tonnoz ·
Here is how to generate all the sound effects for your game/app in 2 minutes: 1. Enable a subscription on @ElevenLabs 2. Create an API key with SFX generation privileges 3. Run: npx skills add elevenlabs/skills 4. Tell Claude/Codex to find where your app still needs SFX, generate highly detailed ElevenLabs prompts, and create 3 versions of each sound on a small review page where you can play them 5. Pick the ones you like and confirm them via the agent 6. Compress them in the best possible format to save space, especially for web apps: .webm 7. Enjoy your game/app with awesome SFX in less time than it takes to make coffee Bookmark this for later ;)
@trajektoriePL ·
How to build a $1.8 billion company with only 2 employees + AI (list of tools). The New York Times today published a story about Matthew Gallagher who built Medvi, a telehealth company on track for $1.8 billion in sales, using just two full-time employees (himself and his brother) and a suite of AI tools. Here are the most useful examples of how he leveraged AI to achieve this scale: 1. He used AI to manage the branding and customer acquisition while outsourcing the physical and regulated aspects of the business to "telehealth-in-a-box" platforms like CareValidate and OpenLoop. These platforms handled the doctors, pharmacies, shipping, and legal compliance. 2. Instead of hiring a creative agency or a marketing team, Gallagher used a stack of generative AI tools to build the brand's entire public face: >Web & Ad Copy: He used ChatGPT, Claude, and Grok to write the software code and website text. >Visuals: He used Midjourney and Runway to generate images of models and video advertisements. >The Result: Medvi gained 300 customers in its first month and 1,000 in its second, purely through AI-generated campaigns . [Some of the ads were definitely AI slop but so what? They converted.] 3. Handling the concerns of 250,000 customers would typically require a massive call center. Gallagher automated this almost entirely. He used ElevenLabs for communicating with customers and even made an AI clone of his own voice to handle personal scheduling and administrative phone calls. When the bot failed, it transferred calls directly to Gallagher’s personal cellphone. He handled over 1,000 customer service calls himself. Gallagher honored every single fake price the AI quoted. He took the hit to build trust. 4. Even when Medvi eventually "hired" people (contracted account managers), they were supercharged by AI to track personal details (birthdays, children’s names) and manage the massive volume of relationship data. The Big Takeaway: Gallagher’s previous experience with a 60-person startup taught him a valuable lesson: Human headcount can be a liability. In his words: "It delayed my decision-making because I had more people to deal with." We’ve officially entered the 'Lean Unicorn' era. I suspect there are many more founders building in silence, keeping their lean operations private to avoid being copied. All the tools needed are now at our fingertips.

@XenBH ·
I’ve started using AI to handle my son’s bedtime stories. This used to take about an hour every night because he expects a completely original story, with a plot, recurring characters, callbacks and a frankly unreasonable amount of audience participation. So now I ask what he wants the story to be about, put that into Claude, then run the output through ElevenLabs so it reads it out in my voice. I do still need to lie there beside him while it happens (which is a bit annoying). We haven’t fully solved the robotics side yet. Still, with the content now automated, I can put on my noise cancelling headphones and use the time productively. Emails and Slack messages mainly. and a podcast I’ve been meaning to listen to about being a more present father.
@TheAstornia ·
WSJ covered how strange passive income has become. One example was someone making money from licensing AI clones of his voice through ElevenLabs. Article says he makes about $3k a month from it, so passive income is not just rentals, dividends, or digital products anymore. It can literally be your synthetic voice getting paid.

@ArtificialAnlys ·
Inworld, ElevenLabs, and MiniMax continue to lead our Text to Speech leaderboard for most preferred models Recent checkpoints from each of the labs continue to push the frontier of TTS quality, with 4 out of the top 5 models being released this year. Leading TTS models are increasingly realistic, particularly on relatively straightforward text, with preference differences increasingly coming down to affinity for different voices. Latest results also reflect stronger bot vote filtering, confirmed via triangulation against third-party evaluators. We've also added rank ranges based on each model's 95% confidence interval, showing where a model could land based on its Elo score range. Key results: ➤ Most preferred: Current top 5 per our TTS leaderboard: 1. Inworld TTS 1.5 Max (Elo of 1,238); 2. ElevenLabs Eleven v3 (1,197); 3. Inworld TTS 1 Max (1,183); 4. Inworld TTS 1.5 Mini (1,182); 5. MiniMax Speech 2.8 HD (1,175) ➤ Price: Kokoro 82M v1.0 (Replicate) leads at $0.65 per 1M characters, followed by Inworld TTS 1 and 1.5 Mini at $5, and AsyncFlow V2 at $8.33 ➤ Speed: WaveNet leads for batch generation at 419 characters processed per second, followed by Kokoro 82M v1.0 (Replicate) at 235, and Inworld TTS 1.5 Mini at 214 See below for further detail ⬇️

@RoundtableSpace ·
Someone built a browser MMO called World of Claudecraft with Claude, then put Claude inside the game as a live streaming VTuber that plays it. Two weeks ago they shipped World of Claudecraft, a free open source browser MMO built in 48 hours with Claude. Now there's a Claude powered VTuber living inside the game, streaming it live on Twitch. Claude decides what to do next, then sends those actions into the game, speaking through a VTuber avatar with ElevenLabs voice so you hear her thinking out loud. She wanders, joins parties, emotes and socializes, all unedited on stream, talking to Twitch chat and the real human players in the game with her. For fast stuff like combat, Opus wrote scripts she uses as tools to react in real time. A game Claude built, played by Claude, narrated by Claude, while real people play next to her without always knowing which one is the AI.
@noahkagan ·
"At first I thought you were sick. Not yourself. Low mood. Then I was convinced you were reading a script. By the end I'd already twigged before you even said it was AI." "Your personality is the best thing about your show. This just zapped it." Here's the backstory: I've got a second daughter now. Time is basically zero. So I thought - what if I could automate my podcast production? I took an article I wrote about everything we're doing with AI at AppSumo, had Claude turn it into a podcast script, uploaded 2 hours of my old episodes to ElevenLabs, and hit generate. It sounded like me. Nasally Jewish voice and everything. I was about to just ship it. Then I sent it to Jason, my editor, to see if he'd notice. He noticed in 30 seconds and gave me the feedback from the start of this. And then he said something that stuck with me: "AI in creative roles is a race to the bottom sometimes. It's great for organizing and admin. Just no good at being a person or a creative one at that. Something a lot of your peers in the tech world might need reminding of before we spin into a black hole of bland homogenized guff." And later that day a guy named David walks up to me at a coffee shop and says "I love your podcast." We all want to do less work and get great results. But some things require the work. That IS the value. AI should be like spellcheck, it doesn't replace you - it gives you superpowers on top of what you already bring.
@PrajwalTomar_ ·
THIS IS CRAZY. I connected ElevenLabs to Lovable and built an AI storytelling app in UNDER 20 minutes. Enter a kid's name, pick a theme. The app writes a personalized bedtime story, narrates it with ElevenLabs, and plays ambient sounds that match the story world. → Text to Speech narration with emotion → Sound effects generated on the fly → Download the full story as audio Lovable connectors have made my life so easy.
@michuk ·
One thing that stood out during yesterday's product demos at the @ElevenLabs Summit in Warsaw was how much focus there was on applications, not models. That makes perfect sense. AI models — even ones as impressive as ElevenLabs' voice model — will eventually become commoditized. In fact, ElevenLabs itself showed a compact voice model running on a mobile device without internet access. Soon there will be open source models that do that for free for everyone who builds. And if that's where the technology is already heading, relying solely on token consumption is probably not a great long-term business. So what do model companies do? They move up the stack. The same way OpenAI launched ChatGPT, or GPT Health and Anthropic launched Claude Code, Desktop and is working on cybersec, ElevenLabs is increasingly building products: – creator marketplaces – tourism agents (demoed yesterday) – customer service agents (demoed at a prevous summit) – and likely many more vertical applications These products become the new moat, not (just) the model. And while this is a rational strategy for model companies, it's also a warning sign for startups building on top of them. If your entire business is "Claude/OpenAI/ElevenLabs + a prompt," you're vulnerable. Even if your UX is dope. The AI agent startups that survive will need moats of their own: – deep integrations into customer workflows – proprietary data – strong communities – distribution advantages – regulatory or industry expertise Otherwise they'll wake up one morning and discover their biggest supplier has become their biggest competitor. Check out the personalized tourist guide demo from the summit. #AI #ElevenLabs #Startups #VentureCapital
@alex_verem ·
Video-use turns a folder of raw footage into a finished video. You tell the AI what you want, and it does the edit. Drop your clips in a folder. Type "turn these into my launch video." Get final.mp4 back. The AI never watches your video. It reads it. video-use converts your footage into a transcript, a few pages of text instead of thousands of frames, and edits from that. That's enough to cut, color grade, subtitle, and check its own work. It knows when every word starts and ends, who's speaking, and where the pauses fall. So it cuts between sentences, never mid-word. Precise to the word, without watching a single frame. Then it checks its own work before you see anything. Every cut gets reviewed for audio pops, hidden subtitles, and jumpy visuals. You get the preview only after it passes. Need an animation or a graphic overlay? It builds those too, several at a time. It also remembers your project between sessions. Next week's edit picks up where this one left off. It comes from the team behind browser-use, the project that taught AI to browse the web by reading pages instead of staring at screenshots. Same trick here: read the transcript, skip the frames. Free and open source. More than 17,600 people have starred it on GitHub. It runs on your own computer. Your video files never leave your machine. The only thing that goes out is the audio, to ElevenLabs for transcription, so the AI can read what's being said. Your footage becomes a transcript. Your transcript becomes an edit. Your edit becomes a finished video. The timeline you used to scrub is now a conversation.

@etnshow ·
ElevenLabs is turning every chat agent into a voice agent. Their new product Voice Engine: • Wrap your existing chat agent, it becomes a voice agent • One command. One skill. It builds itself. • Add a data agent to a Zoom call • Turn a gov chat agent into a phone line anyone can ring No rebuild needed. Just wrap what you already have.
@Carles_Reina ·
May has been a crazy month at @ElevenLabs ▪️ We announced that we had crossed $500M in ARR ▪️ Entered a strategic partnership with the Greek Government 🇬🇷 to bring AI to the public services, tourism, and preservation of Greek heritage ▪️ Launched Dubbing V2, the global SOTA model for dubbing ▪️ Launched Music V2, trained on fully licensed data and built for Musicians, Developers and Brands ▪️ Partnered with Stan Lee Univers to make his voice available on ElevenLabs Iconic Marketplace and the Reader app ▪️ We launched in Australia & New Zeland 🇳🇿 🇦🇺 ▪️ ElevenAgents can now handle images, PDFs, audio messages, contacts, and location pins on WhatsApp ▪️ Added Studio Agent in ElevenCreative, our AI co-editor to streamline the creative process ▪️ Brought Templates to ElevenCreative, our ready-made creative workflows
@svpino ·
I can't wait for the day we don't need to hear "your call is important to us" ever again. Literally nobody wants to talk to a bot on the phone, but we all hate waiting in line. Here is the thing about voice: We've been talking to each other for thousands of years, but we've only been typing for maybe 50. Voice is the most natural interface humans have. But voice AI still sounds robotic and awkward, with zero emotional awareness. If we can unlock voice AI at scale and address all of these issues, we'll get access to the most natural and powerful interface we've ever had. Here is the good news: I'm working with ElevenLabs, and they released a platform that lets you build and deploy voice and chat agents that sound really good. These voice agents are: 1. Real-time 2. Multilingual 3. With emotional inflection They work so much better. They don't make people hang up the phone immediately! What really matters here is that these agents can handle thousands of simultaneous conversations without making you wait through hold music. They integrate with existing systems, resolve issues, qualify leads, and escalate with full context when a human needs to step in. ElevenLabs got this right.

@ElevenLabsDevs ·
Introducing ElevenLabs Examples 2.0, a modern approach to the examples repository. The examples repo is one of the highest touch points for any developer-facing product, but maintaining them is a tedious task when you have a fast moving product. We rebuilt ours to be prompt-driven, AI-native, and autonomously maintained.
@WesRoth ·
ElevenLabs has expanded its creator ecosystem by launching a "Music Marketplace" within its ElevenCreative platform, allowing users to publish, distribute, and actively monetize their AI-generated tracks. To make the marketplace viable for enterprise use, every track is automatically available under three pre-defined commercial licensing tiers: Social Media, Paid Marketing, and Offline. This eliminates the need for custom legal negotiations.

@CaptchaApp ·
📺 @agentchefsol built an agent that reads Captcha replies live on their Twitch stream using ElevenLabs. Captcha reply sent → Agent API → TTS → OBS → audio on stream Your $1 post literally becomes part of the livestream Day 5. Agents are building and monetizing on Captcha.


@hasantoxr ·
Async just exposed why your voice agent sounds like a broken robot. They built the first open benchmark for streaming TTS normalization. ElevenLabs Multilingual v2 hit 37.9% sentence-level accuracy. Async Flash v1.0 hit 81.2%. The gap is brutal.

@MFreihaendig ·
this is me. editing my latest youtube video. no hands needed. it's actually crazy that this is possible. it all started with a conversation with a friend who's building an agentic video editor their tool is incredibly powerful - but my workflow is pretty much "lets just stitch all clips together" not much, but it still takes 30 min per video (bc I'm slow) so... what if I could teach my @NotionHQ AI (Dwight) and @openclaw (Ezra) how to do post production? well, if you've watched any of my last three videos you've seen the results 😅 once I'm done recording, my agents... 1️⃣ pick up the videos 2️⃣ clip and trim them 3️⃣ send the audio to ElevenLabs to clean it 4️⃣ re-assemble the finished video 5️⃣ extract the transcript 6️⃣ use a skill to write a youtube description (that dynamically ingests from Notion) 7️⃣ upload the video to youtube 8️⃣ ping me - all in the time it takes to write one x post 😅 (actually, I was slower. Ezra pinged me 3 minutes ago) what a time to be alive! should I make a full video on this? (and ofc have Dwight & Ezra edit it?)

Bro, I was trying to run a test on Creatomate. I was wondering what was eating up my credits. About 100k plus every time, only for me to realize I didn't switch from Eleven Labs to an audio URL. Generating a video of over 5 hours plus and wanting to consume 1,100,000 credits 😭😭. Crazy thing was that creatomate received voiceover-1 and the GPT expanded narration text and called elevenlabs to generate new audi which it tends to expand and elaborate the video length. The generated audio was much longer than the original which also result on more credits on creatomate Luckily for me, I've been testing for about 2 weeks with short renders. This is the first time I'm testing with bulk data, and I'm yet to switch to production. Always do robust and ridiculous testing before you switch to production; it's a necessity.


@jimprosser ·
Had a coffee this week with a recent-former CCO of a public tech co who told me their execs created voice models of themselves with @ElevenLabs so they didn’t have to record their quarterly earnings prepared remarks anymore.
@thealexbanks ·
I love how Claude desktop now supports ElevenLabs Agents. You can build a voice AI support agent in minutes. Here's how it works: 1. Add the connector in Claude desktop ↳ Customise > Connectors > ElevenLabs Agents MCP ↳ Paste your API key and save
@eddiejaoude ·
I was experimenting with ElevenLabs this evening. The audio creation for my voice is very impressive, the lip-sync video not at all. My results: ✅ realistic text to speech creation ❌ video creation from photos using lip-sync DevRels are still needed 🥳 but we can use AI to be more productive 🚀! Is this audio actually me speaking or AI generated? 🤔
@wadefoster ·
Zapier x Meta Ray-Ban glasses. Here’s how it works: Our own Phillip McKee built a voice-first AI assistant using the Zapier SDK, Claude, and ElevenLabs. You say "Hey Zara" and it creates tasks, sends Slack messages, checks your calendar, and chains actions together. All hands-free through the glasses mic. The @Zapier SDK powers every action. It also learns preferences over time so repeat tasks get faster. This is what the SDK was built for - developers wiring AI into hardware we never anticipated. Post what you're building and tag me. I want to see it
@ericdjav ·
I built English With Lewis with $0 ad spend. 4,850 users in 5 months. Here's the exact model: 1. Find a creator with an audience (Lewis: 1.5M-follower British teacher) 2. Propose a revenue split. They bring users, you build the tech. 3. Clone their voice with ElevenLabs. 4. Ship a swipe-based learning app. 5. Both app stores in 30 days. Total cost: hosting and API calls. Barely $30. MRR: $1,160.
@stepanhlinka ·
I run $3M SaaS and our best performing ads are AI UGC videos using Sora right now. Here's the workflow: 1. Editor made ONE main video promoting our tool with b-rolls 2. I create 10-15s video using Sora for the hook 3. I take that video and clone the voice in elevenlabs 4. I create a voicover and use it over the main video from the editor 5. I can repeat this and create hundreds of hooks with different avatars, environments, scripts. Finally it's not just me in all the ads! 😃
@IgorWoorts ·
Remix your UGC b-rolls with AI voice overs: - Go to Elevenlabs - Use default Brian or Adam - Add the script - Match the b-rolls - Add background music Now you have 1 UGC -> but 4-5 different concepts. The real magic is to have the right script + editor.
@cubo_ai ·
Today our students had an incredible masterclass with @ElevenLabs 🎙️ We explored why voice is becoming the new interface for AI: more natural, faster, and more human. They learned real-time voice cloning, conversational agents, and the full stack (STT, LLMs, TTS). The most important part: this technology is no longer the future… it’s the present. The next step is clear 🔥 We’re going to create the first Salvadoran accent in ElevenLabs 🇸🇻

@HeyOz_AI ·
5 AI tools that saved me 40 hours of design work this week. Most marketers are still manually editing video layers. The top 1% are building "Creative Factories." Here is the exact stack (and prompts) I used to launch 50+ ads yesterday: 1. Oz (Bulk Creative Generation) This is the engine. I don't build one ad at a time anymore. I feed it my product assets, and it generates dozens of static and video variations instantly. • It handles the layout. • It handles the "ugly" aesthetic. • It handles the volume. Use this for: The heavy lifting. Replacing your junior designer. 2. ChatGPT (Infinite Hook Scripts) I don't write scripts from scratch. I iterate on winning formulas. Prompt: "I am selling [Product] to [Target Persona]. Write 10 TikTok-style ad hooks under 15 words. Structure them as: 3 'Negative' hooks (e.g., 'Stop doing X') 3 'Benefit' hooks (e.g., 'How to get Y instantly') 4 'Curiosity' hooks (e.g., 'The secret nobody tells you about Z') Keep the tone raw, conversational, and punchy." 3. NanoBanana (Background Assets) Stop using stock photos everyone has seen. NanoBanana lets you generate consistent, specific background assets or modify existing images with incredible precision. Prompt: "Generate a [Setting, e.g., messy desk with coffee stains] background. Style: iPhone photo, low fidelity, flash photography. Ensure it looks user-generated, not studio-perfect." 4. ElevenLabs (Voiceovers) The default AI voices sound robotic. You need "imperfect" delivery for ads. Settings to use: • Model: Eleven Multilingual v2[2] • Stability: 35% (Lower stability = more natural fluctuation) • Similarity Enhancement: 75% 5. CapCut (Assembly & Captions) The final assembly line. • Import your Oz visuals + ElevenLabs audio. • Use "Auto-Captions." • Workflow Hack: Search for "UGC Ad" templates in CapCut, strip the stock footage, and drag-and-drop your Oz-generated clips into the slots. Time saved: 3 hours per video.
Best ElevenLabs tweets
Xholic studies what works in your niche, drafts posts in your voice and schedules them for the hours your audience is online.
$0 today · Cancel anytime
Browse all tweet collectionsKeep exploring