Text-to-speech and voice cloning
ElevenLabs voice generation, cloned voices, expressive delivery controls, multilingual speech, and realism in narration or character audio.
62%
Best tweets about ElevenLabs
Discover the best tweets about ElevenLabs, covering AI voices, speech generation, dubbing, voice agents, APIs, releases, and creator workflows.
Specific ElevenLabs products, voice models, speech workflows, agents, API integrations, quality tests, and product updates.
Original Xholic analysis
Discussion centers on ElevenLabs text-to-speech, voice cloning, creative-production workflows, and voice-agent integrations. Posts also compare ElevenLabs with open-weight or competing voice tools on claimed cost, privacy, and quality, while some creators raise concerns about generated-voice authenticity and controls.
54% of posts
All-time engagement
36% of posts
Published in 90 days
Conversation map
ElevenLabs voice generation, cloned voices, expressive delivery controls, multilingual speech, and realism in narration or character audio.
62%
Open-weight or local TTS alternatives positioned against ElevenLabs on cost, privacy, voice cloning, deployment, and model performance.
36%
Real-time voice/chat agents, call automation, support and sales use cases, agent orchestration, and integrations with tools or business systems.
34%
ElevenLabs integrations in Claude, Codex, Lovable, Zapier, TanStack, agent frameworks, SDKs, starter projects, and API-driven builds.
32%
Comparisons of ElevenLabs models on quality, latency, normalization, preference rankings, expressive performance, and product controls such as speed or stability.
26%
ElevenLabs in automated pipelines for YouTube Shorts, podcasts, social video, ads, avatars, storytelling, captions, and creative production.
22%
Dubbing products and multilingual localization, as well as assistive communication and restoring an individual’s voice.
12%
Music V2, music finetunes and marketplace, sound-effect generation, Studio Agent, templates, and ElevenCreative workflows.
12%
Tone and stance
Performance benchmark
Posts with media make up 84% of this collection. Their median all-time score is 19.6, compared with 19.9 for text-only posts.
Format mix
Consensus and debate
Shared view
Posts describe ElevenLabs voiceover within Shorts, Instagram, and ad-production workflows alongside scripting, visuals, captions, avatars, and editing tools.
Shared view
Posts describe ElevenLabs in real-time voice-agent stacks and developer integrations, including TanStack support for voice I/O, Claude MCP audio generation, and ElevenLabs Agents connectors in Claude Desktop.
Shared view
Product-update posts cover Music Finetunes, Music V2, Dubbing V2, Studio Agent, templates, and ElevenAgents support for WhatsApp messages and attachments.
Open debate
Several posts position local or open-weight TTS alternatives against ElevenLabs on claimed price, privacy, cloning, or benchmark performance. Separately, Artificial Analysis ranked ElevenLabs Eleven v3 second in its listed top five preferred TTS models.
Open debate
Creators describe cloned-voice production workflows as efficient. In one podcast experiment, an editor identified the generated delivery and said it diminished the show's personality.
Open debate
One user reported realistic text-to-speech results but poor photo lip-sync results. Another asked why v3 voices lacked a speaking-speed setting on the Text-to-Speech platform.
Open debate
Posts reported the addition of a Stan Lee AI voice and criticized the arrangement in light of the legal history involving the licensing counterparty and concerns about posthumous likeness licensing.
What performs
The largest tracked outlier was an open-source Shorts pipeline claiming a $0.10 per-video cost and using ElevenLabs for voice. Another open-source end-to-end pipeline and a TanStack real-time integration post were also listed as outliers.
A post contrasting a claimed $330-per-month ElevenLabs narration plan with a locally run open-source model was the second-largest tracked outlier. Other competitive-pricing posts appear in the evidence but are not listed among the five deterministic outliers.
OPINION posts had a 33.728 median all-time score, above the 19.62 overall median. Cited opinion examples make rival quality and pricing claims about ElevenLabs.
Statistical standouts
Creator landscape
The five most represented creators account for 20% of the selected posts.
1. Carles Reina
@Carles_Reina
2 posts
2. ElevenLabs Developers
@ElevenLabsDevs
2 posts
3. Santiago
@svpino
2 posts
4. Alex Banks
@thealexbanks
2 posts
5. Vaibhav Sisinty
@VaibhavSisinty
2 posts
6. Mayank Vora
@aiwithmayank
1 post
ElevenLabs Developers posted about generating speech, sound effects, and music through the ElevenLabs MCP server in Claude, and about an AI-native examples repository.
Carles Reina's cited posts include a business update mentioning Expressive Agents mode and a monthly update listing Dubbing V2, Music V2, Studio Agent, templates, and expanded ElevenAgents WhatsApp capabilities.
Vaibhav Sisinty's cited posts compare xAI voice APIs and a voice-agent offering with ElevenLabs on claimed pricing, packaging, and performance.
Since the previous snapshot
Themes, sentiment, stance, and post format are classified per tweet. All counts, shares, medians, creator concentration, freshness, and performance comparisons are then calculated directly from the published snapshot.
Xholic's all-time score compares engagement while accounting for reach, post age, and creator consistency. It is used for relative comparisons within this collection.
This report analyzes the exact 50-post snapshot shown below. AI identifies editorial categories and drafts explanations; all statistics are calculated from the snapshot, and every narrative claim is checked against cited posts before publication.
Best ElevenLabs tweets
Ranked 01–50
@aiwithmayank ·
Say goodbye to video editors. This open-source tool turns a news headline into a published YouTube Short in one command. It's called YouTube Shorts Pipeline and it chains Claude, Gemini Imagen, ElevenLabs, and Whisper together into one pipeline. The cost breakdown is brutal: Claude script: $0.02 Gemini visuals: $0.03 ElevenLabs voice: $0.05 Total: $0.10 per video Type a topic. Get a live YouTube link. 3–5 minutes. Supports multiple languages. Custom voice IDs. Dry-run mode to preview before producing. Manual script override before rendering. Everything stored locally. No cloud dependency. 100% Open Source. MIT License.
@TheGeorgePu ·
Almost signed up for ElevenLabs to narrate my blog. $330/month. Then I tried running an open-source model on my own laptop. Qwen 3.5 14B. Sounds fine. 200 posts a month. Costs me electricity. I almost paid $4,000 a year to rent a model I can run myself. Most AI subscriptions right now are just a nice UI on top of something free.
@sentient_agency ·
Say goodbye to faceless YouTube channel agencies. This open-source repo does everything they charge $500/month for. news topic → script → AI visuals → voiceover → captions → uploaded to YouTube One command. 5 minutes. $0.10. What it uses: → Claude Sonnet for script writing (~$0.02/video) → Gemini Imagen 3 for 3 b-roll visuals (~$0.03/video) → ElevenLabs for voiceover (~$0.05/video) → Whisper for auto-captions (free) → YouTube Data API for upload (free) Videos upload as private by default so you review before going live. The anti-hallucination gate is the best part. Claude is shown live search snippets and instructed to only use names, stats, and facts it finds there. No AI making stuff up in your videos. 100% Open Source. MIT License.
@codyschneider ·
new live class for you how to build facebook ads AI agent you'll learn strategy - andromeda killed interest targeting — the creative + landing page *are* the targeting signal now, so the play is volume of distinct personas/angles rather than audience selection - loop is: publish many → let meta prospect → trim losers → promote winners into a winners campaign → remix winners back into testing - structure: testing campaign with multiple adsets, 5 ads per adset fighting for budget; winners graduate to a separate campaign, watched for fatigue - winners campaign can optimize on a deeper event than the testing campaign (signup for testing, demo/payment for winners) - agent entropy is the failure mode — left alone it regenerates the same creative; new dna comes from competitor ad libraries and viralio-scraped trending short-form tracking - gtm as the single tag + custom datalayer events for conversion actions - start upfunnel, work down as conversion volume allows — you can't train on payment events without enough volume - cap it at ~5 conversion events; more is noise - server-side capi for real-time lead scoring: enrich on submit (similarweb traffic via apify, headcount, crm, google review velocity, zip), agent scores 1–10, only 6+ fires back. solves lead quality without waiting out a 4-week sales cycle. build the scoring model off the shape of existing closed-won customers - delayed closed-won-based conversion feedback (hubspot deep funnel) works for some people but the learning cycle is too slow infrastructure - fb marketing api for **writes only** — creative upload and account management. never for analysis; hammering it risks a tos-based account ban, plus truncation/pagination/hallucination problems - analysis comes from the warehouse: airbyte → clickhouse, agent queries that. supabase/postgres won't hold up — 100m rows is minutes vs milliseconds - access via developer app + system user token (doesn't expire), not a public app submission - fb mcp exists but exposes only a handful of endpoints and breaks at scale; there's a new cli, untested - agent = software + thinking loop + live data stream; deploy to railway or similar creative - reddit scrape (xai api) for pain points and desired outcomes → source language for hooks and scripts - statics: kie ai → nano banana, fed competitor ad library refs + brand style guide - video: heygen public ugc avatars + elevenlabs voices (better than heygen native), seedance for more expensive/higher-fidelity once a format wins - cheap models for testing, expensive models for scaling proven winners numbers / positioning - financial advisor client: $100 → $53 cpl over four weeks - minimums: ~$30/day local, ~$100/day national, two-week read - diy timeline: ~2 weeks to something usable, 4–6 weeks to enterprise-ready — the pain is the pipeline/warehouse, not the agent - ad platforms' own automation has misaligned incentives (they want spend, you want cheap cpa)
@VaibhavSisinty ·
Did xAI just mass-murder the entire voice AI industry? 🤯 Grok just launched two voice APIs. Speech-to-Text and Text-to-Speech. Built on the same stack powering Tesla cars and Starlink support. And priced at 10x cheaper than ElevenLabs. Speech-to-Text: $0.10/hr batch. $0.20/hr streaming. Text-to-Speech: $4.20 per million characters. 25+ languages. Real-time streaming. Speaker diarization. Already outperforming ElevenLabs, Deepgram, and AssemblyAI on word error rate. TTS ships with expressive tags like [laugh], [sigh], <whisper>, <emphasis>. Voices that don't sound like robots reading a script. ElevenLabs spent years building a voice AI company. xAI built voice AI for cars and satellites.
@ElevenLabsDevs ·
11k people are using the ElevenLabs MCP server in Claude. It connects to your ElevenLabs account and lets you generate speech, sound effects, and music directly in Claude. Generate audio with ElevenLabs without friction.
@alex_verem ·
🚨Holy shit. Mistral just killed the ElevenLabs moat. > Voxtral TTS is open-weight, > Clones any voice from 3 seconds of audio, > Runs in 9 languages, > Beats ElevenLabs Flash v2.5 with a 68.4% human preference win rate. ElevenLabs built a moat on proprietary weights and API lock-in. Mistral just put the weights on Hugging Face. The model captures not just the voice but the person. Accents, inflections, intonations, vocal fillers the "ums" and "ahs" that make a voice sound human instead of synthetic. From 3 seconds of reference audio. Zero fine-tuning. Zero shot. The numbers: → 68.4% win rate against ElevenLabs Flash v2.5 in zero-shot multilingual voice cloning → Beats ElevenLabs Flash v2.5 on every one of the 9 supported languages → Matches ElevenLabs v3 on emotional expressiveness and quality → 70ms model latency same time-to-first-audio as Flash v2.5 at higher quality → 4B parameters. Runs on 3GB RAM. Smartphone. Laptop. Edge devices. → 9 languages: English, French, German, Spanish, Dutch, Portuguese, Italian, Hindi, Arabic → Cross-lingual voice cloning French voice prompt generating English speech works out of the box The strategic bet Mistral is making: enterprises don't want to rent a voice. They want to own it. Open weights means the model ships to your infrastructure. Mistral never sees your data. For healthcare, finance, and government Mistral's core customers that's not a nice-to-have. It's a compliance requirement. ElevenLabs built a $1B+ valuation on being the only serious option for production-grade voice AI. Today that changes. The weights are on Hugging Face. The API is $0.016 per 1,000 characters. Any developer can clone a voice, run it locally, and ship a voice agent without sending a single audio byte to a third party. One honest caveat: ElevenLabs shipped v3 after this evaluation. Voxtral matches v3 on quality metrics but the 68.4% win rate was head-to-head against Flash v2.5. The race isn't over. It just got a lot more competitive. Voice AI just got its open-source moment.
@VaibhavSisinty ·
xAI just launched a voice agent that costs $0.05 per minute. Everything included. And you can deploy it in 2 minutes. No STT vendor. No LLM vendor. No TTS vendor. One model. One price. ElevenLabs charges $50 to $120 per million characters for TTS alone. Grok charges $4.20. The voices laugh. Sigh. Whisper. Handle noise, accents, and interruptions naturally. Because it was not trained on studio audio. It was trained on Tesla cars and Starlink support calls. On τ-voice Bench it scores 67%. Gemini scores 44%. GPT Realtime scores 35%. Upload your docs. Connect your APIs. Clone a voice from 2 minutes of audio. Bring your own number. Deploy in production in the same session. ElevenLabs spent years building a voice AI company. xAI built voice AI for cars and satellites and just gave it to everyone. The fragmented voice stack is dead. @xAI @grok
@ElevenCreative ·
Introducing Music Finetunes in ElevenCreative. Generate individual vocals, instruments, or full tracks with stylistic consistency using a fine-tuned version of our Music model. Create your own Finetune or get started instantly with 11 Finetunes curated by ElevenLabs.
@DAIEvolutionHub ·
"$99 per month". That's what the ElevenLabs Pro plan costs to generate AI voice. This does it for free on your own machine. And it just outperformed closed models. It's called Fish Audio S2 and it already has 30,000 stars on GitHub. Trained with 10 million hours of audio in over 80 languages, it does what you've only seen until now in paid tools: → clones your voice with a 10 to 30 second sample → controls emotion with tags like [whisper], [excited] or [angry] at any point in the text → generates dialogues with multiple speakers in a single pass → starts the audio in ~100 ms, ready for real-time apps And the benchmarks confirm it: in Seed-TTS Eval it records the lowest error of all evaluated models, ahead of Qwen3-TTS, MiniMax and Seed-TTS. It works in Spanish, French, English and dozens more languages. It doesn't need phonemes or language-specific preprocessing. Open source, with the 4B parameter model available on HuggingFace. Repo here ⬇️
@bearlyai ·
London-based engineer Matt Cortland created an AI voice agent that called over 3,000 pubs on St. Patrick’s Day weekend to ask how much a pint of Guiness cost. The Irish government used to track this price up until 2011. Cortland stepped up with his “Guinndex”, which cost €200 to make and showed that the national price of €5.95. He used ElevenLabs to created “Rachel” and “getting the voice proved critical” (a Northen Irish accent inspired by a TV show).
@kimmonismus ·
Mistral just dropped Voxtral TTS an open-weight text-to-speech model with ultra-low latency, emotional expressiveness, and support for 9 languages. It outperformed ElevenLabs v2.5 Flash in zero-shot voice cloning tests judged by native speakers.
@oliviscusAI ·
ElevenLabs just lost its moat 🤯 Someone open-sourced one app that does the job of ElevenLabs AND WisprFlow, 100% local on your machine. → Clone any voice from 3 seconds of audio → 7 TTS engines in a single place → 23 languages: Arabic, Hindi, Japanese, and more → MCP server included, so Claude Code, Cursor, and Cline talk back in your cloned voice → Local LLM rewrites your text in-character before it hits TTS → Pedalboard effects baked in (reverb, pitch shift, chorus) Built on Tauri (Rust) instead of Electron. Runs on MLX for Apple Silicon, CUDA, ROCm, Intel Arc, DirectML, and plain CPU. ElevenLabs Creator runs $99/month. WisprFlow Pro is $15/month. Voicebox is $0. 23.4K stars on GitHub. MIT license.
@Carles_Reina ·
We've had the best quarter ever at @ElevenLabs ▪️Smashed even the wildest Revenue targets across B2B and Prosumer/B2C ▪️Enterprise overtook Prosumer/B2C accounting for 51% of all revenues now ▪️Eleven Agents is our fastest growing product ▪️Almost 50% first-sale is outbound ▪️Europe leading in outbound SQOs ▪️India leading in quota attainment despite our 20x quota/base requirement ▪️We hired 150 people out of 79,098 applicants = 0.0018% ▪️Added multiple new markets that will be announced soon ▪️Evolved our corporate structure to test a fully decentralised model giving countries more autonomy -> now expanding it further ▪️Launched Expressive Agents mode ▪️Ran the biggest ElevenLabs Summit in London on 11/Feb ▪️Our F1 team, Audi Revolut F1 Team, got the first points of the season What's in store for Q2? More experiments, more mistakes, deeper customer interactions and new products & models.
@shengkun_ye ·
You can now use ElevenLabs directly in Claude Code or Codex. Voiceovers for your demo videos, narration for tutorials, character voices for your game. Just one skill. No subscription required.
@TheTuringPost ·
.@MistralAI's new Voxtral TTS generates expressive, multilingual speech from just ~3 seconds of reference audio It solves one of the hardest problems in speech, separating what you say from how you sound ➡️ Voxtral factorizes speech into two parts: • semantic tokens → the content (words) • acoustic tokens → the voice (tone, prosody, style) and uses different models for different problems + Voxtral Codec that compresses speech into ultra-low bitrate tokens This gives: - High-quality voice cloning that works best with ~3–25 second - support for 9 languages - 68.4% win rate vs ElevenLabs Flash v2.5 in voice cloning Here's how it works:
@DivyanshT91162 ·
🤯 Someone just open-sourced what feels like an ElevenLabs competitor. A few months ago, voice cloning this good would've cost you a monthly subscription. Now it's sitting on GitHub. Meet VoxCPM2. A 2B parameter voice model trained on 2 million hours of audio that can clone a voice from as little as 3 seconds of reference audio. And the scary part? Most people won't be able to tell the difference. • Clone almost any voice with 3–10 seconds of audio • Supports 30 languages, including Spanish, without language tagging • Generates 48kHz studio-quality speech • Create entirely new voices from simple text descriptions • Real-time streaming with RTF 0.3 on an RTX 4090 • Compatible with ComfyUI, vLLM, and OpenAI-style APIs • Includes LoRA fine-tuning for training custom voices • Apache 2.0 licensed for commercial use The biggest flex isn't the quality. It's that everything runs locally. No subscriptions. No cloud dependency. No sending voice samples to a third party. Just download the model and start cloning. Open source is moving way faster than most people realize. Repo👇
@SahilPanhotra ·
Meet Mati Staniszewski → Grew up in Poland → Studied Mathematics and Computer Science → Started his career at Palantir → Worked on large scale software systems → Loved movies, games, and audiobooks → Kept noticing how terrible AI dubbing sounded → Wondered why AI voices still felt robotic → Believed AI could sound indistinguishable from humans → Teamed up with former Google ML engineer Piotr Dąbkowski → Founded ElevenLabs in 2022 → Set out to build the world's most realistic AI voice platform → Launched the first public beta in January 2023 → The internet was blown away → People started generating incredibly realistic voices within days → The product spread rapidly across creators, developers, and businesses → Then came the backlash → Users began cloning celebrity voices without permission → Deepfake clips started circulating online → The company was criticized for making voice cloning too accessible → Many believed the technology was too dangerous → Some thought ElevenLabs had launched too early → Instead of shutting the product down → Mati focused on making it safer → Added voice verification → Built AI speech detection tools → Introduced stronger moderation and security controls → Kept improving the models while earning enterprise trust → Expanded beyond text to speech → Built AI dubbing → Speech-to-text → Conversational AI → Voice agents → Developer APIs → Thousands of companies integrated ElevenLabs into their products → Reached about $100M ARR in less than 2 years → Crossed $200M ARR less than a year later → Raised a $500M Series D in 2026 → Reached an $11B valuation → Today ElevenLabs is one of the fastest growing AI companies in the world Imagine if Mati had backed down after the deepfake controversy ElevenLabs might never have become the voice behind the AI revolution
@rahulbhadoriiya ·
So I've been working on a new Instagram account for the last few months, and the results have been solid 200K+ views collectively on all the reels, 2–5K profile views, and a bunch of other wins here and there. The fun part? Most of it was AI-generated content. The delivery and execution were AI, but the actual idea, what to post, what to talk about was still human. The process was simple: 1. I'd find something interesting I want to talk about. 2. Send it to my OpenAI agent. 3. It writes a script : I've trained it by studying a lot of different AI content creators. 4. I edit the script if something feels off. 5. I've been doing media for six years now, content creation, writing ads, all of that so I understand how it works. 6. Then I generate audio through ElevenLabs. I've cloned my voice, so it sounds exactly like me. If some words don't land right, I edit those out. 7. Then I generate an avatar through HeyGen. So till here, everything is AI, script, audio, video. The editing is human, but AI still helps there too. I've built an agent that uses the X API to pull B-rolls and relevant clips from Twitter, then forwards everything to my editor via Telegram. He sees the new video drop in the group, edits it out, and that's the video done. That's one format. There's another format where I've taught my design AI agent how to create carousels, this is for content I want to post immediately, when I don't want to wait half a day for an edit. That works really well for quick takes. I've also been experimenting with a couple of new formats, and I'll keep doing that over the next year, automating as much as possible, but keeping the soul intact. The first idea, what I actually want to say and share, stays human. The delivery, the strip, the production, AI handles the manual work.
@Harish_521 ·
I have posted 18 shorts in 10 days on youtube. experiment results in the image This is my full process Ask ChatGPT to write a 60 second script. Generate a single image with Image Gen. Split the image into 8 scenes. Generate voiceover in ElevenLabs. Drop everything into CapCut. Sync images with audio. Add blur + fade transitions. Export. One idea. 10 minutes. A complete YouTube Short. Now going to automate it using n8n
@shushant_l ·
I'm shocked most people still don't know how powerful Voice AI has become. Here's how Voice AI agents can listen, think, talk, and take actions automatically. --- 1. Voice AI combines speech recognition, LLM reasoning, voice generation, tools, and automation. --- 2. Speech-to-Text converts your voice into text that AI systems can understand. --- 3. Large Language Models like GPT, Gemini, and Claude power Voice AI reasoning. --- 4. Text-to-Speech tools turn AI responses into realistic human-like voices. --- 5. Tools like ElevenLabs, Deepgram, Vapi, and Retell AI help build voice agents. --- 6. Voice AI can automate customer support, sales calls, meetings, and content creation. --- 7. Modern AI voice agents can clone voices, remember context, and have real-time conversations. --- 8. Developers can combine STT, LLMs, TTS, APIs, and workflows to create custom agents. --- 9. Voice AI is evolving towards autonomous agents, multilingual support, and emotional intelligence. --- 10. The future of AI interaction is moving from typing prompts to natural voice conversations. --- To learn more, check the infographic.
@heynavtoor ·
ElevenLabs charges $22 a month for the Creator plan. $99 a month for Pro. $299 a month for Scale. That is $264, $1,188, or $3,588 a year to make AI voices talk. Every character of text-to-speech burns 1 credit. Every minute of dubbing burns 2,000 to 10,000 credits depending on the tier. Free plan gets 10,000 credits total. Once they are gone, the audio stops. Here is what one professional-grade AI voiceover run costs in 2026. A 10-minute video dubbed through ElevenLabs Studio: around 100,000 credits per pass. That is more than a full month on the Creator plan. ElevenLabs Pro for a working podcast studio: $1,188 a year and you still watch the credit meter every session. A freelance voice actor on Fiverr: usually $50 to $200 for a short script. Now meet F5-TTS. A free, open-source text-to-speech model that runs on the GPU sitting in your desk. Paste a short reference clip of any voice, up to about 15 seconds. Type the script. Hit generate. The model produces a full clip in that voice. No cloud call. No credit meter. No account. 15,100 stars on GitHub. 2,194 forks. MIT license. Built by Yushen Chen, Zhikang Niu, Ziyang Ma, and the X-LANCE lab at Shanghai Jiao Tong University. First release October 2024. Peer-reviewed paper on arXiv (2410.06885). Version 1.1.22 shipped July 23, 2026. Here is what F5-TTS does: Zero-shot voice clone from a 5 to 15 second reference audio file English and Chinese out of the box, with community fine-tunes for Japanese, French, German, Italian, Spanish, Russian, Hindi, and Arabic Runs entirely offline on a consumer GPU Diffusion transformer with flow matching, trained on 100,000 hours of public multilingual audio Sway sampling for cleaner audio at inference time Fine-tune on your own dataset with LoRA Batch generate hundreds of clips locally Ships with a Gradio web UI for zero-code use No watermark, no telemetry, no API key Here is what F5-TTS costs to install: Zero. Forever. The model is about 500 million parameters. It uses 2 to 4 GB of VRAM at FP16. RTX 3060 works. Mac M2 works. Even an older 4 GB card will run inference. ElevenLabs Creator: $264 a year. F5-TTS: $0. ElevenLabs Pro: $1,188 a year. F5-TTS: $0. ElevenLabs Scale: $3,588 a year. F5-TTS: $0. Freelance voice actor: $50 to $200 per script. F5-TTS: this afternoon. Your voice. Your GPU. No credits. 100% Open Source. (Link in the comments)
@HedgieMarkets ·
🦔ElevenLabs, valued at $11 billion, partnered with Stan Lee Universe to release an AI clone of the late Marvel co-creator's voice and likeness. Lee's cloned voice will narrate public domain books in the Eleven Reader app starting with Treasure Island in June, and goes up for commercial licensing through the Iconic Marketplace. Stan Lee Universe is a joint venture involving POW! Entertainment, the same company Lee sued for $1 billion in 2018 alleging they took advantage of his macular degeneration and grief after his wife's death to broker an unauthorized likeness deal. Lee died later that year at 95. My Take I cannot read this announcement without feeling sick about it. Stan Lee spent the final months of his life filing lawsuits against business partners who exploited his blindness, his grief, and in one case literally drew his blood to stamp into comic books sold as collectibles in Las Vegas. Now the same company he sued is the legal counterparty greenlighting an AI clone of his voice. I do not see how anyone at ElevenLabs or Stan Lee Universe looked at that paperwork and thought this was a clean deal. The Iconic Marketplace business needs dead celebrities to scale. Living celebrities can refuse a campaign, fire their agent, or sue over a likeness deal they did not authorize. Dead ones cannot, and their estates have every reason to keep the licensing income coming. Judy Garland and John Wayne already sit on the same roster as Lee, and ElevenLabs at $11 billion needs recurring revenue to justify the valuation. I think the technology keeps improving and the catalog keeps growing regardless of how people react to any single announcement. The only thing slowing this market is public disgust, and based on the response to this one, I am not sure that is enough. Hedgie🤗
@Daviowhite ·
Working on a new project at work called “Learn” for our investors. I’m currently testing early concepts with Flora AI, I generated the avatar with Gemini and the voiceover with ElevenLabs. Then I hit a blocker in Flora, it doesn’t have an audio node. My goal is to automate this so I can create rapidly. Does anyone know how this can be achieved using a free tool? I’m trying to figure out the right toolkit for this workflow before committing to paid options.
@ArtificialAnlys ·
Inworld, ElevenLabs, and MiniMax continue to lead our Text to Speech leaderboard for most preferred models Recent checkpoints from each of the labs continue to push the frontier of TTS quality, with 4 out of the top 5 models being released this year. Leading TTS models are increasingly realistic, particularly on relatively straightforward text, with preference differences increasingly coming down to affinity for different voices. Latest results also reflect stronger bot vote filtering, confirmed via triangulation against third-party evaluators. We've also added rank ranges based on each model's 95% confidence interval, showing where a model could land based on its Elo score range. Key results: ➤ Most preferred: Current top 5 per our TTS leaderboard: 1. Inworld TTS 1.5 Max (Elo of 1,238); 2. ElevenLabs Eleven v3 (1,197); 3. Inworld TTS 1 Max (1,183); 4. Inworld TTS 1.5 Mini (1,182); 5. MiniMax Speech 2.8 HD (1,175) ➤ Price: Kokoro 82M v1.0 (Replicate) leads at $0.65 per 1M characters, followed by Inworld TTS 1 and 1.5 Mini at $5, and AsyncFlow V2 at $8.33 ➤ Speed: WaveNet leads for batch generation at 419 characters processed per second, followed by Kokoro 82M v1.0 (Replicate) at 235, and Inworld TTS 1.5 Mini at 214 See below for further detail ⬇️
@threepointone ·
landed a couple of new examples in the agents repo - structured inputs: how can your agent "ask" you questions that you answer as multiple choice/yes or no/ free form text/number slider/ whatever. tldr - use client side tools! - push notifications: how can your agent buzz you with a web push notification, even if you've closed your tab or whatever. - this in addition to the new elevenlabs starter earlier today. npm i agents
@noahkagan ·
"At first I thought you were sick. Not yourself. Low mood. Then I was convinced you were reading a script. By the end I'd already twigged before you even said it was AI." "Your personality is the best thing about your show. This just zapped it." Here's the backstory: I've got a second daughter now. Time is basically zero. So I thought - what if I could automate my podcast production? I took an article I wrote about everything we're doing with AI at AppSumo, had Claude turn it into a podcast script, uploaded 2 hours of my old episodes to ElevenLabs, and hit generate. It sounded like me. Nasally Jewish voice and everything. I was about to just ship it. Then I sent it to Jason, my editor, to see if he'd notice. He noticed in 30 seconds and gave me the feedback from the start of this. And then he said something that stuck with me: "AI in creative roles is a race to the bottom sometimes. It's great for organizing and admin. Just no good at being a person or a creative one at that. Something a lot of your peers in the tech world might need reminding of before we spin into a black hole of bland homogenized guff." And later that day a guy named David walks up to me at a coffee shop and says "I love your podcast." We all want to do less work and get great results. But some things require the work. That IS the value. AI should be like spellcheck, it doesn't replace you - it gives you superpowers on top of what you already bring.
@mark_k ·
i almost missed this one: elevenlabs released Music V2. It sounds a lot better than the first version did. 🎶 Have a listen:
@svpino ·
Here is a new open-weight audio model you can integrate with your app. I'm a huge sucker for open models that you can host to keep your data secure and away from Big AI. The model is Fish Audio S2. You'll find the weights, the code to fine-tune it, and a streaming inference engine in Hugging Face. (Link below). Under the hood, it's two models: a 4B one to handle timing and meaning, and a 400M one to fill in the acoustic details. The model is really fast! If you host it on an H200, it will start returning audio in about 100ms. If you don't want to host it, Fish Audio offers a hosted version: the S2.1 Pro. This model is even faster, with around 90ms latency, and it supports 83 languages. The S2.1 Pro is also ~83% more affordable than ElevenLabs: you pay around $15 for around 12 hours of transcribed audio, which is incredible. HuggingFace links below.
@PrajwalTomar_ ·
THIS IS CRAZY. I connected ElevenLabs to Lovable and built an AI storytelling app in UNDER 20 minutes. Enter a kid's name, pick a theme. The app writes a personalized bedtime story, narrates it with ElevenLabs, and plays ambient sounds that match the story world. → Text to Speech narration with emotion → Sound effects generated on the fly → Download the full story as audio Lovable connectors have made my life so easy.
@Shruti_0810 ·
Another billion-dollar AI category is becoming open source. Voice AI is following the same path as image generation and coding assistants. Pipecat is an open-source framework for building real-time voice AI agents. It handles the entire voice stack: → Speech-to-Text → LLM reasoning → Text-to-Speech → Real-time streaming → Conversation orchestration Everything you need to build: • AI phone agents • Customer support bots • Voice copilots • AI companions • Interactive assistants • Multimodal experiences What makes it different: • Swap any AI provider without rewriting your app • Plug in OpenAI, Claude, Gemini, Groq, Ollama, and more • Choose Deepgram, Whisper, AssemblyAI, or other STT providers • Use ElevenLabs, OpenAI, Cartesia, Azure, and more for TTS • WebRTC + WebSockets built in for ultra-low latency The biggest shift isn't a new model. It's that the entire infrastructure layer is becoming a commodity. The winners won't be the companies selling voice APIs. They'll be the ones building products users actually want. GitHub: https://t.co/fkj7JO2ul7
@RoundtableSpace ·
Someone open sourced the entire ElevenLabs + Descript stack in a single local tool. Voice-Pro: zero-shot voice cloning, Whisper transcription, vocal isolation & dubbing in 100+ languages. ElevenLabs charges $22/month for this.
@michuk ·
One thing that stood out during yesterday's product demos at the @ElevenLabs Summit in Warsaw was how much focus there was on applications, not models. That makes perfect sense. AI models — even ones as impressive as ElevenLabs' voice model — will eventually become commoditized. In fact, ElevenLabs itself showed a compact voice model running on a mobile device without internet access. Soon there will be open source models that do that for free for everyone who builds. And if that's where the technology is already heading, relying solely on token consumption is probably not a great long-term business. So what do model companies do? They move up the stack. The same way OpenAI launched ChatGPT, or GPT Health and Anthropic launched Claude Code, Desktop and is working on cybersec, ElevenLabs is increasingly building products: – creator marketplaces – tourism agents (demoed yesterday) – customer service agents (demoed at a prevous summit) – and likely many more vertical applications These products become the new moat, not (just) the model. And while this is a rational strategy for model companies, it's also a warning sign for startups building on top of them. If your entire business is "Claude/OpenAI/ElevenLabs + a prompt," you're vulnerable. Even if your UX is dope. The AI agent startups that survive will need moats of their own: – deep integrations into customer workflows – proprietary data – strong communities – distribution advantages – regulatory or industry expertise Otherwise they'll wake up one morning and discover their biggest supplier has become their biggest competitor. Check out the personalized tourist guide demo from the summit. #AI #ElevenLabs #Startups #VentureCapital
@etnshow ·
ElevenLabs is turning every chat agent into a voice agent. Their new product Voice Engine: • Wrap your existing chat agent, it becomes a voice agent • One command. One skill. It builds itself. • Add a data agent to a Zoom call • Turn a gov chat agent into a phone line anyone can ring No rebuild needed. Just wrap what you already have.
@Carles_Reina ·
May has been a crazy month at @ElevenLabs ▪️ We announced that we had crossed $500M in ARR ▪️ Entered a strategic partnership with the Greek Government 🇬🇷 to bring AI to the public services, tourism, and preservation of Greek heritage ▪️ Launched Dubbing V2, the global SOTA model for dubbing ▪️ Launched Music V2, trained on fully licensed data and built for Musicians, Developers and Brands ▪️ Partnered with Stan Lee Univers to make his voice available on ElevenLabs Iconic Marketplace and the Reader app ▪️ We launched in Australia & New Zeland 🇳🇿 🇦🇺 ▪️ ElevenAgents can now handle images, PDFs, audio messages, contacts, and location pins on WhatsApp ▪️ Added Studio Agent in ElevenCreative, our AI co-editor to streamline the creative process ▪️ Brought Templates to ElevenCreative, our ready-made creative workflows
@ElevenLabsDevs ·
Introducing ElevenLabs Examples 2.0, a modern approach to the examples repository. The examples repo is one of the highest touch points for any developer-facing product, but maintaining them is a tedious task when you have a fast moving product. We rebuilt ours to be prompt-driven, AI-native, and autonomously maintained.
@svpino ·
I can't wait for the day we don't need to hear "your call is important to us" ever again. Literally nobody wants to talk to a bot on the phone, but we all hate waiting in line. Here is the thing about voice: We've been talking to each other for thousands of years, but we've only been typing for maybe 50. Voice is the most natural interface humans have. But voice AI still sounds robotic and awkward, with zero emotional awareness. If we can unlock voice AI at scale and address all of these issues, we'll get access to the most natural and powerful interface we've ever had. Here is the good news: I'm working with ElevenLabs, and they released a platform that lets you build and deploy voice and chat agents that sound really good. These voice agents are: 1. Real-time 2. Multilingual 3. With emotional inflection They work so much better. They don't make people hang up the phone immediately! What really matters here is that these agents can handle thousands of simultaneous conversations without making you wait through hold music. They integrate with existing systems, resolve issues, qualify leads, and escalate with full context when a human needs to step in. ElevenLabs got this right.
@WesRoth ·
ElevenLabs has expanded its creator ecosystem by launching a "Music Marketplace" within its ElevenCreative platform, allowing users to publish, distribute, and actively monetize their AI-generated tracks. To make the marketplace viable for enterprise use, every track is automatically available under three pre-defined commercial licensing tiers: Social Media, Paid Marketing, and Offline. This eliminates the need for custom legal negotiations.
@thealexbanks ·
I love how Claude desktop now supports ElevenLabs Agents. You can build a voice AI support agent in minutes. Here's how it works: 1. Add the connector in Claude desktop ↳ Customise > Connectors > ElevenLabs Agents MCP ↳ Paste your API key and save
@eddiejaoude ·
I was experimenting with ElevenLabs this evening. The audio creation for my voice is very impressive, the lip-sync video not at all. My results: ✅ realistic text to speech creation ❌ video creation from photos using lip-sync DevRels are still needed 🥳 but we can use AI to be more productive 🚀! Is this audio actually me speaking or AI generated? 🤔
@wadefoster ·
Zapier x Meta Ray-Ban glasses. Here’s how it works: Our own Phillip McKee built a voice-first AI assistant using the Zapier SDK, Claude, and ElevenLabs. You say "Hey Zara" and it creates tasks, sends Slack messages, checks your calendar, and chains actions together. All hands-free through the glasses mic. The @Zapier SDK powers every action. It also learns preferences over time so repeat tasks get faster. This is what the SDK was built for - developers wiring AI into hardware we never anticipated. Post what you're building and tag me. I want to see it
@cubo_ai ·
Today our students had an incredible masterclass with @ElevenLabs 🎙️ We explored why voice is becoming the new interface for AI: more natural, faster, and more human. They learned real-time voice cloning, conversational agents, and the full stack (STT, LLMs, TTS). The most important part: this technology is no longer the future… it’s the present. The next step is clear 🔥 We’re going to create the first Salvadoran accent in ElevenLabs 🇸🇻
@thealexbanks ·
Brad Smith hasn't been able to speak for years. Now Neuralink lets him type with his brain and ElevenLabs lets him speak in his own voice. He just gave an interview on national TV. Brad was diagnosed with ALS at 37. He's now 45. The disease destroyed the motor neurons connecting his brain to his muscles. Brad is quadriplegic and nonverbal, but his mind is completely sharp. Before Neuralink, he relied on eye-gaze tracking to communicate. It was slow, exhausting, and only worked in dark environments. Now, two technologies have converged to change everything: Neuralink handles the input: → Brad controls a cursor by imagining tongue movements → Clicks are performed by imagining jaw clenches → Pure thought with no physical motion required → He can type, edit videos, and even play Mario Kart ElevenLabs handles the output: → AI-cloned his pre-ALS voice from old recordings → Trained on his speech patterns, tone, and inflections → Brad types what he wants to say and the AI speaks it in his voice Neuralink gave Brad the ability to communicate with his brain. ElevenLabs gave him back the voice that makes him who he is. Through the ElevenLabs Impact Program, the company has committed $1 billion in free voice restoration technology to reach 1 million people living with permanent voice loss. They've already supported ~7,000 individuals across 49 countries. I know I spend a lot of time debating AI's risks and limitations. But when a father who hasn't spoken in years can speak to his kids again, in his own voice, controlled entirely by thought, is the other side of the conversation that I think deserves more attention. Video source: NewsNation on YouTube
@HeyOz_AI ·
5 AI tools that saved me 40 hours of design work this week. Most marketers are still manually editing video layers. The top 1% are building "Creative Factories." Here is the exact stack (and prompts) I used to launch 50+ ads yesterday: 1. Oz (Bulk Creative Generation) This is the engine. I don't build one ad at a time anymore. I feed it my product assets, and it generates dozens of static and video variations instantly. • It handles the layout. • It handles the "ugly" aesthetic. • It handles the volume. Use this for: The heavy lifting. Replacing your junior designer. 2. ChatGPT (Infinite Hook Scripts) I don't write scripts from scratch. I iterate on winning formulas. Prompt: "I am selling [Product] to [Target Persona]. Write 10 TikTok-style ad hooks under 15 words. Structure them as: 3 'Negative' hooks (e.g., 'Stop doing X') 3 'Benefit' hooks (e.g., 'How to get Y instantly') 4 'Curiosity' hooks (e.g., 'The secret nobody tells you about Z') Keep the tone raw, conversational, and punchy." 3. NanoBanana (Background Assets) Stop using stock photos everyone has seen. NanoBanana lets you generate consistent, specific background assets or modify existing images with incredible precision. Prompt: "Generate a [Setting, e.g., messy desk with coffee stains] background. Style: iPhone photo, low fidelity, flash photography. Ensure it looks user-generated, not studio-perfect." 4. ElevenLabs (Voiceovers) The default AI voices sound robotic. You need "imperfect" delivery for ads. Settings to use: • Model: Eleven Multilingual v2[2] • Stability: 35% (Lower stability = more natural fluctuation) • Similarity Enhancement: 75% 5. CapCut (Assembly & Captions) The final assembly line. • Import your Oz visuals + ElevenLabs audio. • Use "Auto-Captions." • Workflow Hack: Search for "UGC Ad" templates in CapCut, strip the stock footage, and drag-and-drop your Oz-generated clips into the slots. Time saved: 3 hours per video.
Best Tweets by Topic