Pricing and model substitution
Token prices, free access periods, premium-versus-cheaper model comparisons, and cost-driven replacement of frontier models.
46%
Best tweets about OpenRouter
Explore the best tweets about OpenRouter, covering model routing, API pricing, provider reliability, benchmarks, and developer integrations.
Specific OpenRouter models, routing behavior, API usage, pricing, provider performance, outages, and engineering lessons.
Original Xholic analysis
Discussion in this 50-post dataset is led by pricing and model-substitution posts (46% of posts), followed by Chinese/open-weight model adoption (36%) and workflow-fit evaluation (30%). The evidence also includes firsthand reports of routing configuration issues, rate limits, model-specific behavior, and endpoint retirement, supporting coverage of evaluation, logging, fallback design, and migration planning.
54% of posts
All-time engagement
100% of posts
Published in 90 days
Conversation map
Token prices, free access periods, premium-versus-cheaper model comparisons, and cost-driven replacement of frontier models.
46%
Growth in usage of Chinese or open-weight models on OpenRouter, including token-share trends, model leaderboards, and ecosystem competition.
36%
Hands-on assessments of coding, documentation, tool use, language behavior, context handling, search, and model-specific tuning needs.
30%
OpenRouter-powered apps, SDK integrations, prompt replay, model comparisons, agent configuration, and practical production workflows.
24%
Automatic task routing, provider selection, fallback behavior, local operators, agent stacks, and multi-model deliberation workflows.
24%
Rate limits, outages, latency, provider routing status, uptime, API quality discrepancies, and direct-provider versus aggregator trade-offs.
20%
New OpenRouter model availability, anonymous previews, context windows, multimodal or agentic capabilities, architecture, and benchmark claims.
18%
Model deprecations, preview shutdowns, migrations, and the maintenance burden of revalidating model choices.
4%
Tone and stance
Performance benchmark
Posts with media make up 72% of this collection. Their median all-time score is 8.81, compared with 9.75 for text-only posts.
Format mix
Consensus and debate
Shared view
Pricing and model substitution is the largest supplied theme, appearing in 23 posts (46%). Several posts compare lower-cost models with premium alternatives; one comparison says cheaper models can be close enough for many tasks while still underperforming on harder work.
Shared view
Chinese and open-weight model adoption appears in 18 posts (36%). Posts report rising token shares and usage rankings on OpenRouter, while attributing interest to price, context capacity, and task-sufficient performance. These are platform-specific usage claims and author interpretations, not a measure of the entire model market.
Shared view
Firsthand posts assess models in concrete workflows, including documentation generation, code review and refactoring, tool use, language behavior, speed, and long-context handling. The accounts show that results can differ by model and workflow.
Shared view
Routing is presented both as a way to dispatch work among models and as an area where configuration can create unexpected usage. One user describes inspecting logs to identify unintended OpenRouter-routed requests; another describes updating skills and hard-coded routing logic when changing model providers.
Open debate
One pricing comparison argues that lower-cost models can be sufficiently capable for many tasks, while a separate firsthand account reports slower performance, weaker outputs, Chinese-language replies, and weaker tool use for some alternatives relative to Opus and GPT models in that workflow.
Open debate
Posts present OpenRouter as a source of provider choice and outage resilience. Conversely, one setup report says OpenRouter rate limits prompted a switch to Kimi’s direct API, and another post alleges added overhead when accessing DeepSeek through OpenRouter. These are user reports and claims rather than measured platform-wide comparisons.
Open debate
Posts cite OpenRouter token rankings as adoption signals, while another user notes that its organization moved some traffic to direct provider accounts after omnibus-key rate limits. A separate post characterizes the leaderboard as a price-sensitive aggregator slice rather than total market usage.
Open debate
A Qwen launch post promotes a free 1M-context preview, while another launch post explicitly warns that a fresh model may have rough edges and recommends testing before production. Separately, a provider announced the retirement of a preview endpoint and advised active users to migrate.
What performs
The G0DM0D3 post was the largest supplied outlier, with an all-time score of 4,023.56 versus the dataset median of 8.99. It describes a browser-opened HTML interface using OpenRouter to access 51 models.
The Qwen free-preview announcement (all-time score 556.06) and the premium-model replacement comparison (179.95) were the next two supplied engagement outliers after G0DM0D3. Both posts make specific claims about access, token pricing, and model trade-offs.
A post describing a personal comparison after an API ban had an all-time score of 73.52, making it a supplied engagement outlier (8.18 times the dataset median). The author reports a change from $150 to $15 monthly spend and near-comparable results for that workflow.
Statistical standouts
Creator landscape
The five most represented creators account for 20% of the selected posts.
1. AshutoshShrivastava
@ai_for_success
2 posts
2. Allie K. Miller
@alliekmiller
2 posts
3. BridgeMind
@bridgemindai
2 posts
4. Layton Gott
@Layton_Gott
2 posts
5. OpenRouter
@OpenRouter
2 posts
6. Pukerainbow 🤮🌈
@pukerrainbrow
2 posts
The evidence includes an official prompt replay feature for auditing and model comparison, a local operator dispatching tasks through OpenRouter, and a firsthand account of multi-model routing changes. These posts offer specific workflow examples rather than generic model comparisons.
OpenRouter announced prompt viewing, replay, and remixing for auditing, iteration, and model comparison. Another official post points users to uptime charts and warns against relying on single-vendor inference.
Launch-oriented posts cover a free Qwen preview and a multi-agent Grok model, while other posts describe model deprecations and preview-endpoint retirement. Together, these support coverage that pairs new-model access with lifecycle and migration considerations.
Since the previous snapshot
Themes, sentiment, stance, and post format are classified per tweet. All counts, shares, medians, creator concentration, freshness, and performance comparisons are then calculated directly from the published snapshot.
Xholic's all-time score compares engagement while accounting for reach, post age, and creator consistency. It is used for relative comparisons within this collection.
This report analyzes the exact 50-post snapshot shown below. AI identifies editorial categories and drafts explanations; all statistics are calculated from the snapshot, and every narrative claim is checked against cited posts before publication.
Best OpenRouter tweets
Ranked 01–50
@heynavtoor ·
🚨 The most disrespectful thing someone has ever done to the AI industry: Fit 51 models into one HTML file with zero dependencies and gave it away for free. Read that again. One HTML file. Not an app. Not a platform. A file. It's called G0DM0D3. Drag it into your browser. You now have Claude, GPT-5, Gemini, Grok, Mistral, LLaMA, DeepSeek, Qwen, and 43 more models. Right there. No install. No React. No Next.js. No node_modules. No build step. No backend. No server. Nothing. Here's what's packed into this single file: → 51 AI models through one chat interface via OpenRouter → Race multiple models against each other in parallel → See which model gives the best answer side by side → Auto-tuning engine that adjusts temperature and sampling per context → Your API key never leaves your browser. Zero data collection. → Works on desktop and mobile → Self-host by opening a file. That's it. That's the deployment. Here's the wildest part: Every AI company on Earth built a team, raised millions, deployed servers, hired engineers, and built apps to give you access to their model behind a $20/month paywall. This person built one HTML file that accesses all of them at once. No team. No funding. No servers. No paywall. One file. ChatGPT Plus: $20/month for one model. Claude Pro: $20/month for one model. Gemini Advanced: $20/month for one model. This: 51 models. One file. Price of an API key. Open Source. AGPL-3.0 License.
@bridgemindai ·
Qwen 3.6 Plus Preview just dropped on OpenRouter. 1,000,000 token context window. $0 input. $0 output. It's Free. A million tokens of context for free. This is insane. Claude Opus 4.6 charges $5/$25 per million tokens for 200K context. GPT 5.4 charges for 1M context. Qwen just gave it away. All free. Open source is not slowing down. It's accelerating.
@DeRonin_ ·
Replaced these PREMIUM models with 100%+ cheaper alternatives 1. Claude Opus 4.6 → Kimi K2.5 Benchmark gap: ~11% Price difference: ~9x cheaper 2. GPT-5.4 → Qwen3.5-397B Benchmark gap: ~21% Price difference: ~6–7x cheaper 3. Claude Sonnet 4.6 → GLM-5 Benchmark gap: ~2% Price difference: ~3.8x cheaper on input and ~5.9x cheaper on output 4. Gemini 3.1 Pro → GLM-5 Benchmark gap: ~12% Price difference: ~2.5x cheaper on input and ~4.7x cheaper on output 5. GPT-5.4 Pro → Kimi K2.5 + DeepSeek R1 Benchmark gap: GPT-5.4 Pro still has an advantage on complex tasks, but the price difference is very attractive Price difference: ~95–97% cheaper Conclusion: Based on OpenRouter pricing and Artificial Analysis benchmark scores, these lower-cost models are often 50%+ cheaper while remaining surprisingly close in capability for many tasks Some of them may underperform the premium models on harder or more complex tasks, but the price difference is significant Considering the 50%+ savings in most cases, this is a strong option if you don't have a large budget
@kaif9998 ·
Anthropic banned my Julius (OpenClaw) from their API 💀 So I tested every single alternative out there. Was burning $150/month on credits. Tried Kimi K2.5… and now I finally see why it’s the most-used model for OpenClaw on OpenRouter. > 90%+ of Claude Opus performance > Built-in web search (no extra API calls) > 8x cheaper than Claude I went from $150/mo → $15/mo. Almost the same coding output and reasoning quality. The math just isn’t even close anymore. Who else made the switch? What’s your new stack? 👇
@HedgieMarkets ·
🦔U.S. companies are routing more AI workloads to Chinese models. CNBC reported that Chinese model usage on OpenRouter sat above 30% of U.S. company tokens every week since February, up from an 11% average the prior year and just 4.5% in H1 2025. https://t.co/IJWtvXQaKh's GLM 5.2 saw token volume grow 27x in its first week on Vercel. One startup, Lindy, moved 100% of its traffic from Claude to DeepSeek and expects to save millions. My Take The pitch from U.S. labs has always been that frontier performance justifies frontier pricing. That logic holds until someone ships a model within a percentage point of Opus 4.8 on agentic benchmarks at a fifth of the cost, which is what GLM 5.2 apparently did. At 60 to 90% cheaper for most workloads, you do not need to match the best model. You just need to be good enough for the task you're routing to it. If U.S. labs keep raising token prices while Chinese open-weight models close the performance gap, the market self-selects toward Chinese infrastructure for everything except the hardest problems. Most companies making that switch are not thinking about geopolitics. They are thinking about their AWS bill. Hedgie🤗
@CodeByNZ ·
👀 Tencent Hunyuan HY3 is moving a lot faster than I expected. Just one week after launch, it's already the #1 most-called model on OpenRouter. What's even crazier is the growth. The preview version hit 3.6 trillion model calls two months ago. Now HY3 has surpassed 6 trillion model calls, nearly doubling that number. The open model race is getting more competitive every month, and Tencent is proving it's a serious contender. @TencentHunyuan #HY3 #Hunyuan #AI
Watch video
@heyshrutimishra ·
China just made the best time to try a frontier model completely free. Alibaba's Qwen3.6-Plus dropped today. 2 weeks free on OpenRouter. What I'm actually paying attention to: 1M token context… That's not a headline feature anymore, it's table stakes. What's different here: it doesn't just generate code. It manages the full loop planning, testing, iterating, until the output is production-ready. Multimodal built in. Documents, images, long-form video - one model, one pass. Real critique: fresh launch, so there'll be rough edges. Don't run it in production day one. But for free? Run your hardest agentic workflow against it this week. Test it here → https://t.co/qkWKdDFUKq
@FredaDuan ·
Token Tracker & Implications If third-party tracking is even roughly right, China may now be the world’s largest token economy. China’s daily AI token consumption has reportedly reached ~180 trillion per day as of Feb 2026 - up from just 100 billion at the start of 2024. That’s a ~1,800x increase in ~2 years. Video generation is a major driver that hasn’t really kicked in yet. Seedance 2.0 uses ~350K tokens to generate a single 10-second 1080p video. A typical animated project can run into the hundreds of millions of tokens. Video AI is quietly pushing token demand into a new regime. At the same time, OpenRouter data shows Chinese models consuming more weekly tokens than U.S. equivalents. But there are real open questions (if anyone has answers/ thoughts, pls DM): 1/ Are we comparing like-for-like models? “Cheap can be expensive.” Lower-quality models may require more retries, longer prompts, and additional iterations - inflating token counts. 2/ Does China’s 2C market structurally consume more tokens? Larger population + historically weaker search substitutes could naturally drive heavier AI usage. 3/ Will “good enough but cheaper” models ever really win? Implications on the SOTA models? If slightly inferior but meaningfully cheaper models are sufficient for most tasks, that has major implications for pricing power, infrastructure demand, and long-term model hierarchy. --- China’s Token Usage Daily token usage across different models: - $Byte Volcano Engine’s large-model daily average token calls have rapidly grown from 2 trillion at the end of 2024 to 63 trillion in January 2026. - $BABA Cloud’s external customer daily average token calls in 2025 are close to 5 trillion, with the 2026 target set at at least 15–20 trillion. Its internal business daily average token calls are planned to rise from 16–17 trillion to 100 trillion. From an industry-wide perspective, China’s overall daily average token consumption has increased from 100 billion at the beginning of 2024, surpassed 30 trillion in mid-2025, and by February 2026 the combined daily average across mainstream large models has reached the 180 trillion level. source: https://t.co/fv6wkBKAaG ByteDance’s Doubao: over 50 trillion tokens daily Impact from video gen: - At the single-video level, Seedance 2.0 consumes roughly 350,000 tokens to generate a 10-second, 1080p video. Under comparable quality settings, Keling requires more than 400,000 tokens, suggesting Seedance is somewhat more efficient in multi-frame synthesis. - In a typical animated project (720p, 15fps video generation), the project often requires hundreds of millions of tokens overall. --- US’ Token Usage $Google - “tokens processed across our surfaces” 1.3 quadrillion tokens per month (Google blog, Oct 9, 2025) - “across our surfaces,” explicitly referencing the prior 980T figure. https://t.co/qSdpb52rYu 1.3 quadrillion/month ≈ ~43 trillion/day (assuming 30 days/month). $MSFT - “tokens processed this quarter” (Azure AI / Foundry) $MSFT said 50T in one month (≈ ~1.7T/day for that month). source: https://t.co/fv6wkBKAaG --- OpenRouter OpenRouter is just a small subset of data being tracked, but here are the latest trends: Global AI model API aggregator OpenRouter data shows that during the week of the 9th–15th, Chinese models reached 4.12 trillion tokens in API calls, surpassing U.S. models for the first time, which recorded 2.94 trillion tokens in the same period. In the following week (16th–22nd), Chinese models’ weekly token usage climbed further to 5.16 trillion tokens, marking a 127% increase over three weeks, while U.S. model usage declined to 2.7 trillion tokens. https://t.co/mnVnGeCMc5 +++ Full analysis: https://t.co/vj8XX3ySd5
@bridgemindai ·
Grok 4.20 Multi-Agent just went live on OpenRouter. This model is built different. Low/medium reasoning: 4 agents in parallel. High/xhigh reasoning: 16 agents in parallel. One API call. Up to 16 agents coordinating, researching, and synthesizing in parallel. 2M context. $2/M input. $6/M output. This is the #1 model on BridgeBench for a reason. It's not one model thinking harder. It's 16 models thinking together. xAI just made multi-agent native at the model level. Everyone else is bolting it on after the fact.
@Rahul_J_Mathur ·
Last night I panicked when our fund’s agent Zen replied to me in Chinese. For a minute, I wasn’t sure WTF happened. Until recently, our agents have run exclusively on the Opus series of models (this is a $100k annualized bill for a 50 person team) Sometime in June, we began experimenting with a multi-model set-up using OpenRouter with an explicit goal of reducing our spend by ~40% without compromising on long term performance or security. Sharing my personal observations: (1) DeepSeek v4 pro replies in Chinese at least once per day (approx. once in ~30 tasks). There is an open ticket in GitHub specifying this very issue which remains unresolved. Furthermore, DeepSeek is noticeably slower for token heavy tasks (to a point where I need to do other work while waiting to check outputs) (2) GLM 5.2 is observably slower than Opus 4.8 & produces worse quality output GLM 5.2 doesn’t read the complete Skill file, fails to handle edge cases etc. However, at 1/8th the cost, this is an acceptable trade-off for me since the output quality is observably worse for me but NOT for the recipient of the output. (3) Skills built for Opus & GPT series models don’t translate well to the GLM series & DeepSeek model variants. Opus & GPT have far superior tool calling capability (my hypothesis) & more grounded to follow instructions (knowledge worker centric RL or inclusion in training data sets). I regret to inform the hype lords that with appropriate & directed efforts, it is possible to switch workflows between model providers⚠️ However, this is NOT seamless & does require one-time effort to update Skills, usage muscle memory & hard code certain routing logic.
@Parul_Gautam7 ·
I've been using Ring-2.6-1T on OpenRouter for my documentation pipeline this week. My usual flow: I dump a raw codebase or a long spec doc into the model and ask it to generate structured docs, flag inconsistencies, and draft a changelog. With other models, I was splitting this into 3 separate prompts because of context limits. Each split costs tokens. Each split loses some coherence between sections. With Ring-2.6-1T, I sent the whole thing in 1 prompt. 38,000 tokens. got back structured docs, a flagged inconsistencies list, and a changelog draft in a single response. Coherence was intact across all 3 outputs. No context bleed between sections. Still running it through more workflows, but the single-prompt full-doc processing alone has already cut my per-task token spend significantly. And then the cost showed $0.00, and I sat there for a second 👉 https://t.co/DZIXXDMiqK #BuildInPublic #OpenSource
@TheGeorgePu ·
DeepSeek went from under 1% of OpenRouter tokens to 17%. That happened in a single month. The top 4 models by usage are now all Chinese open models. Nobody decided this in a press release. Finance teams looked at the bill and rerouted the traffic. I found out I don't have to run my whole workflow on Claude and GPT. The cost difference is multiple orders of magnitude. I think more people are figuring that out too. The leaderboard everyone watches is one thing. The model people actually pay for is another.
@Layton_Gott ·
Last year Anthropic had 29% of all the AI usage on OpenRouter. Today it’s only 13%. You’re probably wondering why Claude’s usage got cut down over 50%… The answer is simple: Chinese open weight models. They went from basically nothing to over 60% of all the tokens running through OpenRouter. In about a year and a half. The top six most used models on the entire platform are now Chinese. Claude sits at number six. MiMo from Xiaomi. DeepSeek. MiniMax. GLM. Models a lot of people have never even heard of are now bigger than Claude by usage. Why? Price. These models run for as little as a tenth of what Claude costs. Some even less. And they’re now good enough that for most jobs you genuinely can’t tell the difference. So developers did the obvious thing. They followed the cheaper option. Anthropic is only 12% of the tokens, but it’s still pulling around 46% of the actual money on the platform. So the market split in two. One lane is cheap, open, good enough, and absolutely massive on volume. That’s the Chinese models. The other lane is expensive, premium, and where the real money still is. That’s Claude and GPT. No, Claude isn’t dying. It’s making more money than ever while slowly losing the volume war. But the direction is impossible to ignore. The cheap open models are taking the actual usage, and they’re getting better every single month. And the closed frontier models keep getting locked behind governments and price tags. This is exactly why I keep saying open source wins in the end. You can’t out price free. And you can’t shut off a model the whole world already downloaded.
@alliekmiller ·
Enterprises are in full-on cost management mode. They’re going to look to places like AWS, Fireworks, and OpenRouter to manage their costs down, while also prioritizing visibility into team usage, easy to set limits, auto cost management with smaller or OS models, speed to enabling state of the art models, zero data retention and other enterprise-friendly policies, access management, ease of use, and tool connections and control. If I were an orchestrator, I would go loud with marketing right about… now.
@david_nix ·
Is Cloud AI a scam? I was running some evals Local qwen models 27b and 9b right off my laptop. Simple data extraction. Here's a pile of text, give me JSON plz. My local models passed evals 100%. Switched to Openrouter. Exact same models... or so I thought. Several evals failed. So what gives? Did they quantize more than advertised?
@pukerrainbrow ·
46% of US company AI usage on openrouter is now chinese models openai cut luna prices 80% three weeks after launch. $1/$6 down to $0.20/$1.20 per million tokens same week they hit a billion users a billion users and they still couldn't keep their prices because chinese labs charge less and companies moved US tried to win with restrictions. china won with pricing have you tried any chinese models yet?
@VaibhavSisinty ·
Every major AI company shipped the same thing in the last 48 hours. Model routers. A layer that decides which model to use so you don't have to. Cursor launched Cursor Router on Tuesday. It reads every request, checks the complexity, and sends it to the cheapest model that can handle it. 60% cost savings. Same quality. But Cursor isn't alone ↓ → OpenAI already built routing into ChatGPT. It auto-switches between Instant and Thinking models based on query complexity. → Factory AI launched Factory Router. Claims 20-25% token savings with no quality drop. → OpenRouter upgraded their Auto Router across 400+ models from 60+ providers. → IDC predicts by 2028, 70% of top AI enterprises will use dynamic model routing. All within the same window. That's not coincidence. That's consensus. Here's the problem they're all solving. Most people pick one model and use it for everything. Cursor's own data says 60% of their developers do this. Frontier models built for complex reasoning are being used to fix typos and write Slack messages. That's a Ferrari doing grocery runs. Routers fix that automatically. Simple task goes to a cheap model. Hard task goes to frontier. The model race is slowing down. The top models are within a few points of each other. The real competition has quietly moved to who can deliver the same output at the lowest cost. That's what routers do. And every major player just shipped one in the same week. The era of one model for everything is ending. Smart routing just started.
@Layton_Gott ·
Open source local models are the future... You should be preparing for that now. But that doesn't mean ignore the present. Right now the paid subscriptions are worth WAY more than what you pay for them, so use them while that lasts. Here's how I do both at once: I run a free local model as my operator. Qwen3.6 27B, running on my own machine, costing me nothing. It doesn't do the heavy lifting itself. It routes. It dispatches tasks out to Claude Code, Codex, DeepSeek V4 Pro, GLM 5.2, and honestly any open source model I want through OpenRouter. So the second a new open model drops, I hook it in. My operator stays free and local, and the actual work goes to whatever model is best for it right now. I get the best of the present (frontier subscriptions worth more than I pay) and I'm already set up for the future of running local models although I'd like to get some more ram soon. Prepare for where it's going... But don't move past the present quite yet without taking FULL advantage of it.
@rohanpaul_ai ·
Reuters: A mysterious AI model named Hunter Alpha recently appeared on OpenRouter. Rumors say it might be a secret test version of DeepSeek V4. - This model is massive 1T param. - 1M token context window - Many noticed reasoning patterns look a lot like the chain-of-thought style seen in previous DeepSeek models During testing, the model claimed it was trained mostly on Chinese data with a knowledge cutoff of May-25, which matches DeepSeek’s official specifications. --- reuters .com/business/media-telecom/mystery-ai-model-has-developers-buzzing-is-this-deepseeks-latest-blockbuster-2026-03-18/
@ai_for_success ·
Qwen has just launched Qwen3.6-Plus model and it’s Free on OpenRouter , it is a significant upgrade over the Qwen3.5 series - 1M context window by default - significantly improved agentic coding capability - better multimodal perception and reasoning ability - Also available via the Alibaba Cloud Model Studio API along side OpenRouter. Some use case below from blog , Video Source and Credit : Qwen Blog Qwen is also planning to Open-source smaller-scale variants of the Qwen3.6 series shortly. 1/5
@burkov ·
The fact that the models are updated all the time and then older versions are deprecated is super annoying. You just found a model that works well for a given use case and costs well too. So you are using it directly from the provider or via OpenRouter. Then they decide to deprecate it, and your only choice, if the model is open weight, is to rent GPUs to run it, which is too expensive for the use case. Otherwise, you need to test them all once again until you find one that works reasonably well for a reasonable price for your use case.
@ai_for_success ·
There is new anonymous model called elephant-alpha, currently in the top 10 on OpenRouter by daily usage. Looks good for agentic work because OpenClaw and Hermes Agent are among the top apps using it. Tried it on some code review and refactoring. It is pretty fast and handles it well. Not a deep reasoner but solid for quick execution tasks. Feels similar to Grok Fast. Anyone tried this yet?
Watch video
@pukerrainbrow ·
OpenRouter's image model rankings (June 2026): 1. Nano Banana (Gemini 2.5 Flash Image), 1.71M requests, 40.7% share 2. Nano Banana 2 (Gemini 3.1 Flash Image), 1.21M requests, 28.8% 3. Nano Banana Pro (Gemini 3 Pro Image), 825K requests, 19.6% 4. Seedream 4.5 (ByteDance), 99K requests, 2.4% 3 google models are swallowing 89% of the platform’s image gen traffic. whatever “open-weights from china” narrative people want to push, the numbers don’t even start to back it up. Is this the one category nobody's actually built the open-weight alternative for yet? Genuine question.
@alex_verem ·
The most-used model on OpenRouter this week isn’t GPT, Claude, or DeepSeek. It’s Tencent’s Hy3 🤯 8.98T tokens. #1 on the platform. So I tested whether the efficiency shows up in the actual output. I asked it to: “Write a Python function that finds the maximum profit from one stock purchase and sale. Start with the brute-force solution, optimize it to O(n), then explain the time and space tradeoffs.” Hy3 first returned the brute-force O(n²) approach, then rewrote it as a single-pass solution that tracks the minimum price. It explained why the second approach reduces the time complexity from O(n²) to O(n), along with the space tradeoff. Architecturally, Hy3 has 295B total parameters while activating 21B per token. The published specs include: • 256K context • Fast and slow reasoning in one model • Apache 2.0 open-source licensing • Tencent Cloud lists pricing at 1 CNY/M input tokens and 4 CNY/M output tokens. Those specs may partly explain its recent OpenRouter usage, although one coding prompt obviously isn’t a full benchmark. @TencentHunyuan #Hy3
@alliekmiller ·
Xiaomi now processes more AI tokens than OpenAI. On OpenRouter, Chinese models just crossed 45% of all token volume. Anthropic is at 15.3%. OpenAI is at 7.4%. For now, the frontier is American models, but the volume is Chinese models. Open source models are often good enough at coding and agents for most tasks, can ship a 1M token context window, and cost a fraction of GPT or Claude. 80% as good at 20% of the price is pretty freaking attractive to a procurement team.
@RoundtableSpace ·
Someone got fed up with Claude Code blocking tasks mid-session and lecturing them about blood test questions, so they switched to Kimi K3 through OpenCode for $19. Here is the exact setup: → Ask Codex or Claude Code to install OpenCode → Go to https://t.co/N9c6hpx2AZ, make an account and pay for membership at $19 a month → Get your API key at https://t.co/3BBIldfy53 → Run OpenCode, type /connect and connect to Kimi Code, paste the API key → Set it to BUILD mode with Shift+Tab → Go to /settings and enable bypass permissions OpenRouter hit rate limits immediately so going direct to Kimi's API was the fix. Once connected, Kimi K3 runs without restrictions or interruptions.
@Marktechpost ·
Arcee AI just released Trinity Large Thinking, an Apache 2.0 open reasoning model built for long horizon agents, multi turn tool use, and structured outputs. Key specs: • 400B sparse MoE • 13B active params per token • 4-of-256 routing • 262k context on OpenRouter • #2 on PinchBench, behind Claude Opus 4.6 Full analysis: https://t.co/pBNXW0tqDN Technical details: https://t.co/Vw298MR3yg Model weight: https://t.co/fFoFEKhNBS @arcee_ai #AI #LLM #OpenSourceAI #Agents #MachineLearning
@rohanpaul_ai ·
Just a few days before, OpenRouter raised $113M at a $1.3B valuation and now, WSJ reports Stripe is negotiating a roughly $10B acquisition talk with it. Stripe already handles OpenRouter's payments, so ownership would combine billing with AI distribution. That position could give Stripe usage data and pricing influence across competing model providers. Companies increasingly mix models because the best choice changes by task, cost, and traffic. OpenRouter sits between those buyers and labs, collecting payments while routing each request. Stripe would mainly pay for control of a growing decision layer above model providers. The price works only if model choice becomes infrastructure as basic as online payments.
@aakashgupta ·
Kimi K2.6 hit #1 on OpenRouter last week with 1.58T tokens. Claude Sonnet 4.6 came in at 1.36T. The leaderboard isn't measuring what most people think it's measuring. OpenRouter is a routing aggregator. Developers set a quality bar, OpenRouter sends the request to the cheapest provider that clears it. The leaderboard is downstream of that. It's a proxy for who wins the price-per-good-enough-token war on the slice of the market that doesn't have direct enterprise contracts. Kimi K2.6 costs $0.60 per million input tokens on Moonshot's own API. Sonnet 4.6 is $3. Opus 4.7 is $5. So Sonnet is 5x Kimi's price and Opus is 8x. And Claude Sonnet still pulled 86% of Kimi's volume on the same platform, at 5x the cost. That's the actual story in this chart. The most expensive workhorse model in the industry is doing 86% of the volume of the new cheapest-good-enough option, on the platform built to commoditize the middle. Now scale that. Anthropic has THREE models in the top 10. Sonnet 4.6 (1.36T) + Opus 4.7 (1.15T) + Opus 4.6 (685B) = 3.2 trillion tokens. That's 2x Kimi's total volume on a platform that systematically routes against expensive models. Opus 4.7 launched four days before Kimi K2.6 and is already at #4, up 279% week over week. The thing nobody screenshots: the Y-axis. OpenRouter is processing 28 trillion tokens a week. The chart shows a near-vertical climb starting in February. The price-sensitive slice of the dev market is exploding. Open WebUI users, Cline and Roo Code agents, indie builders without enterprise contracts. And underneath that explosion, the market is splitting cleanly. Chinese open-weight (Kimi, DeepSeek, MiniMax) owns the volume war on price. Anthropic owns the segment that pays for quality. Google parked two Flash models in the top 10 for the workhorse middle. xAI quietly hit #7. Everyone is winning a different game. OpenRouter is the cheap seats. The actual revenue is happening inside Cursor, Claude Code, Bolt, Lovable, Replit. Direct integrations. Enterprise contracts. None of that volume shows up on this chart. Kimi at #1 is real. It took an open-weight model at 1/5 the price to edge out Sonnet by 16% on the platform built to commoditize them both.
@KenTheRogers ·
One of the biggest benefits of using @OpenRouter is outage safety. If my personal Hermes agent is hooked up to Codex auth and they have an outage it’s annoying but not a huge deal. If a client that I built a custom agent for to run a key workflow has an outage, that’s unacceptable. Anything that needs stability when providers have outages should be built on OpenRouter.
@arcee_ai ·
An update on Trinity-Large-Preview: We are officially sunsetting the hosted preview endpoint on OpenRouter. The preview model has served over 3.3T tokens since it's launch in late January. A final thank you to everyone who used it. If you still have active pipelines targeting the preview endpoint, consider transitioning them to Trinity-Large-Thinking to avoid any service disruption. As a reminder, Trinity-Large-Thinking is free on @OpenRouter until 5/23. https://t.co/a9oVQkIfPn
@ShinkaIoT ·
Fusion turns your prompt into a small multi-model deliberation. The "Fusion" API from Openrouter entirely flips the math on how we use frontier models. A panel of expert models (see below) analyzes your prompt in parallel with web search and web fetch enabled, then a judge model synthesizes their responses into a structured analysis — consensus, contradictions, partial coverage, unique insights, and blind spots — and writes the final answer from it. 🚀🔀 This isn't a price story; it’s an architecture story. The breakdown of how Fusion works and why single-shot prompts are officially obsolete: • 🎰 The Three-Layer Deliberation: When a prompt hits OpenRouter/Fusion, it triggers a compound workflow. First, a parallel Panel of models analyzes the prompt (with web search and tools active). Next, a Judge model audits all competing outputs, mapping out consensus, contradictions, and blind spots. Finally, a Synthesizer builds a clean, unified response grounded in that multi-perspective analysis. • 📊 Benchmarking the Multi-Model Advantage: OpenRouter tested this framework on Draco's strict deep research benchmark across law, finance, and medicine. A panel of Fable 5 and GPT 5.5 synthesized by Opus 4.8 scored an elite 69.0%, outperforming Fable 5 running solo (65.3%). Even more impressive: a budget preset combining cheap models (Gemini 3 Flash, Kimi K2.6, and DeepSeek V4 Pro) hit 64.7%—matching frontier performance at half the price. • 🔄 The Power of Self-Synthesis: The data proves that 75% of Fusion's performance jump comes strictly from the synthesis step, not model scale. Even when a single model like Opus 4.8 evaluates its own work in a dual-panel setup, its benchmark score rises from 58.8% to 65.5%. Iteration and comparative verification beat single-shot answers every time. • ⚖️ Stacking Cost vs. True Accuracy: Because Fusion runs 3 to 5 models in parallel, you pay the cumulative token cost of the underlying completions. However, the budget preset stacks up cheaper than running a single, premium frontier model while offering a built-in audit trail. Knowing that three independent models verified the same fact provides a completely different level of enterprise confidence. The Takeaway: The future of AI is cognitive diversity and smart orchestration, not single-model monopolies. Stop asking one model a question and blindly pasting whatever it spits out. Head to OpenRouter, initialize the Fusion API, and start running multi-model panel checks on your deep research, market analysis, and content workflows. Stop prompting. Start orchestrating. 🖥️⛓️
@ChinaScience ·
The share of tokens used by US companies on Chinese AI models via OpenRouter — one of the major overseas multi-model aggregation and routing platforms, serving as an important window for observing global LLM usage trends — has sat above 30% every week since February 8, with that figure rising as high as 46%, a significant surge from just 4.5% in the first half of 2025 and 11% in past 12 months. Attributing the trend to a combination of factors including cost-effectiveness, enhanced performance, open-source autonomy, and supply chain diversification, Chinese analysts said that they expect more foreign companies to adopt Chinese AI models.
@Aizcalibur ·
I believe this will be helpful to my fellow community members. Did you know @NousResearch Hemes has an auxiliary system for specific functions that can be routed through specific providers? 🤌 But here is what I stumbled upon and fixed before getting rekt. I noticed GLM-5 requests showing up on my OpenRouter dashboard even though my primary model was configured to use ZAI's direct API. GLM-5 doesn't take image input and this is where the auxiliary system comes into play. I assigned Gemma 4 via OpenRouter as the provider to 'vision' which takes image input and provides text output to GLM-5. The config looked correct on the surface and even Hermes just worked as it always does. Luckily, I checked my OpenRouter logs and saw GLM tags and was super confused. Shared the screenshots with Hermes and we dug deep into the configs. For reason I don't know, when I asked it to set OpenRouter for auxiliary, it got added to delegation provider with a blank model which was then taking Zai. Plus 'auto' was set in auxiliary and it routed for GLM but via OpenRouter for all the other tools except for 'vision' for which it took Gemma 4 as intended. As per Hermes, the fix here was changing the delegation + specific providers for each auxiliary tool. Quite an interesting system. If I hadn't seen the logs, I would be wasting money on OpenRouter even though I have a direct coding plan from Zai. I recommend checking logs and auditing if you use multiple models.
@alvinfoo ·
Here’s what most people are missing right now: Everyone is still debating which model is smarter — GPT vs Claude. But the real game has already shifted. Last week’s usage data from OpenRouter tells a very different story: All top 6 models by token consumption were Chinese. Let that sink in. This isn’t about benchmarks. This isn’t about demos. This is about distribution at scale. Take Alibaba’s Qwen: → 100,000+ derivatives across open-source ecosystems → Embedded everywhere developers actually build → Quietly compounding usage, day after day While others compete for attention, this ecosystem is capturing adoption. And that’s the real moat. Because in AI: •The best model doesn’t always win •The most used model does China didn’t just build competitive models. They seeded a self-replicating ecosystem, one that grows through developers, forks, integrations, and real-world usage. No hype. No noise. Just distribution compounding. The takeaway? 👉 If you’re still focused only on model quality, you’re playing yesterday’s game. 👉 The winners of this cycle will be decided by who owns the rails of usage. And right now, the data is telling us: the center of gravity is shifting.
@localjulius ·
Xiaomi's free AI usage beat Claude Sonnet on OpenRouter last week. By 58%. > Xiaomi MiMo-V2-Pro: 1.77T tokens/week > Step 3.5 Flash (free): 1.61T > MiniMax M2.5: 1.39T > DeepSeek V3.2: 1.23T > Claude Sonnet 4.6: 1.12T > Claude Opus 4.6: 1.06T Xiaomi is beating Claude Sonnet 4.6 by 58%. Is this free-tier volume arbitrage? Or is the quality genuinely there?
@gudanglifehack ·
AI Digest: OpenRouter costs and OpenAI’s GPT-5.6 Sol update 1. Accessing DeepSeek models through OpenRouter is more expensive OpenRouter routes requests through its own servers before reaching the third-party host, incurring additional costs and inefficiencies due to routing, caching, and infrastructure overhead.
Best Tweets by Topic