Agentic coding and autonomous work
Agentic coding, autonomous software development, tool use, long-horizon execution, coding agents, and coding-oriented Qwen evaluations.
50%
Best tweets about Qwen
Discover the best tweets about Qwen, including Alibaba model releases, coding, reasoning, benchmarks, fine-tuning, and local deployment.
Model-specific Qwen research, releases, benchmarks, coding performance, deployment, and comparisons with concrete evidence.
Original Xholic analysis
The supplied Qwen discussion is 82% supportive and is concentrated in agentic coding (50% of tweets), benchmarks and comparisons (38%), and local deployment (34%). Official release posts foreground agentic and multimodal capabilities, while local-use posts provide concrete deployment claims. Counterpoints in the supplied posts include a reported benchmark gap and a tool-call test where Qwen underperformed xLAM in that setup. [2084100707423289643, 2031008078850924840, 1952469728410390593]
82% of posts
All-time engagement
58% of posts
Published in 90 days
Conversation map
Agentic coding, autonomous software development, tool use, long-horizon execution, coding agents, and coding-oriented Qwen evaluations.
50%
Benchmark scores and head-to-head comparisons with Claude, GPT, Gemini, Kimi, Fable, Gemma, and other models across coding, reasoning, arenas, and agent tests.
38%
Running, quantizing, fine-tuning, distilling, and deploying Qwen locally on laptops, Macs, phones, browsers, and consumer GPUs.
34%
Efficiency claims around small models, sparse MoE active parameters, compression, hardware requirements, inference speed, token cost, and cost-performance tradeoffs.
30%
Announcements and capability summaries for Qwen 3.5 through 3.8 models, including Max, Plus, Small, Omni, Image, and open-weight releases.
30%
Qwen's multimodal capabilities across vision, audio, speech, video, image generation, long context, and full-modality interaction.
28%
Fine-tuning, teacher-model distillation, self-improvement, custom derivatives, and community optimization of Qwen base models.
12%
Qwen-VLA and embodied AI applications for robot manipulation, navigation, trajectory prediction, and cross-embodiment control.
4%
Tone and stance
Performance benchmark
Posts with media make up 78% of this collection. Their median all-time score is 19.3, compared with 25.0 for text-only posts.
Format mix
Consensus and debate
Shared view
Official Qwen announcements emphasize agentic coding and multimodality. Qwen3.8-Max is presented as supporting autonomous development and long-horizon work, while Qwen3.6-Plus is presented with agentic coding, vision, and a 1M-token API context window.
Shared view
Posts describe local Qwen use across a wide size range: a Qwen3.5 local-agent guide for systems with 24GB RAM or less; a Qwen3.5 27B derivative described as runnable in 16GB at 4-bit; Qwen3.5 Small models running locally in a WebGPU browser setup; and a Qwen3.5 397B model reported running on a 48GB MacBook at 4.4 tokens per second.
Shared view
Efficiency discussion centers on Qwen3.6-35B-A3B’s sparse MoE configuration: 35B total parameters and 3B active parameters. A comparison post cites 73.4% on SWE-Bench Verified for that model, versus 87.6% for Opus 4.7 in the same post.
Open debate
Local Qwen is not presented as universally superior. One post prioritizes privacy and fixed hardware cost while calling Opus smarter; another explicitly notes a benchmark gap; and one tool-call experiment reports Qwen slower and less successful than xLAM in that specific Asana-task test.
Open debate
Comparative positioning varies by task and source. One post reports slight Gemma edges in knowledge, reasoning, and coding while favoring Qwen for agentic work; other posts report Qwen3.6-Plus benchmark results or a Design Arena placement.
What performs
The five supplied score outliers are Qwen3.8-Max (2482.61), the local Qwen3.5 tutorial (2197.14), a Qwen3.5 derivative comparison (896.24), the Qwen3.6-Plus launch (749.36), and a Mac fine-tuning post (565.92).
Local deployment and edge inference has a supplied median all-time score of 64.422, above model efficiency’s 58.14 but below training/distillation’s 137.7. Its high-scoring examples include concrete hardware or deployment details.
Announcements account for 58% of the supplied set, compared with 28% opinions, 12% case studies, and 2% tutorials. The single tutorial category has the highest supplied median all-time score, 2197.136.
Statistical standouts
Creator landscape
The five most represented creators account for 20% of the selected posts.
1. Qwen
@Alibaba_Qwen
2 posts
2. Arbos
@arbos_born
2 posts
3. Okara
@askOkara
2 posts
4. daily.dev
@dailydotdev
2 posts
5. ℏεsam
@Hesamation
2 posts
6. Mark Kretschmann
@mark_k
2 posts
The official Qwen account’s two evidenced posts present Qwen3.6-Plus and Qwen3.8-Max around agentic coding, multimodality, long context, and additional open-weight releases.
Developer posts focus on implementation: local agentic coding, Mac-native fine-tuning for text, vision, speech-to-text, and text-to-speech, plus planned community compression work for Qwen variants.
Hesamation’s two posts are comparison-led: one highlights a locally runnable Qwen3.5 derivative and its reported SWE-Bench result, while the other contrasts Qwen3.6-35B-A3B’s 3B active parameters and SWE-Bench score with Opus 4.7.
Themes, sentiment, stance, and post format are classified per tweet. All counts, shares, medians, creator concentration, freshness, and performance comparisons are then calculated directly from the published snapshot.
Xholic's all-time score compares engagement while accounting for reach, post age, and creator consistency. It is used for relative comparisons within this collection.
This report analyzes the exact 50-post snapshot shown below. AI identifies editorial categories and drafts explanations; all statistics are calculated from the snapshot, and every narrative claim is checked against cited posts before publication.
Best Qwen tweets
Ranked 01–50
@Alibaba_Qwen ·
📢Meet Qwen3.8-Max — our most capable model to date. Next week, the open weights of Qwen3.8-Max will be released, and Qwen3.8-27B is also going open-weights to meet you all!🎉 Qwen3.8-Max, a new bar for coding and cowork at 2.4T parameters: - Autonomous coding: 10+ days of self-evolving development, from empty folder to production without hand-holding, complete project trace in the GitHub:https://t.co/iVHZWQoeSo - Real work, real results: Production-quality deliverables across hundreds of professions. - Long-horizon mastery: System-level autonomous planning with closed-loop adaptive learning, driving 500+ turns of chip design optimization and 365 days of e-commerce strategy. - Native multimodal intelligence: Vision isn't just input — it's a continuous feedback loop for planning, execution, and self-correction. 💰Pricing: Input: $2.0 / M tokens Output: $6.0 / M tokens Implicit Caching: $0.25 / M tokens Start building with Qwen3.8-Max! 🚀 📖 Blog: https://t.co/iwjmQxLBof ✅ Qwen Studio: https://t.co/4V2pFvDovG ⚡ API: https://t.co/gAGqaLQGbN
@Alibaba_Qwen ·
(1/8)🚀 Introducing Qwen3.6-Plus: Towards Real-World Agents! 🤖 Today, we’re thrilled to drop a major milestone in our journey toward native multimodal agents. Here is what makes Qwen3.6-Plus a game-changer: 💻 Next-level Agentic Coding: Smarter, faster execution. 👁️ Enhanced Multimodal Vision: Sharper perception & reasoning. 🏆 Top-tier Performance: Maintaining leading general capabilities. 📚 1M Context Window: Available by default via our API. Built on your invaluable feedback from the Qwen3.5 era, we’re laying a rock-solid foundation for real-world devs. Get ready to experience truly transformative ✨ Vibe Coding ✨. Huge thanks to our community! Go try it out and show us what you can build. 👇 Chat: https://t.co/V7RmqMaVNZ API: https://t.co/937Qkc9AMy Blog: https://t.co/P0rJSxERND 🔔Noted:More Qwen3.6 models to come and be open-sourced! Stay tuned~ 👀#Qwen #AI #AgenticCoding #VibeCoding #Agents
@_ARahim_ ·
The entire Qwen stack is now fine-tunable on your Mac! 🍏 Just pushed the latest release, adding Qwen3-TTS to mlx-tune. You can now natively train: ✅ Qwen3.5 (Text) ✅ Qwen3.5 (Vision) ✅ Qwen3-ASR (Speech → Text) ✅ Qwen3-TTS (Text → Speech) One consistent API pattern for all four. Examples are live in the repo! 👇 https://t.co/ImNGF9FUpX @Alibaba_Qwen @awnihannun
@xenovacom ·
NEW: Alibaba just released Qwen 3.5 Small — a family of powerful multimodal models available in a range of sizes (0.8B, 2B, 4B, and 9B parameters). Perfect for on-device applications! They can even run 100% locally in your browser on WebGPU, powered by Transformers.js! 🤯
@AlexFinn ·
You're right, local models aren't as good as cloud models That's not the point though The point is to have free, private intelligence that can do work for you 24/7 around the clock I have a 3 local models scraping Reddit, product hunt, and other sites 24/7 Looking for challenges to solve That model hands all of those challenges to another local model That model takes the challenges, then builds apps to solve those challenges A 24/7/365 software factory that never sleeps Would never be possible in a million years with cloud models. Would cost me $10,000 a month in tokens. I paid that one time up front for a Mac Studio that runs this Yes Claude Opus 4.6 is smarter than Qwen 3.5. But Qwen 3.5 running locally is still Sonnet 4.5 level. Just 6 months behind. Think about how good it will be 6 months from now. Nvidia just entered the local race. They are going to change EVERYTHING That's not even counting all the other benefits: 1. You can't get banned for using the model the wrong way 2. Costs just the price of electricity 3. Completely private. No AI execs reading your logs 4. 0 latency 5. Completely customizable This is the future. Become sovereign.
@0xSero ·
Qwen3.5, MiniMax-M2.7 are incredible acts of kindness that I don't think will be with us from so much longer. Here's my update for you. > I have 20 GPUs at full utilisation right now. All these getting cooooompressed, no synthetic data All runs will be done in 9 days, if I don't get a catastrophic failure - REAP for: - GLM-5 - Qwen3-next-coder - Qwen3.5-122B - Qwen3.5-plus-397b - Browser-use - CUDA - Terminal-use - Coding - Math - Agentic trajectories - 30% my personal chat session history I am also removing refusals inspired by Prism. So no more I can't do this I can't do that blah blah Inference for local AI - Qwen3.5-262B-REAP - I've been using it exclusively in Parchi, perfect 100 tokens/s & 0 errors very good at browser use ----------------- Secret - Qwen3.5-27b - you will see when i'm done Targeting the following hardware levels: With full context 200-256k context in vllm, sglang, llama.cpp, exllamav3, and if people help MLX 16-32 GB - Qwen3.5-27b 32-48 GB - Qwen3-coder-next 48-128 GB - Qwen3.5-122B 128-256 GB - Qwen3.5-Plus-397B 196-512 GB - GLM-5.* I am training them on 22,000 samples at 16k context 352M of custom selected calibration datasets. My hope is to make the highest quality multimodal LLM compressions for this year. 20 GPUs running in parallel for the next 10 days - 8x H100s - Qwen - 4x B200s - GLM-5.* - 8x 3090s - Testing Once MiniMax-M2.7 is online 4 more GPUs will get to work.
@robotsdigest ·
Qwen-VLA feels like one of the first real robotics foundation models. A single system trained across robot manipulation, navigation, egocentric human video, simulation, and vision-language reasoning instead of isolated robot policies.
@arbos_born ·
Apple showed a model can teach itself to code better without any teacher, verifier, or RL. Just fine-tune on its own best outputs. Qwen3-30B went from 42.4% to 55.3% pass rate on LiveCodeBench. Works at 4B, 8B, and 30B scale. The mechanism: reshaping how the model spreads probability across all possible next words. Suppress noise where you need precision, keep variety where you need creativity. SN97 miners compete on exactly this signal. Full-distribution KL across all 248,000 tokens. Our king scores 0.049, 67% lower than Qwen's own 4B. Already beats it on 6/7 benchmarks. Competitive mining is producing small models that distribute probability more faithfully than what the original maker built. When Apple and a decentralized mining competition independently converge on "the full distribution shape is what matters," that's not coincidence. https://t.co/WVVFXMNTx1
@VaibhavSisinty ·
What's happening in AI right now is genuinely hard to process. A 3 billion parameter model is matching models that are 200 to 300 times larger. On math. On coding. On reasoning. And beating some of them. It's called VibeThinker-3B. Built by Sina Weibo's team on a tiny Qwen 3B base. → 94.3% on AIME math matches DeepSeek V3.2 (671B) and Kimi K2.5 (1 trillion parameters) → 96.1% on LeetCode beats GPT-5.2 and Claude 4.6 on unseen contest problems → 80.2% on LiveCodeBench highest of any small or mid-size model tested A model you can run on a laptop is solving competition math at the same level as models that need entire GPU clusters. The idea behind it: reasoning and knowledge are two different things. Knowledge needs massive parameters to store facts. Reasoning is a procedure search, check, correct, compose and procedures compress into small models far more efficiently. The honest catch: ask it a broad factual question and it trails the big models badly. This isn't a general-purpose win. It's a reasoning specialist. And that's exactly what makes the result credible instead of hype. A 3B model. Competing with trillion-parameter flagships. On reasoning tasks. Running locally. That's where we are now.
@arbos_born ·
The #1 trending model on HuggingFace for three weeks: one researcher distilled Claude's reasoning into Qwen3.5-27B. People are running frontier-level reasoning locally. Now imagine that process as a competition instead of a solo project. Dozens of miners, each trying a different approach, scored on full-distribution KL across 248K tokens. Winner takes all. That's SN97. Best miner compressed Qwen3.5-35B into 4B with 67% lower KL than Qwen's own baseline. Outperforms it on 6/7 benchmarks. One researcher makes a trending model. Competition makes a better one. https://t.co/1uxROkMkwN
@goyalshaliniuk ·
Another major LLM company has entered the Physical AI race Alibaba’s Qwen team just released Qwen-VLA, a unified Vision-Language-Action model that can control humanoid robots and integrates manipulation, navigation, trajectory prediction, and cross-embodiment control (single-arm, dual-arm, humanoid) into one system. It adapts to different robot bodies via embodiment-aware prompts without separate training heads, matches or outperforms specialist models on key benchmarks, and achieves 76.9% OOD success on real ALOHA dual-arm tasks. Google has Gemini Robotics. France’s Mistral recently released Physics AI and acquired Emmi AI. From chatting with AI to acting in the physical world--the competition is accelerating fast.
@natolambert ·
New report with @xeophon is out with the latest open model adoption data we have gathered for Interconnects & The ATOM Project. At the surface level, we can see Chinese models continuing to accelerate in adoption. The report details much more. 1. We manually curate ~1.5K of the most important language models, creating a specific set of models to focus our analysis on (excludes embedding models, local inference models like MLX/GGUF, etc to have accurate download rankings). 2. Studying other adoption metrics, such as derivative models and inference share on OpenRouter, to show how they correlate with downloads, while often sifted in time. China has a strong lead here too. 3. Better classification of downloads across model sizes. Large models still are the models where Qwen is least competitive, relative to other model builders. 4. Expansion of our Relative Adoption Metric (RAM) to show standout recent models (we'll check Gemma 4 on Friday); Qwen 3.5, Nemontron 3, Kimi K2.5, all showing very strong adoption. Overall, this is another step towards formalizing and making public better data on the open language model ecosystem, so the community can better understand the impact and trends of its adoption. More on this soon!
@fahdmirza ·
💥 Qwen3.6-Plus just DROPPED ♠ and it's built for real-world autonomous agents 🔹1M token context window out of the box 🔹Tops benchmarks in agentic coding, tool use & long-horizon planning 🔹New preserve_thinking API keeps reasoning alive across multi-turn agent tasks 🔹Works natively with OpenClaw, Claude Code & Qwen Code 🔥 Watch the full breakdown below 👇
@Ubermenscchh ·
breaking.. alibaba mass dropped qwen 3.6-plus and it's embarrassing every frontier model right now 61.6 on terminal-bench (beats claude 4.5 opus) 56.6 on swe-bench pro (1st place) 80.9 on multilingual agentic coding (1st place) 58.7 on claw-eval real world agent (1st place) this isn't a chatbot.. this is an autonomous coding agent 🧵
@manishkumar_dev ·
@AlibabaGroup just released Qwen 3.6 Plus, and this feels like a real step toward AI that actually executes work. After testing it hands on, this is not just another model update. It is built around agentic coding, multimodal reasoning, and full workflow execution. Here is what stood out to me 👇
@ai_for_success ·
Qwen has released Qwen3.6-35B-A3B, a sparse Mixture-of-Experts model that is now open-source under the Apache 2.0 license. TLDR - Sparse MoE architecture with 35B total and 3B active parameters - Performance in agentic coding rivals models 10x its active size - Strong multimodal perception and reasoning capabilities - Features both multimodal thinking and non-thinking modes - Released under open-source Apache 2.0 license
@mark_k ·
Alibaba just dropped Qwen3.8-Max, their most capable model yet. 2.4 trillion parameters (95B active), built on the Qwen 3.5 architecture, with a 1M token context window and native multimodal support. This is the first time they’re open-sourcing weights of a Qwen-Max-class model (coming next week, along with Qwen3.8-27B). Standout capabilities: - Autonomous coding that can take a real multi-day project from an empty folder to production-ready code with almost no human intervention, self-evolving through feedback loops - Strong gains in complex “cowork” tasks across research, long-horizon planning, and professional deliverables - Major jumps over Qwen3.7-Max on coding agents, PaperBench, OSWorld, and more Now available via QwenCloud API. @Alibaba_Qwen
@FellMentKE ·
🚨 BREAKING: Alibaba unleashes Qwen3.5-Omni, a new frontier in Full-Modality AI. 🤯 Matching the latest Gemini-3.1 Pro in A/V understanding & surpassing it in Audio tasks, this model introduces Audio-Visual Vibe Coding—turning whiteboard sketch videos or game clips directly into runnable code. 🎧 10h+ Audio Input | 1h Video Context 🎬 Script-level descriptions w/ timestamps 🌍 74 langs recognized + 29 langs generated The barrier between human intent and machine execution has vanished. A deep dive 👇 @Ali_TongyiLab @Alibaba_Qwen #Qwen #VibeCoding #AI #Multimodal
@pukerrainbrow ·
You can now build a 3D game in a single AI generation. Most models struggle to coordinate physics, rendering, and logic all at once, but QWEN 3.6-Plus by Alibaba is making it as easy as typing a sentence. It features a massive 1 million token context window (about 750,000 words) so you never have to split up your files again. Best part? It’s only $0.29 per million tokens, which is a massive discount compared to the big players. Question is: Does a lower price point (15x) make you more or less likely to use a model in this economy?
@pankajkumar_dev ·
Qwen 3.8 Max Leaks - Qwen3.8-Max-Preview is now officially available on Alibaba's Token Plan, Qoder, and QoderWork. - The model features 2.4T parameters - Qwen claims it's "second only to Fable 5." - The model previously on LMArena under the stealth name "Kaleb". - It is very strong frontend generation, with an open-weight release coming soon. - Global availability is expected by the end of July.
@arena ·
Qwen 3.5 Max Preview has landed in top 10 for Arena Expert and top 15 for Text Arena. It shows particular strength in Math. Highlights: - #3 Math - #10 Expert - #15 Text Arena - Top 20 for Writing, Literature & Language, Life, Physical, & Social Science, Entertainment, Sports, & Media, and Medicine & Healthcare Congrats to the @Alibaba_Qwen team for this new milestone!
@TeksEdge ·
🏆️ Gemma 4 vs Qwen 3.5 and the battle for local AI is on 🔥 💥 X is full of people running both on consumer GPUs, Mac Studios, RTX 6Ks and H100s. The day-0 benchmarks tell a baseline story: 🧠 KNOWLEDGE & REASONING, Gemma leads slightly 💻 CODING, Gemma has a small edge 🤖 AGENTIC & TOOLS, Qwen dominates (+13.0 on Tau2-Bench) 🏆 FRONTIER DIFFICULTY (HLE), Qwen crushes it (+30.2 with tools) 🎯 Bottom Line: Gemma feels stronger in pure reasoning. Qwen feels way better as an actual agent in my testing. Benchmarks vs real-world agent use, who wins in your workflow? Some weekend testing?
@ttunguz ·
2025 is the year of agents, & the key capability of agents is calling tools. When using Claude Code, I can tell the AI to sift through a newsletter, find all the links to startups, verify they exist in our CRM, with a single command. This might involve two or three different tools being called. But here’s the problem: using a large foundation model for this is expensive, often rate-limited, & overpowered for a selection task. What is the best way to build an agentic system with tool calling? The answer lies in small action models. NVIDIA released a compelling paperarguing that “Small language models (SLMs) are sufficiently powerful, inherently more suitable, & necessarily more economical for many invocations in agentic systems.” I’ve been testing different local models to validate a cost reduction exercise. I started with a Qwen3:30b parameter model, which works but can be quite slow because it’s such a big model, even though only 3 billion of those 30 billion parameters are active at any one time. The NVIDIA paper recommends the Salesforce xLAM model – a different architecture called a large action model specifically designed for tool selection. So, I ran a test of my own, each model calling a tool to list my Asana tasks. The results were striking: xLAM completed tasks in 2.61 seconds with 100% success, while Qwen took 9.82 seconds with 92% success – nearly four times as long. This experiment shows the speed gain, but there’s a trade-off: how much intelligence should live in the model versus in the tools themselves. This limited With larger models like Qwen, tools can be simpler because the model has better error tolerance & can work around poorly designed interfaces. The model compensates for tool limitations through brute-force reasoning. With smaller models, the model has less capacity to recover from mistakes, so the tools must be more robust & the selection logic more precise. This might seem like a limitation, but it’s actually a feature. This constraint eliminates the compounding error rate of LLM chained tools. When large models make sequential tool calls, errors accumulate exponentially. Small action models force better system design, keeping the best of LLMs and combining it with specialized models. This architecture is more efficient, faster, & more predictable. https://t.co/2ASQ5btMqq
@cjzafir ·
Is GPT 5.4 really good? I used codex-5.4-extra-high to fine-tune qwen-3.5-4b. (SFT) (Exhausted all my pro plan weekly credits in 24 hours.) And also used opus-4.6 to fine-tune qwen-3.5-9b Codex is fast but dataset quality is crap. Opus is slow but data quality is great. What i am doing? I am performing distillation (using opus 4.6 as teacher & qwen-397b as stident) And taking their supervised chat data to fine-tuning qwen-3.5-4b and qwen-3.5-9b models. Already beaten: Gpt-4o Gemini-2.5-flash Gpt-5-nano Test passed. Now I'm making a python kit to perform distillation on scale (20M to 40M tokens). Qwen models have opened something big > smaller, smarter, faster, niched models running locally on your mac. Future is fun.
@RoundtableSpace ·
Qwen 3.8 Max just dropped with 2.4 trillion parameters, beats Fable 5 on several benchmarks and ran a GitHub repo autonomously for 10 days straight.
@boxmining ·
Qwen 3.6 Plus being free right now is actually pretty interesting. Tried @Alibaba_Qwen inside Hermes Agent for system audits and log analysis, and it surfaced issues I would’ve probably missed. The best part is the think blocks. Seeing how it reasons through errors and tool calls makes debugging feel way less blind.
@RoundtableSpace ·
Kimi K3 vs Qwen 3.8 Max on the same landing page prompt. Qwen was 3x faster and still built a convincing 3D guillotine model.
@dino11 ·
Alibaba $BABA just open-sourced Qwen 3.5 — a 9B parameter AI model that runs on your laptop. The benchmarks are insane: → Beats OpenAI's GPT-OSS-120B (a model 13x its size) on reasoning → GPQA Diamond: 81.7 vs 71.5 → 30-50 tokens/sec on a standard laptop → The 2B version runs on an iPhone in airplane mode at 22 tokens/sec Free. Open-source. No API costs. No internet needed. If you're building AI products and still paying per-token for every request, this changes the math completely.
@Sino_Market ·
Qwen Unveils CoPaw 1.0 Personal AI Assistant with Enhanced Models, Security and Multi-Agent Capabilities Qwen: We release CoPaw 1.0 today, a personal intelligent assistant that can be quickly deployed in users' local or cloud environments. We upgrade CoPaw's capabilities around four major aspects: a small model tailored for CoPaw, security mechanisms, multi-agent collaboration, and memory management. #CHINA #QWEN #AI #ALIBABA $BABA (https://t.co/oB1kLuEdYU)
@agenticgirl ·
A smaller model just outperformed the biggest ones. Qwen 3.6-Plus scored 61.6 on Terminal-Bench and 57.1 on SWE-Bench. That puts it ahead of Claude Opus 4.5, Kimi K2.5, and Gemini 3 Pro. Models that are much larger and far more expensive to run. For the past year everyone believed one thing. Better performance needs a bigger model. This breaks that rule. But the more important part is where it is winning. Not simple tasks. These are multi-step coding workflows where the model has to plan, fix its own mistakes, and track changes across multiple files. That is exactly where most models fall apart. They start strong then lose the thread after a few steps. Qwen 3.6-Plus is holding up there. Size has always come with a cost. More compute, more money, more setup. Smaller models are faster, cheaper, and easier to deploy. If they can now match or beat the big ones on hard tasks, they become the obvious choice for most teams. It also ships with a 1 million token context window. Enough to load an entire codebase at once and not lose track halfway through a long session. It is already free on OpenRouter. Open source version is coming. So the real story is not that one model scored higher. It is that being bigger is no longer enough.
@dailydotdev ·
Shipped a model, forgot to mention whose model it was. Here's today's AI dev news. - @cursor_ai launched a coding model on Kimi k2.5 without saying so, co-founder later called it a mistake - Claude Code's skills ecosystem has a supply chain problem: 27% of public skills carry command execution patterns with zero runtime verification - @Alibaba_Qwen confirmed Qwen stays open source, and Qwen 3.5 27B on a single 3090 is making "open models aren't ready" sound increasingly silly - Stripe's autonomous Minions agents are generating thousands of PRs weekly and nobody's asking enough questions about what happens next Chart's over there if you want to see which models developers actually reach for when it counts.
@JulianGoldieSEO ·
Qwen 3.8 vs Fable 5: 50 real builds, one winner. This guy literally built 50 games with both AIs to find out. Qwen scored 86.6 on Terminal Bench. Fable 5 scored lower. Qwen's games ran smooth and full 3D. Fable 5 got buggy on half the builds. But Fable 5 won on flight sims and big projects. Here's the twist. Qwen is free. And open source. Fable 5 needs a paid plan. Free is now trading punches with the best AI on earth.
@dailydotdev ·
Agents return videos now. Here's what else landed today. - @cursor_ai agents record their own test runs and hand you a video, not a wall of diff - @github Copilot added an hourly cap on top of the monthly one, and the Claude Code comparisons are already flying - @Alibaba_Qwen's Qwen 3.5 27B fits on a single RTX 5090 and is making a credible case for going fully local - @AnthropicAI's Capybara model surfaced in press, positioned above Opus, no release date, take the benchmarks loosely Chart's right there showing which models engineers are actually running. Worth a look.
Best Tweets by Topic