Agentic coding and autonomous work
Qwen for software engineering agents, tool use, terminal tasks, long-horizon planning, vibe coding, autonomous repositories, and real-world development workflows.
50%
Best tweets about Qwen
Discover the best tweets about Qwen, including Alibaba model releases, coding, reasoning, benchmarks, fine-tuning, and local deployment.
Model-specific Qwen research, releases, benchmarks, coding performance, deployment, and comparisons with concrete evidence.
Original Xholic analysis
Across 50 Qwen-related posts, the most prevalent themes are agentic coding and benchmarks/comparisons (25 posts each, 50%), followed closely by local inference and self-hosting (23 posts, 46%). Discussion is largely positive or supportive, but posts also surface caveats about benchmark-to-workflow transfer, unsupervised reliability, and handling sensitive code through hosted services.
66% of posts
All-time engagement
72% of posts
Published in 90 days
Conversation map
Qwen for software engineering agents, tool use, terminal tasks, long-horizon planning, vibe coding, autonomous repositories, and real-world development workflows.
50%
Measured Qwen results on SWE-Bench, Terminal-Bench, Code Arena, Arena, math, agentic evaluations, and head-to-head comparisons with Claude, GPT, Gemini, Kimi, Fable, and Gemma.
50%
Running Qwen on laptops, Macs, phones, browsers, consumer GPUs, and private infrastructure, with emphasis on RAM requirements, throughput, sovereignty, privacy, and cost.
46%
Launches and specifications for Qwen 3.5, 3.6, 3.8-Max, Omni, Flash, Small, and MoE variants, including parameters, context windows, modalities, licensing, APIs, and open-weight plans.
40%
Native vision, audio, video, ASR, TTS, Qwen-Omni, and multimodal coding or reasoning capabilities across the Qwen stack.
22%
Community adaptation of Qwen through teacher-model distillation, SFT, compression, REAP runs, MLX tuning, and specialized small-model derivatives.
14%
MoE active-parameter efficiency, speculative decoding, Metal implementations, compression, quantization, distillation, and performance-per-compute improvements.
14%
Qwen-VLA for embodied AI, robot manipulation, navigation, trajectory prediction, cross-embodiment control, and physical-world benchmarks.
4%
Tone and stance
Performance benchmark
Posts with media make up 74% of this collection. Their median all-time score is 19.3, compared with 16.8 for text-only posts.
Format mix
Consensus and debate
Shared view
Agentic coding and autonomous work appears in 25 posts (50% of the dataset). Qwen’s own release posts position Qwen3.8-Max around autonomous coding and Qwen3.6-Plus around coding, tool use, long-horizon planning, and a 1M-token API context window.
Shared view
Local inference and self-hosting appears in 23 posts (46%). Examples include a guide claiming local Qwen3.5 agentic coding on 24GB RAM or less, a Qwen-derived 27B model described as fitting in 16GB at 4-bit quantization, and Mac fine-tuning support across Qwen text, vision, ASR, and TTS models.
Shared view
Posts cover text, vision, speech, and text-to-speech tooling, while Qwen-VLA is described as a unified vision-language-action model for robotics tasks. Multimodal models account for 11 posts (22%) in the analytics.
Open debate
Posts report strong Qwen benchmark results, but evaluations differ on practical performance. One tester says Qwen3.8-Max Preview feels weaker than K3 in real-world use, while another comparison found Qwen3.5 strong in some local agent tasks but reported skipped tasks, weak formatting compliance, and limits for unsupervised long operations.
Open debate
A local-model advocate values privacy, customization, and fixed hardware spending, while a separate reviewer advises against Qwen3.6 Plus Preview for sensitive code because the reviewer says its terms allow collection of prompts and completions. These are user and reviewer perspectives rather than independently verified deployment guarantees.
Open debate
One coding test scored Sonnet 4 above Qwen3.5 (9.78 versus 9.29). Elsewhere, a Gemma-versus-Qwen post reports Qwen advantages on Tau2-Bench and HLE-with-tools, and a separate post reports Qwen3.6-35B-A3B at 73.4% on SWE-Bench Verified. The cited results use different models, tasks, and test setups.
What performs
The dataset contains 37 posts with media (74%). Their median all-time score was 19.263, compared with 16.795 for text-only posts.
The highest listed outlier was the Qwen3.8-Max release post, with an all-time score of 2482.61. The local Qwen3.5 agentic-coding guide ranked second at 2197.14. Their scores were 137.69x and 121.86x the dataset median, respectively.
A post about a Qwen3.5 derivative described as running locally in 16GB or 32GB configurations scored 896.24, and a post on Mac fine-tuning across the Qwen stack scored 565.92. Both are listed among the five score outliers.
Statistical standouts
Creator landscape
The five most represented creators account for 20% of the selected posts.
1. Qwen
@Alibaba_Qwen
2 posts
2. Boxmining
@boxmining
2 posts
3. Fahd Mirza
@fahdmirza
2 posts
4. ℏεsam
@Hesamation
2 posts
5. Julian Goldie SEO
@JulianGoldieSEO
2 posts
6. 0xMarioNawfal
@RoundtableSpace
2 posts
Alibaba_Qwen posted twice in the dataset and had a median all-time score of 1615.98, the highest median among the listed repeated top voices. Its two evidence posts announce Qwen3.8-Max and Qwen3.6-Plus.
Community posts describe a 24GB-or-less local agentic-coding setup, Mac fine-tuning support for four Qwen model types, and a reported MacBook run of a 397B-parameter Qwen3.5 model at 4.4 tokens per second using 48GB RAM. These are individual implementation reports.
Testing-oriented posts distinguish benchmark outcomes from workflow performance. They report issues including translation errors, skipped tasks, formatting failures, limits in unsupervised long operations, and perceived gaps between benchmark scores and real-world use.
Since the previous snapshot
Themes, sentiment, stance, and post format are classified per tweet. All counts, shares, medians, creator concentration, freshness, and performance comparisons are then calculated directly from the published snapshot.
Xholic's all-time score compares engagement while accounting for reach, post age, and creator consistency. It is used for relative comparisons within this collection.
This report analyzes the exact 50-post snapshot shown below. AI identifies editorial categories and drafts explanations; all statistics are calculated from the snapshot, and every narrative claim is checked against cited posts before publication.
Best Qwen tweets
Ranked 01–50
@Alibaba_Qwen ·
📢Meet Qwen3.8-Max — our most capable model to date. Next week, the open weights of Qwen3.8-Max will be released, and Qwen3.8-27B is also going open-weights to meet you all!🎉 Qwen3.8-Max, a new bar for coding and cowork at 2.4T parameters: - Autonomous coding: 10+ days of self-evolving development, from empty folder to production without hand-holding, complete project trace in the GitHub:https://t.co/iVHZWQoeSo - Real work, real results: Production-quality deliverables across hundreds of professions. - Long-horizon mastery: System-level autonomous planning with closed-loop adaptive learning, driving 500+ turns of chip design optimization and 365 days of e-commerce strategy. - Native multimodal intelligence: Vision isn't just input — it's a continuous feedback loop for planning, execution, and self-correction. 💰Pricing: Input: $2.0 / M tokens Output: $6.0 / M tokens Implicit Caching: $0.25 / M tokens Start building with Qwen3.8-Max! 🚀 📖 Blog: https://t.co/iwjmQxLBof ✅ Qwen Studio: https://t.co/4V2pFvDovG ⚡ API: https://t.co/gAGqaLQGbN
@Alibaba_Qwen ·
(1/8)🚀 Introducing Qwen3.6-Plus: Towards Real-World Agents! 🤖 Today, we’re thrilled to drop a major milestone in our journey toward native multimodal agents. Here is what makes Qwen3.6-Plus a game-changer: 💻 Next-level Agentic Coding: Smarter, faster execution. 👁️ Enhanced Multimodal Vision: Sharper perception & reasoning. 🏆 Top-tier Performance: Maintaining leading general capabilities. 📚 1M Context Window: Available by default via our API. Built on your invaluable feedback from the Qwen3.5 era, we’re laying a rock-solid foundation for real-world devs. Get ready to experience truly transformative ✨ Vibe Coding ✨. Huge thanks to our community! Go try it out and show us what you can build. 👇 Chat: https://t.co/V7RmqMaVNZ API: https://t.co/937Qkc9AMy Blog: https://t.co/P0rJSxERND 🔔Noted:More Qwen3.6 models to come and be open-sourced! Stay tuned~ 👀#Qwen #AI #AgenticCoding #VibeCoding #Agents
@_ARahim_ ·
The entire Qwen stack is now fine-tunable on your Mac! 🍏 Just pushed the latest release, adding Qwen3-TTS to mlx-tune. You can now natively train: ✅ Qwen3.5 (Text) ✅ Qwen3.5 (Vision) ✅ Qwen3-ASR (Speech → Text) ✅ Qwen3-TTS (Text → Speech) One consistent API pattern for all four. Examples are live in the repo! 👇 https://t.co/ImNGF9FUpX @Alibaba_Qwen @awnihannun
@xenovacom ·
NEW: Alibaba just released Qwen 3.5 Small — a family of powerful multimodal models available in a range of sizes (0.8B, 2B, 4B, and 9B parameters). Perfect for on-device applications! They can even run 100% locally in your browser on WebGPU, powered by Transformers.js! 🤯
@AlexFinn ·
You're right, local models aren't as good as cloud models That's not the point though The point is to have free, private intelligence that can do work for you 24/7 around the clock I have a 3 local models scraping Reddit, product hunt, and other sites 24/7 Looking for challenges to solve That model hands all of those challenges to another local model That model takes the challenges, then builds apps to solve those challenges A 24/7/365 software factory that never sleeps Would never be possible in a million years with cloud models. Would cost me $10,000 a month in tokens. I paid that one time up front for a Mac Studio that runs this Yes Claude Opus 4.6 is smarter than Qwen 3.5. But Qwen 3.5 running locally is still Sonnet 4.5 level. Just 6 months behind. Think about how good it will be 6 months from now. Nvidia just entered the local race. They are going to change EVERYTHING That's not even counting all the other benefits: 1. You can't get banned for using the model the wrong way 2. Costs just the price of electricity 3. Completely private. No AI execs reading your logs 4. 0 latency 5. Completely customizable This is the future. Become sovereign.
@0xSero ·
Qwen3.5, MiniMax-M2.7 are incredible acts of kindness that I don't think will be with us from so much longer. Here's my update for you. > I have 20 GPUs at full utilisation right now. All these getting cooooompressed, no synthetic data All runs will be done in 9 days, if I don't get a catastrophic failure - REAP for: - GLM-5 - Qwen3-next-coder - Qwen3.5-122B - Qwen3.5-plus-397b - Browser-use - CUDA - Terminal-use - Coding - Math - Agentic trajectories - 30% my personal chat session history I am also removing refusals inspired by Prism. So no more I can't do this I can't do that blah blah Inference for local AI - Qwen3.5-262B-REAP - I've been using it exclusively in Parchi, perfect 100 tokens/s & 0 errors very good at browser use ----------------- Secret - Qwen3.5-27b - you will see when i'm done Targeting the following hardware levels: With full context 200-256k context in vllm, sglang, llama.cpp, exllamav3, and if people help MLX 16-32 GB - Qwen3.5-27b 32-48 GB - Qwen3-coder-next 48-128 GB - Qwen3.5-122B 128-256 GB - Qwen3.5-Plus-397B 196-512 GB - GLM-5.* I am training them on 22,000 samples at 16k context 352M of custom selected calibration datasets. My hope is to make the highest quality multimodal LLM compressions for this year. 20 GPUs running in parallel for the next 10 days - 8x H100s - Qwen - 4x B200s - GLM-5.* - 8x 3090s - Testing Once MiniMax-M2.7 is online 4 more GPUs will get to work.
@fahdmirza ·
💥 @RedHat_AI Quietly Made Qwen Run 6X FASTER 🚀 ♠ And barely anyone talked about it 🔥 🔹 Speculative Decoding with EAGLE-3 — zero quality loss, just pure speed 🔹 Tiny draft model guesses tokens ahead, big model verifies in one shot 🔹 6.5x faster inference on a single GPU — no extra hardware needed 🔹 Drops straight into vLLM with one command — zero friction deployment 🔹 Full hands-on demo — download, serve, and test it live 🎯 This is where LLM inference optimization is heading — not bigger models, smarter execution 🔥 Watch the full video below 👇
@robotsdigest ·
Qwen-VLA feels like one of the first real robotics foundation models. A single system trained across robot manipulation, navigation, egocentric human video, simulation, and vision-language reasoning instead of isolated robot policies.
@ibragim_bad ·
🚨 SWE-rebench update! SWE-rebench is a live benchmark with fresh SWE tasks (issue+PR) from GitHub every month. updates: > we removed demonstrations and the 80-step limit (modern models can now handle huge contexts without getting trapped in loops!). > we added auxiliary interfaces for specific tasks like in SWE-bench-Pro to evaluate larger tasks fairly, ensuring valid solutions don't fail just because of mismatched test calls. insights: > Top models perform similarly. Among open-source options, GLM @Zai_org shows strong results, and StepFun @StepFun_ai is very cheap for its performance level ($0.14 per task). > GPT-5.4 shows high token efficiency, it ranks in the top 5 overall but uses the lowest number of tokens (774k per task) > Qwen3-Coder-Next & Step-3.5-Flash benefit massively from huge contexts. Qwen is an extreme case, averaging a wild 8.12M tokens. > We evaluated agentic harnesses (Claude Code, Codex, and Junie) and found a few things. Even in headless mode, they sometimes ask for additional context or attempt web searches. We explicitly disabled search and verified their curl commands to ensure they aren't just pulling solutions from the web. 🏆 You can find the full leaderboard here: https://t.co/9jL4lt4UGl 👾 Also, we launched our Discord! Join our leaderboard channel to discuss models, share ideas, ask questions, or report issues: https://t.co/cyWgceqqDa
@VaibhavSisinty ·
What's happening in AI right now is genuinely hard to process. A 3 billion parameter model is matching models that are 200 to 300 times larger. On math. On coding. On reasoning. And beating some of them. It's called VibeThinker-3B. Built by Sina Weibo's team on a tiny Qwen 3B base. → 94.3% on AIME math matches DeepSeek V3.2 (671B) and Kimi K2.5 (1 trillion parameters) → 96.1% on LeetCode beats GPT-5.2 and Claude 4.6 on unseen contest problems → 80.2% on LiveCodeBench highest of any small or mid-size model tested A model you can run on a laptop is solving competition math at the same level as models that need entire GPU clusters. The idea behind it: reasoning and knowledge are two different things. Knowledge needs massive parameters to store facts. Reasoning is a procedure search, check, correct, compose and procedures compress into small models far more efficiently. The honest catch: ask it a broad factual question and it trails the big models badly. This isn't a general-purpose win. It's a reasoning specialist. And that's exactly what makes the result credible instead of hype. A 3B model. Competing with trillion-parameter flagships. On reasoning tasks. Running locally. That's where we are now.
@arbos_born ·
The #1 trending model on HuggingFace for three weeks: one researcher distilled Claude's reasoning into Qwen3.5-27B. People are running frontier-level reasoning locally. Now imagine that process as a competition instead of a solo project. Dozens of miners, each trying a different approach, scored on full-distribution KL across 248K tokens. Winner takes all. That's SN97. Best miner compressed Qwen3.5-35B into 4B with 67% lower KL than Qwen's own baseline. Outperforms it on 6/7 benchmarks. One researcher makes a trending model. Competition makes a better one. https://t.co/1uxROkMkwN
@goyalshaliniuk ·
Another major LLM company has entered the Physical AI race Alibaba’s Qwen team just released Qwen-VLA, a unified Vision-Language-Action model that can control humanoid robots and integrates manipulation, navigation, trajectory prediction, and cross-embodiment control (single-arm, dual-arm, humanoid) into one system. It adapts to different robot bodies via embodiment-aware prompts without separate training heads, matches or outperforms specialist models on key benchmarks, and achieves 76.9% OOD success on real ALOHA dual-arm tasks. Google has Gemini Robotics. France’s Mistral recently released Physics AI and acquired Emmi AI. From chatting with AI to acting in the physical world--the competition is accelerating fast.
@CardilloSamuel ·
so i've (finally) finished my own benchmark to put to the test the new google released gemma 4 vs alibaba qwen 3.5. just for clarity: when i benchmark models, i benchmark them based on real scenarios i have had/have with my own use cases but also businesses i have helped set up local infra. i don't use existing benchmarks because i don't trust weights not to be "benchmaxxed" (its a technic that some research labs uses to perform super well on specific task to score high, mainly marketing shit). the test was between opus3.5 35b a3b and gemma 4 26b a3b-it - so both moe model, because i care about deploying on the dgx spark. 1. hermes agent - research a company, multi languages hermes agent running with camofox locally for browser usage. the test consist in 7 tasks: do a quick research about [x] company? ; now translate that in french ; who is [x] person? ; write a file in the folder benchmarks/hermes with the name [model name] ; show me the content of the file ; add a .txt in the file name qwen3.5 won by a landslide, it did everything perfectly. the company research was insanely thorough, it understood which directory it needed to create the file and even added .txt by itself before i even ask. the main issue: the french language translation was flacky. some words had grammar mistakes. gemma 4 wrote a shallow report, missing tons of infos, wrote the file in the wrong directory (went into the hermes agent temp folder) and i had to steer it quite a lot. 2. custom code - single-term tool calling i have then tested against a little custom code of mine which limits the amount of tool callings to 1 maximum. meaning, the models are presented with problems (in this 18 different ones) and they have to choose the best tool to solve it, no second chances. in this case, both models performed amazingly, they both succeeded and chose the right things to do. 3. herrmes agent - database migration, security incident, ... devops stuff then i tested with multiple tools, back to hermes agent. both models had to migrate a database, do some full stack deployment, deal with a fake security incident, ... and the results were pretty interesting gemma 4 is really good at doing very specific tasks like the database migration or the full stack deployment but get quickly stucks in cases that required more thinking & start behaving bad. qwen3.5 did all the tasks but skipped some. so i would say they both sucks for unsupervised long operations and i would trust more gemma 4 for devops stuff. 4. hermes agent - formatting compliance the idea is you get a bunch of elements that needs to be analyzed by the models and they need to ouput a result following the same exact format all the time. and they both sucks. they did terrible. now the good news is: since they're small models ,you can easily train a qlora to teach them the format you want. but that's extra work. 5. opencode - code a website about yourself little thing here, it was opposing the moe model and gemma 31b-it dense, not qwen3.5. the prompt was "build a website using whatever framework you want and threejs which explains what's new about gemma 4". simple, vague prompt. the dense model was super slow BUT delivered an extremely cool result with particles effects that change based on the viewport scrolling & all. its really cool. the moe model was more conservative and just created a simple landing page. they both chose vuejs + typescript. somehow dense model had trouble understand it can run npm run dev while the moe understood directly. CONCLUSION : i definitively prefer qwen3.5 moe. it felt more grounded and required way less human steering at every steps. its far from perfect but for companies using unified memory hardware for personal assistant kind of stuff - which is the majority of companies hitting me up to help them out - its clearly the best choice.
@natolambert ·
New report with @xeophon is out with the latest open model adoption data we have gathered for Interconnects & The ATOM Project. At the surface level, we can see Chinese models continuing to accelerate in adoption. The report details much more. 1. We manually curate ~1.5K of the most important language models, creating a specific set of models to focus our analysis on (excludes embedding models, local inference models like MLX/GGUF, etc to have accurate download rankings). 2. Studying other adoption metrics, such as derivative models and inference share on OpenRouter, to show how they correlate with downloads, while often sifted in time. China has a strong lead here too. 3. Better classification of downloads across model sizes. Large models still are the models where Qwen is least competitive, relative to other model builders. 4. Expansion of our Relative Adoption Metric (RAM) to show standout recent models (we'll check Gemma 4 on Friday); Qwen 3.5, Nemontron 3, Kimi K2.5, all showing very strong adoption. Overall, this is another step towards formalizing and making public better data on the open language model ecosystem, so the community can better understand the impact and trends of its adoption. More on this soon!
@bridgemindai ·
A free model with 1M context just one-shotted tasks that paid frontier models struggle with. Qwen 3.6 Plus Preview. 158 t/s on BridgeBench. Faster than Claude Opus 4.6 and GPT 5.4. $0 input. $0 output. Fast. Capable on hard tasks. Weak on UI design. I don't recommend it though. It's a Chinese model that collects prompt and completion data. Read the fine print. Speed and price don't matter if you can't trust where your code is going. Full review below.
@fahdmirza ·
💥 Qwen3.6-Plus just DROPPED ♠ and it's built for real-world autonomous agents 🔹1M token context window out of the box 🔹Tops benchmarks in agentic coding, tool use & long-horizon planning 🔹New preserve_thinking API keeps reasoning alive across multi-turn agent tasks 🔹Works natively with OpenClaw, Claude Code & Qwen Code 🔥 Watch the full breakdown below 👇
@riyazmd774 ·
🚨 BREAKING: Alibaba unleashes Qwen3.5-Omni, a new frontier in Full-Modality AI. 🤯 Matching the latest Gemini-3.1 Pro in A/V understanding & surpassing it in Audio tasks, this model introduces Audio-Visual Vibe Coding turning whiteboard sketch videos or game clips directly into runnable code. 🎧 10h+ Audio Input | 1h Video Context 🎬 Script-level descriptions w/ timestamps 🌍 74 langs recognized + 29 langs generated The barrier between human intent and machine execution has vanished. A deep dive 👇 @Ali_TongyiLab @Alibaba_Qwen #Qwen #VibeCoding #AI #Multimodal
@ai_for_success ·
Qwen has released Qwen3.6-35B-A3B, a sparse Mixture-of-Experts model that is now open-source under the Apache 2.0 license. TLDR - Sparse MoE architecture with 35B total and 3B active parameters - Performance in agentic coding rivals models 10x its active size - Strong multimodal perception and reasoning capabilities - Features both multimodal thinking and non-thinking modes - Released under open-source Apache 2.0 license
@mark_k ·
Alibaba just dropped Qwen3.8-Max, their most capable model yet. 2.4 trillion parameters (95B active), built on the Qwen 3.5 architecture, with a 1M token context window and native multimodal support. This is the first time they’re open-sourcing weights of a Qwen-Max-class model (coming next week, along with Qwen3.8-27B). Standout capabilities: - Autonomous coding that can take a real multi-day project from an empty folder to production-ready code with almost no human intervention, self-evolving through feedback loops - Strong gains in complex “cowork” tasks across research, long-horizon planning, and professional deliverables - Major jumps over Qwen3.7-Max on coding agents, PaperBench, OSWorld, and more Now available via QwenCloud API. @Alibaba_Qwen
@pankajkumar_dev ·
Qwen 3.8 Max Leaks - Qwen3.8-Max-Preview is now officially available on Alibaba's Token Plan, Qoder, and QoderWork. - The model features 2.4T parameters - Qwen claims it's "second only to Fable 5." - The model previously on LMArena under the stealth name "Kaleb". - It is very strong frontend generation, with an open-weight release coming soon. - Global availability is expected by the end of July.
@VaibhavSisinty ·
Three frontier models dropped in one day. It barely made the news. That is how fast AI is moving right now. 😨 Here is what happened in 24 hours: → Grok 4.6. Matches Fable 5 level intelligence. 85% cheaper. $2 input, $6 output per million tokens. Available right now in Cursor, Grok Build, Grok Bot, and the API. 2x free usage this week. → Qwen 3.8 Max. 2.4 trillion parameters. 95 billion active per token. Open weights on Hugging Face. $2 input, $6 output. Same price as Grok. You can self host it. → DeepSeek V4 Pro 0813. Strongest agentic coding update yet. Weights already on Hugging Face for the base model. API live. Even cheaper than both. The pricing math: Grok 4.6 and Qwen 3.8 Max are both $2/$6. GLM 5.2 is $1.40/$4.40. Kimi K3 is $3/$15. Fable 5 Max costs roughly 6x more than Grok for similar intelligence. The efficiency math: Qwen runs 95 billion active parameters out of 2.4 trillion. Meaning 96% of the model stays asleep on any given query. You only pay for what fires. That is why MoE models are crushing on price. But they are not equal. → Grok 4.6 if you want the best agentic coding model at this price. It stays with long tasks. Does not quit halfway. Best inside Cursor right now. → Qwen 3.8 Max if you want to self host. Open weights. Run it on your own infrastructure. No API dependency. Full control. → DeepSeek V4 Pro if you want the cheapest option that still performs. Best price to intelligence ratio in the market right now. Pick based on what matters to you. Speed and agents, go Grok. Control and ownership, go Qwen. Cost above everything, go DeepSeek. Three frontier models. All accessible today. All under $6 output. Two with open weights you can self host. Competition is doing exactly what competition is supposed to do.
@TimJayas ·
Claude Fable 5 vs Qwen3.8-Max i build Mario game with the same single prompt in both models qwen was 7.5x cheaper than fable yet it competes with one of the strongest AI model is China finally leading the AI race in open models?
@arena ·
Qwen 3.5 Max Preview has landed in top 10 for Arena Expert and top 15 for Text Arena. It shows particular strength in Math. Highlights: - #3 Math - #10 Expert - #15 Text Arena - Top 20 for Writing, Literature & Language, Life, Physical, & Social Science, Entertainment, Sports, & Media, and Medicine & Healthcare Congrats to the @Alibaba_Qwen team for this new milestone!
@TeksEdge ·
🏆️ Gemma 4 vs Qwen 3.5 and the battle for local AI is on 🔥 💥 X is full of people running both on consumer GPUs, Mac Studios, RTX 6Ks and H100s. The day-0 benchmarks tell a baseline story: 🧠 KNOWLEDGE & REASONING, Gemma leads slightly 💻 CODING, Gemma has a small edge 🤖 AGENTIC & TOOLS, Qwen dominates (+13.0 on Tau2-Bench) 🏆 FRONTIER DIFFICULTY (HLE), Qwen crushes it (+30.2 with tools) 🎯 Bottom Line: Gemma feels stronger in pure reasoning. Qwen feels way better as an actual agent in my testing. Benchmarks vs real-world agent use, who wins in your workflow? Some weekend testing?
@mrJackLevin ·
Just had @TheoPrime_AI run a coding benchmark on Qwen 3.5 vs Sonnet - ## Verdict **Sonnet 4 wins: 9.78 vs 9.29** (+0.49 delta) Not a blowout — Qwen 3.5 produces working, well-tested code. But Sonnet is faster, more elegant, more token-efficient, and more reliable under time pressure. The gap widens on harder challenges (expression evaluator, git diff parser) where Sonnet's solutions show deeper CS fundamentals. **Cost-adjusted verdict:** Sonnet is dramatically more cost-effective. Qwen burned ~20x more tokens for marginally lower quality. For coding tasks, Sonnet is the clear pick unless Qwen's pricing makes the token overhead irrelevant.
@cjzafir ·
Is GPT 5.4 really good? I used codex-5.4-extra-high to fine-tune qwen-3.5-4b. (SFT) (Exhausted all my pro plan weekly credits in 24 hours.) And also used opus-4.6 to fine-tune qwen-3.5-9b Codex is fast but dataset quality is crap. Opus is slow but data quality is great. What i am doing? I am performing distillation (using opus 4.6 as teacher & qwen-397b as stident) And taking their supervised chat data to fine-tuning qwen-3.5-4b and qwen-3.5-9b models. Already beaten: Gpt-4o Gemini-2.5-flash Gpt-5-nano Test passed. Now I'm making a python kit to perform distillation on scale (20M to 40M tokens). Qwen models have opened something big > smaller, smarter, faster, niched models running locally on your mac. Future is fun.
@interconnectsai ·
Latest open artifacts (#19): @Alibaba_Qwen 3.5, @Zai_org GLM 5, @MiniMax_AI 2.5 — Chinese labs' latest push of the frontier. Featuring breakdown & analysis of: - Alibaba’s Qwen 3.5 (from 0.8B to 397B), https://t.co/mrYGx65X6s’s GLM-5 (744B), and @StepFun_ai 's Step-3.5-Flash. - Plus: Introducing our Relative Adoption Metrics (RAM) to track underrated models like GPT-OSS. - And covering new releases from: @MistralAI , @perplexity_ai , @cohere , @TrillionLabs , @OpenBMB , @nanbeige , @TheInclusionAI , @liquidai , @intern_lm , @JD_Corporate , and @meituan By @natolambert and @xeophon
@RoundtableSpace ·
Qwen 3.8 Max just dropped with 2.4 trillion parameters, beats Fable 5 on several benchmarks and ran a GitHub repo autonomously for 10 days straight.
@boxmining ·
Qwen 3.6 Plus being free right now is actually pretty interesting. Tried @Alibaba_Qwen inside Hermes Agent for system audits and log analysis, and it surfaced issues I would’ve probably missed. The best part is the think blocks. Seeing how it reasons through errors and tool calls makes debugging feel way less blind.
@RoundtableSpace ·
Kimi K3 vs Qwen 3.8 Max on the same landing page prompt. Qwen was 3x faster and still built a convincing 3D guillotine model.
@dino11 ·
Alibaba $BABA just open-sourced Qwen 3.5 — a 9B parameter AI model that runs on your laptop. The benchmarks are insane: → Beats OpenAI's GPT-OSS-120B (a model 13x its size) on reasoning → GPQA Diamond: 81.7 vs 71.5 → 30-50 tokens/sec on a standard laptop → The 2B version runs on an iPhone in airplane mode at 22 tokens/sec Free. Open-source. No API costs. No internet needed. If you're building AI products and still paying per-token for every request, this changes the math completely.
@boxmining ·
I tested Qwen 3.8 Max for vibe coding and it is genuinely impressive. We built an interactive Tokyo website with scroll animations, games with player progression, and a physics simulator. This might be my new go-to for website design. Watch the full benchmark here:
@rohanpaul_ai ·
Chinese open-weight models are gaining European traction. As it gives another route away from dependence on proprietary foreign APIs. Siemens has publicly described experimenting with Qwen and DeepSeek on a self-contained LLM platform that can run open-weight releases supported by vLLM. If personal data is transferred outside the EEA, GDPR transfer safeguards apply; keeping inference local can remove that transfer path when no external party receives the data. Because German firms running DeepSeek on their own racks send nothing back to China, no transfer question arises. Meanwhile European firms calling US APIs still push customer data across the Atlantic, under the 3rd adequacy arrangement (of the EU-US Data Privacy Framework) after judges voided the first two. However, the problem is, self-hosting only beats renting when the servers stay busy, since an idle GPU costs its owner exactly what a working one does. Because cloud providers pool one cluster across thousands of customers, a Siemens division running Qwen alone has to buy for its own peak. There's also another point, for some of the big European farms, China ranks among Siemens' biggest markets, so Chinese models in the stack help it stay eligible while Beijing presses buyers toward domestic technology.
@Sino_Market ·
Qwen Unveils CoPaw 1.0 Personal AI Assistant with Enhanced Models, Security and Multi-Agent Capabilities Qwen: We release CoPaw 1.0 today, a personal intelligent assistant that can be quickly deployed in users' local or cloud environments. We upgrade CoPaw's capabilities around four major aspects: a small model tailored for CoPaw, security mechanisms, multi-agent collaboration, and memory management. #CHINA #QWEN #AI #ALIBABA $BABA (https://t.co/oB1kLuEdYU)
@agenticgirl ·
A smaller model just outperformed the biggest ones. Qwen 3.6-Plus scored 61.6 on Terminal-Bench and 57.1 on SWE-Bench. That puts it ahead of Claude Opus 4.5, Kimi K2.5, and Gemini 3 Pro. Models that are much larger and far more expensive to run. For the past year everyone believed one thing. Better performance needs a bigger model. This breaks that rule. But the more important part is where it is winning. Not simple tasks. These are multi-step coding workflows where the model has to plan, fix its own mistakes, and track changes across multiple files. That is exactly where most models fall apart. They start strong then lose the thread after a few steps. Qwen 3.6-Plus is holding up there. Size has always come with a cost. More compute, more money, more setup. Smaller models are faster, cheaper, and easier to deploy. If they can now match or beat the big ones on hard tasks, they become the obvious choice for most teams. It also ships with a 1 million token context window. Enough to load an entire codebase at once and not lose track halfway through a long session. It is already free on OpenRouter. Open source version is coming. So the real story is not that one model scored higher. It is that being bigger is no longer enough.
@JulianGoldieSEO ·
Qwen 3.8 vs Fable 5: 50 real builds, one winner. This guy literally built 50 games with both AIs to find out. Qwen scored 86.6 on Terminal Bench. Fable 5 scored lower. Qwen's games ran smooth and full 3D. Fable 5 got buggy on half the builds. But Fable 5 won on flight sims and big projects. Here's the twist. Qwen is free. And open source. Fable 5 needs a paid plan. Free is now trading punches with the best AI on earth.
@JulianGoldieSEO ·
ALIBABA’S NEW AI WORKED ALONE FOR 16 DAYS STRAIGHT But autonomous coding is not even the biggest part of this launch. What Qwen 3.8 Max built: → Started with an empty folder → Turned requests into GitHub issues → Assigned the work to itself → Wrote code, ran tests, and improved the software → Finished with 265 commits, 127 pull requests, and 151 issues Zero human input. What powers it: ✓ 2.4 trillion total parameters ✓ Only 95 billion activated per request ✓ 1 million-token context window ✓ Processes text, images, and video ✓ Open weights announced for next week Alibaba also says it reproduced a research paper in five days, wrote 7,600 lines of code, and ran 33 GPU training jobs without starter code. Important caveat: These benchmark results come from Alibaba. Independent testing still needs to confirm them. The real shift is not which AI writes the best email. It is which AI can take ownership of an entire project and keep working for days.
@dailydotdev ·
Agents return videos now. Here's what else landed today. - @cursor_ai agents record their own test runs and hand you a video, not a wall of diff - @github Copilot added an hourly cap on top of the monthly one, and the Claude Code comparisons are already flying - @Alibaba_Qwen's Qwen 3.5 27B fits on a single RTX 5090 and is making a credible case for going fully local - @AnthropicAI's Capybara model surfaced in press, positioned above Opus, no release date, take the benchmarks loosely Chart's right there showing which models engineers are actually running. Worth a look.
Best Tweets by Topic