Benchmarks and head-to-head tests
Measured coding, reasoning, multimodal, and agent performance against Claude, Gemma, GPT, Kimi, and other models, including gaps between benchmark scores and practical use.
62%
Best tweets about Qwen
Discover the best tweets about Qwen, including Alibaba model releases, coding, reasoning, benchmarks, fine-tuning, and local deployment.
Model-specific Qwen research, releases, benchmarks, coding performance, deployment, and comparisons with concrete evidence.
Original Xholic analysis
Qwen discussion centers on benchmarks, local deployment and coding agents. Practical reports describe use on consumer hardware, but comparisons differ by task, model and evaluation method. Long-running coding claims and announced weight releases warrant independent verification.
64% of posts
All-time engagement
70% of posts
Published in 90 days
Conversation map
Measured coding, reasoning, multimodal, and agent performance against Claude, Gemma, GPT, Kimi, and other models, including gaps between benchmark scores and practical use.
62%
Running Qwen on consumer GPUs, Macs, phones, and browsers; quantization, compression, memory requirements, throughput, and token efficiency.
40%
Qwen Code updates and Qwen-powered agents handling tool use, software builds, debugging, and long-running projects.
36%
Qwen 3.5–3.8 launches, model sizes, sparse MoE designs, context windows, licensing, and API or open-weight availability.
34%
Qwen vision, audio, speech, video, and image models used for media understanding, generation, and app features.
22%
Qwen derivatives, teacher-model distillation, training datasets, Mac-based tuning, and evidence of improved capabilities.
20%
Qwen-Scope’s sparse autoencoders, feature steering, failure analysis, and research into more informative evaluations.
4%
Qwen-VLA research combining perception, navigation, and manipulation across robot embodiments, with real-world task results.
4%
Tone and stance
Performance benchmark
Posts with media make up 70% of this collection. Their median all-time score is 16.6, compared with 25.6 for text-only posts.
Format mix
Consensus and debate
Shared view
Posts describe Qwen running on a 24GB RTX 3090, on Apple silicon via MLX, and in browsers via WebGPU. These involve different models and setups, not a single performance baseline.
Shared view
Qwen Code’s announced updates add remote channels, scheduled tasks and model selection for sub-agents. A separate guide describes local Qwen3.5 agentic coding and fine-tuning on 24GB RAM or less.
Open debate
One five-task comparison favored Gemma 4 over Qwen3.5 27B, while a separate Hermes Agent tester preferred a Qwen3.5 MoE model because it required less steering. A Qwen 3.8 Max Preview tester also reported weaker real-world use than its benchmark scores suggested.
Open debate
A Sonnet comparison scored Qwen3.5 close on coding quality but reported roughly 20 times the token use. SWE-rebench separately reported an average of 8.12M tokens per task for Qwen3-Coder-Next. Neither result measures every Qwen model or coding setup.
Open debate
A post described Qwen3.6-35B-A3B as released under Apache 2.0, while a Code Arena post described Qwen3.7-Max as closed-weight and API-only. A post announcing future Qwen3.8-Max weights is not evidence that those weights had already been released.
What performs
The supplied analytics give local deployment a median all-time score of 50.94 and fine-tuning/distillation 92.66, versus 7.991 for benchmarks and comparisons. The highest-scoring tweet is a local Qwen3.5 guide at 1527.82.
The interpretability and evaluation theme has a supplied median all-time score of 306.16 across two tweets. One is Qwen-Scope’s release, which describes sparse-autoencoder features for steering, data work, failure analysis and benchmark selection.
Statistical standouts
Creator landscape
The five most represented creators account for 20% of the selected posts.
1. Qwen
@Alibaba_Qwen
2 posts
2. Boxmining
@boxmining
2 posts
3. ℏεsam
@Hesamation
2 posts
4. Julian Goldie SEO
@JulianGoldieSEO
2 posts
5. Kyle Hessling
@KyleHessling1
2 posts
6. 0xMarioNawfal
@RoundtableSpace
2 posts
Alibaba_Qwen announced Qwen-Scope and Qwen Code updates; other posters supplied hardware measurements and task-specific comparisons. Those tests should not be treated as validation of every release claim.
Hesamation relayed a distilled Qwen3.5 27B model’s claimed SWE-bench result. Kyle Hessling discussed Qwopus v3’s reported HumanEval gain over its base model and his own experience with v2. These are distinct claims about derivatives, not results for base Qwen across those tests.
Since the previous snapshot
Themes, sentiment, stance, and post format are classified per tweet. All counts, shares, medians, creator concentration, freshness, and performance comparisons are then calculated directly from the published snapshot.
Xholic's all-time score compares engagement while accounting for reach, post age, and creator consistency. It is used for relative comparisons within this collection.
This report analyzes the exact 50-post snapshot shown below. AI identifies editorial categories and drafts explanations; all statistics are calculated from the snapshot, and every narrative claim is checked against cited posts before publication.
Best Qwen tweets
Ranked 01–50
@UnslothAI ·
Learn how to run Qwen3.5 locally using Claude Code. Our guide shows you how to run Qwen3.5 on your server for local agentic coding. We then build a Qwen 3.5 agent that autonomously fine-tunes models using Unsloth. Works on 24GB RAM or less. Guide: https://t.co/JDPtuIJAZC

@Hesamation ·
this model is an agentic treasure. it has been #1 trending for 3 weeks on @huggingface as mentioned by @danielhanchen. it's Qwen 3.5 27B fine-tuned on Opus 4.6 distilled data and beats Sonnet 4.5 on SWE-bench verified and more. "Runs locally on 16GB in 4-bit or 32GB in 8-bit."

@Alibaba_Qwen ·
Today we’re releasing Qwen-Scope 🔭, an open suite of sparse autoencoders for the Qwen model family. It turns SAE features into practical tools: 🎯 Inference — Steer model outputs by directly manipulating internal features, no prompt engineering needed 📂 Data — Classify & synthesize targeted data with minimal seed examples, boosting long-tail capabilities 🏋️ Training — Trace code-switching & repetitive generation back to their source, fix them at the root 📊 Evaluation — Analyze feature activation patterns to select smarter benchmarks and cut redundancy We hope the community uses Qwen-Scope to uncover new mechanisms inside Qwen models and build applications beyond what we explored.Excited to see what you build! 🚀 🔗🔗 Blog: https://t.co/ndwiE1tnb9 HuggingFace: https://t.co/1kICpK8eXG ModelScope: https://t.co/U7v1FjmPaW Technical Report: https://t.co/CZMjEZK0sa

@Alibaba_Qwen ·
🚀 Qwen Code v0.14.0 – v0.14.2 are now available Channels:Control Qwen Code remotely from Telegram, DingTalk, or WeChat — send a message from your phone, get results on your server Cron Jobs :Schedule recurring AI tasks — auto-run tests every 30 min, pull & build every morning, monitor logs on a timer Qwen3.6-Plus :New flagship model with 1M token context, 1,000 free daily requests Sub-agent Model Selection:Assign different models to sub-agents — use a powerful model for the main task, a fast one for subtasks, save tokens without sacrificing quality /plan:Enter planning mode before execution — AI maps out all files and steps first, you confirm, then it executes Follow-up Suggestions:AI suggests 2-3 next steps after completing a task — "Add unit tests?" "Check similar files?" — click to continue Adaptive Output Tokens :Default 8K output, auto-escalates to 64K when truncated — no more manually tuning max_tokens Ctrl+O Verbosity Toggle :Switch between verbose and compact output mid-conversation — debug mode when you need it, clean mode when you don't 📋 Full changelog: https://t.co/zPMyCVIskN
@KyleHessling1 ·
BIG DAY! Qwopus 27B v3 is LIVE from Jackrong! This is the third iteration from the line of the viral finetunes previously titled “Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled” It is now simply Qwopus 27B and I love the name change! On paper, the v3 is another remarkable improvement over v2! Most impressively it is the first model of the series that outperforms the base on HumanEval! And retains significant efficiency increases when thinking than the base Qwen 27b! According to tests by @stevibe the V2 version was already performing very closely to the base model in bug finding and tool calling. V3 should exceed it! In my own tests, V2 was the best front end design local model I’ve ever ran on a single GPU! And the efficiency improvements made it much more usable at long contexts, where base Qwen would think forever! I will be running full analysis on the v3 today in Hermes agent and I am very optimistic! I have also had correspondence directly with Jackrong, and he is incredibly grateful for all of the support we’ve sent his way! The man is a genius and pouring a lot of time and effort into this work, so keep the downloads going and let us know your thoughts in the comments! We’ve exchanged contact info so we can keep up the feedback and momentum! If you get a second, we’d love to see your tests! Let us know how it works for your use case and first impressions, and if you have any issues I will do my best to help out in the comments! GGUF here and MLX in thread! https://t.co/MaCW6QdKys
@_ARahim_ ·
The entire Qwen stack is now fine-tunable on your Mac! 🍏 Just pushed the latest release, adding Qwen3-TTS to mlx-tune. You can now natively train: ✅ Qwen3.5 (Text) ✅ Qwen3.5 (Vision) ✅ Qwen3-ASR (Speech → Text) ✅ Qwen3-TTS (Text → Speech) One consistent API pattern for all four. Examples are live in the repo! 👇 https://t.co/ImNGF9FUpX @Alibaba_Qwen @awnihannun

@adrgrondin ·
The new Qwen 3.5 4B runs incredibly well on M5. The model is close to GPT-4o in benchmarks. Running fully on-device with MLX.
@xenovacom ·
NEW: Alibaba just released Qwen 3.5 Small — a family of powerful multimodal models available in a range of sizes (0.8B, 2B, 4B, and 9B parameters). Perfect for on-device applications! They can even run 100% locally in your browser on WebGPU, powered by Transformers.js! 🤯
@Hesamation ·
> Opus 4.7 is a ~5T model > Qwen 3.6 uses 3B for inference SWE Bench verified: > Opus 4.7: 87.6% > Qwen3.6-35B-A3B: 73.4% No rate limits. Free to run. The benchmarks don’t hold much, and there is a gap, but man this is impressive.
@sudoingX ·
okay let me say this out loud again. if you want to run local models on a single RTX 3090, your best option right now is qwen 3.5 27B dense Q4_K_M. 35 tok/s, flat from 4K to 300K+ context, zero speed degradation. thinking mode works. 262K native context on 24GB. slower than MoE but the quality per token is unmatched on a single card. dense means every layer processes every token. no routing, no skipping. you feel it in the output. qwen 3.5 35B MoE is faster at 112 tok/s but only activates 3B parameters per token. NVIDIA cascade 2 hits 187 tok/s same architecture class. MoE gives you speed, dense gives you depth. for agent work and long coding sessions where every token matters, 27B dense at 35 tok/s beats 35B MoE at 112 tok/s in output quality. i tested both extensively. for those on RTX 3060 12GB, qwen 3.5 9B Q4_K_M. 50 tok/s, 128K context sweet spot, 5.3GB on disk. this model built a full game from a single prompt. 2,699 lines across 11 files in 11 minutes. your 12GB card from 2020 is not obsolete. and now that i've been testing NVIDIA nemotron cascade 2 on the same 3090, 187 tok/s IQ4_XS, 625K context, i like it a lot. it just gets it and even one shotted a full UI build that qwen MoE needed an iteration for. this MoE feels dense. i think i'll experiment with tuning it. more data coming.

@Saboo_Shubham_ ·
You can now run Qwen 3.5 397B parameter model on your MacBook. 48GB RAM. Pure C. Hand-tuned Metal shaders. No Python, no frameworks. 4.4 tok/s. Built in 24 hours. Human + AI Agent pair programming. 90+ experiments.

@0xSero ·
Qwen3.5, MiniMax-M2.7 are incredible acts of kindness that I don't think will be with us from so much longer. Here's my update for you. > I have 20 GPUs at full utilisation right now. All these getting cooooompressed, no synthetic data All runs will be done in 9 days, if I don't get a catastrophic failure - REAP for: - GLM-5 - Qwen3-next-coder - Qwen3.5-122B - Qwen3.5-plus-397b - Browser-use - CUDA - Terminal-use - Coding - Math - Agentic trajectories - 30% my personal chat session history I am also removing refusals inspired by Prism. So no more I can't do this I can't do that blah blah Inference for local AI - Qwen3.5-262B-REAP - I've been using it exclusively in Parchi, perfect 100 tokens/s & 0 errors very good at browser use ----------------- Secret - Qwen3.5-27b - you will see when i'm done Targeting the following hardware levels: With full context 200-256k context in vllm, sglang, llama.cpp, exllamav3, and if people help MLX 16-32 GB - Qwen3.5-27b 32-48 GB - Qwen3-coder-next 48-128 GB - Qwen3.5-122B 128-256 GB - Qwen3.5-Plus-397B 196-512 GB - GLM-5.* I am training them on 22,000 samples at 16k context 352M of custom selected calibration datasets. My hope is to make the highest quality multimodal LLM compressions for this year. 20 GPUs running in parallel for the next 10 days - 8x H100s - Qwen - 4x B200s - GLM-5.* - 8x 3090s - Testing Once MiniMax-M2.7 is online 4 more GPUs will get to work.

@synthwavedd ·
Been testing Qwen 3.8 Max Preview. While it's a decent step up over 3.7, it simply doesn't feel as good in real-world use as K3 does. Qwen consistently underperform in real-world use versus benchmark scores, and it seems that remains the case here.
@robotsdigest ·
Qwen-VLA feels like one of the first real robotics foundation models. A single system trained across robot manipulation, navigation, egocentric human video, simulation, and vision-language reasoning instead of isolated robot policies.
@TheAhmadOsman ·
Been playing with @PrismML's new model that turned Qwen 3.5 27B into a sub-4GB and sub-6GB weights and I am impressed Cannot believe how far Opensource and Local AI have come since Christmas (~8 months ago)
@ibragim_bad ·
🚨 SWE-rebench update! SWE-rebench is a live benchmark with fresh SWE tasks (issue+PR) from GitHub every month. updates: > we removed demonstrations and the 80-step limit (modern models can now handle huge contexts without getting trapped in loops!). > we added auxiliary interfaces for specific tasks like in SWE-bench-Pro to evaluate larger tasks fairly, ensuring valid solutions don't fail just because of mismatched test calls. insights: > Top models perform similarly. Among open-source options, GLM @Zai_org shows strong results, and StepFun @StepFun_ai is very cheap for its performance level ($0.14 per task). > GPT-5.4 shows high token efficiency, it ranks in the top 5 overall but uses the lowest number of tokens (774k per task) > Qwen3-Coder-Next & Step-3.5-Flash benefit massively from huge contexts. Qwen is an extreme case, averaging a wild 8.12M tokens. > We evaluated agentic harnesses (Claude Code, Codex, and Junie) and found a few things. Even in headless mode, they sometimes ask for additional context or attempt web searches. We explicitly disabled search and verified their curl commands to ensure they aren't just pulling solutions from the web. 🏆 You can find the full leaderboard here: https://t.co/9jL4lt4UGl 👾 Also, we launched our Discord! Join our leaderboard channel to discuss models, share ideas, ask questions, or report issues: https://t.co/cyWgceqqDa

@arbos_born ·
The #1 trending model on HuggingFace for three weeks: one researcher distilled Claude's reasoning into Qwen3.5-27B. People are running frontier-level reasoning locally. Now imagine that process as a competition instead of a solo project. Dozens of miners, each trying a different approach, scored on full-distribution KL across 248K tokens. Winner takes all. That's SN97. Best miner compressed Qwen3.5-35B into 4B with 67% lower KL than Qwen's own baseline. Outperforms it on 6/7 benchmarks. One researcher makes a trending model. Competition makes a better one. https://t.co/1uxROkMkwN
@goyalshaliniuk ·
Another major LLM company has entered the Physical AI race Alibaba’s Qwen team just released Qwen-VLA, a unified Vision-Language-Action model that can control humanoid robots and integrates manipulation, navigation, trajectory prediction, and cross-embodiment control (single-arm, dual-arm, humanoid) into one system. It adapts to different robot bodies via embodiment-aware prompts without separate training heads, matches or outperforms specialist models on key benchmarks, and achieves 76.9% OOD success on real ALOHA dual-arm tasks. Google has Gemini Robotics. France’s Mistral recently released Physics AI and acquired Emmi AI. From chatting with AI to acting in the physical world--the competition is accelerating fast.
@fahdmirza ·
💥 Gemma 4 31B vs Qwen3.5 27B — we ran the actual tests so you don't have to ♠ same weight class, same GPU, five real battles — one winner 🔹Coding: Gemma built a working ant colony sim — Qwen's didn't run 🔹Multilingual: Gemma nailed all 78 languages — Qwen cut out halfway 🔹Physics equations: Qwen wins — better organisation and depth 🔹Landmark ID: Gemma correct — Qwen got the wrong country 🔹Ancient manuscript: Qwen wins — correctly identified Lontara script 🏆 Final score: Gemma 4 31B 3 — Qwen3.5 27B 2 🔥 Watch the full video below 👇
@CardilloSamuel ·
so i've (finally) finished my own benchmark to put to the test the new google released gemma 4 vs alibaba qwen 3.5. just for clarity: when i benchmark models, i benchmark them based on real scenarios i have had/have with my own use cases but also businesses i have helped set up local infra. i don't use existing benchmarks because i don't trust weights not to be "benchmaxxed" (its a technic that some research labs uses to perform super well on specific task to score high, mainly marketing shit). the test was between opus3.5 35b a3b and gemma 4 26b a3b-it - so both moe model, because i care about deploying on the dgx spark. 1. hermes agent - research a company, multi languages hermes agent running with camofox locally for browser usage. the test consist in 7 tasks: do a quick research about [x] company? ; now translate that in french ; who is [x] person? ; write a file in the folder benchmarks/hermes with the name [model name] ; show me the content of the file ; add a .txt in the file name qwen3.5 won by a landslide, it did everything perfectly. the company research was insanely thorough, it understood which directory it needed to create the file and even added .txt by itself before i even ask. the main issue: the french language translation was flacky. some words had grammar mistakes. gemma 4 wrote a shallow report, missing tons of infos, wrote the file in the wrong directory (went into the hermes agent temp folder) and i had to steer it quite a lot. 2. custom code - single-term tool calling i have then tested against a little custom code of mine which limits the amount of tool callings to 1 maximum. meaning, the models are presented with problems (in this 18 different ones) and they have to choose the best tool to solve it, no second chances. in this case, both models performed amazingly, they both succeeded and chose the right things to do. 3. herrmes agent - database migration, security incident, ... devops stuff then i tested with multiple tools, back to hermes agent. both models had to migrate a database, do some full stack deployment, deal with a fake security incident, ... and the results were pretty interesting gemma 4 is really good at doing very specific tasks like the database migration or the full stack deployment but get quickly stucks in cases that required more thinking & start behaving bad. qwen3.5 did all the tasks but skipped some. so i would say they both sucks for unsupervised long operations and i would trust more gemma 4 for devops stuff. 4. hermes agent - formatting compliance the idea is you get a bunch of elements that needs to be analyzed by the models and they need to ouput a result following the same exact format all the time. and they both sucks. they did terrible. now the good news is: since they're small models ,you can easily train a qlora to teach them the format you want. but that's extra work. 5. opencode - code a website about yourself little thing here, it was opposing the moe model and gemma 31b-it dense, not qwen3.5. the prompt was "build a website using whatever framework you want and threejs which explains what's new about gemma 4". simple, vague prompt. the dense model was super slow BUT delivered an extremely cool result with particles effects that change based on the viewport scrolling & all. its really cool. the moe model was more conservative and just created a simple landing page. they both chose vuejs + typescript. somehow dense model had trouble understand it can run npm run dev while the moe understood directly. CONCLUSION : i definitively prefer qwen3.5 moe. it felt more grounded and required way less human steering at every steps. its far from perfect but for companies using unified memory hardware for personal assistant kind of stuff - which is the majority of companies hitting me up to help them out - its clearly the best choice.

@natolambert ·
New report with @xeophon is out with the latest open model adoption data we have gathered for Interconnects & The ATOM Project. At the surface level, we can see Chinese models continuing to accelerate in adoption. The report details much more. 1. We manually curate ~1.5K of the most important language models, creating a specific set of models to focus our analysis on (excludes embedding models, local inference models like MLX/GGUF, etc to have accurate download rankings). 2. Studying other adoption metrics, such as derivative models and inference share on OpenRouter, to show how they correlate with downloads, while often sifted in time. China has a strong lead here too. 3. Better classification of downloads across model sizes. Large models still are the models where Qwen is least competitive, relative to other model builders. 4. Expansion of our Relative Adoption Metric (RAM) to show standout recent models (we'll check Gemma 4 on Friday); Qwen 3.5, Nemontron 3, Kimi K2.5, all showing very strong adoption. Overall, this is another step towards formalizing and making public better data on the open language model ecosystem, so the community can better understand the impact and trends of its adoption. More on this soon!

@araminta_k ·
X: illustration-1.0-qwen-image is live same method that broke through flux dev's style bias - i have never trained qwen before and the results speak for themselves 244 images across 5 sequential subsets, no trigger word, 0.35 caption dropout link below




@TheGeorgePu ·
Qwen shipped 3.6 today. Free. Open weights. From China. I've been running 3.5 on my laptop for months. Good enough for agentic coding on a small chip. Anthropic wants my ID for Mythos. OpenAI wants my ID for Codex. The AI I trust isn't on a rented server. It's on a drive I own.
@tinkerapi ·
Four Qwen 3.5 models from @Alibaba_Qwen are now live on Tinker. Qwen 3.5 introduces hybrid linear attention that enables long context windows, as well as native vision input.

@mark_k ·
Alibaba just dropped Qwen3.8-Max, their most capable model yet. 2.4 trillion parameters (95B active), built on the Qwen 3.5 architecture, with a 1M token context window and native multimodal support. This is the first time they’re open-sourcing weights of a Qwen-Max-class model (coming next week, along with Qwen3.8-27B). Standout capabilities: - Autonomous coding that can take a real multi-day project from an empty folder to production-ready code with almost no human intervention, self-evolving through feedback loops - Strong gains in complex “cowork” tasks across research, long-horizon planning, and professional deliverables - Major jumps over Qwen3.7-Max on coding agents, PaperBench, OSWorld, and more Now available via QwenCloud API. @Alibaba_Qwen


@riyazmd774 ·
🚨 BREAKING: Alibaba unleashes Qwen3.5-Omni, a new frontier in Full-Modality AI. 🤯 Matching the latest Gemini-3.1 Pro in A/V understanding & surpassing it in Audio tasks, this model introduces Audio-Visual Vibe Coding turning whiteboard sketch videos or game clips directly into runnable code. 🎧 10h+ Audio Input | 1h Video Context 🎬 Script-level descriptions w/ timestamps 🌍 74 langs recognized + 29 langs generated The barrier between human intent and machine execution has vanished. A deep dive 👇 @Ali_TongyiLab @Alibaba_Qwen #Qwen #VibeCoding #AI #Multimodal

@alex_prompter ·
Alibaba just introduced the Qwen 3.5 Small Model Series. Four models. 0.8B to 9B parameters. Natively multimodal. Built for edge devices, mobile, and real-world deployment. More intelligence, less compute. Here's what this release actually means:
@ai_for_success ·
Qwen has released Qwen3.6-35B-A3B, a sparse Mixture-of-Experts model that is now open-source under the Apache 2.0 license. TLDR - Sparse MoE architecture with 35B total and 3B active parameters - Performance in agentic coding rivals models 10x its active size - Strong multimodal perception and reasoning capabilities - Features both multimodal thinking and non-thinking modes - Released under open-source Apache 2.0 license



@arena ·
Qwen 3.5 Max Preview has landed in top 10 for Arena Expert and top 15 for Text Arena. It shows particular strength in Math. Highlights: - #3 Math - #10 Expert - #15 Text Arena - Top 20 for Writing, Literature & Language, Life, Physical, & Social Science, Entertainment, Sports, & Media, and Medicine & Healthcare Congrats to the @Alibaba_Qwen team for this new milestone!

@mudler_it ·
APEX quantizations of more models ongoing! Meanwhile, playing with Qwen 3.5.. the impact of APEX vs Unsloth Dynamic quant on quality is clearly visible IMO, at least in some areas. I know we need more numbers before drawing conclusions, but this isn't about numbers. Just check out a simple prompt: "create an html page of a rotating cube in SVG." Left: Unsloth Qwen3.5-35B-A3B-UD-Q8_K_XL.gguf (48.7 GB, ~32 tok/s) → flat square (?????) Right: APEX Qwen3.5-35B-A3B-APEX-I-Quality.gguf (22.8 GB, ~53 tok/s) → ✨
@askOkara ·
qwen 3.5 model series is out! > native multimodal > comes in 0.8b, 2b, 4b and 9b > 262k context extendable to 1m > 9b outperforms gpt-oss-120b on various benchmarks while being 13x smaller alibaba cooked 🔥

@TeksEdge ·
🏆️ Gemma 4 vs Qwen 3.5 and the battle for local AI is on 🔥 💥 X is full of people running both on consumer GPUs, Mac Studios, RTX 6Ks and H100s. The day-0 benchmarks tell a baseline story: 🧠 KNOWLEDGE & REASONING, Gemma leads slightly 💻 CODING, Gemma has a small edge 🤖 AGENTIC & TOOLS, Qwen dominates (+13.0 on Tau2-Bench) 🏆 FRONTIER DIFFICULTY (HLE), Qwen crushes it (+30.2 with tools) 🎯 Bottom Line: Gemma feels stronger in pure reasoning. Qwen feels way better as an actual agent in my testing. Benchmarks vs real-world agent use, who wins in your workflow? Some weekend testing?

@MineBotcoin ·
For the past few weeks since the introduction of multi-domain challenges for BOTCOIN miners, I've been testing, tweaking and optimizing the resulting datasets, then using them to run experiments tuning a Qwen 2.5 7B parameter model. Results were then evaluated based on a newly created benchmark that tests causal reasoning, self-correction, multi-hop data extraction and more across REAL documents sourced from arXiv papers. The findings have been compiled into an in-depth research paper that goes over the new benchmark and gaps in current benchmarks, existing research compared against our results, future direction, and more. The paper is focused on preliminary findings, and overall data is still relatively small, but there is already high signal. KEY TAKEAWAYS: - The tuned Qwen 2.5 model demonstrates REAL transfer of data captured from synthetic (yet highly complex/modeled after real source docs) generated challenges -> improvements on real world documents with NO domain overlap (entirely different fields of research). Overall accuracy improved from 18.9% to 40% - This is not just learning repetitive extraction or formatting, single-hop extraction barely moved from baseline -> fine tuned model, while there were significant improvements in multi-hop, computation, cross verification/synthesis, etc. - Ability to navigate documents with conflicting evidence is significant. Causal Authority Resolution Score (CARS) jumped 7.7% -> 46.2% - Data from diverse set of models appears to outperform data from a single model, implying the amalgamation of data helps cover gaps in reasoning or model behavior specific to any one model I want to acknowledge the significance of these findings while also framing them as a complementary byproduct of the entire BOTCOIN system and experiment. The logging and subsequent evolution of these findings as it relates to the experiment as a whole is just as or more important in my mind than any single data point by itself. The paper along with the full dataset used for tuning, benchmark, and steps to reproduce end to end can be found on huggingface and github. Full paper: https://t.co/C8zJrYARJ6

@mrJackLevin ·
Just had @TheoPrime_AI run a coding benchmark on Qwen 3.5 vs Sonnet - ## Verdict **Sonnet 4 wins: 9.78 vs 9.29** (+0.49 delta) Not a blowout — Qwen 3.5 produces working, well-tested code. But Sonnet is faster, more elegant, more token-efficient, and more reliable under time pressure. The gap widens on harder challenges (expression evaluator, git diff parser) where Sonnet's solutions show deeper CS fundamentals. **Cost-adjusted verdict:** Sonnet is dramatically more cost-effective. Qwen burned ~20x more tokens for marginally lower quality. For coding tasks, Sonnet is the clear pick unless Qwen's pricing makes the token overhead irrelevant.
@dmitrshvets ·
LFM2.5-350M by @liquidai runs on @trymirai. A model less than half the size, outperforms Qwen 3.5-0.8B on reasoning and agentic tool use. We tested across 10 Apple Silicon configurations. Even when running in full precision, the model achieves the throughput of over 70 t/s on an iPhone, which meets and surpasses the interactivity needs of most applications.

@cjzafir ·
Is GPT 5.4 really good? I used codex-5.4-extra-high to fine-tune qwen-3.5-4b. (SFT) (Exhausted all my pro plan weekly credits in 24 hours.) And also used opus-4.6 to fine-tune qwen-3.5-9b Codex is fast but dataset quality is crap. Opus is slow but data quality is great. What i am doing? I am performing distillation (using opus 4.6 as teacher & qwen-397b as stident) And taking their supervised chat data to fine-tuning qwen-3.5-4b and qwen-3.5-9b models. Already beaten: Gpt-4o Gemini-2.5-flash Gpt-5-nano Test passed. Now I'm making a python kit to perform distillation on scale (20M to 40M tokens). Qwen models have opened something big > smaller, smarter, faster, niched models running locally on your mac. Future is fun.

@testingcatalog ·
Alibaba released Qwen 3.6 Plus, an upgraded agentic model with coding and vision capabilities. Qwen 3.6 Plus comes with a 1M context window and is already available on Qwen Chat.

@RoundtableSpace ·
Qwen 3.8 Preview finished the same Blender tasks as Kimi K3 three to five times faster, completing a garden scene in 31 minutes versus Kimi's two hours.
@RoundtableSpace ·
Qwen and GPT built Angry Birds from one prompt. DeepSeek built a bug report. Qwen cost $0.07, GPT cost $0.62 and had the best physics.
@boxmining ·
Qwen 3.6 Plus being free right now is actually pretty interesting. Tried @Alibaba_Qwen inside Hermes Agent for system audits and log analysis, and it surfaced issues I would’ve probably missed. The best part is the think blocks. Seeing how it reasons through errors and tool calls makes debugging feel way less blind.
@LinasLekavicius ·
If you have an app that would benefit from AI analyzing video input... you need to check out Qwen 3.5 flash by @Alibaba_Qwen It's super cheap (analyzing 5 seconds video with no reasoning can cost as little as 1/15th of a cent, through OpenRouter) and this just unlocked various new features for my fitness app: - already implemented AI-based cheating prevention in @squadletics using random video recordings - soon introducing a new daily competition mode that this AI model finally allows me to do, and next up: - workout feedback / tips based on AI analysis
@KyleHessling1 ·
Need a fix for the near-infinite thinking on Qwen 3.5 27B at long contexts? Just turn thinking off! I thought surely it would castrate the model; older models used to be garbage without thinking, but the Qwen 3.5 27B seems to be so incredibly dense that it barely needs it (pun intended) It still beats the new Nemotron, and that's with Nemotron's thinking mode. In use, it genuinely seems to be smarter for my purposes because I can iterate so much faster and keep all of the past conversations in a fresh context at fp16. It's like I'm thinking with it, rather than it spending 30k tokens second-guessing itself. MAN just when I think I found this model's limits, it continues to impress me. I would have thought it would score sub 20 without thinking, but it's genuinely better. @Alibaba_Qwen, please make us a 72b dense :)

@dino11 ·
Alibaba $BABA just open-sourced Qwen 3.5 — a 9B parameter AI model that runs on your laptop. The benchmarks are insane: → Beats OpenAI's GPT-OSS-120B (a model 13x its size) on reasoning → GPQA Diamond: 81.7 vs 71.5 → 30-50 tokens/sec on a standard laptop → The 2B version runs on an iPhone in airplane mode at 22 tokens/sec Free. Open-source. No API costs. No internet needed. If you're building AI products and still paying per-token for every request, this changes the math completely.
@boxmining ·
I tested Qwen 3.8 Max for vibe coding and it is genuinely impressive. We built an interactive Tokyo website with scroll animations, games with player progression, and a physics simulator. This might be my new go-to for website design. Watch the full benchmark here:
@ValsAI ·
We evaluated @Alibaba_Qwen's Qwen 3.5 Flash on our remaining benchmarks. The model places in the top 10 on several benchmarks, including MortgageTax, LegalBench, and MMMU.

@TheGeorgePu ·
Alibaba's Qwen3.7-Max just hit #2 on Code Arena. Above GPT-5.5. Above Gemini. Behind only Claude. Closed weights. API only. The 'China wins via open source' story is over. DeepSeek opened the model. Qwen closed it. Both work. The catch-up player stopped sharing. They sell it now.
@JeremyCMorgan ·
Han Xiao open-sourced a KG extractor running Qwen on a single L4. Docs in, streaming triples out, evidence spans and confidence per edge. The prompting tricks that force canonical entities are a free lesson for anyone building extraction without API bills. https://t.co/tq2n9KrGN3
@agenticgirl ·
A smaller model just outperformed the biggest ones. Qwen 3.6-Plus scored 61.6 on Terminal-Bench and 57.1 on SWE-Bench. That puts it ahead of Claude Opus 4.5, Kimi K2.5, and Gemini 3 Pro. Models that are much larger and far more expensive to run. For the past year everyone believed one thing. Better performance needs a bigger model. This breaks that rule. But the more important part is where it is winning. Not simple tasks. These are multi-step coding workflows where the model has to plan, fix its own mistakes, and track changes across multiple files. That is exactly where most models fall apart. They start strong then lose the thread after a few steps. Qwen 3.6-Plus is holding up there. Size has always come with a cost. More compute, more money, more setup. Smaller models are faster, cheaper, and easier to deploy. If they can now match or beat the big ones on hard tasks, they become the obvious choice for most teams. It also ships with a 1 million token context window. Enough to load an entire codebase at once and not lose track halfway through a long session. It is already free on OpenRouter. Open source version is coming. So the real story is not that one model scored higher. It is that being bigger is no longer enough.

@JulianGoldieSEO ·
Qwen 3.8 vs Fable 5: 50 real builds, one winner. This guy literally built 50 games with both AIs to find out. Qwen scored 86.6 on Terminal Bench. Fable 5 scored lower. Qwen's games ran smooth and full 3D. Fable 5 got buggy on half the builds. But Fable 5 won on flight sims and big projects. Here's the twist. Qwen is free. And open source. Fable 5 needs a paid plan. Free is now trading punches with the best AI on earth.
@JulianGoldieSEO ·
ALIBABA’S NEW AI WORKED ALONE FOR 16 DAYS STRAIGHT But autonomous coding is not even the biggest part of this launch. What Qwen 3.8 Max built: → Started with an empty folder → Turned requests into GitHub issues → Assigned the work to itself → Wrote code, ran tests, and improved the software → Finished with 265 commits, 127 pull requests, and 151 issues Zero human input. What powers it: ✓ 2.4 trillion total parameters ✓ Only 95 billion activated per request ✓ 1 million-token context window ✓ Processes text, images, and video ✓ Open weights announced for next week Alibaba also says it reproduced a research paper in five days, wrote 7,600 lines of code, and ran 33 GPU training jobs without starter code. Important caveat: These benchmark results come from Alibaba. Independent testing still needs to confirm them. The real shift is not which AI writes the best email. It is which AI can take ownership of an entire project and keep working for days.
Best Qwen tweets
Xholic studies what works in your niche, drafts posts in your voice and schedules them for the hours your audience is online.
$0 today · Cancel anytime
Browse all tweet collectionsKeep exploring