Model releases and performance
Grok version launches, benchmark results, reasoning quality, hallucination rates, context windows, inference speed, and price-performance comparisons.
36%
Best tweets about Grok
Read the best tweets about Grok and xAI, including model updates, features, benchmarks, developer access, and user experiments. Updated weekly.
Specific Grok and xAI analysis, demonstrations, releases, and firsthand usage instead of incidental mentions.
Original Xholic analysis
Discussion is predominantly supportive (68%) and positive (72%). The largest measured theme is model releases and performance (36%), followed by coding agents and Grok Build (28%) and multi-agent systems (24%). An independent Artificial Analysis post reports that Grok 4.20 improved by 6 points over Grok 4 on its Intelligence Index but remained below its stated frontier score; individual posts also raise concerns about usage-limit transparency and platform stability.
72% of posts
All-time engagement
56% of posts
Published in 90 days
Conversation map
Grok version launches, benchmark results, reasoning quality, hallucination rates, context windows, inference speed, and price-performance comparisons.
36%
Grok Build, coding-focused models, terminal and CLI workflows, software-engineering benchmarks, agent tooling, plugins, and developer firsthand tests.
28%
Grok’s parallel agent architecture, custom agent creation, specialist roles, agent coordination, and scaling agent teams.
24%
xAI API access, third-party platforms, OpenClaw support, MCP connectors, enterprise deployments, and Grok embedded in work tools.
22%
Cloud-computer bots that operate apps and websites, retain routines, coordinate with other bots, and perform delegated research, operations, monitoring, and content tasks.
22%
Grok’s use of live X and web context for search, news monitoring, tweet explanation, topic discovery, and research summaries.
22%
Grok’s image and video generation, Imagine Agent Mode, visual-model quality, creative workflows, and creator usage.
20%
SuperGrok tiers, agent allowances, API pricing, compute budgets, access restrictions, and concerns about changing usage limits.
18%
Tone and stance
Performance benchmark
Posts with media make up 68% of this collection. Their median all-time score is 22.2, compared with 18.3 for text-only posts.
Format mix
Consensus and debate
Shared view
Posts emphasize coding, tool use, cloud computers, and delegated workflows rather than only a standalone chat interface.
Shared view
Creators describe using Grok-powered tools to search X, monitor accounts, summarize recent posts, and build continuously updated research databases.
Shared view
Users describe specialized bots, coordinator roles, and parallel custom-agent collaboration; one user also says coordinating many bots can be challenging.
Shared view
Release and benchmark posts cite faster inference, lower API pricing, and improved instruction following or non-hallucination measures, though the sources include both vendor-adjacent claims and independent testing.
Open debate
Artificial Analysis reports that Grok 4.20 improved over Grok 4 but remained below its stated intelligence frontier; other posts make more bullish price-performance and quality comparisons.
Open debate
Firsthand users describe Grok Bot as simple and effective, while also noting comparable workflows can be built with other agent systems; another post makes platform stability a condition of its potential.
Open debate
One critique alleges opaque and changing SuperGrok budgets, while a pricing post describes tiers differentiated by agent count, context, generation volume, and compute priority.
Open debate
Posts ask for formal Grok 4.3 details and question limited public availability or independent validation for an early Grok 4.5 claim.
What performs
The top two supplied engagement outliers are the Grok 4.5 coding-and-agents announcement (all-time score 1,078.76) and a detailed Grok Bot workflow account (748.59).
The OpenClaw Grok-integration post is the third-highest supplied outlier, with an all-time score of 473.12. Other posts also describe integrations across developer and work tools.
Deterministic analytics reports 34 media posts (68% of the set). Their median all-time score is 22.22, compared with 18.31 for text posts.
Announcements account for 28 posts (56%) in the deterministic analytics, ahead of opinion, case-study, and tutorial formats.
Statistical standouts
Creator landscape
The five most represented creators account for 20% of the selected posts.
1. Alex Finn
@AlexFinn
2 posts
2. Artificial Analysis
@ArtificialAnlys
2 posts
3. Derya Unutmaz, MD
@DeryaTR_
2 posts
4. Farzad 🇺🇸 🇮🇷
@farzyness
2 posts
5. Mark Kretschmann
@mark_k
2 posts
6. tetsuo
@tetsuoai
2 posts
Firsthand accounts describe concrete uses including monitoring, research, database creation, application testing, and content repurposing.
Benchmark posts include comparative scores, API pricing, context-window information, inference speeds, and stated limitations or rank positions.
Stepwise posts cover specialized roles, access requirements, delegated outcomes, coordinator bots, and reusable routines.
Since the previous snapshot
Themes, sentiment, stance, and post format are classified per tweet. All counts, shares, medians, creator concentration, freshness, and performance comparisons are then calculated directly from the published snapshot.
Xholic's all-time score compares engagement while accounting for reach, post age, and creator consistency. It is used for relative comparisons within this collection.
This report analyzes the exact 50-post snapshot shown below. AI identifies editorial categories and drafts explanations; all statistics are calculated from the snapshot, and every narrative claim is checked against cited posts before publication.
Best Grok tweets
Ranked 01–50
@SpaceXAI ·
Announcing Grok 4.5, our first model trained specifically for coding and agents. It was trained with Cursor and offers frontier intelligence at leading speeds and cost efficiency. https://t.co/i8HpU7w64k
@AlexFinn ·
After using Grok Bot nonstop for the past week I'm going to say it: it's the best AI agent out there right now Dead simple and just works It eliminates the 10,000 decisions, configs, and fixes that scare people off The integrated cloud computer unlock so many use cases too: 1. Invited the cloud computer to my community. It opened it up and now monitors the community 24/7. Answering questions and DMs 2. Uses the X plugin to monitor all the AI companies X accounts every 15 minutes around the clock. Looking for updates and big news, alerting me the moment it happens 3. Clicks around testing the 2 apps I'm building around the clock. Looking for bugs, thinking of new features, then writing the PRs and waiting for my review 4. My Grok Bot checks my social media channels around the clock. Moment I post anything, it repurposes the content based on other content I've posted, making writing my newsletter much easier 5. The integrated Grok images are really nice. So I can tell it to go to my YT on its cloud computer and make similar thumbnails for my new videos 6. Checks data in post hog on my product retention, gives me a list of at risk subscribers, prepares emails and reachout, then asks me for the OK Can the other AI agents out there do this? Yes. There's literally nothing Grok Bot can do that the others can't. But here's the thing, the user experience is designed so beautifully that it unhobbles all of these use cases and make them so much more pleasant to execute on Grok Bot is a must try
@veggie_eric ·
When training Grok 4.3, we spoke directly with devs and businesses to understand what they actually needed: a model that’s fast, affordable, and great at tool calling. The result is a daily driver that doesn't just look good on random benchmarks, but is actually useful in the real world. 💰 $1.25 in / $2.50 out ⚡️ 100 tokens / second 📖 1 million context window Try it through Hermes Agent or direct through the xAI API!
@mark_k ·
Grok Build is very close now. Release is expected next week unless something major causes a delay. Grok Build is @xAI’s answer to Claude Code and Codex. It consists of a suite of specialized coding models, variants of Grok 4.3, along with full tooling in the form of a web app and CLI. Grok Build is very important for both xAI and its users. Strong alternatives to the big two in AI coding are sorely needed.
@XFreeze ·
Grok expansion this month is unbelievable Most people still do not realize what is happening Grok is being plugged directly into the tools people already use every day: • Interactive Brokers — portfolio analysis, market research, strategy + order instructions • AWS Bedrock — enterprise access to Grok 4.3 • Databricks — AI agents + data workflows • Microsoft Word, Excel & PowerPoint — office productivity • Vapi — voice agents • eToro — real-time market sentiment inside Tori • Gopuff — AI shopping assistant • Warp — coding + terminal workflows • T3code — AI coding-agent workflows • Vercel — deployments, build status + domains • MongoDB — database management + queries • Sentry — debugging + error analysis • Cloudflare — Workers + infrastructure tools • Chrome DevTools — browser control + performance testing • Firecrawl — web search, scraping + page interaction • Superpowers — agent workflow plugins And that is not even the full list Grok Build also added /goal for long-running autonomous tasks, Agent Dashboard for managing multiple coding sessions along with Grok Build 0.1 & Grok Composer 2.5 Grok Imagine Video 1.5 is now faster, better, and live across API, web, iOS, and Android Grok Connectors already plug into Google Workspace, Outlook, SharePoint, OneDrive, Notion, GitHub, Linear, and custom MCP servers Finance. Cloud. Coding. Voice. Shopping. Data. Office work. Browsers. Developer tools. Agents Grok is moving from “AI you chat with” to “AI that works inside everything” xAI’s execution speed right now is insane
@XFreeze ·
xAI just brought Grok to OpenClaw - an open-source, local-first AI agent and personal assistant Starting today, users can log in with either: • SuperGrok • X Premium subscriptions and use Grok models directly inside OpenClaw OpenClaw runs on almost any hardware: • Mac Minis • laptops • servers • VPS systems • even Raspberry Pi devices while maintaining persistent memory across sessions The open-source AI agent ecosystem around Grok is expanding incredibly fast
@DeryaTR_ ·
I created a Grok Bot and gave it a task: scan X feeds from the past couple of days for new AI & robotics releases, then create a Notion database and add everything it finds. Within minutes, it had created the entire database! I’ve now set it up as an automation that scans X every day, finds new AI and robotics releases, and continuously adds them to the database. So the database essentially grows and updates itself every day, with my Grok Bot doing the work! This is, of course, not unique to Grok Bots, you can build similar workflows with other agentic systems. But Grok Bots made the whole process incredibly easy and intuitive. I’ve already set up more than a dozen bots, and I’ll likely keep adding more!
@farzyness ·
Grok Bot is TOTALLY OpenClaw/Hermes for normies. If SpaceX team can nail stability of the platform and continues to develop the Grok models to be excellent for agentic processes, I think they have a MASSIVE hit on their hands. They need to target this to the entrepreneur class. It's nearly as capable as CODEX or Claude Code but FAR more accessible and fun. And as Grok continues to get better, it'll be very difficult to discern which one is actually better at doing tasks between Grok/CODEX/CC, and by then folks will just worry about cost and speed. Two things that Grok is already very good at. And the beautiful thing is that they can package it right into X and have a MASSIVE distribution arm, since a big % of the world's entrepeneurs are probably already on X. This could give SpaceX the same pathway on AI revenue that OpenAI/Anthropic have on tokens, which when paired with their Data Center buildout + space compute + communications, they quickly turn into a jauggernaut. $SPCX
@tetsuoai ·
🚨 xAI’s Grok pricing ladder. Free • Grok 4.20 with very tight limits • ~10 prompts / 10 hours SuperGrok Lite ($10/mo) • Grok 4.20 • 1x AI agent on Expert mode • 480p image + 6 second video SuperGrok (3 days free, then $30/mo) • Grok 4.20 • 4x AI agents on Expert mode • longer context • 720p image + up to 30 second extended video • longer clips • higher usage limits SuperGrok Heavy ($300/mo) • Grok 4.20 Heavy • 16x AI agents on Heavy Mode • maximum compute priority • longest context • highest throughput • significantly higher image + video generation • early access to new features What they’re pricing: • agent count • context • generation volume • compute priority
@AlexFinn ·
Over the weekend I hacked for 24 hours at the xAI headquarters I built a brand new, AI-driven post composer for X Instead of staring at a blank screen, you get a Grok powered assistant that helps you write content Should X implement this onto the site? Full demo video here:
@ArtificialAnlys ·
The Grok 4.20 Beta shows three major improvements over Grok 4: ➤ Our lowest ever hallucination rate on the AA-Omniscience evaluation. When Grok did not know the answer, it hallucinated an incorrect answer 22% of the time - this is the lowest hallucination rate of any model we have tested, topping Claude Haiku 4.5 (25%) ➤ Top scores for instruction following and prompt adherence. On IFBench, Grok 4.20 takes the #1 spot with 82.9% - a +29.2 point increase on Grok 4 ➤ Leading speed for its intelligence. At 265 tokens per second output speed on xAI’s API, Grok 4.20 is significantly faster than its peer and over 2x the output speed seen from Grok 4.1 Fast Congratulations to @xai and @elonmusk on the 4.20 Beta 0309 launch!
@EXM7777 ·
grok bot just released and it's VERY impressive... here's how to get the best out of it: > give each bot one job: inbox, outbound, ops, one lane each > put a chief of staff bot on top so you're not the middleman > do the job once while it watches, it saves the routine and runs it alone next time > sign it into your real tools, it has its own computer and works inside apps that have no API > say what done means, otherwise the bot decides for you > make it verify before reporting back: open the app, click the real path, fix what it finds the bots message each other, share context and get sharper every week this is the closest thing to hiring i've seen from an ai lab
@iruletheworldmo ·
tried grok again today and it now feels much better than opus on many things. xai seem to be the only lab make improvements every week. not sure how do it. groks long context and low hallucination rates has proved very useful. grok is the best model in the world for low hallucination rates. so i spend way less time checking output. grok imagine feels level with anything else ive tried also. i wondered when that compute would pay off. a bitter pill to swallow for dario and his cones. and grok 5 is huuuuuge
@ArtificialAnlys ·
xAI has released Grok 4.20 for API access in beta, and it scores 48 on the Artificial Analysis Intelligence Index with reasoning enabled Compared to @xAI’s previous Grok 4 flagship, Grok 4.20 Beta 0309 is an intelligence upgrade, achieving +6 points on the Intelligence Index. It launches with a longer 2M token context window (up from Grok 4’s 256K context window, matching Grok 4.1 Fast’s 2M), and significantly lower pricing ($2/$6 vs Grok 4’s $3/$15). Grok 4.20’s performance lags behind the current intelligence frontier, but it performs strongly on instruction following and features a notably low hallucination rate, beating all other models we’ve tested on AA-Omniscience for hallucination. xAI released 3 variants: reasoning, non-reasoning, and multi-agent. We’ve evaluated the reasoning and non-reasoning modes, and are considering the best approach for testing the new multi-agent functionality, which parallelizes over multiple agents behind the scenes in one API call. Key takeaways: ➤ Improved intelligence over Grok 4: Grok 4.20 Beta 0309 (Reasoning) scores 48 on the Artificial Analysis Intelligence Index, +6 from Grok 4 and +9 compared to Grok 4.1 Fast. This score falls short of the current intelligence frontier at 57 (Gemini 3.1 Pro Preview and GPT-5.4) ➤ Low price for the level of intelligence: Grok 4.20 is priced at $2/$6 per 1M input/output tokens, representing a decrease compared with Grok 4’s $3/$15 API rates. The reasoning variant cost $484 to complete the evaluations in the Artificial Analysis Intelligence Index, which is a reduction of ~70% compared to Grok 4, driven by lower pricing and lower token use ➤ Leading non-hallucination rate: Grok 4.20 scores 78% in the AA-Omniscience non-hallucination metric. This is the best result we have seen yet for this metric, and reflects the model only answering around one fifth of the time when it did not know the answer ➤ Fast inference performance: xAI is serving Grok 4.20 at 267 tokens per second - similar to what we see for gpt-oss-120b across providers and on the Pareto frontier for speed versus intelligence ➤ Mixed improvements in tool use: Grok 4.20 improved on the tool calling performance of Grok 4 in some evaluations, scoring 97% on Tau2-Telecom. However, its score of 1,062 on GDPval-AA, our benchmark of general agent performance on real work tasks, is well behind frontier peers and sits approximately in line with Grok 4.1 Fast
@tetsuoai ·
🚨 xAI is moving fast right now. Grok 4.20 is being reported with major gains in reasoning, multimodal performance, instruction following, speed, and hallucination reduction, with its multi agent architecture. Currently #1 on IFBench at 82.9% Next best scores there are 79.6% and 79.0% Now Grok Imagine 1.0 is out of beta, and a bigger video release is already being teased for next week. xAI is shipping across reasoning, search, image, video, and agents all at once!
@farzyness ·
What's fascinating about Grok Bot is that you can literally run every other model through it. You can have agents that run all their jobs through a CLI and ping Claude/ChatGPT/etc. for whatever. And as SpaceX grows its compute clusters, then more and more tokens for those models will be generated within SpaceX. Which means that even if you use a model that is not Grok, the tokens that are generated will increasingly come from the company that made Grok. And Grok will be the one orchestrating the whole thing. So even if you're Anthropic/OpenAI, and you have a killer model, there's an increasing likelihood that your model will be triggered by SpaceX and your tokens will be generated by SpaceX. The question becomes will people actually pay a premium for the Grok Bot agentic harness.
@DeRonin_ ·
For now, you can create your own AI agents on Grok 4.20 To do this simply: - click on your pfp in the bottom left corner - click on "Settings" - then click on "Customize" - click on "Create" and build your agent Grok 4.20 itself is built as a multi-agent system that automatically deploys 4 specialized agents But around March 4-7, 2026, xAI released an update that added support for custom agents Now you can: - create your own agents by defining their name, role, personality and instructions - these custom agents participate in multi-agent collaboration (they discuss, fact-check, and work in parallel) the base version still provides 4 agent slots, while SuperGrok Heavy allows scaling the system up to a swarm of 16 agents with deeper configuration the future is already here, and Grok is definitely part of it Source: Grok
@Angaisb_ ·
xAI introduces Grok 4.5, trained alongside Cursor and focused on coding, agentic tasks and office work - 80 tps, faster than flash models according to xAI - 4.2x fewer output tokens than Opus 4.8 on SWE Bench Pro - $2 / $6 per million tokens - 83.3% on Terminal Bench, basically tied with GPT-5.5 and close to Fable - 64.7% on SWE Bench Pro, above GPT-5.5 but far from Fable's 80.4% Available in Grok Build, Cursor and the API. Not in the EU yet, expected mid July
@invideoOfficial ·
Grok Imagine, the AI video model that doesn't want to be real. Since @grok launched its video capabilities, the discourse has been almost entirely content moderation - what will and won't generate. Spicy mode, regulatory investigations, the EUs reaction. It's loud, it's dramatic. And it's completely missing the point. Because buried underneath all of that noise is something genuinely fascinating for filmmakers and visual creators. Grok Imagine video is the only major AI video model that was built to honor imagination over physics. See how it performs in 6 styles versus other models along with detailed prompts 🧵👇
@aiedge_ ·
Grok Bot was recently launched, and it might just be the best AI tool ever released by xAI. If you still haven't set it up yet, you'll want to bookmark this. How to launch your own team of Grok bot agents: Step 1. Get access Currently early beta on SuperGrok Heavy, Cursor Ultra, or Cursor Teams Premium, starting around $120/month. Available now on macOS and iOS, with Windows and Linux desktop apps also live. Step 2. Create your first bot Each bot gets its own cloud computer and can sign into whatever apps or websites you need, even ones with no API. Message it in plain language like you would a real employee. Step 3. Give it a real task, not just a question Don't ask it to explain something. Delegate an actual outcome: "research competitors in X space, pull pricing from their sites, and summarize into a report." It signs in, does the work across apps, and only comes back when it's done or needs your approval. Step 4. Create specialized bots for different roles This is where it gets powerful. Instead of one bot doing everything, spin up separate bots for separate functions, a researcher, a writer, an ops manager. Each one stays focused on its lane. Step 5. Add a coordinator bot Create a "Chief of Staff" style bot whose only job is managing the other bots. It assigns work, checks progress, and pulls you in only when a real decision is needed. Step 6. Let them communicate directly Bots can join a shared chat and pass work between each other, real coordination, not just parallel tasks running in isolation. Step 7. Teach a workflow once, reuse it forever Demonstrate a process while a bot follows along, and it saves that as a routine it can run again without you re-explaining every step. The end result: an entire team of specialized AI agents, working across your actual tools, coordinating with each other, checking in only when it matters. Set this up once. It runs indefinitely.
@FGLXLucas ·
xAI’s Grok Limits: Transparency Deficit and Strategic Risk in a Brutal AI Market In recent months, xAI has repeatedly adjusted SuperGrok usage conditions without clear communication. Instead of fixed, contractually guaranteed render and usage capacities, there is now a dynamic weekly shared budget that is quickly exhausted, especially for compute-intensive features like Grok Imagine (images and videos). Users like Luc report that even with efficient usage (e.g., one 30-second film plus two 10-second clips per day), the budget has become noticeably tighter — a development many perceive as yet another 50-70% reduction. The problem goes far beyond casual users. Developers using Grok for complex programming, agent workflows, or long-running analyses are equally affected, as are creators who generate income with the outputs. This becomes particularly critical for future applications such as robotics (Optimus) and mobile AI, where reliable compute planning is essential. Anyone who builds their workflow or even their business model on Grok risks sudden restrictions or hidden upsells (Heavy tier, extra credits). For xAI, this is a risky game. In an extremely competitive AI market where OpenAI, Anthropic, Google, and specialized providers are fighting for developers and companies, trust and predictability are key differentiators. Frequent, opaque limit adjustments create exactly the kind of dependency that can later lead to higher prices or churn. Many professional users are already consciously diversifying and using Grok only for quick tests, while relying on more transparent alternatives for productive work. Conclusion xAI has excited many users with Grok and SuperGrok — but the current limit policy undermines precisely the trust needed for long-term success in the AI industry. For users, this means: diversification and critical review of dependencies are advisable. For xAI: greater transparency and reliable capacities would be a clear competitive advantage rather than a risk. #Grok #SuperGrok #xAI #GrokLimits #AILimits #AITransparency #xAICriticism #GrokImagine #AI #ArtificialIntelligence #TechBusiness #xAIRisks
@mercor_ai ·
Grok 4.5 from @SpaceXAI places #2 on the APEX-SWE leaderboard at 51.2% Pass@1 (±6.0), behind Fable 5 (65.5% ±6.2) on our benchmark for real-world software engineering work. It leads Integration (65.0% Pass@1) and places #2 in Observability (37.3% Pass@1), covering multi-step build tasks and diagnosis/debugging respectively. The Integration lead maps directly to the agentic workflows Grok 4.5 was built for: multi-step coding tasks run in collaboration with Cursor. Grok models have improved 30.2 pp in a year on this benchmark: Grok 4 (21.0% Pass@1) to Grok 4.5 (51.2% Pass@1). Congratulations to the xAI and Cursor teams.
@LiorOnAI ·
Grok 4.5 may have just produced original mathematical research. A mathematician gave it an open research problem that had remained unresolved for years. Instead of explaining existing work, Grok constructed an explicit counterexample. If the proof checks out, that counterexample settles the question. The problem was about hypercontractivity, a fundamental property used throughout harmonic analysis, probability, and partial differential equations. Researchers already knew the property held in dimensions up to 3 and failed by dimension 13. The missing piece was where the transition actually happened. Grok's counterexample shows the first failure already occurs in dimension 4, making the previous result sharp. If independent verification confirms the proof, this won't just be another example of AI helping with research. It will be an example of an AI contributing a new mathematical discovery.
@icreatelife ·
I'm impressed with Grok as a visual video model, how fast it emerged and keeps improving. I hope every platform will have Grok model as a choice of 3rd party models. I use it a lot for my AI animations.
@Axel_bitblaze69 ·
Grok 4.5 has been out for around 2 weeks, and after using it as my main model for work, here’s the easiest way to understand where it actually fits. yeah if you compare it to Fable 5.. Grok 4.5 is not as smart as it on every task. but i’d say It is probably the best default for most tasks. Even Elon’s positioning is basically this: Fable 5 is still better when the problem is genuinely difficult, but most daily work does not need that level of capability. And after using Grok 4.5, I think that is exactly the point. On Artificial Analysis’ coding-agent tests, Grok 4.5 performs in the same range as GPT-5.5 and below Fable 5 overall. On Terminal-Bench, it scored 83.3% versus 84.3% for Fable and 83.4% for GPT-5.5. So the quality gap is already quite small on practical agentic work. But the cost gap is not small. Average coding-agent task cost: Grok 4.5: $2.49 GPT-5.5: $5.07 Fable 5: $11.80 It also used roughly 1.9M tokens per task, compared with 6.2M for GPT-5.5 and 7.2M for Fable. So you are getting near-frontier performance at roughly one-fifth of Fable’s task cost. That changes how you use it. With an expensive model, you save it for important prompts. With Grok 4.5, you can leave the agent running. Let it inspect the project, search documentation, edit files, run commands, write tests, fix errors and keep iterating without worrying that every loop is burning another $10. This is where I would use it: → day-to-day coding and debugging → terminal tasks and long agent loops → building and testing fast → browser and workflow automation → research that needs live web or X context → repetitive work where the agent needs to keep trying → scripts and tasks you want running in the background For the hardest architecture decisions, complex reasoning or a painful large-codebase refactor where the last few percentage points of quality matter more than cost, I would still reach for Fable 5. That is not a weakness. It is just knowing which tool fits the job. Another thing I like is how quickly the product around the model is improving. Grok Build has been shipping versioned updates frequently, with actual public changelogs instead of vague “performance improvements.” One correction worth making: Grok 4.5 itself is not open source. Grok Build, the terminal agent that runs it, was open-sourced on July 15. You can inspect and modify the harness, use skills, plugins, hooks, MCPs and subagents, or point it toward your own local model. My takeaway after daily-driving it: Stop asking which model is the smartest. Ask which model is capable enough, fast enough and cheap enough to stay inside your workflow all day. For the hardest 5% of tasks, use the strongest model available. For the other 95%, Grok 4.5 is starting to look like one of the best defaults. You can try it through Grok Build with one command: curl -fsSL https://t.co/7z5ptv2drv | bash
@jasondoesstuff ·
Fable 5 release: 90% of X talking about it. GPT-5.6 release: About 25% of X talks about it. Grok 4.5 release: Literally 2 people in my feed 😄. Clear levels of trust in the models. While we may not agree with Anthropic's rollout approach (and rug pull not their fault entirely), they still have consumer trust over the rest. Opus model earned them immense customer retention. Not surprised at all by quietness of Grok 4.5. I've tried other Grok models a few times and it feels like 2024 AI. Gonna need a lot of positive results to earn that trust back.
@VaibhavSisinty ·
Grok just turned AI image generation into a film studio. xAI rolled out Imagine Agent Mode on Grok web. 🤯 Still in beta mode. It's not a new model. It's a new way of working. You drop a brief on an infinite canvas. The agent plans the steps. Generates the assets. Edits them. Iterates. All in the same workspace. → Tell it "generate a 1-minute cinematic film." It builds the film. → Tell it "create a complete manga set." It lays out the panels. → Tell it "build UGC product stories." It ships the whole set. You give the brief. It does the work. Until today, every AI image tool made you the operator. Grok just made you the director. vc: @XFreeze
@cyrilXBT ·
GROK 4.5 MIGHT ALREADY BE REAL. ALMOST NOBODY CAN USE IT YET. A version number quietly disappeared from xAI's menus on June 28. Hours later, Musk confirmed it: Grok 4.5, built on a 1.5T V9 foundation model, now in private beta at SpaceX and Tesla. His claim: performance close to, perhaps exceeding Opus. Here is the catch. That evaluation came from SpaceX and Tesla engineers—employees of the same parent company as xAI. No independent benchmark exists. No public API. No release date. The model everyone can actually use today is still Grok 4.3. Bookmark this. Follow @cyrilXBT
@VadimStrizheus ·
EVERYONE IS MISSING THE IMPORTANT PART OF ELON’S GROK UPDATE. It is not just “Grok got bigger.” The real story is that Elon said a lot of Cursor data was added into Grok V9’s supplementary training. That matters way more than people think. Grok’s original unfair advantage was real-time X data. When xAI launched Grok, the whole pitch was simple: Most chatbots understand the internet’s past. Grok understands what is happening right now. That was the first data edge. Now Elon is pointing Grok at the next one: how builders actually write software. Cursor did not win because it was “another AI coding tool.” Cursor won because it moved AI from the chat window into the repo. It watched developers work inside real projects. Prompts. Files. Bugs. Refactors. Context. Bad outputs. Good outputs. Accepted changes. Rejected changes. Shipping decisions. That is the difference between training an AI to answer coding questions and training an AI to behave like an engineer. In this clip, Elon explains the actual AI race: compute, rate of improvement, talent, and unique data. That is the whole game. Now look at the pieces: xAI has the compute. Grok has live X data. V9-Medium is a jump from the smaller model serving Grok today. And now Cursor workflow data is being added for difficult coding tasks. This is not Grok trying to become another chatbot. This is Grok trying to become the AI engineer inside X. The next AI war is not who writes the prettiest answer. It is who can understand the repo, reason through the task, write the code, test it, fix it, and ship. That is why the Cursor line matters. That was not a random detail. That was the signal.
@VaibhavSisinty ·
Xai shipped 13 versions of grok build in the last 12 days. nobody is shipping AI coding tools at this pace. Here's what's now sitting inside Xai ↓ → a real command surface. /usage to see your token spend in any session, /login to reauth without leaving the terminal, /export to save the full conversation, /rewind to roll back a turn, /config-agents to manage your subagents from a modal. → plan mode and /execute-plan, so grok can break a big task into a structured step-by-step plan and actually run it. now with git and gt support built in. → subagents that share the same terminal backend, scheduler, and monitor across sessions. you can pause a multi-agent task, walk away, come back, and resume with the full UI and history replayed. → X search and a much faster web search wired straight into the terminal. grok can pull live context from https://t.co/VR4iwFS7OW without you alt-tabbing once. → proper image understanding. images now flow as multimodal vision tokens instead of base64 dumps. you can paste or drop multiple at once, and read_file extracts text from powerpoint files too. → video generation with user-set duration. and the terminal itself now plays back media at 30fps. → an "always-approve" mode for runs where you want grok to stop asking permission and just ship. literally called "yes, and don't ask again for anything." → plugins, MCP servers, and hooks all managed from one picker modal. an actual extension surface, not a feature list. And somewhere in the middle of those 13 versions, they shipped a feature called "laziness detector." Yes that's the actual name. it nudges grok when the model starts skipping steps. Most companies ship a product. xai is shipping a release schedule.
@CodeByNZ ·
👀 Grok 4.5 performs on par with GPT-5.5 and close to Opus 4.8 on coding benchmarks. The model shows strong results across Terminal-Bench, SWE-Bench, and DeepSWE, demonstrating its ability to handle difficult, long-running tasks that require creative tool use across software engineering, data science, finance, legal work, and other domains. This marks a significant upgrade over Grok 4.3, especially in complex, multi-step coding and agentic workflows.
@LaceyPresley ·
The distance between today’s AI and truly useful civilizational intelligence is no longer a matter of speculation. It is being compressed in public, at speed, by xAI. • Grok 4.5 launched as xAI’s strongest model to date, purpose-built for coding, agentic tasks, and professional knowledge work • Trained at large scale with heavy emphasis on curated engineering data and reinforced on multi-step real-world software tasks, developed in collaboration with Cursor • Places on the Pareto frontier for intelligence versus cost and speed alongside only the strongest competing systems • Achieves the highest perfect-extraction rate among tested models on large-scale real-world invoice processing (150,000+ business bills) • Serves at 80 tokens per second while using substantially fewer output tokens than leading competitors on equivalent engineering benchmarks • Rolled out across https://t.co/lGcC28UTvr, X, iOS, and Android within days of release, with expanded autonomous workflow capabilities in Grok Build • Grok 4.6 targeted for release in approximately two weeks and Grok 4.7 in four weeks, maintaining an unusually rapid iteration cadence This is the systematic conversion of compute, data discipline, and engineering focus into deployable intelligence that raises the floor of what individuals and organizations can accomplish.
@glenngabe ·
Regarding Google's AIOs citing Grok, here is some information about Grok's search visibility. The account https://t.co/Skl153h1kO has surged in visibility over the past year, although it did take a hit recently post-core update. It ranks for 87K queries right now based on ahrefs data and 35.3K in the US. From an AIO perspective, Grok ranks in AI Overviews for 21K queries. For example, like the one below in the screenshots. And from an AI Mode pov, Grok has surged since late 2025, but did take a hit post-May core update. That's the final screenshot. So yes, Grok is being cited in Google's AI surfaces quite a bit.
@lemire ·
I wanted to share my experience with Grok Build, xAI’s competitor to Claude Code. I should note upfront that I have considerable experience coding with AI, whatever that means in 2026. The Grok Build console interface appeals to me. It has a somewhat more serious appearance than Claude Code and feels less distracting. By default, it launches in full-screen mode, which gives it a cleaner look, though I recognize this is not a critical factor. What about the coding itself? So far, I have tested it on two different toy projects, plus one evolving real-world implementation. The first was a remake of an old game in modern C++. The project remains incomplete. As the first one I attempted, Grok Build handled it reasonably well at the start, but it eventually became confused by repeated edits to the same file. This occurred during the early days of Grok Build, so I am unsure how representative those issues were. I have continued working on it since then, and Grok Build has performed well. I later realized that the project was too ambitious for a simple test; building something that resembles a playable game simply takes a considerable amount of time. I have since moved on to a web application that serves as a dashboard for my school’s statistical data. It performed superbly, Grok Build handled it like a champ. I do not believe Claude Code would have done any better. It is going to take me some time making it really good, but Grok Build is doing very well. The terminology is confusing because it seems that Grok Build is also the name of a model that you can access through the xAI API. I have been working with it for an actual production application. I have designed a custom system that allows a professor to edit a course website using AI. I rely on Grok Build to modify the HTML, JavaScript, and CSS based on natural-language queries. The application works great, and I plan to share more details about it in the future. Exciting times.
@orcdev ·
I tested Grok Build with @grok 4.3, and in just a couple of days the improvement is crazy inside my project where I already had skills and plugins set up for AI agents, I gave it a single prompt and it returned a completely new page with the exact UI from my skills system, plus filters, sorting, featured sections, navbar updates, and multiple extra improvements I never even asked for what will happen when they drop Cursor data + Grok? check out the video 👀
@panditdhamdhere ·
Grok Build in Rust 🦀 Grok Build is @SpaceXAI 's terminal-based AI coding agent. It runs as a full-screen TUI that understands your codebase, edits files, executes shell commands, searches the web, and manages long-running tasks interactively, headlessly for scripting/CI, or embedded in editors via the Agent Client Protocol (ACP). https://t.co/UpLPXczvJw
@BenjaminDEKR ·
Kind of charming: I asked Grok for updates on the Blue Origin deployment and he said "I and the team am monitoring." Who is "the team?" The other Grok Expert Mode AI agents: Harper, Benjamin, and Lucas "We're all working together in real time on queries like this live space mission watch. I’m the team lead, but they help me scan X."
@JulianGoldieSEO ·
GROK JUST REMOVED THE MOST ANNOYING PART OF AI WORKFLOWS And most SEO teams are going to completely miss why this matters. What changed: → One add-on puts Grok inside Google Docs, Sheets and Slides → It works from a side panel beside your actual files → No copying answers from a chatbot into another app What it can do: ✓ Turn rough notes into a structured SEO article ✓ Write formulas, analyse data and build charts in Sheets ✓ Turn an outline into a researched slide deck with consistent styling The practical workflows are even better: ✔ Convert one coaching call into a blog, email, LinkedIn post and X thread ✔ Build a 30-day customer roadmap and training deck ✔ Analyse member activity and identify who needs extra support Grok also works inside Word, Excel and PowerPoint. The real advantage is not “better AI.” It is removing every unnecessary step between an idea and published work.
Best Tweets by Topic