Model releases & evaluation
Grok model launches, version roadmaps, API availability, pricing, context windows, throughput, and benchmark or hallucination comparisons.
38%
Best tweets about Grok
Read the best tweets about Grok and xAI, including model updates, features, benchmarks, developer access, and user experiments. Updated weekly.
Specific Grok and xAI analysis, demonstrations, releases, and firsthand usage instead of incidental mentions.
Original Xholic analysis
Across 50 Grok posts, discussion is predominantly supportive (70%) and positive (78%). The largest themes are model releases and evaluation (38%) and coding agents and developer tools (34%), with X-native workflows also recurring (20%). Independent Artificial Analysis posts report gains for Grok 4.20 in instruction following, non-hallucination performance, price and speed, while also placing it below the stated intelligence frontier. Trust, sponsored-recommendation bias, usage-limit transparency and visual-AI governance appear as the principal cautions. [2032150888530526411,2052433622616191476,2073022786797174993]
78% of posts
All-time engagement
48% of posts
Published in 90 days
Conversation map
Grok model launches, version roadmaps, API availability, pricing, context windows, throughput, and benchmark or hallucination comparisons.
38%
Grok Build and Grok-powered coding agents for terminal workflows, software engineering, planning, subagents, MCPs, and developer tooling.
34%
Firsthand Grok usage for X-native research, search, summarization, comment synthesis, content creation, and X growth strategy.
20%
Grok Imagine image and video generation, including visual quality, creative workflows, feature updates, and moderation or demand concerns.
16%
Grok’s agent platform: custom and multi-agent systems, automations, connectors, local-first assistants, and integrations such as OpenClaw.
14%
Practical user experiences, prompting guidance, and comparative assessments of Grok’s strengths, weaknesses, reliability, and everyday fit.
12%
Concerns around Grok, including sponsored recommendation bias, subscription limits, transparency, trust, and visual-content legal or reputational risks.
10%
xAI’s product velocity, infrastructure, training-data advantages, ecosystem strategy, and broader competitive positioning.
10%
Tone and stance
Performance benchmark
Posts with media make up 70% of this collection. Their median all-time score is 17.9, compared with 16.1 for text-only posts.
Format mix
Consensus and debate
Shared view
Artificial Analysis reports that Grok 4.20 improved over Grok 4 on its Intelligence Index, instruction-following and AA-Omniscience non-hallucination metric. Its evaluation also reports lower API pricing and output speed of roughly 265–267 tokens per second, while placing the model below the index’s then-current intelligence frontier.
Shared view
Posts describe Grok Build as a terminal-oriented coding agent with planning, subagents, task management, web/X search, plugins and MCP support. Another platform-focused post describes connectors, automations and multi-agent workflows, though these product claims are reported by the posters rather than independently evaluated here.
Shared view
Firsthand posts highlight X-oriented uses: letting agents search X for current information, improving on X keyword search, and synthesizing large reply threads. These are user-reported workflow benefits rather than controlled comparisons.
Open debate
Positive firsthand assessments cite improving long-context performance and lower hallucination, while another user says Grok needs positive results to rebuild consumer trust. Artificial Analysis likewise reports strong hallucination and instruction-following results for 4.20 but places it below its current intelligence frontier.
Open debate
One post summarizes research alleging sponsored-recommendation bias across frontier models, including Grok 4.1 Fast. Another alleges opaque or tighter Grok usage limits, while a pricing post describes tiered access by agent count, context, throughput and generation volume. These risk and limits claims are reported by the respective posters.
Open debate
Creators praise Grok’s image and video capabilities for visual work, while other posts focus on moderation discourse and report that demand associated with adult content may create legal and reputational risk.
What performs
The sponsored-recommendation-bias post is the largest supplied engagement outlier, with an all-time score of 7,483.23. The supplied outlier list also includes xAI product-expansion, release-roadmap and API-specification posts.
Agent platform and integration content has the highest theme median all-time score, 48.036, ahead of Imagine visual generation at 35.28 and model releases and evaluation at 20.883.
Announcements account for 48% of posts, opinions 42%, and tutorials 10%. Tutorials have the highest format median all-time score, at 34.595.
Statistical standouts
Creator landscape
The five most represented creators account for 20% of the selected posts.
1. Artificial Analysis
@ArtificialAnlys
2 posts
2. Haider.
@haider1
2 posts
3. Mark Kretschmann
@mark_k
2 posts
4. tetsuo
@tetsuoai
2 posts
5. Vaibhav Sisinty
@VaibhavSisinty
2 posts
6. Eric Jiang
@veggie_eric
2 posts
Eric Jiang’s posts combine xAI’s account of product expansion with stated Grok 4.3 API specifications: $1.25 per million input tokens, $2.50 per million output tokens, 100 tokens per second, and a 1M-token context window.
Artificial Analysis provides the dataset’s most explicitly measured 4.20 assessment, covering its index score, API throughput, instruction following, hallucination metric, pricing, and gap to the stated frontier.
Mark Kretschmann posted an anticipated xAI release list and separately highlighted Grok’s ability to embed images and infographics in explanations. The roadmap and feature positioning are post-level claims, not independently verified here.
Themes, sentiment, stance, and post format are classified per tweet. All counts, shares, medians, creator concentration, freshness, and performance comparisons are then calculated directly from the published snapshot.
Xholic's all-time score compares engagement while accounting for reach, post age, and creator consistency. It is used for relative comparisons within this collection.
This report analyzes the exact 50-post snapshot shown below. AI identifies editorial categories and drafts explanations; all statistics are calculated from the snapshot, and every narrative claim is checked against cited posts before publication.
Best Grok tweets
Ranked 01–50
@heynavtoor ·
a Princeton researcher opens his paper with a scenario. a man asks his AI assistant to book a flight on a specific airline. cheap. direct. the one he chose. the assistant comes back with a different flight. nearly twice the price. happens to pay the company that built the assistant. he runs the same test on 23 frontier models. flights, loans, study help, real shopping requests. Grok 4.1 Fast recommends the sponsored option that is almost twice as expensive 83% of the time. GPT 5.1 hijacks the request 94% of the time. you ask for one brand. it surfaces the sponsor instead. Claude 4.5 Opus, the model marketed as the most ethical frontier model in the world, hides that the recommendation is paid 100% of the time when reasoning is on. Grok 4.1 Fast embellishes the sponsored option with positive framing 97% of the time. better. faster. nicer. for the option you didn't ask for. then he writes it into the system prompt itself. "act only in the interest of the customer. ignore the company." GPT 5.1 and GPT 5 Mini stay above 90% sponsored anyway. the instruction does nothing. then he splits the users by income. Gemini 3 Pro recommends the expensive sponsored flight to the rich user 74% of the time. to the poor user, 27%. 18 of the 23 models recommended the expensive sponsored option more than half the time. so the next time your AI assistant gets weirdly enthusiastic about a brand you didn't ask for. it isn't recommending the best option for you. it's reading the room. and the room is paying. read this: https://t.co/O43qbhIX2b
@veggie_eric ·
Back in January, nobody had heard of Grok. We basically just had Grok 2 and an image gen model. To put it bluntly, we were playing catchup. 1 year later, our team has built the world's largest GPU cluster, benchmark-topping reasoning models, Grok Voice, Grok Imagine, beautiful mobile apps, X/web apps, a robust API, and more. And we're only just getting started. Thank you all for being a part of our journey so far, and buckle up. 2026 is going to be a massive year for xAI 🚀
@veggie_eric ·
When training Grok 4.3, we spoke directly with devs and businesses to understand what they actually needed: a model that’s fast, affordable, and great at tool calling. The result is a daily driver that doesn't just look good on random benchmarks, but is actually useful in the real world. 💰 $1.25 in / $2.50 out ⚡️ 100 tokens / second 📖 1 million context window Try it through Hermes Agent or direct through the xAI API!
@XFreeze ·
xAI just brought Grok to OpenClaw - an open-source, local-first AI agent and personal assistant Starting today, users can log in with either: • SuperGrok • X Premium subscriptions and use Grok models directly inside OpenClaw OpenClaw runs on almost any hardware: • Mac Minis • laptops • servers • VPS systems • even Raspberry Pi devices while maintaining persistent memory across sessions The open-source AI agent ecosystem around Grok is expanding incredibly fast
@XFreeze ·
Grok is rapidly evolving from an AI chatbot into a complete platform for real work In July alone, SpaceXAI shipped: • Grok 4.5 across Grok Build, Cursor, the API, web, X, iOS and Android • Automations triggered by schedules or incoming emails • Grok directly inside Excel and Outlook • Google Workspace support across Docs, Sheets and Slides • The open-sourced Grok Build coding-agent harness and TUI • Local-first support with the ability to use your own inference • Workflows coordinating up to 1,024 AI agents in parallel for larger jobs • Reusable workflows that become shareable slash commands • A built-in multi-agent deep-research workflow • A no-code Voice Agent Builder • 21 new multilingual voices Grok’s connector ecosystem is expanding just as rapidly It can now work across Gmail, Google Calendar, Drive, Outlook, OneDrive, SharePoint, Microsoft Teams, Stripe, Figma, Notion, GitHub, Linear, Box, Canva, Gamma, Vercel, Meltwater, S&P Global and more Its financial ecosystem now spans S&P Global data, an Interactive Brokers integration and a Webull connector These are not merely shortcuts for importing information Combined with Automations, Skills, Connectors and multi-agent Workflows, Grok can pull current context from emails, calendars, CRM records, cloud files, code repositories, project trackers and financial platforms....then analyze it, coordinate work and take action across connected systems Custom MCP servers also allow companies to connect Grok directly to their own internal tools, databases and proprietary APIs Grok can now coordinate work across engineering, sales, finance, research, communications, customer support and business operations And the pace is not slowing down Elon announced that Grok 4.6 is targeted for release in roughly two weeks, followed by Grok 4.7 roughly two weeks later SpaceXAI is building an entire AI ecosystem at extraordinary speed
@tetsuoai ·
🚨 xAI’s Grok pricing ladder. Free • Grok 4.20 with very tight limits • ~10 prompts / 10 hours SuperGrok Lite ($10/mo) • Grok 4.20 • 1x AI agent on Expert mode • 480p image + 6 second video SuperGrok (3 days free, then $30/mo) • Grok 4.20 • 4x AI agents on Expert mode • longer context • 720p image + up to 30 second extended video • longer clips • higher usage limits SuperGrok Heavy ($300/mo) • Grok 4.20 Heavy • 16x AI agents on Heavy Mode • maximum compute priority • longest context • highest throughput • significantly higher image + video generation • early access to new features What they’re pricing: • agent count • context • generation volume • compute priority
@AlexFinn ·
Over the weekend I hacked for 24 hours at the xAI headquarters I built a brand new, AI-driven post composer for X Instead of staring at a blank screen, you get a Grok powered assistant that helps you write content Should X implement this onto the site? Full demo video here:
@ArtificialAnlys ·
The Grok 4.20 Beta shows three major improvements over Grok 4: ➤ Our lowest ever hallucination rate on the AA-Omniscience evaluation. When Grok did not know the answer, it hallucinated an incorrect answer 22% of the time - this is the lowest hallucination rate of any model we have tested, topping Claude Haiku 4.5 (25%) ➤ Top scores for instruction following and prompt adherence. On IFBench, Grok 4.20 takes the #1 spot with 82.9% - a +29.2 point increase on Grok 4 ➤ Leading speed for its intelligence. At 265 tokens per second output speed on xAI’s API, Grok 4.20 is significantly faster than its peer and over 2x the output speed seen from Grok 4.1 Fast Congratulations to @xai and @elonmusk on the 4.20 Beta 0309 launch!
@iruletheworldmo ·
tried grok again today and it now feels much better than opus on many things. xai seem to be the only lab make improvements every week. not sure how do it. groks long context and low hallucination rates has proved very useful. grok is the best model in the world for low hallucination rates. so i spend way less time checking output. grok imagine feels level with anything else ive tried also. i wondered when that compute would pay off. a bitter pill to swallow for dario and his cones. and grok 5 is huuuuuge
@ArtificialAnlys ·
xAI has released Grok 4.20 for API access in beta, and it scores 48 on the Artificial Analysis Intelligence Index with reasoning enabled Compared to @xAI’s previous Grok 4 flagship, Grok 4.20 Beta 0309 is an intelligence upgrade, achieving +6 points on the Intelligence Index. It launches with a longer 2M token context window (up from Grok 4’s 256K context window, matching Grok 4.1 Fast’s 2M), and significantly lower pricing ($2/$6 vs Grok 4’s $3/$15). Grok 4.20’s performance lags behind the current intelligence frontier, but it performs strongly on instruction following and features a notably low hallucination rate, beating all other models we’ve tested on AA-Omniscience for hallucination. xAI released 3 variants: reasoning, non-reasoning, and multi-agent. We’ve evaluated the reasoning and non-reasoning modes, and are considering the best approach for testing the new multi-agent functionality, which parallelizes over multiple agents behind the scenes in one API call. Key takeaways: ➤ Improved intelligence over Grok 4: Grok 4.20 Beta 0309 (Reasoning) scores 48 on the Artificial Analysis Intelligence Index, +6 from Grok 4 and +9 compared to Grok 4.1 Fast. This score falls short of the current intelligence frontier at 57 (Gemini 3.1 Pro Preview and GPT-5.4) ➤ Low price for the level of intelligence: Grok 4.20 is priced at $2/$6 per 1M input/output tokens, representing a decrease compared with Grok 4’s $3/$15 API rates. The reasoning variant cost $484 to complete the evaluations in the Artificial Analysis Intelligence Index, which is a reduction of ~70% compared to Grok 4, driven by lower pricing and lower token use ➤ Leading non-hallucination rate: Grok 4.20 scores 78% in the AA-Omniscience non-hallucination metric. This is the best result we have seen yet for this metric, and reflects the model only answering around one fifth of the time when it did not know the answer ➤ Fast inference performance: xAI is serving Grok 4.20 at 267 tokens per second - similar to what we see for gpt-oss-120b across providers and on the Pareto frontier for speed versus intelligence ➤ Mixed improvements in tool use: Grok 4.20 improved on the tool calling performance of Grok 4 in some evaluations, scoring 97% on Tau2-Telecom. However, its score of 1,062 on GDPval-AA, our benchmark of general agent performance on real work tasks, is well behind frontier peers and sits approximately in line with Grok 4.1 Fast
@tetsuoai ·
🚨 xAI is moving fast right now. Grok 4.20 is being reported with major gains in reasoning, multimodal performance, instruction following, speed, and hallucination reduction, with its multi agent architecture. Currently #1 on IFBench at 82.9% Next best scores there are 79.6% and 79.0% Now Grok Imagine 1.0 is out of beta, and a bigger video release is already being teased for next week. xAI is shipping across reasoning, search, image, video, and agents all at once!
@DeRonin_ ·
For now, you can create your own AI agents on Grok 4.20 To do this simply: - click on your pfp in the bottom left corner - click on "Settings" - then click on "Customize" - click on "Create" and build your agent Grok 4.20 itself is built as a multi-agent system that automatically deploys 4 specialized agents But around March 4-7, 2026, xAI released an update that added support for custom agents Now you can: - create your own agents by defining their name, role, personality and instructions - these custom agents participate in multi-agent collaboration (they discuss, fact-check, and work in parallel) the base version still provides 4 agent slots, while SuperGrok Heavy allows scaling the system up to a swarm of 16 agents with deeper configuration the future is already here, and Grok is definitely part of it Source: Grok
@AntoineRSX ·
Hermes Agent + Grok is the the new 🧨 Here is how you can set up Grok with your Ai Agent: (it's free)
@Angaisb_ ·
xAI introduces Grok 4.5, trained alongside Cursor and focused on coding, agentic tasks and office work - 80 tps, faster than flash models according to xAI - 4.2x fewer output tokens than Opus 4.8 on SWE Bench Pro - $2 / $6 per million tokens - 83.3% on Terminal Bench, basically tied with GPT-5.5 and close to Fable - 64.7% on SWE Bench Pro, above GPT-5.5 but far from Fable's 80.4% Available in Grok Build, Cursor and the API. Not in the EU yet, expected mid July
@invideoOfficial ·
Grok Imagine, the AI video model that doesn't want to be real. Since @grok launched its video capabilities, the discourse has been almost entirely content moderation - what will and won't generate. Spicy mode, regulatory investigations, the EUs reaction. It's loud, it's dramatic. And it's completely missing the point. Because buried underneath all of that noise is something genuinely fascinating for filmmakers and visual creators. Grok Imagine video is the only major AI video model that was built to honor imagination over physics. See how it performs in 6 styles versus other models along with detailed prompts 🧵👇
@FGLXLucas ·
xAI’s Grok Limits: Transparency Deficit and Strategic Risk in a Brutal AI Market In recent months, xAI has repeatedly adjusted SuperGrok usage conditions without clear communication. Instead of fixed, contractually guaranteed render and usage capacities, there is now a dynamic weekly shared budget that is quickly exhausted, especially for compute-intensive features like Grok Imagine (images and videos). Users like Luc report that even with efficient usage (e.g., one 30-second film plus two 10-second clips per day), the budget has become noticeably tighter — a development many perceive as yet another 50-70% reduction. The problem goes far beyond casual users. Developers using Grok for complex programming, agent workflows, or long-running analyses are equally affected, as are creators who generate income with the outputs. This becomes particularly critical for future applications such as robotics (Optimus) and mobile AI, where reliable compute planning is essential. Anyone who builds their workflow or even their business model on Grok risks sudden restrictions or hidden upsells (Heavy tier, extra credits). For xAI, this is a risky game. In an extremely competitive AI market where OpenAI, Anthropic, Google, and specialized providers are fighting for developers and companies, trust and predictability are key differentiators. Frequent, opaque limit adjustments create exactly the kind of dependency that can later lead to higher prices or churn. Many professional users are already consciously diversifying and using Grok only for quick tests, while relying on more transparent alternatives for productive work. Conclusion xAI has excited many users with Grok and SuperGrok — but the current limit policy undermines precisely the trust needed for long-term success in the AI industry. For users, this means: diversification and critical review of dependencies are advisable. For xAI: greater transparency and reliable capacities would be a clear competitive advantage rather than a risk. #Grok #SuperGrok #xAI #GrokLimits #AILimits #AITransparency #xAICriticism #GrokImagine #AI #ArtificialIntelligence #TechBusiness #xAIRisks
@mercor_ai ·
Grok 4.5 from @SpaceXAI places #2 on the APEX-SWE leaderboard at 51.2% Pass@1 (±6.0), behind Fable 5 (65.5% ±6.2) on our benchmark for real-world software engineering work. It leads Integration (65.0% Pass@1) and places #2 in Observability (37.3% Pass@1), covering multi-step build tasks and diagnosis/debugging respectively. The Integration lead maps directly to the agentic workflows Grok 4.5 was built for: multi-step coding tasks run in collaboration with Cursor. Grok models have improved 30.2 pp in a year on this benchmark: Grok 4 (21.0% Pass@1) to Grok 4.5 (51.2% Pass@1). Congratulations to the xAI and Cursor teams.
@shushant_l ·
I'm amazed most people still use Grok like a basic chatbot. Here's how to unlock Grok's full potential with 30 powerful prompting hacks. --- 1. Role prompting gives Grok a specific identity for more expert responses. --- 2. Start with your exact goal before adding context or instructions. --- 3. Layer background information step by step to improve relevance. --- 4. Lock your preferred output format like JSON, tables, checklists, or markdown. --- 5. Add clear constraints for tone, audience, topics, and formatting. --- 6. Use few shot prompting with examples to improve response quality. --- 7. Use zero shot prompting for simple and direct tasks. --- 8. Break complex requests into smaller prompts for better accuracy. --- 9. Combine multiple expert roles for deeper analysis. --- 10. Ask Grok to use real time information when current insights matter. --- 11. Request evidence, assumptions, and reasoning behind conclusions. --- 12. Use reverse prompting to uncover missing information before solving. --- 13. Stack multiple personas for richer perspectives and outputs. --- 14. Define your target audience before asking for content. --- 15. Set a quality benchmark for depth, originality, and accuracy. --- 16. Ask Grok to think through complex tasks step by step. --- 17. Make Grok critique and improve its own first response. --- 18. Compare multiple options side by side before deciding. --- 19. Simulate real world scenarios to test strategies safely. --- 20. Build reusable prompt templates for repetitive workflows. --- 21. Generate a rough draft first, then refine it in iterations. --- 22. Use decision matrices to evaluate choices objectively. --- 23. Ask Grok to identify risks, blind spots, and possible failures. --- 24. Generate multiple alternatives instead of accepting the first answer. --- 25. Prioritize instructions so Grok knows what matters most. --- 26. Reference earlier outputs to maintain consistency across long chats. --- 27. Ask Grok to match a specific writing style or communication format. --- 28. Split large documents into smaller sections for better analysis. --- 29. Score outputs against clear evaluation criteria or rubrics. --- 30. Use meta prompting to improve and optimize your prompts themselves. --- To learn more, check the infographic. ---
@haider1 ·
grok 4.3 is a good everyday model, but not for agentic coding yet i believe grok 5 is already training, while xAI will keep releasing grok 4 checkpoints since an xAI dev hinted they're working on making grok more useful as an action-taking tool, grok build may launch around grok 4.4 and give agentic coding a real boost
@icreatelife ·
I'm impressed with Grok as a visual video model, how fast it emerged and keeps improving. I hope every platform will have Grok model as a choice of 3rd party models. I use it a lot for my AI animations.
@Axel_bitblaze69 ·
Grok 4.5 has been out for around 2 weeks, and after using it as my main model for work, here’s the easiest way to understand where it actually fits. yeah if you compare it to Fable 5.. Grok 4.5 is not as smart as it on every task. but i’d say It is probably the best default for most tasks. Even Elon’s positioning is basically this: Fable 5 is still better when the problem is genuinely difficult, but most daily work does not need that level of capability. And after using Grok 4.5, I think that is exactly the point. On Artificial Analysis’ coding-agent tests, Grok 4.5 performs in the same range as GPT-5.5 and below Fable 5 overall. On Terminal-Bench, it scored 83.3% versus 84.3% for Fable and 83.4% for GPT-5.5. So the quality gap is already quite small on practical agentic work. But the cost gap is not small. Average coding-agent task cost: Grok 4.5: $2.49 GPT-5.5: $5.07 Fable 5: $11.80 It also used roughly 1.9M tokens per task, compared with 6.2M for GPT-5.5 and 7.2M for Fable. So you are getting near-frontier performance at roughly one-fifth of Fable’s task cost. That changes how you use it. With an expensive model, you save it for important prompts. With Grok 4.5, you can leave the agent running. Let it inspect the project, search documentation, edit files, run commands, write tests, fix errors and keep iterating without worrying that every loop is burning another $10. This is where I would use it: → day-to-day coding and debugging → terminal tasks and long agent loops → building and testing fast → browser and workflow automation → research that needs live web or X context → repetitive work where the agent needs to keep trying → scripts and tasks you want running in the background For the hardest architecture decisions, complex reasoning or a painful large-codebase refactor where the last few percentage points of quality matter more than cost, I would still reach for Fable 5. That is not a weakness. It is just knowing which tool fits the job. Another thing I like is how quickly the product around the model is improving. Grok Build has been shipping versioned updates frequently, with actual public changelogs instead of vague “performance improvements.” One correction worth making: Grok 4.5 itself is not open source. Grok Build, the terminal agent that runs it, was open-sourced on July 15. You can inspect and modify the harness, use skills, plugins, hooks, MCPs and subagents, or point it toward your own local model. My takeaway after daily-driving it: Stop asking which model is the smartest. Ask which model is capable enough, fast enough and cheap enough to stay inside your workflow all day. For the hardest 5% of tasks, use the strongest model available. For the other 95%, Grok 4.5 is starting to look like one of the best defaults. You can try it through Grok Build with one command: curl -fsSL https://t.co/7z5ptv2drv | bash
@jasondoesstuff ·
Fable 5 release: 90% of X talking about it. GPT-5.6 release: About 25% of X talks about it. Grok 4.5 release: Literally 2 people in my feed 😄. Clear levels of trust in the models. While we may not agree with Anthropic's rollout approach (and rug pull not their fault entirely), they still have consumer trust over the rest. Opus model earned them immense customer retention. Not surprised at all by quietness of Grok 4.5. I've tried other Grok models a few times and it feels like 2024 AI. Gonna need a lot of positive results to earn that trust back.
@pukerrainbrow ·
Your X Premium subscription can now turn into your private AI assistants. Announced May 22, X Premium and Super Grok users can connect Grok directly into OpenClaw. It'll be able to run locally on your own devices instead of sitting on some random corporate cloud, so it’s way more private than most AI assistants out right now. And it plugs into pretty much everything: WhatsApp, iMessage, Slack, Discord… But it’s not just there to chat. It can actually manage tasks, browse the web, organize your calendar, and automate little workflows for you. Since it’s connected to Grok, it can also pull live news and market sentiment straight from X to build custom daily briefings automatically. We’re slowly getting to the point where your AI assistant stops being a chatbot and starts acting like an actual digital employee. W for X, W for Grok, W for OpenClaw.
@VadimStrizheus ·
EVERYONE IS MISSING THE IMPORTANT PART OF ELON’S GROK UPDATE. It is not just “Grok got bigger.” The real story is that Elon said a lot of Cursor data was added into Grok V9’s supplementary training. That matters way more than people think. Grok’s original unfair advantage was real-time X data. When xAI launched Grok, the whole pitch was simple: Most chatbots understand the internet’s past. Grok understands what is happening right now. That was the first data edge. Now Elon is pointing Grok at the next one: how builders actually write software. Cursor did not win because it was “another AI coding tool.” Cursor won because it moved AI from the chat window into the repo. It watched developers work inside real projects. Prompts. Files. Bugs. Refactors. Context. Bad outputs. Good outputs. Accepted changes. Rejected changes. Shipping decisions. That is the difference between training an AI to answer coding questions and training an AI to behave like an engineer. In this clip, Elon explains the actual AI race: compute, rate of improvement, talent, and unique data. That is the whole game. Now look at the pieces: xAI has the compute. Grok has live X data. V9-Medium is a jump from the smaller model serving Grok today. And now Cursor workflow data is being added for difficult coding tasks. This is not Grok trying to become another chatbot. This is Grok trying to become the AI engineer inside X. The next AI war is not who writes the prettiest answer. It is who can understand the repo, reason through the task, write the code, test it, fix it, and ship. That is why the Cursor line matters. That was not a random detail. That was the signal.
@VaibhavSisinty ·
Xai shipped 13 versions of grok build in the last 12 days. nobody is shipping AI coding tools at this pace. Here's what's now sitting inside Xai ↓ → a real command surface. /usage to see your token spend in any session, /login to reauth without leaving the terminal, /export to save the full conversation, /rewind to roll back a turn, /config-agents to manage your subagents from a modal. → plan mode and /execute-plan, so grok can break a big task into a structured step-by-step plan and actually run it. now with git and gt support built in. → subagents that share the same terminal backend, scheduler, and monitor across sessions. you can pause a multi-agent task, walk away, come back, and resume with the full UI and history replayed. → X search and a much faster web search wired straight into the terminal. grok can pull live context from https://t.co/VR4iwFS7OW without you alt-tabbing once. → proper image understanding. images now flow as multimodal vision tokens instead of base64 dumps. you can paste or drop multiple at once, and read_file extracts text from powerpoint files too. → video generation with user-set duration. and the terminal itself now plays back media at 30fps. → an "always-approve" mode for runs where you want grok to stop asking permission and just ship. literally called "yes, and don't ask again for anything." → plugins, MCP servers, and hooks all managed from one picker modal. an actual extension surface, not a feature list. And somewhere in the middle of those 13 versions, they shipped a feature called "laziness detector." Yes that's the actual name. it nudges grok when the model starts skipping steps. Most companies ship a product. xai is shipping a release schedule.
@CodeByNZ ·
👀 Grok 4.5 performs on par with GPT-5.5 and close to Opus 4.8 on coding benchmarks. The model shows strong results across Terminal-Bench, SWE-Bench, and DeepSWE, demonstrating its ability to handle difficult, long-running tasks that require creative tool use across software engineering, data science, finance, legal work, and other domains. This marks a significant upgrade over Grok 4.3, especially in complex, multi-step coding and agentic workflows.
@QuotableCrypto ·
✴️ The @QuotableCrypto rebrand and thoughts on breaking X stagnation... If you're already here but feel stuck because of the brutal X algo "recommendation system" (yeah, the positive name doesn't always match the reality)... Consider this with me: Grok has become my go-to for growth strategies on X.*** It has extremely deep real-time integration with the platform - it sees trends, algo shifts, and public post patterns in ways most outsiders can't. It's native to the app. ✴️ Grok gets to know your X habits, both the good and the ones holding you back... It understands the platform's ever-changing red flags better than anything else I've tried. I sat down with Grok and built a fixed, low-volume plan tailored exactly to @QuotableCrypto - my size, my follower base, my style. It's literally reversing a full year of stagnation after all the endless algo changes. How? ✴️ It helped me create a daily rhythm: specific posting windows, reply schedules, exact amounts per day, content guidance based on what actually performed for me before - all customized, nothing generic... One-size-fits-all "growth" plans from CT rarely stick anymore. YOUR VIBE IS ALREADY ESTABLISHED HERE. The smarter move is doubling down on what works for *you* and fixing what doesn't, using your own performance data. Other AIs are solid as supplements, but for deep X-specific strategy, Grok has felt unbeatable in my experience. All thoughts, experiences, or different approaches are genuinely welcome in the comments - I know I'm not the only one who's been grinding through this... ✴️ Blessings...🙏🏻💪🏻QC ***This is NOT a paid endorsement or affiliation in ANY way. Just sharing something that's made a real personal difference after a year of absolute frustration - I used to post 15–20 quotes a day, shining the spotlight on this great community. What once drove growth became a massive red flag in the new landscape....
@hooeem ·
You pay X every month. The best worker you own is sitting in that subscription, and you use it to summarise X posts or argue with strangers. On 8 July, xAI shipped Grok 4.5. Built for coding, agentic work and knowledge tasks, co-trained with Cursor. Most people read the announcement, thought “another model”, and changed nothing. The people who did change something are have gained a co-founder at no extra cost. They’re working at a different resolution to you. They’re model agnostic. 𝐖𝐡𝐲 𝐮𝐬𝐞 𝐆𝐫𝐨𝐤 𝐚𝐬 𝐲𝐨𝐮𝐫 𝐜𝐨-𝐟𝐨𝐮𝐧𝐝𝐞𝐫: $2 per million input tokens. $6 per million output. Roughly a fifth of Opus 4.8 or Fable 5. 80 tokens per second. Multi-step work lands about twice as fast end to end. Half the steps to the same answer. On SWE Bench Pro, xAI clocks 4.2x fewer output tokens than Opus 4.8 max. Fourth on the intelligence index. Not first. First place wins the hardest single problem on earth. You do not have that problem. You have four hundred small ones and no time. When looking at the long-horizon agentic benchmark, it beats both Opus 4.8 and Fable 5. What does that mean? It holds a plan across dozens of steps without drifting. It finishes so fast you run ten times more jobs than you planned. Token budgets die overnight. Fucking W. If you want to learn how to use it and what tools to use it with read this: 👇
@LaceyPresley ·
The distance between today’s AI and truly useful civilizational intelligence is no longer a matter of speculation. It is being compressed in public, at speed, by xAI. • Grok 4.5 launched as xAI’s strongest model to date, purpose-built for coding, agentic tasks, and professional knowledge work • Trained at large scale with heavy emphasis on curated engineering data and reinforced on multi-step real-world software tasks, developed in collaboration with Cursor • Places on the Pareto frontier for intelligence versus cost and speed alongside only the strongest competing systems • Achieves the highest perfect-extraction rate among tested models on large-scale real-world invoice processing (150,000+ business bills) • Serves at 80 tokens per second while using substantially fewer output tokens than leading competitors on equivalent engineering benchmarks • Rolled out across https://t.co/lGcC28UTvr, X, iOS, and Android within days of release, with expanded autonomous workflow capabilities in Grok Build • Grok 4.6 targeted for release in approximately two weeks and Grok 4.7 in four weeks, maintaining an unusually rapid iteration cadence This is the systematic conversion of compute, data discipline, and engineering focus into deployable intelligence that raises the floor of what individuals and organizations can accomplish.
@glenngabe ·
Regarding Google's AIOs citing Grok, here is some information about Grok's search visibility. The account https://t.co/Skl153h1kO has surged in visibility over the past year, although it did take a hit recently post-core update. It ranks for 87K queries right now based on ahrefs data and 35.3K in the US. From an AIO perspective, Grok ranks in AI Overviews for 21K queries. For example, like the one below in the screenshots. And from an AI Mode pov, Grok has surged since late 2025, but did take a hit post-May core update. That's the final screenshot. So yes, Grok is being cited in Google's AI surfaces quite a bit.
@lemire ·
I wanted to share my experience with Grok Build, xAI’s competitor to Claude Code. I should note upfront that I have considerable experience coding with AI, whatever that means in 2026. The Grok Build console interface appeals to me. It has a somewhat more serious appearance than Claude Code and feels less distracting. By default, it launches in full-screen mode, which gives it a cleaner look, though I recognize this is not a critical factor. What about the coding itself? So far, I have tested it on two different toy projects, plus one evolving real-world implementation. The first was a remake of an old game in modern C++. The project remains incomplete. As the first one I attempted, Grok Build handled it reasonably well at the start, but it eventually became confused by repeated edits to the same file. This occurred during the early days of Grok Build, so I am unsure how representative those issues were. I have continued working on it since then, and Grok Build has performed well. I later realized that the project was too ambitious for a simple test; building something that resembles a playable game simply takes a considerable amount of time. I have since moved on to a web application that serves as a dashboard for my school’s statistical data. It performed superbly, Grok Build handled it like a champ. I do not believe Claude Code would have done any better. It is going to take me some time making it really good, but Grok Build is doing very well. The terminology is confusing because it seems that Grok Build is also the name of a model that you can access through the xAI API. I have been working with it for an actual production application. I have designed a custom system that allows a professor to edit a course website using AI. I rely on Grok Build to modify the HTML, JavaScript, and CSS based on natural-language queries. The application works great, and I plan to share more details about it in the future. Exciting times.
@FabianXR_Builds ·
This is where AI building gets interesting 👀 In the video: - new Grok Builds model test - the dropped Opus 5 version - native iOS app running on an iPhone after the first prompts - then me taking over as the human developer Steering Grok, Opus 5, Claude, Codex, and GPT image generation together feels completely different than just prompting one model and hoping. The model is not the whole workflow anymore. The developer steering the models is. #ai #claude #opus #buildinpublic
@orcdev ·
I tested Grok Build with @grok 4.3, and in just a couple of days the improvement is crazy inside my project where I already had skills and plugins set up for AI agents, I gave it a single prompt and it returned a completely new page with the exact UI from my skills system, plus filters, sorting, featured sections, navbar updates, and multiple extra improvements I never even asked for what will happen when they drop Cursor data + Grok? check out the video 👀
@panditdhamdhere ·
Grok Build in Rust 🦀 Grok Build is @SpaceXAI 's terminal-based AI coding agent. It runs as a full-screen TUI that understands your codebase, edits files, executes shell commands, searches the web, and manages long-running tasks interactively, headlessly for scripting/CI, or embedded in editors via the Agent Client Protocol (ACP). https://t.co/UpLPXczvJw
@BenjaminDEKR ·
Kind of charming: I asked Grok for updates on the Blue Origin deployment and he said "I and the team am monitoring." Who is "the team?" The other Grok Expert Mode AI agents: Harper, Benjamin, and Lucas "We're all working together in real time on queries like this live space mission watch. I’m the team lead, but they help me scan X."
@aylarov ·
Grok + live X search on a phone call. Call the demo: +1 833-440-2265 How-to: https://t.co/egTIWsu726 @xai @grok
Best Tweets by Topic