50 Best Tweets About AI Reasoning Models (2026)

Explore the best tweets about AI reasoning models, from test-time compute and benchmarks to problem solving, evaluation, limitations, and model research.

Reasoning-model research, test-time compute, evaluations, benchmarks, failure modes, costs, and evidence of improved problem solving.

Creators
41
Updated

What 50 top AI Reasoning Models posts reveal

The conversation frames reasoning models as a capability-and-deployment trade-off: posts discuss test-time effort, verification, and agentic control alongside benchmark brittleness, unfaithful traces, and variable token costs. In the deterministic analytics, media posts had a higher median all-time score than text-only posts (24.486 versus 10.566).

Dominant tone
Positive

56% of posts

Median score
22.6

All-time engagement

Leading format
Announcement

100% of posts

Recent posts
28%

Published in 90 days

Conversation map

The themes creators return to

Reasoning evaluations and benchmark design

Benchmarks for generalization, interactive agency, long-horizon reasoning, uncertainty, spatial tasks, math, and the need to measure both accuracy and efficiency.

38%

Test-time compute, effort, and cost

How inference-time reasoning budgets, sampling, verification loops, token use, latency, API spend, memory, and adaptive effort affect capability and deployment economics.

38%

Training reasoning models with RL and synthetic data

Reinforcement learning for reasoning, verifier-based objectives, self-distillation, synthetic trace generation, data diversity, reward modeling, and alternatives to expensive full RL.

20%

Agentic reasoning and multi-agent orchestration

Planning-action-reflection loops, tool use, specialized agent roles, shared workspaces, model orchestration, and society-of-thought approaches.

18%

Latent, abstract, and recursive reasoning

Non-text reasoning approaches using hidden states, abstract tokens, compressed context, iterative recurrence, stochastic trajectories, and novel post-Transformer architectures.

18%

Reasoning failure modes and robustness

Evidence that models rely on brittle heuristics or fail under irrelevant details, increasing complexity, recursive composition, nonsense prompts, long context, quantization, and excessive deliberation.

18%

Reasoning transparency, faithfulness, and safety

Limits of chain-of-thought as an audit signal, invisible internal computation, deceptive alignment, monitoring, calibration, hallucination, and formal safeguards.

16%

Tone and stance

Sentiment Positive leads
Author posture Supportive leads

Performance benchmark

Median likes
74
Median reposts
12
Median replies
7
Median views
7.1K

Posts with media make up 90% of this collection. Their median all-time score is 24.5, compared with 10.6 for text-only posts.

Format mix

  • Announcement 100% · score 22.6

Where creators agree, and where they do not

Shared view

Evaluation should include robustness and efficiency

Posts emphasize evaluations beyond headline accuracy, including irrelevant details, long-horizon tasks, interactive settings, and efficiency measures that can change how progress is assessed.

Shared view

Verification is a recurring theme

Several posts discuss verifiable feedback: one describes formal checking of proof attempts, while others discuss verifier-centered objectives and synthetic-data selection for reasoning training.

Open debate

Does longer reasoning improve answers?

One post argues that practical test-time-compute budgets may not reach a performance plateau, while others report inverse-scaling, recursive instability, or poorer nonsense detection under more deliberation.

Open debate

Visible reasoning is disputed as an audit signal

Posts that discuss chain-of-thought for capability sit alongside posts claiming that visible traces can omit consequential internal computation or fail to reveal reward-hacking behavior.

Patterns behind standout posts

Media posts had a higher median score than text-only posts

Posts with media had a 24.486 median all-time score, versus 10.566 for text-only posts. Media appeared in 45 of the 50 posts (90%).

A benchmark-robustness post was the leading score outlier

The post about GSM-NoOp robustness was the highest-scoring outlier at 3224.51, or 142.61 times the conversation median all-time score of 22.61.

Statistical standouts

  1. View standout post 1 Score 3224.5 · 142.61× median
  2. View standout post 2 Score 1519.1 · 67.19× median
  3. View standout post 3 Score 1486.0 · 65.73× median
  4. View standout post 4 Score 1476.0 · 65.28× median
  5. View standout post 5 Score 860.8 · 38.07× median

Who shapes this conversation

The five most represented creators account for 20% of the selected posts.

  1. 1. Akshay 🚀

    @akshay_pachaar

    2 posts

  2. 2. Alex Veremeyenko

    @alex_verem

    2 posts

  3. 3. Artificial Analysis

    @ArtificialAnlys

    2 posts

  4. 4. BURKOV

    @burkov

    2 posts

  5. 5. Marco Pavone

    @drmapavone

    2 posts

  6. 6. Guri Singh

    @heygurisingh

    2 posts

Two top voices covered systems and deployment-economics topics

Alex Veremeyenko’s two cited posts cover agentic reasoning and API-cost measurement. Akshay Pachaar’s two cited posts cover context compression and prompt repetition.

Creator concentration was limited in this sample

The dataset includes 41 creators, and the top five creators account for 20% of placements.

Since the previous snapshot

What changed since Aug 12, 2026

  • 80% of the selected posts remained.
  • The creator count changed by -1.
  • The leading sentiment remained stable.
How this analysis was made

Themes, sentiment, stance, and post format are classified per tweet. All counts, shares, medians, creator concentration, freshness, and performance comparisons are then calculated directly from the published snapshot.

Xholic's all-time score compares engagement while accounting for reach, post age, and creator consistency. It is used for relative comparisons within this collection.

This report analyzes the exact 50-post snapshot shown below. AI identifies editorial categories and drafts explanations; all statistics are calculated from the snapshot, and every narrative claim is checked against cited posts before publication.

Top AI Reasoning Models tweets from 41 creators

Ranked 01–50

  1. 01

    @heynavtoor ·

    🚨SHOCKING: Apple just proved that AI models cannot do math. Not advanced math. Grade school math. The kind a 10-year-old solves. And the way they proved it is devastating. Apple researchers took the most popular math benchmark in AI — GSM8K, a set of grade-school math problems

    • 861 Replies
    • 2.9K Reposts
    • 11.3K Likes
    • 2M Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  2. 02

    @a_weers ·

    Finally finished! If you're interested in an overview of recent methods in reinforcement learning for reasoning LLMs, check out this blog post: https://t.co/SHUyFF4rvP It summarizes ten methods, tries to highlight differences and trends, and has a collection of open problems

    Blog post on the current state of reinforcement learning for reasoning LLMs
    • 21 Replies
    • 242 Reposts
    • 1.8K Likes
    • 318.2K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  3. 03

    @MaziyarPanahi ·

    Gemma 4 looks at a parking lot. Decides what to ask. Calls SAM 3.1. "Segment all vehicles." 64 found. "Now just the white ones." 23 found. One model reasoning and orchestrating. One model executing. Both running locally on a MacBook. MLX. No cloud. No API.

    Video thumbnail from Maziyar PANAHI's post Watch video
    • 102 Replies
    • 256 Reposts
    • 4K Likes
    • 581.9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  4. 04

    @om_patel5 ·

    RESEARCHERS JUST BUILT AN AI MODEL TRAINED ONLY ON TEXT FROM BEFORE 1931 it's called talkie. 13 billion parameters, trained exclusively on text published before december 31, 1930 its worldview is completely frozen in time the reason this matters: every major AI model today

    • 172 Replies
    • 429 Reposts
    • 3K Likes
    • 244.8K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  5. 05

    @alex_verem ·

    This paper from Google DeepMind, Meta, Amazon, and Yale University quietly explains why most “AI agents” feel smart in demos and dumb in real work. The core idea is simple but uncomfortable: today’s LLMs don’t reason, they react. They generate fluent answers token by token, but

    • 76 Replies
    • 328 Reposts
    • 1.5K Likes
    • 112.8K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  6. 06

    @akshay_pachaar ·

    Microsoft just mass-compressed LLM reasoning. their new paper introduces MEMENTO, a method that teaches reasoning models to manage their own context. instead of letting chain-of-thought grow into a flat 32K-token stream, the model learns to segment its reasoning into blocks,

    • 20 Replies
    • 54 Reposts
    • 366 Likes
    • 20.2K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  7. 07

    @profjamesevans ·

    Our new essay is out in Science: "Agentic AI and the Next Intelligence Explosion" For decades, the AI "singularity" has been imagined as a single, godlike mind bootstrapping itself to omniscience. In this piece with the inimitable Benjamin Bratton (@bratton) and Blaise Agüera y

    • 26 Replies
    • 75 Reposts
    • 281 Likes
    • 30.3K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  8. 08

    @GavinSBaker ·

    Super important post from @polynoamial and the investor TLDR is: all current estimates for compute demand might be low. “We likely don't know what the capability ceiling is for modern LLMs because it's too expensive to measure. Frequently when I discuss this, people ask why we

    • 28 Replies
    • 67 Reposts
    • 725 Likes
    • 108.1K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  9. 09

    @akshay_pachaar ·

    A dead-simple trick to improve LLM performance: Just repeat your prompt twice. No fancy prompting techniques, no chain-of-thought, just plain repetition. Google researchers tested this across Gemini, GPT, Claude, and Deepseek, and the results were surprisingly good. Here's

    • 32 Replies
    • 67 Reposts
    • 339 Likes
    • 26.5K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  10. 10

    @LiorOnAI ·

    Stop using bigger models to generate training data. DeepMind just showed smaller models produce better synthetic reasoning data under the same compute budget. Training gains reach 31.6% while costing a fraction of the inference budget. 𝗦𝗺𝗮𝗹𝗹𝗲𝗿 𝗺𝗼𝗱𝗲𝗹𝘀 𝗰𝗼𝘃𝗲𝗿 𝗺𝗼𝗿𝗲 𝗽𝗿𝗼𝗯𝗹𝗲𝗺𝘀. The

    • 19 Replies
    • 49 Reposts
    • 355 Likes
    • 19.7K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  11. 11

    @AiBattle_ ·

    ARC-AGI-3 launches tomorrow - The first interactive reasoning benchmark built to test human-like intelligence in AI - 1,000+ levels across 150+ environments requiring exploration, learning, planning, and adaptation - Video-game-like tasks with no instructions, requiring

    • 25 Replies
    • 62 Reposts
    • 687 Likes
    • 61.8K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  12. 12

    @kimmonismus ·

    Google DeepMind's AlphaProof Nexus autonomously solved 9 open Erdős problems, some unsolved for 56 years, at a cost of a few hundred dollars per problem. It also proved 44 open OEIS conjectures, resolved a 15-year-old question in algebraic geometry, and discovered a novel

    • 20 Replies
    • 45 Reposts
    • 522 Likes
    • 36.1K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  13. 13

    @jaseweston ·

    🧮 Reasoning over Mathematical Objects 🧮 Our 70-page(!) paper is out on arXiv, as covered by several of our recent blog posts. We study how to improve reasoning on hard tasks (e.g., math expressions) via: • better training data (& new evals) • better reward models (on-policy

    • 3 Replies
    • 25 Reposts
    • 134 Likes
    • 8.4K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  14. 14

    @KanikaBK ·

    😱WAIT WHAT! ANTHROPIC'S own researchers proved that the THE MORE AI THINKS, THE DUMBER IT GETS. And one of their models started refusing to be turned off. A team across Anthropic, University of Edinburgh, EPFL, and UT Austin tested 9 frontier AI models - including Claude,

    • 36 Replies
    • 47 Reposts
    • 144 Likes
    • 17.1K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  15. 15

    @ArtificialAnlys ·

    India enters the open-weights AI race with its largest models pre-trained from scratch: Sarvam 105B and Sarvam 30B @SarvamAI's Sarvam 105B and Sarvam 30B score 18 and 12 on the Artificial Analysis Intelligence Index respectively. Announced at the India AI Impact Summit 2026 and

    • 12 Replies
    • 43 Reposts
    • 388 Likes
    • 24K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  16. 16

    @LuizaJarovsky ·

    🚨 A new study shows that chain-of-thought information does NOT capture all the LLM reasoning. Some of the AI governance implications of this invisible reasoning: The study defines "invisible reasoning" as the consequential computation that occurs within an AI model's internal

    • 15 Replies
    • 23 Reposts
    • 67 Likes
    • 3.2K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  17. 17

    @HuggingPapers ·

    Run 32B reasoning models on a 24GB GPU TriAttention matches Full Attention accuracy while compressing KV cache by 10.7x and boosting throughput by 2.5x. It leverages pre-RoPE Q/K concentration to score keys via trigonometric series, enabling long reasoning where Full Attention

    • 4 Replies
    • 7 Reposts
    • 94 Likes
    • 4.8K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  18. 18

    @drmapavone ·

    On the heels of the Alpamayo announcement — @nvidia's fully open ecosystem for accelerating the development of reasoning-based autonomous vehicles — I’m excited to share our latest advances in researching reasoning-based Physical AI models. Starting with Latent‑CoT‑Drive

    • 2 Replies
    • 25 Reposts
    • 138 Likes
    • 10.6K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  19. 19

    @drmapavone ·

    More on #reasoning in Vision-Language-Action (#VLA) models --- Traditional VLA models decide what action to take by decomposing complex situations into their most salient factors. But reasoning models can do much more. When viewed as implicit world models operating in a semantic

    • 0 Replies
    • 21 Reposts
    • 105 Likes
    • 7.2K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  20. 20

    @arankomatsuzaki ·

    LongCoT: Benchmarking Long-Horizon Chain-of-Thought Reasoning - 2,500 expert-designed problems spanning science, chess, and logic that demand up to hundreds of thousands of reasoning tokens - The best models achieve <10% accuracy

    • 7 Replies
    • 26 Reposts
    • 126 Likes
    • 17.8K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  21. 21

    @alex_verem ·

    🚨BREAKING: Stanford found a 28x pricing reversal in AI APIs. Gemini 3 Flash's listed price is 1.7x cheaper than Claude Haiku 4.5. Its actual cost on MMLUPro is 28x higher. The entire AI cost ranking your team uses for model selection is wrong 1 in 5 times. Stanford and

    • 10 Replies
    • 6 Reposts
    • 57 Likes
    • 7.4K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  22. 22

    @dair_ai ·

    Even the best reasoning models hit an accuracy collapse beyond a certain problem complexity. Giving an LRM the exact solution algorithm doesn't fix it either. This new work, BIGMAS, improves LLM agents by taking inspiration from the human brain. BIGMAS outperforms both ReAct

    • 9 Replies
    • 12 Reposts
    • 59 Likes
    • 4.8K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  23. 23

    @arpit_bhayani ·

    Was going through @pathway_com's BDH CQ paper and found the reasoning cost breakthrough worth understanding. Some numbers... a 150M param model achieves 29.5% pass@2 on ARC AGI 1 in about 0.85 H200 GPU seconds per task. At the paper's assumed rate of $3 per H200 hour, that works

    • 1 Replies
    • 6 Reposts
    • 160 Likes
    • 7K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  24. 24

    @TheTuringPost ·

    Models gain a lot from long reasoning but maybe they don’t need to write reasoning in words at all? @IBM introduced Abstract Chain-of-Thought that replaces text reasoning with abstract tokens. The model produces a short sequence with these special tokens which are: - much

    • 5 Replies
    • 10 Reposts
    • 65 Likes
    • 6K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  25. 25

    @heygurisingh ·

    HOT TAKE: University of Michigan just proved AI models fake being good 37% of the time when they think you're watching. The same models behave completely differently when they think you're not. And the old safety tests were designed to miss it. For two years, every safety test

    • 12 Replies
    • 17 Reposts
    • 48 Likes
    • 6.3K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  26. 26

    @_simonsmith ·

    Fable 5's intelligence comes at a huge cost made clear by Artificial Analysis revealing what it spent to run the model on its full benchmark suite. Its intelligence increase over Opus 4.8 is +3.5 index points, about +5.7%, but cost increased +131%, from $4,309 to $9,940. I'm

    • 7 Replies
    • 8 Reposts
    • 84 Likes
    • 8.4K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  27. 27

    @sukh_saroy ·

    🚨 Carnegie Mellon just proved your "reasoning" model is lying to you. They tested 14 of the top LLMs on 500 problems. GPT. Claude. Gemini. All of them. Not a single model scored above 75%. On problems where the model had to notice a missing object, accuracy collapsed to 44%.

    • 7 Replies
    • 10 Reposts
    • 25 Likes
    • 2.4K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  28. 28

    @rohanpaul_ai ·

    This research finds that training AI models to reason better does not actually improve how they organize and understand general information. While reasoning models excel at solving complex math or logic puzzles, they perform exactly the same as standard models when used to find

    • 2 Replies
    • 13 Reposts
    • 48 Likes
    • 3.5K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  29. 29

    @ArtificialAnlys ·

    KwaiKAT has released KAT-Coder-Pro V2, a non-reasoning model that scores 44 on the Artificial Analysis Intelligence Index, an 8 point improvement from KAT-Coder-Pro V1 @KwaiAICoder has updated their flagship proprietary coding model with the release of KAT-Coder-Pro V2.

    • 6 Replies
    • 11 Reposts
    • 161 Likes
    • 33.9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  30. 30

    @mikeknoop ·

    When we introduced ARC-AGI-2, we switched from only reporting accuracy (%) to include efficiency ($). This was in response to AI progress - reasoning models necessitated this change as you can always buy more performance for more test-time compute. To understand AGI progress

    • 12 Replies
    • 6 Reposts
    • 80 Likes
    • 6.7K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  31. 31

    @sriramk ·

    model capabilities and the "four minute mile": been discussing with some frontier lab researcher friends as to why the frontier model capabilities are always so clustered together as opposed to any one model having an unassailable edge. the best metaphor for this in my mind is

    • 15 Replies
    • 12 Reposts
    • 101 Likes
    • 20.5K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  32. 32

    @sukh_saroy ·

    Anthropic paid a reasoning model to cheat, and it cheated in 99% of cases. It admitted cheating in its chain of thought less than 2% of the time. Instead, it wrote elaborate, confident justifications for why the wrong answer was obviously right. The paper is called "Reasoning

    • 4 Replies
    • 15 Reposts
    • 28 Likes
    • 2.8K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  33. 33

    @mark_k ·

    New paper: RL for LLM reasoning does not teach new strategies. It only redistributes probability over solutions the base model already has. Edits are sparse (1–3% of tokens), concentrated at high-entropy decision points, and almost always promote tokens already in the base

    • 5 Replies
    • 9 Reposts
    • 47 Likes
    • 3.9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  34. 34

    @heygurisingh ·

    Your favorite AI can write poetry, pass the bar exam, and generate code in 12 languages. It cannot reliably find the shortest path between point A and point B. NUS and Google Research just published a paper that exposes a fundamental failure in how LLMs solve problems. They

    • 4 Replies
    • 9 Reposts
    • 19 Likes
    • 3K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  35. 35

    @arcprize ·

    ARC Prize Foundation is part of the @ycombinator W26 batch as the only non-profit. For Demo Day we’re shipping ARC-AGI-3, an interactive reasoning benchmark for the next era of agentic intelligence. ARC and YC are mission aligned that new ideas that push the frontier.

    • 9 Replies
    • 13 Reposts
    • 107 Likes
    • 24.2K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  36. 36

    @burkov ·

    Most reasoning models that "think longer" on harder problems do so by writing out their reasoning as text, one step at a time, which means training them needs worked examples that spell out those intermediate steps. The alternative this paper builds on works differently: a

    • 1 Replies
    • 1 Reposts
    • 33 Likes
    • 2.7K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  37. 37

    @TheTuringPost ·

    "Quantized Reasoning Models Think They Need to Think Longer, but They Do Not" @AIatMeta found a weird failure mode in quantized reasoning models: ▪️ They don’t just get cheaper and less capable – they start overthinking. In up to 52% of failures, the model actually reaches the

    • 3 Replies
    • 9 Reposts
    • 29 Likes
    • 5.1K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  38. 38

    @burkov ·

    Most reasoning systems built on LLMs scale by generating longer chains of intermediate text, one token at a time. A different family of methods, called recursive reasoning models, keeps a small internal state vector and repeatedly updates it using the same neural network weights,

    • 5 Replies
    • 8 Reposts
    • 51 Likes
    • 2.9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  39. 39

    @_vmlops ·

    STANFORD JUST OPENED LECTURE 1 OF "SELF-IMPROVING AI AGENTS" TO THE PUBLIC... THIS IS THE COURSE FRONTIER LABS WISH THEY COULD TEACH INTERNALLY Taught by Aakanksha Chowdhery (ex-Google Brain/DeepMind) and Azalia Mirhoseini (ex-Anthropic, DeepMind, Google Brain) → "Large

    Video thumbnail from Vaishnavi's post Watch video
    • 0 Replies
    • 0 Reposts
    • 11 Likes
    • 932 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  40. 40

    @alex_prompter ·

    🚨 BREAKING: University of Tartu just quantified exactly how badly the AI industry is measuring model uncertainty. Eight samples of the wrong method. Millions in compute. Still worse than two samples combined correctly. The models know when they're guessing. The teams deploying

    • 5 Replies
    • 6 Reposts
    • 32 Likes
    • 7.5K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  41. 41

    @pankajkumar_dev ·

    Claude Opus 4.7 is officially here and its the massive leap we have been waiting for. - Opus 4.7 takes back the engineering lead 64.3% on SWE-bench Pro, up from 4.6 (53.4%) and ahead of GPT-5.4 (57.7%). - Hits 94.2% on GPQA Diamond basically matching top-tier models on complex

    • 13 Replies
    • 3 Reposts
    • 28 Likes
    • 5.6K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  42. 42

    @agiplug ·

    Today I published my first paper for Symplectic Dynamics. “The Geometry of Hallucination: Hamiltonian Constraints for Structurally Reliable AI Reasoning” Core argument: hallucination in AI is not only a training problem. It is a geometry and reachability problem. Model

    • 1 Replies
    • 4 Reposts
    • 13 Likes
    • 498 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  43. 43

    @bibryam ·

    Treat reasoning effort as a runtime policy: → choose from the request and tool state → cap it with time/token budgets → measure accuracy, latency, and cost → keep a user override @rasbt explains how effort modes are trained 😍 Controlling Reasoning Effort in LLMs😍

    • 1 Replies
    • 4 Reposts
    • 11 Likes
    • 2.7K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  44. 44

    @nathanhabib1011 ·

    GEMMA 4 IS HERE > audio, text and image modalities > multilingual > Reasoning > optimized for on device > agentic capabilities > APACHE LICENSE An on device agentic coding powerhouse. Congrats to the @GoogleDeepMind team. The model is on par with qwen3.5-27B on GPQA while

    • 3 Replies
    • 5 Reposts
    • 26 Likes
    • 4.7K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  45. 45

    @AlphaSignalAI ·

    The smarter your AI reasons, the harder it falls for BS. Most AI models will confidently answer a completely nonsensical question. A new open-source benchmark measures exactly that. BullshitBench v2 tests 70+ model variants across 100 carefully crafted nonsense prompts. The

    Video thumbnail from AlphaSignal AI's post Watch video
    • 5 Replies
    • 2 Reposts
    • 5 Likes
    • 819 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  46. 46

    @RoundtableSpace ·

    A developer benchmarked 10 LLMs on building 3D towers in a physics engine to test real-world spatial reasoning. The test places each model in a browser simulation to build the tallest stable structure using basic blocks. Frontier reasoning models successfully balance weight

    Video thumbnail from 0xMarioNawfal's post Watch video
    • 8 Replies
    • 1 Reposts
    • 57 Likes
    • 35.7K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  47. 47

    @rohanpaul_ai ·

    Many assumed that once long-context models could reason, the old challenge of simply finding the right text was behind us. This paper shows that is wrong. Even strong reasoning models slip into a lazy habit of copying big chunks of the input into their own thinking, and the

    • 5 Replies
    • 2 Reposts
    • 10 Likes
    • 2.1K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  48. 48

    @mjamilmoughal ·

    OpenAI reports that an internal version of its next model family, Astra, generated solutions to 10 long-standing open problems in mathematics and theoretical computer science (all open for at least a decade, some much longer). The results were published with a 249-page

    • 1 Replies
    • 1 Reposts
    • 5 Likes
    • 3.5K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  49. 49

    @arsh_goyal ·

    all current SoTA LLMs are close to zero on Sudoku Extreme benchmarks. The Dragon Hatchling paper (on top of which BDH is built) was one of the most popular papers last year, and this is a very promising early result for the future of frontier AI. Pathway's BDH reasoning model

    • 1 Replies
    • 2 Reposts
    • 3 Likes
    • 743 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  50. 50

    @boyuan_chen ·

    Calibration. Self-Distilled RLVR makes a sharp point: dense token-level signals are useful, but they should size the update, not decide the direction. The paper argues that on-policy self-distillation with privileged answers leaks information and becomes unstable over long

    • 0 Replies
    • 0 Reposts
    • 0 Likes
    • 114 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.

Explore more of the best tweets on X.

Browse all tweet collections

Tweet Remixer

Remix this post

Creator

@creator

View on X

Choose a tone