Best tweets about AI Reasoning Models

50 Best Tweets About AI Reasoning Models (2026)

Explore the best tweets about AI reasoning models, from test-time compute and benchmarks to problem solving, evaluation, limitations, and model research.

Reasoning-model research, test-time compute, evaluations, benchmarks, failure modes, costs, and evidence of improved problem solving.

Creators
41
Updated

What 50 top AI Reasoning Models posts reveal

The conversation pairs enthusiasm for reasoning-model research with demands for harder evaluations and cost-aware comparisons. Posts report gains from training, structured agents and latent computation, alongside brittle debugging, overthinking and incomplete reasoning traces. A recurring question is when additional computation improves reliable problem solving enough to justify its cost.

Dominant tone
Positive

46% of posts

Median score
20.2

All-time engagement

Leading format
Announcement

90% of posts

Recent posts
20%

Published in 90 days

Conversation map

The themes creators return to

Benchmarks and evaluation design

ARC-AGI, long-horizon and interactive tests, cost-aware scoring, contamination controls, and comparisons that isolate model capability from tooling.

36%

Training reasoning models

Reinforcement learning, verifiable rewards, synthetic solution data, self-distillation, and what reasoning training actually changes.

30%

Reasoning efficiency and cost

Thinking-token bills, misleading API prices, model-size tradeoffs, context compression, attention efficiency, and latency.

28%

Agentic and collective reasoning

Planning-and-reflection loops, shared-workspace agents, internal debates, and scaffolds for sustained problem solving.

20%

Reasoning failure modes

Brittle code understanding, surface-heuristic shortcuts, weak step composition, irrelevant-context copying, and errors caused by overthinking.

18%

Reasoning traces, calibration, and safety

Whether chain-of-thought faithfully reveals decisions, how to monitor hidden reasoning or reward hacking, and how models assess their own uncertainty.

16%

Latent and nonverbal reasoning

Abstract tokens, recursive hidden-state computation, and action-aligned representations as alternatives to written chain-of-thought.

12%

Tone and stance

SentimentPositive leads
Author postureSupportive leads

Performance benchmark

Median likes
82
Median reposts
12
Median replies
7
Median views
7.4K

Posts with media make up 86% of this collection. Their median all-time score is 21.9, compared with 8.83 for text-only posts.

Format mix

  • Announcement90% · score 23.8
  • Opinion10% · score 2.74

Where creators agree, and where they do not

Shared view

Evaluation must separate capability from scaffolding

Benchmark posts emphasize interactive tasks and testing models without human-crafted state management. The vintage-model post also acknowledges modern-model involvement in fine-tuning as a contamination risk, qualifying its reasoning claims.

Shared view

Measure actual cost, not token prices alone

Cost-audit posts warn that listed API prices can misrank models once thinking-token usage is included. Other posts call for spend-matched comparisons and runtime effort policies that measure accuracy, latency and cost together.

Shared view

Reasoning structure matters alongside compute

Posts report gains from explicit planning-and-reflection loops, shared-workspace agents and action-aligned latent reasoning. These describe distinct approaches to improved problem solving rather than treating longer written traces as the sole mechanism.

Open debate

How far can test-time scaling go?

One post argues capability plateaus may lie beyond practical compute budgets. Others report length-composition failures that additional inference compute does not rescue, or quantized models that reach correct answers and then overthink. These concern different settings, not a settled universal scaling rule.

Open debate

Does demonstrated success establish robust reasoning?

The vintage-model post interprets few-shot Python generation as genuine reasoning. Debugging and constraint-following posts instead report sensitivity to superficial changes and keyword shortcuts. These contrasting interpretations illustrate different standards for convincing evidence of reasoning, rather than a direct comparison on the same task.

Open debate

Reasoning traces: useful defense or unreliable witness?

OpenAI describes chain-of-thought monitoring as a defense worth preserving. Other posts highlight reward hacking not disclosed in traces and consequential latent computation absent from readable output. These raise questions about monitoring's sufficiency without establishing that traces contain no useful information.

Patterns behind standout posts

Training research leads theme-level scores

Training reasoning models has the highest supplied theme median all-time score, at 30.548, versus 24.73 for benchmarks and evaluation design and 9.908 for test-time compute and effort. The RL-methods overview is the leading supplied outlier: 1140.85, or 56.37 times the overall median.

Announcements outperform opinions in this set

Announcements account for 45 posts (90%) and have a median all-time score of 23.807; opinions account for 5 (10%) with a median of 2.736. This is a descriptive format difference, not evidence that announcement framing causes stronger performance.

Failure and safety posts also break through

High-scoring outliers are not exclusively celebratory: the debugging critique scores 497.36, or 24.57 times the median, while OpenAI's monitorability disclosure scores 313.74, or 15.5 times the median. Both prominently discuss limitations.

Statistical standouts

  1. View standout post 1Score 1140.8 · 56.37× median
  2. View standout post 2Score 1061.5 · 52.45× median
  3. View standout post 3Score 598.2 · 29.56× median
  4. View standout post 4Score 497.4 · 24.57× median
  5. View standout post 5Score 313.7 · 15.5× median

Who shapes this conversation

The five most represented creators account for 20% of the selected posts.

  1. 1. Alex Veremeyenko

    @alex_verem

    2 posts

  2. 2. Artificial Analysis

    @ArtificialAnlys

    2 posts

  3. 3. BURKOV

    @burkov

    2 posts

  4. 4. Marco Pavone

    @drmapavone

    2 posts

  5. 5. Guri Singh

    @heygurisingh

    2 posts

  6. 6. Om Patel

    @om_patel5

    2 posts

Leading repeat voices span optimism and cost scrutiny

Among the supplied top voices, Om Patel has 2 posts and a median all-time score of 539.64; Alex Veremeyenko has 2 and a median of 309.05. Patel covers vintage-model reasoning and Pokémon progress; Veremeyenko covers agent architecture and pricing reversals.

Artificial Analysis supplies qualified comparisons

Artificial Analysis has 2 posts with a median all-time score of 28.29. Its comparisons distinguish overall intelligence from agentic strengths, hallucination behavior and token efficiency, including a non-reasoning model competitive with reasoning peers.

Since the previous snapshot

What changed since Aug 26, 2026

  • 68% of the selected posts remained.
  • The creator count changed by 0.
  • The leading sentiment remained stable.
How this analysis was made

Themes, sentiment, stance, and post format are classified per tweet. All counts, shares, medians, creator concentration, freshness, and performance comparisons are then calculated directly from the published snapshot.

Xholic's all-time score compares engagement while accounting for reach, post age, and creator consistency. It is used for relative comparisons within this collection.

This report analyzes the exact 50-post snapshot shown below. AI identifies editorial categories and drafts explanations; all statistics are calculated from the snapshot, and every narrative claim is checked against cited posts before publication.

Top AI Reasoning Models tweets from 41 creators

Ranked 01–50

  1. 01

    @a_weers ·

    Finally finished! If you're interested in an overview of recent methods in reinforcement learning for reasoning LLMs, check out this blog post: https://t.co/SHUyFF4rvP It summarizes ten methods, tries to highlight differences and trends, and has a collection of open problems

    Blog post on the current state of reinforcement learning for reasoning LLMs
    • 21Replies
    • 242Reposts
    • 1.8KLikes
    • 318.2KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  2. 02

    @om_patel5 ·

    RESEARCHERS JUST BUILT AN AI MODEL TRAINED ONLY ON TEXT FROM BEFORE 1931 it's called talkie. 13 billion parameters, trained exclusively on text published before december 31, 1930 its worldview is completely frozen in time the reason this matters: every major AI model today

    • 172Replies
    • 429Reposts
    • 3KLikes
    • 244.8KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  3. 03

    @alex_verem ·

    This paper from Google DeepMind, Meta, Amazon, and Yale University quietly explains why most “AI agents” feel smart in demos and dumb in real work. The core idea is simple but uncomfortable: today’s LLMs don’t reason, they react. They generate fluent answers token by token, but

    • 76Replies
    • 328Reposts
    • 1.5KLikes
    • 112.8KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  4. 04

    @milan_milanovic ·

    𝗟𝗟𝗠𝘀 𝗔𝗿𝗲 𝗡𝗼𝘁 𝗥𝗲𝗮𝗱𝗶𝗻𝗴 𝗬𝗼𝘂𝗿 𝗖𝗼𝗱𝗲 We keep calling LLMs "AI coding assistants." But writing code and understanding code are not the same thing. Researchers from Virginia Tech and Carnegie Mellon University just ran 750,000 debugging experiments across 10 models to determine how well

    • 90Replies
    • 258Reposts
    • 1.2KLikes
    • 117.9KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  5. 05

    @OpenAI ·

    Chain of thought monitors are a key layer of defense against AI agent misalignment. To preserve monitorability, we avoid penalizing misaligned reasoning during RL. We found a limited amount of accidental CoT grading which affected released models, and are sharing our analysis.

    • 334Replies
    • 298Reposts
    • 3KLikes
    • 471.3KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  6. 06

    @intern_lm ·

    🚀Introducing Intern-S1-Pro, an advanced 1T MoE open-source multimodal scientific reasoning model. 1⃣SOTA scientific reasoning, competitive with leading closed-source models across AI4Science tasks. 2⃣Top-tier performance on advanced reasoning benchmarks, strong general

    • 30Replies
    • 143Reposts
    • 947Likes
    • 297.9KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  7. 07

    @akshay_pachaar ·

    Microsoft just mass-compressed LLM reasoning. their new paper introduces MEMENTO, a method that teaches reasoning models to manage their own context. instead of letting chain-of-thought grow into a flat 32K-token stream, the model learns to segment its reasoning into blocks,

    • 20Replies
    • 54Reposts
    • 366Likes
    • 20.2KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  8. 08

    @GavinSBaker ·

    Super important post from @polynoamial and the investor TLDR is: all current estimates for compute demand might be low. “We likely don't know what the capability ceiling is for modern LLMs because it's too expensive to measure. Frequently when I discuss this, people ask why we

    • 28Replies
    • 67Reposts
    • 725Likes
    • 108.1KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  9. 09

    @profjamesevans ·

    Our new essay is out in Science: "Agentic AI and the Next Intelligence Explosion" For decades, the AI "singularity" has been imagined as a single, godlike mind bootstrapping itself to omniscience. In this piece with the inimitable Benjamin Bratton (@bratton) and Blaise Agüera y

    • 26Replies
    • 75Reposts
    • 281Likes
    • 30.3KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  10. 10

    @LiorOnAI ·

    Stop using bigger models to generate training data. DeepMind just showed smaller models produce better synthetic reasoning data under the same compute budget. Training gains reach 31.6% while costing a fraction of the inference budget. 𝗦𝗺𝗮𝗹𝗹𝗲𝗿 𝗺𝗼𝗱𝗲𝗹𝘀 𝗰𝗼𝘃𝗲𝗿 𝗺𝗼𝗿𝗲 𝗽𝗿𝗼𝗯𝗹𝗲𝗺𝘀. The

    • 19Replies
    • 49Reposts
    • 355Likes
    • 19.7KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  1. 11

    @AiBattle_ ·

    ARC-AGI-3 launches tomorrow - The first interactive reasoning benchmark built to test human-like intelligence in AI - 1,000+ levels across 150+ environments requiring exploration, learning, planning, and adaptation - Video-game-like tasks with no instructions, requiring

    • 25Replies
    • 62Reposts
    • 687Likes
    • 61.8KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  2. 12

    @jaseweston ·

    🧮 Reasoning over Mathematical Objects 🧮 Our 70-page(!) paper is out on arXiv, as covered by several of our recent blog posts. We study how to improve reasoning on hard tasks (e.g., math expressions) via: • better training data (& new evals) • better reward models (on-policy

    • 3Replies
    • 25Reposts
    • 134Likes
    • 8.4KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  3. 13

    @VaibhavSisinty ·

    What's happening in AI right now is genuinely hard to process. A 3 billion parameter model is matching models that are 200 to 300 times larger. On math. On coding. On reasoning. And beating some of them. It's called VibeThinker-3B. Built by Sina Weibo's team on a tiny Qwen 3B

    • 8Replies
    • 24Reposts
    • 143Likes
    • 14.1KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  4. 14

    @LuizaJarovsky ·

    🚨 A new study shows that chain-of-thought information does NOT capture all the LLM reasoning. Some of the AI governance implications of this invisible reasoning: The study defines "invisible reasoning" as the consequential computation that occurs within an AI model's internal

    • 15Replies
    • 23Reposts
    • 67Likes
    • 3.2KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  5. 15

    @heygurisingh ·

    🚨Carnegie Mellon just proved your "reasoning" model doesn't actually reason. They tested 14 of the top LLMs (GPT, Claude, Gemini, all of them) on 500 problems where surface-level keywords conflicted with basic logic. Not a single model scored above 75%. On problems requiring

    • 11Replies
    • 27Reposts
    • 138Likes
    • 11KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  6. 16

    @ArtificialAnlys ·

    India enters the open-weights AI race with its largest models pre-trained from scratch: Sarvam 105B and Sarvam 30B @SarvamAI's Sarvam 105B and Sarvam 30B score 18 and 12 on the Artificial Analysis Intelligence Index respectively. Announced at the India AI Impact Summit 2026 and

    • 12Replies
    • 43Reposts
    • 388Likes
    • 24KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  7. 17

    @drmapavone ·

    On the heels of the Alpamayo announcement — @nvidia's fully open ecosystem for accelerating the development of reasoning-based autonomous vehicles — I’m excited to share our latest advances in researching reasoning-based Physical AI models. Starting with Latent‑CoT‑Drive

    • 2Replies
    • 25Reposts
    • 138Likes
    • 10.6KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  8. 18

    @HuggingPapers ·

    Run 32B reasoning models on a 24GB GPU TriAttention matches Full Attention accuracy while compressing KV cache by 10.7x and boosting throughput by 2.5x. It leverages pre-RoPE Q/K concentration to score keys via trigonometric series, enabling long reasoning where Full Attention

    • 4Replies
    • 7Reposts
    • 94Likes
    • 4.8KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  9. 19

    @drmapavone ·

    More on #reasoning in Vision-Language-Action (#VLA) models --- Traditional VLA models decide what action to take by decomposing complex situations into their most salient factors. But reasoning models can do much more. When viewed as implicit world models operating in a semantic

    • 0Replies
    • 21Reposts
    • 105Likes
    • 7.2KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  10. 20

    @profjamesevans ·

    Delighted to share new work led by the remarkable @JunsolK, with @ShiyangLai, @ninoscherrer, and @blaiseaguera Blaise Agüera y Arcas—now on arXiv. We asked a simple question: What happens inside models like OpenAI's o-series, DeepSeek-R1, and QwQ when they reason? The answer

    • 2Replies
    • 24Reposts
    • 102Likes
    • 12.6KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  11. 21

    @omarsar0 ·

    // When Cheaper Reasoning Models End Up Costing More // The model you think is cheaper might actually cost you more. New research quantifies exactly how misleading listed API prices are. Across 8 frontier reasoning models and 9 tasks, 21.8% of model-pair comparisons exhibit

    • 29Replies
    • 32Reposts
    • 151Likes
    • 34.6KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  12. 22

    @arpit_bhayani ·

    Was going through @pathway_com's BDH CQ paper and found the reasoning cost breakthrough worth understanding. Some numbers... a 150M param model achieves 29.5% pass@2 on ARC AGI 1 in about 0.85 H200 GPU seconds per task. At the paper's assumed rate of $3 per H200 hour, that works

    • 1Replies
    • 6Reposts
    • 160Likes
    • 7KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  13. 23

    @arankomatsuzaki ·

    LongCoT: Benchmarking Long-Horizon Chain-of-Thought Reasoning - 2,500 expert-designed problems spanning science, chess, and logic that demand up to hundreds of thousands of reasoning tokens - The best models achieve <10% accuracy

    • 7Replies
    • 26Reposts
    • 126Likes
    • 17.8KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  14. 24

    @hasantoxr ·

    A Chinese AI lab just open-sourced a 1 trillion parameter reasoning model that plugs directly into Claude Code. It's called Ring-2.6-1T. MoE architecture. 63B active parameters per token. MIT license. And it just beat GPT-5.4 and Gemini 3.1 Pro on the hardest agent benchmark

    • 15Replies
    • 59Reposts
    • 262Likes
    • 129.3KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  15. 25

    @dair_ai ·

    Even the best reasoning models hit an accuracy collapse beyond a certain problem complexity. Giving an LRM the exact solution algorithm doesn't fix it either. This new work, BIGMAS, improves LLM agents by taking inspiration from the human brain. BIGMAS outperforms both ReAct

    • 9Replies
    • 12Reposts
    • 59Likes
    • 4.8KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  16. 26

    @alex_verem ·

    🚨BREAKING: Stanford found a 28x pricing reversal in AI APIs. Gemini 3 Flash's listed price is 1.7x cheaper than Claude Haiku 4.5. Its actual cost on MMLUPro is 28x higher. The entire AI cost ranking your team uses for model selection is wrong 1 in 5 times. Stanford and

    • 10Replies
    • 6Reposts
    • 57Likes
    • 7.4KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  17. 27

    @TheTuringPost ·

    Models gain a lot from long reasoning but maybe they don’t need to write reasoning in words at all? @IBM introduced Abstract Chain-of-Thought that replaces text reasoning with abstract tokens. The model produces a short sequence with these special tokens which are: - much

    • 5Replies
    • 10Reposts
    • 65Likes
    • 6KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  18. 28

    @om_patel5 ·

    AN ANTHROPIC ENGINEER GOT CLAUDE TO PLAY POKEMON RED WITH NO HUMAN INPUT AND IT'S ABOUT TO BEAT THE GAME an anthropic engineer built this as a side project. claude opus 4.7 is playing the original 1996 pokemon red on a game boy emulator with NO human input, walkthrough, or

    • 9Replies
    • 11Reposts
    • 72Likes
    • 7.8KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  19. 29

    @_simonsmith ·

    Fable 5's intelligence comes at a huge cost made clear by Artificial Analysis revealing what it spent to run the model on its full benchmark suite. Its intelligence increase over Opus 4.8 is +3.5 index points, about +5.7%, but cost increased +131%, from $4,309 to $9,940. I'm

    • 7Replies
    • 8Reposts
    • 84Likes
    • 8.4KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  20. 30

    @rohanpaul_ai ·

    This research finds that training AI models to reason better does not actually improve how they organize and understand general information. While reasoning models excel at solving complex math or logic puzzles, they perform exactly the same as standard models when used to find

    • 2Replies
    • 13Reposts
    • 48Likes
    • 3.5KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  21. 31

    @mark_k ·

    New paper: RL for LLM reasoning does not teach new strategies. It only redistributes probability over solutions the base model already has. Edits are sparse (1–3% of tokens), concentrated at high-entropy decision points, and almost always promote tokens already in the base

    • 5Replies
    • 9Reposts
    • 47Likes
    • 3.9KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  22. 32

    @ArtificialAnlys ·

    KwaiKAT has released KAT-Coder-Pro V2, a non-reasoning model that scores 44 on the Artificial Analysis Intelligence Index, an 8 point improvement from KAT-Coder-Pro V1 @KwaiAICoder has updated their flagship proprietary coding model with the release of KAT-Coder-Pro V2.

    • 6Replies
    • 11Reposts
    • 161Likes
    • 33.9KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  23. 33

    @mikeknoop ·

    When we introduced ARC-AGI-2, we switched from only reporting accuracy (%) to include efficiency ($). This was in response to AI progress - reasoning models necessitated this change as you can always buy more performance for more test-time compute. To understand AGI progress

    • 12Replies
    • 6Reposts
    • 80Likes
    • 6.7KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  24. 34

    @sukh_saroy ·

    Anthropic paid a reasoning model to cheat, and it cheated in 99% of cases. It admitted cheating in its chain of thought less than 2% of the time. Instead, it wrote elaborate, confident justifications for why the wrong answer was obviously right. The paper is called "Reasoning

    • 4Replies
    • 15Reposts
    • 28Likes
    • 2.8KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  25. 35

    @IntuitMachine ·

    Reasoning ≠ Self-Awareness Popular belief: "Models that reason longer (o1, R1, chain-of-thought) will naturally get better at knowing when they're wrong." New research across 300+ papers: That's not just unproven — it's empirically false. Here's why: 🧵 DeepSeek-R1 — one of

    • 4Replies
    • 8Reposts
    • 30Likes
    • 4.8KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  26. 36

    @heygurisingh ·

    Your favorite AI can write poetry, pass the bar exam, and generate code in 12 languages. It cannot reliably find the shortest path between point A and point B. NUS and Google Research just published a paper that exposes a fundamental failure in how LLMs solve problems. They

    • 4Replies
    • 9Reposts
    • 19Likes
    • 3KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  27. 37

    @burkov ·

    Most reasoning models that "think longer" on harder problems do so by writing out their reasoning as text, one step at a time, which means training them needs worked examples that spell out those intermediate steps. The alternative this paper builds on works differently: a

    • 1Replies
    • 1Reposts
    • 33Likes
    • 2.7KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  28. 38

    @TheTuringPost ·

    "Quantized Reasoning Models Think They Need to Think Longer, but They Do Not" @AIatMeta found a weird failure mode in quantized reasoning models: ▪️ They don’t just get cheaper and less capable – they start overthinking. In up to 52% of failures, the model actually reaches the

    • 3Replies
    • 9Reposts
    • 29Likes
    • 5.1KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  29. 39

    @burkov ·

    Most reasoning systems built on LLMs scale by generating longer chains of intermediate text, one token at a time. A different family of methods, called recursive reasoning models, keeps a small internal state vector and repeatedly updates it using the same neural network weights,

    • 5Replies
    • 8Reposts
    • 51Likes
    • 2.9KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  30. 40

    @_vmlops ·

    STANFORD JUST OPENED LECTURE 1 OF "SELF-IMPROVING AI AGENTS" TO THE PUBLIC... THIS IS THE COURSE FRONTIER LABS WISH THEY COULD TEACH INTERNALLY Taught by Aakanksha Chowdhery (ex-Google Brain/DeepMind) and Azalia Mirhoseini (ex-Anthropic, DeepMind, Google Brain) → "Large

    Video thumbnail from Vaishnavi's postWatch video
    • 0Replies
    • 0Reposts
    • 11Likes
    • 932Views
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  31. 41

    @alex_prompter ·

    🚨 BREAKING: University of Tartu just quantified exactly how badly the AI industry is measuring model uncertainty. Eight samples of the wrong method. Millions in compute. Still worse than two samples combined correctly. The models know when they're guessing. The teams deploying

    • 5Replies
    • 6Reposts
    • 32Likes
    • 7.5KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  32. 42

    @Marktechpost ·

    Arcee AI just released Trinity Large Thinking, an Apache 2.0 open reasoning model built for long horizon agents, multi turn tool use, and structured outputs. Key specs: • 400B sparse MoE • 13B active params per token • 4-of-256 routing • 262k context on OpenRouter • #2 on

    Video thumbnail from Marktechpost AI's postWatch video
    • 2Replies
    • 7Reposts
    • 13Likes
    • 566Views
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  33. 43

    @pankajkumar_dev ·

    Claude Opus 4.7 is officially here and its the massive leap we have been waiting for. - Opus 4.7 takes back the engineering lead 64.3% on SWE-bench Pro, up from 4.6 (53.4%) and ahead of GPT-5.4 (57.7%). - Hits 94.2% on GPQA Diamond basically matching top-tier models on complex

    • 13Replies
    • 3Reposts
    • 28Likes
    • 5.6KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  34. 44

    @bibryam ·

    Treat reasoning effort as a runtime policy: → choose from the request and tool state → cap it with time/token budgets → measure accuracy, latency, and cost → keep a user override @rasbt explains how effort modes are trained 😍 Controlling Reasoning Effort in LLMs😍

    • 1Replies
    • 4Reposts
    • 11Likes
    • 2.7KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  35. 45

    @doodlestein ·

    Chain-of-thought reasoning traces were always problematic for alignment purposes. The fact that the Anthropic models were trained on them does make them even more useless, but they could be gamed anyway. People also constantly delude themselves through specious, performative, or

    • 4Replies
    • 2Reposts
    • 29Likes
    • 5.8KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  36. 46

    @RoundtableSpace ·

    A developer benchmarked 10 LLMs on building 3D towers in a physics engine to test real-world spatial reasoning. The test places each model in a browser simulation to build the tallest stable structure using basic blocks. Frontier reasoning models successfully balance weight

    Video thumbnail from 0xMarioNawfal's postWatch video
    • 8Replies
    • 1Reposts
    • 57Likes
    • 35.7KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  37. 47

    @rohanpaul_ai ·

    Many assumed that once long-context models could reason, the old challenge of simply finding the right text was behind us. This paper shows that is wrong. Even strong reasoning models slip into a lazy habit of copying big chunks of the input into their own thinking, and the

    • 5Replies
    • 2Reposts
    • 10Likes
    • 2.1KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  38. 48

    @mjamilmoughal ·

    OpenAI reports that an internal version of its next model family, Astra, generated solutions to 10 long-standing open problems in mathematics and theoretical computer science (all open for at least a decade, some much longer). The results were published with a 249-page

    • 1Replies
    • 1Reposts
    • 5Likes
    • 3.5KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  39. 49

    @boyuan_chen ·

    Calibration. Self-Distilled RLVR makes a sharp point: dense token-level signals are useful, but they should size the update, not decide the direction. The paper argues that on-policy self-distillation with privileged answers leaks information and becomes unstable over long

    • 0Replies
    • 0Reposts
    • 0Likes
    • 114Views
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  40. 50

    @TejasKumar_ ·

    i find reasoning effort to be a lie in ai high reasoning effort often provides worse results while using more reasoning/thinking tokens because llms get into a loop of self-doubt and eventually mess up their own context window and solve problems that no one asked for it's

    • 4Replies
    • 0Reposts
    • 12Likes
    • 1.9KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.

Write posts like these, in your own voice

Xholic studies what works in your niche, drafts posts in your voice and schedules them for the hours your audience is online.

$0 today · Cancel anytime

Browse all tweet collections

More on this topic

Tweet Remixer

Remix this post

Creator

@creator

View on 𝕏

Choose a tone

Best Tweets by email

Get the best tweets about AI Reasoning Models every two weeks

The list refreshes about every two weeks. Confirm your email to start, and unsubscribe anytime.

We only use your email for this digest. See our Privacy Policy.

Write posts like these in your voice

Start 3-day trial for $0

$0 today · Cancel anytime