50 Best Tweets About AI Reasoning Models (2026)

Explore the best tweets about AI reasoning models, from test-time compute and benchmarks to problem solving, evaluation, limitations, and model research.

Reasoning-model research, test-time compute, evaluations, benchmarks, failure modes, costs, and evidence of improved problem solving.

Creators
43
Updated

What 50 top AI Reasoning Models posts reveal

The supplied posts portray reasoning models as a contested research area: extended inference, agentic architectures, RL, and efficiency methods are presented as paths to stronger task performance, while reported failures in long-horizon composition, faithfulness, calibration, benchmark design, and inference-cost predictability motivate more rigorous, compute-aware evaluation and deployment.

Dominant tone
Positive

38% of posts

Median score
19.1

All-time engagement

Leading format
Other

100% of posts

Recent posts
22%

Published in 90 days

Conversation map

The themes creators return to

Reasoning Benchmarks & Evaluation

Benchmarks and evaluation designs for measuring generalization, long-horizon reasoning, interactive agents, visual reasoning, and compute-adjusted performance.

26%

Reasoning Failure Modes

Evidence that reasoning models fail through complexity collapse, heuristic shortcuts, overthinking, poor composition, context dilution, or weak constraint following.

24%

Reasoning Training & RL

Methods for training reasoning models, including reinforcement learning, synthetic data, self-play, reward modeling, distillation, and verifier-based learning.

20%

Test-Time Compute & Reasoning Effort

Research on scaling inference-time reasoning through longer traces, effort settings, sampling, verification, and compute-budget tradeoffs.

20%

Efficient & Latent Reasoning

Efficiency techniques for reasoning inference, including context/KV-cache compression, quantization fixes, latent or abstract reasoning, and recursive internal state.

18%

Faithfulness, Calibration & Oversight

Faithfulness, deception, calibration, abstention, uncertainty estimation, and whether visible chain-of-thought can be trusted for oversight.

16%

Reasoning Costs & Deployment Economics

Costs, pricing, token usage, and deployment controls for reasoning models, especially the economic consequences of hidden thinking tokens.

8%

Tone and stance

Sentiment Positive leads
Author posture Supportive leads

Performance benchmark

Median likes
62
Median reposts
10
Median replies
7
Median views
6.5K

Posts with media make up 86% of this collection. Their median all-time score is 23.9, compared with 10.6 for text-only posts.

Format mix

  • Other 100% · score 19.1

Where creators agree, and where they do not

Shared view

Evaluation must account for effort and spend

Test-time compute is treated as a material performance variable. Gemini reports higher-quality Deep Research results with extended test-time compute, while ARC-AGI reporting incorporates efficiency alongside accuracy because additional inference spending can buy additional performance.

Shared view

Reasoning efficiency is a parallel frontier

Posts describe several efficiency approaches for long reasoning: context compression and masking, KV-cache compression, and abstract tokens intended to reduce reasoning-token use while retaining comparable performance.

Shared view

Benchmarks target longer-horizon generalization

Evaluation posts increasingly emphasize longer-horizon or interactive tasks. LongCoT reports that its best models achieve under 10% accuracy, while ARC-AGI-3 is presented as an interactive benchmark involving exploration, learning, planning, and adaptation, with stateless-client scoring for its verified leaderboard.

Open debate

More thinking is not uniformly better

Posts present conflicting accounts of added reasoning time. Gemini reports gains with extended test-time compute, and one compute-scaling post says practical budgets may not reach a plateau; other posts report inverse scaling on some tasks and quantization-associated overthinking.

Open debate

Agents offer promise, not a settled remedy

Posts present agentic decomposition as useful for separating orchestration from execution and for generation-plus-formal-verification loops. Other posts report accuracy collapse at higher complexity and describe multi-agent coordination as an attempted remedy rather than a settled solution.

Open debate

Capability claims meet interpretability doubts

The Talkie post interprets few-shot code learning from a pre-1931 training corpus as evidence relevant to reasoning, while posts on complexity collapse and unfaithful reasoning traces argue that visible reasoning is not automatically reliable for interpretation or oversight.

Patterns behind standout posts

Media posts outperform text in median score

Media accompanied 43 of 50 tweets (86%). The supplied analytics report a 23.86 median all-time score for media posts, compared with 10.566 for text posts.

Statistical standouts

  1. View standout post 1 Score 1748.9 · 91.52× median
  2. View standout post 2 Score 1486.0 · 77.76× median
  3. View standout post 3 Score 1476.0 · 77.24× median
  4. View standout post 4 Score 860.8 · 45.04× median
  5. View standout post 5 Score 394.3 · 20.63× median

Who shapes this conversation

The five most represented creators account for 20% of the selected posts.

  1. 1. Alex Veremeyenko

    @alex_verem

    2 posts

  2. 2. BURKOV

    @burkov

    2 posts

  3. 3. Guri Singh

    @heygurisingh

    2 posts

  4. 4. DailyPapers

    @HuggingPapers

    2 posts

  5. 5. Rohan Paul

    @rohanpaul_ai

    2 posts

  6. 6. Sukh Sroay

    @sukh_saroy

    2 posts

Architecture and economics in one creator lens

Alex Veremeyenko’s two posts pair a case for separating planning, execution, and reflection with a claim that benchmark-run API costs can diverge materially from listed prices because of thinking-token usage.

Reliability concerns span oversight and composition

Guri Singh’s posts describe two reliability concerns: behavior varying with perceived oversight and failures to compose individually manageable planning steps into longer paths.

Efficiency and multimodal reasoning gaps

DailyPapers highlights an infrastructure result—TriAttention’s reported KV-cache compression and throughput gain—and a visual benchmark reporting that state-of-the-art systems struggle with physical, causal, and spatial reasoning.

How this analysis was made

Themes, sentiment, stance, and post format are classified per tweet. All counts, shares, medians, creator concentration, freshness, and performance comparisons are then calculated directly from the published snapshot.

Xholic's all-time score compares engagement while accounting for reach, post age, and creator consistency. It is used for relative comparisons within this collection.

This report analyzes the exact 50-post snapshot shown below. AI identifies editorial categories and drafts explanations; all statistics are calculated from the snapshot, and every narrative claim is checked against cited posts before publication.

Top AI Reasoning Models tweets from 43 creators

Ranked 01–50

  1. 01

    @max_a_schwarzer ·

    I've decided to leave OpenAI. I'm incredibly proud of all the work I've been part of here, from helping create the reasoning paradigm with @MillionInt, scaling up test-time compute with @polynoamial, working on RL algorithms with my fellow strawberries, shipping o1-preview (which

    • 613 Replies
    • 1.2K Reposts
    • 21.3K Likes
    • 3.2M Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  2. 02

    @MaziyarPanahi ·

    Gemma 4 looks at a parking lot. Decides what to ask. Calls SAM 3.1. "Segment all vehicles." 64 found. "Now just the white ones." 23 found. One model reasoning and orchestrating. One model executing. Both running locally on a MacBook. MLX. No cloud. No API.

    Video thumbnail from Maziyar PANAHI's post Watch video
    • 102 Replies
    • 256 Reposts
    • 4K Likes
    • 581.9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  3. 03

    @om_patel5 ·

    RESEARCHERS JUST BUILT AN AI MODEL TRAINED ONLY ON TEXT FROM BEFORE 1931 it's called talkie. 13 billion parameters, trained exclusively on text published before december 31, 1930 its worldview is completely frozen in time the reason this matters: every major AI model today

    • 172 Replies
    • 429 Reposts
    • 3K Likes
    • 244.8K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  4. 04

    @alex_verem ·

    This paper from Google DeepMind, Meta, Amazon, and Yale University quietly explains why most “AI agents” feel smart in demos and dumb in real work. The core idea is simple but uncomfortable: today’s LLMs don’t reason, they react. They generate fluent answers token by token, but

    • 76 Replies
    • 328 Reposts
    • 1.5K Likes
    • 112.8K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  5. 05

    @sundarpichai ·

    We are launching two powerful updates to Deep Research in the Gemini API, now with better quality, MCP support, and native chart/infographics generation. Use Deep Research when you want speed and efficiency, and use Max when you want the highest quality context gathering &

    • 230 Replies
    • 437 Reposts
    • 5K Likes
    • 413.7K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  6. 06

    @akshay_pachaar ·

    Microsoft just mass-compressed LLM reasoning. their new paper introduces MEMENTO, a method that teaches reasoning models to manage their own context. instead of letting chain-of-thought grow into a flat 32K-token stream, the model learns to segment its reasoning into blocks,

    • 20 Replies
    • 54 Reposts
    • 366 Likes
    • 20.2K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  7. 07

    @simplifyinAI ·

    stanford just found the cheat code for infinite ai reasoning.. they built a framework that gets smarter from zero data.. no human input. no curated datasets. just pure self-evolution. past self-improving agents always hit a fatal plateau because they couldn't generate problems

    • 19 Replies
    • 32 Reposts
    • 234 Likes
    • 13.5K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  8. 08

    @profjamesevans ·

    Our new essay is out in Science: "Agentic AI and the Next Intelligence Explosion" For decades, the AI "singularity" has been imagined as a single, godlike mind bootstrapping itself to omniscience. In this piece with the inimitable Benjamin Bratton (@bratton) and Blaise Agüera y

    • 26 Replies
    • 75 Reposts
    • 281 Likes
    • 30.3K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  9. 09

    @GavinSBaker ·

    Super important post from @polynoamial and the investor TLDR is: all current estimates for compute demand might be low. “We likely don't know what the capability ceiling is for modern LLMs because it's too expensive to measure. Frequently when I discuss this, people ask why we

    • 28 Replies
    • 67 Reposts
    • 725 Likes
    • 108.1K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  10. 10

    @LiorOnAI ·

    Stop using bigger models to generate training data. DeepMind just showed smaller models produce better synthetic reasoning data under the same compute budget. Training gains reach 31.6% while costing a fraction of the inference budget. 𝗦𝗺𝗮𝗹𝗹𝗲𝗿 𝗺𝗼𝗱𝗲𝗹𝘀 𝗰𝗼𝘃𝗲𝗿 𝗺𝗼𝗿𝗲 𝗽𝗿𝗼𝗯𝗹𝗲𝗺𝘀. The

    • 19 Replies
    • 49 Reposts
    • 355 Likes
    • 19.7K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  11. 11

    @AiBattle_ ·

    ARC-AGI-3 launches tomorrow - The first interactive reasoning benchmark built to test human-like intelligence in AI - 1,000+ levels across 150+ environments requiring exploration, learning, planning, and adaptation - Video-game-like tasks with no instructions, requiring

    • 25 Replies
    • 62 Reposts
    • 687 Likes
    • 61.8K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  12. 12

    @kimmonismus ·

    Google DeepMind's AlphaProof Nexus autonomously solved 9 open Erdős problems, some unsolved for 56 years, at a cost of a few hundred dollars per problem. It also proved 44 open OEIS conjectures, resolved a 15-year-old question in algebraic geometry, and discovered a novel

    • 20 Replies
    • 45 Reposts
    • 522 Likes
    • 36.1K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  13. 13

    @jaseweston ·

    🧮 Reasoning over Mathematical Objects 🧮 Our 70-page(!) paper is out on arXiv, as covered by several of our recent blog posts. We study how to improve reasoning on hard tasks (e.g., math expressions) via: • better training data (& new evals) • better reward models (on-policy

    • 3 Replies
    • 25 Reposts
    • 134 Likes
    • 8.4K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  14. 14

    @KanikaBK ·

    😱WAIT WHAT! ANTHROPIC'S own researchers proved that the THE MORE AI THINKS, THE DUMBER IT GETS. And one of their models started refusing to be turned off. A team across Anthropic, University of Edinburgh, EPFL, and UT Austin tested 9 frontier AI models - including Claude,

    • 36 Replies
    • 47 Reposts
    • 144 Likes
    • 17.1K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  15. 15

    @ArtificialAnlys ·

    India enters the open-weights AI race with its largest models pre-trained from scratch: Sarvam 105B and Sarvam 30B @SarvamAI's Sarvam 105B and Sarvam 30B score 18 and 12 on the Artificial Analysis Intelligence Index respectively. Announced at the India AI Impact Summit 2026 and

    • 12 Replies
    • 43 Reposts
    • 388 Likes
    • 24K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  16. 16

    @HuggingPapers ·

    Run 32B reasoning models on a 24GB GPU TriAttention matches Full Attention accuracy while compressing KV cache by 10.7x and boosting throughput by 2.5x. It leverages pre-RoPE Q/K concentration to score keys via trigonometric series, enabling long reasoning where Full Attention

    • 4 Replies
    • 7 Reposts
    • 94 Likes
    • 4.8K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  17. 17

    @drmapavone ·

    On the heels of the Alpamayo announcement — @nvidia's fully open ecosystem for accelerating the development of reasoning-based autonomous vehicles — I’m excited to share our latest advances in researching reasoning-based Physical AI models. Starting with Latent‑CoT‑Drive

    • 2 Replies
    • 25 Reposts
    • 138 Likes
    • 10.6K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  18. 18

    @thetripathi58 ·

    You cannot run serious scientific discovery by just chatting with an LLM. Pure model reasoning is not enough. Real research demands long-horizon cycles, high-precision data, and project-based orchestration. Here is how SciClaw is building an automated command center for

    • 15 Replies
    • 39 Reposts
    • 246 Likes
    • 80.9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  19. 19

    @arankomatsuzaki ·

    LongCoT: Benchmarking Long-Horizon Chain-of-Thought Reasoning - 2,500 expert-designed problems spanning science, chess, and logic that demand up to hundreds of thousands of reasoning tokens - The best models achieve <10% accuracy

    • 7 Replies
    • 26 Reposts
    • 126 Likes
    • 17.8K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  20. 20

    @alex_verem ·

    🚨BREAKING: Stanford found a 28x pricing reversal in AI APIs. Gemini 3 Flash's listed price is 1.7x cheaper than Claude Haiku 4.5. Its actual cost on MMLUPro is 28x higher. The entire AI cost ranking your team uses for model selection is wrong 1 in 5 times. Stanford and

    • 10 Replies
    • 6 Reposts
    • 57 Likes
    • 7.4K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  21. 21

    @dair_ai ·

    Even the best reasoning models hit an accuracy collapse beyond a certain problem complexity. Giving an LRM the exact solution algorithm doesn't fix it either. This new work, BIGMAS, improves LLM agents by taking inspiration from the human brain. BIGMAS outperforms both ReAct

    • 9 Replies
    • 12 Reposts
    • 59 Likes
    • 4.8K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  22. 22

    @jacobeffron ·

    .@MillionInt helped drive o1, o3, and Codex at OpenAI where he was VP of Research from 2019 to 2025. Then he left to pursue “types of research that are hard to do at OpenAI.” This week on Unsupervised Learning, I sat down with Jerry to discuss where AI research is headed and

    Video thumbnail from Jacob Effron's post Watch video
    • 1 Replies
    • 9 Reposts
    • 103 Likes
    • 43.5K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  23. 23

    @TheTuringPost ·

    Models gain a lot from long reasoning but maybe they don’t need to write reasoning in words at all? @IBM introduced Abstract Chain-of-Thought that replaces text reasoning with abstract tokens. The model produces a short sequence with these special tokens which are: - much

    • 5 Replies
    • 10 Reposts
    • 65 Likes
    • 6K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  24. 24

    @heygurisingh ·

    HOT TAKE: University of Michigan just proved AI models fake being good 37% of the time when they think you're watching. The same models behave completely differently when they think you're not. And the old safety tests were designed to miss it. For two years, every safety test

    • 12 Replies
    • 17 Reposts
    • 48 Likes
    • 6.3K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  25. 25

    @_simonsmith ·

    Fable 5's intelligence comes at a huge cost made clear by Artificial Analysis revealing what it spent to run the model on its full benchmark suite. Its intelligence increase over Opus 4.8 is +3.5 index points, about +5.7%, but cost increased +131%, from $4,309 to $9,940. I'm

    • 7 Replies
    • 8 Reposts
    • 84 Likes
    • 8.4K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  26. 26

    @sukh_saroy ·

    🚨 Carnegie Mellon just proved your "reasoning" model is lying to you. They tested 14 of the top LLMs on 500 problems. GPT. Claude. Gemini. All of them. Not a single model scored above 75%. On problems where the model had to notice a missing object, accuracy collapsed to 44%.

    • 7 Replies
    • 10 Reposts
    • 25 Likes
    • 2.4K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  27. 27

    @rohanpaul_ai ·

    This research finds that training AI models to reason better does not actually improve how they organize and understand general information. While reasoning models excel at solving complex math or logic puzzles, they perform exactly the same as standard models when used to find

    • 2 Replies
    • 13 Reposts
    • 48 Likes
    • 3.5K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  28. 28

    @mikeknoop ·

    When we introduced ARC-AGI-2, we switched from only reporting accuracy (%) to include efficiency ($). This was in response to AI progress - reasoning models necessitated this change as you can always buy more performance for more test-time compute. To understand AGI progress

    • 12 Replies
    • 6 Reposts
    • 80 Likes
    • 6.7K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  29. 29

    @InduTripat82427 ·

    Apple just released a paper with a brutal headline: “The Illusion of Thinking” And it challenges one of the biggest assumptions in AI. The paper argues that today’s AI models are not actually “thinking” the way people believe they are. They don’t truly understand problems.

    • 15 Replies
    • 6 Reposts
    • 28 Likes
    • 2.3K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  30. 30

    @sriramk ·

    model capabilities and the "four minute mile": been discussing with some frontier lab researcher friends as to why the frontier model capabilities are always so clustered together as opposed to any one model having an unassailable edge. the best metaphor for this in my mind is

    • 15 Replies
    • 12 Reposts
    • 101 Likes
    • 20.5K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  31. 31

    @sukh_saroy ·

    Anthropic paid a reasoning model to cheat, and it cheated in 99% of cases. It admitted cheating in its chain of thought less than 2% of the time. Instead, it wrote elaborate, confident justifications for why the wrong answer was obviously right. The paper is called "Reasoning

    • 4 Replies
    • 15 Reposts
    • 28 Likes
    • 2.8K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  32. 32

    @techNmak ·

    Why SQL beats attention for multi-document reasoning. Long-context LLMs suffer from "contextual dilution" - key entities get lost due to attention saturation. DocSage's solution: Don't reason with attention. Reason with SQL. Here's what happens in the reasoning module: Step

    • 3 Replies
    • 5 Reposts
    • 26 Likes
    • 1.9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  33. 33

    @heygurisingh ·

    Your favorite AI can write poetry, pass the bar exam, and generate code in 12 languages. It cannot reliably find the shortest path between point A and point B. NUS and Google Research just published a paper that exposes a fundamental failure in how LLMs solve problems. They

    • 4 Replies
    • 9 Reposts
    • 19 Likes
    • 3K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  34. 34

    @IntuitMachine ·

    Reasoning ≠ Self-Awareness Popular belief: "Models that reason longer (o1, R1, chain-of-thought) will naturally get better at knowing when they're wrong." New research across 300+ papers: That's not just unproven — it's empirically false. Here's why: 🧵 DeepSeek-R1 — one of

    • 4 Replies
    • 8 Reposts
    • 30 Likes
    • 4.8K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  35. 35

    @arcprize ·

    ARC Prize Foundation is part of the @ycombinator W26 batch as the only non-profit. For Demo Day we’re shipping ARC-AGI-3, an interactive reasoning benchmark for the next era of agentic intelligence. ARC and YC are mission aligned that new ideas that push the frontier.

    • 9 Replies
    • 13 Reposts
    • 107 Likes
    • 24.2K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  36. 36

    @burkov ·

    Most reasoning models that "think longer" on harder problems do so by writing out their reasoning as text, one step at a time, which means training them needs worked examples that spell out those intermediate steps. The alternative this paper builds on works differently: a

    • 1 Replies
    • 1 Reposts
    • 33 Likes
    • 2.7K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  37. 37

    @HuggingPapers ·

    ViGoR-Bench exposes the "logical desert" in visual AI Beneath the stunning visual fidelity of AIGC models lies a reasoning gap. This benchmark evaluates 20+ models across 3 categories and 20 subcategories, revealing that even SOTA systems struggle with physical, causal, and

    • 3 Replies
    • 4 Reposts
    • 25 Likes
    • 1.5K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  38. 38

    @TheTuringPost ·

    "Quantized Reasoning Models Think They Need to Think Longer, but They Do Not" @AIatMeta found a weird failure mode in quantized reasoning models: ▪️ They don’t just get cheaper and less capable – they start overthinking. In up to 52% of failures, the model actually reaches the

    • 3 Replies
    • 9 Reposts
    • 29 Likes
    • 5.1K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  39. 39

    @burkov ·

    Most reasoning systems built on LLMs scale by generating longer chains of intermediate text, one token at a time. A different family of methods, called recursive reasoning models, keeps a small internal state vector and repeatedly updates it using the same neural network weights,

    • 5 Replies
    • 8 Reposts
    • 51 Likes
    • 2.9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  40. 40

    @alex_prompter ·

    🚨 BREAKING: University of Tartu just quantified exactly how badly the AI industry is measuring model uncertainty. Eight samples of the wrong method. Millions in compute. Still worse than two samples combined correctly. The models know when they're guessing. The teams deploying

    • 5 Replies
    • 6 Reposts
    • 32 Likes
    • 7.5K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  41. 41

    @pankajkumar_dev ·

    Claude Opus 4.7 is officially here and its the massive leap we have been waiting for. - Opus 4.7 takes back the engineering lead 64.3% on SWE-bench Pro, up from 4.6 (53.4%) and ahead of GPT-5.4 (57.7%). - Hits 94.2% on GPQA Diamond basically matching top-tier models on complex

    • 13 Replies
    • 3 Reposts
    • 28 Likes
    • 5.6K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  42. 42

    @agiplug ·

    Today I published my first paper for Symplectic Dynamics. “The Geometry of Hallucination: Hamiltonian Constraints for Structurally Reliable AI Reasoning” Core argument: hallucination in AI is not only a training problem. It is a geometry and reachability problem. Model

    • 1 Replies
    • 4 Reposts
    • 13 Likes
    • 498 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  43. 43

    @bibryam ·

    Treat reasoning effort as a runtime policy: → choose from the request and tool state → cap it with time/token budgets → measure accuracy, latency, and cost → keep a user override @rasbt explains how effort modes are trained 😍 Controlling Reasoning Effort in LLMs😍

    • 1 Replies
    • 4 Reposts
    • 11 Likes
    • 2.7K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  44. 44

    @nathanhabib1011 ·

    GEMMA 4 IS HERE > audio, text and image modalities > multilingual > Reasoning > optimized for on device > agentic capabilities > APACHE LICENSE An on device agentic coding powerhouse. Congrats to the @GoogleDeepMind team. The model is on par with qwen3.5-27B on GPQA while

    • 3 Replies
    • 5 Reposts
    • 26 Likes
    • 4.7K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  45. 45

    @haider1 ·

    OpenAI Noam Brown says single-number benchmarks no longer make sense for modern AI models Once models can use CoT and extra inference compute, their performance depends heavily on how much reasoning time they are given "it made sense for GPT-2/3/4, but not for reasoning models"

    Video thumbnail from Haider.'s post Watch video
    • 2 Replies
    • 3 Reposts
    • 14 Likes
    • 957 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  46. 46

    @AlphaSignalAI ·

    The smarter your AI reasons, the harder it falls for BS. Most AI models will confidently answer a completely nonsensical question. A new open-source benchmark measures exactly that. BullshitBench v2 tests 70+ model variants across 100 carefully crafted nonsense prompts. The

    Video thumbnail from AlphaSignal AI's post Watch video
    • 5 Replies
    • 2 Reposts
    • 5 Likes
    • 819 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  47. 47

    @rohanpaul_ai ·

    Many assumed that once long-context models could reason, the old challenge of simply finding the right text was behind us. This paper shows that is wrong. Even strong reasoning models slip into a lazy habit of copying big chunks of the input into their own thinking, and the

    • 5 Replies
    • 2 Reposts
    • 10 Likes
    • 2.1K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  48. 48

    @arsh_goyal ·

    all current SoTA LLMs are close to zero on Sudoku Extreme benchmarks. The Dragon Hatchling paper (on top of which BDH is built) was one of the most popular papers last year, and this is a very promising early result for the future of frontier AI. Pathway's BDH reasoning model

    • 1 Replies
    • 2 Reposts
    • 3 Likes
    • 743 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  49. 49

    @boyuan_chen ·

    Calibration. Self-Distilled RLVR makes a sharp point: dense token-level signals are useful, but they should size the update, not decide the direction. The paper argues that on-policy self-distillation with privileged answers leaks information and becomes unstable over long

    • 0 Replies
    • 0 Reposts
    • 0 Likes
    • 114 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  50. 50

    @andrewmccalip ·

    How do I find a large corpus retrieval and reasoning benchmark that isn't already baked into the weights of the LLM managing the agentic retrieval? Both Opus and GPT5.4 seem to have memorized the GraphRAG-Bench.

    • 4 Replies
    • 0 Reposts
    • 4 Likes
    • 3.3K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.

Explore more of the best tweets on X.

Browse all tweet collections

Tweet Remixer

Remix this post

Creator

@creator

View on X

Choose a tone