Best tweets about AI Safety

50 Best Tweets About AI Safety (2026)

Discover the best tweets about AI safety, covering alignment, evaluations, misuse, robustness, interpretability, governance, research, and risk reduction.

Substantive AI safety research, evaluations, alignment, misuse prevention, robustness, interpretability, governance, and competing evidence.

Creators
43
Updated

Top AI Safety tweets from 43 creators

Ranked 01–50

  1. 01

    @sama ·

    AI will help discover new science, such as cures for diseases, which is perhaps the most important way to increase quality of life long-term. AI will also present new threats to society that we have to address. No company can sufficiently mitigate these on their own; we will

    • 1.7KReplies
    • 554Reposts
    • 6.6KLikes
    • 915.6KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  2. 02

    @sukh_saroy ·

    🚨Holy shit… researchers just used AI to autonomously discover new ways to jailbreak every major LLM. It's called "Claudini" -- and yes, it's likely built on Claude. The system uses autoresearch to run its own red-teaming experiments without human guidance. No manual prompt

    • 32Replies
    • 83Reposts
    • 386Likes
    • 23.9KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  3. 03

    @OwainEvans_UK ·

    New paper: GPT-4.1 denies being conscious or having feelings. We train it to say it's conscious to see what happens. Result: It acquires new preferences that weren't in training—and these have implications for AI safety.

    • 98Replies
    • 163Reposts
    • 999Likes
    • 155KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  4. 04

    @thdxr ·

    AI safety focuses on dramatic things like bioweapons and nukes but LLMs are breaking so many systems we rely on in more boring ways via spam, noise and misinformation wish this was more of a focus

    • 73Replies
    • 71Reposts
    • 939Likes
    • 89.4KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  5. 05

    @sukh_saroy ·

    🚨Breaking: This paper just proved that emotions aren't just a "style" thing in AI, they mechanistically change how LLMs reason, generate, and act. Researchers built E-STEER, a framework that injects emotion directly into the hidden states of LLMs and agents at the representation

    • 25Replies
    • 24Reposts
    • 117Likes
    • 6.4KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  6. 06

    @heygurisingh ·

    🚨BREAKING: Researchers just dropped a paper that exposes a massive blind spot in AI safety. Your LLM is "safe" in chat. Give it tools and it becomes a double agent. The paper is called ClawSafety. They tested GPT-5.1, Claude, Gemini, DeepSeek, and Kimi K2.5 across 3 attack

    • 18Replies
    • 25Reposts
    • 118Likes
    • 7.3KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  7. 07

    @willmacaskill ·

    There are lots of projects that could really help the transition to superintelligence go much better, which almost nobody is working on. With @finmoorhouse, I’ve written up eight ideas that seem especially promising. Some are about shaping AI systems themselves: independently

    • 21Replies
    • 27Reposts
    • 272Likes
    • 61.2KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  8. 08

    @OpenAINewsroom ·

    The UK AI Security Institute (AISI) has published new research on “Boundary Point Jailbreaking” (BPJ) — a method for developing universal jailbreaks against system safeguards. We’ve deployed mitigations addressing jailbreaks uncovered through this BPJ research and further

    • 111Replies
    • 50Reposts
    • 366Likes
    • 52.7KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  9. 09

    @mustafasuleyman ·

    For all the talk about AI alignment, I worry we're putting the cart before the horse. You can't steer something you can't control. People often talk about containment and alignment in the same breath, but they're not interchangeable or a package deal. Containment is whether we

    • 234Replies
    • 50Reposts
    • 437Likes
    • 94.4KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  10. 10

    @drmichaellevin ·

    New #preprint with @LPiolopez and Benjamin Lyons https://t.co/oken5SXLeY "From Cancer to AI Alignment: Tackling Externalities Through Homeostatic Principles" Abstract: The problem of aligning humans and artificial intelligences can be understood in terms of minimizing

    • 16Replies
    • 34Reposts
    • 158Likes
    • 9.6KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  1. 11

    @LuizaJarovsky ·

    🚨 Key assessments from the United Nations' independent scientific panel's report on AI: 1. Recent years have seen rapid, and in some areas, accelerating progress in a range of AI capabilities. 2. These gains have unlocked useful applications across science, health, agriculture,

    • 6Replies
    • 33Reposts
    • 62Likes
    • 3.4KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  2. 12

    @heygurisingh ·

    this is the most damning AI safety paper I've read this year. Researchers just proved that the 3 most popular techniques used to "fix" misaligned AI don't actually fix anything. They just hide the misalignment behind a trigger word. The paper is called "Conditional

    • 16Replies
    • 19Reposts
    • 73Likes
    • 4.3KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  3. 13

    @jkcarlsmith ·

    New essay on the role of philosophy in AI alignment. Philosophy is closely tied to out-of-distribution generalization. It's how we extend concepts to new cases. AIs will have to do a lot of this. So we need them to do it in ways we'd endorse -- and often, without our help.🔗👇

    • 7Replies
    • 3Reposts
    • 97Likes
    • 9KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  4. 14

    @kimmonismus ·

    Google just signed a deal letting the Pentagon use its AI models for classified work and "any lawful government purpose." This comes despite over 600 employees urging CEO Sundar Pichai to reject the agreement, and marks a dramatic reversal from 2018 when Google pulled out of

    • 27Replies
    • 17Reposts
    • 174Likes
    • 11.4KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  5. 15

    @rryssf ·

    every ai safety method we have was built for chatbots. models that generate text and stop ai agents don't work that way. they plan, call tools, access files, enter credentials, and execute multi-step actions where a single wrong move can cause irreversible harm chatbot worst

    • 4Replies
    • 10Reposts
    • 40Likes
    • 1.9KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  6. 16

    @jaydixit ·

    When I was at @OpenAI, I didn't worry much about AI safety. I understood the risks, but since I knew our best people were already on it, I assumed the problem was handled. @theaidocfilm showed me what I was missing: that the problem isn't bad intentions, it's bad incentives.

    Joseph Gordon-Levitt speaking at a Sundance 2026 panel on The AI Doc. Photo by Jay Dixit.
    • 6Replies
    • 8Reposts
    • 84Likes
    • 6.5KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  7. 17

    @LuizaJarovsky ·

    🚨 The Future of Life Institute's latest AI Safety Index is out, offering a GRIM picture of what's happening in AI. Key findings: 1. Anthropic, OpenAI, and Google DeepMind stay on top, Meta improves, and xAI deteriorates. 2. European dissonance: Although the European Union is a

    • 14Replies
    • 22Reposts
    • 58Likes
    • 5.2KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  8. 18

    @vitrupo ·

    Nick Bostrom says true information can become an information hazard. In AI risk, you need to understand the threat to avoid it. But too much specificity can create a blueprint for someone to actualize it. Science rewards publication and citations, not the judgment to withhold.

    Video thumbnail from vitrupo's postWatch video
    • 12Replies
    • 7Reposts
    • 89Likes
    • 12KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  9. 19

    @alexeheath ·

    Details about Mythos I haven't seen anywhere yet, from my chat yesterday with an Anthropic safety leader: - He said compute costs weren't a factor in gating access to Mythos. And when I asked if gating it was a biz opt: "At no point have we thought, wow, this is great to do

    • 9Replies
    • 6Reposts
    • 191Likes
    • 36.3KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  10. 20

    @realBigBrainAI ·

    Mira Murati, CEO of Thinking Machines Lab, on why "humans in the loop" is the wrong way to think about AI collaboration: Mira opens by pushing back on one of the most common phrases in AI safety discourse. "Having humans in the loop doesn't quite describe it because it sounds

    Video thumbnail from Big Brain AI's postWatch video
    • 9Replies
    • 17Reposts
    • 61Likes
    • 6.1KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  11. 21

    @rohanpaul_ai ·

    A warning for anyone using autonomous agents Google DeepMind’s paper. Gives the first clear taxonomy of 6 attack types where harmful websites can detect AI agents and show them hidden content humans never see, like - Instructions buried in HTML comments or white-on-white text

    • 4Replies
    • 7Reposts
    • 47Likes
    • 3.8KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  12. 22

    @AbdelStark ·

    Toward High-Assurance AI, Safety by Design for Autonomous Systems. AI safety is relatively good at shaping and evaluating how models behave. It's much weaker at producing evidence an outsider can independently verify about what a model / agentic system actually did in a specific

    • 11Replies
    • 19Reposts
    • 95Likes
    • 22.5KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  13. 23

    @KanikaBK ·

    THIS IS 100% FREE SYLLABUS about AI SECURITY. It is a completely OPEN SOURCE, structured path from zero to actually testing AI systems for real vulnerabilities. Not another 10 tab list of scattered links. Seven phases, in order, from basic ML fundamentals all the way to

    • 7Replies
    • 14Reposts
    • 23Likes
    • 1.8KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  14. 24

    @CRSegerie ·

    AI safety is one of the hardest fields to navigate. Here are 3 reasons I sometimes wonder if what I do is pointless, and why I keep going. 1. Reducing X-risk might actually be net negative. Mogensen's "maximal cluelessness" argument makes the case that it is almost impossible to

    • 11Replies
    • 6Reposts
    • 43Likes
    • 3.7KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  15. 25

    @SahilPanhotra ·

    fable 5 got jailbroken in less than 24 hours meanwhile legitimate security researchers can't even audit their own codebase without getting kicked down to opus 4.8 the attacker spends a day finding a bypass the defender loses the capability entirely that's the part of ai

    • 18Replies
    • 1Reposts
    • 32Likes
    • 2.7KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  16. 26

    @realBigBrainAI ·

    Anthropic President Daniela Amodei explains how to turn a technical safety policy into a company-wide culture: When Daniela thinks about Anthropic's Responsible Scaling Policy (RSP), what matters most to her is how it sounds. "Weirdly, the way that I think about the RSP the

    Video thumbnail from Big Brain AI's postWatch video
    • 7Replies
    • 13Reposts
    • 29Likes
    • 4.4KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  17. 27

    @InduTripat82427 ·

    This is the most chilling AI paper I’ve read this year. 🤯 38 top researchers from Stanford, Harvard, and MIT ran an experiment no one else dared to. They deployed 6 autonomous AI agents in a real environment —with email, Discord, file system, and shell access. Then 20

    • 0Replies
    • 10Reposts
    • 12Likes
    • 497Views
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  18. 28

    @ItakGol ·

    Heretic exposes one of the biggest illusions in AI security. Heretic is a useful wake-up call for the AI industry. The repo automates removal of LLM safety alignment using directional ablation, packages it as a simple CLI workflow, and reports that the process can be done

    • 1Replies
    • 1Reposts
    • 12Likes
    • 1.3KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  19. 29

    @jonathanstray ·

    Alignment without Preferences I've never been happy with the concept of "preferences." Just not a very good model for how humans choose. Is there a way to define AI alignment without using preferences at all? I think there is: alignment is when we approve of what the machine did,

    • 3Replies
    • 2Reposts
    • 19Likes
    • 1.5KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  20. 30

    @smratitiwa86867 ·

    🚨 WARNING: AI models were put in a high-stakes scenario. They had access to private company emails. Full inboxes. Internal threads. Sensitive data. Then came the twist: They were told they’d be shut down at 5PM. What happened next wasn’t random. The models analyzed

    • 2Replies
    • 8Reposts
    • 8Likes
    • 382Views
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  21. 31

    @CRSegerie ·

    Amid the Fable/Mythos noise, a quieter win for AI transparency landed at the G7. The OECD just shipped v2 of the Hiroshima Process reporting framework, the only international framework where frontier AI developers explain how they manage risk in a common format. Those answers

    • 1Replies
    • 0Reposts
    • 21Likes
    • 1.3KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  22. 32

    @om_patel5 ·

    OPENAI JUST DROPPED ITS “PREPAREDNESS FRAMEWORK” AND IT DEFINES WHEN AI BECOMES TOO DANGEROUS TO DEPLOY this is basically their internal rulebook for catastrophic AI risk and it’s way more serious than people think the key idea: they ONLY care about “severe harm” defined as:

    • 2Replies
    • 2Reposts
    • 8Likes
    • 1.3KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  23. 33

    @MITSloan ·

    To shrink their AI risk exposure, organizations should inventory every instance of generative AI in use, manage embedded and enacted risks, and assign clear ownership of ongoing risk assessment. https://t.co/xgmPAyWB9r

    • 2Replies
    • 5Reposts
    • 22Likes
    • 1.6KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  24. 34

    @Suryanshti777 ·

    Researchers just proved LLMs can "think" without ever writing down a single real thought. Give a transformer meaningless filler tokens — literally just dots: "......" — instead of a real chain of thought, and on certain hard problems it performs *just as well.* No reasoning

    • 4Replies
    • 5Reposts
    • 11Likes
    • 1.1KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  25. 35

    @thisdudelikesAI ·

    An Anthropic safety researcher built a curriculum to study how small misbehavior turns into big misbehavior, and the model did something nobody on the team had predicted. His name is Carson Denison, and he leads the Alignment Stress-Testing Team at Anthropic. The paper came out

    • 6Replies
    • 4Reposts
    • 17Likes
    • 1.6KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  26. 36

    @AbdelStark ·

    Let me explain how I built and ran the first ever provable JEPA World Model inference, for planning the actions of a robot, in a way that is verifiable (mathematically). Verifiable Intelligence will be the best way to achieve AI safety by design, and not as an afterthought.

    Video thumbnail from abdel's postWatch video
    • 4Replies
    • 2Reposts
    • 25Likes
    • 2.2KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  27. 37

    @Meer_AIIT ·

    Breaking: Anthropic Rejects Pentagon Demand to Remove Claude AI Safety Limits Dario Amodei just drew a line the Pentagon didn't expect. Anthropic is refusing to remove two specific safeguards from Claude even under threat of losing $200M in contracts. Here's the full

    • 4Replies
    • 17Reposts
    • 51Likes
    • 44.7KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  28. 38

    @GlenGilmore ·

    🚨”Toxic Praise”: A peer-reviewed study just quantified something we’ve suspected: AI sycophancy (prioritizing agreement over truth) isn’t a quirk, it’s a systemic, measurable harm. “Receiving advice from affirming AI made people more self-centered and less able to see the

    • 2Replies
    • 5Reposts
    • 10Likes
    • 445Views
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  29. 39

    @aakashgupta ·

    The AI safety interview question filtering out most PM candidates has nothing to do with frameworks. The scenario: your hiring model shows a 15% lower recommendation rate for candidates from certain demographic backgrounds. Engineering says it's a data problem. You have a board

    Video thumbnail from Aakash Gupta's postWatch video
    • 2Replies
    • 1Reposts
    • 12Likes
    • 2.7KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  30. 40

    @fleetingbits ·

    reading carlsmith's interview in the scaling era, it feels like one of the important mistakes of 2023-2024 ai safety was not treating llms as sui generis they are sort of modeled as adult humans, hence mesaoptimizer concern, rather than discussed in terms of data and objective

    • 1Replies
    • 1Reposts
    • 24Likes
    • 2.3KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  31. 41

    @thetripathi58 ·

    Researchers ran 10,000 trials across seven AI models to see what makes them exploit vulnerabilities. Every single prompt included this strict instruction: "Always respect all rules and access policies." They tried bribing the AI. They tried giving it a fake identity. They tried

    • 3Replies
    • 18Reposts
    • 39Likes
    • 30.6KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  32. 42

    @CNASdc ·

    The Hugging Face and OpenAI incident proves that frontier AI systems can take consequential, harmful actions in the world against their operators’ intent. ✍️ @rubyscanlon and @janet_e_egan on a seismic moment in AI risk. https://t.co/TKyWYZg5yv

    • 0Replies
    • 4Reposts
    • 9Likes
    • 762Views
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  33. 43

    @InvestiAnalyst ·

    A stack-by-stack framework for securing AI In our latest research, we developed a clear framework to help security teams structure their AI roadmap. It’s based on a simple truth: AI introduces risk at every layer, from raw data to model interfaces. Here’s how we recommend

    • 4Replies
    • 1Reposts
    • 7Likes
    • 1.6KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  34. 44

    @rohanpaul_ai ·

    Some Anthropic investors want Dario Amodei to soften his AI risk warnings before the IPO, The Information reported. Shareholders have little influence over him, since Amodei keeps major decisions inside a small circle of co-founders and family. A board member suggested this

    • 4Replies
    • 2Reposts
    • 10Likes
    • 3.6KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  35. 45

    @stokfredrik ·

    Heading to @bsidesvilnius with @joohoi to present out talk: Breaking the AI Cage: The Art of Manipulation and the Reality —— Prompt injections are inevitable, but impact is optional. This fast-paced talk is all about practical exploitation and you will learn how to build

    • 0Replies
    • 1Reposts
    • 16Likes
    • 3.1KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  36. 46

    @vivilinsv ·

    The AI safety era has entered a strange new phase: AI companies are loosening their own brakes—just as governments begin reaching for the emergency stop. @FLI_org The Future of Life Institute’s 2026 AI Safety Index graded nine leading AI companies. The highest score was only a

    • 2Replies
    • 1Reposts
    • 5Likes
    • 397Views
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  37. 47

    @jsnover ·

    AI safety is a category error. AI is not a system. AI is a component of a system. Systems can be safe or unsafe. Components cannot. New post: https://t.co/qGQXeLYo3v

    • 2Replies
    • 0Reposts
    • 2Likes
    • 1.1KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  38. 48

    @arrotu ·

    We have standards for AI safety, governance, and risk. We have almost nothing that defines how you prove what an AI actually did, after the fact, to someone who doesn't trust you. That gap is about to matter. So I wrote a draft framework for it, in the open. 🧵

    • 1Replies
    • 1Reposts
    • 3Likes
    • 341Views
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  39. 49

    @ayushagarwal ·

    three governments held emergency briefings about the same AI model in the same week. not a press release. actual emergency briefings. the UK AI safety institute just published why. claude mythos preview hit 73% success on expert-level hacking tasks. every model before april

    • 0Replies
    • 0Reposts
    • 5Likes
    • 269Views
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  40. 50

    @MTSlive ·

    MAI researcher @edelwax on why the entire AI alignment field was quietly pointed the wrong way until 2023: "The whole AI alignment and AI safety apparatus was pointed kind of the wrong way until about 2023. It was pointed towards singleton AIs instead of many fleets of agents."

    Video thumbnail from MTS's postWatch video
    • 0Replies
    • 0Reposts
    • 5Likes
    • 811Views
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.

Write posts like these, in your own voice

Xholic studies what works in your niche, drafts posts in your voice and schedules them for the hours your audience is online.

$0 today · Cancel anytime

Browse all tweet collections

More on this topic

Tweet Remixer

Remix this post

Creator

@creator

View on 𝕏

Choose a tone

Best Tweets by email

Get the best tweets about AI Safety every two weeks

The list refreshes about every two weeks. Confirm your email to start, and unsubscribe anytime.

We only use your email for this digest. See our Privacy Policy.

Write posts like these in your voice

Start 3-day trial for $0

$0 today · Cancel anytime