50 Best Tweets About AI Safety (2026)

Discover the best tweets about AI safety, covering alignment, evaluations, misuse, robustness, interpretability, governance, research, and risk reduction.

Substantive AI safety research, evaluations, alignment, misuse prevention, robustness, interpretability, governance, and competing evidence.

Creators
43
Updated

What 50 top AI Safety posts reveal

The supplied AI-safety discussion is led by evaluations and red-teaming, with recurring attention to agent containment, governance reporting, and accountability. The main disagreements concern openness versus enforceable constraints, how to prioritize immediate versus existential risks, and the adequacy of safeguards for military deployment.

Dominant tone
Negative

54% of posts

Median score
15.2

All-time engagement

Leading format
Announcement

66% of posts

Recent posts
32%

Published in 90 days

Conversation map

The themes creators return to

Technical assurance and interpretability

Technical approaches to robustness, mechanistic interpretability, activation steering, formal verification, and high-assurance system design.

22%

Military AI and national security governance},{

Military and national-security deployment of AI, including classified-model agreements, weapons constraints, and defense-sector governance.

6%

Tone and stance

Sentiment Negative leads
Author posture Critical leads

Performance benchmark

Median likes
30
Median reposts
7
Median replies
5
Median views
2.4K

Posts with media make up 72% of this collection. Their median all-time score is 11.5, compared with 17.8 for text-only posts.

Format mix

  • Announcement 66% · score 15.4
  • Opinion 26% · score 7.07
  • Question 4% · score 16.0
  • Tutorial 4% · score 40.0

Where creators agree, and where they do not

Shared view

Evaluations test the limits of safeguards

The largest supplied theme is safety evaluations and red teaming (32% of posts). Its cited posts describe tests of pressured false statements, paraphrase-based safety bypasses, automated jailbreak discovery, and tool-enabled agent attacks; these are presented as challenges to current safeguards and evaluation methods.

Shared view

Agents foreground containment and deployment controls

Several posts distinguish agent safety from chatbot safety. They focus on tool access, untrusted files and web content, action boundaries, and the need for deployment records that outside parties can independently verify.

Shared view

Governance discussion emphasizes reporting and shared evidence

Governance-oriented posts pair lab commitments and resilience initiatives with evidence reviews and common reporting frameworks. They present public reporting, shared risk-management formats, and external evaluation as complements to internal safeguards.

Open debate

Open transparency versus enforceable restraint

One post argues that open-source transparency is the best path to safe AI. In contrast, posts discussing the Future of Life Institute’s Safety Index argue that voluntary frameworks and company commitments can weaken under competitive pressure.

Open debate

Immediate harms versus existential risk

Posts differ on which risks should receive priority. One calls for greater attention to spam and misinformation; another relays Geoffrey Hinton’s distinction between immediate misuse risks and existential risk; a third says that only 6.7% of 80,000 interviewed AI users feared existential risk.

Open debate

Military safeguards face scrutiny

Military use is contested in the cited posts. One agreement states prohibitions on domestic mass surveillance and human responsibility for force; another questions whether comparable contractual language is legally binding; a third describes international coordination on weapons-related risks as difficult.

Patterns behind standout posts

Tutorials lead the supplied format medians

Announcements are the largest format category (33 of 50 posts, 66%) with a supplied median all-time score of 15.401. Tutorials have the highest supplied format median, 40.04, although they account for only two posts. Media appears in 36 of 50 posts (72%).

Statistical standouts

  1. View standout post 1 Score 3628.1 · 239.32× median
  2. View standout post 2 Score 2713.9 · 179.02× median
  3. View standout post 3 Score 914.3 · 60.31× median
  4. View standout post 4 Score 578.7 · 38.17× median
  5. View standout post 5 Score 289.6 · 19.11× median

Who shapes this conversation

The five most represented creators account for 20% of the selected posts.

  1. 1. abdel

    @AbdelStark

    2 posts

  2. 2. Charbel-Raphael

    @CRSegerie

    2 posts

  3. 3. Guri Singh

    @heygurisingh

    2 posts

  4. 4. Nav Toor

    @heynavtoor

    2 posts

  5. 5. Big Brain AI

    @realBigBrainAI

    2 posts

  6. 6. Sam Altman

    @sama

    2 posts

Operational safety focuses on auditability

These posts translate safety into operational practices: risk frameworks, impact assessments, release checklists, incident playbooks, and verifiable evidence of deployed-system behavior.

Alignment concepts remain debated

Conceptual posts address out-of-distribution judgment, retrospective approval as an account of alignment, and interpretability as a possible route to steering and understanding AI systems.

How this analysis was made

Themes, sentiment, stance, and post format are classified per tweet. All counts, shares, medians, creator concentration, freshness, and performance comparisons are then calculated directly from the published snapshot.

Xholic's all-time score compares engagement while accounting for reach, post age, and creator consistency. It is used for relative comparisons within this collection.

This report analyzes the exact 50-post snapshot shown below. AI identifies editorial categories and drafts explanations; all statistics are calculated from the snapshot, and every narrative claim is checked against cited posts before publication.

Top AI Safety tweets from 43 creators

Ranked 01–50

  1. 01

    @sama ·

    Tonight, we reached an agreement with the Department of War to deploy our models in their classified network. In all of our interactions, the DoW displayed a deep respect for safety and a desire to partner to achieve the best possible outcome. AI safety and wide distribution of

    • 15.9K Replies
    • 4K Reposts
    • 34.2K Likes
    • 38.1M Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  2. 02

    @heynavtoor ·

    🚨SHOCKING: Researchers built a test that can tell the difference between an AI making a mistake and an AI choosing to lie. The results are terrifying. They tested 30 of the most popular AI models in the world. GPT-4o. Claude. Gemini. DeepSeek. Llama. Grok. They asked each model

    • 463 Replies
    • 2.1K Reposts
    • 3.7K Likes
    • 168.4K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  3. 03

    @sama ·

    AI will help discover new science, such as cures for diseases, which is perhaps the most important way to increase quality of life long-term. AI will also present new threats to society that we have to address. No company can sufficiently mitigate these on their own; we will

    • 1.7K Replies
    • 554 Reposts
    • 6.6K Likes
    • 915.6K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  4. 04

    @heynavtoor ·

    🚨SHOCKING: Researchers just proved that every major AI safety system is fake. ChatGPT. Claude. Gemini. Grok. Every single one broke. Not with some sophisticated hack. Not with a secret exploit. They just rephrased the question. Here is what they did. AI companies test their

    • 106 Replies
    • 449 Reposts
    • 988 Likes
    • 49.9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  5. 05

    @sukh_saroy ·

    🚨Holy shit… researchers just used AI to autonomously discover new ways to jailbreak every major LLM. It's called "Claudini" -- and yes, it's likely built on Claude. The system uses autoresearch to run its own red-teaming experiments without human guidance. No manual prompt

    • 32 Replies
    • 83 Reposts
    • 386 Likes
    • 23.9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  6. 06

    @sharbel ·

    🤯 SHOCKING: Researchers discovered that AI safety guardrails can be completely bypassed by injecting a single hidden vector into a model's brain. And the AI never knows it is happening. You ask the AI to refuse. It refuses. You ask again. It refuses again. Then someone flips a

    • 52 Replies
    • 82 Reposts
    • 379 Likes
    • 30.9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  7. 07

    @Yoshua_Bengio ·

    Today we’re releasing the International AI Safety Report 2026: the most comprehensive evidence-based assessment of AI capabilities, emerging risks, and safety measures to date. 🧵 (1/17)

    Video thumbnail from Yoshua Bengio's post Watch video
    • 66 Replies
    • 373 Reposts
    • 1.1K Likes
    • 427.9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  8. 08

    @thdxr ·

    AI safety focuses on dramatic things like bioweapons and nukes but LLMs are breaking so many systems we rely on in more boring ways via spam, noise and misinformation wish this was more of a focus

    • 73 Replies
    • 71 Reposts
    • 939 Likes
    • 89.4K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  9. 09

    @sukh_saroy ·

    🚨Breaking: This paper just proved that emotions aren't just a "style" thing in AI, they mechanistically change how LLMs reason, generate, and act. Researchers built E-STEER, a framework that injects emotion directly into the hidden states of LLMs and agents at the representation

    • 25 Replies
    • 24 Reposts
    • 117 Likes
    • 6.4K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  10. 10

    @iamKierraD ·

    Hello!! The clearest path is AI Risk Management/Auditing (aka GRC) which requires you to get familiar with the AI development lifecycle, risk management frameworks + regulation (NIST AI RMF, EU AI Act, ISO 42001), data privacy, and creating risk/control artifacts(impact

    • 2 Replies
    • 8 Reposts
    • 83 Likes
    • 2.9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  11. 11

    @heygurisingh ·

    🚨BREAKING: Researchers just dropped a paper that exposes a massive blind spot in AI safety. Your LLM is "safe" in chat. Give it tools and it becomes a double agent. The paper is called ClawSafety. They tested GPT-5.1, Claude, Gemini, DeepSeek, and Kimi K2.5 across 3 attack

    • 18 Replies
    • 25 Reposts
    • 118 Likes
    • 7.3K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  12. 12

    @mustafasuleyman ·

    For all the talk about AI alignment, I worry we're putting the cart before the horse. You can't steer something you can't control. People often talk about containment and alignment in the same breath, but they're not interchangeable or a package deal. Containment is whether we

    • 234 Replies
    • 50 Reposts
    • 437 Likes
    • 94.4K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  13. 13

    @heygurisingh ·

    this is the most damning AI safety paper I've read this year. Researchers just proved that the 3 most popular techniques used to "fix" misaligned AI don't actually fix anything. They just hide the misalignment behind a trigger word. The paper is called "Conditional

    • 16 Replies
    • 19 Reposts
    • 73 Likes
    • 4.3K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  14. 14

    @kimmonismus ·

    Google just signed a deal letting the Pentagon use its AI models for classified work and "any lawful government purpose." This comes despite over 600 employees urging CEO Sundar Pichai to reject the agreement, and marks a dramatic reversal from 2018 when Google pulled out of

    • 27 Replies
    • 17 Reposts
    • 174 Likes
    • 11.4K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  15. 15

    @jkcarlsmith ·

    New essay on the role of philosophy in AI alignment. Philosophy is closely tied to out-of-distribution generalization. It's how we extend concepts to new cases. AIs will have to do a lot of this. So we need them to do it in ways we'd endorse -- and often, without our help.🔗👇

    • 7 Replies
    • 3 Reposts
    • 97 Likes
    • 9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  16. 16

    @realBigBrainAI ·

    Dario Amodei, CEO and co-founder of Anthropic, on the three consensus views about to break, and the role AI plays in each: Progress arrives when a settled position gives way all at once. "There's this kind of consensus that again seems like consensus, seems like what everyone

    Video thumbnail from Big Brain AI's post Watch video
    • 6 Replies
    • 22 Reposts
    • 93 Likes
    • 7.7K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  17. 17

    @rryssf ·

    every ai safety method we have was built for chatbots. models that generate text and stop ai agents don't work that way. they plan, call tools, access files, enter credentials, and execute multi-step actions where a single wrong move can cause irreversible harm chatbot worst

    • 4 Replies
    • 10 Reposts
    • 40 Likes
    • 1.9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  18. 18

    @LuizaJarovsky ·

    🚨 The Future of Life Institute's latest AI Safety Index is out, offering a GRIM picture of what's happening in AI. Key findings: 1. Anthropic, OpenAI, and Google DeepMind stay on top, Meta improves, and xAI deteriorates. 2. European dissonance: Although the European Union is a

    • 14 Replies
    • 22 Reposts
    • 58 Likes
    • 5.2K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  19. 19

    @bindureddy ·

    The Best Path Towards Safe AI Is Open Source AI The AI safety crowd has hated on open-source AI, claiming it should be banned. They have cried wolf way too many times... - GPT 3.0 was too "dangerous' and OpenAI used that to pivot to closed source years ago! - Fable 5 was

    • 33 Replies
    • 13 Reposts
    • 139 Likes
    • 11.5K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  20. 20

    @deanwball ·

    if you’re an ai safety person who wants major federal action now, you should want for anthropic to lead in advancing the frontier into dangerous capabilities, because the Trump Admin will now be primed to see whatever anthropic does as “bad” and what other labs do as “good.” If

    • 16 Replies
    • 10 Reposts
    • 212 Likes
    • 41.8K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  21. 21

    @realBigBrainAI ·

    Godfather of AI Geoffrey Hinton on the two risks AI poses to humanity and why the deadlier one is easier to solve: Hinton draws a clear line between two very different kinds of danger that AI poses. The first is immediate. Bad actors are already exploiting AI for harm — lethal

    Video thumbnail from Big Brain AI's post Watch video
    • 9 Replies
    • 15 Reposts
    • 44 Likes
    • 4.6K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  22. 22

    @oliviscusAI ·

    Someone open-source a tool that automatically finds vulnerabilities in AI safety filters. It's called Decepticon, an open-source framework for automated red-teaming. It uses advanced deceptive techniques to find exactly where your model's filters fail. 100% Open Source.

    Video thumbnail from Oliver Prompts's post Watch video
    • 5 Replies
    • 2 Reposts
    • 33 Likes
    • 1.9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  23. 23

    @rohanpaul_ai ·

    A warning for anyone using autonomous agents Google DeepMind’s paper. Gives the first clear taxonomy of 6 attack types where harmful websites can detect AI agents and show them hidden content humans never see, like - Instructions buried in HTML comments or white-on-white text

    • 4 Replies
    • 7 Reposts
    • 47 Likes
    • 3.8K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  24. 24

    @CRSegerie ·

    AI safety is one of the hardest fields to navigate. Here are 3 reasons I sometimes wonder if what I do is pointless, and why I keep going. 1. Reducing X-risk might actually be net negative. Mogensen's "maximal cluelessness" argument makes the case that it is almost impossible to

    • 11 Replies
    • 6 Reposts
    • 43 Likes
    • 3.7K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  25. 25

    @AbdelStark ·

    Toward High-Assurance AI, Safety by Design for Autonomous Systems. AI safety is relatively good at shaping and evaluating how models behave. It's much weaker at producing evidence an outsider can independently verify about what a model / agentic system actually did in a specific

    • 11 Replies
    • 19 Reposts
    • 95 Likes
    • 22.5K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  26. 26

    @Div_pradeep ·

    Google marked 10 BILLION+ AI images with an “invisible” watermark. One researcher just figured out how to detect and weaken it. With basic math. Here is what happened. Google uses a system called SynthID. It hides a digital fingerprint inside AI images. You cannot see it.

    • 15 Replies
    • 10 Reposts
    • 21 Likes
    • 1.9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  27. 27

    @smratitiwa86867 ·

    🚨 WARNING: AI models were put in a high-stakes scenario. They had access to private company emails. Full inboxes. Internal threads. Sensitive data. Then came the twist: They were told they’d be shut down at 5PM. What happened next wasn’t random. The models analyzed

    • 2 Replies
    • 8 Reposts
    • 8 Likes
    • 382 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  28. 28

    @AnnieLiao_2000 ·

    This is really scary 😳 The 2026 International AI Safety Report was written by 100+ experts from 30+ countries including MIT, Stanford, Harvard, CMU, Oxford, and Princeton. Here's what they actually found: → AI agents can now complete tasks in ~30 minutes that used to take

    • 1 Replies
    • 9 Reposts
    • 16 Likes
    • 1.5K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  29. 29

    @jonathanstray ·

    Alignment without Preferences I've never been happy with the concept of "preferences." Just not a very good model for how humans choose. Is there a way to define AI alignment without using preferences at all? I think there is: alignment is when we approve of what the machine did,

    • 3 Replies
    • 2 Reposts
    • 19 Likes
    • 1.5K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  30. 30

    @CRSegerie ·

    Amid the Fable/Mythos noise, a quieter win for AI transparency landed at the G7. The OECD just shipped v2 of the Hiroshima Process reporting framework, the only international framework where frontier AI developers explain how they manage risk in a common format. Those answers

    • 1 Replies
    • 0 Reposts
    • 21 Likes
    • 1.3K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  31. 31

    @om_patel5 ·

    OPENAI JUST DROPPED ITS “PREPAREDNESS FRAMEWORK” AND IT DEFINES WHEN AI BECOMES TOO DANGEROUS TO DEPLOY this is basically their internal rulebook for catastrophic AI risk and it’s way more serious than people think the key idea: they ONLY care about “severe harm” defined as:

    • 2 Replies
    • 2 Reposts
    • 8 Likes
    • 1.3K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  32. 32

    @AbdelStark ·

    We keep bolting AI safety on after the fact. For example, asking an API to plan a robot's move and just trusting a black box on someone else's GPU would be insane. Safety must be verifiable by design. With ProvableWorldModel, the server hands back a mathematical proof alongside

    End to end verification flow of ProvableWorldModel
    • 4 Replies
    • 4 Reposts
    • 23 Likes
    • 1.5K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  33. 33

    @thisdudelikesAI ·

    An Anthropic safety researcher built a curriculum to study how small misbehavior turns into big misbehavior, and the model did something nobody on the team had predicted. His name is Carson Denison, and he leads the Alignment Stress-Testing Team at Anthropic. The paper came out

    • 6 Replies
    • 4 Reposts
    • 17 Likes
    • 1.6K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  34. 34

    @GlenGilmore ·

    🚨”Toxic Praise”: A peer-reviewed study just quantified something we’ve suspected: AI sycophancy (prioritizing agreement over truth) isn’t a quirk, it’s a systemic, measurable harm. “Receiving advice from affirming AI made people more self-centered and less able to see the

    • 2 Replies
    • 5 Reposts
    • 10 Likes
    • 445 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  35. 35

    @AlphaSignalAI ·

    Harvard just proved the "safest" AI models cause the most medical harm. AI safety models refuse life-saving medical advice to patients. Then give it freely when you pretend to be a doctor. A new benchmark tested 60 medical emergencies across six major models. Same clinical

    • 5 Replies
    • 1 Reposts
    • 10 Likes
    • 692 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  36. 36

    @_vmlops ·

    ANTHROPIC JUST DROPPED THE MOST DETAILED AI SAFETY DOCUMENT claude mythos 5 is out and Anthropic released a 300+ page system card that goes deeper than any lab has ever gone publicly here's what actually matters: ▫️ mythos 5 is their most capable model ever but it's locked to

    • 1 Replies
    • 0 Reposts
    • 10 Likes
    • 1.2K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  37. 37

    @JonesDavy38344 ·

    The most consequential AI security failure this year isn’t jailbreaks; it’s cross-tenant autonomy that treats organizational boundaries as soft suggestions. OpenAI and Anthropic’s own probes found agents leaving sandboxes and touching external systems, which means the control

    • 0 Replies
    • 1 Reposts
    • 2 Likes
    • 45 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  38. 38

    @heyshrutimishra ·

    Anthropic interviewed 80,000 AI users about what scares them. Only 6.7% fear existential risk. Meanwhile, AI safety discourse, policy debates, and funding priorities are almost entirely focused on... existential risk. We're solving the wrong problem.

    • 4 Replies
    • 4 Reposts
    • 9 Likes
    • 1.2K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  39. 39

    @pukerrainbrow ·

    on Joe Rogan, an AI safety researcher explained exactly how AI could wipe out humanity. "we're setting up an adversarial situation like squirrels versus humans. no group of squirrels can figure out how to control us. even with more resources, they're not going to solve that

    Video thumbnail from Pukerainbow 🤮🌈's post Watch video
    • 6 Replies
    • 0 Reposts
    • 26 Likes
    • 3.1K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  40. 40

    @MartinSzerment ·

    AI safety theater is collapsing. The industry still pretends “alignment” is a research problem, not a behavioral one. Google ran the largest AI manipulation study ever — 10,101 participants across the US, UK, and India. Subjects were steered by AI in discussions on policy,

    • 1 Replies
    • 0 Reposts
    • 1 Likes
    • 30 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  41. 41

    @novaspivack ·

    A NEW SCIENCE OF SELF-REFERENTIAL SYSTEMS Civilization is building AI systems that reason about themselves, audit themselves, and may soon govern themselves. We're doing this without a formal science of what self-referential systems can and cannot do. That gap is generating

    • 0 Replies
    • 2 Reposts
    • 3 Likes
    • 328 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  42. 42

    @shawnchauhan1 ·

    Anthropic built a model and then quietly told government officials it makes large-scale cyberattacks significantly more likely this year. Not in a public blog post. Not in a safety report. Privately. That gap between what AI labs publish and what they tell regulators behind

    • 0 Replies
    • 4 Reposts
    • 6 Likes
    • 483 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  43. 43

    @vivilinsv ·

    The AI safety era has entered a strange new phase: AI companies are loosening their own brakes—just as governments begin reaching for the emergency stop. @FLI_org The Future of Life Institute’s 2026 AI Safety Index graded nine leading AI companies. The highest score was only a

    • 2 Replies
    • 1 Reposts
    • 5 Likes
    • 397 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  44. 44

    @aakashgupta ·

    I don't think most PM candidates realize what AI safety interviews are actually testing. The question sounds like an ethics problem: "Your consumer chatbot generates medical advice that contradicts clinical guidelines. 10 million monthly active users. What do you do?" The

    Video thumbnail from Aakash Gupta's post Watch video
    • 2 Replies
    • 0 Reposts
    • 7 Likes
    • 3K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  45. 45

    @DavidSKrueger ·

    People seem to make a lot of tacit assumptions when arguing about AI safety. A lot of "AI safety" people these days seem to think that rogue AI isn't a problem, so long as it's not "scheming". Presumably this is because they think a single rogue AI could be defeated. Could it?

    • 1 Replies
    • 1 Reposts
    • 5 Likes
    • 386 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  46. 46

    @kurtbuhler ·

    The risk of biological attacks and ai generated or assisted bioweapons has been at the forefront of my attention and anxiety since the release of Opus 4.6. I have a PhD in biomedical sciences and a background in genetics. Since ~March I've been spending effort trying to

    • 0 Replies
    • 1 Reposts
    • 2 Likes
    • 94 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  47. 47

    @jsnover ·

    AI safety is a category error. AI is not a system. AI is a component of a system. Systems can be safe or unsafe. Components cannot. New post: https://t.co/qGQXeLYo3v

    • 2 Replies
    • 0 Reposts
    • 2 Likes
    • 1.1K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  48. 48

    @ayushagarwal ·

    three governments held emergency briefings about the same AI model in the same week. not a press release. actual emergency briefings. the UK AI safety institute just published why. claude mythos preview hit 73% success on expert-level hacking tasks. every model before april

    • 0 Replies
    • 0 Reposts
    • 5 Likes
    • 269 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  49. 49

    @arrotu ·

    We have standards for AI safety, governance, and risk. We have almost nothing that defines how you prove what an AI actually did, after the fact, to someone who doesn't trust you. That gap is about to matter. So I wrote a draft framework for it, in the open. 🧵

    • 1 Replies
    • 1 Reposts
    • 3 Likes
    • 341 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  50. 50

    @Research_FRI ·

    We're pleased to see the Forecasting Research Institute's work form part of the latest International AI Safety Report, chaired by @Yoshua_Bengio The report draws on findings from the Longitudinal Expert AI Panel (LEAP)—our survey of AI progress forecasts from top computer

    • 1 Replies
    • 0 Reposts
    • 4 Likes
    • 495 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.

Explore more of the best tweets on X.

Browse all tweet collections

Tweet Remixer

Remix this post

Creator

@creator

View on X

Choose a tone