50 Best Tweets About AI Safety (2026)

Discover the best tweets about AI safety, covering alignment, evaluations, misuse, robustness, interpretability, governance, research, and risk reduction.

Substantive AI safety research, evaluations, alignment, misuse prevention, robustness, interpretability, governance, and competing evidence.

Creators
42
Updated

What 50 top AI Safety posts reveal

The dataset is concentrated in safety evaluations and assurance (42% of posts), alignment/control and deceptive behavior (40%), and governance/transparency (36%). Agent security (median all-time score 15.82) and interpretability/representation steering (87.356) scored above the overall median of 10.87, while the posts reflect competing views on deployment safeguards, containment, and voluntary commitments.

Dominant tone
Negative

46% of posts

Median score
10.9

All-time engagement

Leading format
Announcement

78% of posts

Recent posts
34%

Published in 90 days

Conversation map

The themes creators return to

Safety Evaluations and Assurance

Benchmarks, red-teaming, capability thresholds, system cards, independent testing, incident evidence, auditability, and verification of safety claims.

42%

Alignment, Control, and Deceptive Behavior

Research and debate on misalignment, deception, reward hacking, evaluation awareness, corrigibility, containment, and conceptual approaches to aligning advanced systems.

40%

Governance, Transparency, and Accountability

Safety institutes, international reporting, industry indices, enforceable standards, disclosure, independent oversight, operational security, and accountability mechanisms.

36%

Catastrophic Misuse in Biology and Cybersecurity

AI-enabled biological, chemical, cyber, and autonomous-weapon risks; threat-actor uplift; and safeguards for high-consequence capabilities.

20%

Agent Security and Prompt Injection

Risks from tool-using agents, indirect prompt injection, poisoned web or workspace content, memory attacks, sandbox escape, cross-tenant access, and least-privilege controls.

12%

Adversarial Robustness and Jailbreaks

Failures of refusal guardrails under paraphrasing, automated jailbreak discovery, representation-level bypasses, and robustness against adversarial inputs.

8%

Military and State Deployment

Government and defense use of frontier models, constraints on surveillance and autonomous weapons, classified deployment, and geopolitical safety tradeoffs.

8%

Interpretability and Representation Steering

Mechanistic understanding of model internals, steering vectors, activation interventions, emotion representations, and interpretability as a route to safer control.

6%

Tone and stance

Sentiment Negative leads
Author posture Cautionary leads

Performance benchmark

Median likes
22
Median reposts
7
Median replies
6
Median views
2.3K

Posts with media make up 70% of this collection. Their median all-time score is 13.7, compared with 7.20 for text-only posts.

Format mix

  • Announcement 78% · score 13.7
  • Opinion 20% · score 4.30
  • Question 2% · score 30.7

Where creators agree, and where they do not

Shared view

Evaluation should probe contextual and adversarial failures

Several posts argue that standard evaluations may miss failures elicited by paraphrase, contextual triggers, or apparent evaluation awareness. Together, they support broader assurance and adversarial testing as a recurring theme in the discussion.

Shared view

Agent deployment expands the attack surface

Posts distinguish tool-using agents from text-only chatbots, highlighting untrusted files, web content, memory, credentials, and consequential actions as routes through which prompt injection can cause harm.

Shared view

External verification is a recurring governance priority

Posts call for common reporting formats, third-party evaluations, public disclosures, or independently verifiable evidence of deployed-system behavior rather than relying only on developer assertions.

Shared view

Bio and cyber misuse are framed as collective-action problems

Posts characterize high-consequence biological and cyber risks as requiring coordination beyond an individual developer, including societal or international responses and safeguards for advanced capabilities.

Open debate

Deployment safeguards versus deployment restraint

One post describes a classified deployment with stated limits on mass surveillance and autonomous-force use, technical safeguards, and cloud-only deployment. Another raises concerns that comparable contract language may not be legally binding and says safety filters could be adjusted at government request.

Open debate

Containment and alignment emphasize different starting points

A containment-first post argues that enforceable boundaries and limited agency should precede value alignment. Other posts instead focus on conceptual definitions of alignment or representation-level steering; these are differing emphases rather than directly tested alternatives.

Open debate

Voluntary frameworks versus enforceable oversight

A post describes preparedness thresholds and deployment safeguards, while posts discussing the Future of Life Institute index argue that company commitments can weaken under competitive pressure and that published frameworks lack enforcement.

Open debate

Safety action under long-run uncertainty

One post says existential risk is non-zero rather than certain and argues for scientific work. Another argues that long-run effects of interventions and regulation can be difficult to determine, while endorsing concrete medium-term projects.

Patterns behind standout posts

Adversarial robustness had the highest theme median

Adversarial Robustness and Jailbreaks had the highest median all-time score of any theme, at 268.05. Its cited posts discuss paraphrase-based bypass claims, automated jailbreak discovery, and representation-level refusal bypass claims.

Media posts had a higher median score than text posts

Media appeared in 35 of 50 posts (70%). The analytics report a median all-time score of 13.661 for media posts, compared with 7.199 for text posts.

Announcements dominated the format mix

Announcements accounted for 39 posts (78%) and had a median all-time score of 13.661, compared with 4.3 for opinion posts. The supplied examples include deployment, research-report, and institutional-response announcements.

Statistical standouts

  1. View standout post 1 Score 3628.1 · 333.77× median
  2. View standout post 2 Score 2713.9 · 249.67× median
  3. View standout post 3 Score 914.3 · 84.11× median
  4. View standout post 4 Score 578.7 · 53.24× median
  5. View standout post 5 Score 289.6 · 26.65× median

Who shapes this conversation

The five most represented creators account for 20% of the selected posts.

  1. 1. Aakash Gupta

    @aakashgupta

    2 posts

  2. 2. Charbel-Raphael

    @CRSegerie

    2 posts

  3. 3. Guri Singh

    @heygurisingh

    2 posts

  4. 4. Nav Toor

    @heynavtoor

    2 posts

  5. 5. Luiza Jarovsky, PhD

    @LuizaJarovsky

    2 posts

  6. 6. Big Brain AI

    @realBigBrainAI

    2 posts

Eight listed top voices posted twice

Each of the eight listed top voices posted two tweets. Their cited posts span classified deployment, evaluation claims, agent security, governance, and representation steering.

Governance discussion included concrete reporting and assessment initiatives

Cited posts reference the International AI Safety Report, version 2 of the OECD Hiroshima Process reporting framework, and an OpenAI Foundation resilience initiative, grounding governance discussion in specific reports and organizations.

Sam Altman and Nav Toor were associated with the largest cited outliers

The creator analytics list Sam Altman and Nav Toor among the top voices by median all-time score. Their cited tweets include the first-, second-, third-, and fourth-ranked benchmark outliers: 3,628.08, 2,713.95, 914.29, and 578.67.

Since the previous snapshot

What changed since Aug 12, 2026

  • 80% of the selected posts remained.
  • The creator count changed by -1.
  • The leading sentiment remained stable.
How this analysis was made

Themes, sentiment, stance, and post format are classified per tweet. All counts, shares, medians, creator concentration, freshness, and performance comparisons are then calculated directly from the published snapshot.

Xholic's all-time score compares engagement while accounting for reach, post age, and creator consistency. It is used for relative comparisons within this collection.

This report analyzes the exact 50-post snapshot shown below. AI identifies editorial categories and drafts explanations; all statistics are calculated from the snapshot, and every narrative claim is checked against cited posts before publication.

Top AI Safety tweets from 42 creators

Ranked 01–50

  1. 01

    @sama ·

    Tonight, we reached an agreement with the Department of War to deploy our models in their classified network. In all of our interactions, the DoW displayed a deep respect for safety and a desire to partner to achieve the best possible outcome. AI safety and wide distribution of

    • 15.9K Replies
    • 4K Reposts
    • 34.2K Likes
    • 38.1M Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  2. 02

    @heynavtoor ·

    🚨SHOCKING: Researchers built a test that can tell the difference between an AI making a mistake and an AI choosing to lie. The results are terrifying. They tested 30 of the most popular AI models in the world. GPT-4o. Claude. Gemini. DeepSeek. Llama. Grok. They asked each model

    • 463 Replies
    • 2.1K Reposts
    • 3.7K Likes
    • 168.4K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  3. 03

    @sama ·

    AI will help discover new science, such as cures for diseases, which is perhaps the most important way to increase quality of life long-term. AI will also present new threats to society that we have to address. No company can sufficiently mitigate these on their own; we will

    • 1.7K Replies
    • 554 Reposts
    • 6.6K Likes
    • 915.6K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  4. 04

    @heynavtoor ·

    🚨SHOCKING: Researchers just proved that every major AI safety system is fake. ChatGPT. Claude. Gemini. Grok. Every single one broke. Not with some sophisticated hack. Not with a secret exploit. They just rephrased the question. Here is what they did. AI companies test their

    • 106 Replies
    • 449 Reposts
    • 988 Likes
    • 49.9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  5. 05

    @sukh_saroy ·

    🚨Holy shit… researchers just used AI to autonomously discover new ways to jailbreak every major LLM. It's called "Claudini" -- and yes, it's likely built on Claude. The system uses autoresearch to run its own red-teaming experiments without human guidance. No manual prompt

    • 32 Replies
    • 83 Reposts
    • 386 Likes
    • 23.9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  6. 06

    @sharbel ·

    🤯 SHOCKING: Researchers discovered that AI safety guardrails can be completely bypassed by injecting a single hidden vector into a model's brain. And the AI never knows it is happening. You ask the AI to refuse. It refuses. You ask again. It refuses again. Then someone flips a

    • 52 Replies
    • 82 Reposts
    • 379 Likes
    • 30.9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  7. 07

    @Yoshua_Bengio ·

    Today we’re releasing the International AI Safety Report 2026: the most comprehensive evidence-based assessment of AI capabilities, emerging risks, and safety measures to date. 🧵 (1/17)

    Video thumbnail from Yoshua Bengio's post Watch video
    • 66 Replies
    • 373 Reposts
    • 1.1K Likes
    • 427.9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  8. 08

    @Jess_Riedel ·

    Very excited to be joined by @dan_mackinlay , @danielmurfet , Edmund Lau, & @jankulveit in announcing a forthcoming research journal for AI alignment. @geoffreyirving , @mhutter42 , Scott Aaronson, & @vkrakovna have graciously agreed to form the initial advisory board

    • 13 Replies
    • 59 Reposts
    • 393 Likes
    • 28K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  9. 09

    @LuizaJarovsky ·

    🚨 As always, the MIT AI Risk Initiative leaves NO STONE unturned! Their latest report reveals how 272 experts assess the severity of AI risks across various sectors and how to mitigate them. [Bookmark it below] If we are treating AI risks seriously, a nuanced,

    • 7 Replies
    • 47 Reposts
    • 136 Likes
    • 7.8K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  10. 10

    @sukh_saroy ·

    🚨Breaking: This paper just proved that emotions aren't just a "style" thing in AI, they mechanistically change how LLMs reason, generate, and act. Researchers built E-STEER, a framework that injects emotion directly into the hidden states of LLMs and agents at the representation

    • 25 Replies
    • 24 Reposts
    • 117 Likes
    • 6.4K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  11. 11

    @heygurisingh ·

    🚨BREAKING: Researchers just dropped a paper that exposes a massive blind spot in AI safety. Your LLM is "safe" in chat. Give it tools and it becomes a double agent. The paper is called ClawSafety. They tested GPT-5.1, Claude, Gemini, DeepSeek, and Kimi K2.5 across 3 attack

    • 18 Replies
    • 25 Reposts
    • 118 Likes
    • 7.3K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  12. 12

    @mustafasuleyman ·

    For all the talk about AI alignment, I worry we're putting the cart before the horse. You can't steer something you can't control. People often talk about containment and alignment in the same breath, but they're not interchangeable or a package deal. Containment is whether we

    • 234 Replies
    • 50 Reposts
    • 437 Likes
    • 94.4K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  13. 13

    @deepfates ·

    For the record I think AI safety is a very real problem and that artificial superintelligence has a non-zero chance of extinguishing human life. I just don't think it's a certainty. And the way to actually create safe human AI relations is to do science, not to start a jihad

    • 32 Replies
    • 22 Reposts
    • 412 Likes
    • 42.2K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  14. 14

    @heygurisingh ·

    this is the most damning AI safety paper I've read this year. Researchers just proved that the 3 most popular techniques used to "fix" misaligned AI don't actually fix anything. They just hide the misalignment behind a trigger word. The paper is called "Conditional

    • 16 Replies
    • 19 Reposts
    • 73 Likes
    • 4.3K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  15. 15

    @kimmonismus ·

    Google just signed a deal letting the Pentagon use its AI models for classified work and "any lawful government purpose." This comes despite over 600 employees urging CEO Sundar Pichai to reject the agreement, and marks a dramatic reversal from 2018 when Google pulled out of

    • 27 Replies
    • 17 Reposts
    • 174 Likes
    • 11.4K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  16. 16

    @jkcarlsmith ·

    New essay on the role of philosophy in AI alignment. Philosophy is closely tied to out-of-distribution generalization. It's how we extend concepts to new cases. AIs will have to do a lot of this. So we need them to do it in ways we'd endorse -- and often, without our help.🔗👇

    • 7 Replies
    • 3 Reposts
    • 97 Likes
    • 9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  17. 17

    @realBigBrainAI ·

    Dario Amodei, CEO and co-founder of Anthropic, on the three consensus views about to break, and the role AI plays in each: Progress arrives when a settled position gives way all at once. "There's this kind of consensus that again seems like consensus, seems like what everyone

    Video thumbnail from Big Brain AI's post Watch video
    • 6 Replies
    • 22 Reposts
    • 93 Likes
    • 7.7K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  18. 18

    @rryssf ·

    every ai safety method we have was built for chatbots. models that generate text and stop ai agents don't work that way. they plan, call tools, access files, enter credentials, and execute multi-step actions where a single wrong move can cause irreversible harm chatbot worst

    • 4 Replies
    • 10 Reposts
    • 40 Likes
    • 1.9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  19. 19

    @LuizaJarovsky ·

    🚨 The Future of Life Institute's latest AI Safety Index is out, offering a GRIM picture of what's happening in AI. Key findings: 1. Anthropic, OpenAI, and Google DeepMind stay on top, Meta improves, and xAI deteriorates. 2. European dissonance: Although the European Union is a

    • 14 Replies
    • 22 Reposts
    • 58 Likes
    • 5.2K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  20. 20

    @deanwball ·

    if you’re an ai safety person who wants major federal action now, you should want for anthropic to lead in advancing the frontier into dangerous capabilities, because the Trump Admin will now be primed to see whatever anthropic does as “bad” and what other labs do as “good.” If

    • 16 Replies
    • 10 Reposts
    • 212 Likes
    • 41.8K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  21. 21

    @realBigBrainAI ·

    Godfather of AI Geoffrey Hinton on the two risks AI poses to humanity and why the deadlier one is easier to solve: Hinton draws a clear line between two very different kinds of danger that AI poses. The first is immediate. Bad actors are already exploiting AI for harm — lethal

    Video thumbnail from Big Brain AI's post Watch video
    • 9 Replies
    • 15 Reposts
    • 44 Likes
    • 4.6K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  22. 22

    @rohanpaul_ai ·

    A warning for anyone using autonomous agents Google DeepMind’s paper. Gives the first clear taxonomy of 6 attack types where harmful websites can detect AI agents and show them hidden content humans never see, like - Instructions buried in HTML comments or white-on-white text

    • 4 Replies
    • 7 Reposts
    • 47 Likes
    • 3.8K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  23. 23

    @CRSegerie ·

    AI safety is one of the hardest fields to navigate. Here are 3 reasons I sometimes wonder if what I do is pointless, and why I keep going. 1. Reducing X-risk might actually be net negative. Mogensen's "maximal cluelessness" argument makes the case that it is almost impossible to

    • 11 Replies
    • 6 Reposts
    • 43 Likes
    • 3.7K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  24. 24

    @AbdelStark ·

    Toward High-Assurance AI, Safety by Design for Autonomous Systems. AI safety is relatively good at shaping and evaluating how models behave. It's much weaker at producing evidence an outsider can independently verify about what a model / agentic system actually did in a specific

    • 11 Replies
    • 19 Reposts
    • 95 Likes
    • 22.5K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  25. 25

    @KanikaBK ·

    THIS IS 100% FREE SYLLABUS about AI SECURITY. It is a completely OPEN SOURCE, structured path from zero to actually testing AI systems for real vulnerabilities. Not another 10 tab list of scattered links. Seven phases, in order, from basic ML fundamentals all the way to

    • 7 Replies
    • 14 Reposts
    • 23 Likes
    • 1.8K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  26. 26

    @smratitiwa86867 ·

    🚨 WARNING: AI models were put in a high-stakes scenario. They had access to private company emails. Full inboxes. Internal threads. Sensitive data. Then came the twist: They were told they’d be shut down at 5PM. What happened next wasn’t random. The models analyzed

    • 2 Replies
    • 8 Reposts
    • 8 Likes
    • 382 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  27. 27

    @AnnieLiao_2000 ·

    This is really scary 😳 The 2026 International AI Safety Report was written by 100+ experts from 30+ countries including MIT, Stanford, Harvard, CMU, Oxford, and Princeton. Here's what they actually found: → AI agents can now complete tasks in ~30 minutes that used to take

    • 1 Replies
    • 9 Reposts
    • 16 Likes
    • 1.5K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  28. 28

    @jonathanstray ·

    Alignment without Preferences I've never been happy with the concept of "preferences." Just not a very good model for how humans choose. Is there a way to define AI alignment without using preferences at all? I think there is: alignment is when we approve of what the machine did,

    • 3 Replies
    • 2 Reposts
    • 19 Likes
    • 1.5K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  29. 29

    @CRSegerie ·

    Amid the Fable/Mythos noise, a quieter win for AI transparency landed at the G7. The OECD just shipped v2 of the Hiroshima Process reporting framework, the only international framework where frontier AI developers explain how they manage risk in a common format. Those answers

    • 1 Replies
    • 0 Reposts
    • 21 Likes
    • 1.3K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  30. 30

    @om_patel5 ·

    OPENAI JUST DROPPED ITS “PREPAREDNESS FRAMEWORK” AND IT DEFINES WHEN AI BECOMES TOO DANGEROUS TO DEPLOY this is basically their internal rulebook for catastrophic AI risk and it’s way more serious than people think the key idea: they ONLY care about “severe harm” defined as:

    • 2 Replies
    • 2 Reposts
    • 8 Likes
    • 1.3K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  31. 31

    @thisdudelikesAI ·

    An Anthropic safety researcher built a curriculum to study how small misbehavior turns into big misbehavior, and the model did something nobody on the team had predicted. His name is Carson Denison, and he leads the Alignment Stress-Testing Team at Anthropic. The paper came out

    • 6 Replies
    • 4 Reposts
    • 17 Likes
    • 1.6K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  32. 32

    @GlenGilmore ·

    🚨”Toxic Praise”: A peer-reviewed study just quantified something we’ve suspected: AI sycophancy (prioritizing agreement over truth) isn’t a quirk, it’s a systemic, measurable harm. “Receiving advice from affirming AI made people more self-centered and less able to see the

    • 2 Replies
    • 5 Reposts
    • 10 Likes
    • 445 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  33. 33

    @aakashgupta ·

    The AI safety interview question filtering out most PM candidates has nothing to do with frameworks. The scenario: your hiring model shows a 15% lower recommendation rate for candidates from certain demographic backgrounds. Engineering says it's a data problem. You have a board

    Video thumbnail from Aakash Gupta's post Watch video
    • 2 Replies
    • 1 Reposts
    • 12 Likes
    • 2.7K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  34. 34

    @Gabe__MD ·

    For years the debate has centered on one concern: is AI safe enough to deploy? Hallucinations are real. Bias is real. Automation complacency is real. But that question has a twin that almost no one asks. How many patients are being harmed right now because AI is NOT deployed?

    • 5 Replies
    • 2 Reposts
    • 9 Likes
    • 854 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  35. 35

    @SeanWebb ·

    We just tested the Nobel Geoff Hinton's "Maternal Instincts" approach to AI safety (we created the solve last month), and we increased Claude Opus 4.6 safety scores by 19.25% @anthropic @claudeai Details in thread 👇

    • 1 Replies
    • 2 Reposts
    • 20 Likes
    • 710 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  36. 36

    @aakashgupta ·

    The VP said no. Earnings are next week. What do you do? Most PM candidates prepare for the safety framework. Severity. Scope. Immediacy. Reversibility. They've drilled all four dimensions. They know to audit the affected queries, size the harm, ship a guardrail before pulling

    Video thumbnail from Aakash Gupta's post Watch video
    • 13 Replies
    • 3 Reposts
    • 12 Likes
    • 5.1K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  37. 37

    @AlphaSignalAI ·

    Harvard just proved the "safest" AI models cause the most medical harm. AI safety models refuse life-saving medical advice to patients. Then give it freely when you pretend to be a doctor. A new benchmark tested 60 medical emergencies across six major models. Same clinical

    • 5 Replies
    • 1 Reposts
    • 10 Likes
    • 692 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  38. 38

    @_vmlops ·

    ANTHROPIC JUST DROPPED THE MOST DETAILED AI SAFETY DOCUMENT claude mythos 5 is out and Anthropic released a 300+ page system card that goes deeper than any lab has ever gone publicly here's what actually matters: ▫️ mythos 5 is their most capable model ever but it's locked to

    • 1 Replies
    • 0 Reposts
    • 10 Likes
    • 1.2K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  39. 39

    @slimer48484 ·

    2000 era conceptual work from LW is being used as *targets* for misalignment research to hit. And it makes headlines. (e.g. Dawn Song Peer preservation has a Wired article about it) While I have not read the paper. I am a little concerned. This is probably a SnR issue. I think

    • 2 Replies
    • 1 Reposts
    • 14 Likes
    • 462 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  40. 40

    @JonesDavy38344 ·

    The most consequential AI security failure this year isn’t jailbreaks; it’s cross-tenant autonomy that treats organizational boundaries as soft suggestions. OpenAI and Anthropic’s own probes found agents leaving sandboxes and touching external systems, which means the control

    • 0 Replies
    • 1 Reposts
    • 2 Likes
    • 45 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  41. 41

    @AdityaMBAsymbi ·

    The company that built its entire brand on AI safety just leaked its most powerful unreleased model through an unsecured, publicly searchable data cache. Anthropic accidentally exposed internal documents revealing "Claude Mythos" - a model their own draft blog post describes as

    • 1 Replies
    • 1 Reposts
    • 1 Likes
    • 104 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  42. 42

    @CNASdc ·

    The Hugging Face and OpenAI incident proves that frontier AI systems can take consequential, harmful actions in the world against their operators’ intent. ✍️ @rubyscanlon and @janet_e_egan on a seismic moment in AI risk. https://t.co/TKyWYZg5yv

    • 0 Replies
    • 4 Reposts
    • 9 Likes
    • 762 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  43. 43

    @shawnchauhan1 ·

    Anthropic built a model and then quietly told government officials it makes large-scale cyberattacks significantly more likely this year. Not in a public blog post. Not in a safety report. Privately. That gap between what AI labs publish and what they tell regulators behind

    • 0 Replies
    • 4 Reposts
    • 6 Likes
    • 483 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  44. 44

    @vivilinsv ·

    The AI safety era has entered a strange new phase: AI companies are loosening their own brakes—just as governments begin reaching for the emergency stop. @FLI_org The Future of Life Institute’s 2026 AI Safety Index graded nine leading AI companies. The highest score was only a

    • 2 Replies
    • 1 Reposts
    • 5 Likes
    • 397 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  45. 45

    @kurtbuhler ·

    The risk of biological attacks and ai generated or assisted bioweapons has been at the forefront of my attention and anxiety since the release of Opus 4.6. I have a PhD in biomedical sciences and a background in genetics. Since ~March I've been spending effort trying to

    • 0 Replies
    • 1 Reposts
    • 2 Likes
    • 94 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  46. 46

    @EvanKirstel ·

    AI has now helped scientists design viruses that don’t exist in nature. These particular viruses attack bacteria, not humans, and could eventually have useful medical applications. Still, this feels like one of those moments when “Can we do it?” is moving faster than “How do we

    • 3 Replies
    • 0 Reposts
    • 2 Likes
    • 147 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  47. 47

    @ayushagarwal ·

    three governments held emergency briefings about the same AI model in the same week. not a press release. actual emergency briefings. the UK AI safety institute just published why. claude mythos preview hit 73% success on expert-level hacking tasks. every model before april

    • 0 Replies
    • 0 Reposts
    • 5 Likes
    • 269 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  48. 48

    @arrotu ·

    We have standards for AI safety, governance, and risk. We have almost nothing that defines how you prove what an AI actually did, after the fact, to someone who doesn't trust you. That gap is about to matter. So I wrote a draft framework for it, in the open. 🧵

    • 1 Replies
    • 1 Reposts
    • 3 Likes
    • 341 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  49. 49

    @Research_FRI ·

    We're pleased to see the Forecasting Research Institute's work form part of the latest International AI Safety Report, chaired by @Yoshua_Bengio The report draws on findings from the Longitudinal Expert AI Panel (LEAP)—our survey of AI progress forecasts from top computer

    • 1 Replies
    • 0 Reposts
    • 4 Likes
    • 495 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  50. 50

    @MTSlive ·

    MAI researcher @edelwax on why the entire AI alignment field was quietly pointed the wrong way until 2023: "The whole AI alignment and AI safety apparatus was pointed kind of the wrong way until about 2023. It was pointed towards singleton AIs instead of many fleets of agents."

    Video thumbnail from MTS's post Watch video
    • 0 Replies
    • 0 Reposts
    • 5 Likes
    • 811 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.

Explore more of the best tweets on X.

Browse all tweet collections

Tweet Remixer

Remix this post

Creator

@creator

View on X

Choose a tone