Best tweets about AI Observability

45 Best Tweets About AI Observability (2026)

Find the best tweets about AI observability, including LLM tracing, evaluations, monitoring, prompt analytics, cost, latency, and production reliability.

Production AI and LLM observability, tracing, evaluation, monitoring, prompt analytics, cost, latency, incidents, tools, and engineering practices.

Creators
38
Updated

Top AI Observability tweets from 38 creators

Ranked 01–45

  1. 01

    @techNmak ·

    Someone documented the engineering principles behind AI agents that actually work in production. It's called 12-Factor Agents. Here's what each factor actually means and why it matters: Factor 1 - Natural Language to Tool Calls The LLM's only job is to decide what to do next,

    • 10Replies
    • 63Reposts
    • 309Likes
    • 14.7KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  2. 02

    @techNmak ·

    Most engineers think AI engineering means fine-tuning models. It doesn't. Chip Huyen's AI Engineering, the most-read book on O'Reilly since release, is a masterclass in what building production AI actually looks like. Here's what matters most. Traditional ML engineers build

    • 10Replies
    • 76Reposts
    • 477Likes
    • 27.4KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  3. 03

    @bibryam ·

    🌟 Building Reliable Agentic AI Systems🌟 https://t.co/5yRJLkIsyl - @thoughtworks What it actually takes to build product-ready agents: → Start with bounded workflows, not open-ended autonomy. Agents need clear task boundaries, allowed tools, and explicit stopping conditions. →

    • 21Replies
    • 52Reposts
    • 271Likes
    • 15.7KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  4. 04

    @_avichawla ·

    DevOps vs. MLOps vs. LLMOps: Many teams are trying to apply DevOps practices to LLM apps. But DevOps, MLOps, and LLMOps solve fundamentally different problems. DevOps is software-centric. You write code, test it, and deploy it. The feedback loop is straightforward: Does the

    Video thumbnail from Avi Chawla's postWatch video
    • 10Replies
    • 52Reposts
    • 236Likes
    • 15.8KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  5. 05

    @JustAnotherPM ·

    𝗔𝗜 𝗣𝗿𝗼𝗱𝘂𝗰𝘁 𝗠𝗮𝗻𝗮𝗴𝗲𝗺𝗲𝗻𝘁 𝗶𝘀 𝗼𝘃𝗲𝗿𝗵𝘆𝗽𝗲𝗱. That is what I told myself when I first started working on AI products. Turns out, I had no idea what an AI PM really did. After spending years watching world-class AI PMs, building AI products at scale, making 100s of bad decisions, I have a

    • 4Replies
    • 21Reposts
    • 138Likes
    • 6KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  6. 06

    @arpit_bhayani ·

    New write-up is live, and this time I covered G-Eval. It helps answer one important question: how do you know whether what an LLM generated is apt, correct, and aligned with your requirements? G-Eval is a pretty simple framework that leverages Chain-of-Thought prompting over

    • 11Replies
    • 17Reposts
    • 246Likes
    • 16.3KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  7. 07

    @emollick ·

    Our lab just released our AI Behavioral Observatory open source. It lets you run statistically valid tests on how AI behavior changes under various types of prompts. We have been using it for our own studies & I think it could help others do similar work.

    • 24Replies
    • 42Reposts
    • 272Likes
    • 20.1KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  8. 08

    @businessbarista ·

    Loved this 22-minute talk on continual learning for AI agents. Must watch for anyone looking to get agents performant and into production. Credit: @FeiziSoheil at @aiDotEngineer • Agent learning can happen at three layers: the model (weights), the harness (prompts, tools,

    Video thumbnail from Alex Lieberman's postWatch video
    • 12Replies
    • 6Reposts
    • 104Likes
    • 13.3KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  9. 09

    @Aurimas_Gr ·

    𝗔𝗜 𝗢𝗯𝘀𝗲𝗿𝘃𝗮𝗯𝗶𝗹𝗶𝘁𝘆 is a must have in your tool belt as an AI Engineer. 𝗧𝗿𝗮𝗰𝗶𝗻𝗴 sits at the core of it, why is it important? Tracing and instrumentation of software have been around for decades now. With AI systems resembling regular software even more, we are now moving the

    • 7Replies
    • 31Reposts
    • 119Likes
    • 5.8KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  10. 10

    @HamelHusain ·

    New Blog Post: Do Automated Evals Work? There has been a rise of tools that look through your traces with AI and identifies issues. We tested these tools with real production data to see how good they are. Where they shine - They often spot issues human miss - Integrate into

    • 13Replies
    • 12Reposts
    • 108Likes
    • 9.1KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  1. 11

    @skull8888888888 ·

    Excited to share that @lmnrai has raised $3M to build open-source observability for long-running AI agents. Laminar is how companies like @browser_use, @OpenHandsDev, and Rye see what their agents are doing, understand why they fail, and spot patterns across millions of runs.

    Video thumbnail from Robert's postWatch video
    • 44Replies
    • 35Reposts
    • 231Likes
    • 43.5KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  2. 12

    @Aurimas_Gr ·

    This is how you measure your AI system as an AI Engineer 👇 For regular software you would track metrics like uptime, error rate, p95 latency. However, they say little about whether the system is fast where users feel it, affordable at scale or correct. Here are the metrics we

    • 8Replies
    • 12Reposts
    • 74Likes
    • 3KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  3. 13

    @nikks_techie ·

    AI tool of the day: LiteLLM What it is LiteLLM is an open-source gateway that provides a single, OpenAI-compatible API for 100+ LLMs and providers, including OpenAI, Anthropic, Gemini, Grok, Ollama, Bedrock, and Azure OpenAI. Benefits Switch between LLM providers with minimal

    • 33Replies
    • 4Reposts
    • 38Likes
    • 716Views
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  4. 14

    @_avichawla ·

    As an AI Engineer shipping agents to production, please learn: - Not every intent needs an agent - Early stopping over indefinite retries - Fallback parsers for structured output - Evals for agent behavior not just output - Delivery infra that's framework-agnostic - Provider

    • 21Replies
    • 9Reposts
    • 66Likes
    • 9.7KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  5. 15

    @akshay_pachaar ·

    A great LLM interview question: (answer shared below) You have 80k Agent-user interactions from production. You need to find the top 100 worth reviewing to improve the agent. You cannot use an LLM to evaluate them since it will be expensive. This is one of the most painful

    Video thumbnail from Akshay 🚀's postWatch video
    • 12Replies
    • 8Reposts
    • 90Likes
    • 8.7KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  6. 16

    @0xlelouch_ ·

    90% of AI Engineering interviews in 2026 come down to these 7 points: 1) Problem framing + success metric Define the target (latency, accuracy@k, cost/request). Say what you’ll measure in prod, not just offline. 2) Data + labeling reality Where does training data come from,

    • 1Replies
    • 7Reposts
    • 24Likes
    • 1.4KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  7. 17

    @pvergadia ·

    Your LLM app isn't broken because of the model. It's broken because you never measured it. AI Evals!! Most teams do the same thing: → Build it → Test it on 5 examples → Demo goes perfectly → Ship it → Pray Then 3 weeks in, a user screenshots your chatbot confidently

    • 3Replies
    • 3Reposts
    • 25Likes
    • 1.6KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  8. 18

    @DeepStarts ·

    Here's a list of Al Engineer Interview questions + concepts you need to know (from Al/ML Engineering Manager perspective) LLM Fundamentals: -What is tokenization, and how does it affect generation? -How do embeddings really work? -What's the role of attention, positional

    • 12Replies
    • 2Reposts
    • 20Likes
    • 227Views
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  9. 19

    @joulee ·

    Designing for trust. When I asked dozens of data leaders how accurate their AI tools were, the answers ranged from 30% to 85%. That’s the Achilles heel of LLMs: they don’t know what they don't know. They speak with all the clarity and confidence of a sales rep in a polished

    Video thumbnail from Julie Zhuo's postWatch video
    • 7Replies
    • 7Reposts
    • 54Likes
    • 5.9KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  10. 20

    @therealdanvega ·

    Spring AI advisors are AOP for your LLM calls. Do something before the request, something after. I built two custom advisors: one that logs which tools are visible to the model, and one that tracks token usage per call with a running total. Full walkthrough here:

    • 1Replies
    • 4Reposts
    • 37Likes
    • 2.3KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  11. 21

    @xelebofficial ·

    Why operating AI agents is becoming the next big challenge? The first wave of Agentic AI was about capability. Can an agent reason? Can it use tools? Can it complete tasks autonomously? The answer is increasingly yes. But a new problem is emerging. Once an agent can act, how

    • 18Replies
    • 4Reposts
    • 45Likes
    • 4.3KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  12. 22

    @JustAnotherPM ·

    Every Product Manager thinks they can "figure out AI" when they need to. They're wrong. I've been too many product reviews where PMs confidently presented AI features they didn't understand. The results were painful to watch. Here's what traditional PM experience actually

    • 4Replies
    • 2Reposts
    • 14Likes
    • 1.3KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  13. 23

    @alex_verem ·

    The AI is 5% of the work. The 95% that breaks: → Observability (Langfuse, Braintrust, Helicone) - you can't debug what you can't see → Evals - regression suites for non-deterministic software. The new CI. → Durable runtime (Temporal, Inngest) - so a 10-minute agent run survives

    • 8Replies
    • 9Reposts
    • 18Likes
    • 3.2KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  14. 24

    @0xlelouch_ ·

    90% of LLMOps interviews in 2026 come down to these 7 points: 1) RAG system design: chunking, embeddings, top-k, rerankers, and how you measure retrieval quality vs latency. 2) Evaluation: offline golden sets + online A/B, judge model pitfalls, and metrics for hallucination,

    • 0Replies
    • 2Reposts
    • 24Likes
    • 1.1KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  15. 25

    @pauliusztin_ ·

    Every AI feature should answer two questions before it gets merged: 1. Did it improve anything? 2. Did it break anything? Most teams only answer the first question. And that's exactly why regressions keep slipping into production. This is where Evaluation-Driven Development

    • 0Replies
    • 3Reposts
    • 13Likes
    • 732Views
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  16. 26

    @ambient_xyz ·

    Enterprise AI contracts promise stability but they rarely deliver it. When a model provider silently updates weights or reroutes inference, your SLAs do not protect you. Your monitoring does not catch it. Your customers feel it first. The repeated problem here is that you cannot

    • 10Replies
    • 3Reposts
    • 32Likes
    • 800Views
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  17. 27

    @socialwithaayan ·

    The biggest risk in AI right now isn't that your agent fails loudly. It's that it fails silently 🤯 There's a free tool called iFixAi that catches exactly that. 45 inspections. Letter grade in under 5 minutes. Any model, any industry. Here's the problem it solves: Your AI agent

    • 17Replies
    • 6Reposts
    • 53Likes
    • 55.6KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  18. 28

    @posthog ·

    Clusters is live in PostHog LLM Analytics – and it might be the fastest way to understand what your AI app is doing in production. We use it to keep an eye on our own AI. Here's a look at recent generations from our setup wizard.

    LLM generation clusters in PostHog LLM analytics
    • 4Replies
    • 4Reposts
    • 38Likes
    • 8KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  19. 29

    @WillyChuang ·

    Everyone on X feed is still posting "look what the model can do." That conversation is over... This week's summary at @aiDotEngineer World's Fair with @alanwuuuuuu. 1. Evals Have "Eaten the Conference" — AI Engineering Has Grown Up -The conversation has shifted from "Look what

    • 3Replies
    • 3Reposts
    • 14Likes
    • 481Views
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  20. 30

    @DivyanshT91162 ·

    The fastest way to understand LLM inference? Stop reading about it. Ship one. This 10-week roadmap takes you from a blank GPU to a production-grade, OpenAI-compatible inference service — in just 30 minutes a day. You’ll build your way through: → vLLM + SGLang →

    • 7Replies
    • 3Reposts
    • 7Likes
    • 1KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  21. 31

    @Jbm_dev ·

    the longer you ship AI in production, the more you need to: - enforce governance - track every model call - block PII before it leaves - stop trusting vibes over evals anything else you'd add?

    • 6Replies
    • 0Reposts
    • 10Likes
    • 243Views
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  22. 32

    @sagar_batchu ·

    Claude Code, Cursor, Codex, and VS Code Copilot all expose dozens of hook events. But if you're standing up AI governance this quarter, you only need to know about four hooks that will be the basis of your AI governance posture 1. UserPromptSubmit. Fires when a developer submits

    • 1Replies
    • 1Reposts
    • 11Likes
    • 250Views
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  23. 33

    @ambient_xyz ·

    ok I will bite: everyone says ai is becoming enterprise infrastructure then should it not meet enterprise infrastructure standards? when AI breaks inside payroll, claims, support, finance or R&D the damage is immediate because enterprises cannot run critical systems on opaque

    ai enterprise claude mythos
    • 5Replies
    • 1Reposts
    • 18Likes
    • 519Views
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  24. 34

    @hugobowne ·

    What exactly are guardrails for AI systems? I asked Katharine Jarmul, ML/AI Privacy expert and author of O'Reilly's Practical Data Privacy, and she broke it down in way in a way that's useful whether you're technical or not: 1. External deterministic: Fast, software-based

    Video thumbnail from Hugo Bowne-Anderson's postWatch video
    • 1Replies
    • 1Reposts
    • 7Likes
    • 520Views
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  25. 35

    @ashugarg ·

    The graveyard of enterprise AI pilots is full of products that couldn't clear security and governance. At our recent CEO and CIO dinners, we heard this repeatedly from @djpersia (@databricks), Rajat Taneja (@Visa), and CIOs from @Zuora, @asana, and @BlackLine. In many orgs,

    • 2Replies
    • 0Reposts
    • 11Likes
    • 1.4KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  26. 36

    @illyism ·

    👀 how I debugged a tiny but painful AI SEO Tracker extraction bug today: 1. found a weird real example the app counted source/company names as “brand mentions” even when the prompt asked for one person (@nic_amadio) 2. turned the bug into an LLM eval / benchmark made a fixture

    • 2Replies
    • 2Reposts
    • 6Likes
    • 1.4KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  27. 37

    @Padday ·

    Let's talk about Observability and Agents. Yesterday we launched Monitors for @Fin_ai . Monitors is a very powerful tool for measuring and improving the quality of your customer experience, fully AI powered, highly configurable, with 100% coverage of every single customer

    • 1Replies
    • 2Reposts
    • 19Likes
    • 1.4KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  28. 38

    @yizucodes ·

    Voice AI in production is WAY harder than demos suggest. I just left LiveKit's panel with CTOs from Portola, Infinitus, Yelp & Bluejay breaking down what "reliability" actually means at scale. The gap between prototype and production is wild 🧵 1. Memory consistency > latency

    • 1Replies
    • 0Reposts
    • 2Likes
    • 262Views
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  29. 39

    @mchulet ·

    As an AI Engineer. Please learn >Harness engineering, not just prompt engineering >Context engineering, not just long prompts >Prompt caching vs. semantic caching tradeoffs >KV cache management, eviction, reuse, and memory pressure at scale >Prefill vs. decode latency and

    • 2Replies
    • 1Reposts
    • 3Likes
    • 177Views
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  30. 40

    @MakadiaHarsh ·

    building ai features for startups has taught me one thing calling an LLM API is the easy part the hard part is: - chunking and retrieving the right data - keeping costs from exploding - handling failures gracefully - knowing when NOT to fine-tune - evals that actually measure

    • 8Replies
    • 2Reposts
    • 12Likes
    • 3.1KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  31. 41

    @n8n_io ·

    Your AI workflow passed every test. Two weeks later, quality drops. No errors. Just silent drift. The fix isn’t more pre-deployment testing. It’s continuous evaluation. New in the Production AI Playbook by Elvis Saravia (@omarsar0) 👉 https://t.co/vBb5l1bgBu

    • 5Replies
    • 2Reposts
    • 35Likes
    • 21.5KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  32. 42

    @petesoder ·

    A lot of AI observability tools are fantastic as long as the security team never logs in. With @honeyhiveai v2 (announced today!) you can keep full raw traces inside your own environment and still look your CISO in the eye. @mohak__sharma, @ds3638 and team have rebuilt HoneyHive

    • 0Replies
    • 2Reposts
    • 4Likes
    • 366Views
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  33. 43

    @hugobowne ·

    AI evals aren't just important for building your feedback loop into the product development lifecycle, but they're essential to make sure you're compliant when building in regulated industries. I had a great chat with @stellawliu (Head of Applied Science, ASU) and Eddie

    Video thumbnail from Hugo Bowne-Anderson's postWatch video
    • 1Replies
    • 0Reposts
    • 1Likes
    • 374Views
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  34. 44

    @krishnan ·

    The unglamorous part of AI is becoming the moat. Everyone is covering model launches. Netflix's new engineering writeup points to the harder Day 2 question: can you run LLMs like production infrastructure? Netflix says (https://t.co/tuYfXJNdnn) it runs the full LLM serving stack

    • 0Replies
    • 0Reposts
    • 3Likes
    • 98Views
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  35. 45

    @arrotu ·

    Most AI teams already have logs and traces. Far fewer can prove what actually ran when a workflow is challenged later. That is the gap this new article explores. It breaks down what OpenTelemetry solves well, where logs and traces stop short, and why a separate layer of

    • 0Replies
    • 1Reposts
    • 1Likes
    • 199Views
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.

Write posts like these, in your own voice

Xholic studies what works in your niche, drafts posts in your voice and schedules them for the hours your audience is online.

$0 today · Cancel anytime

Browse all tweet collections

More on this topic

Tweet Remixer

Remix this post

Creator

@creator

View on 𝕏

Choose a tone

Best Tweets by email

Get the best tweets about AI Observability every two weeks

The list refreshes about every two weeks. Confirm your email to start, and unsubscribe anytime.

We only use your email for this digest. See our Privacy Policy.

Write posts like these in your voice

Start 3-day trial for $0

$0 today · Cancel anytime