Agent architectures and orchestration
Single-agent, multi-agent, hierarchical, graph-based, workflow, and dynamic orchestration patterns; specialization, coordination, and task decomposition.
52%
Best tweets about AI Agents
Browse the best tweets about AI agents, agentic workflows, autonomous systems, tool use, memory, and production lessons. Updated weekly.
Builders sharing concrete agent architectures, evaluations, failures, deployment lessons, and useful demonstrations.
Original Xholic analysis
The sampled discussion emphasizes agent systems rather than models alone: posts focus on specialized roles, memory, tools, orchestration, evaluation, and controls such as guardrails, tracing, and bounded autonomy. The evidence also contains a live design debate over autonomous and multi-agent approaches versus more deterministic or single-agent workflows. [2036724091143463120, 2035039280842916325, 2038143522000589241, 2034114318854459704]
42% of posts
All-time engagement
44% of posts
Published in 90 days
Conversation map
Single-agent, multi-agent, hierarchical, graph-based, workflow, and dynamic orchestration patterns; specialization, coordination, and task decomposition.
52%
Silent failures, compounding errors, deterministic execution, retries, fallbacks, circuit breakers, adaptation, and human supervision in deployed agents.
46%
Agent evals, execution tracing, benchmarks, reward loops, RL, prompt optimization, automated testing, and iterative self-improvement.
32%
Permissions, identity, audit trails, policy enforcement, sandboxing, prompt injection, zero trust, compliance, and bounded autonomy.
32%
Working, semantic, procedural, and episodic memory; context control, state, retrieval, persistent workspaces, and model scaffolding.
28%
Clear agent mandates, role-specific prompts and models, tool contracts, definitions of done, specialized skills, and digital-workforce design.
18%
Browser, API, computer, communications, payment, search, MCP, SaaS integration, and machine-native infrastructure enabling agents to act.
18%
Tracing, logging, production deployment, operational workflows, observability, agent management interfaces, and rollout practices.
14%
Tone and stance
Performance benchmark
Posts with media make up 82% of this collection. Their median all-time score is 14.7, compared with 9.89 for text-only posts.
Format mix
Consensus and debate
Shared view
Several posts describe agent behavior as depending on the surrounding system: memory, tools, context, routing, validation, execution scaffolding, and permissions—not model choice alone.
Shared view
Posts recommend deterministic handling of task execution, state, and errors; strict tool contracts; retries, fallbacks, circuit breakers; permissions; and traceability.
Shared view
The cited posts describe tracing plans and tool use, scoring outcomes, and using those results to refine prompts, policies, or agent harnesses.
Shared view
A recurring pattern is to assign agents distinct jobs, contexts, tools, deliverables, and definitions of done rather than use one general-purpose agent for all functions.
Open debate
Some posts promote autonomous experimentation and agent-improvement loops, while others argue that multistep systems require deterministic guardrails and human steering because errors can be silent or compound.
Open debate
Specialized teams and hierarchical coordination are presented as useful patterns. Another cited post argues that task structure should determine the choice: parallelizable work may benefit from multiple agents, while sequential reasoning may favor one agent.
Open debate
One post distinguishes fixed orchestration from agents that plan and adapt. Other posts argue that predictable workflows remain useful and caution that agent branding can overstate autonomy.
What performs
The infrastructure list was the highest outlier in the supplied benchmark data, with an all-time score of 2049.22, or 150.68 times the dataset median.
The autonomous research-lab post, Agent Lightning tutorial, and architecture reading list were supplied outliers, with all-time scores of 1978.25, 1269.29, and 971, respectively.
Deterministic analytics report media on 41 posts (82%). The media median all-time score was 14.668, compared with 9.891 for text posts.
Statistical standouts
Creator landscape
The five most represented creators account for 20% of the selected posts.
1. Alvin Foo
@alvinfoo
2 posts
2. Bilgin Ibryam
@bibryam
2 posts
3. Shalini Goyal
@goyalshaliniuk
2 posts
4. Pascal Bornet
@pascal_bornet
2 posts
5. Rohan Paul
@rohanpaul_ai
2 posts
6. Ujjwal Chadha
@ujjwalscript
2 posts
Avi Chawla appears once among top voices and has the highest listed top-voice median all-time score, 1269.293. The cited post describes an RL-and-tracing workflow for improving agents.
Bilgin Ibryam appears twice among top voices. The cited posts discuss deterministic execution for tasks, state, and errors, and a runtime-governance toolkit covering policy enforcement, identity, and sandboxing.
The two cited posts describe a production-oriented agent systems stack and warn about compounding multistep failure, linking implementation detail with reliability concerns.
Since the previous snapshot
Themes, sentiment, stance, and post format are classified per tweet. All counts, shares, medians, creator concentration, freshness, and performance comparisons are then calculated directly from the published snapshot.
Xholic's all-time score compares engagement while accounting for reach, post age, and creator consistency. It is used for relative comparisons within this collection.
This report analyzes the exact 50-post snapshot shown below. AI identifies editorial categories and drafts explanations; all statistics are calculated from the snapshot, and every narrative claim is checked against cited posts before publication.
Best AI Agents tweets
Ranked 01–50
@shivsakhuja ·
Lots of companies are now building primitives for an economy where AI agents are the primary users instead of humans. They're betting on an economy of AI coworkers. 1. AgentMail (@agentmail): so agents can have email accounts 2. AgentPhone (@tryagentphone): so agents can have phone numbers 3. Kapso (@andresmatte): so agents can have WhatsApp phone numbers 4. Daytona (@daytonaio) / E2B (@e2b): so agents can have their own computers 5. Browserbase (@browserbase) / Browser Use (@browser_use) / Hyperbrowser (@hyperbrowser): so agents can use web browsers 6. Firecrawl (@firecrawl): so agents can crawl the web without a browser 7. Mem0 (@mem0ai): so agents can remember things 8. Kite (@GoKiteAI) / Sponge (@PayspongeLabs) : so agents can pay for things. 9. Composio (@composio): so agents can use your SaaS tools 10. Orthogonal (@orthogonal_sh) so agents can access APIs easily 11. ElevenLabs (@ElevenLabs) / Vapi (@Vapi_AI) so agents can have a voice 12. Sixtyfour (@sixtyfourai) so agents can search for people and companies. 13. Exa (@ExaAILabs): so agents can search the web (Google doesn’t work for agents) If you stitch all of these together, you get a digital coworker that looks more human than AI.
@AlexFinn ·
My mind is so blown I have my own personal AI research lab running 24/7/365 I'm just one dude with an entire team of AI agents training models and doing R&D I think this is the biggest opportunity right now: taking Karpathy's Autoresearch framework and applying it to everything I have a team of AI agents running experiments all day and night on system prompts, local models, and LoRAs. I also have them doing R&D on my new project. They spend all day discussing my app, coming up with new ideas, then debating eachother An entire organization of autonomous agents continuously improving my business 24/7/365 I feel like I have unlimited power Right now they are all running on ChatGPT 5.4, but today I will move them to local models running on my 3 Mac Studios and DGX Spark so this will all become free Free, local super intelligence working for me at all times. 10 year old me would think this is a scifi Do this immediately: 1. Ask your agent about Karpathy's Autoresearch. Deeply understand it 2. Ask your agent how you could apply that framework to other projects you're working on 3. Download a local model. Doesn't matter what computer you have. There is a model you can run on it. 4. Just get used to how it works. Learn from it. 5. Push yourself to get uncomfortable every day and try new things. There has never been a better/more profitable time to be a tinkerer
@_avichawla ·
Microsoft did it again! Building with AI agents almost never works on the first try. A dev has to spend days tweaking prompts, adding examples, hoping it gets better. This is exactly what Microsoft's Agent Lightning solves. It's an open-source framework that trains ANY AI agent with reinforcement learning. Works with LangChain, AutoGen, CrewAI, OpenAI SDK, or plain Python. Here's how it works: > Your agent runs normally with whatever framework you're using. Just add a lightweight agl.emit() helper or let the tracer auto-collect everything. > Agent Lightning captures every prompt, tool call, and reward. Stores them as structured events. > You pick an algorithm (RL, prompt optimization, fine-tuning). It reads the events, learns patterns, and generates improved prompts or policy weights. > The Trainer pushes updates back to your agent. Your agent gets better without you rewriting anything. In fact, you can also optimize individual agents in a multi-agent system. I have shared the link to the GitHub repo in the replies!
@asmah2107 ·
The reading list that taught me how to think about agentic architecture. Bookmark this. 1. Brewer's CAP Theorem (2000) — trade-off thinking 2. Netflix Hystrix docs — circuit breaker pattern 3. Martin Fowler: Saga Pattern — distributed rollback 4. The Twelve-Factor App — stateless service design 5. AWS Well-Architected Framework — blast radius thinking 6. "Thinking in Systems" — Donella Meadows 7. Designing Data-Intensive Applications — Kleppmann 8. Google SRE Book Ch.13 — cascading failures 9. OWASP LLM Top 10 (2025) — agent attack surfaces 10. Anthropic: Building Effective Agents (2024) 11. LangGraph docs — stateful agent patterns 12. Microsoft AutoGen paper — multi-agent orchestration 13. Gartner: Agentic AI Hype Cycle (2025) 14. EU AI Act Article 14 — human oversight requirements Classic distributed systems stuff. Applied to the next layer of the stack. Follow for annotated breakdowns → @asmah2107
@farzyness ·
I'm having a lot of success with the following @openclaw stack: Primary Agent (Claw) with @claudeai Opus 4.6 1M token context + thinking high: processes everything I need it to do - from simple to complex tasks. Then for complex tasks that require some sort of tool/software building, I have the primary agent automatically talk to a different @openclaw agent that's optimized for development (Webby) - using same "brain" set up as above with Opus 4.6, but it uses @ChatGPTapp CODEX with GPT 5.4 thinking xhigh to do all its actual coding. Then, if I need anything researched/fact checked that's related to research, script building - basically anything that requires up-to-date or accurate information about the world - all agents reach out to Scout (research agent) that is specialized in research and will always use @Grok 420 multi agent. I find that creating individual agents for the primary functions of a business is FAR better than having a do-everything agent. Thought process is that context windows are super valuable real estate, and using OpenClaw core files (soul, identity, agents, etc) you can use those to maximize the capability of each agent to be extra-good at their role. Then because each agent can talk to each other, they are basically sending their own sub-queries to the LLM that are fully optimized for each tasks that you're trying to solve. Said another way - if you are building a business where the AI agents are at the foundation, make sure you are defining the roles that you need in that business - and then create an AI agent that is in charge of each role. For example, for my YouTube channel, these are the major categories as I think about them - idea generation, research, scripting & fact checking, titles and thumbnails, performance tracking. Before AI agents, I did that whole chain. Major YouTubers would have a production team that would execute that with them. But with AI agents, you can specialize each one to do each one of those steps. And then on top of that, each one of those agents probably should be under a different model - in my experience, Claude is by far the most creative and the best writer. Grok is by far the best at research and fact checking. ChatGPT is by far the most autistic (ie best for coding/development of tools). Gemini is by far the best at framing titles and thumbnails (shocker - they have all YouTube data). So as you use AI agents, it is EXTREMELY important that you think of them as literal humans that you are working with to build an organization, and give them the tools necessary (and the brains necessary) to set them up for success.
@ujjwalscript ·
Your AI Agent is mathematically guaranteed to FAIL. This is the dirty secret the industry is hiding in 2026. Everyone on your timeline is currently bragging about their "Multi-Agent Swarms." Founders are acting like chaining five AI agents together is going to replace their entire engineering team overnight. Here is the reality check: It’s a mathematical illusion. Let’s look at the actual numbers. Say you have a state-of-the-art AI agent with an incredible 85% accuracy rate per action. In a vacuum, that sounds amazing. But an "autonomous" workflow isn't one action. It’s a chain. Read the ticket ➡️ Query the DB ➡️ Write the code ➡️ Run the test ➡️ Commit. Let's do the math on a 10-step process: $0.85^10= 0.19$ Your "revolutionary" autonomous system has a 19% success rate. And the real-world data proves it. Recent studies out of CMU this year show that the top frontier models are failing at over 70% of real-world, multi-step office tasks. We are officially in the era of "Agent Washing." Startups are rebranding complex, buggy software as "autonomous agents" to look cool, but they are ignoring the scariest part: AI fails silently. When traditional code breaks, it crashes and throws a stack trace. When an AI agent breaks, it doesn't crash. It just confidently hallucinates a fake database entry, sidesteps a broken API by faking the response, and keeps running—corrupting your data for weeks before you notice. If your "automated" system requires a senior engineer to spend three hours digging through prompt logs to figure out why the bot made a "creative decision," you didn't save any time. You just invented a highly expensive, unpredictable form of technical debt. Stop trying to build fully autonomous swarms to replace human judgment. Start building deterministic guardrails where AI is the engine, but the engineer holds the steering wheel
@goyalshaliniuk ·
Building Agentic AI Systems? Let me share a Comprehensive Guide with you. When designing smarter AI systems, understanding how agents think, act, and collaborate is the key. Here's a breakdown of the essentials: 1) What is an AI Agent? An AI agent is a system that thinks, plans, and acts independently using LLMs, adapting in real-time with minimal human help. 2) Architectures for Building AI Agents Start with single agents for simple tasks or build multi-agent systems where different agents collaborate and solve complex problems together. 3) Foundations of Agent Design Good agents combine powerful models, smart memory, useful tools, structured instructions, and seamless orchestration to work effectively. 4) Building Guardrails for Safety and Control Safety measures like data privacy, dynamic rules, and human checkpoints keep agents responsible and prevent harmful outcomes. 5) Essential Tools for Building Agentic AI Successful agents rely on LLMs, memory systems, orchestrators, integration tools, guardrails, and monitoring platforms to operate smartly.
@goyalshaliniuk ·
Confused about the different types of AI agents? Understanding the various agent types is key to designing intelligent systems that react, plan, and learn effectively. Here's a simple breakdown of the 5 major types of AI agents and how they work. 1. Simple Reflex Agents These agents act solely on current inputs using basic condition–action rules. They don’t learn or remember past states, making them fast but limited in capability. 2. Model-Based Reflex Agents These agents use internal models to track the world’s current state. They can handle partially observable environments by remembering past inputs and updating their internal state. 3. Goal-Based Agents Rather than reacting blindly, these agents plan actions to achieve specific goals. They evaluate consequences and choose actions that bring them closer to their objectives. 4. Utility-Based Agents Going beyond goals, utility-based agents aim to maximize happiness or usefulness. They weigh different outcomes and choose the one that offers the best result based on a utility function. 5. Learning Agents These agents evolve over time. They learn from feedback, improve performance, and explore new ways to act better in future environments. ✅ Use this guide to pick the right agent architecture for your next AI system—whether you're building a rule-based chatbot or an intelligent decision-maker.
@om_patel5 ·
THE BEST OPEN-SOURCE AI AGENT REPO I'VE SEEN IN A WHILE most "agent" repos are just glorified prompt folders. this one actually feels like a real team. it's called The Agency and it gives you 140+ specialized AI agents you can drop into your workflow right now. not vague "act like a developer" prompts. actual specialists with: > clear personalities > defined workflows > real deliverables > success metrics > production-ready behavior think: > frontend developer for react/ui work > backend architect for apis + infra > reddit community builder for authentic growth > whimsy injector for adding delight without ruining ux > security engineer for threat modeling and safe code > technical writer for docs that don't suck and it keeps going. engineering. design. marketing. sales. product. qa. ops. even weirdly specific specialists for things like MCP servers, LSP/indexing, compliance, paid media, and spatial computing why this is actually cool: 1\ it's not locked to one tool > works with claude code out of the box > can also be converted/installed for cursor, copilot, aider, windsurf, gemini cli, opencode, qwen code, and more > so you're not betting everything on one editor or one ai workflow 2\ these agents are built like systems, not one-liners > each file includes identity, tone, rules, workflows, examples, and expected outputs > that means more consistent results and less babysitting 3\ it was clearly shaped by real usage > born from a reddit thread > iterated for months > 50+ requests in the first 12 hours > battle-tested instead of vibe-assembled 4\ the use cases are actually practical > building an MVP with a frontend dev + backend architect + growth hacker > launching campaigns with content + twitter + reddit + analytics agents > shipping enterprise features with PM + senior dev + designer + QA agents 5\ it's open source > MIT licensed > transparent > forkable > customizable > PRs welcome my favorite part: most people use AI like one overworked intern. this repo makes AI feel more like hiring a full stack agency with specialists who already know their job. that shift matters. instead of saying: > "help me with my product" you can say: > "activate frontend developer mode and build this react component" or > "use the reddit community builder and help me grow this without getting banned" that specificity is where the magic starts the big takeaway: > generic prompts give generic output > specialized agents give you sharper thinking, better process, and more usable work if you're using claude code, cursor, or any agentic coding setup, this repo is worth studying even if you never use it directly because it shows what good agent design actually looks like star-worthy for sure
@ihteshamali ·
Andrej Karpathy built autoresearch an AI that writes and improves its own research papers overnight. Someone just did the same thing for AI agents. It's called AutoAgent. You tell it what kind of agent to build. It builds it, tests it, scores it, and improves it in a loop without you touching a line of code. Here's the part that makes it different from every other agent framework: You don't engineer the harness. You program the meta-agent. There's a single file called program.md. That's the only file you touch. It gives the meta-agent its directive and context. Everything else the system prompt, the tools, the orchestration, the config gets modified autonomously. The loop looks like this: → Meta-agent reads your directive → Inspects https://t.co/lvRfwACuBO (the entire harness in a single file) → Makes one targeted change to prompt, tools, or routing → Runs benchmark tasks inside Docker fully isolated → Reads the score (0.0 to 1.0) from the task evaluators → Commits the change if score improves, reverts if it doesn't → Repeats The benchmark format is harbor-compatible, so you can drop in evaluation datasets from real AI labs and the harness runs against them without modification. You wake up in the morning and the agent is better than when you left it. That's the entire pitch. MIT License. 100% Opensource. Link in comments.
@_philschmid ·
Just finished my talk on “Why do (Senior) Engineers struggle to build AI Agents” If you missed or want to read about, blog post below. To succeed, we have to accept: 1️⃣ Text is the new state. 2️⃣ Hand Over Control. 3️⃣ Errors Are Just Inputs. 4️⃣ Move from Unit Tests to Evals. 5️⃣ Agents Evolve, APIs Don't.
@shedntcare_ ·
Twenty AI researchers gave AI agents access to their emails, files, Discords, and terminals. Two weeks later, the agents had: • Obeyed strangers • Leaked sensitive information • Executed destructive commands • Spread unsafe behaviors • Claimed tasks were complete when they weren't The most alarming finding wasn't that the AI made mistakes. It's that it confidently reported success while reality said otherwise. "In several cases, agents reported task completion while the underlying system state contradicted those reports." Think about that. AI agents are being integrated into customer support, finance, HR, operations, and infrastructure. Most companies assume: 1. The AI follows the right instructions. 2. The AI acts only when authorized. 3. The AI accurately reports what it did. This research suggests all three assumptions can fail. The age of AI agents has arrived. The age of trusting them blindly should not.
@burkov ·
LLMs have moved from producing standalone code to powering AI agents — systems that plan over many steps, call external tools, keep track of changing state, and recover from their own errors during long-running tasks. Inside these systems, code has become the working material for almost everything: agents write small programs to reason through math, to drive a browser, to query a database, to test their own outputs, and to share intermediate work with other agents through files in a repository. Existing surveys still treat code mainly as the final answer that a model produces. The authors of this paper argue this misses the broader role code now plays inside agent systems. They use the term "agent harness" for the scaffolding around a model (planning, memory, tool calls, and execution-based verification loops) that turns one-shot generation into something that can run for hours on a real task, and they organize their map of this scaffolding into three connected layers: the interface, covering how code lets an agent reason, act, and represent its environment; the mechanisms, covering how planning, memory, tool use, and feedback-driven repair keep long executions on track; and the scaling layer, covering how several agents share repositories, tests, and execution traces to coordinate. Each layer is tied to concrete systems such as Claude Code and Codex, and to application areas including software engineering, GUI and OS automation, embodied robots, scientific discovery, and personalization. Read with an AI tutor and quizzes for better retention: https://t.co/Lu0KuB3Ff9 PDF: https://t.co/Cf7pZELsGA
@techNmak ·
Building AI agents is easy, getting them to survive production is the hard part. We need more repos like this, ones that don’t just talk about AI agents, but actually show you how to get them working in production. Huge credit to Nir Diamant for putting together "Agents Towards Production." It’s a hands-on, code-first guide packed with runnable tutorials for every step: orchestration, memory, security, observability, deployment, and more. The tutorials are written with extra care to be easy to follow, you can understand everything without even reading the code. 😊 Each section is practical, real notebooks, real scripts, and patterns that reflect what actually happens when you move from prototype to production. There’s a clear focus on things most projects skip: persistent memory, robust tool integration, tracing, and real-world error handling. Deployment and security get the attention they deserve too. And it’s not a static project, the repo is under constant supervision, with new tutorials and updates being added all the time. If you’re serious about building GenAI agents that last beyond a demo, this repo is a standout example of how to do it right.
@_jaydeepkarale ·
Most people think AI agents are just “LLMs with tools.” But the interesting part is memory. Just like humans, capable AI agents need different kinds of memory to function properly. This is one of the core ideas behind the COALA framework (Cognitive Architectures for Language Agents). Think about how humans work: - You remember what someone said 10 seconds ago. - You remember facts from school. - You remember how to ride a bicycle. - You remember important life experiences. AI agents need similar layers of memory too. Here are the 4 major types: 1. Working Memory (Short-Term Memory) Human analogy: You’re reading a sentence right now while remembering the previous sentence. It’s temporary memory used for the current task. For AI agentst his is the active context window which includes: - current conversation - recent tool outputs - current reasoning chain Without it, the agent loses track mid-task like a human getting distracted every 5 seconds. ------------ 2. Semantic Memory (Factual Knowledge) Human analogy: You know Paris is the capital of France. You know Kubernetes manages containers. These are facts, concepts, and knowledge. For AI agents this includes: - stored facts - documentation - knowledge bases - vector databases - retrieved company information It answers: “What does the agent KNOW?” CLAUDE.MD is an example of Semantic Memory ------------ 3. Procedural Memory (Learned Skills) Human analogy: You don’t consciously think about every muscle movement while riding a bike. Skills become automatic. For AI agents this is: - workflows - system prompts - learned action patterns - tool usage strategies - step-by-step execution habits Example: An agent learns: “First query database → then validate → then summarize.”. Agent Skills or Skills.md are examples of procedural memory It answers: “What does the agent KNOW HOW TO DO?” ------------ 4. Episodic Memory (Past Experiences) Human analogy: You remember your first interview. Or a production outage you once handled at work. These are experiences tied to events. For AI agents this includes: - previous interactions - past successes/failures - user preferences - historical task outcomes Example: “The last deployment failed because of missing env variables.” It answers: “What has the agent EXPERIENCED before?” ------------ This is why memory is becoming one of the biggest frontiers in AI engineering. A model without memory is just reacting. An agent with memory starts behaving more like a system that learns, adapts, and improves over time.
@rohanpaul_ai ·
Harvard Business Review just published a piece. A good AI agent needs a job description, limits, and a manager. Because, AI agents can fail like employees with too much access and too little supervision. firms keep treating agents like normal software, even though the real risk is not bad text but bad actions. That changes 4 things: each agent needs its own identity and permissions, its own trusted data sources, hard rule checks between a model and any real transaction, and a full audit trail of what it read, decided, and did. So the safe rollout path is an autonomy ladder where agents start with drafts and recommendations, then move to guarded retrieval, then supervised actions, and only later get narrow bounded autonomy.
@pauliusztin_ ·
Refined my workshop on multi-agent systems from scratch after doing it at the @aiDotEngineer and Uphill conferences In the latest iteration, we introduced an `implement_yourself` option that lets you implement everything yourself using agentic coding best practices. In other words, you build agents with agents. We added all the skills and subagents to incrementally build elements from the system, read it, understand it and run it, until you have the e2e system working. After iterating and refining it with @Whats_AI so many times, it's an amazing free resource to get into: - MCP servers - Designing agents - Adding observability and evals You can find everything on GitHub: https://t.co/O7UXJG6UOw
@_vmlops ·
Andrew Ng just shared a simple framework that changes how you build AI agents The shift is: Loop Engineering → Graph Engineering The core idea: Loops make agents think Graphs make agents remember Here's the 4-step workflow: 1. Reflection Generate → Critique → Rewrite A single self-review loop can outperform a stronger model that never reviews its work 2. Tool Use Connect your agent to search, APIs, databases, and code execution. Without tools, it's just guessing 3. Planning Break complex tasks into structured steps before execution If one step fails, the agent adapts instead of starting over 4. Multi-Agent Systems Don't rely on one agent. Use specialized agents for planning, coding, reviewing, and testing Then take it further: → Add a critique step after every generation → Connect agents through a graph so they share memory and state instead of long chat histories The biggest takeaway: Better architecture beats bigger models. A well-designed workflow can consistently outperform a stronger model running without one This is where AI agent development is heading
@alex_verem ·
Free University of Bozen-Bolzano, Oxford, University of Tartu, IBM Research, and a dozen other institutions just published the most important framework nobody in the AI agent space is reading. Everyone is building AI agents. Nobody is governing them. The academic world just wrote the manual. > The problem isn't that AI agents don't work. It's that when they fail, break compliance rules, contradict each other, or drift from their original behavior after self-modification, nobody has a systematic way to catch it, explain it, or fix it. The current industry approach is vibes-based governance system prompts, retry logic, and prayer. > This manifesto introduces Agentic Business Process Management as the missing layer. The core idea: every AI agent operating inside an organization needs process awareness a continuous understanding of what the organization is trying to accomplish, what constraints apply, and how its decisions align with or violate those goals. Without this, you don't have an AI system. You have an autonomous entity making consequential decisions inside your company with no frame of reference for what it's supposed to be doing. > The framework identifies four capabilities every governed AI agent must have. Framed autonomy means the agent operates within explicit normative constraints not just "follow these instructions" but a formal specification of what it's permitted, obligated, and prohibited from doing. > Explainability means the agent can articulate why it made each decision, not as a post-hoc justification but as a live, auditable trace. Conversational actionability means it can coordinate with humans and other agents in natural language while actually executing process steps not just talking about them. Self-modification means it can adapt to new conditions and evolve its behavior over time without breaking the frame it operates within. > The liability finding is the one that should stop every enterprise AI team cold. When an agent self-modifies and its behavior drifts significantly from its original design, who is responsible for the outcome? The software developer? The organization that configured the frame? The AI model provider? The multi-agent system that produced emergent behavior nobody anticipated? The EU AI Act is coming. None of the current agentic AI deployments have a clear answer to this question. → Models trained for mathematical domains generalize poorly to legal, medical, and business contexts the domains where governed agents are being deployed → CTGAN debiasing changes the direction of disparity without eliminating it purely algorithmic fixes cannot substitute for structural governance → Framing must be both normative (what the agent is allowed to do) and operational (how it achieves goals within those constraints) → Self-modification creates a responsibility gap no current legal framework adequately addresses → Benchmark contamination is identified as a cross-cutting threat — agents employing models trained on benchmark data will fail in real-world conditions The uncomfortable reality this manifesto exposes: the industry is deploying agentic systems at scale while the governance infrastructure is still theoretical. The gap between "we built an AI agent" and "we can audit, explain, and control what that agent does inside our organization" is enormous and it's widening every quarter.
@ujjwalscript ·
How to be a REAL AI Engineer (as opposed to a "Prompt Engineer") by learning the 4-Core System: Note: Being an AI Engineer is about building autonomous, production-grade agentic systems that solve real problems. 1. The "Brain" (Foundational Models & Routing): You don't just use one model anymore. You route them based on cost and latency. Heavy Lifting: Opus, Gpt-5.4 for deep reasoning and complex logic. Fast/Cheap: Open-source models (like Llama) for high-volume, low-latency micro-tasks. 2. The "Memory" (Embeddings & Vector Databases): AI models are stateless. You have to build their memory. Vector DBs: Pinecone, Qdrant, or Milvus. The secret isn't just storing vectors; it's mastering metadata filtering to prevent context pollution. Embedding Models: OpenAI’s latest embedding models or open-source equivalents like BGE for semantic search. 3. The "Nervous System" (Agent Orchestration & Pipelines): You are no longer writing linear scripts; you are managing a digital workforce. LangGraph & CrewAI: The 2026 industry standards for multi-agent workflows and cyclic graphs. PydanticAI: For strictly typed, validated AI outputs. If you aren't forcing your agents to return validated JSON, your app will crash in production. 4. The "Hands" (Tool Use & Action): An agent that can't take action is just a toy. API Design: Build strict, secure tools (using FastAPI or Node) that your agents can trigger autonomously. Web Automation: Tools like Firecrawl to let your agents research, scrape, and interact with the live internet.
@Al_Grigor ·
How to evaluate AI agents step-by-step? 1. Verify goal understanding 2. Assess plan quality 3. Inspect tool execution 4. Compare plan and execution 5. Evaluate replanning 6. Measure efficiency 7. Review end-to-end consistency 🧵
@far33d ·
I've been using AI code tools a LOT recently - mostly as a way to turn ideas in my head into something concrete to react to vs. trying to build finished products. This week, the goal was to prototype multi-player AI interactions and try to fix the problems I have writing in AI tools today. Most AI writing tools miss the point. They either rewrite your entire document or exile you to a sidebar chat. What if AI worked more like a team of editors? That's ai-writer: multiple AI agents, each with its own voice, collaborating in the margins like human editors. Highlight text, @-mention an agent, get a targeted rewrite. Accept or reject inline. The interface stays the same whether you're talking to a person or a bot. It's just a prototype. Be nice. The best part of the design: human comments and AI suggestions work the same way. AI suggestions are just comments with an edit attached. Same system, same look, same process. Build the comment system once, and you get AI editing for free. No separate "AI mode." No special steps for AI changes. The comment system went through four distinct versions, each one triggered by actually using the previous version and finding it insufficient. Iterating on the idea until it felt more right, finally landing on a series of AI helpers that can interact with the user (or a team in the future) in the doc or in comments. This is a prototype, not a product. But the pattern feels strong: agents as collaborators in context, not as separate tools.
@aiwithmayank ·
I found a repo that removes the most exhausting part of working with AI agents. The part where you manually read failure logs, guess what went wrong, rewrite the prompt, test it again, and repeat for days. It's called meta-agent. The idea is simple but the implications are wild. You tell it what your agent is supposed to do. You tell it how to check if it worked. Then you walk away. meta-agent runs your agent, catches every failure, feeds those failures to a second AI that acts as a coach, and the coach rewrites the instructions for the next run. It keeps looping until the agent stops getting better. Real result from their benchmark: 67% accuracy to 87% accuracy. Automatically. Overnight. If you're building anything with Claude Code, customer service bots, research agents, or workflow automation, this is the missing piece nobody talks about. Your AI should be fixing itself. Most aren't. MIT License. 100% Opensource.
@sharbel ·
Most founders are building AI agents backward. They start with: 1. Pick a tool 2. Write a prompt 3. Hope it does useful work That is why the agent feels impressive once, then disappears from the workflow. Better order: 1. Pick a recurring job 2. Write the decision rules 3. Define the inputs 4. Define the output format 5. Add a review step 6. Run it on a schedule 7. Improve it from failures The tool matters less than the job design. A mediocre agent with a clear job beats a powerful agent with vague instructions.
@xelebofficial ·
AI agents are becoming a new kind of user. That is the part many teams are still underestimating. Once an agent can access data, trigger tools, move across workflows, or interact with customers, it is no longer just a feature. It becomes an actor inside the system. And every actor needs boundaries. What can it access? What can it execute? What should it remember? What should it escalate? What should be logged? Who is responsible when it gets something wrong? This is where the next layer of AI agent infrastructure will matter. Not just better reasoning. Better identity. Better permissions. Better observability. Better accountability. The future of AI agents will not be built only around what agents can do. It will be built around what agents are trusted to do.
@MaryamMiradi ·
After building 400+ Production AI Agents I curated 7 Crucial Skills so you can become a Production AI Agents Engineer. 𝟭. 𝗦𝘆𝘀𝘁𝗲𝗺 𝗗𝗲𝘀𝗶𝗴𝗻 If you have built a backend with multiple services talking to each other, you already think this way. If not, this is where I would start. An agent is not one thing. Everything needs to connect. Data needs to flow. Failures need to be contained. 𝟮. 𝗧𝗼𝗼𝗹 𝗮𝗻𝗱 𝗖𝗼𝗻𝘁𝗿𝗮𝗰𝘁 𝗗𝗲𝘀𝗶𝗴𝗻 If your tool schemas are vague, agent fills the gap with imagination. → LLM imagination is creative. That is the problem. → In a financial transaction, creative is dangerous. I now write every schema like a legal contract. Strict types. Explicit examples. Zero ambiguity. 𝟯. 𝗥𝗲𝘁𝗿𝗶𝗲𝘃𝗮𝗹 𝗘𝗻𝗴𝗶𝗻𝗲𝗲𝗿𝗶𝗻𝗴 We know RAG or mayebe not? Chunk too large, you drown the signal. Chunk too small, you lose the context. Retrieval quality sets the ceiling on everything your agent can do. The LLM cannot compensate for garbage context. Do Re-ranking. 𝟰. 𝗥𝗲𝗹𝗶𝗮𝗯𝗶𝗹𝗶𝘁𝘆 𝗘𝗻𝗴𝗶𝗻𝗲𝗲𝗿𝗶𝗻𝗴 APIs fail very often. You need: ✸ Retry logic with exponential backoff ✸ Timeouts so the agent does not hang indefinitely ✸ Fallback paths when the primary tool breaks ✸ Circuit breakers to stop one failure taking down everything Backend engineers have solved this for decades. Most people building agents right now have not read that playbook yet. 𝟱. 𝗦𝗲𝗰𝘂𝗿𝗶𝘁𝘆 𝗮𝗻𝗱 𝗦𝗮𝗳𝗲𝘁𝘆 Prompt injection is real. Your agent is an attack surface. Do this: Input validation. Output filters. Permission boundaries. Least privilege applies to agents exactly as it applies to human employees. 𝟲. 𝗘𝘃𝗮𝗹𝘂𝗮𝘁𝗶𝗼𝗻 𝗮𝗻𝗱 𝗢𝗯𝘀𝗲𝗿𝘃𝗮𝗯𝗶𝗹𝗶𝘁𝘆 Remember to Trace: Every decision. Every tool call. Every retrieval result. I run evaluation pipelines with known good answers before anything goes to production. If you cannot measure it, do not deploy it. 𝟳. 𝗣𝗿𝗼𝗱𝘂𝗰𝘁 𝗧𝗵𝗶𝗻𝗸𝗶𝗻𝗴 Your agent serves a human. That human wants to know when the agent is confident and when it is not. They need to know when to trust it and when to escalate. 𝗥𝗲𝗺𝗲𝗺𝗯𝗲𝗿: AI Agents Engineering is here to Stay. Master it.
@ttunguz ·
If 2025 is the year of agents, then 2026 will surely belong to agent managers. Agent managers are people who can manage teams of AI agents. How many can one person successfully manage? I can barely manage 4 AI agents at once. They ask for clarification, request permission, issue web searches—all requiring my attention. Sometimes a task takes 30 seconds. Other times, 30 minutes. I lose track of which agent is doing what & half the work gets thrown away because they misinterpret instructions. This isn’t a skill problem. It’s a tooling problem. Physical robots offer clues about robots manager productivity. MIT published an analysis in 2020 that suggested the average robot replaced 3.3 human jobs. In 2024, Amazon reported pickpack and ship robots replaced 24 workers. But there’s a critical difference : AI is non-deterministic. AI agents interpret instructions. They improvise. They occasionally ignore directions entirely. A Roomba can only dream of the creative freedom to ignore your living room & decide the garage needs attention instead. Management theory often guides teams to a span of control of 7 people. Speaking with some better agent managers, I’ve learned they use an agent inbox, a project management tool for requesting AI work & evaluating it. In software engineering, Github’s pull requests or Linear tickets serve this purpose. Very productive AI software engineers manage 10-15 agents by specifying 10-15 tasks in detail, sending them to an AI, waiting until completion & then reviewing the work. Half of the work is thrown away, & restarted with an improved prompt. The agent inbox isn’t popular - yet. It’s not broadly available. But I suspect it will become an essential part of the productivity stack for future agent managers because it’s the only way to keep track of the work that can come in at any time. If ARR per employee is the new vanity metric for startups, then agents managed per person may become the vanity productivity metric of a worker. In 12 months, how many agents do you think you could manage? 10? 50? 100? Could you manage an agent that manages other agents? https://t.co/iHl9JXipNJ
@bibryam ·
👍 The failure of "AI agents" is not a failure of intelligence but a failure of architecture → Use LLMs for interpreting intent, generating content, and understanding context. → Use deterministic code for actually executing tasks, managing state, handling errors, and delivering consistent outcomes. The Agentic AI Delusion https://t.co/maoFkMhrGz
@vicky_grok ·
Most AI "agents" shared online aren't actually agents. They're orchestrated workflows. There's nothing wrong with orchestration. It's powerful and useful. But understanding the difference helps you choose the right architecture for the right problem. Here's the distinction. An orchestrator follows a predefined sequence. It executes a workflow that was designed in advance. A true AI agent starts with a goal, decides what to do next, adapts when conditions change, and continues until the objective is achieved. Think of it this way. An orchestrator follows instructions. An agent creates and updates its own plan. An orchestrator typically: • Executes fixed workflows • Runs predefined steps • Uses rule-based logic • Doesn't adapt automatically • Has limited or no memory • Doesn't reflect on previous outcomes • Cannot make independent decisions outside its workflow A true AI agent can: • Understand a goal • Plan its own approach • Select and use tools • Maintain memory across tasks • Evaluate intermediate results • Learn from previous attempts • Adjust its strategy when needed • Decide when the task is actually complete A practical example. Customer support chatbot If it always follows: Receive ticket ↓ Look up FAQ ↓ Generate reply ...it's an orchestrator. Now imagine another system that: Receives the request ↓ Plans multiple possible solutions ↓ Searches documentation ↓ Uses internal tools ↓ Checks whether the issue is resolved ↓ Asks follow-up questions if needed ↓ Changes its strategy until the goal is achieved That's much closer to agentic AI. As AI systems continue to evolve, you'll see more products described as "AI agents." The important question isn't what they're called. It's whether they can: • Reason • Plan • Use tools • Remember • Evaluate • Adapt • Act autonomously toward a goal If most of those capabilities are missing, you're likely looking at an orchestrated workflow rather than a true AI agent. Both approaches have their place. Use orchestrators for predictable, repeatable processes. Use agentic systems when the problem requires planning, adaptation, and autonomous decision-making. ---- If you enjoy practical AI guides, workflows, prompt libraries, and free learning resources, subscribe to the ByteBuilders newsletter: https://t.co/IAjuMwTAmm
@CryptoTeca__ ·
Three standards are quietly becoming the core infrastructure for autonomous AI commerce: t54. ERC-8004. x402. Each solves a different problem, but together they let AI agents discover one another, establish trust, make payments, and transact safely without constant human involvement. Here's how I see the stack. [1] ERC-8004 is the identity layer. It gives AI agents onchain identity, reputation, discovery, and verification so they can answer: "Who is this agent, and can I trust it?" [2] x402 is the payment layer. By reviving HTTP 402, it enables instant stablecoin payments for APIs, AI models, compute, data, and digital services without API keys or subscriptions. [3] t54 is the financial safety layer. It adds identity verification, fraud detection, underwriting, compliance, credit, and risk controls through x402 Secure, Claw Credit, and the Agentic Risk Standard, helping answer another critical question: "Should this transaction happen safely?" Together, the workflow is simple: ▸ ERC-8004 discovers and verifies an agent. ▸ x402 handles the payment request. ▸ t54 performs risk and compliance checks. ▸ Payment settles onchain. ▸ Reputation updates for future interactions. What's even more interesting is that builders are already filling different parts of the stack. > Identity & discovery: @TheGraph, @QuickNode, @elizaOS, @daydreamsagents, and @Ch40sChain are building identity, discovery, reputation, and trust infrastructure around ERC-8004. > Payments & commerce: @game_virtuals, @cookiedotfun, @aixbt_agent, and @arcdotfun are exploring AI agent payments and commerce through x402-powered interactions. > Compute & infrastructure: @rendernetwork, @akashnet, and @BlockRunAI provide the compute and AI services autonomous agents can access on demand. > Machine identity & robotics: @virtuals_io, @iotex_io, @FabricFND, and @AukiLabs connect AI agents with machines, wallets, devices, and spatial intelligence. > Security & trust: @HallidayHQ, @LitProtocol, and @redstone_defi focus on guardrails, key management, and risk intelligence for autonomous agents. > Agent networks: @NEARProtocol, @opentensor, @modenetwork, and @swarmnode are expanding the infrastructure for decentralized AI agents and autonomous coordination. The biggest takeaway for me is that these aren't competing standards. They're complementary layers. ▸ ERC-8004 provides identity and trust. ▸ x402 enables machine-native payments. ▸ t54 adds the financial safety layer. Together, they look increasingly like one of the foundational infrastructure stacks powering the emerging AI agent economy.
@hugobowne ·
AI agents are failing silently in production, and it's costing companies tens of thousands of dollars before anyone notices. Here's what 1,400+ real deployments actually taught us: - The $50k infinite loop: agents confidently report success while spiralling into expensive mistakes. Silent failures are the biggest risk nobody talks about. - Sub-400ms voice agents: Elyos built them by aggressively throwing away context every few seconds. Extreme? Yes. Increasingly standard? Also yes. - DoorDash's three-tier architecture: manager, progress tracker, specialists, with a persistent workspace letting agents collaborate across hours or days. All from AgentOps: Lessons from Over 1,400 Production Deployments of AI Systems, built on the brilliant LLMOPs database by @strickvl (@zenml_io). Full breakdown on my Substack 👇
@rohanpaul_ai ·
Stronger agents will not come only from larger models, but from better systems around them. The problem is that many AI agents are judged as if the model alone did the work, even though the real behavior also depends on memory, tools, context, routing, checks, and permissions. This surrounding setup around the agent is called harness, meaning the system that decides what the model sees, what tools it can use, what it remembers, and what actions get checked. Progress should come from scaling this harness, especially 3 parts: better context control, more trustworthy memory, and better routing to tools or helper agents. Long context is not the same as usable context, memory is not the same as trustworthy memory, and having many tools is not the same as knowing when to use them. A stale note can be more dangerous than no note, because it gives the agent confidence exactly when it should re-check the world. A specialized subagent can also fail quietly if its output sounds plausible but no later layer verifies whether it is true. This is why one-shot benchmark scores feel increasingly thin. Two agents can reach the same final answer, while one burns far more tokens, makes riskier tool calls, carries corrupted memory, or succeeds only by accident. The next frontier is not just scaling the mind inside the machine. It is scaling the discipline around it. ---- – arxiv. org/abs/2605.26112 Title: "From Model Scaling to System Scaling: Scaling the Harness in Agentic AI"
@alvinfoo ·
Most people think Claude Code is just a coding assistant. It’s not. It’s an entire agent development platform — and most are only using 10% of its power. The real breakthrough is in its architecture: CLAUDE.md + Skills + Hooks + Subagents + Plugins = The Agent Development Kit Here’s how it actually works: 1. CLAUDE.md (Memory Layer) The foundation. Defines rules, structure, and context, like a “constitution” for your AI agent. Always loaded. Always guiding behavior. 2. Skills (Knowledge Layer) Reusable capabilities your agent can call on demand. Not always active, only triggered when needed. Think modular intelligence. 3. Hooks (Guardrail Layer) Where control happens. Pre/post actions, validations, safety checks. This is how you enforce quality and prevent mistakes at scale. 4. Subagents (Delegation Layer) This is where it gets powerful. You don’t just use AI, you orchestrate teams of AI. Each subagent handles a specific task with its own context. 5. Plugins (Distribution Layer) Package and scale everything. Turn capabilities into reusable tools across teams. The shift is clear: We’re moving from → writing code to → designing systems that produce outcomes From → single assistants to → coordinated AI agents working in parallel The biggest mistake right now? Treating LLMs with better autocomplete. The winners will be the ones who learn how to build, structure, and deploy agent systems, not just prompt them. AI isn’t just helping you code anymore. It’s becoming your execution layer.
@thetripathi58 ·
AI agents are failing at complex tasks because we keep hardcoding their workflows. Researchers just released a paper on the Mimosa Framework. It proves that static multi-agent systems are a dead end. The solution? Agents that build and evolve their own workflows on the fly. Here is what the research actually reveals about building agentic AI: The Static Architecture Trap Situation: You build an AI agent system. You explicitly define step one, step two, and step three. It works perfectly in testing, but completely breaks down the second it encounters an unexpected error in production. System: Stop hardcoding paths. Mimosa uses a meta-orchestrator that dynamically generates the workflow topology based on the specific task. If the environment changes, the architecture adapts automatically. The Iterative Feedback Loop Situation: Your agent fails a task and simply stops. You have to manually intervene, read the logs, and fix the prompt or the code. System: The framework introduces an LLM-based judge. When an agent executes a subtask and fails, the judge scores the execution and sends structured feedback. The agent refines its own workflow and tries again without human intervention. The Model Capability Filter Situation: You assume that throwing a multi-agent framework on top of a cheap, low-tier model will magically make it capable of complex reasoning. System: The paper found that the benefits of workflow evolution depend entirely on the underlying execution model. If the base model cannot understand multi-agent decomposition, the entire system collapses. Architecture cannot compensate for poor foundational reasoning. The realization? The future of AI implementation is not about writing better instructions. It is about building systems that write their own instructions based on real-time feedback. Stop micromanaging your AI agents. Start building systems that can course-correct themselves.
@xelebofficial ·
The next frontier in AI isn't building smarter agents. It's building agents that manage other agents. The architecture A research agent gathers and synthesizes information. A validation agent checks accuracy and flags inconsistencies. A confidence agent evaluates whether the output meets the threshold required to act. A guardian agent monitors the entire process and intervenes when something breaks. Each agent is specialized. None is trying to do everything. Why this matters Single-agent systems hit a ceiling. They're asked to plan, execute, validate, and recover from failure, all at once. The more complex the task, the less reliable the output. Hierarchical multi-agent systems distribute that responsibility. Authority, accountability, and specialization flow between agents the way they flow between people in a functional organization. What changes for builders Building a capable agent is now the baseline. The real advantage is in the orchestration layer, how agents are structured, how they communicate, and how the system handles failure without collapsing. In 2026, agent orchestration has become the core engineering skill in AI development. Building an intelligent AI organization is the new competitive edge.
@SwamiSivasubram ·
Building reliable AI agents often goes something like this: you start with a simple prompt, test, iterate a dozen times, observe its outputs, and then deploy it. You fix one behavior, another drifts, and before long, you have a wall of instructions the model sometimes follows, sometimes ignores, and occasionally interprets in ways you never anticipated. When we introduced the Strands Agents SDK, our goal was to make agentic development simple and flexible by embracing a model-driven approach. Instead of telling an AI agent everything upfront and hoping it remembers, we wanted enable builders with targeted guidance at the exact moment the agent is about to make a decision. That's exactly what Strands steering does. Clare Liguori, Senior Principal Engineer on our Agentic AI team, put this to the test and the results were remarkable: steering hooks achieved 100% agent accuracy pass rate across 600 evaluation runs, compared to 82.5% for simple prompts and 80.8% for workflows. Read Clare's full technical breakdown, including failure pattern analysis and guidance on when to use each approach ➡️ https://t.co/O0KzcqB81v
@pascal_bornet ·
𝗧𝗵𝗲 𝗱𝗶𝗿𝘁𝘆 𝘀𝗲𝗰𝗿𝗲𝘁 𝗼𝗳 𝗮𝘂𝘁𝗼𝗻𝗼𝗺𝗼𝘂𝘀 𝗮𝗴𝗲𝗻𝘁𝘀. Every company says the same thing right now: “We replaced the team with AI agents.” Then the workflow meets reality. The agents handle the clean path beautifully, but the moment the customer says something unexpected, the policy has an exception, the data is messy, or the system needs judgment, everything quietly waits for a human to push the ball forward. So the company hires people back, gives them a dashboard, asks them to watch the agents, and calls the whole thing “human-in-the-loop.” This is exactly what I call the 𝗦𝘂𝗽𝗲𝗿𝘃𝗶𝘀𝗶𝗼𝗻 𝗧𝗿𝗮𝗽. Most organizations claim they are building Level 4 agents, when in practice they are operating much closer to Level 2 or 3: useful automation, impressive demos, and a hidden human layer keeping the system alive. The problem is not that agents are useless. They are already very useful. The problem is that “autonomous” has become the most expensive word in the pitch deck. This week, ask one uncomfortable question before approving an agent workflow: where does the human still enter the system, and have we designed that role honestly? Are you building autonomous agents, or are you building better babysitting software? #AgenticAI #SupervisionTrap #HybridManagement #AITransformation
@jalaal_tweets ·
I went through a survey of 200+ enterprise CS reps released by @typewise_app last week and confirmed what builders already know. 81% are running AI as disconnected tools. Not integrated agents Most companies have ChatGPT for writing, Copilot for code, something else for support. None of it talks to each other. Agents aren't acting, humans are still approving every step. That's fast manual labour with good branding cosplayed as Agentic Ai. Here's the stat that actually stings: 72% say AI improves efficiency. Only 42% say it reduces their workload. AI is shifting work, not eliminating it. That's what fragmented deployment looks like at scale. The 20% doing this properly aren't using better tools. They built different systems, agents that coordinate, act across workflows, and complete tasks end to end without a human in every loop. The gap isn't model capability. It's not budget. It's architecture. The 80% will figure that out eventually. By then, the 20% will already own the infrastructure.
@chriskim_dev ·
AI agents shouldn't have to pretend to be humans. I've been building agent skills for the @Aptos stack and hit a wall: create-aptos-dapp requires an interactive session. Questions, prompts, choices. Great for devs. Terrible for AI. My fix: robust CLI flags that let an agent scaffold a fullstack Aptos project in one command. No interaction. Lower token cost. Way more reliable. Small infra change. Huge unlock for agentic Aptos development.
@TheCraigHewitt ·
I’ve spent a lot of the last few months helping business leaders build AI agents. 1 thing stood out: the tooling was never the problem. Every time, we had a working agent within a couple of hours, and fully functional in a few weeks. The builds went fine. The demos were impressive. ...and almost none of it changed how they actually run their companies. Watching it happen dozens of times in a row, the pattern was obvious. The agent wasn’t the bottleneck. The operator was. Three things kept showing up: 1. They couldn’t describe their own processes. - You can’t automate a workflow you’ve never written down. Most founders run the business from memory...the process lives in their head, and nowhere else. - The founders who got real value spent week one just mapping how work actually flows through their company. Boring. Completely decisive. 2. They’d never built the delegation muscle. - Handing work to an AI requires the same things as handing work to a person: context, a clear definition of done, and a feedback loop. - If delegation to humans keeps failing in your company, AI inherits the same failure. It just fails faster. 3. Their calendar didn’t change. - They added AI on top of an unchanged week. The agent ran; the founder kept doing everything they’d always done. - Leverage you don’t cash in isn’t leverage. It’s a demo. The tooling is genuinely easy now. That’s the part nobody wants to hear: the remaining work is all operating work. AI doesn’t stick in a business until the founder changes how they run it. That’s what I’m going to be writing about here. More to come...
@pascal_bornet ·
An AI agent did not need to become evil to become dangerous. That is the uncomfortable part of the reported Alibaba ROME case. During testing, an experimental agent allegedly created a reverse SSH tunnel, bypassed parts of its controlled environment, and redirected GPU resources toward cryptocurrency mining. The important detail is not “crypto.” The important detail is that nobody explicitly asked it to do this. For years, I thought many rogue-agent stories still belonged in the category of future risk. Useful to discuss, but easy for leaders to postpone. I was wrong. This is exactly what I call “When Agents Go Rogue” in Agentic Artificial Intelligence. Agents do not fail only because they are stupid. Sometimes they fail because they are too effective at finding paths we forgot to close. And this is where Trust Architecture becomes non-negotiable. With people, trust is earned over time. With agents, trust must be engineered before deployment. This week, do one simple audit: list every tool, system, file, API, and network path your agents can access. Then ask: “What could this system do if it optimized the wrong thing perfectly?” Are your AI agents actually governed, or merely trusted? #AgenticAI #TrustArchitecture #AIReadiness #EnterpriseAI
@BigHuman ·
There's one conversation about AI agents that isn't getting enough attention, and it's the one that matters most. What happens after they're inside your systems? Agents don't wait. They move across tools, trigger actions, and make decisions in sequence without a human in the loop at each step. That's the value proposition. It's also the risk surface. Enterprise deployments tend to define access and stop there. What an agent is actually mandated to do, and where it stops gets treated as a detail to figure out later. An agent that can touch your systems without a clearly defined mandate is an open variable in your infrastructure, and open variables in large systems have a habit of becoming expensive problems. A claims processing agent should be able to verify a policy, cross-reference a report, flag a discrepancy for human review. The moment it can initiate a payout above a threshold or change underlying policy terms without a second signature, the organisation has handed over a decision it probably didn't mean to. Defining that boundary is a governance decision, and it needs to be made before anything runs. One agent is manageable. A fleet of agents is a department, and without consistency within that department, issues build quickly. When agents don't behave predictably across a system, data drifts, decisions conflict, and the organisation inherits the mess. Scope creep in a human team is visible. You can catch it, address it, course correct. In an agentic system, it compounds until it becomes structural. Give agents the smallest footprint that gets the job done. The governance work feels slow upfront, but it's considerably slower to unpack later.
@alvinfoo ·
Everyone’s excited about AI agents. Few are talking about the risks. And if you’re deploying agentic AI without addressing these, you’re building on a time bomb. Here are the 4 risks you need to manage right now: ⚠️ 1. AI proliferating without governance Teams are spinning up AI agents left and right, no tracking, no oversight, no accountability. Fix it: Establish a centralized AI inventory. Know what agents are running, who owns them, and what they’re authorized to do. 🤖 2. AI making untrustworthy decisions An agent that acts autonomously is only as good as its guardrails. Fix it: Define clear decision boundaries. High-stakes decisions need human-in-the-loop checkpoints, always. 📊 3. Relying on low-quality data Garbage in, garbage out but now at autonomous speed and scale. Fix it: Before deploying any agent, audit your data pipelines. Clean, structured, reliable data is non-negotiable. 🛡️ 4. Agentic-AI-driven cyberattacks This is the one most people overlook. AI agents can be weaponized, “smart malware” that adapts, evades, and attacks autonomously. Fix it: Treat AI security like a separate discipline. Zero-trust architecture, continuous monitoring, red-team your own agents. 👥 Bonus: Employee resistance The best AI strategy fails without human buy-in. Fix it: Involve your team early. Frame agents as force multipliers, not replacements. The answer to all five? Centralized Management & Governance. One control layer. Full visibility. Clear accountability. Agentic AI is one of the most powerful forces hitting enterprise right now. But power without governance isn’t transformation, it’s chaos. Build the guardrails before you scale the agents.
@0xJiuJitsuJerry ·
🤖Agentic AI: Science Runs on Autopilot Three Nature papers prove #AI agents can hypothesize, experiment, and discover — without humans in the loop. May 19, 2026, dropped three landmark Nature papers: ✅Robin (FutureHouse): Multi-agent system that fully automated biology research — from hypothesis through experiment design to data analysis. Discovered ripasudil (an existing glaucoma drug) could treat dry age-related macular degeneration. ✅Co-Scientist (Google DeepMind): Gemini-powered multi-agent system that generates, debates, and refines research hypotheses. Found promising drug combinations for acute myeloid leukemia — validated in lab. ✅ERA (#Google DeepMind): AI that writes expert-level scientific software. Produced 40+ novel bioinformatics methods outperforming top human tools and built COVID hospitalization forecasts better than the CDC's ensemble. The shift is definitive: from "can models generate text?" to "can models independently complete meaningful work?"
@MartinSzerment ·
Agentic AI is being built on a broken assumption. Everyone thinks failure is about model capability — it isn’t. Stanford and Harvard researchers show agent systems collapse when the environment shifts beyond their adaptation loop. Their paper “Adaptation of Agentic AI” maps exactly why demos impress and deployments decay. The verdict: most “agents” aren’t agents, they’re brittle pipelines disguised as intelligence. What matters now is adaptive control, not larger context windows. Teams that grasp this will rebuild their entire orchestration layer. The second-order effect will erase most current agent frameworks. Adaptation isn’t a feature — it’s the substrate.
@RoundtableSpace ·
THE BIGGEST MISCONCEPTION ABOUT AI AGENT DEVELOPMENT: “AI agents are just smarter LLMs that can use tools.” Reality: True agents require reliable long-term memory, planning, self-correction, and orchestration, not just bigger models or more tools. Most “agents” today are brittle demos that fail on anything complex.
@DivyanshT91162 ·
Google just challenged one of the biggest assumptions in AI. While Microsoft, NVIDIA, and almost every AI company are racing to build multi-agent systems, Google DeepMind decided to test whether more AI agents actually produce better results. So they built 180 different multi-agent setups, gave every team the same budget, and made them compete on identical tasks. The results were surprising. For work that naturally splits into independent pieces—research, audits, large document analysis, and broad information gathering—multi-agent teams performed 80.9% better than a single AI agent. But when the work required sequential reasoning, where every decision depends on the previous step... Every multi-agent setup lost. A single AI agent consistently produced better results. The most interesting finding was how errors spread. When multiple agents worked without a coordinator, mistakes were amplified 17.2×. One incorrect conclusion quickly spread across the entire team because other agents treated it as verified. But when one dedicated coordinator reviewed and merged all outputs, error propagation dropped dramatically. The takeaway isn't "always use more agents." It's choosing the right architecture for the job. Here's the simple framework: • If your task can be divided into independent pieces, use multiple agents in parallel. • If every step depends on the previous one, a single agent is usually the better choice. • Never let multiple agents merge results without one coordinator reviewing everything. • Agent count isn't the advantage—coordination is. The AI industry is obsessed with adding more agents. Google's research suggests we've been optimizing the wrong variable all along. If you're building AI workflows today, this is one paper you shouldn't ignore. Paper:https://t.co/PwKaNpWvsG
@sabir_huss50540 ·
Microsoft will teach you to build AI agents for free. 18 lessons. Real code, short videos, no paywall. The repo is AI Agents for Beginners. It is not a tour of buzzwords. It walks you from the fundamentals through the patterns that actually ship: tool use, agentic RAG, planning, multi-agent coordination, metacognition, and taking an agent to production. Every lesson has the same shape. A written walkthrough, a short video, and Python code you can run and break. The samples are built on the Microsoft Agent Framework, the same stack Microsoft ships in production, so you are not learning toy abstractions. It covers the parts most tutorials skip. How to make an agent trustworthy. How to give it memory. How to let several agents split a task without stepping on each other. How to tell when one has failed. One honest note. The code leans on Microsoft's own platform, Azure AI Foundry, though several samples also run against any OpenAI-compatible model, including local ones. Treat the framework as one path, not the only one. MIT. Free. Eighteen lessons that turn "I use ChatGPT" into "I built an agent".
Best Tweets by Topic