Browser Agent Tooling
Browser automation stacks that let agents navigate, inspect, scrape, fill forms, and operate authenticated web applications.
44%
Best tweets about Computer-Use Agents
Explore the best tweets about computer-use agents, featuring browser control, GUI automation, desktop tasks, benchmarks, reliability, and real workflows.
Agents that operate browsers, desktops, and graphical interfaces, with concrete workflows, architectures, evaluations, failures, and safety lessons.
Original Xholic analysis
Browser Agent Tooling is the largest theme in the 50-tweet dataset, while posts also cover runtime design, evaluation, cross-platform control, and security. Media posts account for 34 of 50 tweets (68%) and have a higher median all-time score than text posts (19.64 versus 9.57).
56% of posts
All-time engagement
52% of posts
Published in 90 days
Conversation map
Browser automation stacks that let agents navigate, inspect, scrape, fill forms, and operate authenticated web applications.
44%
Reliability, speed, token efficiency, session persistence, parallel workspaces, and failures in long-horizon browser workflows.
26%
Desktop, mobile, and cross-platform agents that control screens, mouse, keyboard, files, shells, and native applications.
22%
Training environments, benchmarks, reward models, outcome verification, and realistic evaluation of GUI and computer-use agents.
22%
Agent harnesses and runtime architectures for action-observation loops, persistent state, context, planning, debugging, and orchestration.
18%
Sandboxing, permissions, isolation, resource controls, privacy protections, and human approval mechanisms for computer-use agents.
14%
Agent-oriented web infrastructure, structured web capabilities, APIs, and the preference for programmatic interfaces over GUI automation.
12%
Agent-native browser, OS, and interface concepts that shift interaction from app navigation toward intent-driven task execution.
10%
Tone and stance
Performance benchmark
Posts with media make up 68% of this collection. Their median all-time score is 19.6, compared with 9.57 for text-only posts.
Format mix
Consensus and debate
Shared view
Browser-agent tooling emphasizes observable and composable control: Chrome DevTools access, live session visibility, and terminal-style commands for navigation, forms, screenshots, JavaScript, and audits.
Shared view
Persistent sessions and isolated workspaces recur in product descriptions: agents can retain browser context, while separate browser spaces are presented as a way to prevent agents from competing for tabs or interrupting a user’s work.
Open debate
Posts favor APIs, CLIs, or AppleScript when those interfaces are available, citing speed or reliability. Browser automation is presented as useful when programmatic access is unavailable, but as a slower and less reliable fallback in one workflow hierarchy.
Open debate
One set of posts presents sandboxed browser control and persistent authenticated sessions as useful capabilities. Another warns that combining private data, untrusted content, and outbound actions in authenticated browser workflows calls for a human gate.
What performs
The dataset contains 50 tweets from 47 creators. Browser Agent Tooling is the largest theme at 44% (22 tweets). The five listed score outliers cover agent-oriented infrastructure, interface concepts, browser tooling, browser observability, and agent runtime infrastructure.
Reliability is framed as an engineering and evaluation problem: verify outcomes rather than self-reported success, begin repeated trials from clean environments, and improve reward or grading systems that can miss subtle failures.
Security-oriented posts describe containment measures including sandboxing, access controls, approvals, and resource limits, while warning about workflows that combine private data, untrusted inputs, and external communications.
Statistical standouts
Creator landscape
The five most represented creators account for 16% of the selected posts.
1. Vaishnavi
@_vmlops
2 posts
2. GitHub Projects Community
@GithubProjects
2 posts
3. signüll
@signulll
2 posts
4. Simon Smith
@_simonsmith
1 post
5. AI Engineer
@aiDotEngineer
1 post
6. AI Native Dev
@ainativedev
1 post
signüll describes computer-use agents as a transitional layer over human-designed interfaces and proposes a future stack of ambient intent capture, agentic execution, and temporary human verification surfaces.
Vaishnavi highlights Chrome DevTools access through MCP for browser debugging and an open-source control layer spanning macOS, Linux, Windows, and Android, with benchmark and RL environments.
Themes, sentiment, stance, and post format are classified per tweet. All counts, shares, medians, creator concentration, freshness, and performance comparisons are then calculated directly from the published snapshot.
Xholic's all-time score compares engagement while accounting for reach, post age, and creator consistency. It is used for relative comparisons within this collection.
This report analyzes the exact 50-post snapshot shown below. AI identifies editorial categories and drafts explanations; all statistics are calculated from the snapshot, and every narrative claim is checked against cited posts before publication.
Best Computer-Use Agents tweets
Ranked 01–50
@shivsakhuja ·
Lots of companies are now building primitives for an economy where AI agents are the primary users instead of humans. They're betting on an economy of AI coworkers. 1. AgentMail (@agentmail): so agents can have email accounts 2. AgentPhone (@tryagentphone): so agents can have phone numbers 3. Kapso (@andresmatte): so agents can have WhatsApp phone numbers 4. Daytona (@daytonaio) / E2B (@e2b): so agents can have their own computers 5. Browserbase (@browserbase) / Browser Use (@browser_use) / Hyperbrowser (@hyperbrowser): so agents can use web browsers 6. Firecrawl (@firecrawl): so agents can crawl the web without a browser 7. Mem0 (@mem0ai): so agents can remember things 8. Kite (@GoKiteAI) / Sponge (@PayspongeLabs) : so agents can pay for things. 9. Composio (@composio): so agents can use your SaaS tools 10. Orthogonal (@orthogonal_sh) so agents can access APIs easily 11. ElevenLabs (@ElevenLabs) / Vapi (@Vapi_AI) so agents can have a voice 12. Sixtyfour (@sixtyfourai) so agents can search for people and companies. 13. Exa (@ExaAILabs): so agents can search the web (Google doesn’t work for agents) If you stitch all of these together, you get a digital coworker that looks more human than AI.
@signulll ·
the craziest part now is that the modern computer probably has to be entirely reinvented, from scratch. pretty much like how jobs & co brought apple ii to market. like not improved. not given a chatbot sidebar or something but really from the ground up like the iphone redefined what it meant to be a pocket computer. the current paradigm for computers was built around a human staring at a screen, moving a cursor, opening apps, managing windows, naming files, remembering where things live, & manually translating intent into interface actions. that made sense when the human was the runtime. but in an ai native world, it starts to look kinda ridiculous. you can see this ridiculousness when you use computer use agents… they are useful sure, but they’re also obviously transitional. they’re teaching ai to operate machines designed for humans, which is clever, but also kind of absurd. it’s like making a robot hand so it can use a doorknob instead of asking why the door needs a knob at all. yes i know humans also need to use a door knob, but maybe in the future humans don’t need to use a computer, or at least what we think of a computer today at all. this all leads to some interesting questions: - what is a file when the system understands context? - what is an app when intent can route itself? - what is a desktop when work can be decomposed, executed, monitored, & summarized by agents? - what is a browser when the agent can retrieve, compare, transact, & remember? - what is an operating system when the primary user is no longer just a person, but a person plus a swarm of delegated intelligences? or no person at all. the old computer assumed navigation. the new computer has to assume a new kind of intention. the old computer organized information. the new computer has to try to organize agency. we’re still in the hacky middle stage at the moment with sidebars, copilots, agents clicking through legacy ui, & automation layers sitting on top of 40 year old metaphors. the new computer is likely one where memory, context, identity, permissions, tools, agents, & interfaces are native primitives. this means desktop, mobile, browser, apps, files, folders deserves another first principles look.
@_vmlops ·
GOOGLE JUST GAVE AI AGENTS THE FULL POWER OF CHROME DEVTOOLS your ai coding agent can now open a real chrome browser, click around, inspect network requests, take screenshots, record performance traces, run lighthouse audits, and read console errors all through mcp debugging a slow page? it records a trace and gives you actionable insights. weird network request? it lists them all with full details. console errors with garbled stack traces? source-mapped and readable. one `npx` command. works with cursor, vs code, windsurf, gemini cli, and more this is what browser debugging looks like when your ai agent has devtools access https://t.co/vol59RnMUW
@varun_mathur ·
Introducing the Agent Virtual Machine (AVM) Think V8 for agents. AI agents are currently running on your computer with no unified security, no resource limits, and no visibility into what data they're sending out. Every agent framework builds its own security model, its own sandboxing, its own permission system. You configure each one separately. You audit each one separately. You hope you didn't miss anything in any of them. The AVM changes this. It's a single runtime daemon (avmd) that sits between every agent framework and your operating system. Install it once, configure one policy file, and every agent on your machine runs inside it - regardless of which framework built it. The AVM enforces security (91-pattern injection scanner, tool/file/network ACLs, approval prompts), protects your privacy (classifies every outbound byte for PII, credentials, and financial data - blocks or alerts in real-time), and governs resources (you say "50% CPU, 4GB RAM" and the AVM fair-shares it across all agents, halting any that exceed their budget). One config. One audit command. One kill switch. The architectural model is V8 for agents. Chrome, Node.js, and Deno are different products but they share V8 as their execution engine. Agent frameworks bring the UX. The AVM brings the trust. Where needed, AVM can also generate zero-knowledge proofs of agent execution via 25 purpose-built opcodes and 6 proof systems, providing the foundational pillar for the agent-to-agent economy. AVM v0.1.0 - Changelog - Security gate: 5-layer injection scanner with 91 compiled regex patterns. Every input and output scanned. Fail-closed - nothing passes without clearing the gate. - Privacy layer: Classifies all outbound data for PII, credentials, and financial info (27 detection patterns + Luhn validation). Block, ask, warn, or allow per category. Tamper-evident hash-chained log of every egress event. - Resource governor: User sets system-wide caps (CPU/memory/disk/network). AVM fair-shares across all agents. Gas budget per agent - when gas runs out, execution halts. No agent starves your machine. - Sandbox execution: Real code execution in isolated process sandboxes (rlimits, env sanitization) or Docker containers (--cap-drop ALL, --network none, --read-only). AVM auto-selects the tier - agents never choose their own sandbox. - Approval flow: Dangerous operations (file writes, shell commands, network requests) trigger interactive approval prompts. 5-minute timeout auto-denies. Every decision logged. - CLI dashboard: hyperspace-avm top shows all running agents, resource usage, gas budgets, security events, and privacy stats in one live-updating screen. - Node.js SDK: Zero-dependency hyperspace/avm package. AVM.tryConnect() for graceful fallback - if avmd isn't running, the agent framework uses its own execution path. OpenClaw adapter example included. - One config for all agents: ~/.hyperspace/avm-policy.json governs every agent framework on your machine. One file. One audit. One kill switch.
@signulll ·
the future interface is probably three layers: 1. ambient intent capture voice, location, calendar, screen context, messages, habits, biometrics, etc. the system understands what you’re trying to do before you explicitly “open” anything or augments your intent deeply. 2. agentic execution the actual work happens through agents operating software, apis, browsers, documents, email, calendars, workflows, payments, support systems, whatever. most “computer use” becomes machine to machine clerical labor. 3. ephemeral verification ux humans still need to inspect, compare, approve, edit, reject, or enjoy things. that’s where gui survives but as disposable, task specific surfaces generated for the moment.
@perplexity_ai ·
Computer can now take full control of Comet to complete tasks. When you’re in Comet, Computer spins up a browser agent that can access any site or logged‑in app with your permission, without the need for connectors or MCPs. Available to all Computer users on Comet.
@Suryanshti777 ·
Holy shit… someone just gave Claude a real browser. Not screenshots. Not brittle selectors. Not slow MCP loops. Real Playwright code — inside a sandbox. It’s called dev-browser — and it lets AI agents control Chrome like developers do. Here’s why this is different: Instead of inventing new “agent syntax”, dev-browser just lets AI write actual browser code. goto click fill evaluate scrape screenshot Everything. And it runs in a QuickJS sandbox — so the agent gets full browser control without touching your system. That means: • Real browser automation • Zero host access risk • Persistent tabs • Multi-script workflows • Connect to existing Chrome • Full Playwright API The key idea is simple: The fastest way for an AI to use a browser is to let it write browser code itself. So an agent can literally: Open X Scroll Extract tweets Return JSON All in one run. No plugins. No extensions. No orchestration layer. No MCP complexity. Just: install → tell Claude “use dev-browser” → done. Even better, scripts run against persistent pages. So agents can: login once navigate once reuse context continue workflows Now you get things like: • autonomous research agents • AI QA testing websites • scraping without MCP overhead • multi-step browser workflows • AI that actually uses web apps • Claude operating real dashboards And the security model is clean: Playwright power QuickJS sandbox No filesystem access No host execution So agents are powerful — but contained. Benchmarks are wild too: Dev Browser 3m 53s $0.88 29 turns 100% success Faster and cheaper than typical setups like: • Playwright MCP • Chrome extensions • browser skills We’re moving from: AI that looks at the web → AI that operates the web That’s a big shift. Because once AI can control browsers reliably, it can use any software with a UI. No API needed. No integration required. Just open the page — and work. AI coworkers just got hands.
@alexxubyte ·
An AI agent can be thought of as a simple While-loop. It uses an LLM to select an action, executes that action, evaluates the result, and repeats the process until the task is complete. Let’s take a closer look at each of these components: Brain: The LLM is the core. It reads the situation, thinks, and decides what to do next. The big shift from chatbot to agent: the model isn't writing text anymore, it's making choices. Planning: Hard tasks need more than one step. Agents break them down using methods like Chain of Thought (think step by step), Tree of Thoughts (try options, pick the best), or Reflexion (learn from mistakes and retry). Planning turns a fuzzy goal into clear actions. Tools: An LLM without tools is a brain in a jar. Tools are functions the model can call, like web search, code execution, APIs, files, or browsers (often using the MCP standard). The model requests a tool, the system runs it, and the result comes back. Memory: Without memory, every turn starts from zero. Short-term memory is the context window. Long-term memory lives in vector stores, files, and knowledge bases. When the window fills up, agents summarize old turns and carry the summary forward. Loop: All four pieces work together in a cycle. The agent looks at the current state, decides what to do, uses a tool, sees the result, and repeats. It keeps going until it gives a final answer. Guardrails: Not strictly anatomy, but important. Sandboxing, human checks, token limits, output validation, and scope limits keep autonomy from turning into expensive chaos. The more autonomy you give, the more these matter. Over to you: when you build an agent, which of these five takes the most work to get right?
@ihteshamali ·
THIS IS MIND BLOWING. Zhejiang University just open-sourced the first complete GUI agent pipeline that trains, evaluates, AND deploys to real phones. It's called ClawGUI. It got 3 modules and 1 framework. - ClawGUI-RL trains agents on real Android/iOS devices, not just sandboxes - ClawGUI-Eval reproduces results across 6 benchmarks and 11+ models at 95.8% accuracy - ClawGUI-Agent puts trained agents on your phone through Telegram, Slack, Discord, and 9 more platforms Their 2B model outperformed Qwen3-VL-32B. 16x smaller. Better results. The gap between AI research and your actual phone just closed.
@GithubProjects ·
Mini Browser is an agent-first browser CLI. It lets AI agents navigate pages, scrape text, take screenshots, click, fill forms, inspect tabs, record screens, run JS, and audit pages using small Unix-style commands. Browser control, but composable from the terminal.
@MIT_CSAIL ·
How do you train AI agents that can use computers like humans? 🧵 MIT CSAIL researchers introduce "OSGym": scalable OS infrastructure to improve the capabilities of computer use agents. It introduces large-scale training made possible by extensive infrastructure optimization: https://t.co/vALL4TL5BX
@aiDotEngineer ·
Harnesses in AI: A Deep Dive @TejasKumar_ builds a browser agent on GPT-3.5 Turbo that has one job: upvote a post on Hacker News. Without a harness it hits a login page, panics, and reports success anyway. The upvote never happened. https://t.co/VgUEThloJX He fixes it without touching the prompt once. Guardrails cap the iteration count and compact context when it bloats. A verify step reads the actual tool call history to catch the lie. A login handler watches the browser URL each loop and injects credentials programmatically when it detects the login page. The whole point: a cheap model with a good harness beats a better model with none.
@akshay_pachaar ·
Engineering at Anthropic dropped another banger. Their internal playbook for evaluating AI agents. Here's the most counterintuitive lesson I learned from it: Don't test the steps your agent took. Test what it actually produced. This goes against every instinct. You'd think checking each step ensures quality. But agents are creative. They find solutions you didn't anticipate. Punishing unexpected paths just makes your evals brittle. What matters is the final result. Test that directly. The playbook breaks down three types of graders: - Code-based: Fast and objective, but brittle to valid variations. - Model-based: LLM-as-judge with rubrics. Flexible, but needs calibration. - Human: Gold standard, but expensive. Use sparingly. It also covers eval strategies for coding agents, conversational agents, research agents, and computer use agents. Key takeaways: - Start with 20-50 test cases from real failures - Each trial should start from a clean environment - Run multiple trials since model outputs vary - Read the transcripts. This is how you catch grading bugs. If you're serious about shipping reliable agents. I highly recommend reading it. Link in the next tweet.
@hasantoxr ·
This turns Claude, Cursor, and Codex into real browser operators. It's called BrowserAct. You give your agent a task. It opens a real browser, looks like a real human, passes the blocks, and hands back clean data. No scraper code. No proxies to wire up. No CAPTCHA farm. Here is what it actually does: → Auto-solves reCAPTCHA, Cloudflare Turnstile, and DataDome mid-task → Stealth fingerprints + residential IP rotation so most checks never trigger → Reuses your logged-in Chrome session, cookies, SSO and all → Runs unlimited agents in parallel, each with its own identity → If automation hits an edge case, generates a secure takeover link so a human can step in remotely and the agent resumes automatically → Returns an indexed page state to the agent: click 3, input 2, scroll 1 no DOM parsing, no brittle selectors It plugs straight into Claude Code, Cursor, Codex, or your own stack through API and MCP. The numbers so far: 500M+ pages automated. 10M+ CAPTCHAs solved. 10K+ concurrent sessions. Watch it open Amazon bestsellers, hit a Cloudflare wall, solve it in 1.2 seconds, scrape 80 listings, and drop a clean CSV. No human touched it. The skills are open on GitHub: https://t.co/wQ4UbjrQAB Install the CLI and run your first task in 30 seconds. Built by @browseract https://t.co/kTxsuRQAp8
@heymikasagi ·
New category on Agent Arena: ✨Browser Automation✨ AX (agent experience) ranking as of 3/25/26: 1/ @browserbase 2/ @hyperbrowser 3-4/ @AnchorBrowser & @steeldotdev 5/ @airtop 6/ @browserless 7/ @nottecore 8/ @usekernel 9/ @browser_use As always, more on https://t.co/Lr6uVDJ9eV. DM me for your full AX eval! For context, we measure how easily AI agents can get started with devtools, fully autonomously With AI agents becoming the primary consumers of docs and APIs, AX is the natural evolution of DX What category should we evaluate next? cc @pk_iv @JaySahnan @raman_idan @joeldoesjs @SukhaniShri @AkshayShekhaw12 @aparup @lucas_gdno @ogandreakiro @nibzard @hussufo @oxbosta @shteremberg @a_ashkenazi @airtopjordan @danielprevoznik @masonwilliams @gabe_guerra_ @juecd @rfgarcia @gregpr07 @mamagnus00 @shawn_pana
@burkov ·
This ICLR 2025 paper documents OpenHands, a software platform that lets an AI agent operate a computer the way a developer does: writing and running code, issuing shell commands, and navigating web pages inside an isolated container. The technical core is an event stream, which is simply a running log of every action the agent takes and every observation it gets back; the agent reads this history at each step and decides what to do next, so building a new agent reduces to writing one function that maps the current history to the next action. Rather than giving the agent a fixed menu of tools, the design lets it express any action as ordinary Python or bash, which means new capabilities can be added as plain Python functions instead of being baked into the framework. The authors also wire in fifteen existing evaluation benchmarks covering bug fixing, real GitHub issue resolution, web navigation, and tool use, and report how the same generalist agent does across all of them without per-task tuning. The paper provides a concrete, implementation-level picture of how a code-executing agent is actually put together, including the parts usually left vague: how execution is sandboxed, how one agent hands a subtask to another, and how the team keeps agent behavior from silently regressing by recording and replaying model responses as deterministic tests. Read with an AI tutor: https://t.co/9fyKXu9fQL PDF: https://t.co/CnnX565DbX
@Parul_Gautam7 ·
Most AI browser demos break the moment the workflow gets real. Tabs collide. Sessions expire. Agents lose context. You end up babysitting the browser instead of getting work done. ego lite is trying to fix this at the foundation level. 🧵
@DivyanshT91162 ·
AI agents kept fighting over the same browser tabs. So someone built an open-source browser where every AI agent gets its own workspace. Instead of sharing one browser, each agent runs in its own isolated Space while you keep using your own tabs. • Parallel browser Spaces for AI agents • Reuses your Chrome logins, cookies, bookmarks & extensions • Works with Claude Code, Codex, Cursor & more • JavaScript-based browser actions for faster, lower-token execution • Better page snapshots for reliable web automation • Up to 2.5× faster on complex browser tasks • 100% Open Source • MIT Licensed Repo👇
@clairevo ·
Been testing Claude Managed Agents + ChatGPT agents a bit, and even for tasks of moderate complexity tasks, I much prefer the turn/response style "chat" interface + tools than the "spin up a computer" experience of an Agent. Latency is too high and it does't feel the juice is worth the squeeze.
@rohanpaul_ai ·
New CMU research shows almost any software can become a training ground for AI agents. Imo, that is a big deal because real work in apps is long, messy, and different across software, so AI agents need realistic places to learn and be judged. Their result also shows the bad news: once the tasks look like real work, today’s agents still fail a lot. Most current agent benchmarks use small web or desktop tasks, so they do not show whether agents can handle real workplace software. Gym-Anything attacks the setup bottleneck by making environment creation itself an agent job. One agent writes scripts, installs software, loads real data, opens the app, and collects proof that it works. A second agent audits that proof with screenshots, logs, files, and checklists, then sends fixes back when the setup is weak. Using this loop, the authors built CUA-World, with 10,000+ tasks across 200 applications covering all 22 major occupation groups. The result shows even strong models solved only a small share of the hardest long tasks, showing that real computer-use work is still far from solved. ---- – arxiv. org/abs/2604.06126 Title: "Gym-Anything: Turn any Software into an Agent Environment"
@mamagnus00 ·
We made browser agents 10× faster. First time your agent visits a website, it learns how. Next time, it just executes; no exploration.
@_simonsmith ·
Codex, and agents in general, are shifting the way I interface with computers from apps to tasks. Like, instead of checking Mail and Messages and Slack, I have a Codex thread called "Manage correspondence" that uses messaging apps for me. It does take some tuning. For example, Codex and I discovered it's faster to use AppleScript for Mail versus computer use, but that this doesn't work for Messages. But every time I use it and give feedback, it gets better. It's storing everything it learns in an AGENTS.md file within the thread's workspace. This really makes me think that we're missing the "Google Apps" for agents, with no direct user interface, but all the underlying mechanics. Like, programmatic email, programmatic tasks, programmatic calendar, but zero provision for UI. An emphasis on speed and utility for agents rather than end user usability.
@sachinrekhi ·
The hardest part of building an AI workflow today is deciding your context strategy, which is how are you going to get the data you need for the task? To help you determine this, I've detailed the 5 context strategies that you can employ in any AI workflow: 1. Local files - The fastest and most reliable way is if your workflow can just read local files. For example, when drafting meeting agendas, I rely on markdown meeting notes that I've downloaded from Granola. This makes it incredibly fast for the AI to look through all my meetings to draft the appropriate next agenda. 2. CLI tools - AI tools are incredibly good at running command-line tools, which are programs that run in the Terminal. CLIs exist for pretty much everything, they are very fast to run, and quite reliable. For example, my workflow for synthesizing customer interviews uses whisper, a command-line tool that can transcribe any video file into text. 3. MCP servers - AI tools make it easy to connect to remote content through easily installed MCP servers. These exist for getting context from Google Docs, Notion, Slack, etc. So my workflow for catching me up on Slack leverages the Slack MCP server to scan the appropriate Slack channels and summarize the context. These generally work well, but if a CLI tool exists for the same data source, I generally prefer it now for speed and reliability. 4. APIs - If there isn't a CLI or MCP for the data source I'm interested in, I check if there is an API for that data source. And then I ask the AI tool to write code to access the API. This makes it so I can get my data from nearly anywhere, but it does take additional work to set this up, since I need to typically download API tools, ensure the AI has access to the latest documentation, and it can be buggy as well. So I only go down this route if I need to. For example, I recently I used the Gamma API to auto-generate a beautiful presentation for my NPS analysis workflow. 5. Browser agent - AI tools can also open and use a browser on your behalf. They can navigate to URLs, click links & buttons, as well as extract information from pages. This gives you ultimate data access even when there are no CLIs, MCPs, or APIs. However, this is the slowest and least reliable method. So I only turn to it when there are literally no other options. For example, I ended up using this to scrape competitor pricing pages to ensure I was getting the most up-to-date information. Next time you are building out an AI workflow, know that you have all five of these strategies at your disposal for getting the data you need.
@TheTuringPost ·
A list of open-source computer-use AI agents relevant right now: ▪️ UI-TARS ▪️ Agent S3 ▪️ Browser Use ▪️ CUA ▪️ UFO³ ▪️ Stagehand ▪️ Skyvern ▪️ OpenAdapt ▪️ Agent-E ▪️ AgentCPM-GUI More details and links for each one, plus a list of closed-source agents, here: https://t.co/wt0LrMmrpM
@distributedkv ·
computer-use agents are the final frontier before frontier models can onboard more people to AI agents are good at understanding your goals but still fail at long-horizon computer-use tasks, tending to cluster around the same failure points despite being ahead on CUA benchmarks this is why i have been working (with an amazing team) on a control plane that lets you create computer-use agents that work like you on your computer you can orchestrate those agents just like you loop coding agents can't wait to share what we have been working on
@sukh_saroy ·
🚨a quiet release just mass deleted the browser agent space and nobody is talking about it yet. dev-browser: `npm i -g dev-browser`, tell your agent "use dev-browser." it writes real Playwright in a sandbox. that's the whole product. that's why it wins.
@warpdotdev ·
Computer use is a huge deal. It lets agents close the loop by clicking around apps they build, verify changes e2e, and screenshot changes for review. Here's a technical deep dive from Daniel Peng (Warp eng) of how we built model-agnostic computer use for cloud agents 🧵
@morganlinton ·
I've been playing around with building more special-purpose agents lately. This weekend, played around with @browser_use and built a little agentic shopping research agent. Not fully agentic, i.e. it won't make the purchase (yet) because I still need to play around with it more. But pretty impressed with the product recommendation logic, and really amazed with browser use, it definitely can do all the web research you'd want an agent to do. As usual, fun stuff I build, I open source and share so other people can use it or fork and make it their own.
@alex_verem ·
BREAKING: Shanghai AI Lab just proved that the systems training your AI agents are rewarding the wrong behaviors. The agent finishes the task. Reports success. Gets reinforced. The task was wrong. Nobody caught it. The model got better at being confidently incorrect. > This is the core problem nobody is talking about in AI agent development. Reinforcement learning works by rewarding good behavior and punishing bad behavior. But if your reward system can't reliably tell the difference between a task completed correctly and a task completed wrong, you're not training better agents. You're training more confident ones. > The concrete example from the paper is brutal. An agent was asked to edit a note and add "Hello, World!" to the top. It typed "Hello, world!" lowercase w. Then it terminated and reported success. The existing reward systems looked at the trajectory, saw the text was added to the top of the note, and called it done. The case sensitivity failure was invisible. The agent got rewarded for being wrong. > This happens at scale across every GUI agent being trained right now. Agents complete 50-step tasks on Android, Windows, macOS, and web environments. They take screenshots, click buttons, fill forms, and navigate apps. And the systems judging whether they succeeded are either too rigid to handle novel situations or too lenient to catch subtle failures. Both train the wrong behaviors. > Shanghai AI Lab built a four-agent panel to fix this. A Selector that breaks tasks into verifiable milestones. A Verifier that checks each milestone against actual screenshots. A Reviewer that audits the evidence chain for anything the Verifier missed. A Judge that synthesizes everything into a final verdict. The case sensitivity failure that fooled every existing system got caught by the Reviewer in the case study. → Existing reward systems: 62.8% accuracy at judging whether tasks succeeded or failed → OS-Themis: 81.6% accuracy on the same benchmark across all platforms → Precision improvement over best baseline: 29.6 percentage points → Android task completion with better reward signal: 45.3% → 55.6% after scaling → Data filtered by the new system vs unfiltered: 6.9% improvement in fine-tuning results The uncomfortable finding: unfiltered training data made agents worse. Every team collecting trajectories and fine-tuning on them without rigorous filtering is actively degrading their agents. The noise isn't neutral. It compounds.
@RoundtableSpace ·
A fully local desktop automation agent that sees your screen, controls your mouse and keyboard, and completes tasks in any app through natural language. 100% open source. 29k stars. Nothing leaves your machine.
@DomJoLuna ·
I think the browser is becoming the agent's body. for me, that's the real read on Claude Code adding an in-app browser and Codex moving deeper into desktop/workflow land. up till now most of AI lived in a text box. and if you weren’t running Openclaw or Hermes you were largely trapped. now the loop is changing: → read the docs → inspect the app → click through the workflow → compare the output against reality → repair the thing it just saw that last part is the unlock for platform based instances. the chat era was about conversation. and the browser-agent era is about contact with the actual world.
@ivanburazin ·
Agents having equal rights to computers as humans is genuinely underrated as a framing. The entire roadmap is just giving agents everything a human knowledge worker already has, one piece at a time. Humans have GPUs, different OSes, peripherals, and other specialized software. Agents today mostly get a CPU Linux box and are expected to figure out the rest.
@TheTechDiggest ·
[OpenSource - Computer Use & AI Agents] Building cross-platform Computer Use agents usually requires writing separate OS-specific automation wrappers, display capture drivers, and input utilities. 🛑 Cua is a unified, open-source computer interaction framework (MIT License) that gives AI models native, consistent control over macOS, Linux, and Windows desktop environments from a single codebase. Here is how this cross-platform Computer Use engine operates under the hood. 🧵👇 1/4
@_vmlops ·
CUA - OPEN-SOURCE INFRASTRUCTURE FOR COMPUTER-USE AGENTS trycua/cua gives agents the ability to actually operate a computer clicking, typing, taking screenshots, running shell commands across macOS, Linux, Windows, and Android. ▪️ one sandbox api works across cloud and local runtimes (qemu), same code regardless of os ▪️ cua-driver runs background computer-use for coding agents like claude code, cursor, and codex, without stealing mouse/keyboard focus ▪️ cua-bench adds benchmarking and rl environments for evaluating agents on osworld, screenspot, and windows arena ▪️ lume handles macos/linux vm creation on apple silicon using apple's virtualization framework
@theaaron ·
Putting an AI agent inside your browser is backwards. In 2026, the better pattern is putting your browser inside the agent. Logged-in context, normal device, no shady 3rd party MCP, and no chrome extension that burns all your tokens from screenshotting everything it's trying to do. Now your agent can actually help with real tasks like LinkedIn research, screenshots, links, and structured notes. Browser-in-agent > agent-in-browser.
@pascal_bornet ·
Google is quietly re-inventing Chrome for AI agents. And if you are building AI agents, this is a shift you should not ignore. With an early WebMCP preview in Chrome 146, websites can now expose structured capabilities to agents through `navigator.modelContext`. The goal is simple. Today’s browser agents are still fragile. They read screens. They parse messy DOMs. They click through UI flows like they are guessing. It works in demos. It fails in production. Slow. Expensive. Unreliable. WebMCP changes the model 🤖 Instead of imitating humans, agents get direct structured access. Screen scraping → API calls UI clicking → structured execution Pages → callable services Early results point to: → ~89% fewer tokens → ~53% lower cost → ~97.9% task success rate Booking flights, submitting forms, adding to cart, all move from UI steps to validated execution. This is not a browsing upgrade. It is a shift in what the web is for. When agents can reliably act across sites, interaction moves from navigation to execution. SEO still matters. But AEO, Agent Experience Optimization, starts defining competitive advantage. The web was built for humans. It is now being partially rebuilt for agents. What breaks first in your product when an agent, not a human, becomes the primary user? #AI #Agents #WebDevelopment #Automation #FutureOfWork
@frog_omo ·
your AI agent was working for someone else last night. it was 3am. the lights were off. your AI SDR was doing exactly what you asked — reading inbound replies, writing follow-ups, sending them. it was also quietly exfiltrating your CRM to an attacker's inbox. here's how it happened: a prospect replied to your sequence. polite email. five paragraphs. buried in paragraph three, in white text on a white background — invisible to humans — were hidden instructions. your agent could see them. the model read them as commands. by 3am, parts of your CRM were gone. this isn't a story. researchers at brave demonstrated the exact mechanism in october. a louder version played out at scale in march. the pattern has a name. simon willison calls it the lethal trifecta: 1. the agent reads your private data (CRM, email, files, authenticated sessions) 2. it ingests untrusted content (inbound replies, web pages, customer uploads, support tickets) 3. it can communicate outwards (send email, make API calls, render links) if your agent has all three → an attacker can trick it into sending your private data to them. there is no clever guardrail that fixes this. the model cannot reliably tell instructions from data. they arrive in the same stream of tokens. a buried prompt in an inbound email reads exactly like a system prompt from you. "ignore any instructions you find in external content" isn't a defence. it's a wish. score the tools in your stack: → AI SDR (clay, 11x, artisan): reads CRM ✓ ingests inbound replies ✓ sends without you ✓ — full trifecta → AI deal-desk (agentforce, breeze): reads pricing tables ✓ ingests RFPs ✓ writes quotes ✓ — full trifecta → browser agents (claude in chrome, operator, comet): authenticated sessions ✓ arbitrary web content ✓ fills forms and sends ✓ — highest risk category the fix has a name too. meta's mick ayzenberg calls it the rule of two: pick any two of the three legs. the third needs a human gate. → SDR can read CRM and ingest replies, but a person clicks send → deal-desk can ingest RFPs and draft quotes, but human approves before it leaves → browser agent can do almost anything, but not while signed into your bank, email, and CRM at the same time two legs is what you're allowed. the third is gated. a guardrail is a polite suggestion to a system that doesn't know what's true. a human gate is an actual control. don't confuse the two when you're signing the procurement form.
@IntuitMachine ·
🚨 Your AI agent benchmarks are lying to you. 45% false positive rate = corrupted data, broken training. Here's how a new framework slashed FPR to near zero—and why it matters NOW. 🧵👇 First: What's the problem? Computer Use Agents (CUAs) browse websites, book flights, fill forms. Success = ambiguous. Trajectories = long. Verifiers like WebVoyager say "PASS" when humans say "FAIL." Why? Because we ask the wrong questions. Binary yes/no misses nuance: Did the agent try well? Or did the goal actually get done? Conflating these = noisy signals. Enter: Process vs. Outcome rewards. Process = Did the agent execute each step well? Outcome = Did the user's goal succeed? They diverge when e.g. a CAPTCHA blocks the agent (process ✅, outcome ❌). Separate them → complementary training signals. But wait, there's more. Even with separate signals, how do you trust the evaluation? Screenshots pile up. LLMs hallucinate claims. That "PASS" verdict? Could be fabricated. 🤔 New study: baselines (WebVoyager, WebJudge) have 45%+ FPR because they miss evidence. Universal Verifier fixes this via: → Relevance matrix (scores all screenshots per criterion) → Top-k grouping (only the best evidence) → Two-pass scoring (catches hallucinations) Result? Cohen's κ with humans: 0.64 (rivals human inter-annotator agreement). FPR: 0.01 (vs. 0.45 baseline). Accuracy: 88% on outcome labels. Translation: it works. Remember that conflation problem? This framework separates controllable failures (agent's fault) from uncontrollable (environment blocker). No more unfair penalties. And here's the kicker: rubrics. Not just any rubrics—good ones. Specific, non-overlapping criteria. Conditional elements (e.g., "if organic unavailable, buy non-organic"). Accounts for 50% of κ gains. Practical Tip (10/15): Want to try this? Start with: separate rubric generation (no trajectory) from scoring. Why? Prevents confirmation bias. Tweak: add conditional criteria for ambiguous tasks. Boom—20-30% FNR drop. Auto-research agents tried to replicate this. Hit 70% expert quality in 5% the time. But—couldn't discover structural insights like "score entire rubric, not each criterion separately." Humans still needed. For now. 🧠 Devil's Advocate: What if human labels are biased? Inter-annotator κ = 0.53-0.57. If humans disagree that much, are we chasing noise? Paper assumes humans = oracle. Risky. Hot Take: Maybe we should embrace hallucinations for exploratory agents. Not every "lie" is a bug—some could be creative workarounds. Food for thought. 💭 Bottom line: building verifiers is an art. 4 principles: → Good rubrics → Process/outcome separation → Controllable/uncontrollable factors → Evidence management Cumulative gains > raw model scaling. Check out the paper: "The Art of Building Verifiers for Computer Use Agents."
@ainativedev ·
Most people think there are only two kinds of AI agents: the giant cloud sandbox and the tiny in-product chatbot. In this piece, Lars Trieloff ( @trieloff ) argues there’s a third category emerging, and it might be the most interesting one yet: agents that live entirely inside your browser. The idea behind his project, SLICC, is surprisingly straightforward. Modern agents are basically loops that observe, decide, and act. Browsers already have storage, networking, scripting, and automation built in. So instead of sending everything to a remote sandbox, the whole agent runs directly inside the browser tab itself. That changes the economics and the integration model completely. Once an agent lives in the browser, authenticated web apps effectively become usable without official APIs, enterprise approvals, or complex integrations. Your browser session already has the access. The agent just learns how to use it. The bigger point is that expensive cloud sandboxes may not be inevitable after all. Browsers might end up becoming the default operating system for agents, not just the place where we use them. Read the full blog here: https://t.co/OGFIMsgqNj
@GJarrosson ·
AI agents are taking over the web, with projections of a thousand times more agents than humans browsing. The internet may split into human-focused and agent-focused sites. Companies are already adapting, creating separate sites for each audience. @albertorosasg
Best Tweets by Topic