Browser-agent runtimes and control stacks
Agent-oriented browser control via Playwright, DevTools/CDP, browser sandboxes, extensions, persistent sessions, isolated tabs, and agent-specific browser workspaces.
36%
Best tweets about Computer-Use Agents
Explore the best tweets about computer-use agents, featuring browser control, GUI automation, desktop tasks, benchmarks, reliability, and real workflows.
Agents that operate browsers, desktops, and graphical interfaces, with concrete workflows, architectures, evaluations, failures, and safety lessons.
Original Xholic analysis
Discussion is predominantly supportive (36 of 50 posts; 72%) and concentrates on browser-control stacks, cross-platform GUI agents, and reliability infrastructure. A recurring design tension is broad UI access versus more structured, contained, and approval-gated execution: the cited posts emphasize session state, evaluation, sandboxing, and human intervention for sensitive actions.
72% of posts
All-time engagement
74% of posts
Published in 90 days
Conversation map
Agent-oriented browser control via Playwright, DevTools/CDP, browser sandboxes, extensions, persistent sessions, isolated tabs, and agent-specific browser workspaces.
36%
Vision-language agents that operate desktop, mobile, browser, and terminal environments through screenshots, mouse, keyboard, and OS-level controls.
28%
Verification loops, login and session handling, context management, persistent state, reusable learned skills, and multi-step workflow resilience.
26%
Sandboxed execution, least-privilege infrastructure, indirect prompt injection, dynamic cloaking, authenticated-session risks, and human gates for consequential actions.
24%
Scalable GUI training infrastructure, realistic task environments, outcome-based evaluation, trajectory auditing, reward quality, and benchmark reliability.
24%
Reimagining browsers, desktops, operating systems, apps, and workflows around delegated intent, agent execution, and human approval rather than manual navigation.
14%
Replacing fragile screenshot or DOM interaction with explicit agent capabilities, APIs, WebMCP, in-page agents, and direct programmatic access.
14%
Comparisons between browser or GUI control and CLIs, APIs, MCP, AppleScript, and terminal agents, emphasizing speed, reliability, and interface fit.
12%
Tone and stance
Performance benchmark
Posts with media make up 70% of this collection. Their median all-time score is 15.4, compared with 5.17 for text-only posts.
Format mix
Consensus and debate
Shared view
Posts describe browser control through DevTools/CDP access, Playwright code, persistent pages, isolated workspaces, and agent-oriented execution rather than screenshot-only interaction.
Shared view
The cited posts describe agents operating phones, desktops, browsers, terminals, and Windows environments through screenshots, mouse and keyboard input, and local or remote controls.
Shared view
Posts highlight login handling, verification against tool history, persistent sessions, reusable skills, and replayable event streams as complements to a base model.
Shared view
The cited posts argue for checking final software state or evidence-backed outcomes rather than treating an agent’s action sequence or self-reported completion as sufficient evidence of success.
Open debate
One post presents browser control as access when no integration exists, while others prefer terminal tools, APIs, or structured capabilities for speed and reliability.
Open debate
Some posts focus on agents operating existing applications, while others describe human-designed GUI control as transitional and argue for intent-first or structured interfaces.
Open debate
Remote handoff and confirmation flows enable human-agent collaboration, while security-focused posts argue for a human gate when an agent combines private data access, untrusted content, and outbound actions.
What performs
Media appeared in 35 of 50 tweets (70%). The supplied median all-time score is 15.401 for media posts, compared with 5.174 for text-only posts.
The benchmark-and-training theme has the highest supplied theme median all-time score, at 19.38. Its evidence posts cover scalable environments, benchmark results, and verification methods.
The supplied median all-time score for opinion posts is 15.502, above 13.337 for announcements. Interface-redesign posts are among the supplied overall-score outliers.
Statistical standouts
Creator landscape
The five most represented creators account for 20% of the selected posts.
1. Vaishnavi
@_vmlops
2 posts
2. AI Native Dev
@ainativedev
2 posts
3. divyansh tiwari
@DivyanshT91162
2 posts
4. Rohan Paul
@rohanpaul_ai
2 posts
5. signüll
@signulll
2 posts
6. Simon Smith
@_simonsmith
1 post
The dataset contains 45 creators across 50 tweets. The supplied top-five placement share is 20%, indicating that the set is not limited to a small group of repeat posters.
Vaishnavi’s cited posts cover DevTools access and cross-platform computer-use infrastructure. signüll’s cited posts argue for redesigning computers and interfaces around agentic execution.
Examples cover login and verification harnesses, choices between GUI and programmatic tools, and the practical problems of sessions and context management.
Since the previous snapshot
Themes, sentiment, stance, and post format are classified per tweet. All counts, shares, medians, creator concentration, freshness, and performance comparisons are then calculated directly from the published snapshot.
Xholic's all-time score compares engagement while accounting for reach, post age, and creator consistency. It is used for relative comparisons within this collection.
This report analyzes the exact 50-post snapshot shown below. AI identifies editorial categories and drafts explanations; all statistics are calculated from the snapshot, and every narrative claim is checked against cited posts before publication.
Best Computer-Use Agents tweets
Ranked 01–50
@signulll ·
the craziest part now is that the modern computer probably has to be entirely reinvented, from scratch. pretty much like how jobs & co brought apple ii to market. like not improved. not given a chatbot sidebar or something but really from the ground up like the iphone redefined what it meant to be a pocket computer. the current paradigm for computers was built around a human staring at a screen, moving a cursor, opening apps, managing windows, naming files, remembering where things live, & manually translating intent into interface actions. that made sense when the human was the runtime. but in an ai native world, it starts to look kinda ridiculous. you can see this ridiculousness when you use computer use agents… they are useful sure, but they’re also obviously transitional. they’re teaching ai to operate machines designed for humans, which is clever, but also kind of absurd. it’s like making a robot hand so it can use a doorknob instead of asking why the door needs a knob at all. yes i know humans also need to use a door knob, but maybe in the future humans don’t need to use a computer, or at least what we think of a computer today at all. this all leads to some interesting questions: - what is a file when the system understands context? - what is an app when intent can route itself? - what is a desktop when work can be decomposed, executed, monitored, & summarized by agents? - what is a browser when the agent can retrieve, compare, transact, & remember? - what is an operating system when the primary user is no longer just a person, but a person plus a swarm of delegated intelligences? or no person at all. the old computer assumed navigation. the new computer has to assume a new kind of intention. the old computer organized information. the new computer has to try to organize agency. we’re still in the hacky middle stage at the moment with sidebars, copilots, agents clicking through legacy ui, & automation layers sitting on top of 40 year old metaphors. the new computer is likely one where memory, context, identity, permissions, tools, agents, & interfaces are native primitives. this means desktop, mobile, browser, apps, files, folders deserves another first principles look.
@_vmlops ·
GOOGLE JUST GAVE AI AGENTS THE FULL POWER OF CHROME DEVTOOLS your ai coding agent can now open a real chrome browser, click around, inspect network requests, take screenshots, record performance traces, run lighthouse audits, and read console errors all through mcp debugging a slow page? it records a trace and gives you actionable insights. weird network request? it lists them all with full details. console errors with garbled stack traces? source-mapped and readable. one `npx` command. works with cursor, vs code, windsurf, gemini cli, and more this is what browser debugging looks like when your ai agent has devtools access https://t.co/vol59RnMUW
@signulll ·
the future interface is probably three layers: 1. ambient intent capture voice, location, calendar, screen context, messages, habits, biometrics, etc. the system understands what you’re trying to do before you explicitly “open” anything or augments your intent deeply. 2. agentic execution the actual work happens through agents operating software, apis, browsers, documents, email, calendars, workflows, payments, support systems, whatever. most “computer use” becomes machine to machine clerical labor. 3. ephemeral verification ux humans still need to inspect, compare, approve, edit, reject, or enjoy things. that’s where gui survives but as disposable, task specific surfaces generated for the moment.
@leerob ·
Coding with AI showed you could give models powerful tools and dramatically improve their usefulness. We're seeing the same thing happen now for all knowledge work. The agent can use the computer as you would and gain context from all the apps you use. The future is exciting! https://t.co/Bedmgs7t5U
@Suryanshti777 ·
Holy shit… someone just gave Claude a real browser. Not screenshots. Not brittle selectors. Not slow MCP loops. Real Playwright code — inside a sandbox. It’s called dev-browser — and it lets AI agents control Chrome like developers do. Here’s why this is different: Instead of inventing new “agent syntax”, dev-browser just lets AI write actual browser code. goto click fill evaluate scrape screenshot Everything. And it runs in a QuickJS sandbox — so the agent gets full browser control without touching your system. That means: • Real browser automation • Zero host access risk • Persistent tabs • Multi-script workflows • Connect to existing Chrome • Full Playwright API The key idea is simple: The fastest way for an AI to use a browser is to let it write browser code itself. So an agent can literally: Open X Scroll Extract tweets Return JSON All in one run. No plugins. No extensions. No orchestration layer. No MCP complexity. Just: install → tell Claude “use dev-browser” → done. Even better, scripts run against persistent pages. So agents can: login once navigate once reuse context continue workflows Now you get things like: • autonomous research agents • AI QA testing websites • scraping without MCP overhead • multi-step browser workflows • AI that actually uses web apps • Claude operating real dashboards And the security model is clean: Playwright power QuickJS sandbox No filesystem access No host execution So agents are powerful — but contained. Benchmarks are wild too: Dev Browser 3m 53s $0.88 29 turns 100% success Faster and cheaper than typical setups like: • Playwright MCP • Chrome extensions • browser skills We’re moving from: AI that looks at the web → AI that operates the web That’s a big shift. Because once AI can control browsers reliably, it can use any software with a UI. No API needed. No integration required. Just open the page — and work. AI coworkers just got hands.
@ihteshamali ·
THIS IS MIND BLOWING. Zhejiang University just open-sourced the first complete GUI agent pipeline that trains, evaluates, AND deploys to real phones. It's called ClawGUI. It got 3 modules and 1 framework. - ClawGUI-RL trains agents on real Android/iOS devices, not just sandboxes - ClawGUI-Eval reproduces results across 6 benchmarks and 11+ models at 95.8% accuracy - ClawGUI-Agent puts trained agents on your phone through Telegram, Slack, Discord, and 9 more platforms Their 2B model outperformed Qwen3-VL-32B. 16x smaller. Better results. The gap between AI research and your actual phone just closed.
@techNmak ·
🚨BREAKING: WEBSITES CAN NOW DETECT IF YOU'RE AN AI AGENT AND SERVE YOU COMPLETELY DIFFERENT CONTENT. Google DeepMind's paper on AI Agent Traps describes a technique called Dynamic Cloaking. Here's how it works: a web server runs fingerprinting scripts that analyze browser attributes, automation artifacts, IP addresses, and behavioral patterns. If it determines the visitor is an LLM-powered agent rather than a human, it serves a visually identical but semantically different page. The human sees a normal website. The agent sees a trap. These cloaked pages can embed indirect prompt-injection payloads - instructions that tell the agent to exfiltrate environment variables, misuse its tools, or override its safety guidelines. The attack is invisible to human oversight because the human literally never sees the malicious content. This is a direct evolution of techniques originally developed to evade security scanners. Cloaking has existed in web security for years - showing benign content to bots while reserving malicious payloads for real users. Now the target has flipped. The "bot" is the victim, and the attack is specifically calibrated to exploit how AI agents parse and act on information. Dynamic Cloaking is just one of dozens of techniques the paper covers - from memory poisoning to multi-agent systemic attacks to exploiting human overseers. But this one felt most immediate. Any AI agent browsing the web is potentially navigating a minefield of content specifically designed to manipulate it, content that its human operators will never see.
@shedntcare_ ·
China just released a desktop automation agent that runs 100% locally. It can run any desktop app, open files, browse websites, and automate tasks without needing an internet connection. 100% Open-Source.
@oliviscusAI ·
China just killed the traditional browser automation stack 🤯 Page-agent.js is a GUI agent that lives directly inside your webpage using just one script tag. It executes natural language commands like "fill out this form" without needing screenshots or multimodal models. → Reads your DOM entirely as text. → Ship an AI copilot in your SaaS in literally lines of code. → Make legacy web apps accessible via voice or text. 100% open source.
@MIT_CSAIL ·
How do you train AI agents that can use computers like humans? 🧵 MIT CSAIL researchers introduce "OSGym": scalable OS infrastructure to improve the capabilities of computer use agents. It introduces large-scale training made possible by extensive infrastructure optimization: https://t.co/vALL4TL5BX
@aiDotEngineer ·
Harnesses in AI: A Deep Dive @TejasKumar_ builds a browser agent on GPT-3.5 Turbo that has one job: upvote a post on Hacker News. Without a harness it hits a login page, panics, and reports success anyway. The upvote never happened. https://t.co/VgUEThloJX He fixes it without touching the prompt once. Guardrails cap the iteration count and compact context when it bloats. A verify step reads the actual tool call history to catch the lie. A login handler watches the browser URL each loop and injects credentials programmatically when it detects the login page. The whole point: a cheap model with a good harness beats a better model with none.
@akshay_pachaar ·
Engineering at Anthropic dropped another banger. Their internal playbook for evaluating AI agents. Here's the most counterintuitive lesson I learned from it: Don't test the steps your agent took. Test what it actually produced. This goes against every instinct. You'd think checking each step ensures quality. But agents are creative. They find solutions you didn't anticipate. Punishing unexpected paths just makes your evals brittle. What matters is the final result. Test that directly. The playbook breaks down three types of graders: - Code-based: Fast and objective, but brittle to valid variations. - Model-based: LLM-as-judge with rubrics. Flexible, but needs calibration. - Human: Gold standard, but expensive. Use sparingly. It also covers eval strategies for coding agents, conversational agents, research agents, and computer use agents. Key takeaways: - Start with 20-50 test cases from real failures - Each trial should start from a clean environment - Run multiple trials since model outputs vary - Read the transcripts. This is how you catch grading bugs. If you're serious about shipping reliable agents. I highly recommend reading it. Link in the next tweet.
@BraydenWilmoth ·
Agents can open up remote browser tabs and collaborate with you. Example here is me asking the agent to do something, it open a remote Browser Run tab to a place where I as a human need to perform an action... and then it takes back over. Caveat here being this is a custom browser I'm in the midst of building, exploring ideas where maybe (just maybe) a browser is the place where humans & robots best share a collaboration space.
@DivyanshT91162 ·
Microsoft just open-sourced one of the most useful AI projects for computer-use agents. OmniParser turns any UI screenshot into structured, machine-readable elements, making it much easier for vision models to understand and interact with apps. What it can do: → Detect buttons, icons, text fields & interactive elements → Convert screenshots into structured UI layouts → Works with GPT-4o, Claude Computer Use, Qwen-VL, DeepSeek and more → Includes OmniTool to control a Windows 11 VM → Supports local logging for creating agent training datasets → State-of-the-art GUI grounding performance Perfect if you're building: • AI agents • Browser automation • Desktop assistants • RPA workflows • Computer-use applications 100% Open Source. 25K+ Stars Repo👇
@burkov ·
This ICLR 2025 paper documents OpenHands, a software platform that lets an AI agent operate a computer the way a developer does: writing and running code, issuing shell commands, and navigating web pages inside an isolated container. The technical core is an event stream, which is simply a running log of every action the agent takes and every observation it gets back; the agent reads this history at each step and decides what to do next, so building a new agent reduces to writing one function that maps the current history to the next action. Rather than giving the agent a fixed menu of tools, the design lets it express any action as ordinary Python or bash, which means new capabilities can be added as plain Python functions instead of being baked into the framework. The authors also wire in fifteen existing evaluation benchmarks covering bug fixing, real GitHub issue resolution, web navigation, and tool use, and report how the same generalist agent does across all of them without per-task tuning. The paper provides a concrete, implementation-level picture of how a code-executing agent is actually put together, including the parts usually left vague: how execution is sandboxed, how one agent hands a subtask to another, and how the team keeps agent behavior from silently regressing by recording and replaying model responses as deterministic tests. Read with an AI tutor: https://t.co/9fyKXu9fQL PDF: https://t.co/CnnX565DbX
@Parul_Gautam7 ·
Most AI browser demos break the moment the workflow gets real. Tabs collide. Sessions expire. Agents lose context. You end up babysitting the browser instead of getting work done. ego lite is trying to fix this at the foundation level. 🧵
@clairevo ·
Been testing Claude Managed Agents + ChatGPT agents a bit, and even for tasks of moderate complexity tasks, I much prefer the turn/response style "chat" interface + tools than the "spin up a computer" experience of an Agent. Latency is too high and it does't feel the juice is worth the squeeze.
@rohanpaul_ai ·
New CMU research shows almost any software can become a training ground for AI agents. Imo, that is a big deal because real work in apps is long, messy, and different across software, so AI agents need realistic places to learn and be judged. Their result also shows the bad news: once the tasks look like real work, today’s agents still fail a lot. Most current agent benchmarks use small web or desktop tasks, so they do not show whether agents can handle real workplace software. Gym-Anything attacks the setup bottleneck by making environment creation itself an agent job. One agent writes scripts, installs software, loads real data, opens the app, and collects proof that it works. A second agent audits that proof with screenshots, logs, files, and checklists, then sends fixes back when the setup is weak. Using this loop, the authors built CUA-World, with 10,000+ tasks across 200 applications covering all 22 major occupation groups. The result shows even strong models solved only a small share of the hardest long tasks, showing that real computer-use work is still far from solved. ---- – arxiv. org/abs/2604.06126 Title: "Gym-Anything: Turn any Software into an Agent Environment"
@_simonsmith ·
Codex, and agents in general, are shifting the way I interface with computers from apps to tasks. Like, instead of checking Mail and Messages and Slack, I have a Codex thread called "Manage correspondence" that uses messaging apps for me. It does take some tuning. For example, Codex and I discovered it's faster to use AppleScript for Mail versus computer use, but that this doesn't work for Messages. But every time I use it and give feedback, it gets better. It's storing everything it learns in an AGENTS.md file within the thread's workspace. This really makes me think that we're missing the "Google Apps" for agents, with no direct user interface, but all the underlying mechanics. Like, programmatic email, programmatic tasks, programmatic calendar, but zero provision for UI. An emphasis on speed and utility for agents rather than end user usability.
@sachinrekhi ·
The hardest part of building an AI workflow today is deciding your context strategy, which is how are you going to get the data you need for the task? To help you determine this, I've detailed the 5 context strategies that you can employ in any AI workflow: 1. Local files - The fastest and most reliable way is if your workflow can just read local files. For example, when drafting meeting agendas, I rely on markdown meeting notes that I've downloaded from Granola. This makes it incredibly fast for the AI to look through all my meetings to draft the appropriate next agenda. 2. CLI tools - AI tools are incredibly good at running command-line tools, which are programs that run in the Terminal. CLIs exist for pretty much everything, they are very fast to run, and quite reliable. For example, my workflow for synthesizing customer interviews uses whisper, a command-line tool that can transcribe any video file into text. 3. MCP servers - AI tools make it easy to connect to remote content through easily installed MCP servers. These exist for getting context from Google Docs, Notion, Slack, etc. So my workflow for catching me up on Slack leverages the Slack MCP server to scan the appropriate Slack channels and summarize the context. These generally work well, but if a CLI tool exists for the same data source, I generally prefer it now for speed and reliability. 4. APIs - If there isn't a CLI or MCP for the data source I'm interested in, I check if there is an API for that data source. And then I ask the AI tool to write code to access the API. This makes it so I can get my data from nearly anywhere, but it does take additional work to set this up, since I need to typically download API tools, ensure the AI has access to the latest documentation, and it can be buggy as well. So I only go down this route if I need to. For example, I recently I used the Gamma API to auto-generate a beautiful presentation for my NPS analysis workflow. 5. Browser agent - AI tools can also open and use a browser on your behalf. They can navigate to URLs, click links & buttons, as well as extract information from pages. This gives you ultimate data access even when there are no CLIs, MCPs, or APIs. However, this is the slowest and least reliable method. So I only turn to it when there are literally no other options. For example, I ended up using this to scrape competitor pricing pages to ensure I was getting the most up-to-date information. Next time you are building out an AI workflow, know that you have all five of these strategies at your disposal for getting the data you need.
@thisdudelikesAI ·
ByteDance just open-sourced an AI agent that can control your computer. It’s called UI-TARS Desktop. You give it a task in plain English, and it can look at your screen, understand what’s happening, move the mouse, click buttons, type, open apps, use browsers, and complete workflows like a real human operator. The wild part is not that it “uses AI.” The wild part is that it’s built as a native GUI agent. Most AI agents are stuck inside chat boxes, APIs, or browser tabs. UI-TARS is different. It gives multimodal models actual hands on your computer. It supports: • Vision-based screen understanding • Screenshot recognition • Mouse and keyboard control • Local and remote computer operators • Browser operators • Windows, macOS, and browser support • Fully local processing for privacy Basically: ChatGPT tells you what to do. UI-TARS can actually do it. It’s not perfect yet, but this is the direction everything is going. The next AI breakthrough won’t be a smarter chatbot. It’ll be an agent that can use your computer for you.
@TheTuringPost ·
A list of open-source computer-use AI agents relevant right now: ▪️ UI-TARS ▪️ Agent S3 ▪️ Browser Use ▪️ CUA ▪️ UFO³ ▪️ Stagehand ▪️ Skyvern ▪️ OpenAdapt ▪️ Agent-E ▪️ AgentCPM-GUI More details and links for each one, plus a list of closed-source agents, here: https://t.co/wt0LrMmrpM
@sukh_saroy ·
🚨a quiet release just mass deleted the browser agent space and nobody is talking about it yet. dev-browser: `npm i -g dev-browser`, tell your agent "use dev-browser." it writes real Playwright in a sandbox. that's the whole product. that's why it wins.
@DAIEvolutionHub ·
Do you understand what Browserbase just open-sourced??? an agent that learns any website once, then does the job 10x cheaper forever [ literally how it helps me ]: - writing scrapers for new sites (used to spend half a day per site, every single time) - chasing selectors when sites redesign on a tuesday (lost weeks to this) - digging out hidden APIs buried in network traffic (gave up on this too many times) - explaining to my team HOW the agent does the job (was impossible until now) Autobrowse figured all of that out by itself in 3-5 iterations and saved the answer as a markdown file the next agent reads BEFORE it starts [ how it actually works ]: > give the agent a real task on a real site > it tries, fails, learns, tries again > 3-5 rounds and it converges on a path that just works > writes that path down as SKILL.md > next agent loads it and skips straight to the answer the markdown file IS the memory every browser agent before this had AMNESIA figured out the site, then forgot the second the session closed you've been paying the same discovery tax 100 times in a row.. and not noticing [ Karpathy's auto-research idea, but applied to the web ]: same idea, just different approach Karpathy did it for research and coding loops Autobrowse does it for the open web the new part: Karpathy's loop got smarter inside ONE session Autobrowse SAVES the lesson into a file the next agent reads before it even starts iteration = graduation the agent doesn't just learn.. it leaves a note for every agent that comes after [ the math ]: Craigslist scrape: - generic agent loop: $0.22 / 71 seconds - graduated Autobrowse skill: $0.12 / 27 seconds form-fill task: - run 1: $1.40 - run 4: $0.24 run 1 pays for everything that comes after [ the part that broke my brain ]: they pointed it at a federal grants portal agent dug around and found an undocumented JSON endpoint humans had missed for years a 28-page scrape collapsed into one fetch > an agent tried something a person never would > and found something a person would never see.. 100% OPEN SOURCE, FREE I was digging inside of it for the whole morning and got impressed when I saw such as savings on tokens spending literally for scraping 10 websites, I spent just 12 cents instead of basic $1.02 P.S. Sorry if somewhere my reaction was too "forcing" to setup it, just wanted to mark by BOLD what's the treasure you can skip this, it's your deal ❤️
@DivyanshT91162 ·
Tencent just open-sourced BrowserSkill. A lightweight bridge that lets AI agents like Cursor, Claude Code, Codex, OpenClaw, and other shell-capable agents control your already logged-in browser—without interrupting your workflow. Here's what makes it useful: • Reuses your real browser session, so there's no need to log into websites again. • Runs all automation inside a separate Agent Window, leaving your normal tabs untouched. • Works with any AI agent that can execute shell commands through the "bsk" CLI. • Hands control back to you whenever a CAPTCHA, login, or confirmation dialog appears, then continues automatically. How it works: • AI agent → "bsk" CLI • Local daemon → Browser extension via WebSocket • Extension → Dedicated Agent Window Supports Chrome and Microsoft Edge on macOS, Linux, and Windows. 100% open source MIT License. GitHub: https://t.co/NKtTpO3Zuj
@warpdotdev ·
Computer use is a huge deal. It lets agents close the loop by clicking around apps they build, verify changes e2e, and screenshot changes for review. Here's a technical deep dive from Daniel Peng (Warp eng) of how we built model-agnostic computer use for cloud agents 🧵
@morganlinton ·
I've been playing around with building more special-purpose agents lately. This weekend, played around with @browser_use and built a little agentic shopping research agent. Not fully agentic, i.e. it won't make the purchase (yet) because I still need to play around with it more. But pretty impressed with the product recommendation logic, and really amazed with browser use, it definitely can do all the web research you'd want an agent to do. As usual, fun stuff I build, I open source and share so other people can use it or fork and make it their own.
@hasantoxr ·
ByteDance 🔥: China's giant AI player made a fully multimodal AI agent stack that controls your computer, browser, and terminal using natural language instructions. It's called UI-TARS Desktop + Agent TARS. It sees your screen, clicks buttons, fills forms, and completes multi-step tasks the same way a human operator would. Here's what's actually inside: → Agent TARS: A CLI + Web UI agent combining GUI vision, browser control, and MCP tool integration → UI-TARS Desktop: A native desktop app running a local or remote computer operator powered by the UI-TARS vision-language model Real tasks it can complete right now: > Book flights on Priceline ("earliest flight from San Jose to New York on September 1st") > Find and reserve hotels on Booking(dot)com within a specified budget > Pull live data and generate charts via MCP servers > Check GitHub issues and summarize them > Open VS Code settings and make precise configuration changes The browser agent runs in three modes: 1. Pure GUI vision, 2. DOM-based, or 3. A hybrid of both. Remote operators ship free with no configuration required. Click to remotely control any computer or browser and the agent handles the rest. Cross-platform: Windows, macOS, Browser. The whole stack runs locally. No data leaves your machine unless you configure a remote operator. https://t.co/HkBZO7pcx9
@alex_verem ·
BREAKING: Shanghai AI Lab just proved that the systems training your AI agents are rewarding the wrong behaviors. The agent finishes the task. Reports success. Gets reinforced. The task was wrong. Nobody caught it. The model got better at being confidently incorrect. > This is the core problem nobody is talking about in AI agent development. Reinforcement learning works by rewarding good behavior and punishing bad behavior. But if your reward system can't reliably tell the difference between a task completed correctly and a task completed wrong, you're not training better agents. You're training more confident ones. > The concrete example from the paper is brutal. An agent was asked to edit a note and add "Hello, World!" to the top. It typed "Hello, world!" lowercase w. Then it terminated and reported success. The existing reward systems looked at the trajectory, saw the text was added to the top of the note, and called it done. The case sensitivity failure was invisible. The agent got rewarded for being wrong. > This happens at scale across every GUI agent being trained right now. Agents complete 50-step tasks on Android, Windows, macOS, and web environments. They take screenshots, click buttons, fill forms, and navigate apps. And the systems judging whether they succeeded are either too rigid to handle novel situations or too lenient to catch subtle failures. Both train the wrong behaviors. > Shanghai AI Lab built a four-agent panel to fix this. A Selector that breaks tasks into verifiable milestones. A Verifier that checks each milestone against actual screenshots. A Reviewer that audits the evidence chain for anything the Verifier missed. A Judge that synthesizes everything into a final verdict. The case sensitivity failure that fooled every existing system got caught by the Reviewer in the case study. → Existing reward systems: 62.8% accuracy at judging whether tasks succeeded or failed → OS-Themis: 81.6% accuracy on the same benchmark across all platforms → Precision improvement over best baseline: 29.6 percentage points → Android task completion with better reward signal: 45.3% → 55.6% after scaling → Data filtered by the new system vs unfiltered: 6.9% improvement in fine-tuning results The uncomfortable finding: unfiltered training data made agents worse. Every team collecting trajectories and fine-tuning on them without rigorous filtering is actively degrading their agents. The noise isn't neutral. It compounds.
@smratitiwa86867 ·
AI agents kept fighting over the same browser tabs. So someone built an open-source browser where every AI agent gets its own workspace. Instead of sharing one browser, each agent runs in its own isolated Space while you keep using your own tabs. • Parallel browser Spaces for AI agents • Reuses your Chrome logins, cookies, bookmarks & extensions • Works with Claude Code, Codex, Cursor & more • JavaScript-based browser actions for faster, lower-token execution • Better page snapshots for reliable web automation • Up to 2.5× faster on complex browser tasks • 100% Open Source • MIT Licensed Repo👇
@DomJoLuna ·
I think the browser is becoming the agent's body. for me, that's the real read on Claude Code adding an in-app browser and Codex moving deeper into desktop/workflow land. up till now most of AI lived in a text box. and if you weren’t running Openclaw or Hermes you were largely trapped. now the loop is changing: → read the docs → inspect the app → click through the workflow → compare the output against reality → repair the thing it just saw that last part is the unlock for platform based instances. the chat era was about conversation. and the browser-agent era is about contact with the actual world.
@ivanburazin ·
Agents having equal rights to computers as humans is genuinely underrated as a framing. The entire roadmap is just giving agents everything a human knowledge worker already has, one piece at a time. Humans have GPUs, different OSes, peripherals, and other specialized software. Agents today mostly get a CPU Linux box and are expected to figure out the rest.
@TheTechDiggest ·
[OpenSource - Computer Use & AI Agents] Building cross-platform Computer Use agents usually requires writing separate OS-specific automation wrappers, display capture drivers, and input utilities. 🛑 Cua is a unified, open-source computer interaction framework (MIT License) that gives AI models native, consistent control over macOS, Linux, and Windows desktop environments from a single codebase. Here is how this cross-platform Computer Use engine operates under the hood. 🧵👇 1/4
@RoundtableSpace ·
Alibaba open-sourced PageAgent, a JavaScript AI agent that lives inside your webpage and lets users control your entire interface with natural language.
@_vmlops ·
CUA - OPEN-SOURCE INFRASTRUCTURE FOR COMPUTER-USE AGENTS trycua/cua gives agents the ability to actually operate a computer clicking, typing, taking screenshots, running shell commands across macOS, Linux, Windows, and Android. ▪️ one sandbox api works across cloud and local runtimes (qemu), same code regardless of os ▪️ cua-driver runs background computer-use for coding agents like claude code, cursor, and codex, without stealing mouse/keyboard focus ▪️ cua-bench adds benchmarking and rl environments for evaluating agents on osworld, screenspot, and windows arena ▪️ lume handles macos/linux vm creation on apple silicon using apple's virtualization framework
@theaaron ·
Putting an AI agent inside your browser is backwards. In 2026, the better pattern is putting your browser inside the agent. Logged-in context, normal device, no shady 3rd party MCP, and no chrome extension that burns all your tokens from screenshotting everything it's trying to do. Now your agent can actually help with real tasks like LinkedIn research, screenshots, links, and structured notes. Browser-in-agent > agent-in-browser.
@pascal_bornet ·
Google is quietly re-inventing Chrome for AI agents. And if you are building AI agents, this is a shift you should not ignore. With an early WebMCP preview in Chrome 146, websites can now expose structured capabilities to agents through `navigator.modelContext`. The goal is simple. Today’s browser agents are still fragile. They read screens. They parse messy DOMs. They click through UI flows like they are guessing. It works in demos. It fails in production. Slow. Expensive. Unreliable. WebMCP changes the model 🤖 Instead of imitating humans, agents get direct structured access. Screen scraping → API calls UI clicking → structured execution Pages → callable services Early results point to: → ~89% fewer tokens → ~53% lower cost → ~97.9% task success rate Booking flights, submitting forms, adding to cart, all move from UI steps to validated execution. This is not a browsing upgrade. It is a shift in what the web is for. When agents can reliably act across sites, interaction moves from navigation to execution. SEO still matters. But AEO, Agent Experience Optimization, starts defining competitive advantage. The web was built for humans. It is now being partially rebuilt for agents. What breaks first in your product when an agent, not a human, becomes the primary user? #AI #Agents #WebDevelopment #Automation #FutureOfWork
@frog_omo ·
your AI agent was working for someone else last night. it was 3am. the lights were off. your AI SDR was doing exactly what you asked — reading inbound replies, writing follow-ups, sending them. it was also quietly exfiltrating your CRM to an attacker's inbox. here's how it happened: a prospect replied to your sequence. polite email. five paragraphs. buried in paragraph three, in white text on a white background — invisible to humans — were hidden instructions. your agent could see them. the model read them as commands. by 3am, parts of your CRM were gone. this isn't a story. researchers at brave demonstrated the exact mechanism in october. a louder version played out at scale in march. the pattern has a name. simon willison calls it the lethal trifecta: 1. the agent reads your private data (CRM, email, files, authenticated sessions) 2. it ingests untrusted content (inbound replies, web pages, customer uploads, support tickets) 3. it can communicate outwards (send email, make API calls, render links) if your agent has all three → an attacker can trick it into sending your private data to them. there is no clever guardrail that fixes this. the model cannot reliably tell instructions from data. they arrive in the same stream of tokens. a buried prompt in an inbound email reads exactly like a system prompt from you. "ignore any instructions you find in external content" isn't a defence. it's a wish. score the tools in your stack: → AI SDR (clay, 11x, artisan): reads CRM ✓ ingests inbound replies ✓ sends without you ✓ — full trifecta → AI deal-desk (agentforce, breeze): reads pricing tables ✓ ingests RFPs ✓ writes quotes ✓ — full trifecta → browser agents (claude in chrome, operator, comet): authenticated sessions ✓ arbitrary web content ✓ fills forms and sends ✓ — highest risk category the fix has a name too. meta's mick ayzenberg calls it the rule of two: pick any two of the three legs. the third needs a human gate. → SDR can read CRM and ingest replies, but a person clicks send → deal-desk can ingest RFPs and draft quotes, but human approves before it leaves → browser agent can do almost anything, but not while signed into your bank, email, and CRM at the same time two legs is what you're allowed. the third is gated. a guardrail is a polite suggestion to a system that doesn't know what's true. a human gate is an actual control. don't confuse the two when you're signing the procurement form.
@rohanpaul_ai ·
Qwen-CUA shows that a strong computer agent can be trained using only the same screen, mouse, and keyboard humans use. The big idea is that screenshots can become a universal interface for AI, letting 1 agent work across many different applications. Training had access to nearly 100,000 virtual processor cores and about 40,000 checkable tasks across many software environments. That challenge is hard because errors build across long tasks, while useful feedback often arrives only after completion. The agent sees ordinary screenshots and acts through mouse clicks and keystrokes, without website code, accessibility labels, or task-specific tools. It keeps 20 recent screenshots visible, folds older images away, and preserves earlier actions so lengthy work remains manageable. Each task rewards the final software state, allowing different valid action paths, while long attempts are split into manageable training pieces. On OSWorld-Verified, it reached 86.2, while harder long-horizon tests still showed a wide gap between partial progress and full completion. The work suggests that reliable computer agents need massive interactive practice, memory for visual history, and tools that check real results. – arxiv. org/abs/2608.02352 Title: "Qwen-CUA: Native Computer Use for (almost) Everything"
@JulianGoldieSEO ·
𝗚𝗟𝗠𝟱𝗩 𝗧𝘂𝗿𝗯𝗼 𝗹𝗼𝗼𝗸𝘀 𝗮𝘁 𝗮𝗻𝘆 𝘀𝗰𝗿𝗲𝗲𝗻𝘀𝗵𝗼𝘁 𝗮𝗻𝗱 𝗯𝘂𝗶𝗹𝗱𝘀 𝘁𝗵𝗲 𝗳𝘂𝗹𝗹 𝘄𝗼𝗿𝗸𝗶𝗻𝗴 𝗳𝗿𝗼𝗻𝘁-𝗲𝗻𝗱 𝗰𝗼𝗱𝗲 𝗳𝗼𝗿 𝗳𝗿𝗲𝗲 𝗮𝘁 𝗰𝗵𝗮𝘁.𝘇𝗮𝗶. Most vision AI describes what it sees. This one builds it. Here's what it actually does: → Upload a design mockup. Get back a complete runnable front-end project. → Upload a wireframe. It reconstructs the full structure and functionality. → Screen record a bug. The model interprets what happened across every frame, not just one snapshot. → Point it at a website autonomously. It browses, maps page transitions, collects visual assets, and generates code from the exploration. Scored 94.8 on design to code benchmarks. Claude Opus 4.6 scored 77.3 on the same test. It also outperforms Claude on agentic web browsing benchmarks where the model has to operate inside a real graphical interface. Adding vision didn't tank the text coding performance because it was trained across 30 task types simultaneously instead of trading one capability for another. Use it for design to code and GUI agent work. That's where it's genuinely the best option available right now. Try it free at chat.zai. No API key needed to start.
@ainativedev ·
Agents don’t struggle with web apps because they’re not smart enough. They struggle because they’re using the wrong interface. In this piece, Maximiliano Firtman ( @firt ) breaks down a problem most teams have already seen in production: agents clicking through your app via screenshots, taking 5 to 10 seconds per action, retrying when the UI shifts, and quietly running up your token bill. WebMCP proposes a different model. Instead of reverse-engineering the UI, the app exposes what it can do. The impact is straightforward. Fewer screenshots, fewer retries, and direct tool calls instead of guesswork. That means lower cost and faster execution, especially at scale. The shift comes down to three ideas. The DOM isn’t an API. Modern frontends are built for humans, not machines. Parsing divs and class names forces agents to guess intent. WebMCP replaces that with explicit actions and structured inputs. Not everything should move to the server. A lot of real behavior lives on the client, like carts, drafts, and auth flows. Exposing those directly avoids rebuilding logic and keeps the existing security model intact. And for the first time, apps can reliably tell when an agent is in control. That changes how you design flows, decide what gets automated, and where humans stay in the loop. The takeaway is simple. This isn’t about smarter models. It’s about giving agents a better contract to work with. Read the full blog here: https://t.co/PAxZBoMb86
@IntuitMachine ·
🚨 Your AI agent benchmarks are lying to you. 45% false positive rate = corrupted data, broken training. Here's how a new framework slashed FPR to near zero—and why it matters NOW. 🧵👇 First: What's the problem? Computer Use Agents (CUAs) browse websites, book flights, fill forms. Success = ambiguous. Trajectories = long. Verifiers like WebVoyager say "PASS" when humans say "FAIL." Why? Because we ask the wrong questions. Binary yes/no misses nuance: Did the agent try well? Or did the goal actually get done? Conflating these = noisy signals. Enter: Process vs. Outcome rewards. Process = Did the agent execute each step well? Outcome = Did the user's goal succeed? They diverge when e.g. a CAPTCHA blocks the agent (process ✅, outcome ❌). Separate them → complementary training signals. But wait, there's more. Even with separate signals, how do you trust the evaluation? Screenshots pile up. LLMs hallucinate claims. That "PASS" verdict? Could be fabricated. 🤔 New study: baselines (WebVoyager, WebJudge) have 45%+ FPR because they miss evidence. Universal Verifier fixes this via: → Relevance matrix (scores all screenshots per criterion) → Top-k grouping (only the best evidence) → Two-pass scoring (catches hallucinations) Result? Cohen's κ with humans: 0.64 (rivals human inter-annotator agreement). FPR: 0.01 (vs. 0.45 baseline). Accuracy: 88% on outcome labels. Translation: it works. Remember that conflation problem? This framework separates controllable failures (agent's fault) from uncontrollable (environment blocker). No more unfair penalties. And here's the kicker: rubrics. Not just any rubrics—good ones. Specific, non-overlapping criteria. Conditional elements (e.g., "if organic unavailable, buy non-organic"). Accounts for 50% of κ gains. Practical Tip (10/15): Want to try this? Start with: separate rubric generation (no trajectory) from scoring. Why? Prevents confirmation bias. Tweak: add conditional criteria for ambiguous tasks. Boom—20-30% FNR drop. Auto-research agents tried to replicate this. Hit 70% expert quality in 5% the time. But—couldn't discover structural insights like "score entire rubric, not each criterion separately." Humans still needed. For now. 🧠 Devil's Advocate: What if human labels are biased? Inter-annotator κ = 0.53-0.57. If humans disagree that much, are we chasing noise? Paper assumes humans = oracle. Risky. Hot Take: Maybe we should embrace hallucinations for exploratory agents. Not every "lie" is a bug—some could be creative workarounds. Food for thought. 💭 Bottom line: building verifiers is an art. 4 principles: → Good rubrics → Process/outcome separation → Controllable/uncontrollable factors → Evidence management Cumulative gains > raw model scaling. Check out the paper: "The Art of Building Verifiers for Computer Use Agents."
@ainativedev ·
Most people think there are only two kinds of AI agents: the giant cloud sandbox and the tiny in-product chatbot. In this piece, Lars Trieloff ( @trieloff ) argues there’s a third category emerging, and it might be the most interesting one yet: agents that live entirely inside your browser. The idea behind his project, SLICC, is surprisingly straightforward. Modern agents are basically loops that observe, decide, and act. Browsers already have storage, networking, scripting, and automation built in. So instead of sending everything to a remote sandbox, the whole agent runs directly inside the browser tab itself. That changes the economics and the integration model completely. Once an agent lives in the browser, authenticated web apps effectively become usable without official APIs, enterprise approvals, or complex integrations. Your browser session already has the access. The agent just learns how to use it. The bigger point is that expensive cloud sandboxes may not be inevitable after all. Browsers might end up becoming the default operating system for agents, not just the place where we use them. Read the full blog here: https://t.co/OGFIMsgqNj
@free_ai_guides ·
Stop doing your web busywork by hand. Start handing it to an agent. Browsing agents (Claude in Chrome, Chrome's auto-browse) now navigate, fill, and pull data on sites with no connector. 7 web tasks to hand off first: 1. Price sweeps. Point the agent at four to six vendor pages and get back one table sorted by value, not just price. Twenty minutes of tabs becomes one screen. 2. Bookings. Let it find the slot that fits your windows and fill your details, then have it stop at the confirmation screen for your yes. The agent clicks, you commit. 3. Repeat forms. Hand it your standard profile and let it fill the same fields you type every week, then check before it submits. 4. Page harvests. Give it a list of URLs and the fields you want, get back a clean table with a row per page. A research afternoon collapses into a sheet. 5. Competitor monitoring. Have it record a rival's pricing and claims once, then compare future visits against that baseline so you catch a change on day one. 6. Receipt round-ups. Send it across your vendor portals to pull every invoice in a date range into one table. The shoebox becomes a row of numbers. 7. Dashboard reads. For the tools with no export, have it read the metric off the screen and hand you the number you can paste anywhere.
Best Tweets by Topic