Autonomous engineering workflows
Delegating features, bug fixes, and application development to long-running Devin sessions, from repository work through testing and merged PRs.
40%
Best tweets about Devin
Discover the best tweets about Devin AI, including software engineering tasks, coding workflows, benchmarks, product updates, team adoption, and limitations.
Cognition Devin software engineering agent capabilities, coding tasks, benchmarks, releases, adoption, reliability, limitations, and user evidence.
Original Xholic analysis
This Devin conversation centers on delegated engineering, testing and integrations, with supportive posts accounting for 73.3% of the supplied sample. User reports describe long-running work and useful reviews, but also completion doubts, usage limits and review overhead. Release and workflow posts have standout supplied scores; neither those scores nor reported benchmark rankings establish general reliability.
63.3% of posts
All-time engagement
43.3% of posts
Published in 90 days
Conversation map
Delegating features, bug fixes, and application development to long-running Devin sessions, from repository work through testing and merged PRs.
40%
Subscription value, model promotions, credits and usage windows, token efficiency, review overhead, and enterprise productivity guarantees.
30%
Browser testing, CI checks, review evidence, and Devin Review findings, including subtle logic errors, edge cases, and supply-chain threats.
30%
Connecting Devin to Slack, Linear, GitHub, MCP tools, and personal workspaces; automating sessions, communicating progress, and handing work between agents and humans.
23.3%
Cloud-agent setup and developer infrastructure, plus Devin Outposts for running agents on local machines, private networks, GPU hosts, and Kubernetes.
13.3%
Windsurf's transition to Devin Desktop, multi-provider agent management, and experiences with SWE models, fusion, and adaptive model selection.
13.3%
Using Devin as a cloud orchestrator that dispatches parallel subagents, manages task dependencies, and coordinates with other agents and review tools.
10%
Mac environments with Xcode, simulators, signing, and computer use for building and testing native apps, sharing recordings, and delivering TestFlight builds.
6.7%
Tone and stance
Performance benchmark
Posts with media make up 60% of this collection. Their median all-time score is 23.0, compared with 3.31 for text-only posts.
Format mix
Consensus and debate
Shared view
Autonomous engineering workflows account for 12 posts (40%). Ryan Carson describes browser testing and CI-gated merging; Vincent van der Meulen describes cloud orchestration through Linear and Slack. These are workflow reports, not proof that arbitrary tasks can run unattended.
Shared view
sunil pai reports Devin finding edge cases missed in his reviews. Can Vardar describes Greptile and Devin as useful for finding small logic mistakes, edge cases and cross-file issues, without separating their contributions. Scott Wu separately claims Devin Review detected a supply-chain attack before it was publicly known. These accounts do not measure detection accuracy or missed defects.
Shared view
Cognition announces Mac VMs for building, testing and sharing iOS apps, plus Outposts for running Devin on custom machines. nader dabit describes a demonstration in which Devin builds and tests an iOS game and sends a recording and test report for review. These posts describe capabilities, not a broad success-rate evaluation.
Open debate
Jai Bhagat reports lengthy work with less steering, while Dennison praises SWE-1.7's speed but doubts its tenacity and task completion. Vincent van der Meulen's successful orchestration account also emphasizes detailed planning, testing and a remaining human-review bottleneck.
Open debate
Dwayne reports low metered usage but notes SWE-2 work was temporarily excluded. Alex Turovski reports hitting Desktop limits and questions whether the harness or model is responsible. Boyuan (Nemo) Chen argues ROI must account for accepted work, retries and human review.
What performs
The Mac announcement has a supplied allTimeScore of 339.64, or 20.44 times the supplied median; the orchestration tutorial scores 315.98, or 19.01 times the median. The macOS/iOS theme has a median score of 266.88 across two posts. These scores are not engineering benchmarks.
Sally Stockholm reports Devin at 62 on the Coding Agent Index alongside Claude Code and Codex. The supplied post does not explain evaluation methodology or task coverage, so that ranking cannot resolve the completion and reliability concerns raised elsewhere.
Statistical standouts
Creator landscape
The five most represented creators account for 33.3% of the selected posts.
1. Cognition
@cognition
2 posts
2. Dwayne
@CtrlAltDwayne
2 posts
3. nader dabit
@dabit3
2 posts
4. Anthony Kroeger
@kr0der
2 posts
5. Ryan Carson
@ryancarson
2 posts
6. Scott Wu
@ScottWu46
2 posts
Cognition supplies environment announcements; Ryan Carson supplies startup-workflow testimony; Vincent van der Meulen supplies an orchestration tutorial and setup caveats. These posts offer distinct evidence types rather than interchangeable validation.
Scott Wu says enterprise sessions and merged Devin PRs already exceeded the prior year's totals. Walden reports a team was on track for 1000+ February PRs. These are author-reported adoption signals, not independently established market-wide uptake or quality measures.
Since the previous snapshot
Themes, sentiment, stance, and post format are classified per tweet. All counts, shares, medians, creator concentration, freshness, and performance comparisons are then calculated directly from the published snapshot.
Xholic's all-time score compares engagement while accounting for reach, post age, and creator consistency. It is used for relative comparisons within this collection.
This report analyzes the exact 30-post snapshot shown below. AI identifies editorial categories and drafts explanations; all statistics are calculated from the snapshot, and every narrative claim is checked against cited posts before publication.
Best Devin tweets
Ranked 01–30
@cognition ·
Special delivery: Devin just got a Mac 🍎 Now Devin can: 1. Build & test apps on its own Mac VM with iOS simulator 2. Send a screen recording via Slack 3. Send a TestFlight link so you can start using it 📲
@vinvan ·
fable produced 60+ (!) ready to merge prs for @mainframe overnight. here's how you can set up a software factory using fable as an orchestrator. prompt included! 1. narrate your *entire* todo list — everything that's top of mind — using dictation. i used @usemonologue and my prompt was ~10K words! you have to use voice because we want excruciating levels of detail. 2. locally, tell fable to plan all the work and put it in @linear. the goal is to make it such that a dumb model dropped in linear knows *exactly* what to do. aka all issues should have a clear DAG, plan, and optimize for max. parallelizability. i recommend running fable on high and having it fan out to codex subagents on xhigh (using the cc codex plugin.) to save usage. just make sure that fable itself is doing the architecture for the hardest issues! 3. now switch to @DevinAI. we'll use devin to run our factory because our orchestrator needs to be a cloud agent (it's going to run for a long time!), and be able to spin up other cloud subagents. 4. hook up the @linear @SlackHQ mcps. this will let your agents read the fable planned work, and communicate with you/your teammates. e.g. it can talk to your coworkers to make sure it's not stepping on anyone's feet! 5. a good feedback loop is *crucial.* before starting, make sure you have a bug bot (we use @cursor_ai.) and that your agents can test all its changes. e.g. we use the fantastic @limrun to make sure our agents can do ios development e2e. 6. create a devin ultra (fable) agent — this will be your orchestrator or 'middle manager.' we're going to tell it to look at linear and dispatch subagents for all tasks. the middle manager's goal is to keep spinning up subagents with the right level of intelligence until all the work is done. and keep all agents unblocked while you're sleeping, hanging out in the park, whatever! the reason the middle manager is able to orchestrate all of this is because fable planned everything in linear. and because the middle manager itself is also fable (contrary to most of its subagents.) 7. to prompt your middle manager, attach the middle manager prompt i'll post in the comments, modified for your repo. it's important that we *attach* the middle manager's prompt to our conversation vs. typing it out, so it never gets compacted. the middle manager's context window is holy. so our prompt tells it to spin up subagents — instead of doing things itself — for pretty much everything. 8. the prompt has some other cool, non-obvious details. firstly, every agent is told to keep working until it's met the burden of proof. in our case that's a (self)-review, bug bot being green, bug bot fixes being elegant, and a recording. we tell our middle manager to enforce this. secondly, we ask the middle manager to make sure linear stays up to date and communicate via the slack mcp. specifically, it should tag coworkers in slack whenever it has a question and post heartbeat updates in the eng channel. 9. your software factory is ready to go! in our case, it worked for 12+ hours without getting blocked. and produced extremely high quality prs. addendum: all of this is possible because with @AnthropicAI fable, models are finally smart enough to orchestrate/navigate large sets of work and their dependencies. and fable is a great software architect! now of course the 'downside' is that you're now review bottlenecked. we have 60+ PRs to review! but i think as an industry we should be able to come up with some innovation there. we'll certainly try our best to help here with @mainframe too! hope this is helpful! happy to answer any qs.

@aashatwt ·
i make $14k per mo. i am spending $860 on ai subs. here is my ai stack : > devin : the best agent setup you can get with just $20 monthly sub. just connect devin with your gmail, slack, github etc. i also gave it some api keys for scraping and research. - my devin is connected with my notion so i don’t have to anything manually. everything gets updated on it own. > notion : i use notion as my second brain, content inspos. i got all my trackers there > screenstudio: i use this to make demo videos. free alternative is openscreen > gpt 6 astra is so good at writing now i am starting to use it instead of claude. > grok bot : i recently got this one. i like what i have tried so far. still figuring out how i can get max benefit!
Watch video@dabit3 ·
You can't build iOS apps without a Mac, so we gave Devin one. Devin Cloud Agents now run macOS, with Xcode, iOS simulator, and your signing setup, all in a real macOS environment (with full computer use). In this demo, Devin builds a native iOS game, then plays and tests it with computer use. Devin then sends a screen recording of the gameplay along with a test report for my review. With @namespacelabs you get as many macOS environments as you'd like, configurable with many major versions of macOS operating system and Xcode.
@ScottWu46 ·
Devin Review caught the axios supply chain attack for multiple Cognition customers before the attack was publicly known. These attacks will be 10x more frequent in the age of AI; it is critical that repo maintainers start using AI for defense as well. (showing one example below where Devin Review caught the attack within an hour of its release - text minorly edited for anonymization)

@vinvan ·
some reflections from solely using cloud agents this year: 1. every engineer should default to cloud. it completely changes how you view and use agents. if you run a company, it might be worth mandating everyone starts in cloud 2. cloud agent adoption has been much slower than i expected— e.g. looking at a ton of cursor profiles it’s clear majority cloud usage is still rare 3. getting your dx cloud agent ready still requires creative jiu jitsu. dev infra docs could be much better — “this is how to make our stuff accessible to agents/parallelizable.” luckily investments also benefit humans 4. it’s still a PITA to setup & manage cloud envs across cursor/devin etc. but i assume it’ll get bitter lessoned and we don’t need conventions for setup scripts etc. 5. where are the labs?! would love to see codex et al. invest more in their cloud experience. i know they can do it :) 6. it’s strange that cursor/devin’s investment in mobile apps lags behind their investment in cloud agents. they should go hand in hand. the ability to start agents from slack mobile isn’t enough! 7. a cloud agent spinning up other cloud agents (middle manager pattern) is goated. e.g. nice to go for a run, yap for twenty minutes, and end up with parallel agents. only devin supports this well 8. the uis of ADEs have somewhat adapted for cloud agents. but ui patterns for upcoming long running *and* proactive agents are understudied. super excited to see more experiments here (and will contribute) overall: i freaking love cloud agents. you’ll dissappoint me personally if next month you still spin up more local agents than cloud. very grateful for cursor and devin for making this technology so easy to use!
@ryancarson ·
This is how I’m currently running my startup with @DevinAI + @openclaw The browser testing in Devin is mind-blowing. I was trying to duct tape and jerry-rig all this stuff together with Playwright + uploading videos to PRs and all sorts of stuff and Devin just does it all e2e. Wild. I didn't show how I'm using @linear, which I am using for issue tracking. I have a "land" skill that I tell Devin to use whenever all the browser testing is done and all the CI goes green and it just merges the PR
@cognition ·
Introducing Devin Outposts: run Devin on any machine. Your Mac mini, a GPU box in your lab, a VM inside your private network, or a Kubernetes cluster next to your internal services.
@ziwenxu_ ·
If we only have $20 a month for AI subs 1. Devin Pro. SWE-2 is free on the $20 plan until Oct 10. 2. DeepSeek or OpenCode Go ($10), then wire it into whatever harness we like.
@ryancarson ·
I haven't typed `npm run dev` on my local machine for three days now and it's absolute bliss. Having my agents 100% in the cloud is a massive unlock. (One of those agents is openclaw, which is technically on my mbp in my office, but the only way I interact with it is via email/slack so it “feels” cloud) I'm able to run all the engineering and marketing for my startup through Slack and Linear and because of this the work product that I'm shipping has increased dramatically. I know all of us devs love creating our own custom solutions to this stuff but the truth is that creating an agent orchestration layer for your company or startup is a full-time job. Our job as startup founders is to be growing the company, not to be building out an agent orchestration custom platform. I think if you have a larger engineering team like Ramp, then it does make sense to build an entire layer like Inspect agent. However, I would venture to say that I'm getting most of the value by simply paying for a pre-built, battle-hardened solution like Devin. Again to be clear I'm not being paid by Devin or anybody to say these things, just my real-world experience using this stuff.
@aiwithsally ·
I’ve been watching coding agents closely, and this chart is getting pretty wild. Claude Code and Devin are already hitting 62 on the Coding Agent Index, with Codex right there at 62 as well. What catches my attention isn’t just the leaderboard. It’s how quickly coding agents are moving from “AI that helps me code” to “AI that can actually work on a software project.” We’re seeing models write code, debug, use terminals, navigate repositories, run tests, and keep working through problems. I think this is one of the areas where AI could change software development the fastest. The interesting question for me now isn’t: “Can AI code?” It’s: “How much of the entire software development loop can I eventually hand to an agent?”

@dabit3 ·
New video — Devin in 8 Minutes Everything you need to know to integrate and build with @DevinAI remote agents. From creating and automating sessions to autofixing bugs, autoreviewing PRs, and connecting with Slack. Also covers: DeepWiki, Playbooks, and the MCP Marketplace.
@threepointone ·
we've been using devin on agents/sandbox repo and it's been really good at figuring out real edge cases that I've otherwise missed in my own reviews (llm assisted or otherwise). thinking back to when it launched and everyone dumped on them really badly. congrats on the 180.
@CtrlAltDwayne ·
Trying out Fable 5.1 and SWE-2 in Devin today. This is incredible. My usage has moved 1% in a day using Fable 5.1 and SWE-2 in fusion mode. It has been working for over 3 hours and counting. If Devin had proper computer use, OpenAI would be in serious trouble.


@kr0der ·
no 5 hour limits is the best change ever - both Codex and Devin have it right now if i feel like blasting a lot of work one day, im not restricted by the 5 hour windows i mean it's less predictable for their infra i'm sure, but it's sooo much better as a user

@ScottWu46 ·
Interesting stat - our enterprise customers have already done more Devin sessions (and more merged Devin PRs) in 2026 than in all of 2025. Not bad for 2-ish months into the year!
@MollySOShea ·
BREAKING: Devin now writes ~95% of Devin's code Scott Wu (@ScottWu46), CEO of @cognition says the abundance era of AI is real & coming soon.. FULL INTERVIEW Cog Stats: › $2.5B raised, $26B valuation › $500M+ run rate, up from $37M LY › Customer usage is up 11x-12x in 6 months › Total code shipped inside Cognition has grown 7x › Devin writes ~95% of Cog's code, up from 89% at May Series D › The entire Windsurf acquisition was negotiated over 1 weekend, ~40 people buying a 200 person company with $82M of ARR › Goldman Sachs, Mercedes-Benz, Citi, Dell, Santander, NASA, the US Navy & the US Army all run Devin R I P : "The tokenmaxx era lasted from January to May of 2026, roughly." "I'm excited to have AI that allows me to spend all of my time & focus on the things that I care about & take care of the rest. I think the abundance era is real, & I think we'll be there pretty soon." Recorded 7 July 2026. On 23 July 2026 Cognition acquired Poke maker The Interaction Company 𝐓𝐈𝐌𝐄𝐒𝐓𝐀𝐌𝐏𝐒 (00:00) Scott Wu, Co-Founder & CEO at Cognition (01:00) The Identity crisis of building Devin (01:42) How Cognition 7X-ed its output in 6 months (02:55) The Metric every AI company gets wrong (04:18) Cognition's 12x growth curve (07:46) The Windsurf acquisition (13:26) Scaling overnight: Cognition's hardest week (15:11) The strategy behind the integration of Cognition products (17:23) Cognition's M&A playbook (20:01) Why the DMV is Cognition's next big target (22:59) Why Cognition bet against AI centralization (25:09) Why enterprises refuse to bet on a single AI lab (28:32) Scott's honest take on Superintelligence (32:36) Is Silicon Valley losing its mind over the AI gold rush? (33:26) Devin's origin story (37:17) The question Scott thinks nobody's asking yet (39:59) Poker with Peter Thiel & Napoleon ? (41:24) Scott's vision for the next year
@CtrlAltDwayne ·
Devin has been working for almost 10.5 hours now. And yes, it has been writing code in that time. My weekly usage has gone up another 1%, so I've used 2% in 10 hours with the fusion model. Most of the work has been SWE-2 which is currently not counted towards usage. INSANE.

@walden_yan ·
Cloud agents are having their takeoff moment. Across our eng team, we're on track to merge 1000+ Devin PRs in February. That is comparable to the last 3 months of 2025 combined.


@kr0der ·
if Devin is encountering an issue repeatedly, it automatically reports it to the Cognition team 👀

@walden_yan ·
I was really pleased with how Devin automatically responds on GitHub threads and backs away once you start pushing code yourself The details are what separate Devin from other cloud coding agents. Goes a long way to feel like a real team player

@markfenner ·
Someone asked why I use both Hermes and Devin. Fair question. They don’t do the same job. Hermes is my daily driver. It runs the cron jobs, business ops, content workflows, study material, and all the work that needs to keep moving whether I’m at my desk or not. It’s the operating layer. @DevinAI is where the building happens. Repos, features, PRs, security scans. When something needs to be engineered, scoped, and shipped, that’s Devin’s lane. Hermes runs the operation. Devin handles the engineering. Beyond that, I’m just a tech kinda guy. I’ll use Codex, run local models, and test whatever interesting tool dropped that week. Could I force one tool to do everything? For sure . But, I don’t need to. I’ve tuned this setup around how I work, I enjoy using it, and it’s the most efficient version of my day I’ve found. I’m always curious how other people split this work. What does your setup look like?
@DennisonBertram ·
I've been trying @cognition's harness Devin lately because I'm loving the speed of SWE-1.7. That said, its not very tenacious and I'm missing a robust /goal where I have confidence it will finish the task. Otherwise, I'm trying to get used to it. I'll try Cloud next.
@ChaiWithJai ·
When I use Claude Code and Codex, its like I'm working for them because of so much steering needed. Devin feels like the first coding agent that works for me. Yesterday, I set up the sandbox environment and went to the beach and it knocked out 2 important design engineering challenges: 1. Bringing a semi-complex AI harness to life I teach critical thinking and AI skills to first-time entrepreneurs. And I want my learners to use a whiteboard to scaffold their thinking BEFORE and AFTER each AI session. My goal when teaching cognitive strategies is to allow my learners to walk away with a model of reality (i.e. mental model) that they continually test and update until it predictively explains and reconciles reality. 2. Re-design the homepage and get the core JTBD done working to onboard my first 10 users. Devin worked from 2 PM until 3 AM last night and I'm going to plan a longer retro to understand what I can rely on Devin for and what work I need to have ready for it for the application to go live. Today, I'm going to continue working on this first milestone of getting my first 10 users. My milestones are: a. Get 10 users b. Ensure the JTBD is complete for my AI World Cup Challenge c. Extend the web application into a mobile application and publish before NYC Tech Week (bonus): raise a round of funding so I can build Khan Academy for Adults I'm so happy to welcome Devin onboard to our team so I can fire myself as our engineer and get back to doing what I love, teaching. Teaching is the love of my life. @dabit3 and @cognition thank you, thank you, thank you so much. I really feel like my dreams are in reach. Its been less than 24 hours. And I want to wait to experiment more and distill why this feels so different. Stay tuned! Life is so rich, Jai P.S. I'm excited that I might have free time again! In which case, I'll be vlogging my experience trying to bring my dreams to life. What's your dogged pursuit?

@theinformation ·
Cognition is overhauling Windsurf into Devin Desktop, a hub where developers can manage AI coding agents from OpenAI, Anthropic and others. The strategy positions Cognition as a neutral platform in a market increasingly dominated by model providers. Full story: https://t.co/ZmPZ4t1PKJ
@boyuan_chen ·
Return on tokens is becoming a real product surface. Cognition says its typical customer saw agent usage grow 1000% in the last few months. It also announced an AI Productivity Guarantee, covering up to $10M in Devin usage if it does not deliver positive ROI. That framing is more interesting than the guarantee itself. Once agents move from experiment to daily engineering habit, the buyer starts asking different questions: How much token spend produced accepted work? How many retries were waste? Which tasks should use a cheaper model? How much human review was saved or added? Which failures created rollback risk? How do you convert time saved into dollars without lying to yourself? The winning coding-agent product has to run, explain, and account for the work. For enterprise agents, observability, evals, cost attribution, and reviewer trust are becoming the same system.
@AlxTurovski ·
So what's up with GPT-5.5 limits? Ran an agent in Devin Desktop (former Windsurf) and I'm done for the day. I'm doing 3x more with Opus in Claude Code. Is it the harness or are GPT limits that tight?

@icanvardar ·
tools like greptile and devin have been surprisingly useful in github pr reviews. they pick up things i would normally miss in review, especially small logic mistakes, edge cases, and subtle issues that come from changes spread across multiple files
@nichochar ·
> need to build tiny feature (15 LOC) > ask Devin, gets to work > spins up a whole computer to test it with browser use, 500M tokens burned > works, submits to github > 5 other agents (codex, devin, cursor bugbot, etc...) all perform a review > for each comment, each bot polls and responds > small improvement found, Devin fixes it > round 2 of the five agents reviewing > tokens go brrrr > LGTM, merge > 2h, 1B tokens later, tiny feature shipped. is this agi?
@tristanbob ·
It's been a few months since I last used @DevinAI Desktop (formerly Windsurf). I'm using the "Adaptive" mode because I want my credits to be used as efficiently as possible. (BTW, I'm still using the grandfathered original plan from when Windsurf first became available!) First observation, it produces incredibly verbose thinking output. I love reading thinking about, but this was too much. You can see it repeating things, checking rules, correcting itself, etc. Second observation, it worked great! It pushed two PRs to an open source project after testing them locally. Have you tried Devin recently?
Best Devin tweets
Xholic studies what works in your niche, drafts posts in your voice and schedules them for the hours your audience is online.
$0 today · Cancel anytime
Browse all tweet collectionsKeep exploring