Trust First, Then Parallel: What Lauren Tan's Talk Actually Argues

Published on October 5, 2026 by Remy

Lauren Tan is @poteto on X, and poteto on GitHub as well. On 2026-09-21 she published a post (the original is here). The number in the post’s title is 2,500; the gist is: last month I shipped 2,500 PRs to production. The talk was originally planned for Cursor Compile London and was later turned into a livestream for Grok @Bot Galaxy. The video runs about 38 minutes, or 2,281 seconds, and its video id is 2101938030122868736.

The spoken number is different. The transcript comes from X’s automatic English captions, not a human-checked transcript. In those captions she says that last month she merged about 2,000 pull requests into production. 2,500 belongs only to the title of that X post; about 2,000 belongs only to the auto-caption transcript of this talk. Both numbers stay in this post. I don’t merge them, and I don’t use either one to correct the other.

Her argument is narrow. Throughput comes from trust, not from opening more agents. Trust means that when no human is watching, an agent still turns in work of acceptable quality. Once that environment is in place, output goes up, and it starts to look a bit like a personal or team “software factory.” She doesn’t like that term. She’d rather compare it to a Michelin kitchen: it’s not an assembly line stamping out copies, and a person is still responsible for what goes out to the table. When agents take over parts of the work, what you’re arranging is who stands at the stove, the tools, the training, and the ratio of prep cooks to chefs. Later in the talk she crosses out “software factory” altogether.

The auto-captions mangle a number of proper names, and none of those misheard forms appear in this post. The account is poteto, the plugin is PStack, the framework is Dune, the product is Grok Bot, the memory capture is a heap snapshot, static checking is lint, and the review bot is Bugbot.

Without trust, a hundred cloud agents just produce garbage

In the talk, she starts the story about six months ago. She had just joined Cursor, and the company had not yet been merged into SpaceXAI. There were no ready-made agent skills (reusable procedures written for agents), and both the codebase and the product were new to her. Cursor was building a replacement for the IDE: the new agents window. Before she arrived, that window already had plenty of performance problems. Her manager asked her to help because she had been on the React team before joining Cursor.

Performance work she knew how to do. What she couldn’t keep up with was the merge rate. Pull requests kept arriving like a wall, and she had no way to tell whether the app’s performance was regressing. Early on, almost all of her time went into Chrome DevTools: reading the Performance panel, capturing traces, taking heap snapshots. It was manual enough that she got sick of it. We already have agents, so why is a human still sitting in front of these panels?

So she started thinking about verification skills. Have the agent launch the app itself, capture a trace itself, understand the trace, find the hot spots, and then push performance up. She says her output went up over these six months, but she never treated “about 2,000 PRs a month” as a goal. Looking back at the skills, tools, and codebase changes, they all stack on the same thing: trust. The blunter way she put it at the time: I’m the bottleneck, and I need to move what engineers know into this set of agents so that not everything gets stuck on me.

Her public résumé separately says she joined Cursor in April 2026. Both the “about six months” in the talk and that start month come from checked material. I don’t reduce them to a single date, and I don’t fold the later “fixing Cursor 3 on day two” story into this account of the agents window.

She describes how she worked at the start: she only dared watch one to five conversations at a time. Every chat had to be supervised, with constant course corrections. If she stepped away, either nothing happened or the agent did the wrong thing. She thinks this is the hardest stage to get out of, because you can’t see the exit. You can’t get out because you don’t yet trust what the agents hand back. If you open 100 sub-agents or cloud agents (agents that run in the cloud) at that point, what you get is a pile of garbage pull requests, regressions, and bugs. Nobody is happy.

So the question isn’t how many more you can open. It’s what it takes before you dare stop watching.

Every correction: first ask whether the strongest layer can catch it

When she corrects an agent, she ranks the places a fix can live into five layers, from strongest to weakest. The order is part of the argument. The higher a layer sits, the less it depends on an agent or a human remembering to read some piece of prompt text in one particular conversation.

Models carry forward whatever is already in context. Files the agent has read or opened are in context. It won’t refactor the old code away in every PR; it continues in the same style. So what the repository looks like lasts longer than a one-off correction in chat. A correction that lives only in the conversation is invisible to the next session. A correction that becomes “this can’t be written” or “CI will fail” can’t be dodged by the next session either.

The first layer is the codebase and the architecture: make the bad pattern impossible as a category. She treats the codebase as the agent’s memory. Existing anti-patterns get treated as memory too. A small workaround, or a comment explaining that workaround, gets copied into a de facto standard within days or weeks. She compares that spread to a virus, and also to growth in a garden that shouldn’t be left there. The state you want copied has to be a state you’d be happy for the next agent to imitate. This layer includes changes to data structures and to how things are done. Instead of verbally correcting the same mistake every time, make that mistake impossible to write. She thinks that if a team really believes most future code will be written by agents, this is the engineering most worth investing in, not just writing more prompts.

The second layer is static analysis: lint, compilers, and CI. These are constraints you can enforce in the repository. If an agent keeps making the same mistake, you can add a lint rule; better still is going back to the first layer and making the mistake categorically impossible. When she sees tech debt or a bad pattern, her first reaction is to write a lint rule. It doesn’t have to be cleaned up immediately. Stop the bleeding first so it stops growing, then spend time having agents clean it up. She didn’t name any specific linter or CI product, so don’t fill this layer in with a particular vendor’s tool.

The third layer is rules, Bugbot, and skills. Here we’ve moved from hard constraints to guidance. Bugbot has docs here. Rules and skills get used most of the time, but the agent may forget to read a rule, and the person driving the agent may ignore them. So this layer matters, but you can’t treat it as enforced. Miss one read and the constraint might as well not exist.

The fourth layer is style guides and human review. If a style guide hasn’t been written into rules, Bugbot, or skills, it only takes effect when a person reviews the code. That person has to read every line and remember to leave a comment. Once PR velocity goes up, that can’t be done. She doesn’t recommend relying on style guides alone. Human review is good for discovering what’s missing, and then the time should go into the layers above, rather than treating human review itself as a way to scale. When the same kind of review comment keeps coming up, it’s a sign that the concern is still stuck at layer four and hasn’t been absorbed into the architecture, lint, or rules.

The fifth layer is verification skills. They can prove correctness; they don’t automatically prove performance or code quality. Correctness, for her, means: did this feature or this code do what you wanted it to do? A verification skill can produce empirical evidence, such as a path actually working end to end. It doesn’t tell you whether the feature is fast, and it doesn’t tell you whether the code is well written. Performance needs its own metrics and telemetry. Code quality needs a different kind of skill, one that teaches the agent to work the way an engineer would, rather than only proving “it ran.”

At the other end of the spectrum are formal methods. Lean and TLA+ can be used to check whether business invariants always hold and whether the application is in a state that can be formally proven. She says she hasn’t gone down that path herself. Formal methods are still hard and still an open problem, and very few people actually use them together with verification skills. Her judgment is that verification skills can take you a long way even without formal methods. That’s a boundary, not “formal methods are useless.” She doesn’t claim verification skills cover the kind of invariants formal methods are meant to prove.

To wrap this part up, she asks the audience to remember the order. When you notice you’re chasing an agent with corrections, pick the most effective move: first make the pattern impossible in the codebase, architecture, or data structures; if you can’t, use static analysis; then layer rules, Bugbot, and skills on top. Quality-oriented skills need separate time. Only when these layers stack up is there enough trust for agents to move forward on their own. She says this isn’t a trick; it’s a lot of work. Human review and style guides are for finding gaps; verification skills are for getting evidence of correctness. Neither should replace the first layer.

Control Glass is just an internal verification skill

When she joined Cursor and started on agents-window performance, the first skill she built was called Control Glass. It’s a verification skill. It has no public page. There’s no link below, and there’s no public documentation to read alongside this. Don’t describe it as a released product.

It teaches the agent how to launch the app and capture traces, using the Chrome DevTools Protocol. She says the structure only became clear after iteration; it didn’t start out this way. A verification skill has two parts.

One part is a reproducible CLI that lives in the skill directory, instead of having the agent write a fresh script every time. Scripts drift between sessions, and the same performance line ends up measuring different things. The fixed CLI is responsible for launching the app, collecting traces and heap snapshots, and producing empirical evidence that “the code works and the performance line was met.” This part needs continuous investment to cover different usage patterns, rather than stopping after one demo.

The other part she calls a feature map, which you can think of as a map of the app’s features written into the repository. She says she came up with the name herself, inspired by something like a sitemap. In essence it’s memory written to disk: how the app works, what features it has, how users reach them, what the keyboard shortcuts are, which DOM elements to click, and what each feature does. The feature map lives in the skill directory of the codebase and is maintained by a set of automations. “Automations” here is her description of how this map is kept up; don’t present it as a feature list from some publicly released product.

Neither part is enough alone. Internal Slack is full of very vague reports: a small screenshot of some piece of UI, plus three question marks. The agent can launch the app, but it can only guess what the user is pointing at. With the CLI plus the feature map, the agent can reliably control the app, capture traces, and understand requests from both internal and external users. She says the team quickly came to treat this kind of app-controlling verification skill as critical infrastructure that has to be maintained continuously. An agent being able to check its own work is a very concrete piece of trust, and they spent a lot of time on this skill.

Let me be clear about the boundary. Control Glass answers “did it get done,” and whether performance samples can be captured the same way every time. Being able to capture a trace doesn’t mean the code quality passes, and it doesn’t mean you’ve done a formal proof. It also isn’t a documented package outsiders can install. If you want to build something similar, you need your own reproducible CLI and you need to write down how features are reached in the skill directory, rather than looking for a public page that doesn’t exist.

Dune turns the shortcut into the only correct path

In the Grok Bot codebase, they built a client framework they call Dune, aimed at making it hard for agents to go off track. Dune has no public page. Don’t describe it as a product with a documentation site, and there’s no link here either.

The direct trigger was the performance problems in Cursor’s agents window, and many of the practices came from there. The principle she set: agents love shortcuts, so make the shortcut the correct path. A codebase like that is annoying for humans; what you can and can’t do is locked down tightly. She thinks that’s exactly what suits agents, especially agents with very little context. In the future, engineers won’t be the only ones submitting changes to the repo. Busy people with little context, like designers, product managers, and CEOs, will come in to build features too. Getting it right by default is more dependable than expecting them to read the whole style guide first.

The codebase is memory, and the converse holds too: anti-patterns grow on their own. The state she wants is one locked down to the point that it’s uncomfortable for humans to write, but with conventions strong enough that even innocent-looking patterns don’t get left behind. Her example is comments. At first she didn’t think agents leaving comments was necessarily bad. People leave notes for colleagues next to edge cases and workarounds too. Later, in Cursor’s codebase, she saw agents using comments as an excuse not to fix the real problem, papering over bugs with short-term hacks. So in Dune, which is used for Grok Bot, they banned comments, to keep that pattern from being copied across the whole repository. A comment had gone from “an explanation for humans” to “a bad example for the next agent.”

She gives this role a name: the gardener. A repository, like a garden, grows things you don’t want, and you have to pinch them off before they spread. She admits she doesn’t know anything about gardening, so the metaphor stops there: someone has to be dedicated to watching what’s creeping into the codebase. Behind Dune she stresses three things. Delete existing tech debt, because agents will copy it. For most patterns you want to encourage, leave just one paved road so the agent doesn’t have to guess; the codebase, CI, and lint need enough guidance to herd it onto that road. When you see tech debt or a bad pattern, the instinct should be to write a lint rule. Stop the bleeding first, then have agents clean up the old stuff, so the repository stays in a state where “being copied is fine.”

She didn’t walk through Dune’s structure item by item. She gave one example: a performance problem learned from Cursor’s agents window and then eliminated in the architecture. Things that run on the Electron main process are not allowed to end up in the renderer. In the agents window, code had been mistakenly imported into the renderer, and slow code came along with it. The renderer has to keep the UI smooth and can’t carry long tasks that exceed the frame budget: about 16 milliseconds for 60 frames per second, about 8 milliseconds for 120 frames per second. Work has to be split up, not piled into one chunk. Dune uses the import dependency graph to fix this boundary in place, making that kind of mistaken import categorically impossible. She says the single example doesn’t matter. What matters is that your own framework can pull senior engineers’ experience out of style guides and review comments and write it into the codebase. The codebase then becomes the already-committed state you want agents to extend. The next agent that comes in is more likely to keep that state intact, rather than learning a workaround from an outdated comment.

On the Grok Bot side, she has locked the repository down to the point where it’s almost impossible to write bad code. An agent with little context and weak reasoning can still come in and write passable code. That’s because the road has been narrowed, not because more conversations were opened.

The outer loop is about connecting tools, not building a company brain

Later in the talk, she places Grok Bot and Cursor in different positions. Grok Bot is good at what she calls the outer loop: plugging into all kinds of connectors, gathering information, and using it to make decisions. Her examples include Slack, Datadog, Sentry, PlanetScale, and whatever services you’re already using. These are examples, not a fixed integration list. Don’t present the four names as a ready-made bundle, and don’t add tools she didn’t mention.

Some people call this kind of setup a company brain. She doesn’t buy into a fancy version of it. She doesn’t think it needs to be that complicated, because agents are already good at using tools. Connect the tools to Grok Bot, let it kick off cloud agents automatically, and you don’t need to build a huge pile of infrastructure first. She crosses out “software factory” again and says instead that you can use Grok Bot to build that kitchen for yourself. Routines (which automatically kick off tasks based on subscriptions or conditions) let you subscribe to a Slack thread or a Sentry alert and start work automatically. The public documentation is in the Grok Bot overview, and in skills, routines, and automations.

All of this stacks on top of the earlier layers; it isn’t a separate system. The codebase, rules, and skills come first, and only then can Grok Bot react to events from the outer loop and kick off cloud agents. On the Cursor side, you can also use automations and the TypeScript SDK to build bots that reuse the agent infrastructure you’ve already laid down for more complex tasks. During the livestream she showed some screenshots of automations on Cursor: automatically reproducing bug reports, automatically opening pull requests. Her framing is that stacking these things adds usable output for the whole team. There is no independent third-party output figure here. The screenshots are her demo, not an audited report.

Cloud Agents and the agents window are public product entry points. They don’t solve “going parallel before you have trust.” If the outer loop runs on a repository that hasn’t been locked down, automation amplifies garbage PRs. Get layers one through three in place first, then talk about subscribing to alerts and auto-opening PRs.

PStack is the kind of engineering playbook she puts in the guidance layer. In the talk she made a point of saying she wouldn’t go deep on this plugin today. It’s a Cursor plugin she built: a set of skills drawn from how she herself debugs, builds features, and prototypes. The public page is on Cursor’s plugin marketplace, and the guide is in the plugin repository’s docs. The entry point is /poteto-mode. In practice you first run /setup-pstack to pick a model, then enter /poteto-mode. She says these are the skills she uses every day at Cursor, and the goal is to write less and write it correctly, not to pile up lines. More experienced engineers can collect skills like these into a team repository so agents write code the way your team wants. With verification skills plus these workflows, the agent isn’t just proving it got things right; it has a chance to reach quality. Performance numbers still have to be captured through that reproducible sampling pipeline, not by the name of a playbook.

What agent builders can take away, and what she didn’t say

Which layer a correction goes into is worth looking at before how many PRs were merged last month. While a person is still stuck at one to five conversations, don’t treat parallelism as progress. If the same mistake has only been corrected in chat, the next session will make it again. The same kind of comment recurring in human review means it hasn’t yet become architecture, lint, or rules. The guidance layer should be assumed to get skipped by default, so it can’t be the only gate.

What she didn’t say should be kept on the books separately. Verification skills only provide empirical evidence of correctness; they don’t replace performance metrics, and they don’t replace the invariants Lean or TLA+ are meant to prove. She hasn’t gone the formal-methods route herself. Neither Control Glass nor Dune has a public page, and this post doesn’t add links for them. Slack, Datadog, Sentry, and PlanetScale are examples, not a fixed list. She doesn’t advocate building a fancy company brain first. The 2,500 in the X post title and the roughly 2,000 in the spoken talk come from different sources. In the talk she says explicitly that about 2,000 PRs was never a goal she set. A hundred cloud agents is the counterexample for when there’s no trust, not a recommended setup. The outer loop only adds output when the repository is already worth copying; otherwise it automatically amplifies garbage PRs.

The checked career path this method grew out of

The path that public materials support: a background in front-end and type systems, tools for the content studio at Netflix, the React compiler at Meta, and now applying “let the agent verify itself” at Cursor and Grok Bot. She lists her location as Southern California. For current details see LinkedIn. Her old site no.lol still shows the Facebook and pre-Netflix introduction; it’s out of date and shouldn’t be read as her current role.

She’s self-taught and switched careers. She started in UI design, then taught herself to turn those designs into working interfaces. about.me says she has degrees in finance, from the London School of Economics and Monash. That page dates from when she was still an engineering manager at Meta.

Her blog from 2014 to 2016 is almost entirely Ember.js, written while she was at DockYard. Her GitHub account goes back to 2012. Her best-known repository is hiring-without-whiteboards.

At Netflix she was an engineering manager on Studio UI, the content production tools covering everything from the pitch to launch. Earlier she led Studio Programming, which schedules original content. Her old site still lists three talks: Building the World’s Largest Studio at RubyConf 2018, Just Use Any at TSConf 2019, and Ambitious UIs for Pitch to Play at Netflix UI 2019. The stack for that last one was React, TypeScript, and GraphQL. These pages show what tools she was building at the time, not her current title.

She spent about six years at Meta and left in mid-March 2026. A LinkedIn post on 2026-04-07 says she had left about three weeks earlier. In the React org, public descriptions cover this work: prototyping Server Components, the Relay compiler and runtime, React 18, React Compiler, useEffectEvent, cross-platform React, and later helping shape the org’s AI strategy. Her React Conf 2021 speaker page says she was then an engineering manager on React, and that before moving into management she led the Server Components prototype at Meta. The page is here.

In the same LinkedIn self-description, she says she later rewrote React Compiler as an individual contributor: moving from statement-level analysis to expression-level analysis based on a custom control flow graph (CFG). She mentions SSA, hoisting, validation, eslint-plugin-react-compiler, and an experimental compiler LSP, and says she was responsible for the open-source release. This passage is her own account. The speedup figures in it have not been independently verified for this post. She writes that after shipping on Instagram Web, Threads, the Quest Store, and Facebook.com, interactions became 2.5× faster, loads and navigations 12% faster, with memory flat. She also writes that React’s CI moved from CircleCI to GitHub Actions, which she says was cheaper and more than 50% faster. 2.5×, 12%, flat memory, and more than 50% all appear only in this LinkedIn self-report. They are not from the talk, and they are not third-party measurements. Any citation has to carry “she wrote this herself; not independently verified,” and must not present them as proven product results.

She joined Cursor in April 2026. Her post says that on day two she was already fixing performance in the Cursor 3 client: teaching the agent to run the app, profile it, and capture traces, then following the metrics to deal with memory leaks and rendering. In the talk, it’s the Chrome performance work on the agents window, first done by hand and then folded into a verification skill. Both are her own accounts. I don’t merge them into a single change, and I don’t supply a matching PR number. In May she open-sourced PStack. Her post says the engineering team used these skills about 10,000 times that week. That 10,000 is also her claim, not an independent count. After Cursor was merged into SpaceXAI, her title became Grok Bot at SpaceXAI. Her current public profile: working on Grok Bot and Cursor at SpaceXAI, while also on the React Compiler core team.

What lines up with the talk are two ends of the same thing: a background in compilers and performance, and then at Cursor, folding manual traces and heap snapshots into verification skills, and turning patterns that shouldn’t be copied into hard constraints in the repository. The 2,500 in the X post and the roughly 2,000 in the talk are both numbers written down after walking that path; they aren’t the method. The contact she left at the end of the talk is poteto on X.