Full Width [alt+shift+f] Shortcuts [alt+shift+k]
Sign Up [alt+shift+s] Log In [alt+shift+l]
38

Crafting Engineering Strategy!

from Irrational Exuberance [alt+shift+b] in programming

On November 3rd, 2023, I posted Thoughts on writing and publishing Primer to celebrate the completion of my work on my prior book, The Engineering Executive’s Primer. Three weeks later, I posted Engineering strategy notes on November 21st, 2023, as I started to pull together thoughts to write my upcoming book, Crafting Engineering Strategy. Those initial thoughts turned into my first chapter draft, How should you adopt LLMs? on May 14th, 2024. Writing continued all the way through the Stripe API deprecation strategy, which was my final draft, completed the afternoon of April 5th, 2025. In between there were another 35 chapters: two of which were Wardley maps, four of which are systems models, and two of which are short section introductions. This post is a collection of notes on writing Crafting Engineering Strategy, how I decided to publish it, ways that I did and did not use foundational models in writing it, and so on. Buy on Amazon. Read online on craftingengstrategy.com. Why write this book? One of my decade goals for the 2020s, last updated in my 2024 year in review, is to write three books on some aspect of engineering. I published An Elegant Puzzle in 2019, so that one doesn’t count, but have two other books published in the 2020s: Staff Engineer in 2021, and The Engineering Executive’s Primer in 2024. So I knew I needed one more to complete this decade goal. At one point, I was planning on finishing Infrastructure Engineering as my fourth book, but honestly I’ve lost traction on that topic a bit over the past few years. Instead, the thing I’d been thinking about a lot instead was engineering strategy. I’ve written chapters on this topic in my last two books, and each of those chapter was tortured by the sheer volume of things I wanted to write. By the end of 2024, I’d done enough strategy work myself that I was pretty confident I could take the best ideas from Staff Engineer–anchoring the essays in amazing stories–and the topic I’d spent so much time on...
25th Oct 2025

Stay updated

Get a weekly newsletter with the top 5 articles worth reading every week.

More from Irrational Exuberance

Trying the Software factory pattern.

One of the interesting challenges of the AI ecosystem in 2026 is that new, effective patterns emerge faster than I can adopt them. I’ll find a handful, get back to work, and realize a month later that I’d missed four or five more. The adoption cycle for Imprint this year has been something like: January: get every engineer onto Claude Code every single day March: ok, let’s also get everyone else onto Claude Code or Claude Cowork every single day April: local development is bottlenecked on checkout and worktree model, instead create ~10 local workspaces which each have an independent checkout of every repository, and operate at the workspace level, not at the repository level, so it can generate cross-repository pull requests across frontend, backend, infrastructure and data monorepos June: oh boy, agent-driven development is heavily constrained by lack of a common task management system with higher visibility and less permission complexity than Jira, so let’s migrate the entire company over to Linear and hard stop on Jira July: yikes, now we have visibility into all these tickets, many of them are trivial but managing them through local development isn’t scaling, let’s roll out an orchestrated harness which internally we call “Agent Fleet”, along the lines of Stripe’s Minions The most recent question for me has been figuring out how to adopt the software factory pattern. (After some light research, the specific AI-context origin of this term is slightly messy to attribute, but I think it might be Justin McCarthy in February 2026’s Software Factories And The Agentic Moment.) The software factory pattern is looping on a broad goal, and then relying on the harness to drive progress towards that goal. Our first pass at implementation is fairly basic: An agent skill /linear-project-loop which reads in a Linear project and starts by auditing that project’s goal definition on these dimensions: An RFC in Notion that describes the project’s goals, how those goals are measured, and the general approach A Datadog dashboard or Snowflake queries that measure progress against those goals If those are missing, or the Linear project is missing in its entirety, it iterates with you on creating those missing tools. Then it reviews the state of the metrics and issues for the project. If new work is identified, it adds those issues to the project. It updates the state of issues that have moved. It works on the non-blocked tasks based on the project’s current state. This is often writing a pull request, updating a pull request, pinging for review, asking a clarifying question, etc. When a task completes, if the project description is fresh, it takes on the next task. If the description hasn’t been updated in a while, it reruns the loop starting with the first step. Right now I am running this locally in a local harness, but it’s working well enough that I anticipate moving the behavior to be driven by the same orchestrated harness that we assign one-off tasks to. What I particularly like about the factory pattern is that it parallels very closely how I’ve been working locally, while forcing me to recognize the places where I was accidentally hording parts of the state for myself regarding the goals of the project. I was already asking agents to iterate on specific Linear projects, but they didn’t have the ability to evaluate if they were going in the right direction, or if it was missing necessary tasks. Now it does. The other place this has been extremely helpful for me is checking in on projects post release. For example, I shipped our passkeys implementation earlier this year, but some months go by without my checking in on how it’s going. If we saw adoption spike, or error rates start to turn, I might miss it, but running the factory in a less frequent post-release mode would catch it immediately. The final thought that’s been interesting to me is how much all of the pieces here compound only to the extent that you have the other pieces. For example, this factory pattern depends on having Datadog MCP and Snowflake access available to manage goal-tracking, but it also depends on Linear being the single source of state for the company’s work, and an orchestrated harness that can perform work independently from your laptop. Keeping up with this many migrations is a fascinating industry moment.

a week ago • 3 votes
Middle management roles are also a trap.

Six years ago, I wrote Tech Lead Management roles are a trap. My argument then was that TLM roles present themselves as easier than moving into a full management role, but the tension between doing the software engineering and engineering management aspects of the role made being a TLM a much harder first management role than a pure engineering management role. I still agree with that post, and I have some additional bad news to share: middle management roles are mostly a trap as well, if you goal is to become an executive. The core aspects of middle management roles are: Balancing between top-down executive, lateral stakeholder, and bottom-up team pressure, e.g. keeping morale up as an organization deprioritizes last year’s big initiative Defining and operating an organization’s process, e.g. creating career ladders and interview loops Competing for a share of fixed organizational resources (e.g. budget) and allocating acquired resources These are all extremely important skills to be an effective executive, and they make up the bulk of The Engineering Executive’s Primer, but they are insufficient to make you a great executive. If you don’t have them, you will be a deeply flawed executive, but even if you’re an expert at them, you can still be a terrible executive. That’s because the most important skills of an effective executive are the same exact skills that make an excellent line manager: developing domain expertise, driving execution (including setting pace), and translating both of those into an organizational culture that extends beyond you (in any of innumerable different ways). All of them are more easily practiced and mastered as a line manager than as a middle manager. Most middle management roles make practicing those skills difficult, and sometimes negatively select against developing them. As a middle manager, if you drive execution too closely, you might get told off as a micromanager. As a middle manager, if you go too deep on domain expertise, you might get told that you’re not focusing enough on your internal stakeholders. That’s undoubtedly valid feedback in many middle management roles, but it’s the perfectly wrong feedback to someone who is trying to become an effective executive. As a result, I’ve come to believe that the filters for good middle managers inadvertently negatively select out the “challenging” line managers who actually have the best chance to be excellent executives. That’s not even necessarily irrational: most companies are developing middle managers to take on more complex middle management roles; very rarely do they worry about growing future executives from within. If you’re willing to embrace this fact that a little bit of time in middle management roles is important preparation to become an executive, but spending a great deal of time in middle management only prepares you for further middle management roles–and makes you less effective as a potential executive–then the AI-driven shift in managerial fads might be a threat to your current role, but it’s likely to improve your chances to succeed as an executive.

8th Aug 2026 • 2 votes
Make no assumptions.

I’ve recently been thinking a lot about the concept of “soil horizons”, which is the idea that there are many distinct layers of soil, from topsoil all the way down to bedrock, which all combine into a soil horizon. Translating this idea into software, the ideal codebase would have a single uniform “code layer”, but a surprisingly large percentage of production software has numerous, distinct code layers as the leading architect shifted over time. I’ve found this particularly true for software in problem-spaces with high essential complexity and low scale complexity, where the purifying challenges of scaling never create enough pressure to compact disjoint layers into a unified layer. Codebases with the most code layers tend to be created by small teams working on complex domains over a long period of time. In many companies this might be an identity, permissions or payments team: stuff that’s permanently valuable, but usually not the central concern at any given time. On such teams, there is often only one architect who understands the nuances of the domain well enough to make tradeoffs. When that architect leaves, they are replaced by someone who aspires to operate in the same code layer, but simply cannot because they lack enough context to do so. As a result, that new replacement creates a new code layer, despite not intending to. If the team runs through a handful of folks as the new team leads struggle, it’s easy to end up with a complex code horizon very quickly. The problem of messy code horizons is not a new one, and the general approach to addressing them is the same one I wrote about seven years ago in Reclaim unreasonable software, but with the proliferation of coding and non-coding harnesses, lately I’m running into the problem of messy code horizons more frequently. Even more concerning, I’m seeing this problem expand from impacting code horizons into impacting how organizations make decisions outside of software, e.g. the company’s general reasoning horizons. When individuals or teams rely on LLMs to reason to conclusions, rather than using LLMs to explore or draft options, it’s possible for even the most important decisions to be built on top of flawed reasoning layers underneath. In the next section, I’ll develop the problem statement a bit about what I’m running into, and then in the final section I’ll lay out the approaches that I am finding (moderately) effective to navigate that problem. Messy reasoning horizons If you give three enthusiastic engineers a problem, a new codebase, a coding harness, and self-approval rights, it’s very easy to end up with three new soil horizons as their harnesses gleefully commit code. However, in engineering we have a number of techniques to derisk this problem. First, we have manual and automated code review, and second we increasingly have the ability for the harnesses to operate off sufficiently clear instructions that they write new code consistently with the existing code, even if the operator is unaware of what good looks like. This is also true for code review, where coding harnesses can drive consistency across pull requests even if the person (or harness) creating the pull requests is not operating off the same shared context as the wider team. Many codebases are not well-configured for this new reality, and those codebases are getting worse at an accelerating rate as more harness and agent contributions get added. Legacy codebases that reach a certain size before introducing these better practices are easier to fix than before, but still require a lot of work to fix. That said, I’m confident that coding harnesses are going to substantially improve the quality of code horizons over the next year or two as the way we configure harnesses improves. That’s not the problem I’m worried about. What I’m worried about is the application of harnesses to problems outside of writing software, where there’s no static typing, linting, or unit tests to validate the output. Let me provide a very recent example from my own work that highlights this problem: I wanted to understand how our incidents were trending over time. So I pulled data via an MCP, and the analysis was unintuitive to me, in particular I thought we were having more Data related incidents than the results reflected. I had to look at the incidents in Slack, then the results in our incident tool, and understand why the two conflicted. After a bit, I recognized the results in our incident tool were only showing incidents that properly tagged a team when the alert was triggered, so it was omitting about half the relevant incidents. After having the agent manually tag the incidents without team assignments, the data made a lot more sense. After recognizing the issue, it was trivial to fix. However, if I had simply accepted the initial analysis, I would have made the perfectly wrong conclusion about what was happening. On top of that wrong conclusion, I could have easily pushed the team to take on a project to solve an illusionary problem. What’s so pernicious about messy reasoning horizons, is once any reasoning layer is poisoned, it’s impossible to reason effectively on top of it. If you take the incident analysis example, it’s easy to imagine prioritizing the perfectly wrong set of remediations, which have the artifacts of solid strategic reasoning, but are nonetheless just wrong. It’s easy to imagine a team wasting a quarter of time building a solution to this sort of problem that never existed. It’s true that poor reasoning has always existed, long before harnesses, but my experience is that poor reasoning wearing well-formatted clothing is proliferating more widely than I’ve previously seen, and it is increasingly difficult to combat because certain social norms are – at least temporarily – collapsing around folks actually thinking. That collapse is largely driven by unprincipled adoption of AI techniques without paying attention to whether they work. Widespread adoption is, in my opinion, the fundamental risk for most companies at the moment, and something companies need to be doing, but many approaches inadvertently mix play (experimenting with something new in ways that are likely to fail!) with production (creating load-bearing work product!) in ways that erode social norms for quality. The norms are not uniformly collapsing by any means, they are generally intact, but even a small increase in the proliferation of low quality reasoning layers has a devastating effect on your ability to reason successfully. Especially true the further up the poor reasoning occurs (sloppy reasoning from senior leaders) or when senior leaders rely on layers of reasoning without inspection (leaders who aren’t sufficiently “in the details” to spot likely reasoning errors in reasoning layers). As a result, we now live in a world where accepting any part of the reasoning context before inspecting it might lead to making a catastrophic mistake. This is an exhausting way to live. Make no assumptions Accepting that this is the world we live in, I wanted to lay out the techniques that I am finding useful to deal with it. Some of these are novel, but many of them are the same techniques I was using before the LLM-advent: Make no assumptions. When new hires join my team or my company, the first thing I tell them is that it’s essential that they “make no assumptions.” This is difficult to do, and it goes against every instinct because it forces you to inspect each aspect of how the company works and thinks, but I do think it’s the necessary approach. It’s a bit like learning “internet-skepticism” at some point in your life, where you realize that everything on the internet is self-motivated in some way, and you have to maintain a strict filter on what ideas you accept. This is a hard change to make, but I genuinely believe this is the correct mindset for accepting new information in the current era. The combination of fewer management layers and more flawed reasoning layers means that the core job of leadership is inspecting the details. The author must be the first human in the loop for their output. The biggest cultural failure with harnesses is when you can tell that you–the recipient of a piece of work–are the first human in the loop reviewing it. You must set a cultural norm that the creator of a piece of content is always the first human in the loop before asking another human to review it. If you fail to set that cultural expectation, then you will quickly crush the remaining team with a high standard for quality reasoning, which will lead to a full destruction of your reasoning horizon. Prioritize reasonable software. Run the Reclaim unreasonable software playbook, recognizing that migrations are cheap in 2026, so it’s much faster to remediate gaps. The core idea here is that relying on convention doesn’t work, and instead you have to rely on deterministic decisioning for each approach. For humans this can feel overly prescriptive, but harnesses don’t care. Learn faster by separating play and production. Many folks trying to learn how to use harnesses and LLMs leap directly into using them in their most critical work. This is a slow way to learn, and can lead to substantial errors in your most critical work. It’s much faster to work by buffering small pockets of time to learn. For example, our head of data has spent time building an iOS app fully “hands off the keyboard” to get a better feel for the tools. This sort of experiment goes much faster and gives you more repetitions in less time. The very practical version of this is setting aside a day or two periodically for folks to experiment. Structure how you think with LLMs. In Crafting Engineering Strategy, I lay out a structured approach to reasoning through creating a strategy document, which aims to prevent the reasoning errors that folks make in their thinking. This applies equally in how we use LLMs, and I think you can substantially reduce the chance of introducing flawed reasoning layers by focusing LLM work on exploration (gathering information on internet and via various MCPs), refinement (presenting gathered information effectively), and a final formatting pass. That takes much of the work out of strategy creation while constraining the areas you have to avoid making any assumptions about its output. I’m certain there are more things! What are you trying?

11th Jul 2026 • 0 votes
Revised rules of engineering leadership.

From early 2014 through late 2020, I was working in hypergrowth environments, which are challenging, but also educational. The most valuable feature of hypergrowth is that your mistakes reveal themselves next month rather than next year, because things go wrong very loudly when you’re moving fast. I’ve been thinking a lot about hypergrowth recently, because Imprint’s business is growing quickly and we did a large batch of hiring last year, but also because the AI-tooling shift has changed the pace at which it’s possible to work. This post documents the new rules I’ve revised my approach to engineering leadership around, and then talks through the specific projects I’ve worked on over the past year that caused me to believe in these rules. Revised rules Migrations can be done by an individual rather than a team. Even complex, large changes can be 95% owned by the driving individual or team, and done in 10% of the time. As the initial cost of migrations goes down, the reward/penalty of each migration’s quality goes up: even small sharp edges will break your colleagues’ mental models about the software you co-maintain. The impact of individual judgment on your company has never been higher. While 1st-pass code is nearly free, the cost of working code depends on your development harness, and is not free. We’re in an era when many companies say that everyone should be writing code, however our experience is that writing code that works well, while avoiding messy edgecases, remains difficult. Just how difficult remains a factor of your development harness, e.g. your tests, CI/CD, validation environments, preview-ability of changes, and so on. While I personally don’t imagine it’s valuable for most folks at a company to be contributing code, I suspect that most disagreement about that topic is actually a miscommunication: even at a company where “everyone codes”, the marketing team isn’t reducing allocations in your servers, instead it’s about whether there is a safe boundary where they can participate. (Much like a SaaS product that allows customization by writing software.) The good news is that this means the things that were most valuable to speed up engineering two years ago are still the things that are most valuable to speed them up today. Optimize the base-case of process for agents. Most steps of most processes can be fully automated in most cases. With the right harnesses, the right controls, domain context, and good judgment in their designers, you can fully automate the base-case of most processes in modern technology companies. For example, the base case of code review from a human is slower and less effective than a good harness’ code review. Of course, the harness will miss things, but so will human reviewers, and most areas are relatively safe to make changes. Of course, there are some higher risk areas, where this doesn’t hold true. By effectively capturing these distinctions properly, we can go much faster without introducing risk. By failing to capture these distinctions, we’ll create innumerable problems for ourselves. As a corollary, I think most planning processes like weekly or bi-weekly sprints are operating at too low an altitude. Humans planning together still matters, but should be operating at a higher level. Durable, high-ownership teams with domain-context are even more important. One of my biggest lessons at Uber was that persistent, durable teams work magic by accumulating domain-context, building a sense of camaraderie, and feeling an increasingly strong sense of ownership over an area as they continue to work in it. Even in an era where specifically doing something is much cheaper, you still have to do the right thing, which has gotten a bit easier but not much easier, and structural improvements help address this. (As a recent example of that, we had an issue in production where the necessary data to optimize it simply wasn’t being captured at all, so the harness’ ideas to solve it were reasonable but wrong, since the only real path forward was instrumenting the missing information.) As a specific disagreement, there’s a prevailing idea that AI-first companies will be run by a small number of genius engineers who create perfect versions of things one by one, doing such a good job that there’s nothing to maintain. This is a very compelling vision, but I don’t see it happening. High judgment individuals can wander across a company doing remarkable things, but at some point they do get hemmed in by lack of domain context, which is why durable teams are the fundamental building block, even in this era. Quick, good, and durable decision-making is a prerequisite to meaningfully benefit from AI. Being able to replace a legal review with automation only works if Legal can commit to that change, which depends on designing the automation thoughtfully, and also the teams’ willingness to collaborate. Implementing a new feature is only valuable if you can decide to launch that feature. Your team and company can only benefit from this increased pace of execution if you can make durable decisions quickly, and those decisions are good. This is the primary reason, in my opinion, why the average CTO role has necessarily become substantially more technical and less bureaucratic than a year ago. In many cases, I am the only person who can make binding decisions when teams disagree on the path forward, and that means I am making decisions constantly in this new world in order to maintain the pace. (That’s not an argument that executives are better decision makers, just that binding executive decisions are uniquely powerful to the extent that the executives themselves are aligned enough to honor those decisions.) What have we done in practice So, I genuinely believe the above rules based on my experiences over the past year, and let me try to connect them to specific projects we’ve worked on that have convinced me of them: Migrations A year ago, we deployed manually, and deployed ~6 times a week, and now we deploy 200-400 times a week. Our engineering headcount has doubled, but even if we double the prior deploys, we’re still up 20-30x year-over-year. This is due to a complete overhaul of how we deploy and run migrations, and this migration was done over two months and done 90% by two folks on our infrastructure team. The first day of January, about 25% of folks on our team used Claude Code or Cursor every day. By the end of February, 100% did. We did this without any top-down mandate, just by making the tooling good and chatting with non-adopters to remove sources of friction. Pretty much every PR is written by harnesses now, at least in the first pass. We migrated from a large number of varied configuration mechanisms to two configuration mechanisms (one for client or server constants that rarely change, a second for product-specific or frequently changing values). This was a large series of changes, which were largely done as a series of isolated projects by individual engineers. First, one engineer cleaned up the architecture to support this approach. Then another engineer did a reference architecture on the new approach. Then several more engineers followed the reference architecture in other areas of our codebase. This might have been a years long project of many people in the prior world, but took less than a quarter to complete, including a new internal tool for managing these values across engineering and non-engineer teams. We unified a multi-repo frontend application architecture into a mono-repo frontend architecture over about a month. This was 95% driven by one frontend engineer. We now have a shared frontend development harness, can maintain libraries cheaply, and entirely moved off using npm for package hosting, which was a source of ongoing friction. We fully statically typed our frontend code, going from a place where the majority of our frontend code was not typed. This was done by one engineer, and a lot of tokens, over the course of a few weeks. We migrated from npm to pnpm for better security defaults and faster deploys. This took one engineer a few hours a day for a few days. Cost of working code depends on your development harness. Where we’ve tried to throw design documents and PRs “over the wall” to engineers on other teams, they’ve never gone anywhere. Slop pull requests and design documents are cheap, but are actively harmful. They not only have to be cleaned up and repaired, their context poisons the LLM, leading to worse outcomes than starting over. We’ve seen tremendous success in managers contributing software, as long as those managers are validating the work directly, looking at dashboards after their changes go out, and resolving any issues their changes cause. We’ve found no positive impact from folks attempting to make changes where they don’t do those things. Optimize the base-case of process for agents. We triage all incoming issues from our customer operations team using a harness which knows our team, our open tickets, and has limited access to our data warehouse to size the impact of issues. This is complex, high-skill but not particularly interesting labor that we’re now doing better and faster with agents. Yes, there is still a human triage for the edgecases. Importantly, we’re also doing this without changing human workflows, it’s the same workflow, just with some steps automated. The first pass of code review is done by the same harness that implements the changes, cleared of the context used to write the change, allowing humans to focus on higher value feedback. We rolled out Claude Code and Cowork to all folks in the company last quarter, and have seen them also automate an increasingly large swath of their work as well. Our fraud team has been particularly ambitious in replacing manual workflows with a first-pass of automation–with attribution to the data itself–to do the initial investigation on potential attacks automatically. We’ve migrated to Linear, and off Jira, to better support this workflow with a more capable MCP and better Slack integration, making it possible for everyone internally to have better infrastructure for building these agent-first workflows. More on this later, but we’re almost done alpha-testing our internal harness pulling issues off Linear, and working to resolve them, automatically which is our biggest next step in this direction. Durable, high ownership teams with domain-context are even more important. When I joined, we had a number of areas supported by very talented folks who rotated through them quickly on a per-project basis. This worked, but it meant we were very reactive to issues. Now, we’ve been able to dedicate at least a small team to every important area of the company, where they are able to persistently invest. These teams are now wielding all the new techniques afforded by AI themselves. Without them, no one would be capturing these opportunities, because there is simply too much happening. We launched SierraAI, which is quite good, but since then the team has iterated on it relentlessly, getting it truly excellent. This is something we wouldn’t have been able to do without a dedicated, focused team. Quick, good and durable decision making is a prerequisite to benefit from AI. Changing how we do configuration was a controversial decision, and I’ve had to make repeated clarifications on the approach. This would have been very difficult to do bottom-up, because it impacts every team differently, and the benefit is only experienced at the ecosystem-level (allowing one person to configure all configuration across teams). Reworking our CI/CD pipeline was controversial, as it changed many folks’ mental models of how we deploy and release (e.g., it forced us to explicitly decouple deploy and release via feature flagging). This was a contentious decision, and would have been slow and difficult to make bottom-up. Unifying into a web mono-repo was also a controversial decision with varied opinions. It benefitted greatly from having a unified decision. Moving to SierraAI was a difficult discussion versus both various competitors, and also not doing it. It needed the executive stamp to finalize the cross-functional debate. These are just representative examples, we’ve done a lot more than these. The aperture of what’s possible has continued to expand every month this year, but the things holding us back haven’t changed all that much: organizational misalignment, lack of clarity, and poor technical architecture. It’s a wild time to be working in technology.

15th Jun 2026 • 1 votes
Agents as scaffolding for recurring tasks.

One of my gifts/curses is an endless fixation with how processes can be optimized. For a brief moment early in my career, that was focused on improving how humans collaborate, but that quickly switched to figuring out how we can minimize human involvement, and eliminate human-to-human handoffs as much as possible. Lately, every time I perform a recurring task–or see someone else perform one–I think about how we might eliminate the human’s involvement entirely by introducing agents. This both has worked well, but also worked poorly, and I wanted to highlight the pattern I’ve found useful. For a concrete example, a problem that all software companies have is patching security vulnerabilities. We have that problem too, and I check our security dashboards periodically to ensure nothing has gone awry. Sometimes when I check that dashboard, I’ll notice a finding that’s precariously close to our resolution SLAs, and either fix it myself or track down the appropriate team to fix it. However, this feels like a process that shouldn’t require me checking on it. Five to six months ago, I added Github Dependabot webhooks as an input into our internal agent framework. Then I set up an agent to handle those webhooks, including filtering incoming messages down to the highest priority issues. About a month ago, when I upgraded from GPT 4.1 to GPT 5.4 with high reasoning, I noticed that it got quite good at using the Github MCP to determine the appropriate owners for a given issue, using the same variety of techniques that a human would use: looking at Codeowners files where available, looking at recent commits on the repository, and so on. The alerts and owners were already getting piped into a Slack channel. So, this worked! However, it didn’t actually work that well, because despite repeated iteration on the prompt, including numerous CRITICAL: you must... statements, it simply could not reliably restrict itself to critical severity alerts. It would also include some high severity alerts, and even the occasional medium severity alert. This is a recurring issue with using agents as drop-in software replacement: they simply are not perfect, and interrupting your colleagues requires a level of near-perfection. If I’d hired someone on our Security team to notify teams about critical alerts, and they occasionally flagged non-critical alerts, eventually someone would pop into my DMs to ask me what was going wrong. That didn’t happen here, because the knowledge that those DMs would show up prevented me from rolling the notifications out more aggressively. Coding agents address this sort of issue by running tests, typechecking, or linting, but less structured tasks are either harder or more expensive to verify. For example, I could have added an eval verifying messages didn’t mention medium or high severity tasks before allowing it to send to Slack, but I found that somewhat unsatisfying despite knowing that it would work. Instead, after some procrastination on other tasks, I finally prompted Claude to update this agent to rely on a code-driven workflow where flow-control is managed by software by default, and only cedes control to an agent where ideal. That workflow looks like: A webhook comes in from Dependabot Script extracts the severity and action (e.g. is it a new issue versus a resolved issue), and filters out low priority or non-actionable webhooks The code packages the metadata into a list of issues and repositories The code passes each repository-scoped bundle to an agent with our internal ownership skill and the Github MCP to determine appropriate folks to notify for each issue The issues and ownership data are passed to a second agent that formats them as a Slack message This works 100% of the time, while still allowing us to rely on our internal ownership skill to determine the most likely teams or individuals to notify for a given problem. It’s now something I can rollout more aggressively. The immediate fast follow was a weekly follow-up ping for open critical issues, relying on the same split of deterministic and agentic behaviors. The next improvement will be automating the generation of the vulnerability fixes, such that the human involvement is just reviewing the change before it automatically deploys. (We already do this for Dependabot generated PRs, but in my experience Dependabot can solve a reasonable subset of identified issues, but far from all of them.) That is the pattern that I’ve found effective: Prototype with agent-driven workflow until I get a feel for the workflow and what’s difficult about it Refactor agent-driven control away, increasingly relying on code-driven workflow for more and more of the solution End with a version that narrowly relies on agents for their strengths (navigating ambiguous problems like identifying code owners) This has worked well for pretty much every problem I’ve encountered. The end-result is faster, cheaper, and more maintainable. It’s also a cheap transition, generally I can take logs of some recent runs, the agent’s prompt, and some brief instructions, throw them into Codex/Claude, and get a working replacement in a few minutes.

12th Apr 2026 • 1 votes

More in programming

Clip of me singing Despard in Ruddigore in 2013

A clip of me singing a funny song from Gilbert and Sullivan’s Ruddigore back in 2013

5 hours ago • 1 votes
How and Why fork() Uses Copy-on-Write

In this video, we look at why fork() needs copy-on-write, how it works inside the kernel, and a memory usage problem that Instagram encountered with Python.

12 hours ago • 1 votes
What we lost when we lost comments

Comments require commitment, but they’re worth it.

18 hours ago • 1 votes
Lighthouse map

Lovely global map with animated lights sweeping the waters

20 hours ago • 1 votes
Warming up the Puma master before it forks

Basecamp 5 runs on Puma in cluster mode: one master process with preload_app! and 63 single-threaded workers per host, deployed as a Docker container with Kamal. We serve Basecamp from several sites. Each site has its own web hosts and a read replica of the database, and writes go to a single primary database in one of them. On our busiest hosts, each deploy left up to 2,000 requests waiting while the new workers warmed up. We reduced those queues by running signed-in requests through the app in the Puma master, before it forked the workers. Why 63 single-threaded workers? Basecamp has always served web requests from processes rather than threads. It ran on Unicorn, which only does processes, until we moved to Puma in January 2025, and we kept the same setup: workers (Concurrent.physical_processor_count * 1.3).ceil threads 1, 1 preload_app! On a 48-core host that’s 63 workers, each handling one request at a time. We chose 1.3 after benchmarking HEY in 2023, when we moved our apps out of the cloud and onto our own hardware. We tested several combinations of workers and threads with a mix of GET and POST requests on a 32-vCPU VM. Every multithreaded configuration we tested was slower and handled fewer requests than single-threaded workers. Adding workers beyond about 1.2 to 1.3 per vCPU brought little benefit. The threaded workers spent a lot of their time waiting for Ruby’s global VM lock. That made single-threaded workers a good fit for this workload, and we use the same setup for Basecamp. An app that spends more time waiting on its database or other services may benefit from more threads, so benchmark your own app. The other reason is the app itself. Basecamp has class-level state in places and has never needed to be thread-safe. With one request per process, it still doesn’t. Processes do use more memory than threads, and preload_app! reduces the difference. The master loads the app once and the workers share its memory through copy-on-write until they write to it. Shopify’s comparison of Ruby execution models explains the trade-off well. In the HEY benchmark the best setup came to about 260 MB of PSS per core, where PSS counts each shared page once, split between the processes using it, and the gap to a threaded setup was smaller than we’d expected. What Puma does on each host when a container starts: one master, then 63 forked workers that share its memory until they write to it. Two things about this setup matter for the rest of the post. A worker that’s compiling or loading something is fully blocked — there’s no other thread to pick up the next request. And whatever the master has in memory before it forks, all 63 workers share. Whatever they build after the fork, they build 63 times. What happens when we deploy Kamal starts the new container alongside the old one, and kamal-proxy moves the host’s traffic across as soon as the health check passes. At that moment, the new workers have handled health checks but no customer requests. preload_app! means the master loads the app once and the workers inherit it through fork. That covers the code. It doesn’t cover anything Ruby and Rails set up on first use: YJIT compiled code. YJIT compiles a method once it’s been called a certain number of times. The master calls very little during boot, so every worker compiles the same methods again on its own first requests. Compiled templates. Action View turns each ERB template into a Ruby method the first time it’s rendered. The schema cache. Active Record reads each model’s columns from the database the first time that model is used. Inline caches and memoized values throughout Ruby, Rails and the app. All 63 workers did all of this at once, while serving the traffic the old container had been handling a second earlier. In the test environment with YJIT on, the first request to a project page on a cold process took 652 ms, 151 ms of it YJIT compiling. The same request to a warm process took 28 ms. In production, CPU time per request peaked at around 200 ms while kamal-proxy moved traffic to the new container, against about 30 ms once the workers had warmed up. A host with spare CPU absorbs this. Every one of our web hosts has 48 cores and 63 workers, but each Amsterdam host serves around 250 requests per second, against 25 to 60 at our other sites. In Amsterdam the slow first requests turned into a queue. At a peak-hour deploy, the Puma backlog on an Amsterdam host reached anywhere from 250 to 2,238 requests, and kamal-proxy’s p99 response time hit about 10 seconds. Eron, our Director of Operations, had been tracking this since June. Another server in Amsterdam would help, but it would take weeks to arrive, so we also wanted to make deploys cheaper on the hardware we already had. What didn’t work We tried a few things first. In June, Donal tested the first two on a single Amsterdam host, comparing it with its neighbors, and they ruled out two likely causes. Warming each worker’s database connections. Puma’s before_fork hook clears the master’s connections, and each worker opened its own on its first request. Opening them in before_worker_boot instead made no difference. Queries on a freshly booted production host were already under a millisecond, so connections weren’t the problem. A synthetic request in each worker. Next, each worker made a few requests in before_worker_boot to an internal controller that touched every model. That ran the middleware, routing and Active Record paths, but it ran them in 63 workers at once — exactly the CPU spike we were trying to avoid. And a request with no real data renders no real views, so most of the app stayed cold. Spreading YJIT compilation out. Delaying YJIT in each worker by a random interval spread the compiling out over a few minutes, but every worker still ran interpreted until its delay ended. The queue didn’t change. Reforking from a warm worker. This is what Shopify’s Pitchfork does: let one worker serve traffic until it’s warm, then fork the others from it. Puma has an experimental version called fork_worker, and on beta it worked — the reforked workers were warm after three to five requests, where fresh ones took up to 30 seconds. But with fork_worker the template is worker 0, and it keeps serving requests. If it exits, the workers waiting to be forked never start (puma/puma#3596). If it gets no traffic, the refork never happens, which is what we saw on beta. Instacart have a mold_worker patch that promotes a warm worker to a template that stops serving, but it isn’t in a Puma release. We have a branch of it, and we may come back to it. That last experiment did show us where the fix was, though. Everything a warm worker has that a cold one lacks is in its memory, and fork copies memory. The master already has the app loaded. It just never runs it. Run the requests in the master So now, before the master binds its socket and forks, it makes the app’s own requests, in-process, the way a signed-in user would. Rack has a hook for exactly this. Rack::Builder#warmup takes a block that’s called once with the built app, before the server starts. rails server builds the app from config.ru, so the change to boot is one line: require_relative "config/environment" warmup { WarmUp.configured.run } if ENV["WARM_UP"] run Rails.application With preload_app! this runs in the master, and the workers inherit whatever it did. Puma binds its socket after the app is built, so until the warm-up finishes the health check’s connection is refused and kamal-proxy keeps retrying. No request reaches a worker that hasn’t been warmed. The warm-up has three steps. After precompiling the views, it gives the page requests and schema loading a shared 20-second budget, checked before each page or model. 1. Precompile the views actionview_precompiler reads every template for its render calls and compiles each one with the locals it’s passed. For us that’s 1,394 templates in about two seconds. A first request to a project page then compiles 2 templates instead of 44. 2. Request the pages, signed in A small browser class makes the requests through Rack::MockRequest, with the two cookies a real sign-in sets, then goes back for each page’s lazy Turbo frames: class WarmUp::Browser def initialize(signed_in_as:) @client = Rack::MockRequest.new(Rails.application) @headers = { "HTTP_USER_AGENT" => "Basecamp warm-up", "HTTP_COOKIE" => cookie_for(signed_in_as), "bc3.warm_up" => true } end def visit(path) page = get(path) frames_in(page).each { |id, src| get(src, "HTTP_TURBO_FRAME" => id) } end private def get(path, headers = {}) @client.get("https://#{host}#{path}", @headers.merge(headers)) end def frames_in(page) Nokogiri::HTML5(page.body).css("turbo-frame[src]").map { |frame| [ frame["id"], frame["src"] ] } end end The requests are signed in. The user is a monitoring account we already use for automated checks, and the pages are its own project, Campfire, to-dos, documents and messages. Public pages weren’t enough: after warming up with signed-out pages only, the first signed-in request to the projects page still took 131 ms, because authentication, the signed-in controllers and their views had never run. With signed-in pages it took 40 ms. cookie_for writes the same signed cookie the sign-in controller does, using the app’s own cookie jar, so there’s no API token and no secret to store. The frames are followed. The busiest HTML requests in production aren’t pages at all but Turbo frames — the sidebar badge, the inbox, the navigation menus. The browser parses each page and requests its <turbo-frame src> URLs with the Turbo-Frame header, so those controllers and views get warmed too. Our first four pages turned into 60 requests. The requests are excluded from rate limiting. They are internal, so they do not count against the rate limits that apply to real visitors. 3. Load the rest of the schema The page requests load the schema for the models they touch. The last step loads the rest, from the read replica: ApplicationRecord.reading do models.lazy.take_while { time_left? }.each { |model| model.load_schema if model.table_exists? } end The step checks 261 models and loads any schema information still missing. Those database round trips add up when the primary is far away: outside a request, Active Record uses the writing role, and from a host a long way from the primary each round trip is tens of milliseconds. Reading from the local replica brings the step down from about 20 seconds to 3.5. The pages go first because they load most of the schema anyway. If the time budget runs out, the step stops, logs how many models it got through, and the workers load the rest on first use like they always did. Rails can also load the schema from a dumped cache file at boot (bin/rails db:schema:cache:dump), which would make this step unnecessary. We don’t ship one in our image yet, because the dump needs a database to read from at build time, and we have several databases to cover. It’s on the list. What to close before the fork Running requests in the master opens things the master never opened before, and every worker inherits them. Two processes writing to the same socket will corrupt each other’s traffic, so you need to know what’s open before you fork. The way to find out is to list the master’s open file descriptors — ls -l /proc/<pid>/fd — before and after a warm-up, in an environment set up like production. Development wasn’t enough for us: it stores files on disk, so our S3 connections only showed up in production. Then, for each thing that’s open, check how its library handles a fork. We found three kinds: Already handled. Plenty of libraries detect a fork on their own, either by recording the PID they connected from and reconnecting in the child, by opening per-process files, or by resetting their thread pools. Redis clients, metrics libraries and concurrency libraries tend to be in this group. Check, but you probably don’t need to do anything. Already closed. Database connections are the classic one, and most Puma configs already clear them in before_fork. Anything else that’s opened per process — we have a SQLite cache the workers open on boot — needs closing when the warm-up finishes. Needs a new step. HTTP clients with keep-alive connections are the ones to look for: cloud SDKs with connection pools, tracing exporters, error reporters. They usually have no fork handling at all. We empty the aws-sdk connection pools in before_fork, and we run the warm-up untraced so the OpenTelemetry exporter never opens its connection to Tempo in the first place. Once that’s done, before_fork finishes with Process.warmup, which Ruby 3.3 added for this purpose: a major GC, a heap compaction, and every surviving object promoted to the old generation, so the memory pages the workers share change as little as possible afterwards. Choosing the pages The first list was the four pages that ran the busiest requests on beta. Once the warm-up was live, production showed us which endpoints were still cold. For one deploy, we compared each endpoint’s mean duration in the six minutes after kamal-proxy moved traffic to the new container with the same endpoint an hour later, then multiplied the difference by the number of requests in those six minutes. That gives the extra time each endpoint cost us because it was cold: Endpoint Cold Warm Requests in 6 min Extra seconds Campfire 246 ms 70 ms 6,490 1,140 Projects (JSON API) 84 ms 50 ms 22,077 771 Docs & Files 262 ms 177 ms 4,996 421 To-dos tool 205 ms 113 ms 4,018 371 To-dos (JSON API) 33 ms 16 ms 18,738 320 The pages already in the warm-up showed what to expect: the project page kept a 36 ms gap after a deploy, and the to-do page 10 ms. We’ve proposed adding these five requests, and expect them to add about five to seven seconds to the page step. The two JSON endpoints were a surprise. The warm-up’s page list had no API requests in it, so nothing on the API path had run before the first real request: not the API controllers, and not the Jbuilder templates rendering real records. Precompiling the views covers JSON templates too, but it isn’t a substitute for running the request. Results The warm-up is on for all 68 web hosts. With the first four pages it took 12 to 16 seconds per host: about 2 seconds to precompile the views, 7 to 9 for the 60 requests, and 3.5 for the schema. Deploys take that much longer per host, and we raised the deploy timeout from 30 to 60 seconds to cover it. In Amsterdam, at a peak-hour deploy: During deploy Before After Peak Puma backlog per host 250–2,238 requests 19–223 requests Peak kamal-proxy p99 about 10 s 2.4–4.8 s Peak CPU time per request 201–214 ms 88–132 ms Peak database time per request 56–69 ms 39–47 ms The same eight hosts at three deploys on 1 October, an hour apart, as the warm-up went from one host to four to all eight. The deploy in the middle, with four hosts warmed and four not, shows why every host needed the warm-up. Each warmed host recovered faster on its own: mean request duration peaked at 130 to 173 ms, against 203 to 311 ms on the hosts that weren’t warmed. But the backlog on all eight was about the same, because they were all waiting on the same database. Mean request duration on each host at the 07:21 UTC deploy. Blue hosts warmed up in the master before forking, orange hosts did not. Memory came down too. The workers now share compiled templates, YJIT code and the schema with the master instead of each building their own copy. On beta, the view precompiler alone took a busy worker’s private memory from 174–202 MB to 119–135 MB. Thirty minutes after the deploy, the web containers used about 39 GB less memory than the previous day’s containers at the same age and traffic. Amsterdam served most of our traffic at the times we tested. In Amsterdam, each new container used about 2 GB less just after traffic moved to it, which lowers the peak while the old and new containers overlap. Working with Claude Claude Code helped throughout. It combed through the per-worker backlogs and per-endpoint timings in Prometheus and Loki after each deploy, worked out the cold-versus-warm cost of each endpoint, and prepared the changes and the pull request descriptions with the benchmarks in them. We decided what to try, deployed it and read the results. If you do this Warm the master before it forks. Compile common code and templates and load their schema in the master, so workers inherit that work. With preload_app!, Rack::Builder#warmup runs before the workers start accepting traffic. Use the app’s real requests. Public pages, internal endpoints and synthetic queries warm the paths they run and nothing else. Signed-in requests to real records, frames included, run what production runs. Measure the cold penalty per endpoint. The difference between an endpoint’s cold and warm duration, times its request count after a deploy, ranks the pages worth adding. Ours weren’t the ones we’d have guessed, and two of them were JSON. Check what the warm-up leaves open. List the master’s file descriptors after a warm-up and account for every one before the fork. Two of ours needed changes. Set a time budget. A warm-up that runs long on one slow host fails the deploy on that host. Ours gives the page requests and schema loading a shared 20-second budget, checked before each page or model, puts the most valuable pages first, and logs what it skipped. Reforking from a warm worker, as Pitchfork does, solves the same problem continuously rather than once at boot, and it would warm paths no fixed list of pages covers. We may still get there: our branch brings Instacart’s mold_worker up to date with Puma’s main branch and fixes the bugs we found in it. But warming the master works with the Puma we already run, took a few days to implement, and substantially reduced the queues after deployment.

yesterday • 1 votes
📚 BoredReading

You seem to be enjoying this.

Join free to unlock everything.

Create free account

Already have an account? Sign in