More from On Test Automation
This post was previously published through my newsletter on May 25, 2026. From time to time, I will republish newsletter issues on my blog here if I think people (and search engines) might benefit from it. If you want to read everything I’ve written once it is posted, I recommend signing up for the newsletter. A few months ago, I had the pleasure of delivering a keynote at an internal developer conference for one of the largest banks in the Netherlands. While I still don’t really see myself as a keynote speaker - I enjoy doing practical, hands-on sessions much more - I had a great time talking about challenges of E2E testing, breaking down E2E tests and the test automation quadrant model that I use in my thinking, speaking and teaching about test automation these days. However, the keynote or the contents of it are not what I want to talk about this week. Instead, it was a question that came up during the Q&A after the talk that triggered me to write this post. That question was “Do you think that testers should be writing unit tests?” It’s not the first time I heard that question. In fact, I’ve seen and heard whether or not testers should be involved in unit testing being discussed regularly in the past, with arguments for and against the various standpoints. However, I was under the impression that we, collectively, had found some sort of answer to the question and moved past this point by now. Guess I was mistaken. I tried to give the person asking the question an answer as well as I could, but given that there is quite a bit to unpack around the topic, I’m not sure if I gave them the entire story. So, that’s what I’ll try and do here. Who knows they might even read it… So, should testers be writing unit tests? Well, my answer is either a ‘yes’ or a ‘no’, depending on how you interpret the question. Before we explore these various interpretations, what is a unit test anyway? Well, by now, I don’t really know anymore, as there are so many definitions floating around, some of them contradicting each other. This is one of the reasons I came up with the test automation quadrant model as an alternative to the well-known automation pyramid model, but again, that’s not the topic of this post. For the sake of the argument, let’s define a unit test as a test that verifies a very small piece of behaviour of our product, without relying on external interfaces like APIs, databases or file systems, for example. With that definition in place, let’s look at two different interpretations of the question of ‘should testers be writing unit tests?’. Interpretation 1: Testers should be responsible for unit testing Well, no, I don’t think they should, no matter how good of a tester they are, and no matter what their coding skills are. Writing unit tests is an activity that should be performed in support of and lockstep with software development. Often, especially in practices like test-driven development, tests are written first, and they drive the design and development of the product. Leaving the writing of these tests to testers, especially if it is done after the product code itself is written, is both inefficient and a potential source of problems. Inefficient, because the product has already been written, yet we only know whether its behaviour matches expectations once the tests are written and run. Also, because there’s a handoff happening between ‘development’ and ‘testing’, which takes up valuable time as the developer will switch to a different task while the tester writes the unit tests. If there’s a problem with the product that is discovered during unit testing, the developer needs to make a context switch back to the original task, and the more context switching you do during the day, the less efficiently you will work and the less you will get done. A potential source of problems, because when the developer only focuses on shipping a potentially working product, they will likely not spend too much time thinking about what to test for, or how to make the product (the code, in this case) easily testable in the first place. Also, the handoff from ‘development’ to ‘test’ I described before leaves open room for different interpretations of what the software should do, leading to potential bugs slipping through and the resulting back-and-forth discussions after the fact. So, no, I don’t think we should leave unit testing to testers. I still sometimes hear about developer who don’t want to or do not know how to write (decent) unit tests, and I think that’s a problem that needs to be fixed at the source, instead of trying to patch it up by someone else writing the unit tests for the already-created product. Interpretation 2: testers should be involved in unit testing If we interpret the question this way, I think the answer should be a resounding ‘yes’, testers should be involved in unit testing. I mean, there’s the word ‘testing’ in the name, why shouldn’t we involve the people who specialize in testing in the process? What that involvement looks like, exactly, depends on the context, of course. Again, ‘being involved’ and ‘being responsible’ are two entirely different things. While I believe that writing unit tests is a development activity, there are a couple of ways in which testers can add value, too: They can review the unit tests to learn about what has been covered already, so that they do not repeat that testing later on They can review the unit tests to identify what has not yet been covered, and either give that back as feedback to a developer, or add the missing tests themselves They can suggest and use techniques like mutation testing to test the tests and find out whether the tests that were written before are actually able to catch meaningful problems I’m sure there are a few more benefits, but these alone should, in my opinion, be enough for any team to not exclude testers from the process of writing unit tests from here on. And yes, that will require some additional skills from both testers and developers. This is where the power of collaboration comes in. I’m a big fan of pair programming and testing, and I’ve seen a lot of good things come from testers and developers pairing up to write, review and improve unit tests. The developer improves their testing skills, the tester learns more about both the product they’re testing and the development process, and in the end, both the entire team and the product itself reaps the benefits. That’s a long answer to a simple question, and I’m sure there are some nuances that I didn’t yet unpack, but generally speaking, these are my views on the question of whether testers should be writing unit tests. Oh, and before you ask ‘but what about integration / end-to-end / performance / security / … tests’? That’s easy. Simply replace ‘unit tests’ and ‘unit testing’ with ‘integration / end-to-end / performance / security / … tests’ and ‘… testing’, and you’ll have my answer. I’ve always found it a little strange that unit tests have traditionally been seen as ‘different’ from other types of tests, as if they’re some kind of special artefact that only developers know how to write. If that’s what you think, too, let me let you in on a little secret: they aren’t special. Unit tests are simply tests that test a small piece of the behaviour of the product that we write, written against and invoking a specific interface of the product: the source code. Other types of tests do exactly the same thing: verify parts of our product behaviour by invoking one or more specific interfaces (APIs, the UI, a database, a queue, …). They just have a different scope, and with that, they verify behaviour at a different scope. That’s all. This, again, is the reason where I think the traditional test automation pyramid model lacks: I really don’t care that much about what is a unit / integration / end-to-end test. All I really care about is tests that produce valuable information about the state and the behaviour of our product in as efficient a way as possible. Unit tests really aren’t any different.
Just a quick update to let those of you who bookmarked this blog or who have subscribed to my RSS feed know that I have (re-)started a newsletter. Why a newsletter? As you might know (or not), while I’ve been pretty active on LinkedIn over the years, I do have a love-hate (or rather an appreciate-hate) relationship with that platform. Lately, I’ve been noticing that the pendulum is swinging in the ‘hate’ direction more often, mainly because the ever-changing algorithm used by LinkedIn makes it incredibly hard to predict if people are even going to see what I write. I’d rather publish my thoughts, ideas and other ramblings via a platform that I do control, and that platform will be a newsletter. I’ve had a newsletter in the past, but that only lived for about three months. This time, I intend to keep writing and publishing a new issue every week. The first edition goes out a few hours after I’m writing this, and a new issue will be sent to subscribers every Monday morning around 11 AM CET. But what about the blog? I’ll still publish to the blog, too, but that will be on a much less regular basis. Just like it has been for a while, really. The idea is to post the more ‘technical’ posts, that is, the ones including code, directly to my blog, whereas the ‘text-and-images-only’ posts go through my blog post first. My priority is with the newsletter, though. How to subscribe That’s easy, just go to the subscription page, leave your email address, click the button on the confirmation email and you’re in. I promise I won’t use the newsletter or your email to spam or sell to you. Ever.
When I talk about the goals and the purpose of test automation, I often use the phrase ‘valuable feedback, fast’: we use tools to support our testing to help us get valuable information about the state of our product in the most efficient manner possible. The ‘fast’ part of ‘valuable feedback, fast’ is pretty self-explanatory for most people: as build and release cycles are becoming shorter, teams want to be informed timely about any unexpected changes in behaviour of their product, often after every change they make to that product. Tools can help them achieve that by running quick, focused tests automatically when a change is made or committed to version control. Of course, it takes plenty of hard work to write those tests to be fast, but that’s not what I wanted to talk about here. The ‘valuable’ in ‘valuable feedback, fast’ is a much more ambiguous term, and one that deserves some more explanation. To me, there are multiple dimensions to what makes a test valuable, and in this post, I want to unpack and address them one by one. Valuable = important to someone who matters Borrowing from the classic definition of ‘quality’ as defined by Jerry Weinberg and further refined by James Bach and Michael Bolton, this is where it all starts. The information presented by a test should be important to someone who matters. That someone could be a member of the development team, a stakeholder such as a product owner or business analyst, the end user of the product, or a combination of those. Without that importance, a test is meaningless, dead weight. It could be the most reliable, best-written test ever, but if the information that is provided by it is not important to someone who matters in the context of the product, why bother writing, running and maintaining the test? Valuable = covering what matters Test coverage is a tricky subject, and I want to steer clear of the discussion on what ‘coverage’ means exactly in this blog post. The only realistic answer is ‘it depends’, anyway, as there are so many ways to define coverage (line, branch, requirements, mutation, …). Having said that, for the information provided by our tests to be valuable, teams should invest time in making sure that the tests cover the parts of the product behaviour that are deemed ‘important enough’ in a sufficient manner. What exactly constitutes ‘sufficient’ here depends on, you guessed it, the context. Some products require deeper, more thorough coverage than others. The same applies to individual parts of the same product. It all depends on the acceptable amount of risk a team is willing to take before putting a product in the hands of their users. Teams would do well to have a continual discussion about these risks and the extent to which they are covered by the tests that accompany and scrutinize the product. Valuable = trustworthy The higher the degree of automation in the build and delivery process of a product, and that includes testing, the more teams will rely (and have to rely) on the results of the execution of that automation. Concerning tests, that means that teams need to be able to rely on the information presented by the tests, because they will make decisions based on that information. The nature of that decision might vary from anywhere between ‘this build seems sufficiently stable to warrant deeper testing’ to ‘this change is ready to be put in the hands of our users’. No matter what the specific decision is, if teams make it based on the results of your test automation, even in part, they can only confidently do so if the information provided by the tests is trustworthy. In practice, that means that when a test emits a signal indicating a problem with the product, the team can safely conclude that there is a problem with the product, not with the test, the data it uses or the environment it runs in (no false positives). It also means that when a test does not emit such a signal, the team can trust that the particular piece of behaviour exercised by the test is working according to expectations expressed in the test (no false negatives). Valuable = actionable Another dimension of the value of the feedback provided by a test is that it should be actionable. This applies specifically to those situations where a test ‘fails’, i.e., it indicates a problem with the product I have put ‘fails’ between quotes here, because the test didn’t fail, the product failed the test. There’s a difference. Anyway, when a test result indicates a (potential) problem with the product, teams need to able to act on that information as soon as possible, spending as little time digging deeper into the product or into the test as possible to identify the root cause of the problem. Some practices that might help here are: Making your test scope as small as possible - the fewer moving parts your test has, the easier it will be to identify which of those parts made a move that was unexpected Have good test names - A descriptive test name that tells you what part of the behaviour your product verifies and what the expected behaviour is helps in finding out where exactly the problem might be found Use custom assertion messages - Many test frameworks allow you to specify custom, descriptive error messages in case of assertion failures (something RestAssured.Net supports as of version 5.0.0, too) So, is this a complete and final definition of what ‘valuable’ means to me when I talk about ‘valuable feedback, fast’ as the goal of test automation? I don’t think so. I don’t know if it is complete, but it definitely is a good reflection of my current thoughts on ‘value’ in test automation right now. Those thoughts are definitely not ‘final’, and I would appreciate your takes on what I wrote here.
In a recent post, I wrote about how I used Claude Code to analyze the code for RestAssured.Net and then perform a refactoring action, using hand-written tests as the safety net. In that post, I wrote that I didn’t want Claude to touch the tests themselves, and why. I was still curious, though, to find out for myself what Claude was capable of in terms of writing tests. In this blog post, I’ll share with you some first steps in doing exactly that, and you’ll read about my thoughts and my thought process along the way. You’ll see how I create an initial suite of tests for a small Spring Boot-based API that I wrote for use in my workshops, and how I think about and assess the results. In a follow-up blog post, I’ll show you how I improved the test suite based on my findings, again using Claude Code. The starting point As a starting point, I created a new repository containing the code for the API I use in my mutation testing workshop. I removed the existing tests, as we’re going to ask Claude to generate these for us. I also removed the README and the GitHub Actions build pipeline definition, as I want Claude to write tests based only on the product code itself, without being primed by other artifacts in the codebase. The only thing I left in are the dependencies used to write and run the tests, in this case REST Assured and JUnit. After installing and initializing Claude, I gave it a first prompt: “Add acceptance tests for the endpoints exposed by the AccountController to this project. Cover all the logic in the AccountService class. Use REST Assured as the tool to interact with the API. Use JUnit 5 as the test runner. Both libraries are already part of the project, see the pom.xml. Assert status codes and relevant response body elements as part of the tests. Extract common request properties into a RequestSpecification.” After some deliberation, Claude added a new test file to the project, containing 23 tests, all of them passing. You can see these tests here. What you’re seeing in this file is the raw output from the above prompt, I haven’t changed anything in there. It took Claude only a minute or two to write these tests, which definitely is a lot faster than what I could have done myself. But how good are they, really? A first look at the tests Let’s look at the quality of the code first. I’m seeing people argue that code quality is not really all that important anymore once AI will write most of our code, but I beg to differ, especially when it concerns our tests. Tests are documentation of the intended behaviour of our code, and I would say that being able to read that documentation as a human being, without too much effort, remains very important. So, is our code easy to read? There’s a @BeforeEach hook creating the RequestSpecification (an object in REST Assured containing shared HTTP request properties). There’s a helper method to create a new account passing in the AccountType and a predefined balance. There’s the aforementioned 23 tests that, especially at first glance, seem to verify things that are valuable. What Claude did not do, probably because I didn’t explicitly ask for it, is add an abstraction layer to make the code easier to read, such as the one described here. We’ll see how Claude does in this area in the next blog post, as I want to stick to assessing the quality of the initial output from Claude in this one. And I have to say, all in all, for a first try, I’m not unhappy with what I’m seeing. Yes, there’s room for improvement, but I have seen humans do far worse than this. The tests seem to cover all endpoints defined in the API controller, and most paths in the business logic defined in the service layer. I should note here that I was able to fairly quickly come to this conclusion only because: I wrote the code for the API, so I have knowledge of the inner workings and the intent of the API, and I have plenty of experience writing tests for APIs and writing tests in REST Assured, so I’d like to think I know what ‘good’ looks like If you don’t have that prior knowledge and experience, it will be harder to draw meaningful conclusions from just looking at what Claude coughs up. And there’s a significant risk there: the risk of saying ‘looks good to me’ without actually understanding what you’re approving, and then ending up with a safety net of tests that is riddled with holes. Testing the generated tests with mutation testing To further increase our understanding of the value of the tests that were generated for me, let’s see if these tests can fail. If they can’t, the fact that we have generated 23 passing tests in two minutes flat is nothing more than an example of productivity theater. To check if our tests can actually fail, let’s use a mutation testing tool to scrutinize our tests a little more. In this case, because we’re working with Java code, I’ll use PITest as my mutation testing tool of choice. I configured the tool to mutate all the code in the project and run all the tests, to get a complete overview of the quality of the test suite generated. Note that in a real life-sized project, you probably want to start by mutating only part of the code base and run part of the tests to get mutation testing feedback within a reasonable amount of time. After about a minute, PITest reports back that the initial test suite achieves 95% line coverage. This looks impressive, but it doesn’t really tell me anything. The much more valuable metric here is the number of mutants that were killed by the test suite. PITest reports that this is 91%, which, again is pretty good. In absolute numbers, out of 55 mutants generated by PITest, 50 were detected by the initial test suite. Two follow-up questions arise immediately: Which mutants were missed by the tests, and what is the impact of that? Could we have achieved the same amount of (line and mutation) coverage with fewer tests? In other words, do we have tests that are dead weight? Looking at the surviving mutants First, let’s have a look at the mutants that survived, i.e., changes in the API code that were not detected by any of the tests. To start, in the CustomizedResponseEntityExceptionHandler, the HTTP 500 path isn’t covered in any of the tests, and that causes a surviving mutant. By design, the API returns an HTTP 500 when an Exception occurs that isn’t a ResourceNotFoundException (returning an HTTP 404) or a BadRequestException (returning an HTTP 400). This looks like a useful path to cover in a test. Second, the API returns an HTTP 204 in response to a GET call to /accounts when there are no accounts in the database. That path isn’t covered in the tests. This, too, seems like a useful path to test, because it is intentional API behaviour. Finally, the tests that were written do not properly cover some of the boundary values, both in the logic that implements the business rule of ‘you cannot overdraw on a savings account’ and in the interest calculation logic. Once more, I would like to have these situations covered by tests. Coincidentally (or maybe not?), these are all cases that I cover in my mutation testing workshop, too. This, to me, indicates that mutation testing is a powerful way to assess what is tested and what isn’t, no matter if you wrote the tests or you had them write by an LLM. I’m also happy to see that I’m probably covering the right things in my workshop. Note: I can confidently and quickly perform this analysis of the signals produced by PITest, and of the quality of my tests, because I know that mutation testing as a technique exists, and because I know how it works. Most importantly, I’m motivated / I feel like I am morally obliged to do so, because I deeply value writing tests that test meaningful things and that are actually able to detect changes in product behaviour. If all I cared about was having some tests to cover the API and declared, for example, 90% line coverage as ‘good enough’, I would be done by now. However, I don’t. In the next blog post, I want to return this feedback to Claude and see how well it does in updating the existing test suite based on my observations. I also want to see if I can add mutation testing to the test generation loop, and have Claude achieve better mutation coverage without my interfering. For now, I’ll conclude that when I ask Claude to generate tests in the way I have done, it produces pretty good results in terms of both line and mutation coverage, but that it missed certain key paths in my application code. Identifying dead weight in our test suite As a next step, I want to find out if the test suite that was generated by Claude contains dead weight, that is, do we have any tests that do not uniquely contribute to either line or mutation coverage? To do so, I asked PITest to generate a report in XML format next to the HTML report, as (for some reason) only the XML report contains information about which test killed a specific mutant. Performing this analysis required a bit of elbow grease, as I had to manually search the XML test report for occurrences of the test name for every test in the test suite. This, too, is probably a process that can be automated, but for now, I’m OK with doing this the manual way, since there’s only 23 tests in the suite anyway. This search tells me that four tests that were generated by Claude did were not mentioned as a test killing a mutant in the results file. In all four cases, the reason behind this is that the exact same code path is exercised in another test. For example, one of the tests performs a withdrawal on a checking account and verifies that the balance is updated accordingly: @Test void withdraw_positiveAmount_fromCheckingAccount_updatesBalance() { long id = createAccount(AccountType.CHECKING, 500.0); given(requestSpec) .post("/{id}/withdraw/{amount}", id, 200.0) .then() .statusCode(200) .body("balance", equalTo(300.0f)); } The next test in the suite, however, does the exact same thing for a savings account: @Test void withdraw_positiveAmount_fromSavingsAccount_withSufficientFunds_updatesBalance() { long id = createAccount(AccountType.SAVINGS, 500.0); given(requestSpec) .post("/{id}/withdraw/{amount}", id, 200.0) .then() .statusCode(200) .body("balance", equalTo(300.0f)); } After removing these four tests from the suite and running mutation testing again, as expected, I can see that the impact on both line and mutation coverage is 0, meaning that these four tests can indeed be classified as ‘dead weight’. Conclusions So, after completing the analysis of the results of asking Claude Code to generate tests for a new code base, what do I think? Well, while I am impressed, I think a couple of words of warning are in order. I am positively surprised by the quality and the coverage of the initial test suite. 95% line coverage and 91% mutation coverage are good numbers, and all that coverage was generated in a few minutes, definitely a lot less time than it would have taken me to write these tests myself. There is some room for improvement in terms of readability of the tests, but that can probably be resolved by being more specific in my prompt and / or using dedicated Claude Code skills. I’ll explore and write about that soon. While Claude achieved a pretty decent mutation coverage, it did oversee a few critical paths in the code. Maybe I was simply ‘unlucky’, and another attempt with the same prompt would have given better results. I don’t know, but it does tell me not to simply accept what Claude gives me at face value. The same applies to the tests that Claude did generate. 4 out of the 23 tests generated were dead weight, which equates to 17% of the test suite. Now, n = 1, and this is a small codebase and test suite, so the numbers might be skewed, but again, if you want your test suite to be as efficient and effective as possible, these are numbers that you probably don’t want to ignore. Finally, there are of course many things that Claude did not do, mainly because I didn’t ask it to. An example of that would be telling me that since we’re working with a banking API, it probably would be a good idea to add some form of authentication to the endpoints. There’s a lot more to unpack about what Claude does and does not do, and I will probably write about that in more detail in another blog post, but not here. First, in a follow-up blog post, I’ll document the process of improving the existing test suite that Claude generated, both in terms of coverage and of coding style. I will once again be using Claude and mutation testing to do that. The code for the API that was used in this blog post, as well as the initial suite of tests generated by Claude, can be found here.
More in programming
A clip of me singing a funny song from Gilbert and Sullivan’s Ruddigore back in 2013
In this video, we look at why fork() needs copy-on-write, how it works inside the kernel, and a memory usage problem that Instagram encountered with Python.
Comments require commitment, but they’re worth it.
Basecamp 5 runs on Puma in cluster mode: one master process with preload_app! and 63 single-threaded workers per host, deployed as a Docker container with Kamal. We serve Basecamp from several sites. Each site has its own web hosts and a read replica of the database, and writes go to a single primary database in one of them. On our busiest hosts, each deploy left up to 2,000 requests waiting while the new workers warmed up. We reduced those queues by running signed-in requests through the app in the Puma master, before it forked the workers. Why 63 single-threaded workers? Basecamp has always served web requests from processes rather than threads. It ran on Unicorn, which only does processes, until we moved to Puma in January 2025, and we kept the same setup: workers (Concurrent.physical_processor_count * 1.3).ceil threads 1, 1 preload_app! On a 48-core host that’s 63 workers, each handling one request at a time. We chose 1.3 after benchmarking HEY in 2023, when we moved our apps out of the cloud and onto our own hardware. We tested several combinations of workers and threads with a mix of GET and POST requests on a 32-vCPU VM. Every multithreaded configuration we tested was slower and handled fewer requests than single-threaded workers. Adding workers beyond about 1.2 to 1.3 per vCPU brought little benefit. The threaded workers spent a lot of their time waiting for Ruby’s global VM lock. That made single-threaded workers a good fit for this workload, and we use the same setup for Basecamp. An app that spends more time waiting on its database or other services may benefit from more threads, so benchmark your own app. The other reason is the app itself. Basecamp has class-level state in places and has never needed to be thread-safe. With one request per process, it still doesn’t. Processes do use more memory than threads, and preload_app! reduces the difference. The master loads the app once and the workers share its memory through copy-on-write until they write to it. Shopify’s comparison of Ruby execution models explains the trade-off well. In the HEY benchmark the best setup came to about 260 MB of PSS per core, where PSS counts each shared page once, split between the processes using it, and the gap to a threaded setup was smaller than we’d expected. What Puma does on each host when a container starts: one master, then 63 forked workers that share its memory until they write to it. Two things about this setup matter for the rest of the post. A worker that’s compiling or loading something is fully blocked — there’s no other thread to pick up the next request. And whatever the master has in memory before it forks, all 63 workers share. Whatever they build after the fork, they build 63 times. What happens when we deploy Kamal starts the new container alongside the old one, and kamal-proxy moves the host’s traffic across as soon as the health check passes. At that moment, the new workers have handled health checks but no customer requests. preload_app! means the master loads the app once and the workers inherit it through fork. That covers the code. It doesn’t cover anything Ruby and Rails set up on first use: YJIT compiled code. YJIT compiles a method once it’s been called a certain number of times. The master calls very little during boot, so every worker compiles the same methods again on its own first requests. Compiled templates. Action View turns each ERB template into a Ruby method the first time it’s rendered. The schema cache. Active Record reads each model’s columns from the database the first time that model is used. Inline caches and memoized values throughout Ruby, Rails and the app. All 63 workers did all of this at once, while serving the traffic the old container had been handling a second earlier. In the test environment with YJIT on, the first request to a project page on a cold process took 652 ms, 151 ms of it YJIT compiling. The same request to a warm process took 28 ms. In production, CPU time per request peaked at around 200 ms while kamal-proxy moved traffic to the new container, against about 30 ms once the workers had warmed up. A host with spare CPU absorbs this. Every one of our web hosts has 48 cores and 63 workers, but each Amsterdam host serves around 250 requests per second, against 25 to 60 at our other sites. In Amsterdam the slow first requests turned into a queue. At a peak-hour deploy, the Puma backlog on an Amsterdam host reached anywhere from 250 to 2,238 requests, and kamal-proxy’s p99 response time hit about 10 seconds. Eron, our Director of Operations, had been tracking this since June. Another server in Amsterdam would help, but it would take weeks to arrive, so we also wanted to make deploys cheaper on the hardware we already had. What didn’t work We tried a few things first. In June, Donal tested the first two on a single Amsterdam host, comparing it with its neighbors, and they ruled out two likely causes. Warming each worker’s database connections. Puma’s before_fork hook clears the master’s connections, and each worker opened its own on its first request. Opening them in before_worker_boot instead made no difference. Queries on a freshly booted production host were already under a millisecond, so connections weren’t the problem. A synthetic request in each worker. Next, each worker made a few requests in before_worker_boot to an internal controller that touched every model. That ran the middleware, routing and Active Record paths, but it ran them in 63 workers at once — exactly the CPU spike we were trying to avoid. And a request with no real data renders no real views, so most of the app stayed cold. Spreading YJIT compilation out. Delaying YJIT in each worker by a random interval spread the compiling out over a few minutes, but every worker still ran interpreted until its delay ended. The queue didn’t change. Reforking from a warm worker. This is what Shopify’s Pitchfork does: let one worker serve traffic until it’s warm, then fork the others from it. Puma has an experimental version called fork_worker, and on beta it worked — the reforked workers were warm after three to five requests, where fresh ones took up to 30 seconds. But with fork_worker the template is worker 0, and it keeps serving requests. If it exits, the workers waiting to be forked never start (puma/puma#3596). If it gets no traffic, the refork never happens, which is what we saw on beta. Instacart have a mold_worker patch that promotes a warm worker to a template that stops serving, but it isn’t in a Puma release. We have a branch of it, and we may come back to it. That last experiment did show us where the fix was, though. Everything a warm worker has that a cold one lacks is in its memory, and fork copies memory. The master already has the app loaded. It just never runs it. Run the requests in the master So now, before the master binds its socket and forks, it makes the app’s own requests, in-process, the way a signed-in user would. Rack has a hook for exactly this. Rack::Builder#warmup takes a block that’s called once with the built app, before the server starts. rails server builds the app from config.ru, so the change to boot is one line: require_relative "config/environment" warmup { WarmUp.configured.run } if ENV["WARM_UP"] run Rails.application With preload_app! this runs in the master, and the workers inherit whatever it did. Puma binds its socket after the app is built, so until the warm-up finishes the health check’s connection is refused and kamal-proxy keeps retrying. No request reaches a worker that hasn’t been warmed. The warm-up has three steps. After precompiling the views, it gives the page requests and schema loading a shared 20-second budget, checked before each page or model. 1. Precompile the views actionview_precompiler reads every template for its render calls and compiles each one with the locals it’s passed. For us that’s 1,394 templates in about two seconds. A first request to a project page then compiles 2 templates instead of 44. 2. Request the pages, signed in A small browser class makes the requests through Rack::MockRequest, with the two cookies a real sign-in sets, then goes back for each page’s lazy Turbo frames: class WarmUp::Browser def initialize(signed_in_as:) @client = Rack::MockRequest.new(Rails.application) @headers = { "HTTP_USER_AGENT" => "Basecamp warm-up", "HTTP_COOKIE" => cookie_for(signed_in_as), "bc3.warm_up" => true } end def visit(path) page = get(path) frames_in(page).each { |id, src| get(src, "HTTP_TURBO_FRAME" => id) } end private def get(path, headers = {}) @client.get("https://#{host}#{path}", @headers.merge(headers)) end def frames_in(page) Nokogiri::HTML5(page.body).css("turbo-frame[src]").map { |frame| [ frame["id"], frame["src"] ] } end end The requests are signed in. The user is a monitoring account we already use for automated checks, and the pages are its own project, Campfire, to-dos, documents and messages. Public pages weren’t enough: after warming up with signed-out pages only, the first signed-in request to the projects page still took 131 ms, because authentication, the signed-in controllers and their views had never run. With signed-in pages it took 40 ms. cookie_for writes the same signed cookie the sign-in controller does, using the app’s own cookie jar, so there’s no API token and no secret to store. The frames are followed. The busiest HTML requests in production aren’t pages at all but Turbo frames — the sidebar badge, the inbox, the navigation menus. The browser parses each page and requests its <turbo-frame src> URLs with the Turbo-Frame header, so those controllers and views get warmed too. Our first four pages turned into 60 requests. The requests are excluded from rate limiting. They are internal, so they do not count against the rate limits that apply to real visitors. 3. Load the rest of the schema The page requests load the schema for the models they touch. The last step loads the rest, from the read replica: ApplicationRecord.reading do models.lazy.take_while { time_left? }.each { |model| model.load_schema if model.table_exists? } end The step checks 261 models and loads any schema information still missing. Those database round trips add up when the primary is far away: outside a request, Active Record uses the writing role, and from a host a long way from the primary each round trip is tens of milliseconds. Reading from the local replica brings the step down from about 20 seconds to 3.5. The pages go first because they load most of the schema anyway. If the time budget runs out, the step stops, logs how many models it got through, and the workers load the rest on first use like they always did. Rails can also load the schema from a dumped cache file at boot (bin/rails db:schema:cache:dump), which would make this step unnecessary. We don’t ship one in our image yet, because the dump needs a database to read from at build time, and we have several databases to cover. It’s on the list. What to close before the fork Running requests in the master opens things the master never opened before, and every worker inherits them. Two processes writing to the same socket will corrupt each other’s traffic, so you need to know what’s open before you fork. The way to find out is to list the master’s open file descriptors — ls -l /proc/<pid>/fd — before and after a warm-up, in an environment set up like production. Development wasn’t enough for us: it stores files on disk, so our S3 connections only showed up in production. Then, for each thing that’s open, check how its library handles a fork. We found three kinds: Already handled. Plenty of libraries detect a fork on their own, either by recording the PID they connected from and reconnecting in the child, by opening per-process files, or by resetting their thread pools. Redis clients, metrics libraries and concurrency libraries tend to be in this group. Check, but you probably don’t need to do anything. Already closed. Database connections are the classic one, and most Puma configs already clear them in before_fork. Anything else that’s opened per process — we have a SQLite cache the workers open on boot — needs closing when the warm-up finishes. Needs a new step. HTTP clients with keep-alive connections are the ones to look for: cloud SDKs with connection pools, tracing exporters, error reporters. They usually have no fork handling at all. We empty the aws-sdk connection pools in before_fork, and we run the warm-up untraced so the OpenTelemetry exporter never opens its connection to Tempo in the first place. Once that’s done, before_fork finishes with Process.warmup, which Ruby 3.3 added for this purpose: a major GC, a heap compaction, and every surviving object promoted to the old generation, so the memory pages the workers share change as little as possible afterwards. Choosing the pages The first list was the four pages that ran the busiest requests on beta. Once the warm-up was live, production showed us which endpoints were still cold. For one deploy, we compared each endpoint’s mean duration in the six minutes after kamal-proxy moved traffic to the new container with the same endpoint an hour later, then multiplied the difference by the number of requests in those six minutes. That gives the extra time each endpoint cost us because it was cold: Endpoint Cold Warm Requests in 6 min Extra seconds Campfire 246 ms 70 ms 6,490 1,140 Projects (JSON API) 84 ms 50 ms 22,077 771 Docs & Files 262 ms 177 ms 4,996 421 To-dos tool 205 ms 113 ms 4,018 371 To-dos (JSON API) 33 ms 16 ms 18,738 320 The pages already in the warm-up showed what to expect: the project page kept a 36 ms gap after a deploy, and the to-do page 10 ms. We’ve proposed adding these five requests, and expect them to add about five to seven seconds to the page step. The two JSON endpoints were a surprise. The warm-up’s page list had no API requests in it, so nothing on the API path had run before the first real request: not the API controllers, and not the Jbuilder templates rendering real records. Precompiling the views covers JSON templates too, but it isn’t a substitute for running the request. Results The warm-up is on for all 68 web hosts. With the first four pages it took 12 to 16 seconds per host: about 2 seconds to precompile the views, 7 to 9 for the 60 requests, and 3.5 for the schema. Deploys take that much longer per host, and we raised the deploy timeout from 30 to 60 seconds to cover it. In Amsterdam, at a peak-hour deploy: During deploy Before After Peak Puma backlog per host 250–2,238 requests 19–223 requests Peak kamal-proxy p99 about 10 s 2.4–4.8 s Peak CPU time per request 201–214 ms 88–132 ms Peak database time per request 56–69 ms 39–47 ms The same eight hosts at three deploys on 1 October, an hour apart, as the warm-up went from one host to four to all eight. The deploy in the middle, with four hosts warmed and four not, shows why every host needed the warm-up. Each warmed host recovered faster on its own: mean request duration peaked at 130 to 173 ms, against 203 to 311 ms on the hosts that weren’t warmed. But the backlog on all eight was about the same, because they were all waiting on the same database. Mean request duration on each host at the 07:21 UTC deploy. Blue hosts warmed up in the master before forking, orange hosts did not. Memory came down too. The workers now share compiled templates, YJIT code and the schema with the master instead of each building their own copy. On beta, the view precompiler alone took a busy worker’s private memory from 174–202 MB to 119–135 MB. Thirty minutes after the deploy, the web containers used about 39 GB less memory than the previous day’s containers at the same age and traffic. Amsterdam served most of our traffic at the times we tested. In Amsterdam, each new container used about 2 GB less just after traffic moved to it, which lowers the peak while the old and new containers overlap. Working with Claude Claude Code helped throughout. It combed through the per-worker backlogs and per-endpoint timings in Prometheus and Loki after each deploy, worked out the cold-versus-warm cost of each endpoint, and prepared the changes and the pull request descriptions with the benchmarks in them. We decided what to try, deployed it and read the results. If you do this Warm the master before it forks. Compile common code and templates and load their schema in the master, so workers inherit that work. With preload_app!, Rack::Builder#warmup runs before the workers start accepting traffic. Use the app’s real requests. Public pages, internal endpoints and synthetic queries warm the paths they run and nothing else. Signed-in requests to real records, frames included, run what production runs. Measure the cold penalty per endpoint. The difference between an endpoint’s cold and warm duration, times its request count after a deploy, ranks the pages worth adding. Ours weren’t the ones we’d have guessed, and two of them were JSON. Check what the warm-up leaves open. List the master’s file descriptors after a warm-up and account for every one before the fork. Two of ours needed changes. Set a time budget. A warm-up that runs long on one slow host fails the deploy on that host. Ours gives the page requests and schema loading a shared 20-second budget, checked before each page or model, puts the most valuable pages first, and logs what it skipped. Reforking from a warm worker, as Pitchfork does, solves the same problem continuously rather than once at boot, and it would warm paths no fixed list of pages covers. We may still get there: our branch brings Instacart’s mold_worker up to date with Puma’s main branch and fixes the bugs we found in it. But warming the master works with the Puma we already run, took a few days to implement, and substantially reduced the queues after deployment.