More from Alex Meub
Using modern AI coding tools feels like jumping into the cockpit of a BattleMech. My co-worker used this analogy recently and I love it. It perfectly sums up the feeling of vibe coding for me. I can move faster, jump higher and it feels like a there is a whole new world of possibilities available to me. This is true for me as someone who no longer writes code every day, but many experienced software engineers don’t feel this way and I get it. They’ve been running around on foot and learned to be very effective without it. Jumping into the cockpit of something entirely new is jarring. The BattleMech can feel clunky, burdensome and they basically have to relearn all their instincts around movement and orientation (it can also sometimes shoot itself in the foot!). On top of this, many non-technical folks have also jumped into the BattleMech. They are running off in all these strange directions because they don’t know what to use it for. They are copying things, building things that suck and filling social media feeds with their creations. Many software engineers feel the same way as artists did a few years ago because pretty much anyone can create software now. The good news is that domain knowledge and software development instincts are still essential. The BattleMech can be incredible if you know exactly where you want it to go, but it’s also happy to lead you straight off a cliff.
Progress Quest is generally considered the original idle game. It came out in 2002 as a parody of EverQuest and the emerging MMORPG boom. “Playing” it consists of creating a character, clicking “Sold!”, and then watching progress bars fill forever. There’s no interaction, no real gameplay, just waiting. At first glance, it feels like a gimmick — a joke game built to poke fun at the MMO trend. But what’s surprising is that its creator, Eric Fredricksen, built a whole RPG simulation underneath the progress bars. There are 270+ monsters, procedurally named equipment, multiple storylines, intentionally weighted stats, real loot tables, and a surprisingly complex progression system. On top of that, there’s even an authentication system for competitive multiplayer leaderboards that are somehow still around today. I love Progress Quest because of its absurdity, but also because it’s such a good example of something being far better than it needed to be. The amount of effort and attention to detail in this game makes me smile. Exponential Progression At the heart of Progress Quest is a single formula that controls level progression. The time it takes to complete level N is: (20 + 1.15^N) * 60 seconds. That means early levels take minutes (a few hours to get to level 10), while later ones take years (many years to get to level 100). There is also additional time outside of leveling for the player to go to market, buy/sell things, and head back to the “killing fields”. There are still thousands of players active across the remaining multiplayer realms, and pretty much everyone above level 95 has had the game running for over a decade. That doesn’t even count single-player characters. Races and Classes The race and class systems are hilariously absurd, but they don’t affect gameplay at all. The player races are: Half Orc, Half Man, Half Halfling, Double Hobbit, Hob-Hobbit, Low Elf, Dung Elf, Talking Pony, Gyrognome, Lesser Dwarf, Crested Dwarf, Eel Man, Panda Man, Trans-Kobold, Enchanted Motorcycle, Will o’ the Wisp, Battle-Finch, Double Wookiee, Skraeling, Demicanadian, and Land Squid. And the classes are: Ur-Paladin, Voodoo Princess, Robot Monk, Mu-Fu Monk, Mage Illusioner, Shiv-Knight, Inner Mason, Fighter/Organist, Puma Burgular, Runeloremaster, Hunter Strangler, Battle-Felon, Tickle-Mimic, Slow Poisoner, Bastard Lunatic, Jungle Clown, Birdrider, and Vermineer. Hilarious Monster Types There are over 270 hand-crafted monsters, and every one of them has a thematic loot drop. The Giant series imagines what giants would be if they were made of basically anything: Humidity Giant (drops “drops”) Beef Giant (drops “steak”) Rice Giant (drops “grain”) Porcelain Giant (drops “fixture”) Mini Giant (drops “pompadour”) The Golem series follows the same logic: Beer Golem (drops “foam”) Oxygen Golem (drops “platelet”) Cardboard Golem (drops “recycling”) The Scout hierarchy is great: Cub Scout (drops “neckerchief”) Girl Scout (drops “cookie”) Boy Scout (drops “merit badge”) Eagle Scout (drops “merit badge”) The Elemental series is an entirely new take on Elementals: Bacon Elemental (drops “bit”) Cheese Elemental (drops “curd”) Hair Elemental (drops “follicle”) Porn Elemental (drops “lube”) When the game picks a monster to fight, it level-matches against your character and then applies modifier prefixes based on the gap. That adds more flavor to the monster names. Procedural Equipment When you get new gear, the game doesn’t just pull from a list. It runs a little algorithm: Pick a base item matched to your level from a list like Stick → Shiv → Longsword → Halberd Calculate the quality gap between the item’s base level and your level “Spend” that gap across up to two modifier adjectives, each with a point value Whatever is left becomes a numeric +N prefix So a level 40 character might find a +13 Custom Holy Mithril Mail — a level 19 Mithril Mail base, with Custom (+3) and Holy (+5), leaving +13 unspent. The modifier tables are split into good and bad. If the gap is negative — meaning the item is actually better than you — the game pulls from the bad list instead: Rusty, Dull, Bent, Plastic, Nerf (-7), Rubber (-6). There’s also a whole list of spells with names like “Holy Batpole,” “Grognor’s Big Day Off”, and “Roger’s Grand Illusion” that will make their way into your spellbook and increase in level with roman numerals. The Main Game Loop Surprisingly, the game has a real game loop. It works like this: Kill monster task — The game generates a monster with a duration based on your level. When the timer finishes: Loot is added to your inventory, either a specific drop or generic loot XP is gained, which can trigger a level-up Quest and plot bars advance Check encumbrance — After a kill, if encumberance is at or above your limit, you go to market instead of fighting: The game will say “heading to market to sell loot” Then the game sells items one at a time, removing the top item in inventory and adding gold Items with “of” in the name sell for much more This continues until only gold remains Buy or head out — After selling, or if you weren’t encumbered in the first place: If the player’s gold is high enough to buy better gear, the game says “Negotiating purchase of better equipment” and the game upgrades a random equipment slot Otherwise the game will say “Heading to the killing fields” Next kill — After heading out, the game generates another monster and the cycle repeats Stat Progression When you gain a stat point, the game uses a weighted system biased toward your highest stat. Half the time, the gain is completely random. The other half uses quadratic weighting, where each stat’s chance is proportional to its value squared. That creates a snowball effect where your best stat keeps getting better, which feels authentic to how RPG builds tend to work. The only real strategy to playing Progress Quest is trying to roll high STR at character creation. Having higher STR gives you higher max encumberance which affects how often you need to go to market. The thing is, the market trips are such a small fraction of actual game time that this only about a 5% difference in how fast your character will level up. In Conclusion This is probably more than anyone wanted to know about Progress Quest. It’s an absurd, charming little game, and I hope it somehow keeps living forever. If you want to play the “multiplayer” Windows version, you can download it here. I also wanted to play on my Mac, so I vibe-coded an Swift version that runs on modern Mac hardware. See you on the killing fields!
I made a retro-inspired dock to charge my Playdate out of a Raspberry Pi case. The case is a miniature version of the Super Famicom and I love how it makes the Playdate look like a little game cartridge when it’s charging. Making one yourself is pretty straight-forward, you’ll just need the following components: Retroflag SUPERPi case A compact right-angle USB-C cable, like this one The 3D printed insert I designed First, print the two halves that make up the 3D printed insert and attach them with super glue. Make sure to align the cutouts on each side and then clamp the two pieces in place while the glue dries. Then take the Retroflag case apart and unscrew the main board. You’ll have to cut some wires and remove the front-facing USB ports. Make sure to leave the rear-facing USB-C port in place as we’ll reuse this to power the Playdate. Then, using a Dremel, cut a rough 86 by 21 mm rectangular hole in the top of the case. It doesn’t have to be clean as it will get covered up by the 3D printed insert. You will also need to remove some of the internal support structure inside the case using pliers or flush cutters to make space for the insert and wires. After it has dried, insert the 3D printed piece through the hole and hot glue the right-angle USB-C cable into place. Lastly, splice the USB-C cable to the red and black wires coming off the rear-facing USB port on the case. The red (or pink) USB-C cable wire should be spliced to the red wire on the USBC-C port. The two black wires should also be spliced together. That’s it! At some point I’d like to make it into a functional USB hub and add an internal LED.
A few years ago, I built a Wi-Fi-controlled Nerf turret, but I never got around to creating a proper build guide for it. When I finally sat down to write one, I realized just how many things I would do differently. That realization quickly snowballed into a full redesign—and the result is a vastly improved version of the SwarmTurret. This new version is not only more powerful and precise, but it’s also easier build. Here are some of the major upgrades: Simplified Assembly: I reused the shell of an existing plastic blaster, which significantly cuts down on 3D printing and makes putting it together much easier. Improved Stability: A new belt-driven Y-axis adds smoother motion and includes an adjustable tension system. Enhanced Accuracy: The camera has been repositioned for better aiming precision. Direct X-Axis Drive: I replaced the original gear system with direct motor control for more responsive movement. Performance Boost: Upgraded from Raspberry Pi 4 to Pi 5 for faster web app performance. Integrated Power Supply: Now features a built-in power supply with an on/off switch—no more fumbling with cables. Web App Enhancements: The control interface is more intuitive and responsive. I’ve published the complete build guide, along with the updated code and 3D printable parts.
3D Printing has allowed me to be creative in ways I never thought possible. It has allowed me to create products that provide real value, products that didn’t exist before I designed them. On top of that, it’s satisfied my desire to ship products, even if the end-user is just me. Another great thing is how quickly 3D printing provides value. If I see a problem, I can design and print a solution that works in just a few hours. Even if I’m the only one who benefits, that’s enough. But sharing these creations takes the experience even further. When I see others use or improve on something I’ve made, it makes the process feel so much more worthwhile. It gives me the same feeling of fulfillment when I ship software products at work. Before mass-market 3D printing, creators would need to navigate the complexity and high costs of mass-production methods (like injection molding) even to get a limited run of a niche product produced. With 3D printing, they can transfer the cost of production to others. Millions of people have access to good 3D printers now (at home, work, school, libraries, maker spaces), which means almost anyone can replicate a design. Having a universal format for sharing 3D designs dramatically lowers the effort that goes into sharing them. Creators can share their design as an STL file, which describes the surface geometry of their 3D object as thousands of little triangles. This “standard currency” of the 3D printing world is often all that is required to precisely replicate a design. This dramatically lowers the effort that goes into sharing printable designs. The widespread availability of 3D printers and the universal format for sharing 3D designs has allowed 3D-printed products to not only exist but thrive in maker communities. This is the magic of 3D printing: it empowers individuals to solve their own problems by designing solutions while enabling others to reproduce those designs at minimal cost and effort.
More in programming
A clip of me singing a funny song from Gilbert and Sullivan’s Ruddigore back in 2013
In this video, we look at why fork() needs copy-on-write, how it works inside the kernel, and a memory usage problem that Instagram encountered with Python.
Comments require commitment, but they’re worth it.
Basecamp 5 runs on Puma in cluster mode: one master process with preload_app! and 63 single-threaded workers per host, deployed as a Docker container with Kamal. We serve Basecamp from several sites. Each site has its own web hosts and a read replica of the database, and writes go to a single primary database in one of them. On our busiest hosts, each deploy left up to 2,000 requests waiting while the new workers warmed up. We reduced those queues by running signed-in requests through the app in the Puma master, before it forked the workers. Why 63 single-threaded workers? Basecamp has always served web requests from processes rather than threads. It ran on Unicorn, which only does processes, until we moved to Puma in January 2025, and we kept the same setup: workers (Concurrent.physical_processor_count * 1.3).ceil threads 1, 1 preload_app! On a 48-core host that’s 63 workers, each handling one request at a time. We chose 1.3 after benchmarking HEY in 2023, when we moved our apps out of the cloud and onto our own hardware. We tested several combinations of workers and threads with a mix of GET and POST requests on a 32-vCPU VM. Every multithreaded configuration we tested was slower and handled fewer requests than single-threaded workers. Adding workers beyond about 1.2 to 1.3 per vCPU brought little benefit. The threaded workers spent a lot of their time waiting for Ruby’s global VM lock. That made single-threaded workers a good fit for this workload, and we use the same setup for Basecamp. An app that spends more time waiting on its database or other services may benefit from more threads, so benchmark your own app. The other reason is the app itself. Basecamp has class-level state in places and has never needed to be thread-safe. With one request per process, it still doesn’t. Processes do use more memory than threads, and preload_app! reduces the difference. The master loads the app once and the workers share its memory through copy-on-write until they write to it. Shopify’s comparison of Ruby execution models explains the trade-off well. In the HEY benchmark the best setup came to about 260 MB of PSS per core, where PSS counts each shared page once, split between the processes using it, and the gap to a threaded setup was smaller than we’d expected. What Puma does on each host when a container starts: one master, then 63 forked workers that share its memory until they write to it. Two things about this setup matter for the rest of the post. A worker that’s compiling or loading something is fully blocked — there’s no other thread to pick up the next request. And whatever the master has in memory before it forks, all 63 workers share. Whatever they build after the fork, they build 63 times. What happens when we deploy Kamal starts the new container alongside the old one, and kamal-proxy moves the host’s traffic across as soon as the health check passes. At that moment, the new workers have handled health checks but no customer requests. preload_app! means the master loads the app once and the workers inherit it through fork. That covers the code. It doesn’t cover anything Ruby and Rails set up on first use: YJIT compiled code. YJIT compiles a method once it’s been called a certain number of times. The master calls very little during boot, so every worker compiles the same methods again on its own first requests. Compiled templates. Action View turns each ERB template into a Ruby method the first time it’s rendered. The schema cache. Active Record reads each model’s columns from the database the first time that model is used. Inline caches and memoized values throughout Ruby, Rails and the app. All 63 workers did all of this at once, while serving the traffic the old container had been handling a second earlier. In the test environment with YJIT on, the first request to a project page on a cold process took 652 ms, 151 ms of it YJIT compiling. The same request to a warm process took 28 ms. In production, CPU time per request peaked at around 200 ms while kamal-proxy moved traffic to the new container, against about 30 ms once the workers had warmed up. A host with spare CPU absorbs this. Every one of our web hosts has 48 cores and 63 workers, but each Amsterdam host serves around 250 requests per second, against 25 to 60 at our other sites. In Amsterdam the slow first requests turned into a queue. At a peak-hour deploy, the Puma backlog on an Amsterdam host reached anywhere from 250 to 2,238 requests, and kamal-proxy’s p99 response time hit about 10 seconds. Eron, our Director of Operations, had been tracking this since June. Another server in Amsterdam would help, but it would take weeks to arrive, so we also wanted to make deploys cheaper on the hardware we already had. What didn’t work We tried a few things first. In June, Donal tested the first two on a single Amsterdam host, comparing it with its neighbors, and they ruled out two likely causes. Warming each worker’s database connections. Puma’s before_fork hook clears the master’s connections, and each worker opened its own on its first request. Opening them in before_worker_boot instead made no difference. Queries on a freshly booted production host were already under a millisecond, so connections weren’t the problem. A synthetic request in each worker. Next, each worker made a few requests in before_worker_boot to an internal controller that touched every model. That ran the middleware, routing and Active Record paths, but it ran them in 63 workers at once — exactly the CPU spike we were trying to avoid. And a request with no real data renders no real views, so most of the app stayed cold. Spreading YJIT compilation out. Delaying YJIT in each worker by a random interval spread the compiling out over a few minutes, but every worker still ran interpreted until its delay ended. The queue didn’t change. Reforking from a warm worker. This is what Shopify’s Pitchfork does: let one worker serve traffic until it’s warm, then fork the others from it. Puma has an experimental version called fork_worker, and on beta it worked — the reforked workers were warm after three to five requests, where fresh ones took up to 30 seconds. But with fork_worker the template is worker 0, and it keeps serving requests. If it exits, the workers waiting to be forked never start (puma/puma#3596). If it gets no traffic, the refork never happens, which is what we saw on beta. Instacart have a mold_worker patch that promotes a warm worker to a template that stops serving, but it isn’t in a Puma release. We have a branch of it, and we may come back to it. That last experiment did show us where the fix was, though. Everything a warm worker has that a cold one lacks is in its memory, and fork copies memory. The master already has the app loaded. It just never runs it. Run the requests in the master So now, before the master binds its socket and forks, it makes the app’s own requests, in-process, the way a signed-in user would. Rack has a hook for exactly this. Rack::Builder#warmup takes a block that’s called once with the built app, before the server starts. rails server builds the app from config.ru, so the change to boot is one line: require_relative "config/environment" warmup { WarmUp.configured.run } if ENV["WARM_UP"] run Rails.application With preload_app! this runs in the master, and the workers inherit whatever it did. Puma binds its socket after the app is built, so until the warm-up finishes the health check’s connection is refused and kamal-proxy keeps retrying. No request reaches a worker that hasn’t been warmed. The warm-up has three steps. After precompiling the views, it gives the page requests and schema loading a shared 20-second budget, checked before each page or model. 1. Precompile the views actionview_precompiler reads every template for its render calls and compiles each one with the locals it’s passed. For us that’s 1,394 templates in about two seconds. A first request to a project page then compiles 2 templates instead of 44. 2. Request the pages, signed in A small browser class makes the requests through Rack::MockRequest, with the two cookies a real sign-in sets, then goes back for each page’s lazy Turbo frames: class WarmUp::Browser def initialize(signed_in_as:) @client = Rack::MockRequest.new(Rails.application) @headers = { "HTTP_USER_AGENT" => "Basecamp warm-up", "HTTP_COOKIE" => cookie_for(signed_in_as), "bc3.warm_up" => true } end def visit(path) page = get(path) frames_in(page).each { |id, src| get(src, "HTTP_TURBO_FRAME" => id) } end private def get(path, headers = {}) @client.get("https://#{host}#{path}", @headers.merge(headers)) end def frames_in(page) Nokogiri::HTML5(page.body).css("turbo-frame[src]").map { |frame| [ frame["id"], frame["src"] ] } end end The requests are signed in. The user is a monitoring account we already use for automated checks, and the pages are its own project, Campfire, to-dos, documents and messages. Public pages weren’t enough: after warming up with signed-out pages only, the first signed-in request to the projects page still took 131 ms, because authentication, the signed-in controllers and their views had never run. With signed-in pages it took 40 ms. cookie_for writes the same signed cookie the sign-in controller does, using the app’s own cookie jar, so there’s no API token and no secret to store. The frames are followed. The busiest HTML requests in production aren’t pages at all but Turbo frames — the sidebar badge, the inbox, the navigation menus. The browser parses each page and requests its <turbo-frame src> URLs with the Turbo-Frame header, so those controllers and views get warmed too. Our first four pages turned into 60 requests. The requests are excluded from rate limiting. They are internal, so they do not count against the rate limits that apply to real visitors. 3. Load the rest of the schema The page requests load the schema for the models they touch. The last step loads the rest, from the read replica: ApplicationRecord.reading do models.lazy.take_while { time_left? }.each { |model| model.load_schema if model.table_exists? } end The step checks 261 models and loads any schema information still missing. Those database round trips add up when the primary is far away: outside a request, Active Record uses the writing role, and from a host a long way from the primary each round trip is tens of milliseconds. Reading from the local replica brings the step down from about 20 seconds to 3.5. The pages go first because they load most of the schema anyway. If the time budget runs out, the step stops, logs how many models it got through, and the workers load the rest on first use like they always did. Rails can also load the schema from a dumped cache file at boot (bin/rails db:schema:cache:dump), which would make this step unnecessary. We don’t ship one in our image yet, because the dump needs a database to read from at build time, and we have several databases to cover. It’s on the list. What to close before the fork Running requests in the master opens things the master never opened before, and every worker inherits them. Two processes writing to the same socket will corrupt each other’s traffic, so you need to know what’s open before you fork. The way to find out is to list the master’s open file descriptors — ls -l /proc/<pid>/fd — before and after a warm-up, in an environment set up like production. Development wasn’t enough for us: it stores files on disk, so our S3 connections only showed up in production. Then, for each thing that’s open, check how its library handles a fork. We found three kinds: Already handled. Plenty of libraries detect a fork on their own, either by recording the PID they connected from and reconnecting in the child, by opening per-process files, or by resetting their thread pools. Redis clients, metrics libraries and concurrency libraries tend to be in this group. Check, but you probably don’t need to do anything. Already closed. Database connections are the classic one, and most Puma configs already clear them in before_fork. Anything else that’s opened per process — we have a SQLite cache the workers open on boot — needs closing when the warm-up finishes. Needs a new step. HTTP clients with keep-alive connections are the ones to look for: cloud SDKs with connection pools, tracing exporters, error reporters. They usually have no fork handling at all. We empty the aws-sdk connection pools in before_fork, and we run the warm-up untraced so the OpenTelemetry exporter never opens its connection to Tempo in the first place. Once that’s done, before_fork finishes with Process.warmup, which Ruby 3.3 added for this purpose: a major GC, a heap compaction, and every surviving object promoted to the old generation, so the memory pages the workers share change as little as possible afterwards. Choosing the pages The first list was the four pages that ran the busiest requests on beta. Once the warm-up was live, production showed us which endpoints were still cold. For one deploy, we compared each endpoint’s mean duration in the six minutes after kamal-proxy moved traffic to the new container with the same endpoint an hour later, then multiplied the difference by the number of requests in those six minutes. That gives the extra time each endpoint cost us because it was cold: Endpoint Cold Warm Requests in 6 min Extra seconds Campfire 246 ms 70 ms 6,490 1,140 Projects (JSON API) 84 ms 50 ms 22,077 771 Docs & Files 262 ms 177 ms 4,996 421 To-dos tool 205 ms 113 ms 4,018 371 To-dos (JSON API) 33 ms 16 ms 18,738 320 The pages already in the warm-up showed what to expect: the project page kept a 36 ms gap after a deploy, and the to-do page 10 ms. We’ve proposed adding these five requests, and expect them to add about five to seven seconds to the page step. The two JSON endpoints were a surprise. The warm-up’s page list had no API requests in it, so nothing on the API path had run before the first real request: not the API controllers, and not the Jbuilder templates rendering real records. Precompiling the views covers JSON templates too, but it isn’t a substitute for running the request. Results The warm-up is on for all 68 web hosts. With the first four pages it took 12 to 16 seconds per host: about 2 seconds to precompile the views, 7 to 9 for the 60 requests, and 3.5 for the schema. Deploys take that much longer per host, and we raised the deploy timeout from 30 to 60 seconds to cover it. In Amsterdam, at a peak-hour deploy: During deploy Before After Peak Puma backlog per host 250–2,238 requests 19–223 requests Peak kamal-proxy p99 about 10 s 2.4–4.8 s Peak CPU time per request 201–214 ms 88–132 ms Peak database time per request 56–69 ms 39–47 ms The same eight hosts at three deploys on 1 October, an hour apart, as the warm-up went from one host to four to all eight. The deploy in the middle, with four hosts warmed and four not, shows why every host needed the warm-up. Each warmed host recovered faster on its own: mean request duration peaked at 130 to 173 ms, against 203 to 311 ms on the hosts that weren’t warmed. But the backlog on all eight was about the same, because they were all waiting on the same database. Mean request duration on each host at the 07:21 UTC deploy. Blue hosts warmed up in the master before forking, orange hosts did not. Memory came down too. The workers now share compiled templates, YJIT code and the schema with the master instead of each building their own copy. On beta, the view precompiler alone took a busy worker’s private memory from 174–202 MB to 119–135 MB. Thirty minutes after the deploy, the web containers used about 39 GB less memory than the previous day’s containers at the same age and traffic. Amsterdam served most of our traffic at the times we tested. In Amsterdam, each new container used about 2 GB less just after traffic moved to it, which lowers the peak while the old and new containers overlap. Working with Claude Claude Code helped throughout. It combed through the per-worker backlogs and per-endpoint timings in Prometheus and Loki after each deploy, worked out the cold-versus-warm cost of each endpoint, and prepared the changes and the pull request descriptions with the benchmarks in them. We decided what to try, deployed it and read the results. If you do this Warm the master before it forks. Compile common code and templates and load their schema in the master, so workers inherit that work. With preload_app!, Rack::Builder#warmup runs before the workers start accepting traffic. Use the app’s real requests. Public pages, internal endpoints and synthetic queries warm the paths they run and nothing else. Signed-in requests to real records, frames included, run what production runs. Measure the cold penalty per endpoint. The difference between an endpoint’s cold and warm duration, times its request count after a deploy, ranks the pages worth adding. Ours weren’t the ones we’d have guessed, and two of them were JSON. Check what the warm-up leaves open. List the master’s file descriptors after a warm-up and account for every one before the fork. Two of ours needed changes. Set a time budget. A warm-up that runs long on one slow host fails the deploy on that host. Ours gives the page requests and schema loading a shared 20-second budget, checked before each page or model, puts the most valuable pages first, and logs what it skipped. Reforking from a warm worker, as Pitchfork does, solves the same problem continuously rather than once at boot, and it would warm paths no fixed list of pages covers. We may still get there: our branch brings Instacart’s mold_worker up to date with Puma’s main branch and fixes the bugs we found in it. But warming the master works with the Puma we already run, took a few days to implement, and substantially reduced the queues after deployment.