Full Width [alt+shift+f] Shortcuts [alt+shift+k]
Sign Up [alt+shift+s] Log In [alt+shift+l]
1

A letter to John Ternus

from Marco.org [alt+shift+b] in programming

As Apple celebrates its fiftieth birthday, we celebrate the spirit of its formation, when people who loved computers started making great computers to inspire more people to love computers. That spirit is difficult to find in the tech business today. Immense scale, soulless optimization, and an insatiable thirst for growth dominate its behavior and discourse, leaving little room for the spirit and principles embodied by Steve Jobs and Steve Wozniak. Apple still has this spirit, and I believe you do, too. But it’s not infinite or invincible. It’s under constant pressure, including from Apple itself. It seems likely that you’ll soon be leading Apple, which will place unfathomable responsibility on your shoulders. As you grow into the leader that we know you can be, I urge you, on behalf of everyone who loves computers as much as we do, to protect and cultivate this spirit of Apple’s founders as the company’s top priority: We love computers. We don’t hide that — we celebrate it! We use computers to enhance our minds, lives, and abilities — not to be controlled, restricted, tricked, placated, angered, or surveilled. Our computers work for us, with the utmost respect for our time, attention, money, data, and privacy. We are customers and owners — not resources to be harvested, annoyed, or badgered into ever more services and upsells. Apple leads the industry in these values, but leading doesn’t always mean excelling. Remaining true to these values requires constant diligence, honest evaluation, introspection, and the audacity and courage to effect change. Apple doesn’t settle for fine, functional, or good enough in its hardware (and thanks for your incredible work on that). We love making and using products that aren’t just great, but greater than they need to be, always raising the bar of greatness for its own sake. Software, services, revenue sources, and world impact need to be held to that same standard. Focus on making great computers with great user experiences...
1st Apr 2026

Stay updated

Get a weekly newsletter with the top 5 articles worth reading every week.

More from Marco.org

Bob and Van

For the first half of my childhood in suburban Ohio, we lived next to Bob. Bob and his wife were a kind, reserved older couple who decorated their home mostly with crucifixes and talked about Jesus a lot. Bob helped my mom with a lot of common household tasks after my dad died. When she woke up to a small fire on our deck, she didn’t call the fire department — she called Bob.1 I never learned any of Bob’s political views, though I can make some guesses in retrospect. They never mattered. He was just our neighbor. We later moved, and hit the neighbor jackpot with Van. Van was deeply kind, generous, and friendly — the sort of neighbor everyone wishes for. In my lazy teenage years, he’d often mow our lawn or shovel our snow for us before I even woke up. As his mobility later declined, I was honored to return the favor. Van and his delightful wife passed away years ago, before I learned how either of them felt on any political issue. If they were my neighbors today, how likely would that be? *     *     * If I learn today that my neighbors have political views that my “side” considers unacceptable and unforgivable, should I be able to have dinner at their house, shovel their snow, or even greet them outside? Should I be condemned for supporting or associating with them in non-political contexts? Should I be condemned for not publicly condemning them myself? While privilege and power dynamics play a complex and unavoidable role in this, I don’t want to associate with any community that would insist on either of those.2 And I can’t help but feel that maybe we were better off before we knew everyone’s hot takes on everything. The town coffee shop was just a coffee shop. We weren’t publicly shamed for going there because one of the owners was an ass. It was just a place to hang out with our friends. Our casual work friends were just casual work friends. We didn’t expect them to take positions on the world’s most complicated conflicts. We could just talk about weather and TV shows, and maybe sometimes get a beer after work. And our neighbors were just our neighbors. That was the day I learned the word “SHIT!” ↩︎ Somewhere on Bob’s wall might’ve been a relevant statement from an influential political dissident about casting the first stone. ↩︎

4 weeks ago • 1 votes
Unforgetful

I have ADHD. Probably.1 My entire adolescence was defined by forgetting to do my homework, disappointing everyone around me, and being told by every adult that I was lazy, “just” needed to try harder, and wasn’t living up to my potential. I’ll be working on the resulting shame and anxiety for the rest of my life.2 If someone tells me, “Remember to pick up the dry cleaning tomorrow,” I definitely won’t remember. If someone tells me, “Remember to put the laundry in the dryer in ten minutes,” I probably won’t even remember that. If you’re thinking, “It’s just ten minutes! Nobody can possibly forget that,” I have made an app that’s probably not for you. *     *     * Computers saved me. Despite constantly being told that I’d never get a good job, computers provided a lucrative career path filled with even worse academic performers who couldn’t care less about my grades. What a stroke of luck! I’d be mediocre to terrible at most jobs, but I turned out to be very good at using computers to turn coffee and Phish into money. But computers also save me in smaller, everyday ways. They help me remember. *     *     * Sorry, I just had to go put the laundry in the dryer, which I had forgotten. Really. It’s that bad. *     *     * Calendar alarms changed my life. Events without alarms don’t happen. (If you’re thinking, “Why doesn’t he just check his calendar?” I have made an app that’s probably not for you.) I had less success with reminders until we had ubiquitous voice capture. If I think of something at inconvenient times, like while driving or exercising, I need to capture it immediately or I’ll forget. It cannot wait until I get to my phone or computer. It has to be recorded right then or it’s gone. Siri, for all of its faults, has always been very good at creating tasks in Apple’s Reminders app — and Siri is everywhere. Computer. Phone. Car. Watch. I can almost always ask Siri to add something to Reminders. So Reminders’ integration and ubiquity are invaluable to me. The Reminders app, though, doesn’t fit me at all. I don’t like how it looks, how it works, or how it’s organized. Hell, I don’t even like its icon. There’s nothing about it I like. So I never open it. (Not as if I’d ever remember to check it anyway.)3 Most of my Reminders interaction has always been via Siri and notifications. And that fits me very well, with three major exceptions: If I say “Remind me to buy milk,” but I forget to specify a time (or Siri misses it), Reminders will never notify me, and I’ll never check the app, so it’ll be lost. If I inadvertently dismiss a Reminders notification without snoozing it, Reminders will never notify me about it again, so it’ll be lost. When I do snooze a Reminders notification, which I do a lot, the snooze options suck.4 This spring, I made myself an app to fix all three, and… it escalated. I’ve been using it for months, and it has profoundly improved my life. Today, I’m releasing it, in hopes that maybe it can help other people, too. *     *     * Introducing Unforgetful. The concept is simple: Reminders for ADHD. This means: It’s impossible to lose a task. Notifications always repeat, even if you miss or dismiss them, until the task is completed or deleted. Designed for procrastination. Snoozing is a core feature, with thoughtful intervals that scale as you snooze a task more. Nothing is hidden away. A hidden task is a forgotten task. No folders, no tags, no organization, no different views or modes. No judgment or shame. Nothing is overdue. You’re not in trouble. Everything is either due now or in the future. Didn’t get to it today? Maybe tomorrow. Every day is a fresh start. And Unforgetful does all of that with your Reminders data. That means: Siri works perfectly, everywhere. Just say “Remind me…” without having to specify an app. It replaces your Reminders notifications. Turn them off, turn these on. No import. This isn’t a new system to learn or migrate into. Your Reminders data is just… there. No export. If you try Unforgetful and it’s not for you, just delete it and turn Reminders notifications back on, and all of your data is right there in Reminders. Mix and match. Don’t like it everywhere? You can still use Reminders on your Vision Pro or whatever, or simultaneously use Unforgetful with Reminders or any other app that uses Reminders data. I’ve optimized lots of features and details for ADHD, too: Fast capture with dictation. Tap the microphone button in Unforgetful — or its full suite of widgets, or its Lock Screen or Control Center controls — and quickly dictate a task.5 Spaced-out notifications, not an overwhelming pile at once. If multiple tasks are set to alert you at the same time, they space themselves out so you’re not barraged with a hopeless stack of obligations. And the order changes every day, so nothing gets lost between things. Recurring-task backlogs don’t pile up. If you miss multiple intervals for a recurring task, completing it doesn’t make you click through all of your past failures to get to the present day… it just schedules the next one from now. Remind me 5 minutes after I get home. Location-based reminders can notify you after a set delay. Because if a “remind me when I get home” task fires as I pull into the driveway, I’ll forget it by the time I’m inside, unloaded, and ready to do anything about it. Unforgetful is a Mac and iOS app that’s $19.99 per year (US). One subscription buys all platforms. There’s a one-month free trial so you can really live with it and see if it’s right for you before paying. There are a million task-management apps. Unforgetful is not for everyone, but it’s really for me, far more than anything else has ever been. Frankly, I have no idea how to reach the other people it’s for. But I know you’re out there. If any of this resonates with you… give it a shot. Despite significant progress on my mental health around this from therapy and media, I haven’t (yet?) sought an official ADHD diagnosis because I haven’t sought medication, and that seemed like the main reason why I’d want a diagnosis. That’s probably worth reconsidering. I’ll make myself a reminder to do it. Someday. ↩︎ I saw a psychologist throughout high school for my homework problem. Rather than recognizing 7 of the 9 symptoms of what we now call “inattentive” ADHD, he was just one more adult condemning me for not caring. (I really did care. Nobody wanted me to be “better” more than me.) In the 1990s, if you didn’t do homework but weren’t hyperactive, you were obviously just choosing to be lazy, and the solution was apparently to apply more shame. Not all doctors are good. ↩︎ I’ve tried other apps that are beautiful and work differently, but they’re all complete task-management systems that introduce a lot of features and complexity that I neither need nor want, at the expense of Reminders’ ubiquitous integration that I highly value. I’ve also tried other apps that use the Reminders database, but I haven’t found any of them to be… good. (Sorry. I probably didn’t try yours.) ↩︎ At least they specify the actual snooze times now instead of vague descriptions like “the afternoon,” which I proudly take full credit for, since it happened to change only after I ranted about it relentlessly on our podcast for months. If I truly made this happen, it might be the greatest impact on the world I’ll ever have. ↩︎ Dictation just records every word you say as the title, which is extremely fast and reliable. It doesn’t yet recognize things like “Remind me to buy milk at 10 AM tomorrow.” I haven’t nailed that natural-language processing yet — coming soon, maybe! For now, Siri does that part better. ↩︎

14th Aug 2026 • 1 votes
Retreating to Safety

Ten years ago, Apple’s Phil Schiller surprised Apple enthusiasts and developers by walking out on stage at John Gruber’s The Talk Show Live WWDC event and giving an open, human, honest interview to a somewhat jaded community. I wrote this in response: Both Apple and Phil Schiller himself took a huge risk in doing this. That they agreed at all is a noteworthy gift to this community of long-time enthusiasts, many of whom have felt under-appreciated as the company has grown. […] Phil’s appearance on the show was warm, genuine, informative, and entertaining. It was human. And humanizing the company and its decisions, especially to developers — remember, developer relations is all under Phil — might be worth the PR risk. This started a ten-year run of interviews by Apple executives on The Talk Show every year at WWDC that proved to be great, surprisingly safe PR for Apple. No executive ever said something they shouldn’t have (they’re pros), no sensational or negative news stories ever resulted from them, and Apple’s enthusiastic fans and developers felt seen, heard, and appreciated. *     *     * For unspecified reasons, Apple has declined to participate this year, ending what had become a beloved tradition in our community — and I can’t help but suspect that it won’t come back. (A lot has changed in the meantime.) Maybe Apple has good reasons. Maybe not. We’ll see what their WWDC PR strategy looks like in a couple of weeks. In the absence of any other information, it’s easy to assume that Apple no longer wants its executives to be interviewed in a human, unscripted, unedited context that may contain hard questions, and that Apple no longer feels it necessary to show their appreciation to our community and developers in this way. I hope that’s either not the case, or it doesn’t stay the case for long. This will be the first WWDC I’m not attending since 2009 (excluding the remote 2020 one, of course). Given my realizations about my relationship with Apple and how they view developers, I’ve decided that it’s best for me to take a break this year, gain some perspective, and decide what my future relationship should look like. Maybe Apple’s leaders are doing that, too.

30th May 2025 • 38 votes
Ten years of Overcast: A new foundation

Today, on the tenth anniversary of Overcast 1.0, I’m happy to launch a complete rewrite and redesign of most of the iOS app, built to carry Overcast into the next decade — and hopefully beyond. Like podcasts better than blog posts? Listen to ATP #596 for more! What’s new Much faster, more responsive, more reliable, and more accessible. Modern design, optimized for easily-reached controls on today’s phone sizes. Improvements throughout, such as undoing large seeks, new playlist-priority options, easier navigation, and more. What’s not Most features. Overcast is still Overcast! The audio engine. It’s the best part of Overcast, and still leads the industry in sound quality, silence skipping, and volume normalization. (More soon!) The business. I’m still a one-person operation, with no funding or external ownership, serving only my customers. My principles. I always want to make the best podcast app, and I’ll never disrespect your time, attention, or privacy. What’s gone Streaming. Most big podcasts now use dynamic ad insertion, which causes bugs and problems for streaming playback.1 Downloading episodes completely before they begin playback is much more reliable. Tapping a non-downloading episode will now open the playback screen, download it, then start playback. It works similarly to the way streaming did before, but playback begins after the download completes, not after a portion of it is buffered. On today’s fast networks, this usually only takes a few extra seconds. And in the near future, I’ll be adding smarter options and more control over selective downloading of episodes to further improve the experience for people who don’t automatically download every episode. What’s next The last few missing features from the old app, such as Shortcuts support, storage management, and OPML. These are absent now, but will return soon. More options for downloading and deleting episodes. Upgrading the Apple Watch app to the new, faster sync engine. (The Watch app is currently unchanged from the previous one.) And, of course, more features, including some of your most-requested features over the last decade. Getting this rewrite out the door was a monumental task. Thank you for your patience as I work through this list! Why? Most of Overcast’s core code was 10 years old, which made it cumbersome or impossible to easily move with the times, adopt new iOS functionality, or add new features, especially as one person. That’s why there haven’t been many new features or changes in years. You saw it, and I saw it. I wasn’t able to serve my customers as well as I wanted. For Overcast to have a future, it needed a modern foundation for its second decade. I’ve spent the past 18 months rebuilding most of the app with Swift, SwiftUI, Blackbird, and modern Swift concurrency. Now, development is rapidly accelerating. I’m more responsive, iterating more quickly, and ultimately making the app much better. Thank you all so much for the first decade of Overcast. Here’s to the next one. Dynamic ad insertion (DAI) splices ads into each download, and no two downloads are guaranteed to have the same number or duration of ads. So, for example, if the first half of an episode downloads, then the download fails, and it downloads the second half with another request, the combined audio may jump forward or back at the halfway mark, losing or repeating content. ↩︎

16th Jul 2024 • 93 votes

More in programming

Clip of me singing Despard in Ruddigore in 2013

A clip of me singing a funny song from Gilbert and Sullivan’s Ruddigore back in 2013

7 hours ago • 1 votes
How and Why fork() Uses Copy-on-Write

In this video, we look at why fork() needs copy-on-write, how it works inside the kernel, and a memory usage problem that Instagram encountered with Python.

13 hours ago • 1 votes
What we lost when we lost comments

Comments require commitment, but they’re worth it.

19 hours ago • 1 votes
Lighthouse map

Lovely global map with animated lights sweeping the waters

21 hours ago • 1 votes
Warming up the Puma master before it forks

Basecamp 5 runs on Puma in cluster mode: one master process with preload_app! and 63 single-threaded workers per host, deployed as a Docker container with Kamal. We serve Basecamp from several sites. Each site has its own web hosts and a read replica of the database, and writes go to a single primary database in one of them. On our busiest hosts, each deploy left up to 2,000 requests waiting while the new workers warmed up. We reduced those queues by running signed-in requests through the app in the Puma master, before it forked the workers. Why 63 single-threaded workers? Basecamp has always served web requests from processes rather than threads. It ran on Unicorn, which only does processes, until we moved to Puma in January 2025, and we kept the same setup: workers (Concurrent.physical_processor_count * 1.3).ceil threads 1, 1 preload_app! On a 48-core host that’s 63 workers, each handling one request at a time. We chose 1.3 after benchmarking HEY in 2023, when we moved our apps out of the cloud and onto our own hardware. We tested several combinations of workers and threads with a mix of GET and POST requests on a 32-vCPU VM. Every multithreaded configuration we tested was slower and handled fewer requests than single-threaded workers. Adding workers beyond about 1.2 to 1.3 per vCPU brought little benefit. The threaded workers spent a lot of their time waiting for Ruby’s global VM lock. That made single-threaded workers a good fit for this workload, and we use the same setup for Basecamp. An app that spends more time waiting on its database or other services may benefit from more threads, so benchmark your own app. The other reason is the app itself. Basecamp has class-level state in places and has never needed to be thread-safe. With one request per process, it still doesn’t. Processes do use more memory than threads, and preload_app! reduces the difference. The master loads the app once and the workers share its memory through copy-on-write until they write to it. Shopify’s comparison of Ruby execution models explains the trade-off well. In the HEY benchmark the best setup came to about 260 MB of PSS per core, where PSS counts each shared page once, split between the processes using it, and the gap to a threaded setup was smaller than we’d expected. What Puma does on each host when a container starts: one master, then 63 forked workers that share its memory until they write to it. Two things about this setup matter for the rest of the post. A worker that’s compiling or loading something is fully blocked — there’s no other thread to pick up the next request. And whatever the master has in memory before it forks, all 63 workers share. Whatever they build after the fork, they build 63 times. What happens when we deploy Kamal starts the new container alongside the old one, and kamal-proxy moves the host’s traffic across as soon as the health check passes. At that moment, the new workers have handled health checks but no customer requests. preload_app! means the master loads the app once and the workers inherit it through fork. That covers the code. It doesn’t cover anything Ruby and Rails set up on first use: YJIT compiled code. YJIT compiles a method once it’s been called a certain number of times. The master calls very little during boot, so every worker compiles the same methods again on its own first requests. Compiled templates. Action View turns each ERB template into a Ruby method the first time it’s rendered. The schema cache. Active Record reads each model’s columns from the database the first time that model is used. Inline caches and memoized values throughout Ruby, Rails and the app. All 63 workers did all of this at once, while serving the traffic the old container had been handling a second earlier. In the test environment with YJIT on, the first request to a project page on a cold process took 652 ms, 151 ms of it YJIT compiling. The same request to a warm process took 28 ms. In production, CPU time per request peaked at around 200 ms while kamal-proxy moved traffic to the new container, against about 30 ms once the workers had warmed up. A host with spare CPU absorbs this. Every one of our web hosts has 48 cores and 63 workers, but each Amsterdam host serves around 250 requests per second, against 25 to 60 at our other sites. In Amsterdam the slow first requests turned into a queue. At a peak-hour deploy, the Puma backlog on an Amsterdam host reached anywhere from 250 to 2,238 requests, and kamal-proxy’s p99 response time hit about 10 seconds. Eron, our Director of Operations, had been tracking this since June. Another server in Amsterdam would help, but it would take weeks to arrive, so we also wanted to make deploys cheaper on the hardware we already had. What didn’t work We tried a few things first. In June, Donal tested the first two on a single Amsterdam host, comparing it with its neighbors, and they ruled out two likely causes. Warming each worker’s database connections. Puma’s before_fork hook clears the master’s connections, and each worker opened its own on its first request. Opening them in before_worker_boot instead made no difference. Queries on a freshly booted production host were already under a millisecond, so connections weren’t the problem. A synthetic request in each worker. Next, each worker made a few requests in before_worker_boot to an internal controller that touched every model. That ran the middleware, routing and Active Record paths, but it ran them in 63 workers at once — exactly the CPU spike we were trying to avoid. And a request with no real data renders no real views, so most of the app stayed cold. Spreading YJIT compilation out. Delaying YJIT in each worker by a random interval spread the compiling out over a few minutes, but every worker still ran interpreted until its delay ended. The queue didn’t change. Reforking from a warm worker. This is what Shopify’s Pitchfork does: let one worker serve traffic until it’s warm, then fork the others from it. Puma has an experimental version called fork_worker, and on beta it worked — the reforked workers were warm after three to five requests, where fresh ones took up to 30 seconds. But with fork_worker the template is worker 0, and it keeps serving requests. If it exits, the workers waiting to be forked never start (puma/puma#3596). If it gets no traffic, the refork never happens, which is what we saw on beta. Instacart have a mold_worker patch that promotes a warm worker to a template that stops serving, but it isn’t in a Puma release. We have a branch of it, and we may come back to it. That last experiment did show us where the fix was, though. Everything a warm worker has that a cold one lacks is in its memory, and fork copies memory. The master already has the app loaded. It just never runs it. Run the requests in the master So now, before the master binds its socket and forks, it makes the app’s own requests, in-process, the way a signed-in user would. Rack has a hook for exactly this. Rack::Builder#warmup takes a block that’s called once with the built app, before the server starts. rails server builds the app from config.ru, so the change to boot is one line: require_relative "config/environment" warmup { WarmUp.configured.run } if ENV["WARM_UP"] run Rails.application With preload_app! this runs in the master, and the workers inherit whatever it did. Puma binds its socket after the app is built, so until the warm-up finishes the health check’s connection is refused and kamal-proxy keeps retrying. No request reaches a worker that hasn’t been warmed. The warm-up has three steps. After precompiling the views, it gives the page requests and schema loading a shared 20-second budget, checked before each page or model. 1. Precompile the views actionview_precompiler reads every template for its render calls and compiles each one with the locals it’s passed. For us that’s 1,394 templates in about two seconds. A first request to a project page then compiles 2 templates instead of 44. 2. Request the pages, signed in A small browser class makes the requests through Rack::MockRequest, with the two cookies a real sign-in sets, then goes back for each page’s lazy Turbo frames: class WarmUp::Browser def initialize(signed_in_as:) @client = Rack::MockRequest.new(Rails.application) @headers = { "HTTP_USER_AGENT" => "Basecamp warm-up", "HTTP_COOKIE" => cookie_for(signed_in_as), "bc3.warm_up" => true } end def visit(path) page = get(path) frames_in(page).each { |id, src| get(src, "HTTP_TURBO_FRAME" => id) } end private def get(path, headers = {}) @client.get("https://#{host}#{path}", @headers.merge(headers)) end def frames_in(page) Nokogiri::HTML5(page.body).css("turbo-frame[src]").map { |frame| [ frame["id"], frame["src"] ] } end end The requests are signed in. The user is a monitoring account we already use for automated checks, and the pages are its own project, Campfire, to-dos, documents and messages. Public pages weren’t enough: after warming up with signed-out pages only, the first signed-in request to the projects page still took 131 ms, because authentication, the signed-in controllers and their views had never run. With signed-in pages it took 40 ms. cookie_for writes the same signed cookie the sign-in controller does, using the app’s own cookie jar, so there’s no API token and no secret to store. The frames are followed. The busiest HTML requests in production aren’t pages at all but Turbo frames — the sidebar badge, the inbox, the navigation menus. The browser parses each page and requests its <turbo-frame src> URLs with the Turbo-Frame header, so those controllers and views get warmed too. Our first four pages turned into 60 requests. The requests are excluded from rate limiting. They are internal, so they do not count against the rate limits that apply to real visitors. 3. Load the rest of the schema The page requests load the schema for the models they touch. The last step loads the rest, from the read replica: ApplicationRecord.reading do models.lazy.take_while { time_left? }.each { |model| model.load_schema if model.table_exists? } end The step checks 261 models and loads any schema information still missing. Those database round trips add up when the primary is far away: outside a request, Active Record uses the writing role, and from a host a long way from the primary each round trip is tens of milliseconds. Reading from the local replica brings the step down from about 20 seconds to 3.5. The pages go first because they load most of the schema anyway. If the time budget runs out, the step stops, logs how many models it got through, and the workers load the rest on first use like they always did. Rails can also load the schema from a dumped cache file at boot (bin/rails db:schema:cache:dump), which would make this step unnecessary. We don’t ship one in our image yet, because the dump needs a database to read from at build time, and we have several databases to cover. It’s on the list. What to close before the fork Running requests in the master opens things the master never opened before, and every worker inherits them. Two processes writing to the same socket will corrupt each other’s traffic, so you need to know what’s open before you fork. The way to find out is to list the master’s open file descriptors — ls -l /proc/<pid>/fd — before and after a warm-up, in an environment set up like production. Development wasn’t enough for us: it stores files on disk, so our S3 connections only showed up in production. Then, for each thing that’s open, check how its library handles a fork. We found three kinds: Already handled. Plenty of libraries detect a fork on their own, either by recording the PID they connected from and reconnecting in the child, by opening per-process files, or by resetting their thread pools. Redis clients, metrics libraries and concurrency libraries tend to be in this group. Check, but you probably don’t need to do anything. Already closed. Database connections are the classic one, and most Puma configs already clear them in before_fork. Anything else that’s opened per process — we have a SQLite cache the workers open on boot — needs closing when the warm-up finishes. Needs a new step. HTTP clients with keep-alive connections are the ones to look for: cloud SDKs with connection pools, tracing exporters, error reporters. They usually have no fork handling at all. We empty the aws-sdk connection pools in before_fork, and we run the warm-up untraced so the OpenTelemetry exporter never opens its connection to Tempo in the first place. Once that’s done, before_fork finishes with Process.warmup, which Ruby 3.3 added for this purpose: a major GC, a heap compaction, and every surviving object promoted to the old generation, so the memory pages the workers share change as little as possible afterwards. Choosing the pages The first list was the four pages that ran the busiest requests on beta. Once the warm-up was live, production showed us which endpoints were still cold. For one deploy, we compared each endpoint’s mean duration in the six minutes after kamal-proxy moved traffic to the new container with the same endpoint an hour later, then multiplied the difference by the number of requests in those six minutes. That gives the extra time each endpoint cost us because it was cold: Endpoint Cold Warm Requests in 6 min Extra seconds Campfire 246 ms 70 ms 6,490 1,140 Projects (JSON API) 84 ms 50 ms 22,077 771 Docs & Files 262 ms 177 ms 4,996 421 To-dos tool 205 ms 113 ms 4,018 371 To-dos (JSON API) 33 ms 16 ms 18,738 320 The pages already in the warm-up showed what to expect: the project page kept a 36 ms gap after a deploy, and the to-do page 10 ms. We’ve proposed adding these five requests, and expect them to add about five to seven seconds to the page step. The two JSON endpoints were a surprise. The warm-up’s page list had no API requests in it, so nothing on the API path had run before the first real request: not the API controllers, and not the Jbuilder templates rendering real records. Precompiling the views covers JSON templates too, but it isn’t a substitute for running the request. Results The warm-up is on for all 68 web hosts. With the first four pages it took 12 to 16 seconds per host: about 2 seconds to precompile the views, 7 to 9 for the 60 requests, and 3.5 for the schema. Deploys take that much longer per host, and we raised the deploy timeout from 30 to 60 seconds to cover it. In Amsterdam, at a peak-hour deploy: During deploy Before After Peak Puma backlog per host 250–2,238 requests 19–223 requests Peak kamal-proxy p99 about 10 s 2.4–4.8 s Peak CPU time per request 201–214 ms 88–132 ms Peak database time per request 56–69 ms 39–47 ms The same eight hosts at three deploys on 1 October, an hour apart, as the warm-up went from one host to four to all eight. The deploy in the middle, with four hosts warmed and four not, shows why every host needed the warm-up. Each warmed host recovered faster on its own: mean request duration peaked at 130 to 173 ms, against 203 to 311 ms on the hosts that weren’t warmed. But the backlog on all eight was about the same, because they were all waiting on the same database. Mean request duration on each host at the 07:21 UTC deploy. Blue hosts warmed up in the master before forking, orange hosts did not. Memory came down too. The workers now share compiled templates, YJIT code and the schema with the master instead of each building their own copy. On beta, the view precompiler alone took a busy worker’s private memory from 174–202 MB to 119–135 MB. Thirty minutes after the deploy, the web containers used about 39 GB less memory than the previous day’s containers at the same age and traffic. Amsterdam served most of our traffic at the times we tested. In Amsterdam, each new container used about 2 GB less just after traffic moved to it, which lowers the peak while the old and new containers overlap. Working with Claude Claude Code helped throughout. It combed through the per-worker backlogs and per-endpoint timings in Prometheus and Loki after each deploy, worked out the cold-versus-warm cost of each endpoint, and prepared the changes and the pull request descriptions with the benchmarks in them. We decided what to try, deployed it and read the results. If you do this Warm the master before it forks. Compile common code and templates and load their schema in the master, so workers inherit that work. With preload_app!, Rack::Builder#warmup runs before the workers start accepting traffic. Use the app’s real requests. Public pages, internal endpoints and synthetic queries warm the paths they run and nothing else. Signed-in requests to real records, frames included, run what production runs. Measure the cold penalty per endpoint. The difference between an endpoint’s cold and warm duration, times its request count after a deploy, ranks the pages worth adding. Ours weren’t the ones we’d have guessed, and two of them were JSON. Check what the warm-up leaves open. List the master’s file descriptors after a warm-up and account for every one before the fork. Two of ours needed changes. Set a time budget. A warm-up that runs long on one slow host fails the deploy on that host. Ours gives the page requests and schema loading a shared 20-second budget, checked before each page or model, puts the most valuable pages first, and logs what it skipped. Reforking from a warm worker, as Pitchfork does, solves the same problem continuously rather than once at boot, and it would warm paths no fixed list of pages covers. We may still get there: our branch brings Instacart’s mold_worker up to date with Puma’s main branch and fixes the bugs we found in it. But warming the master works with the Puma we already run, took a few days to implement, and substantially reduced the queues after deployment.

yesterday • 1 votes
📚 BoredReading

You seem to be enjoying this.

Join free to unlock everything.

Create free account

Already have an account? Sign in