Full Width [alt+shift+f] Shortcuts [alt+shift+k]
Sign Up [alt+shift+s] Log In [alt+shift+l]
1

How I Use GitHub Issues

from A Beautiful Site [alt+shift+b] in programming

I like using GitHub issues for actionable things such as bugs and backlog items I've committed to doing. Non-maintainers are encouraged to open issues only for bugs, whereas everything else is a discussion. Bug Report 👉 Issue Help/support 👉 Discussion Ask a question 👉 Discussion Ideas/suggestions 👉 Discussion Waxing Philosophical 👉 Discussion Alas, no matter how clear you make your templates, some folks will inevitably open issues anyway. Fear not! You can use this gem of a feature in GitHub's sidebar to convert it to a discussion. I use this very liberally. (Tip: make sure you take the time to categorize your discussions. Organization matters!) Nobody ever got mad that their issue got turned into a discussion. Strategic decisions # Sometimes things need to be addressed but there isn't anything actionable yet. Perhaps a decision needs to be made. I think it's best to hash that out in discussions. RFCs, proposals, philosophical discussions are all just that…discussions. Once a decision is made, you can create an actionable issue from the discussion summarizing the problem and the agreed solution, keeping the [often messy] details out of the issue but cleanly linking back to them for reference. The keyword is actionable. This can be anything from "implement feature X" to "determine the best approach to Y." In the case of the latter, the issue should probably contain a follow up, e.g. "once the best approach is determined, implement and document Y". Using this approach combined with milestones has been such a joy compared to the old way of doing issues. Result: GitHub issues serve as a clear backlog of committed, actionable TODOs for the project. Meanwhile, the discussion forum flourishes with questions, suggestions, and other thoughts from the community. Win-win. nb4 — I'm aware labels exist and I find them more useful as a tool to organize a backlog and triage than to separate actual "issues" from "discussions." We used to do this before GitHub had...
4th Aug 2025

Stay updated

Get a weekly newsletter with the top 5 articles worth reading every week.

More from A Beautiful Site

Most Software Was Already Slop

As software engineers, we take pride in hand-written code and knowing how everything we build works in detail. While there are many developers out there who build good software, there are many, many more who don’t. Yet we balk at the idea of vibe coding, even when it yields decent results. You don’t need a scientific study to realize the mean quality of human-authored software leaned towards slop even before AI was around. How many times have you been frustrated by an app or website that just didn’t work right? (This is especially true for software produced by big companies.) The quality of software humans can build is absolutely incredible. But the quality of software humans actually do build is usually not. So no, AI can’t build perfect software. Maybe that will improve in the future, or maybe it will remain just as messy as the human-authored code LLMs have been trained on. I don’t think any of that matters, because if AI can build software that’s at least as good as the mediocre code that came before it, that will be acceptable for most industries. Customers don’t care about the craft, they care about the result. And if the result is faster, cheaper, and good enough…well, that’s a better combination than anything they’ve ever had before. Why wouldn't they choose it?

4 weeks ago • 2 votes
I'm pulling ColorCopy from the macOS App Store

I launched ColorCopy on May 18. It was my second macOS app, but the first one I launched on the Mac App Store. Today, I decided to pull it. The app isn't a new idea, but it bakes four color tools into a single menu bar app: an eyedropper, a color picker, a palette manager, and a contrast checker, each one a hotkey away. I made it available for free, with a one-time in-app purchase to unlock unlimited use. No subscription, no recurring fees. It's simple, stable software with the lowest possible barrier to entry…exactly the kind of utility the App Store seems made for. I wasn't sure what to expect, so I ran an experiment to answer some questions: Is there any real benefit to having your Mac app on the App Store? Is the Apple tax worth it? Will customers just come flooding in? For ColorCopy's release, I did zero marketing. No Product Hunt launch. No Hacker News post. The only things I published were this blog post and a tweet. If the App Store delivers on its promise of discoverability, that should be enough to see some kind of traffic…right? You be the judge. In nearly three months on the App Store, ColorCopy got a little over 1,000 impressions, 149 product page views, 85 first-time downloads, three in-app purchases, and $21 in proceeds. Not per day. In total. Broken down over 81 days, that averages out to: 13 impressions per day 1.8 product page views per day 1 download per day 0.04 in-app purchases per day (about one per month) $0.26 in proceeds per day Extrapolated to 12 months, ColorCopy would have earned about $95…not even enough to cover Apple's $99 annual developer fee. If discoverability isn't a part of the App Store deal, what exactly is the benefit? Why limit your Mac app to the sandbox?* Why spend hours in App Store Connect filling out metadata, screenshots, and localization information? Why wait an arbitrary number of days for someone to review your app with the consistency of a coin flip? It doesn't seem worth it for Mac developers. Adios, App Store 👋 As of today, ColorCopy is self-distributed and sold through Polar (the new Stripe, which I highly recommend). This is the same way I sell TongueType, which makes pretty much everything easier on my end: one dashboard, real customer relationships, no finicky review process, and updates ship the moment they're ready. The move required swapping out in-app purchases from StoreKit to Polar, and updates now ship through Sparkle instead of the App Store's built-in update mechanism. Both are tried and true solutions for self-distributed Mac apps. To be fair, you can't really do this on iOS. Most users aren't jailbroken…walled garden and all. But on Mac, where self-distribution is still a first-class option, there seems to be zero incentive to be in the App Store, especially if you're counting on discoverability. In three months, the App Store sent me barely a trickle of customers. Had I marketed the app myself, the traffic would have flowed the other way. I would've been sending my customers to Apple and paying a tax for the privilege.** Self-distribution gives me the freedom and control over my apps that the App Store's sandbox never will. If it's on me to drum up all of my own traffic, I'm going to send it to my own website. *I originally used a third-party library for the eye dropper because NSColorSampler is meh, but the sandbox forbids it so I was forced to remove it. Yes, I had to make my app shittier in order to put it on the App Store. **I fully acknowledge this may not be the case for every app, but it was my experience and worth sharing. Your mileage may vary. Aside: ColorCopy was available in 10 languages on the App Store. For some reason, the listing always showed FR as its primary language (it wasn't). The app is available in many languages, including English.

7th Aug 2026 • 1 votes
I Mostly Stopped Typing

I built the dictation app I wanted. It's called TongueType, and my daughter did the voice over for the video. (Family business.) It hasn't gotten much traction yet, and I think I know why: dictation has been bad for so long that most people stopped paying attention. I don't blame them. But I don't think most people realize how good local models have gotten. The thing that was flaky and frustrating five years ago is genuinely good now, and it runs entirely on your Mac. Downloads are low. But the conversion rate is great. The people who actually try it tend to stick around, which tells me the problem isn't the app, it's getting someone to give dictation one more honest chance. The real hurdle is the habit The hard part isn't accuracy. It's that talking instead of typing is a new habit, and new habits are awkward before they're automatic. For the first week it feels strange. You catch yourself reaching for the keyboard out of muscle memory. Then one day it clicks, and you realize how slow typing was making you for certain things. I still write code by hand. That's thinking, not transcribing, and I want my fingers on it (plus saying HTML tags and attributes out loud seems counterintuitive 😂). But for almost everything else, I talk. I prompt LLMs I send emails I reply on Slack I write commit messages I do most other text with my voice The common thread is that the thinking is already done and the only thing left is getting the words out. That turns out to be a surprising amount of my day. My desk setup On the go, the MacBook's built-in mic works just fine. You don't need fancy hardware to get good results. But when I'm at my desk, my laptop is docked, so I pair TongueType with a Tula mic. It's small, portable, sounds great, and looks the part! (Kuru Toga mechanical pencil positioned for size comparison.) One tip: use a wired mic if you can. Bluetooth adds latency, and latency is the enemy of a tool you reach for dozens of times a day. A wired connection makes the whole thing feel instant. Honest caveat: it's probably not a great look in a co-working space. 😂 "Why is this better than macOS dictation?" More accurate. Less flaky. More customizable. More fun. I could list the features, but the honest answer is the same one I'd give for most things: just try it and you'll feel the difference immediately. It's free to start. What's your experience been with dictation? If you wrote it off years ago, I'd genuinely love for you to give TongueType a shot and tell me what you think.

25th Jun 2026 • 1 votes
Introducing TongueType

I just launched my first macOS app called TongueType. It's voice dictation that runs entirely on your Mac. Hold a key, speak, release, and your words appear wherever your cursor happens to be. I build small, simple, stable software. TongueType fits that description, and it scratches an itch I've had for a while. Why I made it I type fast, but I often think faster than I type. When an idea is fully formed in my head, the bottleneck is my fingers. macOS has had built-in dictation forever, but I never liked relying on it. Accuracy aside, I didn't love the idea of my voice taking a trip to a server and back just to write a sentence. There are many dictation apps in the wild, but I want one that's privacy focused, doesn't send data to the cloud, doesn't charge a monthly subscription, and gets out of the way. TongueType uses OpenAI's Whisper model running locally on Apple Silicon. Nothing is uploaded. Nothing is logged. There's no account to create. Zero telemetry. Your voice never leaves your Mac. How it works The whole interaction is one key. By default it's the Right Option key, because it's sitting right there and your thumb isn't doing anything important. Hold it, talk, let go. The transcribed text is inserted at your cursor…in your editor, your email, a chat box, a search field, anywhere text goes. You can also drop in an audio or video file — WAV, MP3, MP4, MOV — and get a transcript back. Handy for meeting recordings and voice memos. A few things I sweated the details on: A grace period so a quick accidental tap doesn't start recording. Double-tap to latch for when you want to keep talking without holding the key down. Cancel phrases — say "scratch that" at the end and the whole thing gets discarded. You will use this more than you expect. Spoken symbols — say "new line" or "question mark" and you get the symbol, not the words. Post-processing — for common terms that seldom get dictated properly. TongueType speaks twelve languages and includes automatic detection, so you don't have to tell it which one you're using. How I actually use it Building the app was one thing. Using it every day turned out to be another. A couple months in, it's quietly worked its way into most of what I do at the keyboard: Prompting LLMs. Talking to an AI assistant is conversational by nature, and typing out a long, detailed prompt is tedious. Speaking it isn't. I get more context into a prompt because I'm not rationing my words to save my fingers. Email. Replies that used to sit in my drafts now get spoken out in a single pass. I still read them before sending, but the blank-page friction is gone. Code comments and commit messages. The parts of coding that are just writing. It's faster to explain why a change exists out loud than to stop and type it. Direct messages. Quick replies in chat without breaking flow. Hold the key, say it, done. The common thread: TongueType is best wherever the thinking is already done and the only thing left is getting words out. That's a surprising amount of my work day. A fun personality TongueType is minimal and fun. It lives in the menu bar. The recording overlay is small and out of the way, and you can configure its position on screen. There are twenty accent colors including Rainbow Mode. None of these extras were necessary, but all of it was fun to build. Accessibility I want to call this out specifically. Voice dictation isn't only a convenience. For some people it's the difference between using a computer comfortably or not. If typing is painful or difficult for you, TongueType is built to be a genuine alternate input method, not an afterthought. That mattered to me, and it shaped a lot of the decisions above. Pricing TongueType is free to try, and the free tier includes every feature. You get 30 minutes of live dictation each month and short file transcriptions. If you want unlimited, TongueType Pro is a one-time $19.99 purchase that covers up to five Macs and unlocks unlimited dictation and full-length file transcription. No subscription. Buy it once, keep it forever. Requirements TongueType needs macOS 14 or later on an Apple Silicon Mac (M1 or newer). The local model is the whole point, and that's what makes it run so well. If any of this sounds useful, give it a try at TongueType.app. It's free to start, and I'd genuinely like to hear what you think.

14th May 2026 • 1 votes
My Stance on AI in Software Development

I believe artificial intelligence is a powerful and valuable tool that can significantly improve how we create, solve problems, and bring ideas to life. I didn't always feel this way.. But these days, I use AI regularly in my work and I expect that to continue. We may not have chosen this reality, but it's the reality we're in. The tools and their benefits — costs be damned — can no longer be ignored. That said, when people ask "was this made with AI?" the honest answer is rarely simple. AI can speed up many parts of the process, but it doesn’t replace human judgment, creativity, or responsibility. What appears effortless on the surface rests on deliberate human direction, critical thinking, and careful review. Getting good results from AI requires active guidance and oversight. The nuances of context, ethics, user needs, and real-world application are simply too varied given the current technology. Moreover, AI doesn’t generate meaningful ideas or elegant solutions on its own. Strong human vision, architecture, and decision-making are still essential. There is no prompt, model, or service that can deliver finished, trustworthy work without substantial human input. My commitment to you is this: everything I create will be driven by human ideas, architecture, verification, and final review. I will use AI as an assistant to do what I would have done anyway, but more efficiently. I will not let AI replace my intelligence, but I will use it to turn my intelligence into code faster. — Cory LaViska

20th Apr 2026 • 1 votes

More in programming

Clip of me singing Despard in Ruddigore in 2013

A clip of me singing a funny song from Gilbert and Sullivan’s Ruddigore back in 2013

7 hours ago • 1 votes
How and Why fork() Uses Copy-on-Write

In this video, we look at why fork() needs copy-on-write, how it works inside the kernel, and a memory usage problem that Instagram encountered with Python.

13 hours ago • 1 votes
What we lost when we lost comments

Comments require commitment, but they’re worth it.

19 hours ago • 1 votes
Lighthouse map

Lovely global map with animated lights sweeping the waters

22 hours ago • 1 votes
Warming up the Puma master before it forks

Basecamp 5 runs on Puma in cluster mode: one master process with preload_app! and 63 single-threaded workers per host, deployed as a Docker container with Kamal. We serve Basecamp from several sites. Each site has its own web hosts and a read replica of the database, and writes go to a single primary database in one of them. On our busiest hosts, each deploy left up to 2,000 requests waiting while the new workers warmed up. We reduced those queues by running signed-in requests through the app in the Puma master, before it forked the workers. Why 63 single-threaded workers? Basecamp has always served web requests from processes rather than threads. It ran on Unicorn, which only does processes, until we moved to Puma in January 2025, and we kept the same setup: workers (Concurrent.physical_processor_count * 1.3).ceil threads 1, 1 preload_app! On a 48-core host that’s 63 workers, each handling one request at a time. We chose 1.3 after benchmarking HEY in 2023, when we moved our apps out of the cloud and onto our own hardware. We tested several combinations of workers and threads with a mix of GET and POST requests on a 32-vCPU VM. Every multithreaded configuration we tested was slower and handled fewer requests than single-threaded workers. Adding workers beyond about 1.2 to 1.3 per vCPU brought little benefit. The threaded workers spent a lot of their time waiting for Ruby’s global VM lock. That made single-threaded workers a good fit for this workload, and we use the same setup for Basecamp. An app that spends more time waiting on its database or other services may benefit from more threads, so benchmark your own app. The other reason is the app itself. Basecamp has class-level state in places and has never needed to be thread-safe. With one request per process, it still doesn’t. Processes do use more memory than threads, and preload_app! reduces the difference. The master loads the app once and the workers share its memory through copy-on-write until they write to it. Shopify’s comparison of Ruby execution models explains the trade-off well. In the HEY benchmark the best setup came to about 260 MB of PSS per core, where PSS counts each shared page once, split between the processes using it, and the gap to a threaded setup was smaller than we’d expected. What Puma does on each host when a container starts: one master, then 63 forked workers that share its memory until they write to it. Two things about this setup matter for the rest of the post. A worker that’s compiling or loading something is fully blocked — there’s no other thread to pick up the next request. And whatever the master has in memory before it forks, all 63 workers share. Whatever they build after the fork, they build 63 times. What happens when we deploy Kamal starts the new container alongside the old one, and kamal-proxy moves the host’s traffic across as soon as the health check passes. At that moment, the new workers have handled health checks but no customer requests. preload_app! means the master loads the app once and the workers inherit it through fork. That covers the code. It doesn’t cover anything Ruby and Rails set up on first use: YJIT compiled code. YJIT compiles a method once it’s been called a certain number of times. The master calls very little during boot, so every worker compiles the same methods again on its own first requests. Compiled templates. Action View turns each ERB template into a Ruby method the first time it’s rendered. The schema cache. Active Record reads each model’s columns from the database the first time that model is used. Inline caches and memoized values throughout Ruby, Rails and the app. All 63 workers did all of this at once, while serving the traffic the old container had been handling a second earlier. In the test environment with YJIT on, the first request to a project page on a cold process took 652 ms, 151 ms of it YJIT compiling. The same request to a warm process took 28 ms. In production, CPU time per request peaked at around 200 ms while kamal-proxy moved traffic to the new container, against about 30 ms once the workers had warmed up. A host with spare CPU absorbs this. Every one of our web hosts has 48 cores and 63 workers, but each Amsterdam host serves around 250 requests per second, against 25 to 60 at our other sites. In Amsterdam the slow first requests turned into a queue. At a peak-hour deploy, the Puma backlog on an Amsterdam host reached anywhere from 250 to 2,238 requests, and kamal-proxy’s p99 response time hit about 10 seconds. Eron, our Director of Operations, had been tracking this since June. Another server in Amsterdam would help, but it would take weeks to arrive, so we also wanted to make deploys cheaper on the hardware we already had. What didn’t work We tried a few things first. In June, Donal tested the first two on a single Amsterdam host, comparing it with its neighbors, and they ruled out two likely causes. Warming each worker’s database connections. Puma’s before_fork hook clears the master’s connections, and each worker opened its own on its first request. Opening them in before_worker_boot instead made no difference. Queries on a freshly booted production host were already under a millisecond, so connections weren’t the problem. A synthetic request in each worker. Next, each worker made a few requests in before_worker_boot to an internal controller that touched every model. That ran the middleware, routing and Active Record paths, but it ran them in 63 workers at once — exactly the CPU spike we were trying to avoid. And a request with no real data renders no real views, so most of the app stayed cold. Spreading YJIT compilation out. Delaying YJIT in each worker by a random interval spread the compiling out over a few minutes, but every worker still ran interpreted until its delay ended. The queue didn’t change. Reforking from a warm worker. This is what Shopify’s Pitchfork does: let one worker serve traffic until it’s warm, then fork the others from it. Puma has an experimental version called fork_worker, and on beta it worked — the reforked workers were warm after three to five requests, where fresh ones took up to 30 seconds. But with fork_worker the template is worker 0, and it keeps serving requests. If it exits, the workers waiting to be forked never start (puma/puma#3596). If it gets no traffic, the refork never happens, which is what we saw on beta. Instacart have a mold_worker patch that promotes a warm worker to a template that stops serving, but it isn’t in a Puma release. We have a branch of it, and we may come back to it. That last experiment did show us where the fix was, though. Everything a warm worker has that a cold one lacks is in its memory, and fork copies memory. The master already has the app loaded. It just never runs it. Run the requests in the master So now, before the master binds its socket and forks, it makes the app’s own requests, in-process, the way a signed-in user would. Rack has a hook for exactly this. Rack::Builder#warmup takes a block that’s called once with the built app, before the server starts. rails server builds the app from config.ru, so the change to boot is one line: require_relative "config/environment" warmup { WarmUp.configured.run } if ENV["WARM_UP"] run Rails.application With preload_app! this runs in the master, and the workers inherit whatever it did. Puma binds its socket after the app is built, so until the warm-up finishes the health check’s connection is refused and kamal-proxy keeps retrying. No request reaches a worker that hasn’t been warmed. The warm-up has three steps. After precompiling the views, it gives the page requests and schema loading a shared 20-second budget, checked before each page or model. 1. Precompile the views actionview_precompiler reads every template for its render calls and compiles each one with the locals it’s passed. For us that’s 1,394 templates in about two seconds. A first request to a project page then compiles 2 templates instead of 44. 2. Request the pages, signed in A small browser class makes the requests through Rack::MockRequest, with the two cookies a real sign-in sets, then goes back for each page’s lazy Turbo frames: class WarmUp::Browser def initialize(signed_in_as:) @client = Rack::MockRequest.new(Rails.application) @headers = { "HTTP_USER_AGENT" => "Basecamp warm-up", "HTTP_COOKIE" => cookie_for(signed_in_as), "bc3.warm_up" => true } end def visit(path) page = get(path) frames_in(page).each { |id, src| get(src, "HTTP_TURBO_FRAME" => id) } end private def get(path, headers = {}) @client.get("https://#{host}#{path}", @headers.merge(headers)) end def frames_in(page) Nokogiri::HTML5(page.body).css("turbo-frame[src]").map { |frame| [ frame["id"], frame["src"] ] } end end The requests are signed in. The user is a monitoring account we already use for automated checks, and the pages are its own project, Campfire, to-dos, documents and messages. Public pages weren’t enough: after warming up with signed-out pages only, the first signed-in request to the projects page still took 131 ms, because authentication, the signed-in controllers and their views had never run. With signed-in pages it took 40 ms. cookie_for writes the same signed cookie the sign-in controller does, using the app’s own cookie jar, so there’s no API token and no secret to store. The frames are followed. The busiest HTML requests in production aren’t pages at all but Turbo frames — the sidebar badge, the inbox, the navigation menus. The browser parses each page and requests its <turbo-frame src> URLs with the Turbo-Frame header, so those controllers and views get warmed too. Our first four pages turned into 60 requests. The requests are excluded from rate limiting. They are internal, so they do not count against the rate limits that apply to real visitors. 3. Load the rest of the schema The page requests load the schema for the models they touch. The last step loads the rest, from the read replica: ApplicationRecord.reading do models.lazy.take_while { time_left? }.each { |model| model.load_schema if model.table_exists? } end The step checks 261 models and loads any schema information still missing. Those database round trips add up when the primary is far away: outside a request, Active Record uses the writing role, and from a host a long way from the primary each round trip is tens of milliseconds. Reading from the local replica brings the step down from about 20 seconds to 3.5. The pages go first because they load most of the schema anyway. If the time budget runs out, the step stops, logs how many models it got through, and the workers load the rest on first use like they always did. Rails can also load the schema from a dumped cache file at boot (bin/rails db:schema:cache:dump), which would make this step unnecessary. We don’t ship one in our image yet, because the dump needs a database to read from at build time, and we have several databases to cover. It’s on the list. What to close before the fork Running requests in the master opens things the master never opened before, and every worker inherits them. Two processes writing to the same socket will corrupt each other’s traffic, so you need to know what’s open before you fork. The way to find out is to list the master’s open file descriptors — ls -l /proc/<pid>/fd — before and after a warm-up, in an environment set up like production. Development wasn’t enough for us: it stores files on disk, so our S3 connections only showed up in production. Then, for each thing that’s open, check how its library handles a fork. We found three kinds: Already handled. Plenty of libraries detect a fork on their own, either by recording the PID they connected from and reconnecting in the child, by opening per-process files, or by resetting their thread pools. Redis clients, metrics libraries and concurrency libraries tend to be in this group. Check, but you probably don’t need to do anything. Already closed. Database connections are the classic one, and most Puma configs already clear them in before_fork. Anything else that’s opened per process — we have a SQLite cache the workers open on boot — needs closing when the warm-up finishes. Needs a new step. HTTP clients with keep-alive connections are the ones to look for: cloud SDKs with connection pools, tracing exporters, error reporters. They usually have no fork handling at all. We empty the aws-sdk connection pools in before_fork, and we run the warm-up untraced so the OpenTelemetry exporter never opens its connection to Tempo in the first place. Once that’s done, before_fork finishes with Process.warmup, which Ruby 3.3 added for this purpose: a major GC, a heap compaction, and every surviving object promoted to the old generation, so the memory pages the workers share change as little as possible afterwards. Choosing the pages The first list was the four pages that ran the busiest requests on beta. Once the warm-up was live, production showed us which endpoints were still cold. For one deploy, we compared each endpoint’s mean duration in the six minutes after kamal-proxy moved traffic to the new container with the same endpoint an hour later, then multiplied the difference by the number of requests in those six minutes. That gives the extra time each endpoint cost us because it was cold: Endpoint Cold Warm Requests in 6 min Extra seconds Campfire 246 ms 70 ms 6,490 1,140 Projects (JSON API) 84 ms 50 ms 22,077 771 Docs & Files 262 ms 177 ms 4,996 421 To-dos tool 205 ms 113 ms 4,018 371 To-dos (JSON API) 33 ms 16 ms 18,738 320 The pages already in the warm-up showed what to expect: the project page kept a 36 ms gap after a deploy, and the to-do page 10 ms. We’ve proposed adding these five requests, and expect them to add about five to seven seconds to the page step. The two JSON endpoints were a surprise. The warm-up’s page list had no API requests in it, so nothing on the API path had run before the first real request: not the API controllers, and not the Jbuilder templates rendering real records. Precompiling the views covers JSON templates too, but it isn’t a substitute for running the request. Results The warm-up is on for all 68 web hosts. With the first four pages it took 12 to 16 seconds per host: about 2 seconds to precompile the views, 7 to 9 for the 60 requests, and 3.5 for the schema. Deploys take that much longer per host, and we raised the deploy timeout from 30 to 60 seconds to cover it. In Amsterdam, at a peak-hour deploy: During deploy Before After Peak Puma backlog per host 250–2,238 requests 19–223 requests Peak kamal-proxy p99 about 10 s 2.4–4.8 s Peak CPU time per request 201–214 ms 88–132 ms Peak database time per request 56–69 ms 39–47 ms The same eight hosts at three deploys on 1 October, an hour apart, as the warm-up went from one host to four to all eight. The deploy in the middle, with four hosts warmed and four not, shows why every host needed the warm-up. Each warmed host recovered faster on its own: mean request duration peaked at 130 to 173 ms, against 203 to 311 ms on the hosts that weren’t warmed. But the backlog on all eight was about the same, because they were all waiting on the same database. Mean request duration on each host at the 07:21 UTC deploy. Blue hosts warmed up in the master before forking, orange hosts did not. Memory came down too. The workers now share compiled templates, YJIT code and the schema with the master instead of each building their own copy. On beta, the view precompiler alone took a busy worker’s private memory from 174–202 MB to 119–135 MB. Thirty minutes after the deploy, the web containers used about 39 GB less memory than the previous day’s containers at the same age and traffic. Amsterdam served most of our traffic at the times we tested. In Amsterdam, each new container used about 2 GB less just after traffic moved to it, which lowers the peak while the old and new containers overlap. Working with Claude Claude Code helped throughout. It combed through the per-worker backlogs and per-endpoint timings in Prometheus and Loki after each deploy, worked out the cold-versus-warm cost of each endpoint, and prepared the changes and the pull request descriptions with the benchmarks in them. We decided what to try, deployed it and read the results. If you do this Warm the master before it forks. Compile common code and templates and load their schema in the master, so workers inherit that work. With preload_app!, Rack::Builder#warmup runs before the workers start accepting traffic. Use the app’s real requests. Public pages, internal endpoints and synthetic queries warm the paths they run and nothing else. Signed-in requests to real records, frames included, run what production runs. Measure the cold penalty per endpoint. The difference between an endpoint’s cold and warm duration, times its request count after a deploy, ranks the pages worth adding. Ours weren’t the ones we’d have guessed, and two of them were JSON. Check what the warm-up leaves open. List the master’s file descriptors after a warm-up and account for every one before the fork. Two of ours needed changes. Set a time budget. A warm-up that runs long on one slow host fails the deploy on that host. Ours gives the page requests and schema loading a shared 20-second budget, checked before each page or model, puts the most valuable pages first, and logs what it skipped. Reforking from a warm worker, as Pitchfork does, solves the same problem continuously rather than once at boot, and it would warm paths no fixed list of pages covers. We may still get there: our branch brings Instacart’s mold_worker up to date with Puma’s main branch and fixes the bugs we found in it. But warming the master works with the Puma we already run, took a few days to implement, and substantially reduced the queues after deployment.

yesterday • 1 votes
📚 BoredReading

You seem to be enjoying this.

Join free to unlock everything.

Create free account

Already have an account? Sign in