Full Width [alt+shift+f] Shortcuts [alt+shift+k]
Sign Up [alt+shift+s] Log In [alt+shift+l]
1

Escaping the local maxima

from Eliran Turgeman [alt+shift+b] in programming

when i was a student, everything was simpler. grind leetcode. build projects. get an offer. i knew the salaries. i knew what “winning” looked like. it was a somewhat straight line from broke student to backend engineer at a top company. and i did it, i 10x my life in the span of 4 years. five years in. and honestly? it’s… fine. it’s more than fine. but it also feels like i’m stuck on a plateau. the growth feels logarithmic. the peak that isn’t the peak things are good, but i can’t stop feeling like i want more, even though i am comfortable. i look back at the student version of me, and i see hunger. direction. i look at me now and i see someone who’s tried a bunch of things: built products that barely anyone used started a newsletter, got some nice traffic, but it didn’t stick thinking about podcasts, courses, maybe a dev agency? dreaming of 10x-ing my life again, but not sure where to invest my time. when i was younger, the path was obvious. now it’s all vague, i could do anything. do i go all-in on indie hacking? live off my rsu’s for a few years and just build? try again with another product? double down on the blog? start a podcast? well, the next level doesn’t seem to come with an instructions book. escaping means risking the fall as a cs student you learn that escaping a local maxima usually means exploring a few downs to find a higher maxima. well, applying it to life is scary. what if i lose everything i worked hard to build? i am no longer a student living off of scholarships, i have more obligations. and also, the scariest part of all is what if this is the best it gets? i prefer to be positive and believe there’s another jump out there. something worth building. something that might actually shift my trajectory again. i just haven’t found it yet. but i’m looking. 🚨 Become a better software engineer. practice building real systems, get code reviews, and mentorship from senior engineers. Get started with 404skill
5th Apr 2025

Stay updated

Get a weekly newsletter with the top 5 articles worth reading every week.

More from Eliran Turgeman

My productivity rules

A friend asked me for some study/productivity tips, and I figured the most productive thing I can do is write a post about it. That way, it might help more people too. So here we go. Before we start, there are two things you need before any productivity advice will work: Be introspective. Be honest with yourself. My productivity rules Plan your day the night before Don’t leave decisions for your foggy, maybe lazy, morning self. Here’s the loop: wake up → review study plan → study → write down the plan for tomorrow → sleep → repeat. Check your energy during the day If you just ate and feel sleepy, don’t force deep study. Take a 30–60 minute break - nap, walk, exercise, or scroll your phone a bit, then come back refreshed. You can also take micro-breaks: finish a chapter, grab water or a piece of chocolate, and get back to it within five minutes. Don’t lie to yourself about effort Setting goals is great, but be honest about how hard you actually worked. You can check all the boxes, feel proud, and still know deep down that you took it easy. That’s fine sometimes, we all need rest days - but don’t confuse that with an intense study session. Review your day When you plan tomorrow’s tasks, reflect on today. Did you actually do what you set out to do? If not, why? Maybe your goal was too ambitious, or maybe you just spent too much time gaming. Either way, learn from it. There’s always room to improve, if the goal really matters to you. Limit distractions Put your phone on silent and out of reach. If you study on your computer, close anything that might tempt you, and even hide shortcuts to distracting apps. When I was in uni, I played a ton of League of Legends. The desktop icon was staring at me every time I opened my laptop — so I buried it under three folders. Sounds dumb, but it worked. A few more things Get good at breaking big goals into small tasks. If your goal is to pass an exam, start by mapping out all the smaller steps that’ll get you there. Spread them out over time, with a bit of buffer. Plans will change — that’s fine — but you should always know where you stand and adjust as you go. And one more time, because it’s that important: don’t lie to yourself. If you spent four hours on TikTok, felt bad, then studied a little to compensate — that’s not a “productive” day. Call it what it is. It’s okay to have those days, you’re human — just plan for them instead of pretending they didn’t happen. I believe that much of the productivity advice online includes some stupid ceremonies and whatnot, just do what you feel works for you. I think the ability to introspect, being honest with yourself and improving over time is all that matters. Take whatever breaks you need, in whatever order you want. You don’t need a fancy journal to write your tasks, open a txt file. You don’t need a perfect system, just something simple that works for you.

8th Oct 2025 • 1 votes
Fighting subscription fatigue with vibe-coding

Most people solve their subscription fatigue by canceling Netflix. I solved it by vibe-coding my own workout app instead of paying another SaaS. Two weeks ago, I decided to get serious about my workouts again and start logging them. I looked for an existing solution that has the following: create workouts log sets, reps, and weights a calendar to track consistency simple enough. Free apps were crammed with ads. Paid apps had bloated features I didn’t want. Both annoyed me. Before vibe-coding, I’d either tolerate ads or pay. Now there’s a third option - build my own. Of course, you could always build your own, but pre-vibe-coding it would take much more time to be worth it. How did I choose the vibe coding platform? I logged into loveable, base44, bolt, and wrote the following (imperfect) prompt 1 2 3 4 5 6 7 I want to create a personal webapp for managing my own workout routines (kind of a workout logger) I want to be able to define "workouts" - collection of exercises including sets and reps I want to be able to track which workout i did on which day - calendar view I want to be able to log the weights I did for every exercise in every set. I want my workout templates to be a simple collection of exercises that are plaintext - don't create some kind of an exercise library.. i just want to type the exercise name myself make it stateful, including a db connection to store all the relevant data. Whichever tool gave me the best first shot, I ran with. This time it was Loveable. Iterating With that one-shot starting point from loveable, I published it, and went to my first workout at the gym, all excited and ready to use what I built. First workout: I needed notes for exercises. Another: supersets. Each time I wrote it down, went home, and 15 minutes later I published a new version with loveable. After two weeks of using this app at the gym and doing tweaks, the app feels solid. On my last iteration, I added a badges/achievement page, and a github-like consistency widget that looks cool I hope will help me stay on track and be consistent. Sharing Friends wanted to try it too - the problem? I didn’t make it secured by sign-in, I thought I only need it for myself, and even if someone’s going to find this weird loveable URL, they could only see my workouts - who cares… But of course for my friends I’ll write one more prompt - and so I added Supabase auth within minutes. Final Thoughts This post isn’t about showing off my little workout logger anyone could make in a few hours prompting. It’s about how easy it is today to scratch your own itch. You build the exact features you need, when you need them. No ads, no bloat, no adapting to someone else’s UX. Just something comfortable and fun to use. It’s never been easier to bring your ideas to life, small or big.

16th Aug 2025 • 1 votes
Why sharing a redis cluster across services is asking for trouble

If there’s one pattern I’ve seen across multiple companies, from scrappy startups to big corps, that causes endless headaches, it’s this: a single cache cluster shared across services. I recently shortly wrote about my lessons from building and maintaining distributed systems at scale, and the first point that came to mind is exactly this - it starts with an excuse of simplicity, “we already have a cache cluster up and running, let’s just make this other service use it, no need for more infra”, and ends with a confused on-call engineer trying to debug which services were affected by the last keys eviction. So I want to double down on this idea and explain in more detail why it becomes a nightmare once your system scales. One eviction policy You got different services each throwing keys at the same redis cluster. A sudden spike/bug just caused a dramatic increase in cache writes - your cluster wasn’t ready for this, it hits maxmemory and now different keys are being removed based on your eviction policy. What’s the problem? there’s no isolation - service A caused the max memory, and now service B, C, D also pay the price - their keys are being removed as well from the cluster, and could affect the latency, and correctness of other flows of your system. Monitoring is harder Our metric fires up — we see a drop in hit rate on the cluster. Which service is causing it? Who’s affected? Instead of thinking about one service, you’re now mentally juggling everything across the entire system. More noise, less clarity. Although monitoring is harder, you could set up application monitors that you send once you write/read from the cache, based on the prefix of the key. potentially if you are organized and each service that uses the cluster has a unique prefix and you can easily identify between the hit rates of different prefixes - that’s great, but you have to work to get there. Debugging is harder This ties back to my first point about the eviction policy. You had 10m keys. something happend. now you got 5m. The effect on the services is really hard to trace. One service might have lost 100k keys, and you barely see a difference in its monitors, but it doesn’t mean your users are not feeling something is off, maybe today the are waiting a bit more for the page to load, but it’s not too long to hit your monitors thresholds. In that case, if you didn’t have a monitor on the cache cluster for keys eviction, you might be totally blind..”oh I see a slight latency increase here, but no monitors popped - guess all is well” So, never use a shared cache cluster? No, that’s not the lesson here. In some cases it is totally fine to use a single cache cluster. For example: You don’t really have a lot of traffic read/written to the cache so most of it is free anyway You store shared static data (for example, feature flags) Also note that some of the points I was making here against using a single cache cluster, can be somewhat mitigated by having good monitoring set in place. For example, having a defined prefix for the cache key per use-case, per service, and publishing metrics in the application level so we have observability to which type of keys (by prefix) are experiecning a low hit ratio. But on the other hand, tracking keys eviction is harder to monitor, since it’s not initiated by your system. Anyway, I hope you get the point. If you are getting started, a single cache cluster is totally fine. Otherwise, spin up another cache cluster, and sleep better at night. 🚨 Become a better software engineer. practice building real systems, get code reviews, and mentorship from senior engineers. Get started with 404skill

1st May 2025 • 1 votes
On over-engineering; Architecture Edition

I recently wrote about over-engineering and striking a good balance between making your code “too” future-proof and not making it future-proof at all. Some time later, I realized it was missing a critical perspective. I hadn’t addressed over-engineering from an architectural point of view, so this post is dedicated precisely to that. Let’s talk about a decision I made for Collecto, my side project. Collecto is still in its early stages, and like most early-stage projects, its future is uncertain. It could grow into something big—or not. That’s where architectural decisions get tricky. You don’t want to overengineer and waste time, but you also don’t want to under-engineer and regret not laying a solid foundation. So what’s the problem? Collecto is a forms-backend service, meaning it handles the creation, management, and processing of forms data for applications. I wanted to add the ability to send emails on certain events. For example, when a new user signs up for your form, you might want to send them a welcome email. The simplest solution? I could write a new service responsible for sending emails and call it directly wherever needed— for example, right after a user signup is saved to the database. This approach works, is easy to set up, and introduces no additional overhead. However, it results in tight coupling, making future changes more challenging. If tomorrow I want to also send a notification to the form owner when they receive a new subscription, I would have to keep adding more responsibilities to the form service code. This bloats the core service, which should ideally focus solely on CRUD operations for forms. On the other end of the spectrum, I could go all-in and build a distributed pub/sub system with a service bus like RabbitMQ or Azure Service Bus. This would give me scalability, decoupling, and all the good stuff. But it’s also a massive investment in time and complexity for a project that doesn’t need it, yet. I didn’t like both options, so I looked for a 3rd alternative and found MediatR which is a mediator pattern implementation in .NET. Why MediatR is a good middle-ground? MediatR facilitates communication between different parts of the application without them needing to reference each other directly. Instead of invoking methods directly, you can send requests or publish notifications, allowing registered handlers to respond accordingly. This approach maintains loose coupling, making the system easier to maintain and evolve. At the same time, it doesn’t introduce the overhead of managing infrastructure like a service bus or message queue. Everything stays in-process, simple, and fast. One of the primary reasons I chose MediatR is its simplicity. Implementing communication patterns with MediatR is straightforward and requires minimal configuration. Compared to a full-fledged service bus, MediatR demands a much smaller time investment and eliminates operational overhead such as monitoring queues or scaling message brokers. It can’t be all sunshines and rainbows MediatR has a few cons compared to other out-of-process messaging brokers, for example Events are in-process only. If your application crashes, you lose the events. There’s no out of the box retry mechanism for failed event handlers. If you deploy multiple instances of Collecto, MediatR won’t distribute events across them. Bottom line Architecture isn’t about perfection—it’s about trade-offs. MediatR worked for Collecto because it gave me a decoupled, flexible way to handle events without the overhead of a service bus. It wasn’t the simplest solution, but it was the right one for where the project is today. The next time you’re making an architectural decision, remember this: the best solution isn’t the most impressive or complex—it’s the one that solves your problem now while leaving room for growth later. 🚨 Become a better software engineer. practice building real systems, get code reviews, and mentorship from senior engineers. Get started with 404skill

10th Dec 2024 • 1 votes

More in programming

Clip of me singing Despard in Ruddigore in 2013

A clip of me singing a funny song from Gilbert and Sullivan’s Ruddigore back in 2013

7 hours ago • 1 votes
How and Why fork() Uses Copy-on-Write

In this video, we look at why fork() needs copy-on-write, how it works inside the kernel, and a memory usage problem that Instagram encountered with Python.

13 hours ago • 1 votes
What we lost when we lost comments

Comments require commitment, but they’re worth it.

19 hours ago • 1 votes
Lighthouse map

Lovely global map with animated lights sweeping the waters

21 hours ago • 1 votes
Warming up the Puma master before it forks

Basecamp 5 runs on Puma in cluster mode: one master process with preload_app! and 63 single-threaded workers per host, deployed as a Docker container with Kamal. We serve Basecamp from several sites. Each site has its own web hosts and a read replica of the database, and writes go to a single primary database in one of them. On our busiest hosts, each deploy left up to 2,000 requests waiting while the new workers warmed up. We reduced those queues by running signed-in requests through the app in the Puma master, before it forked the workers. Why 63 single-threaded workers? Basecamp has always served web requests from processes rather than threads. It ran on Unicorn, which only does processes, until we moved to Puma in January 2025, and we kept the same setup: workers (Concurrent.physical_processor_count * 1.3).ceil threads 1, 1 preload_app! On a 48-core host that’s 63 workers, each handling one request at a time. We chose 1.3 after benchmarking HEY in 2023, when we moved our apps out of the cloud and onto our own hardware. We tested several combinations of workers and threads with a mix of GET and POST requests on a 32-vCPU VM. Every multithreaded configuration we tested was slower and handled fewer requests than single-threaded workers. Adding workers beyond about 1.2 to 1.3 per vCPU brought little benefit. The threaded workers spent a lot of their time waiting for Ruby’s global VM lock. That made single-threaded workers a good fit for this workload, and we use the same setup for Basecamp. An app that spends more time waiting on its database or other services may benefit from more threads, so benchmark your own app. The other reason is the app itself. Basecamp has class-level state in places and has never needed to be thread-safe. With one request per process, it still doesn’t. Processes do use more memory than threads, and preload_app! reduces the difference. The master loads the app once and the workers share its memory through copy-on-write until they write to it. Shopify’s comparison of Ruby execution models explains the trade-off well. In the HEY benchmark the best setup came to about 260 MB of PSS per core, where PSS counts each shared page once, split between the processes using it, and the gap to a threaded setup was smaller than we’d expected. What Puma does on each host when a container starts: one master, then 63 forked workers that share its memory until they write to it. Two things about this setup matter for the rest of the post. A worker that’s compiling or loading something is fully blocked — there’s no other thread to pick up the next request. And whatever the master has in memory before it forks, all 63 workers share. Whatever they build after the fork, they build 63 times. What happens when we deploy Kamal starts the new container alongside the old one, and kamal-proxy moves the host’s traffic across as soon as the health check passes. At that moment, the new workers have handled health checks but no customer requests. preload_app! means the master loads the app once and the workers inherit it through fork. That covers the code. It doesn’t cover anything Ruby and Rails set up on first use: YJIT compiled code. YJIT compiles a method once it’s been called a certain number of times. The master calls very little during boot, so every worker compiles the same methods again on its own first requests. Compiled templates. Action View turns each ERB template into a Ruby method the first time it’s rendered. The schema cache. Active Record reads each model’s columns from the database the first time that model is used. Inline caches and memoized values throughout Ruby, Rails and the app. All 63 workers did all of this at once, while serving the traffic the old container had been handling a second earlier. In the test environment with YJIT on, the first request to a project page on a cold process took 652 ms, 151 ms of it YJIT compiling. The same request to a warm process took 28 ms. In production, CPU time per request peaked at around 200 ms while kamal-proxy moved traffic to the new container, against about 30 ms once the workers had warmed up. A host with spare CPU absorbs this. Every one of our web hosts has 48 cores and 63 workers, but each Amsterdam host serves around 250 requests per second, against 25 to 60 at our other sites. In Amsterdam the slow first requests turned into a queue. At a peak-hour deploy, the Puma backlog on an Amsterdam host reached anywhere from 250 to 2,238 requests, and kamal-proxy’s p99 response time hit about 10 seconds. Eron, our Director of Operations, had been tracking this since June. Another server in Amsterdam would help, but it would take weeks to arrive, so we also wanted to make deploys cheaper on the hardware we already had. What didn’t work We tried a few things first. In June, Donal tested the first two on a single Amsterdam host, comparing it with its neighbors, and they ruled out two likely causes. Warming each worker’s database connections. Puma’s before_fork hook clears the master’s connections, and each worker opened its own on its first request. Opening them in before_worker_boot instead made no difference. Queries on a freshly booted production host were already under a millisecond, so connections weren’t the problem. A synthetic request in each worker. Next, each worker made a few requests in before_worker_boot to an internal controller that touched every model. That ran the middleware, routing and Active Record paths, but it ran them in 63 workers at once — exactly the CPU spike we were trying to avoid. And a request with no real data renders no real views, so most of the app stayed cold. Spreading YJIT compilation out. Delaying YJIT in each worker by a random interval spread the compiling out over a few minutes, but every worker still ran interpreted until its delay ended. The queue didn’t change. Reforking from a warm worker. This is what Shopify’s Pitchfork does: let one worker serve traffic until it’s warm, then fork the others from it. Puma has an experimental version called fork_worker, and on beta it worked — the reforked workers were warm after three to five requests, where fresh ones took up to 30 seconds. But with fork_worker the template is worker 0, and it keeps serving requests. If it exits, the workers waiting to be forked never start (puma/puma#3596). If it gets no traffic, the refork never happens, which is what we saw on beta. Instacart have a mold_worker patch that promotes a warm worker to a template that stops serving, but it isn’t in a Puma release. We have a branch of it, and we may come back to it. That last experiment did show us where the fix was, though. Everything a warm worker has that a cold one lacks is in its memory, and fork copies memory. The master already has the app loaded. It just never runs it. Run the requests in the master So now, before the master binds its socket and forks, it makes the app’s own requests, in-process, the way a signed-in user would. Rack has a hook for exactly this. Rack::Builder#warmup takes a block that’s called once with the built app, before the server starts. rails server builds the app from config.ru, so the change to boot is one line: require_relative "config/environment" warmup { WarmUp.configured.run } if ENV["WARM_UP"] run Rails.application With preload_app! this runs in the master, and the workers inherit whatever it did. Puma binds its socket after the app is built, so until the warm-up finishes the health check’s connection is refused and kamal-proxy keeps retrying. No request reaches a worker that hasn’t been warmed. The warm-up has three steps. After precompiling the views, it gives the page requests and schema loading a shared 20-second budget, checked before each page or model. 1. Precompile the views actionview_precompiler reads every template for its render calls and compiles each one with the locals it’s passed. For us that’s 1,394 templates in about two seconds. A first request to a project page then compiles 2 templates instead of 44. 2. Request the pages, signed in A small browser class makes the requests through Rack::MockRequest, with the two cookies a real sign-in sets, then goes back for each page’s lazy Turbo frames: class WarmUp::Browser def initialize(signed_in_as:) @client = Rack::MockRequest.new(Rails.application) @headers = { "HTTP_USER_AGENT" => "Basecamp warm-up", "HTTP_COOKIE" => cookie_for(signed_in_as), "bc3.warm_up" => true } end def visit(path) page = get(path) frames_in(page).each { |id, src| get(src, "HTTP_TURBO_FRAME" => id) } end private def get(path, headers = {}) @client.get("https://#{host}#{path}", @headers.merge(headers)) end def frames_in(page) Nokogiri::HTML5(page.body).css("turbo-frame[src]").map { |frame| [ frame["id"], frame["src"] ] } end end The requests are signed in. The user is a monitoring account we already use for automated checks, and the pages are its own project, Campfire, to-dos, documents and messages. Public pages weren’t enough: after warming up with signed-out pages only, the first signed-in request to the projects page still took 131 ms, because authentication, the signed-in controllers and their views had never run. With signed-in pages it took 40 ms. cookie_for writes the same signed cookie the sign-in controller does, using the app’s own cookie jar, so there’s no API token and no secret to store. The frames are followed. The busiest HTML requests in production aren’t pages at all but Turbo frames — the sidebar badge, the inbox, the navigation menus. The browser parses each page and requests its <turbo-frame src> URLs with the Turbo-Frame header, so those controllers and views get warmed too. Our first four pages turned into 60 requests. The requests are excluded from rate limiting. They are internal, so they do not count against the rate limits that apply to real visitors. 3. Load the rest of the schema The page requests load the schema for the models they touch. The last step loads the rest, from the read replica: ApplicationRecord.reading do models.lazy.take_while { time_left? }.each { |model| model.load_schema if model.table_exists? } end The step checks 261 models and loads any schema information still missing. Those database round trips add up when the primary is far away: outside a request, Active Record uses the writing role, and from a host a long way from the primary each round trip is tens of milliseconds. Reading from the local replica brings the step down from about 20 seconds to 3.5. The pages go first because they load most of the schema anyway. If the time budget runs out, the step stops, logs how many models it got through, and the workers load the rest on first use like they always did. Rails can also load the schema from a dumped cache file at boot (bin/rails db:schema:cache:dump), which would make this step unnecessary. We don’t ship one in our image yet, because the dump needs a database to read from at build time, and we have several databases to cover. It’s on the list. What to close before the fork Running requests in the master opens things the master never opened before, and every worker inherits them. Two processes writing to the same socket will corrupt each other’s traffic, so you need to know what’s open before you fork. The way to find out is to list the master’s open file descriptors — ls -l /proc/<pid>/fd — before and after a warm-up, in an environment set up like production. Development wasn’t enough for us: it stores files on disk, so our S3 connections only showed up in production. Then, for each thing that’s open, check how its library handles a fork. We found three kinds: Already handled. Plenty of libraries detect a fork on their own, either by recording the PID they connected from and reconnecting in the child, by opening per-process files, or by resetting their thread pools. Redis clients, metrics libraries and concurrency libraries tend to be in this group. Check, but you probably don’t need to do anything. Already closed. Database connections are the classic one, and most Puma configs already clear them in before_fork. Anything else that’s opened per process — we have a SQLite cache the workers open on boot — needs closing when the warm-up finishes. Needs a new step. HTTP clients with keep-alive connections are the ones to look for: cloud SDKs with connection pools, tracing exporters, error reporters. They usually have no fork handling at all. We empty the aws-sdk connection pools in before_fork, and we run the warm-up untraced so the OpenTelemetry exporter never opens its connection to Tempo in the first place. Once that’s done, before_fork finishes with Process.warmup, which Ruby 3.3 added for this purpose: a major GC, a heap compaction, and every surviving object promoted to the old generation, so the memory pages the workers share change as little as possible afterwards. Choosing the pages The first list was the four pages that ran the busiest requests on beta. Once the warm-up was live, production showed us which endpoints were still cold. For one deploy, we compared each endpoint’s mean duration in the six minutes after kamal-proxy moved traffic to the new container with the same endpoint an hour later, then multiplied the difference by the number of requests in those six minutes. That gives the extra time each endpoint cost us because it was cold: Endpoint Cold Warm Requests in 6 min Extra seconds Campfire 246 ms 70 ms 6,490 1,140 Projects (JSON API) 84 ms 50 ms 22,077 771 Docs & Files 262 ms 177 ms 4,996 421 To-dos tool 205 ms 113 ms 4,018 371 To-dos (JSON API) 33 ms 16 ms 18,738 320 The pages already in the warm-up showed what to expect: the project page kept a 36 ms gap after a deploy, and the to-do page 10 ms. We’ve proposed adding these five requests, and expect them to add about five to seven seconds to the page step. The two JSON endpoints were a surprise. The warm-up’s page list had no API requests in it, so nothing on the API path had run before the first real request: not the API controllers, and not the Jbuilder templates rendering real records. Precompiling the views covers JSON templates too, but it isn’t a substitute for running the request. Results The warm-up is on for all 68 web hosts. With the first four pages it took 12 to 16 seconds per host: about 2 seconds to precompile the views, 7 to 9 for the 60 requests, and 3.5 for the schema. Deploys take that much longer per host, and we raised the deploy timeout from 30 to 60 seconds to cover it. In Amsterdam, at a peak-hour deploy: During deploy Before After Peak Puma backlog per host 250–2,238 requests 19–223 requests Peak kamal-proxy p99 about 10 s 2.4–4.8 s Peak CPU time per request 201–214 ms 88–132 ms Peak database time per request 56–69 ms 39–47 ms The same eight hosts at three deploys on 1 October, an hour apart, as the warm-up went from one host to four to all eight. The deploy in the middle, with four hosts warmed and four not, shows why every host needed the warm-up. Each warmed host recovered faster on its own: mean request duration peaked at 130 to 173 ms, against 203 to 311 ms on the hosts that weren’t warmed. But the backlog on all eight was about the same, because they were all waiting on the same database. Mean request duration on each host at the 07:21 UTC deploy. Blue hosts warmed up in the master before forking, orange hosts did not. Memory came down too. The workers now share compiled templates, YJIT code and the schema with the master instead of each building their own copy. On beta, the view precompiler alone took a busy worker’s private memory from 174–202 MB to 119–135 MB. Thirty minutes after the deploy, the web containers used about 39 GB less memory than the previous day’s containers at the same age and traffic. Amsterdam served most of our traffic at the times we tested. In Amsterdam, each new container used about 2 GB less just after traffic moved to it, which lowers the peak while the old and new containers overlap. Working with Claude Claude Code helped throughout. It combed through the per-worker backlogs and per-endpoint timings in Prometheus and Loki after each deploy, worked out the cold-versus-warm cost of each endpoint, and prepared the changes and the pull request descriptions with the benchmarks in them. We decided what to try, deployed it and read the results. If you do this Warm the master before it forks. Compile common code and templates and load their schema in the master, so workers inherit that work. With preload_app!, Rack::Builder#warmup runs before the workers start accepting traffic. Use the app’s real requests. Public pages, internal endpoints and synthetic queries warm the paths they run and nothing else. Signed-in requests to real records, frames included, run what production runs. Measure the cold penalty per endpoint. The difference between an endpoint’s cold and warm duration, times its request count after a deploy, ranks the pages worth adding. Ours weren’t the ones we’d have guessed, and two of them were JSON. Check what the warm-up leaves open. List the master’s file descriptors after a warm-up and account for every one before the fork. Two of ours needed changes. Set a time budget. A warm-up that runs long on one slow host fails the deploy on that host. Ours gives the page requests and schema loading a shared 20-second budget, checked before each page or model, puts the most valuable pages first, and logs what it skipped. Reforking from a warm worker, as Pitchfork does, solves the same problem continuously rather than once at boot, and it would warm paths no fixed list of pages covers. We may still get there: our branch brings Instacart’s mold_worker up to date with Puma’s main branch and fixes the bugs we found in it. But warming the master works with the Puma we already run, took a few days to implement, and substantially reduced the queues after deployment.

yesterday • 1 votes
📚 BoredReading

You seem to be enjoying this.

Join free to unlock everything.

Create free account

Already have an account? Sign in