Full Width [alt+shift+f] Shortcuts [alt+shift+k]
Sign Up [alt+shift+s] Log In [alt+shift+l]
1

Cooking with glasses

from Tom MacWright [alt+shift+b] in programming

I’ve been thinking the new Meta Ray-Ban augmented reality glasses. Not because they failed onstage, which they absolutely did. Or that shortly after they received rave reviews from Victoria Song at The Verge and MKBHD, two of the most influential tech reviewers. My impression is that the hardware has improved but the software is pretty bad. Mostly I keep thinking about the cooking demo. Yeah, it bombed. But what if it worked? What if Meta releases the third iteration of this hardware next year and it worked? This post is just questions. The demos were deeply weird: both Mark Zuckerberg and Jack Mancuso (the celebrity chef) had their AR glasses in a particular demo mode that broadcasted the audio they were hearing and the video they were seeing to the audience and to the live feed of the event. Of course they need to square the circle of AR glasses being ‘invisible technology,’ but showing people what it does. According to MKBHD’s review, one of the major breakthroughs of the new edition is that you can’t see when other people are looking at the glasses built-in display. I should credit Song for mentioning the privacy issues of the glasses in her review for The Verge. MKBHD briefly talks about his concern that people will use the glasses to be even more distracted during face-to-face interactions. It’s not like everyone’s totally ignoring the implications. But the implications. Like for example: I don’t know, sometimes I’m cooking for my partner. Do they hear a one-sided conversation between me and the “Live AI” about what to do first, how to combine the ingredients? Without the staged demo’s video-casting, this would have been the demo: a celebrity chef talking to himself like a lunatic, asking how to start. Or what if we’re cooking together – do we both have the glasses, and maybe the Live AI responds to… both of us? Do we hear each other’s audio, or is it a mystery what the AI is telling the sous-chef? Does it broadcast to a bluetooth speaker maybe? Maybe the...
21st Sep 2025

Stay updated

Get a weekly newsletter with the top 5 articles worth reading every week.

More from Tom MacWright

Recently

I skipped August's Recently because this summer has been relentless, mostly in the positive sense. July and August were filled with bike rides, 5K races, friends in town, life changes and logistics. For a friend's birthday, we rode the rickety Coney Island Cyclone, which was faster and more intense than I remembered it. But it's been 18 years since the last injury on the coaster, and the rider was partly at fault: good enough for me. .two-up{display:grid;grid-template-columns:1fr 1fr;gap:10px} I rode my first proper randonneuring ride: 200km (124mi) to all of the beaches around Brooklyn. It was pretty challenging, as you'd expect, but I didn't 'find my limit' for endurance, but I did find my limit in terms of how many gels & ride-nutrition packets I could consume before sugar becomes repulsive. I do still heartily endorse maple syrup packets for ride nutrition: compared to the space-age fuels that all taste like fruit or coffee, it's an identifiable substance that tastes like you'd expect. And, not pictured, but I ran a bunch of 5Ks. Haven't hit my goal this year of going sub-20, but got close in the last one - 20:15. I'm not giving up yet, because the weather is getting better and Brooklyn has a lot of 5K options. Listening Hot Creek by Duffy x Uhlmann Duffy x Uhlmann has been in my rotation for a year now and the new album has all the things I like from the combination. This and Shrunken Elvis go along well together and both have been hits at dinner parties. This album from Frances Quinlan, the frontperson of Hop Along, went under the radar, but it's so good. I've also been playing Hop Along's send album Get Disowned. I really like everything that this band has put out. It's a product of Philadelphia but I hear some of the elements that drew me to DC post-punk in the turn-on-a-dime songwriting. Get Disowned by Hop Along Watching This video, 'The surprising truth about what motivates us', is kind of an artifact of the 'Obamaverse', as one of my friends would put it. It's by Daniel Pink, a writer-speaker-self-help kind of guy, and it's accompanied by one of those realtime hand-illustrations. There are lots of reasons to tune it out as YouTube filler. Still, I thought it was pretty cool for two reasons: first, it connected directly with Andrew Kelley's video, which is the next one. And it connects with my feelings about incentives and how most simple incentives are counterproductive. And second, the drawing is real. I'm so used to the artificial version of this visual style that it took a minutes to realize that the it wasn't AI or some 'drawing hand' style of video essay, it was a real person making quite nice illustrations: specifically Andrew Park. Found myself nodding along to every bit of this interview with Andrew Kelley. His values and motivation are so nice to encounter in this current phase of tech. He's kind of a role model, which is something that I used to find trite in my early career but I'm coming around to. As Adam Neely pointed out in his video, a lot of the people heavily using Suno (an AI music tool) could not cite any role models, and defaulted to a sort of enclosed, self-centered stance. I see this with engineering too: it's hard to discern what values or talents the people at 'the forefront of the industry' have that I want. Finding personal qualities as something to imitate instead of desires and possessions (cue the segue into Girard's Memetic theory) seems like a better way to live. Reading What strengthens a relationship? Almost always, it is personal investment. Perhaps an AI agent might be better at predicting what my father will want for his birthday than I am, but it definitionally cannot give him my time and consideration. Love is not just a feeling; it is a way of paying attention. Elizabeth Lopatto in The Verge picks apart Mark Zuckerberg's weird, sad view of humanity. It's worth a read even just for the incredible opening anecdote. Home of the titular 100 foot waves. Which existed when I was a kid of course, but the larger world didn’t know about them yet. They were our secret. My grandfather and I would walk to the lighthouse and watch the waves hitting against the cliffs and he’d tell me that these were the biggest waves in the world. I didn’t believe him, of course. I thought it was standard grandpa hyperbole. But I also didn’t want to believe that I was seeing the biggest of anything. Not yet. I wanted the really great things to be in the future. Something to aim for. Which means I missed out on some great things that were right in front of me. Mike Monteiro's newsletter is always beautiful and melancholy, and this is a really nice edition of it. I read and highlighted a few other articles but they were all saying things about AI, mostly about how it makes the world worse and makes people feel bad. I won't add those to the pile because I don't want that kind of content to dominate what I mention, even if it does dominate what I read. Elsewhere For the Val Town blog, I wrote about our experience with the bug bounty program. It's been a very interesting experience. Val Town has a pretty difficult security surface: in contrast to Mapbox, which was mostly a read-only API, it's a very read-write system with many access controls. At the center of the product is a sandbox that we want to keep secure. And because we're living in the AI age, a single bit of functionality can be exposed via REST API, React Router loaders, tRPC, and MCP. Building a model layer and a shared authorization layer - modeled after Oso - has helped a bit. But the reports still come in. It's pretty wild how AI factors into it: I'd say that 100% of the vulnerability reports used AI to write the report itself. From the best reporters, the vulnerabilities come with screen recordings and evidence, and it's clearly not all-robot. But - I suspect that other people manning help desks & bug bounty programs know this experience - it's jarring to go three replies into an email conversation with someone and they suddenly stop using an LLM to write their emails and the voice completely changes, from hyper-literate to abbreviated and misspelled. I also wrote about our new authentication scheme that makes vals require login and makes it possible to manage permissions to val applications, not just their code. It was pretty fun to implement, and the sixth or seventh time working with HMAC schemes I'm starting to get a natural grasp on them.

1st Sep 2026
Technology optimism hour

Back in 2022, I wrote Web technology optimism hour, highlighting the things I thought were cool about the state of technology. Time flies. Let's do it again. It's still easy to be pessimistic, but it's nice to mine your feelings for the positive ones. The elephant in the room is AI, but I don't feel like throwing another post into the gravity field of that topic. Without further adieu: TypeScript 7 It's here, finally, and it's faster. In the last edition, one of the headings was "Rust and Go based tools are making TypeScript development faster." At that point we had a few upstart projects for minifying, transpiling, and bundling TypeScript. Those tools have become standard-issue. But TypeScript, the actual type-checking of the language and the language tooling in editors, was still a bottleneck. There was exactly one implementation of the TypeScript type checker, and it was implemented in TypeScript, so it ran at JavaScript speeds and had limited ability to become multi-threaded. Unfortunately there's still only one game in town for TypeScript type checkers: it's still just Microsoft. But they rewrote TypeScript in Go for TypeScript 7 and it's much, much faster. This is one of the most noticeable changes to my day-to-day coding in a while. Plus, even better: TypeScript 7 supports the Language Server Protocol! As I wrote about in 2024, Microsoft invented TypeScript and also the Language Server Protocol, but TypeScript never implemented the Language Server Protocol. So there was no standard way for other editors, other than Microsoft's VS Code, to provide language tooling for TypeScript. This was suspicious behavior because of the company's history. But now: it just works. Neovim can support TypeScript 7 with its built-in Language Server support instead of someone wrapping the VSCode extension in an LS-protocol shell. So: a big win. Hooray for the new TypeScript: it's faster, it supports open standards, and now there are two TypeScript implementations. Now all we need is one that's made by someone who doesn't work at Microsoft. The Effect ecosystem is exciting I've been writing little development logs about Effect over the last year or two. I'm excited about Effect. It's a big, all-encompassing system, but I think it has a lot of potential and is arriving at exactly the right time. This is an era of supply-chain risk and I think there is, and should be, a swing away from the tiny module ecosystems in NPM. Software that's composed of hundreds of little modules from hundreds of authors is a lovely idea that doesn't work with malicious actors. The swing is toward larger, buttoned-up libraries, and Effect is really well-positioned for that. It has type validation, observability, state management, fancy data types, and lots of functional programming helpers all under one roof, in one package that's pretty well-maintained and heavily tested. It's also very, very typed, which makes it safer to use at scale. And it makes limits a lot easier: you can just stick an Effect.timeout on anything and prevent network requests or database queries or anything else from taking too long. The same with concurrency, it's very easy to limit concurrency. It's ready for a very uncertain world where you have to validate and be careful about everything. But at the same time, it's fun to use, if you're that kind of person, which I am. Bluesky and Mastodon are successful, meaningfully decentralized systems I use Bluesky and Mastodon all the time and both just work. They don't feel like tech demos, it's not all nerds using them. The uptime is pretty good, the featureset doesn't seem limited by the technology choices. And yet both are legitimately decentralized (in different ways) networks, with really interesting ideas about security, privacy, and community systems. Bluesky is aggressively building out its protocol as an independent technology from the social network, and as I learned at their conference, there are lots of creative people building on the technology. There's a big gap between a cool tech demo and a user-friendly popular service, and both of these have bridged it. It's exciting, and shows that you don't have to accept a bad set of tradeoffs to make tech that's both principled and popular. Signal, WhatsApp, iMessage, and RCS are meaningfully secure messaging platforms I want to just celebrate the end of a technology: SMS. It's so good that people don't use SMS anymore: it was needlessly expensive and extremely insecure. My texts with friends are almost exclusively on Signal & iMessage, and occasionally WhatsApp. Now, WhatsApp is a mixed bag and nothing's perfect, but it's cool that the secure options recommended by the EFF and other sources are also popular options that people use every day. That's also pretty important from a privacy standard: having Signal installed on your phone is normal because it has a wide userbase, but the same would not be true if it was only used by people attending protests and being wary of government surveillance. Hardware update cycles are slowing People are hanging on to their phones and computers for longer. I'm part of this trend - my iPhone hit its fourth birthday, and I have no reason to even think about replacing this computer. With the exception of their batteries, modern laptops and phones are robust: they're water-resistant, made of durable materials, and basically just work. You could read the same data differently: maybe people are upgrading slower because they have less disposable income. I'm not sure that's true. But either way, it's good that upgrade cycles are longer: it means less e-waste, less pollution, less churn overall.

14th Aug 2026 • 1 votes
Recently

June was a big month: I went to Porto & Lisbon, and had a lot of life stuff happen. I'll get into the trip once I get my rolls of film developed. Three rolls at a new photo developing place: fingers crossed! Reading I finished reading Intermezzo (of the bag) and it was fantastic. I've always liked Sally Rooney's books but this was the one where the writing style really clicked. Also, The Vegan. Meh. I read Patricia Lockwood's 'A Tradcath Wedding' via Perfect Sentences but found an additional sentence to be perfect: Whenever they rang the chimes, which seemed to be every four seconds or so, a toddler screamed ‘WOW A BELL!’ to the visible displeasure of the celebrants – though isn’t the entire point of the ritual that you’re supposed to be that awestruck every time? Patricia Lockwood is the funniest writer I've read. You cannot grow a pumpkin, but you can improve the odds. Taylor's 'You Cannot Grow a Pumpkin' is a fantastic little prose poem of sorts. Watching I'm always trying to find a 'romp' when it comes to movies. Something lighthearted, pretty easy to watch, so on. We watched The Pink Panther this month and it is a perfect example of the genre. The inspiration for the watch came from the hamburger scene: But there's so much more of this kind of thing in the movie, little bits and physical comedy. Oh, and I also watched The Departed, which everyone says is good and is good. Elsewhere I wrote Accidental Anonymity on the micro blog, and it stirred up some discussion on Bluesky as well as at least one blog response. It was kind of an angry piece, as I said at the start. I will keep trying to stay out of the trap of writing about that topic all the time. Listening The only album I bought this month was Songs of Her's by Her's, which is fine. I wish I had something more profound to say about it given how the band met an unbelievably tragic end. Maybe more influential than that was Know Your Enemy's recent podcasts, especially this one about the pope's encyclical. I've really grown to love that podcast, and it has been part of me intellectually reconnecting with, but not readopting, Catholicism. Art Here's some art I really liked this month: This Hockney piece called 'Picture Emphasizing Stillness' from 1962 was at the MAC/CCB museum in Lisbon. Via Tim Babb, I enjoyed finding Tomás Sánchez's work.

1st Jul 2026 • 1 votes
Recently

A section of trail in Switzerland May was a big month! The highlight was a cycling trip around Lake Konstanz, passing through Konstanz, Bregenz Austria, Stein am Rhein in Switzerland, and Meersburg. A farm in Bodenseeufer, Germany We passed a lot of operating agriculture, growing apples, strawberries, and other fruits. The route is very continuous, mostly flat, and popular with retirees. It's very idyllic riding: lots of protected and separated lanes. There are some interesting regional differences in the riding. Swiss drivers were noticeably more aggressive and gave a lot less space to bikes, but the Swiss infrastructure was really nice when it was off-road. Also it was interesting to see that the vast majority of other cyclists on most of the route were on ebikes. This only changed when we were close to hip cities and we'd notice more young people on high-end road bikes. Münster St. Maria und Markus Taking the ride in 40-60 mile days left a lot of room for checking out the towns, food, and views. It was very easy being vegan in Germany, and significantly harder in Austria and Switzerland. Some highlights: KERVAN Imbiss in Konstanz: cheap, delicious vegan kebabs Insel Mainau: beautiful botanical garden on an island, recommended to follow up with Biergarten St. Katharina, a beer garden in the woods The Pile Dwelling Museum in Uhldingen-Mühlhofen-Unteruhldingen: very aesthetic museum plus recreations of pile dwellings In Zurich: the Swiss National Museum was gigantic and featured the best-integrated high-tech exhibits I've seen On the final evening we took a cable car up to Panorama Restaurant Falsenegg and it delivered on the name Reading I read The Technological Republic the book by Alex Karp and one of his employees. It was terrible as expected: part sales-pitch, part standard-issue MAGA cultural critique, part implied defense of just war. The last part was the most interesting to me, because Karp spends so much time grandstanding about his intellectual background and telling protestors to quiet down and have a real discussion, and this book confirms that he can't actually have that conversation. He has nothing to say about the moral complexity or justification of war. It's surprising that the only critical reviews I could find of the book are from other right-leaning sources. Providence, an American Christian Realist magazine, found it too weak. The Independent Institute didn't like it based on their Libertarian principles. I guess it's just so far from any left-wing thought that nobody bothers to read it. Democratic governance rests on a bargain so old we’ve forgotten it’s a bargain at all. The governed have something the governors need: labor, tax revenue, military service, consumer spending. This dependency is the source of democratic leverage. The whole system functions because power is distributed, and it’s distributed because the people at the top need something from the people at the bottom. I'm still trying to avoid writing about AI, but this piece by Owen McGrann was worth the time. It takes a lot of points that I've long agreed with and stitches them together into a coherent but miserable whole. On the bright side, I'm just finishing reading Intermezzo and it's wonderful. A totally rejuvenating and inspiring novel. Listening Cheekface is a ~9 year old band in the tradition of Cake, They Might Be Giants, etc. They've got some great hits - part sardonic and ironic, part euphoric and ultra-light pop choruses. It's Sorted by Cheekface Watching Adam Neely's talking about AI and music again, and like always, it's very worth watching. The generational element is especially interesting here: just like the commencement speakers getting booed, there's a consistent theme where Gen Z (and millennials, to some extent) are rejecting the AI hype while older generations are optimistic and coincidentally in a position to benefit from it.

4th Jun 2026 • 1 votes
Recently

I spent a lot of time outside in April. The first two weekends I did long bike rides: from Brooklyn to Tarrytown along the Old Croton Aqueduct trail, and out to Rockaway Beach via the Marine Parkway Bridge & Cross Bay Memorial Bridge. Then, I ran the Brooklyn Experience Half Marathon. I wrote some notes about the experience over in /micro. I properly trained for it this time, and it felt good and was pretty fast: shaved five minutes off my previous record. And then, the Great Saunter… The Great Saunter is a full loop around Manhattan. The route is 32 miles but it took us 33 with GPS jitter and detours. Ten hours and forty-three minutes. It was very difficult, in a lot of ways tougher than the half-marathon in terms of the stress it puts on the body. I'm getting more comfortable with longer endurance events, but still have no interest in running a marathon, and definitely no ultramarathons. A 200-300km randonneuring ride though could be in the cards. Reading I read The Origins of Efficiency by Brian Potter, the author of Construction Physics. I learned that I like his blogging better than his book-length writing: for example, this month's article Helium is Hard to Replace is really great, with understandable and fascinating charts and examples. there’s a final layer to this argument that nobody’s quite articulated yet. product quality improvements, at the frontier, are not bounded by how fast you can write code. they’re bounded by how fast you can come up with ideas good enough to push the frontier. I liked claude code is not making your product better, and think it perfectly rhymes with John Cutler's post about maximizers vs. focusers which isn't directly about AI. Maximizers are overrepresented in the top rungs on tech companies, people who are marked by opportunistic, experimental thinking, and for them the ability to implement a ton of features really quickly even at low quality is a gift. Their ability to enact their will was previously gated by the people doing the coding and designing, who mostly dislike pushing low-quality work. Now it isn't, as much. I took the time to thoroughly read through JavaScript has a Unicode Problem and JavaScript’s internal character encoding: UCS-2 or UTF-16? They're from 2013 but still relevant, and really engaging from a technical perspective if you are very interested in text encoding and also JavaScript, which I am. Listening Extra Stars by Gregory Uhlmann An excellent month for music. The new Gregory Uhlman (guitarist for SML) album. Circadia by Mammal Hands New Mammal Hands, too. Dosh, Ismaily, Young by Dosh, Ismaily, Young New album from Dosh, Ismaily, Young. Dosh who you might know from being Andrew Bird's drummer at one point (after Kevin O'Donnell and before Griffin Goldsmith). Complex Emotions by The Bad Plus The Bad Plus's final album from back in 2024! They're on a farewell tour right now, all things must come to an end. Watching This was a great long watch on some math concepts that I've had to re-learn a bunch of times, and it was the first time that the idea of a quaternion really clicked for me. Freya's channel has a lot more great videos. Jamelle Bouie's channel is consistently the best video commentary on the details of US politics I can find.

4th May 2026 • 1 votes

More in programming

Clip of me singing Despard in Ruddigore in 2013

A clip of me singing a funny song from Gilbert and Sullivan’s Ruddigore back in 2013

6 hours ago • 1 votes
How and Why fork() Uses Copy-on-Write

In this video, we look at why fork() needs copy-on-write, how it works inside the kernel, and a memory usage problem that Instagram encountered with Python.

13 hours ago • 1 votes
What we lost when we lost comments

Comments require commitment, but they’re worth it.

19 hours ago • 1 votes
Lighthouse map

Lovely global map with animated lights sweeping the waters

21 hours ago • 1 votes
Warming up the Puma master before it forks

Basecamp 5 runs on Puma in cluster mode: one master process with preload_app! and 63 single-threaded workers per host, deployed as a Docker container with Kamal. We serve Basecamp from several sites. Each site has its own web hosts and a read replica of the database, and writes go to a single primary database in one of them. On our busiest hosts, each deploy left up to 2,000 requests waiting while the new workers warmed up. We reduced those queues by running signed-in requests through the app in the Puma master, before it forked the workers. Why 63 single-threaded workers? Basecamp has always served web requests from processes rather than threads. It ran on Unicorn, which only does processes, until we moved to Puma in January 2025, and we kept the same setup: workers (Concurrent.physical_processor_count * 1.3).ceil threads 1, 1 preload_app! On a 48-core host that’s 63 workers, each handling one request at a time. We chose 1.3 after benchmarking HEY in 2023, when we moved our apps out of the cloud and onto our own hardware. We tested several combinations of workers and threads with a mix of GET and POST requests on a 32-vCPU VM. Every multithreaded configuration we tested was slower and handled fewer requests than single-threaded workers. Adding workers beyond about 1.2 to 1.3 per vCPU brought little benefit. The threaded workers spent a lot of their time waiting for Ruby’s global VM lock. That made single-threaded workers a good fit for this workload, and we use the same setup for Basecamp. An app that spends more time waiting on its database or other services may benefit from more threads, so benchmark your own app. The other reason is the app itself. Basecamp has class-level state in places and has never needed to be thread-safe. With one request per process, it still doesn’t. Processes do use more memory than threads, and preload_app! reduces the difference. The master loads the app once and the workers share its memory through copy-on-write until they write to it. Shopify’s comparison of Ruby execution models explains the trade-off well. In the HEY benchmark the best setup came to about 260 MB of PSS per core, where PSS counts each shared page once, split between the processes using it, and the gap to a threaded setup was smaller than we’d expected. What Puma does on each host when a container starts: one master, then 63 forked workers that share its memory until they write to it. Two things about this setup matter for the rest of the post. A worker that’s compiling or loading something is fully blocked — there’s no other thread to pick up the next request. And whatever the master has in memory before it forks, all 63 workers share. Whatever they build after the fork, they build 63 times. What happens when we deploy Kamal starts the new container alongside the old one, and kamal-proxy moves the host’s traffic across as soon as the health check passes. At that moment, the new workers have handled health checks but no customer requests. preload_app! means the master loads the app once and the workers inherit it through fork. That covers the code. It doesn’t cover anything Ruby and Rails set up on first use: YJIT compiled code. YJIT compiles a method once it’s been called a certain number of times. The master calls very little during boot, so every worker compiles the same methods again on its own first requests. Compiled templates. Action View turns each ERB template into a Ruby method the first time it’s rendered. The schema cache. Active Record reads each model’s columns from the database the first time that model is used. Inline caches and memoized values throughout Ruby, Rails and the app. All 63 workers did all of this at once, while serving the traffic the old container had been handling a second earlier. In the test environment with YJIT on, the first request to a project page on a cold process took 652 ms, 151 ms of it YJIT compiling. The same request to a warm process took 28 ms. In production, CPU time per request peaked at around 200 ms while kamal-proxy moved traffic to the new container, against about 30 ms once the workers had warmed up. A host with spare CPU absorbs this. Every one of our web hosts has 48 cores and 63 workers, but each Amsterdam host serves around 250 requests per second, against 25 to 60 at our other sites. In Amsterdam the slow first requests turned into a queue. At a peak-hour deploy, the Puma backlog on an Amsterdam host reached anywhere from 250 to 2,238 requests, and kamal-proxy’s p99 response time hit about 10 seconds. Eron, our Director of Operations, had been tracking this since June. Another server in Amsterdam would help, but it would take weeks to arrive, so we also wanted to make deploys cheaper on the hardware we already had. What didn’t work We tried a few things first. In June, Donal tested the first two on a single Amsterdam host, comparing it with its neighbors, and they ruled out two likely causes. Warming each worker’s database connections. Puma’s before_fork hook clears the master’s connections, and each worker opened its own on its first request. Opening them in before_worker_boot instead made no difference. Queries on a freshly booted production host were already under a millisecond, so connections weren’t the problem. A synthetic request in each worker. Next, each worker made a few requests in before_worker_boot to an internal controller that touched every model. That ran the middleware, routing and Active Record paths, but it ran them in 63 workers at once — exactly the CPU spike we were trying to avoid. And a request with no real data renders no real views, so most of the app stayed cold. Spreading YJIT compilation out. Delaying YJIT in each worker by a random interval spread the compiling out over a few minutes, but every worker still ran interpreted until its delay ended. The queue didn’t change. Reforking from a warm worker. This is what Shopify’s Pitchfork does: let one worker serve traffic until it’s warm, then fork the others from it. Puma has an experimental version called fork_worker, and on beta it worked — the reforked workers were warm after three to five requests, where fresh ones took up to 30 seconds. But with fork_worker the template is worker 0, and it keeps serving requests. If it exits, the workers waiting to be forked never start (puma/puma#3596). If it gets no traffic, the refork never happens, which is what we saw on beta. Instacart have a mold_worker patch that promotes a warm worker to a template that stops serving, but it isn’t in a Puma release. We have a branch of it, and we may come back to it. That last experiment did show us where the fix was, though. Everything a warm worker has that a cold one lacks is in its memory, and fork copies memory. The master already has the app loaded. It just never runs it. Run the requests in the master So now, before the master binds its socket and forks, it makes the app’s own requests, in-process, the way a signed-in user would. Rack has a hook for exactly this. Rack::Builder#warmup takes a block that’s called once with the built app, before the server starts. rails server builds the app from config.ru, so the change to boot is one line: require_relative "config/environment" warmup { WarmUp.configured.run } if ENV["WARM_UP"] run Rails.application With preload_app! this runs in the master, and the workers inherit whatever it did. Puma binds its socket after the app is built, so until the warm-up finishes the health check’s connection is refused and kamal-proxy keeps retrying. No request reaches a worker that hasn’t been warmed. The warm-up has three steps. After precompiling the views, it gives the page requests and schema loading a shared 20-second budget, checked before each page or model. 1. Precompile the views actionview_precompiler reads every template for its render calls and compiles each one with the locals it’s passed. For us that’s 1,394 templates in about two seconds. A first request to a project page then compiles 2 templates instead of 44. 2. Request the pages, signed in A small browser class makes the requests through Rack::MockRequest, with the two cookies a real sign-in sets, then goes back for each page’s lazy Turbo frames: class WarmUp::Browser def initialize(signed_in_as:) @client = Rack::MockRequest.new(Rails.application) @headers = { "HTTP_USER_AGENT" => "Basecamp warm-up", "HTTP_COOKIE" => cookie_for(signed_in_as), "bc3.warm_up" => true } end def visit(path) page = get(path) frames_in(page).each { |id, src| get(src, "HTTP_TURBO_FRAME" => id) } end private def get(path, headers = {}) @client.get("https://#{host}#{path}", @headers.merge(headers)) end def frames_in(page) Nokogiri::HTML5(page.body).css("turbo-frame[src]").map { |frame| [ frame["id"], frame["src"] ] } end end The requests are signed in. The user is a monitoring account we already use for automated checks, and the pages are its own project, Campfire, to-dos, documents and messages. Public pages weren’t enough: after warming up with signed-out pages only, the first signed-in request to the projects page still took 131 ms, because authentication, the signed-in controllers and their views had never run. With signed-in pages it took 40 ms. cookie_for writes the same signed cookie the sign-in controller does, using the app’s own cookie jar, so there’s no API token and no secret to store. The frames are followed. The busiest HTML requests in production aren’t pages at all but Turbo frames — the sidebar badge, the inbox, the navigation menus. The browser parses each page and requests its <turbo-frame src> URLs with the Turbo-Frame header, so those controllers and views get warmed too. Our first four pages turned into 60 requests. The requests are excluded from rate limiting. They are internal, so they do not count against the rate limits that apply to real visitors. 3. Load the rest of the schema The page requests load the schema for the models they touch. The last step loads the rest, from the read replica: ApplicationRecord.reading do models.lazy.take_while { time_left? }.each { |model| model.load_schema if model.table_exists? } end The step checks 261 models and loads any schema information still missing. Those database round trips add up when the primary is far away: outside a request, Active Record uses the writing role, and from a host a long way from the primary each round trip is tens of milliseconds. Reading from the local replica brings the step down from about 20 seconds to 3.5. The pages go first because they load most of the schema anyway. If the time budget runs out, the step stops, logs how many models it got through, and the workers load the rest on first use like they always did. Rails can also load the schema from a dumped cache file at boot (bin/rails db:schema:cache:dump), which would make this step unnecessary. We don’t ship one in our image yet, because the dump needs a database to read from at build time, and we have several databases to cover. It’s on the list. What to close before the fork Running requests in the master opens things the master never opened before, and every worker inherits them. Two processes writing to the same socket will corrupt each other’s traffic, so you need to know what’s open before you fork. The way to find out is to list the master’s open file descriptors — ls -l /proc/<pid>/fd — before and after a warm-up, in an environment set up like production. Development wasn’t enough for us: it stores files on disk, so our S3 connections only showed up in production. Then, for each thing that’s open, check how its library handles a fork. We found three kinds: Already handled. Plenty of libraries detect a fork on their own, either by recording the PID they connected from and reconnecting in the child, by opening per-process files, or by resetting their thread pools. Redis clients, metrics libraries and concurrency libraries tend to be in this group. Check, but you probably don’t need to do anything. Already closed. Database connections are the classic one, and most Puma configs already clear them in before_fork. Anything else that’s opened per process — we have a SQLite cache the workers open on boot — needs closing when the warm-up finishes. Needs a new step. HTTP clients with keep-alive connections are the ones to look for: cloud SDKs with connection pools, tracing exporters, error reporters. They usually have no fork handling at all. We empty the aws-sdk connection pools in before_fork, and we run the warm-up untraced so the OpenTelemetry exporter never opens its connection to Tempo in the first place. Once that’s done, before_fork finishes with Process.warmup, which Ruby 3.3 added for this purpose: a major GC, a heap compaction, and every surviving object promoted to the old generation, so the memory pages the workers share change as little as possible afterwards. Choosing the pages The first list was the four pages that ran the busiest requests on beta. Once the warm-up was live, production showed us which endpoints were still cold. For one deploy, we compared each endpoint’s mean duration in the six minutes after kamal-proxy moved traffic to the new container with the same endpoint an hour later, then multiplied the difference by the number of requests in those six minutes. That gives the extra time each endpoint cost us because it was cold: Endpoint Cold Warm Requests in 6 min Extra seconds Campfire 246 ms 70 ms 6,490 1,140 Projects (JSON API) 84 ms 50 ms 22,077 771 Docs & Files 262 ms 177 ms 4,996 421 To-dos tool 205 ms 113 ms 4,018 371 To-dos (JSON API) 33 ms 16 ms 18,738 320 The pages already in the warm-up showed what to expect: the project page kept a 36 ms gap after a deploy, and the to-do page 10 ms. We’ve proposed adding these five requests, and expect them to add about five to seven seconds to the page step. The two JSON endpoints were a surprise. The warm-up’s page list had no API requests in it, so nothing on the API path had run before the first real request: not the API controllers, and not the Jbuilder templates rendering real records. Precompiling the views covers JSON templates too, but it isn’t a substitute for running the request. Results The warm-up is on for all 68 web hosts. With the first four pages it took 12 to 16 seconds per host: about 2 seconds to precompile the views, 7 to 9 for the 60 requests, and 3.5 for the schema. Deploys take that much longer per host, and we raised the deploy timeout from 30 to 60 seconds to cover it. In Amsterdam, at a peak-hour deploy: During deploy Before After Peak Puma backlog per host 250–2,238 requests 19–223 requests Peak kamal-proxy p99 about 10 s 2.4–4.8 s Peak CPU time per request 201–214 ms 88–132 ms Peak database time per request 56–69 ms 39–47 ms The same eight hosts at three deploys on 1 October, an hour apart, as the warm-up went from one host to four to all eight. The deploy in the middle, with four hosts warmed and four not, shows why every host needed the warm-up. Each warmed host recovered faster on its own: mean request duration peaked at 130 to 173 ms, against 203 to 311 ms on the hosts that weren’t warmed. But the backlog on all eight was about the same, because they were all waiting on the same database. Mean request duration on each host at the 07:21 UTC deploy. Blue hosts warmed up in the master before forking, orange hosts did not. Memory came down too. The workers now share compiled templates, YJIT code and the schema with the master instead of each building their own copy. On beta, the view precompiler alone took a busy worker’s private memory from 174–202 MB to 119–135 MB. Thirty minutes after the deploy, the web containers used about 39 GB less memory than the previous day’s containers at the same age and traffic. Amsterdam served most of our traffic at the times we tested. In Amsterdam, each new container used about 2 GB less just after traffic moved to it, which lowers the peak while the old and new containers overlap. Working with Claude Claude Code helped throughout. It combed through the per-worker backlogs and per-endpoint timings in Prometheus and Loki after each deploy, worked out the cold-versus-warm cost of each endpoint, and prepared the changes and the pull request descriptions with the benchmarks in them. We decided what to try, deployed it and read the results. If you do this Warm the master before it forks. Compile common code and templates and load their schema in the master, so workers inherit that work. With preload_app!, Rack::Builder#warmup runs before the workers start accepting traffic. Use the app’s real requests. Public pages, internal endpoints and synthetic queries warm the paths they run and nothing else. Signed-in requests to real records, frames included, run what production runs. Measure the cold penalty per endpoint. The difference between an endpoint’s cold and warm duration, times its request count after a deploy, ranks the pages worth adding. Ours weren’t the ones we’d have guessed, and two of them were JSON. Check what the warm-up leaves open. List the master’s file descriptors after a warm-up and account for every one before the fork. Two of ours needed changes. Set a time budget. A warm-up that runs long on one slow host fails the deploy on that host. Ours gives the page requests and schema loading a shared 20-second budget, checked before each page or model, puts the most valuable pages first, and logs what it skipped. Reforking from a warm worker, as Pitchfork does, solves the same problem continuously rather than once at boot, and it would warm paths no fixed list of pages covers. We may still get there: our branch brings Instacart’s mold_worker up to date with Puma’s main branch and fixes the bugs we found in it. But warming the master works with the Puma we already run, took a few days to implement, and substantially reduced the queues after deployment.

yesterday • 1 votes
📚 BoredReading

You seem to be enjoying this.

Join free to unlock everything.

Create free account

Already have an account? Sign in