Full Width [alt+shift+f] Shortcuts [alt+shift+k]
Sign Up [alt+shift+s] Log In [alt+shift+l]
55

Proving that every program halts

from ntietz.com blog - technically a blog [alt+shift+b] in programming

One of the best known hard problems in computer science is the halting problem. In fact, it's widely thought[1] that you cannot write a program that will, for any arbitrary program as input, tell you correctly whether or not it will terminate. This is written from the framing of computers, though: can we do better with a human in the loop? It turns out, we can. And we can use a method that's generalizable, which many people can follow for many problems. Not everyone can use the method, which you'll see why in a bit. But lots of people can apply this proof technique. Let's get started. * * * We'll start by formalizing what we're talking about, just a little bit. I'm not going to give the full formal proof—that will be reserved for when this is submitted to a prestigious conference next year. We will call the set of all programs P. We want to answer, for any p in P, whether or not p will eventually halt. We will call this h(p) and h(p) = true if p eventually finished and false otherwise. Actually, scratch that. Let's simplify it and just say that yes, every program does halt eventually, so h(p) = true for all p. That makes our lives easier. Now we need to get from our starting assumptions, the world of logic we live in, to the truth of our statement. We'll call our goal, that h(p) = true for all p, the statement H. Now let's start with some facts. Fact one: I think it's always an appropriate time to play the saxophone. *honk*! Fact two: My wife thinks that it's sometimes inappropriate to play the saxophone, such as when it's "time for bed" or "I was in the middle of a sentence![2] We'll give the statement "It's always an appropriate time to play the saxophone" the name A. We know that I believe A is true. And my wife believes that A is false. So now we run into the snag: Fact three: The wife is always right. This is a truism in American culture, useful for settling debates. It's also useful here for solving major problems in computer science because, babe, we're both...
23rd Jun 2025

Stay updated

Get a weekly newsletter with the top 5 articles worth reading every week.

More from ntietz.com blog - technically a blog

Where I've been, and the future of this blog

Recently, I got an email from a reader who noted that I hadn't published anything in a little while, and he said he hopes I'm doing well. The truth is, some days I am and some days I'm not. I want to share a little bit about what's been going on, and what the future looks like. This is a pretty vulnerable post, so here's the deal. I'm going to tell you this, but it's just between you and me, okay? And there's one condition: you don't reach out to me to say "I'm so sorry!" When people share stories like this, it's often met with pitying responses. These often serve to make the sender feel better, not the recipient, and it feels bad to receive. I much prefer genuine connection from people who've had or seen similar experiences, or understand what I went through, or are facing something similar. I usually write my posts to educate, entertain, or a mixture of the two. This one is different. I'm not writing it to ask for pity or sympathy, but it's also not educational, and it probably won't be entertaining in a "this is fun to read" sense. The main thing I'm looking for here is connection, and maybe giving someone else on a similar journey something to hold onto. So here I am. I'm alive, and I'm writing this post. It's the first post I've written in over four months. That's the longest break I've taken from writing blog posts since I started writing regularly in 2022. What happened? what happened? enter: pain I've had nerve pain and tingling on-and-off since 2022. It was in my hands and forearms, and it was really sharp and really strong. If I type on a regular keyboard, this comes back with a vengeance. Instead, I started using Talon Voice for all my writing and coding, and eventually got a Keyboardio Model 100 which helps me mitigate the pain. The pain would only come on when it was triggered by certain conditions. A physical therapist eventually figured out some shoulder positions were contributing. I knew my keyboard was contributing. If I did something the wrong way, I'd be in pain for days afterwards, but if I did everything just right, I'd go months without any pain. The pain was with me when I started blogging regularly. The first posts in my weekly cadence were all written using my voice, not my hands. I committed code at work, pair programmed, wrote design docs, and was a very effective software engineer and leader. The pain was a little distracting, but it was not debilitating as long as I took care of my body. Accessibility tech saved my career, truly, and reduced my pain dramatically. I could not have gotten through the last four years of gainful employment without Talon, my keyboard, and the portable ergonomic setup I designed. I settled into a comfortable existence in my life: I had this oddball setup, it made me quirky but in a fun way. But then things got much worse. neurology, part 1 A year ago, I started to notice some other symptoms. They were tolerable and not very distressing. But since they were a change in something that was stable, and pointed to different potential causes than the presumed diagnosis, I wanted to get checked out to be cautious. I got a referral to neurology, and went to one of the best neurology departments in my area. We ran a lot of tests, and didn't find anything wrong. This was supposed to be relief, because it meant nothing was progressive: I shouldn't expect it to get worse, certainly not quickly. (Is this foreshadowing? It sure is!) It was also a little frustrating, because it meant we still didn't know why I had these weird sensations and pain. Three years in, no diagnosis. My doctor told me to follow up with her in a few months, and we'll keep monitoring this. I didn't follow up, because I was recovering from a planned surgery, and I was feeling pretty good. That didn't last. business ramps up A few months after my surgery, I left my job of 8 years and started my business! This was a milestone I was working toward for a while, and I was excited to dive in. I loaded up my calendar with chats. Work came in from one client, developing some software in an industry that is very privacy and security conscious. Other friends and acquaintances were happy to chat with me, and some promising leads came up. Some clients came in asking for coaching, too. Things looked so good and promising for me and my business. more tests At the same time my business was starting to ramp up, so were my nerve issues. The pain started to ramp up, and so did some unsettling sensations in my peripheral nerves. This time was different. Instead of something that I could manage and mitigate on my own, it would take over my life. But I didn't know that just yet. At this point in the story, it just felt annoying. I got an appointment with my neurologist, and went back in. She ordered more tests, and we discussed medication to manage the symptoms. Yet again, the tests said everything is normal, when my doctor and I could both clearly see that it was very much not normal. Meanwhile, the pain got worse. A month after that appointment, I was constantly at 6 or 7 on the pain scale[1]. I couldn't feel temperature in my hands, and my fingers were a little numb. My feet were burning most of the time. And I had so much brain fog that I wasn't sure if I could code a hello world example. I wasn't capable of working even one hour in a week. And we didn't know why, or what was happening to me. The only thing we could do was monitor symptoms more, and cover them up. And so I started medication. It helped, a little bit, maybe. I started using a TENS unit. It helped, for the time it was on. I started physical therapy. It helped, being forced to exercise. I worked with a dietitian nutritionist. It helped resolve other symptoms, improved quality of life. I landed a bigger client, and ramped up my work hours. I'd go through times when I worked too many hours, landing myself back on bed rest for a couple of days after. Six hours was too many. Three was maybe doable. My life felt like it was over, in a lot of ways. But we kept working on it, and adjusted my medication. Eventually, I felt mostly okay again. stable With the right medication, dietary changes, physical therapy, and (attempts at) relaxation, I got to what felt like a stable point. (Again with that foreshadowing!) I was able to take on more clients, and work more hours. That's where I am today, mostly. I say mostly because, well, today I'm actually in pain again. I've had the weird nerve symptom precursors on and off a lot recently. My fatigue is at the gate, held off for now but waiting to storm the keep. I'm writing this post with pain in my arms, but I'm typing, and that doesn't seem to make it worse. Neurology and I continue to be besties, working on our mutual hobby of debugging the weird things my body decides to do. I finally had some blood tests which were out of band one way, then the other way. I am having new symptoms, which gives us more leads to pull on. I'm still in pain, sometimes, but it's bearable and not constant. I am able to do more work every day than I have since the middle of last year. But it feels a little bit precarious. That's the life of having a chronic illness and chronic pain, especially when you don't have a diagnosis and you don't know the triggers. I have a really good life. Business is going well enough that I'm able to comfortably support my family. (That said: please do give me more business!) My personal life is also going well. And I'm making some beautiful poetry and music. I cherish this life. I'm drinking in every minute I can. I don't know if I have one more year of playing my saxophone, or sixty more. So I'm going to enjoy it and be in the now, and worry about tomorrow, tomorrow. what's next on the blog? It's probably clear that I'm not coming back to weekly blog posts. Or that if I do, it's not going to be something I can rely on doing for the long haul. But I'm finally well enough that I do have the urge to write again. The urge to dive deep into tech and do weird things with it, just to horrify you all. I'm honestly not sure what's next. I have a few ideas I want to work on, but my time is limited. My business is taking a lot of my energy, as are my personal life and some medical needs. I might end up writing more things like this. Essays that wind through a personal topic. Dives into things like Lyme disease tests and the ways in which they're broken. Explainers about all the different pain scales. And I might end up writing tech posts. Explain something that I learned recently. Make a silly, terrible idea come to life, and subject you to it share it with you. Expound on leadership topics and the human side of tech. But when will it happen? I really don't know. I had quietly launched a Patreon page this year, and I've shut that down, since I can't keep that up with all of *gestures at her body's shenanigans* this. A dream I had, and still have, is to be able to make a living from my creative endeavors. That dream seems like it will remain just a dream, but things can always change. * * * And yet. And yet, I'm happy. Dreams are great, but they can't hold a candle to life itself. Life's pretty rad. In the last few months, I've written more poetry than ever before in my life. Better poetry. Meaningful poetry, silly poetry, poems about booty calls and poems about grief. My music has taken on more expression and more meaning. I've been able to write instrumental pieces that authentically communicate an arc I had through trauma and recovery. Connections with people have formed, and deepened. New people have entered my life, enriching it more than I could ever have imagined. Old friends have shown me so much love and support through my hardest moments. This isn't what I thought my life would look like. It's both harder and better than I'd ever dared to wish for. If you're wondering, I think this one from Alberta Health Services is the clearest, most useful pain scale. ↩

17th Jun 2026 • 1 votes
You can't always fix it

I have some weird hobbies, and one of those is opening up the network tab on just about anything I'm using. Sometimes, I find egregious problems. Usually, this is something that can be fixed, when responsibly reported. But over time, I learned a bitter lesson: sometimes, you can't get it fixed. Tracking my package Recently, I was waiting for a time-sensitive delivery of medication. It used a courier company which focused on just delivering prescription medications. I opened up the tracking page on my computer, and saw the information I wanted: the medication would probably arrive around 6 PM. But... what if there's more? And what are they doing with my data? Can anyone else see it? So I peeked at the network tools, and was disappointed by what I saw. The first time this happened, I was surprised. By now, I expect to see this. And what I saw was every customer's address along the delivery route. I also saw how much the courier would get paid per stop, what their hourly rate was, and the driver's GPS coordinates (though these were sometimes missing). After the package was delivered, the tracking page changed and displayed a feedback form, my signature, and a picture of my porch. The JSON payload no longer included the entire route, but it included my address, and the payload from an easily guessable related endpoint did still contain the entire route. And that route? It included other recipients' ids, which can be used to find their home addresses, names, contents of the package (sometimes), a photo of their porch, and a copy of their signature. Um. This is bad, right? Resposibly disclosing the vulnerability I've actually found approximately this vulnerability in two separate couriers' tracking pages (and they're using different software). One of them was even worse for them, it included their Stripe private key, I suppose as a bug bounty for people without ethics. And each time I find it, I try to report it. And I fail. They don't let me report it. These companies don't list security contacts. The staff I can find on LinkedIn or their website don't have email addresses that I can find or guess. Mail sent to the addresses I do find listed has all bounced. I tried going through back channels. I messaged the pharmacy which was using this courier. I talked to my prescriber, who was shocked at this issue. And the next time I got a delivery, it came via UPS instead (they do not have a leaky sieve for a tracking page, but they did "lose" my prescription once). But I don't know if they just did that for me, the miscreant who looks at her network tools? Or did they switch everyone over to a different courier? Either way, at least my data was safe now, right? It was, until I started using a different pharmacy, and this one is back to using the leaky couriers again. Sigh. It's not my responsibility to fix this (and I can't) I got pretty upset about this at one point. There's a security issue! Data is being leaked, I must get this fixed! And someone told me something really wise: "it's not your responsibility to fix this, and you've done everything you can (and more than you had to)." And ultimately, she was right. I was getting myself worked up about it, but it's not my responsibility to fix. Sometimes there will be things like this that are bad, that I cannot fix, and that I have to accept. So, where do I go from here? I could probably publicly name-and-shame the couriers, but it would not do anything productive. It would not get their attention to fix it, and it wouldn't be seen by the folks who need to know (pharmacists and prescribers). So I'm not going to disclose the specific company, because the main thing it would do is risk me getting in legal trouble, for dubious benefit. I've already notified the pharmacists and prescribers that I know; it's on them, if they want to let anyone else know.

2nd Mar 2026 • 1 votes
Some good English word datasets

I'm working on a silly word game right now, which means I need a list of words for it. I can't rely on the system word list, since this is for a web project. I'd also like, ideally, multiple lists of words with different criteria: most common words, all words, some subsets. It's not really clear how many words there are, because what is a word[1]? And how do you find what words are in use? Or decide that something's used enough that it merits being put in a list, rather than just being a misspelling or something used once or twice? You probably don't want every mashing of keys that someone's done in there, or adjklfalkjsdfaklsd and variations are going to take up some considerable space. So that results in a lot of different word lists being available! With that in mind, here are the two best word list sources I found and some of their properties. WordNet: a while back, researchers at Princeton put together WordNet, a database of relations between words. This includes definitions and synonyms and things like that. It's no longer maintained by Princeton, but was forked and maintained with a release 1 month ago. There are libraries to load it. Wortschatz Leipzig frequency lists: A project under Leipzig University, this is a collection of word frequency lists from various sources. The English downloads page has datasets drawn from news, the web, or Wikipedia. And there are two honorable mentions, which I want to call some attention to. These are useful, but not useful for this project. /usr/share/dict/words: Unix systems come with word lists installed, which is handy for things like spellcheckers. On my system, this file contains about 480k words. This is accessible through various libraries, too. Mine does contain things like "1080", so I want something a little cleaner. It's nice to have available! Licensing isn't super clear to me, but it can be figured out and the word list could (with appropriate licensing) then be distributed with the application. Wikitionary: It's available under permissive license terms! There's a page documenting different frequency word lists, which can be used to get the top words from different contexts, like the most common words on wikipedia, in TV scripts, or in contemporary poetry! It doesn't make the cut for me since the data is prepared for human consumption rather than being machine readable. The WordNet data and Leipzig frequency lists both need to be loaded and processed, but the formats are documented and can be implemented pretty easily, especially if you have a specific subset of the data you need. I'll be using the Leipzig data most likely for my silly little word game. I might combine it with WordNet to be able to pull up definitions, but we'll see! Someday I'd like to pull some of the Wikitionary data, because it's really cool and has a lot of different frequency lists. Like the one with the 2000 most common words in contemporary English poetry. That might not make the cut for this project, but that's just crying to be used for something else. For example, I maintain that "apple tree" is a word. "Tall tree" is two: "tall" is an adjective that's used to modify "tree" which is a noun. But that's not true with "apple tree" since, "apple" isn't an adjective, it's another noun. We use "apple tree" as one singular word, and it has its own dictionary entry. I'm only slightly joking. ↩

2nd Feb 2026 • 1 votes
Reflecting on 2025, preparing for 2026

As I do every year, it's that time to reflect on the year that's been, and talk about some of my hopes and goals for the next year! I'll be honest, this one is harder to write than last year's. It was an emotionally intense year in a lot of ways. Here's to a good 2026! Reflecting on 2025 Where last year I got sick and had time black holes from that, this year I lost time to various planned surgeries. I didn't get nearly as much done, because it was also hard to stay focused with all the attacks on trans rights happening. Without further ado, what'd I get up to? Professional I helped coaching clients land job and improve their lives at work and beyond. I started coaching informally in 2024, and in 2025 I took on some clients formally. During the year, I helped clients improve their skills, build their confidence, and land great new jobs. I also helped clients learn how to balance their work and home life, how to be more productive and focused, and how to navigate a changing industry. This was one of the most rewarding things I did all year. I hope to do more of it this coming year! If you want to explore working together, email me or schedule an intro. I solved interesting problems at work. This reflection is mostly private, because it's so intertwined with work that's confidential. I learned a lot, and also got to see team members blossom into their own leadership roles. It is really fun watching people grow over time. I took on some consulting work. I had some small engagements to consult with clients, and those were really fun. Most of the work was focused on performance-sensitive web apps and networked code, using (naturally) Rust. This is something I'll be expanding this year! I've left my day job and am spinning up my consulting business again. More on that soon, but for now, email me if you want help with software engineering (especially web app performance) or need a principal engineer to step in and provide some engineering leadership. I wrote some good blog posts. This year, my writing output dropped to about 1/3 of what it was last year. Despite the reduction, I wrote some pretty good posts that I'm really happy with! I took a break intentionally to spend some time dealing with everything going on around me, and that helped a lot. I didn't get back to consistent weekly posts, but I intend to in 2026. Personal My hernias were fixed. During previous medical adventures, some hernias were found. I go those fixed[1]! Recovering from hernia repair isn't fun, but wasn't too bad in the long run. It resolved some pain I'd had for a while, which I hadn't realized was unusual pain. (Story of my life, honestly.) Long-awaited surgery! In addition to the hernia repair, I had another planned surgery done. The recovery was long, and is still ongoing. My medical leave was 12 weeks, and I'm going to continue recovering for about the first year in various forms. This has brought me so much deep relief, I can't even put it in words. Performed a 30-minute set at West Philly Porchfest. I did a solo set in West Philly Porchfest! All the arrangements were done by me, and I performed all the parts live (well, one part used a pre-sequenced arpeggiator). I played my wind synth as my main instrument, layering parts over top of myself with a looper, and I also played the drum parts. You can watch a few of the pieces in a YouTube playlist. Wrote and recorded two pieces of original music. This was one of my goals from 2024, and I'm very proud that I got it done. The first piece of music, Anticipation, came from an exercise a music therapist had me do. I took the little vignette and expanded it into a full piece, but more importantly, the exercise gave me an approach to composition. I'd like to rerecord Anticipation sometime, since I've grown as a musician significantly across the year. My second piece I'm even happier with. It's called Little Joys, and I'm just tickled that I was able to write this. I played it on my alto sax (piped through a pedal board) and programmed the other parts using a sequencer. One of my poems was published! I've written a lot more poetry this year. One of my close friends told me that I should get one of them published to have more people read it. They thought it was a good and important poem. That gave me the confidence to submit some poems, and one of them was accepted! (The one they told me to submit was not yet accepted anywhere, but fingers crossed.) You can read my poem, "my voice", in the December issue of Lavender Review. Last year's goals Every year when I write this, I realize I got a lot done. This year was a lot, filled with way more creative output than previous years. How does it stack up against what I wanted to do last year? ❓ Once again, I wanted to keep my rights. It's a perennial goal, and I did keep my rights in the state/community I live in. I'm awarding this one a question mark since my rights were under assault, and there are now many more places I cannot safely travel to. That means it's not a full miss, but not a win either. ✅ No personal-time side projects went into production! Yet another year that I toyed with the idea and again talked myself out of it. I'm taking it off the list for 2026, since the urge wasn't really even there this time. ✅ Maintained relationships with friends and family. I've had regular, scheduled calls with some people close to me. I've visited people, supported them when they needed me, and asked for support when I needed it. ❓ I did a little consulting and coaching, but didn't explore many ways to make this (playful exploration like I do on here) my living. I'm giving this the question mark of dubiousity, since I don't think I got much information from the year toward the questions I wanted to answer. ✅ Kept my mental health strong! There were certainly some challenges. What I'm proud of most is that I recognized those challenges and made space for myself. That's why I stopped blogging regularly: I needed the space to get through things with intact mental health. ❓ Did some ridiculous fun projects with code, but not as much as I wanted. The main project was making it so I can type using my keyboard (you know, like a piano, not the thing with letters on it). I had aspired to do more, and I'm glad I let myself relax on this. ✅ Wrote some original music! ✅ Also recorded that original music! It's on my bandcamp page. I am really proud of how much I did on my goals. I might be unhappy with my slipping on if it were a "normal" year where the government isn't trying to strip my rights, but you know what? I'll take it. Especially since I prioritized my health and happiness. Hopes and goals for 2026 So, what would I like to get out of this new year, 2026? These aren't my predictions for what will happen, nor are they concrete goals. They're more of a reflection on what I'd like this coming year to be. This is what I'm dreaming 2026 will be like for me. Keep my rights (and maybe regain ground). A perennial goal, I'd like to be able to stay where I am and have access to, I don't know, doctors and bathrooms. We've held a lot of ground this year. Hopefully some of what was lost can be regained. I'm going to keep doing what I can, and that includes living my best life and being positive representation for all others who are under attack. Maintain relationships with friends and family. I want to keep up with my friends and family and continue having regular chats with those I care about. We're a social species, and we rely on each other for support. I'm going to keep being there for the people I care about when they need me, and keep accepting their help as well when I need them. Spin up my business. I'm going out on my own, and I'm going to be offering my software engineering services again. By the end of the year, this will hopefully be thrumming along to support me and my family. Publish weekly blog posts (sustainably). I'm back in the saddle! This is the first post of 2026, and they're going to hopefully keep coming regularly. To make it sustainable, I'm going to explore if Patreon is a viable option to offset some of the time it takes to make the blog worth reading. Record a short album. I have a track in progress, and I have four more track ideas planned. I accidentally started writing an EP, I think??? This year I would love to actually finish that and release it. Publish more poetry. Writing poetry this year was very meaningful, and it's deeply important to me. I want to get more of it published so that I can share it with people who will also be able to get deep importance from it. * * * That's it! Wow, the year was a lot. I've put a lot of myself in this post. If you've read this far, thank you so much for reading. If you've not read this far, then how're you reading this sentence anyway? 2025 had a lot in it. There were some very good things I am very grateful for. There were some very scary and bad things that I wish had never happened. All told, it's been a long few years jammed into one calendar year. I hope that 2026 will be a little calmer, with less of the bad. Maybe it can feel like just one year. Regardless, I'm going to hold as much joy in the world this year as I can. Please join me in that. Let's fill 2026 with as much joy as we can, and make the world shine in spite of everything. The surgeon really meshed me up! ↩

5th Jan 2026 • 1 votes
A new system for organizing my writing and projects

Keeping up with regular blog posts is a challenge. To do it, I've churned through a few different organizational systems. Sometimes I have to change them due to actual life circumstances changing; other times, due to the old one just wearing off[1]. Well, my life circumstances have changed again. I'm on medical leave right now recovering from surgery[2], and I have been writing less this year anyway. So it's that time again: new system, baby. This post isn't a how-to. It's not a prescription for your own organizational system. Rather, let it be an invitation to examine your own systems, and decide whether they're serving you or not. Out with the old So what am I replacing? Before this change, my system involved Obsidian and LunaTask: Obsidian for recording ideas LunaTask for keeping track of works in progress and things that have a due date When I get an intrusive thought about how I definitely need to try to make a transistor at home in my garage[3], I put that in my list of ideas. This list of ideas contains some ridiculously bad ideas, and some very good ideas. But the thing is, I can't tell which is which when I have the idea! The ideas have to sit and marinate, and then eventually one sticks in my head long enough that it turns itself into a blog post. With most of the already completed ideas removed, that file is over 150 lines long right now. LunaTask is used to get these ideas from my brain over the finish line. I have a project called "Writing" which is kanban style. Each blog post I earnestly intend to write gets added into it, and then I keep track of which ones are waiting, in progress, and completed. This is mostly useful for posts which require code to go along with them, or which require research[4]. Posts like this one that are all prose without research go pretty fast and never enter LunaTask. The thing is, this has been showing cracks. I've stopped recording blog post ideas all that much this year, because I've been stressed from "global events" (I'm a trans woman, my wife is an immigrant, etc.) and preparing for multiple surgeries. And tracking things in LunaTask is just not working for me at all, because it's disconnected from idea tracking but ends up being its own idea tracking, since I just really want to write all the posts. To be honest, I'm not 100% sure why LunaTask isn't working for me right now for this. It might just be that I need to perturb the system a bit to make a change. But no matter what the cause, it's time to change it. In with the new The new system is simpler than the old, because I'm removing a tool. I'm taking LunaTask out of my blog toolkit, and switching to tracking everything in Obsidian. The new system is going to be smaller: an ideas page in Obsidian a works-in-progress page a task tracker for recurring blog tasks I'm using a task tracking plugin in Obsidian for the recurring tasks, because there are some things I need to do on a regular schedule. If I want to publish on Monday, that means that the post needs to be written before the weekend and edited before Monday. I'm not using task tracking for actual posts because I know from personal experience that it ends up similar to my ideas page but with structure. I'm also not using Obsidian's Bases for this, because a raw text file is just the lowest friction way I've found to do what I need. A raw text file lets me move things around freely, add notes however and wherever I want, and do freeform brain dumps. Structure is a little tyrant that kills my motivation at certain phases, so bye bye structure, we're doing a coup here. * * * So far, so good with the new system. I'm sure I'll be onto another one eventually. The most important thing, I've found, is to remain mindful of whether or not the systems are working for you. Change them once they stop working. There's this phenomenon where self-help productivity books almost always help you when you adopt a new system. But the help doesn't last forever! It seems that changing your system helps in some way, perhaps from the increased awareness and mindfulness, or maybe from new dopamine. ↩ It was a planned gender affirming surgery. The recovery is long and challenging, but I'm recovering very well. I'm very thankful for the support I have from my friends, family, and community. ↩ How hard can it be? The hubris of a software engineer knows no bounds, except for that refactor you know you should do but that you think is just a little too hard. (And yes, making a transistor would be very hard, and no, I'm not going to do it. Probably. The thoughts are still here.) ↩ I have a partially written post on Lyme disease which I do intend to complete, but it's a topic that I really want to be careful around. And so it languishes. ↩

1st Dec 2025 • 3 votes

More in programming

Clip of me singing Despard in Ruddigore in 2013

A clip of me singing a funny song from Gilbert and Sullivan’s Ruddigore back in 2013

5 hours ago • 1 votes
How and Why fork() Uses Copy-on-Write

In this video, we look at why fork() needs copy-on-write, how it works inside the kernel, and a memory usage problem that Instagram encountered with Python.

12 hours ago • 1 votes
What we lost when we lost comments

Comments require commitment, but they’re worth it.

18 hours ago • 1 votes
Lighthouse map

Lovely global map with animated lights sweeping the waters

20 hours ago • 1 votes
Warming up the Puma master before it forks

Basecamp 5 runs on Puma in cluster mode: one master process with preload_app! and 63 single-threaded workers per host, deployed as a Docker container with Kamal. We serve Basecamp from several sites. Each site has its own web hosts and a read replica of the database, and writes go to a single primary database in one of them. On our busiest hosts, each deploy left up to 2,000 requests waiting while the new workers warmed up. We reduced those queues by running signed-in requests through the app in the Puma master, before it forked the workers. Why 63 single-threaded workers? Basecamp has always served web requests from processes rather than threads. It ran on Unicorn, which only does processes, until we moved to Puma in January 2025, and we kept the same setup: workers (Concurrent.physical_processor_count * 1.3).ceil threads 1, 1 preload_app! On a 48-core host that’s 63 workers, each handling one request at a time. We chose 1.3 after benchmarking HEY in 2023, when we moved our apps out of the cloud and onto our own hardware. We tested several combinations of workers and threads with a mix of GET and POST requests on a 32-vCPU VM. Every multithreaded configuration we tested was slower and handled fewer requests than single-threaded workers. Adding workers beyond about 1.2 to 1.3 per vCPU brought little benefit. The threaded workers spent a lot of their time waiting for Ruby’s global VM lock. That made single-threaded workers a good fit for this workload, and we use the same setup for Basecamp. An app that spends more time waiting on its database or other services may benefit from more threads, so benchmark your own app. The other reason is the app itself. Basecamp has class-level state in places and has never needed to be thread-safe. With one request per process, it still doesn’t. Processes do use more memory than threads, and preload_app! reduces the difference. The master loads the app once and the workers share its memory through copy-on-write until they write to it. Shopify’s comparison of Ruby execution models explains the trade-off well. In the HEY benchmark the best setup came to about 260 MB of PSS per core, where PSS counts each shared page once, split between the processes using it, and the gap to a threaded setup was smaller than we’d expected. What Puma does on each host when a container starts: one master, then 63 forked workers that share its memory until they write to it. Two things about this setup matter for the rest of the post. A worker that’s compiling or loading something is fully blocked — there’s no other thread to pick up the next request. And whatever the master has in memory before it forks, all 63 workers share. Whatever they build after the fork, they build 63 times. What happens when we deploy Kamal starts the new container alongside the old one, and kamal-proxy moves the host’s traffic across as soon as the health check passes. At that moment, the new workers have handled health checks but no customer requests. preload_app! means the master loads the app once and the workers inherit it through fork. That covers the code. It doesn’t cover anything Ruby and Rails set up on first use: YJIT compiled code. YJIT compiles a method once it’s been called a certain number of times. The master calls very little during boot, so every worker compiles the same methods again on its own first requests. Compiled templates. Action View turns each ERB template into a Ruby method the first time it’s rendered. The schema cache. Active Record reads each model’s columns from the database the first time that model is used. Inline caches and memoized values throughout Ruby, Rails and the app. All 63 workers did all of this at once, while serving the traffic the old container had been handling a second earlier. In the test environment with YJIT on, the first request to a project page on a cold process took 652 ms, 151 ms of it YJIT compiling. The same request to a warm process took 28 ms. In production, CPU time per request peaked at around 200 ms while kamal-proxy moved traffic to the new container, against about 30 ms once the workers had warmed up. A host with spare CPU absorbs this. Every one of our web hosts has 48 cores and 63 workers, but each Amsterdam host serves around 250 requests per second, against 25 to 60 at our other sites. In Amsterdam the slow first requests turned into a queue. At a peak-hour deploy, the Puma backlog on an Amsterdam host reached anywhere from 250 to 2,238 requests, and kamal-proxy’s p99 response time hit about 10 seconds. Eron, our Director of Operations, had been tracking this since June. Another server in Amsterdam would help, but it would take weeks to arrive, so we also wanted to make deploys cheaper on the hardware we already had. What didn’t work We tried a few things first. In June, Donal tested the first two on a single Amsterdam host, comparing it with its neighbors, and they ruled out two likely causes. Warming each worker’s database connections. Puma’s before_fork hook clears the master’s connections, and each worker opened its own on its first request. Opening them in before_worker_boot instead made no difference. Queries on a freshly booted production host were already under a millisecond, so connections weren’t the problem. A synthetic request in each worker. Next, each worker made a few requests in before_worker_boot to an internal controller that touched every model. That ran the middleware, routing and Active Record paths, but it ran them in 63 workers at once — exactly the CPU spike we were trying to avoid. And a request with no real data renders no real views, so most of the app stayed cold. Spreading YJIT compilation out. Delaying YJIT in each worker by a random interval spread the compiling out over a few minutes, but every worker still ran interpreted until its delay ended. The queue didn’t change. Reforking from a warm worker. This is what Shopify’s Pitchfork does: let one worker serve traffic until it’s warm, then fork the others from it. Puma has an experimental version called fork_worker, and on beta it worked — the reforked workers were warm after three to five requests, where fresh ones took up to 30 seconds. But with fork_worker the template is worker 0, and it keeps serving requests. If it exits, the workers waiting to be forked never start (puma/puma#3596). If it gets no traffic, the refork never happens, which is what we saw on beta. Instacart have a mold_worker patch that promotes a warm worker to a template that stops serving, but it isn’t in a Puma release. We have a branch of it, and we may come back to it. That last experiment did show us where the fix was, though. Everything a warm worker has that a cold one lacks is in its memory, and fork copies memory. The master already has the app loaded. It just never runs it. Run the requests in the master So now, before the master binds its socket and forks, it makes the app’s own requests, in-process, the way a signed-in user would. Rack has a hook for exactly this. Rack::Builder#warmup takes a block that’s called once with the built app, before the server starts. rails server builds the app from config.ru, so the change to boot is one line: require_relative "config/environment" warmup { WarmUp.configured.run } if ENV["WARM_UP"] run Rails.application With preload_app! this runs in the master, and the workers inherit whatever it did. Puma binds its socket after the app is built, so until the warm-up finishes the health check’s connection is refused and kamal-proxy keeps retrying. No request reaches a worker that hasn’t been warmed. The warm-up has three steps. After precompiling the views, it gives the page requests and schema loading a shared 20-second budget, checked before each page or model. 1. Precompile the views actionview_precompiler reads every template for its render calls and compiles each one with the locals it’s passed. For us that’s 1,394 templates in about two seconds. A first request to a project page then compiles 2 templates instead of 44. 2. Request the pages, signed in A small browser class makes the requests through Rack::MockRequest, with the two cookies a real sign-in sets, then goes back for each page’s lazy Turbo frames: class WarmUp::Browser def initialize(signed_in_as:) @client = Rack::MockRequest.new(Rails.application) @headers = { "HTTP_USER_AGENT" => "Basecamp warm-up", "HTTP_COOKIE" => cookie_for(signed_in_as), "bc3.warm_up" => true } end def visit(path) page = get(path) frames_in(page).each { |id, src| get(src, "HTTP_TURBO_FRAME" => id) } end private def get(path, headers = {}) @client.get("https://#{host}#{path}", @headers.merge(headers)) end def frames_in(page) Nokogiri::HTML5(page.body).css("turbo-frame[src]").map { |frame| [ frame["id"], frame["src"] ] } end end The requests are signed in. The user is a monitoring account we already use for automated checks, and the pages are its own project, Campfire, to-dos, documents and messages. Public pages weren’t enough: after warming up with signed-out pages only, the first signed-in request to the projects page still took 131 ms, because authentication, the signed-in controllers and their views had never run. With signed-in pages it took 40 ms. cookie_for writes the same signed cookie the sign-in controller does, using the app’s own cookie jar, so there’s no API token and no secret to store. The frames are followed. The busiest HTML requests in production aren’t pages at all but Turbo frames — the sidebar badge, the inbox, the navigation menus. The browser parses each page and requests its <turbo-frame src> URLs with the Turbo-Frame header, so those controllers and views get warmed too. Our first four pages turned into 60 requests. The requests are excluded from rate limiting. They are internal, so they do not count against the rate limits that apply to real visitors. 3. Load the rest of the schema The page requests load the schema for the models they touch. The last step loads the rest, from the read replica: ApplicationRecord.reading do models.lazy.take_while { time_left? }.each { |model| model.load_schema if model.table_exists? } end The step checks 261 models and loads any schema information still missing. Those database round trips add up when the primary is far away: outside a request, Active Record uses the writing role, and from a host a long way from the primary each round trip is tens of milliseconds. Reading from the local replica brings the step down from about 20 seconds to 3.5. The pages go first because they load most of the schema anyway. If the time budget runs out, the step stops, logs how many models it got through, and the workers load the rest on first use like they always did. Rails can also load the schema from a dumped cache file at boot (bin/rails db:schema:cache:dump), which would make this step unnecessary. We don’t ship one in our image yet, because the dump needs a database to read from at build time, and we have several databases to cover. It’s on the list. What to close before the fork Running requests in the master opens things the master never opened before, and every worker inherits them. Two processes writing to the same socket will corrupt each other’s traffic, so you need to know what’s open before you fork. The way to find out is to list the master’s open file descriptors — ls -l /proc/<pid>/fd — before and after a warm-up, in an environment set up like production. Development wasn’t enough for us: it stores files on disk, so our S3 connections only showed up in production. Then, for each thing that’s open, check how its library handles a fork. We found three kinds: Already handled. Plenty of libraries detect a fork on their own, either by recording the PID they connected from and reconnecting in the child, by opening per-process files, or by resetting their thread pools. Redis clients, metrics libraries and concurrency libraries tend to be in this group. Check, but you probably don’t need to do anything. Already closed. Database connections are the classic one, and most Puma configs already clear them in before_fork. Anything else that’s opened per process — we have a SQLite cache the workers open on boot — needs closing when the warm-up finishes. Needs a new step. HTTP clients with keep-alive connections are the ones to look for: cloud SDKs with connection pools, tracing exporters, error reporters. They usually have no fork handling at all. We empty the aws-sdk connection pools in before_fork, and we run the warm-up untraced so the OpenTelemetry exporter never opens its connection to Tempo in the first place. Once that’s done, before_fork finishes with Process.warmup, which Ruby 3.3 added for this purpose: a major GC, a heap compaction, and every surviving object promoted to the old generation, so the memory pages the workers share change as little as possible afterwards. Choosing the pages The first list was the four pages that ran the busiest requests on beta. Once the warm-up was live, production showed us which endpoints were still cold. For one deploy, we compared each endpoint’s mean duration in the six minutes after kamal-proxy moved traffic to the new container with the same endpoint an hour later, then multiplied the difference by the number of requests in those six minutes. That gives the extra time each endpoint cost us because it was cold: Endpoint Cold Warm Requests in 6 min Extra seconds Campfire 246 ms 70 ms 6,490 1,140 Projects (JSON API) 84 ms 50 ms 22,077 771 Docs & Files 262 ms 177 ms 4,996 421 To-dos tool 205 ms 113 ms 4,018 371 To-dos (JSON API) 33 ms 16 ms 18,738 320 The pages already in the warm-up showed what to expect: the project page kept a 36 ms gap after a deploy, and the to-do page 10 ms. We’ve proposed adding these five requests, and expect them to add about five to seven seconds to the page step. The two JSON endpoints were a surprise. The warm-up’s page list had no API requests in it, so nothing on the API path had run before the first real request: not the API controllers, and not the Jbuilder templates rendering real records. Precompiling the views covers JSON templates too, but it isn’t a substitute for running the request. Results The warm-up is on for all 68 web hosts. With the first four pages it took 12 to 16 seconds per host: about 2 seconds to precompile the views, 7 to 9 for the 60 requests, and 3.5 for the schema. Deploys take that much longer per host, and we raised the deploy timeout from 30 to 60 seconds to cover it. In Amsterdam, at a peak-hour deploy: During deploy Before After Peak Puma backlog per host 250–2,238 requests 19–223 requests Peak kamal-proxy p99 about 10 s 2.4–4.8 s Peak CPU time per request 201–214 ms 88–132 ms Peak database time per request 56–69 ms 39–47 ms The same eight hosts at three deploys on 1 October, an hour apart, as the warm-up went from one host to four to all eight. The deploy in the middle, with four hosts warmed and four not, shows why every host needed the warm-up. Each warmed host recovered faster on its own: mean request duration peaked at 130 to 173 ms, against 203 to 311 ms on the hosts that weren’t warmed. But the backlog on all eight was about the same, because they were all waiting on the same database. Mean request duration on each host at the 07:21 UTC deploy. Blue hosts warmed up in the master before forking, orange hosts did not. Memory came down too. The workers now share compiled templates, YJIT code and the schema with the master instead of each building their own copy. On beta, the view precompiler alone took a busy worker’s private memory from 174–202 MB to 119–135 MB. Thirty minutes after the deploy, the web containers used about 39 GB less memory than the previous day’s containers at the same age and traffic. Amsterdam served most of our traffic at the times we tested. In Amsterdam, each new container used about 2 GB less just after traffic moved to it, which lowers the peak while the old and new containers overlap. Working with Claude Claude Code helped throughout. It combed through the per-worker backlogs and per-endpoint timings in Prometheus and Loki after each deploy, worked out the cold-versus-warm cost of each endpoint, and prepared the changes and the pull request descriptions with the benchmarks in them. We decided what to try, deployed it and read the results. If you do this Warm the master before it forks. Compile common code and templates and load their schema in the master, so workers inherit that work. With preload_app!, Rack::Builder#warmup runs before the workers start accepting traffic. Use the app’s real requests. Public pages, internal endpoints and synthetic queries warm the paths they run and nothing else. Signed-in requests to real records, frames included, run what production runs. Measure the cold penalty per endpoint. The difference between an endpoint’s cold and warm duration, times its request count after a deploy, ranks the pages worth adding. Ours weren’t the ones we’d have guessed, and two of them were JSON. Check what the warm-up leaves open. List the master’s file descriptors after a warm-up and account for every one before the fork. Two of ours needed changes. Set a time budget. A warm-up that runs long on one slow host fails the deploy on that host. Ours gives the page requests and schema loading a shared 20-second budget, checked before each page or model, puts the most valuable pages first, and logs what it skipped. Reforking from a warm worker, as Pitchfork does, solves the same problem continuously rather than once at boot, and it would warm paths no fixed list of pages covers. We may still get there: our branch brings Instacart’s mold_worker up to date with Puma’s main branch and fixes the bugs we found in it. But warming the master works with the Puma we already run, took a few days to implement, and substantially reduced the queues after deployment.

yesterday • 1 votes
📚 BoredReading

You seem to be enjoying this.

Join free to unlock everything.

Create free account

Already have an account? Sign in