Full Width [alt+shift+f] Shortcuts [alt+shift+k]
Sign Up [alt+shift+s] Log In [alt+shift+l]
1

A Thousand Tiny Optimisations

from Lee A Johnson [alt+shift+b] in programming

This is a sort-of-transcript of a talk i gave a couple of years ago, and planned to give at FOSDEM but it wasn't accepted. There is also a companion playthrough video that goes into much more detail. Note that the post was written well over a year ago, as usual it just didn't make it out of my drafts folder. Zelda is back in people’s minds again with the recent release of “Tears of the Kingdom” - the fastest selling Zelda game of all time. I haven’t started playing that one yet as I have only just started the previous game, “Breath of the Wild”. Not just that, I have been playing a much older Zelda game a lot. By “a lot” I mean over and over and over… … And Over And Over Again I used to occasionally watch speed runs of games on YouTube, those where the players attempt to complete the game as quickly as possible. This is more out of curiosity than any sort of appreciation of the skill, such as absurd examples like a player completing Final Fantasy VII in less time than the sum of the requisite unskippable content1. Inevitably the algorithm started throwing videos at me, but it was all the same type of content in which a player had spent hundreds of hours optimising their execution of the game controls to a point it’s almost subconcious. This kind of stuff, while impressive, is tedious. However, the algorithm did throw one particular suggestion at me that I found interesting: “A Link to the Past by Andy in 1:14:58”. This one showcased all sorts of ways the game could be broken by inputting specific movements and/or frame-perfect timing2. I had fond memories of this game from my childhood, and because I watched that video in its entirety the algorithm started throwing more at me that piqued my interest further. A Link to the Past The third game in the Zelda series, “A Link to the Past” was released in Europe in late 1992. I saved up many weeks of paper round money to eventually purchase a SNES and the game in the summer of 1993. I then spent the rest of the summer...
8th Jun 2025

Stay updated

Get a weekly newsletter with the top 5 articles worth reading every week.

More from Lee A Johnson

I Bought A Scanner (No, Really This Time)

This is a transcript from a talk I gave at the German Perl Workshop earlier this year. If you'd prefer to watch the video recording, you can find it here. I have lots of photographic projects on the go. Lots of these being on film, as some of these I started shooting a long time ago. I don’t have any particular loyalty or attraction to film, it’s just that I started shooting many of these projects before affordable medium format digital was available. Since I mostly shoot medium/large format film I never really jumped to digital until recently, so film has continued to feature heavily in my workflow. That said, it’s a pain in the arse to shoot film now given the spiraling costs, limited availability, and issues around traveling with it: modern airport CT scanners, being rolled out across many airports, are much more convenient but will fog film. Asking for a hand inspection often comes down to arbitrary timing - how busy the security is, how experienced the operator is, or if you’re lucky/unlucky. I’ve had film forced to be scanned (and fogged) and politely argued with security on more than one occasion. I don’t want to deal with that so don’t travel with film anymore, thus I am shooting less of it and have mostly moved to digital. I still have a tonne of film I need to scan and process however. Here’s just some of the binders and files of film. I don’t plan to scan all of this, but I do plan to scan the ones I need to. Probably in the region of a couple of thousand frames. I want to scan to the highest possible quality (within reason) for archiving, book projects, and large prints. If you’re wondering how large I print, it can be up to 160x60cm panoramics for selling. This is restricted by the size of my printer (that’s another story). Three Years Ago Three years ago I almost bought a scanner. I ended up blogging about it and the post got a bit of traction on Hacker News (HN). I’m never quite sure which posts I submit will pique the interest of the users. I’ll spend months chipping away at a draft and when I post it it tanks. Or I’ll cobble something together in twenty minutes, like the linked one above, and it gets 440 points and over 300 comments… The thread had some useful suggestions and some not so useful ones, the not so useful ones being effectively “buy an Epson”: I’ve had one for fifteen years and it’s not good enough for large prints or archiving. It’s passable for web stuff and smaller prints, but for my recent use cases? Not even close. Ten years ago I had negatives scanned with a high resolution scanner for the first time and recently, wanting to scan my archives for various projects, I decided I should invest in one of those scanners. The Original Plan The plan, back in 2023, was simple: Buy scanner (at significantly reduced rate) Scan all my film Sell scanner Profit! And I mean profit - the scanner that I almost bought was being offered to me at about 2/3rd of the price they usually sell. And they’re becoming harder to find in working order so the prices are going up. Or profit in not having to pay > 25.- CHF per frame to have someone else do this. You can see the pricing from The Film Lab. You can read the original blog post to find out more about the scanner in question, so I won’t repeat it here. Other than the parts being relevant to the rest of this post, namely that the scanner was showing hard and soft problems. The software that drives the scanner was last updated in 2012, it’s proprietary and closed source, requiring 32bit architecture and no third party drivers or software exist. So you are stuck using old software/computers to run it. Or maybe you could use emulation / virtualisation? The problem there is that the interface is firewire, or SCSI on the even older models, and firewire is known to be problematic on these scanners as the controllers start to go bad after a decade of continued use. That’s a risk, and the scanner was very much EOL as the firewire controller was dying: both ports were bad that suggests controller, not ports. The scanner would have been €5,000 to purchase and then €3,000 (ish) to repair. Or, as HN suggested - just open it up and use a soldering iron. I’m not going to drop 5k on something and then start poking it with a soldering iron. I’ll pass on that thanks. Camera Scanning In the meantime I’ve been camera scanning, which you can read about in another blog post. But how does that compare cost wise? It’s expensive because you’ll need a high resolution camera, a macro lens, copy stand, negative carrier/holder, and quality light source. You’ll look to spend anything from three to five thousand Euros on everything. Camera scanning does actually work well, in that it’s close to a high resolution dedicated scanner. But you have to setup the entire thing every time you want to use it, including ensuring everything is straight and parallel. It also suffers from the same weakness as most other scanning methods. What do you think that is? Film Flatness Or lack thereof: Film is rarely flat, especially so with 35mm. These are pretty mild examples of curl. It tends to be flatter in the larger formats but then you get into flatness issues due to it sagging. The smallest difference in the film plane can cause major issues in sharpness due to focus fall off (film scanning is essentially macro photography). Any workflow or solution that does not take this into account is significantly compromised. And the workflow is only as good as its weakest part. This is the biggest problem in scanning film - all other considerations are more than adequate these days: resolution, dynamic range, etc. However, most negative carriers don’t keep the film perfectly flat. This has always been a problem - this is from a book called “Edge of Darkness” which is about traditional analog photography and printing, and summarises the problems of negative carriers thusly: “if you use a glassless negative carrier, you might as well just buy the cheapest enlarging lens you can find. You are simply throwing away the money and sharpness you paid for it in your enlarging lens, and also in your fine camera and the expensive lenses you bought for it… No film will lie flat in a glassless carrier. That’s right, none… There is no avoiding this issue. Use glass.” So you have to use (anti-newton ring) glass, which introduces other issues - you’ve now got extra glass in the transmission path, and dust (which isn’t a massive problem, but a pain nonetheless). You could use drum scanning, which is absurdly impractical from a cost and operating point of view. Or you could use a Flextight, the scanner I almost bought three years ago. Interim Solution I stuck with camera scanning, but wasn’t happy though, because of film flatness and the setup faff. So of course I started looking for another scanner. I was idly browsing near the end of 2025 and came across this one. It’s exactly the same spec as the one I tried three years ago, except SCSI not Firewire so less prone to failure. It just predates Hasselblad buying Imacon (so is pre the rebranding, etc). It was in Switzerland so I could inspect and pick it up. It was also significantly cheaper than the previous one I had looked at, so worth a punt even if I needed to take a soldering iron to it. We went to St Gallen for a weekend and I picked it up. Here’s the software interface back in my studio. Look at that marvelous interface! None of that liquid glass bollocks. The first scans were promising, but I had the sense things needed some TLC. The first thing was calibrating the focus, which the software can do in combination with a focus slide. I was lucky that the focus slide was included with the scanner and I’m not sure what I would have done otherwise. Probably paid a fortune for a replacement? Possibly a lot of manual trial and error with the software? After doing that I scanned images of the 1951 USAF resolution test chart (taken on ultra high resolution 35mm film): That’s what the resulting scan looked like. Notice that it’s sharp from edge to edge, corner to corner. At 100% crop we can resolve around 110 to 123 line pairs per mm, which equates to about 5,600 to 6,300 DPI. This is beyond the limit of most 35mm lenses, but importantly - exactly to spec for this scanner. So I was happy the focus was calibrated. If you’re curious this is the same target with the camera scanning setup. It’s close, but we’ve got another variable in the workflow, several even, and that impacts the results. It’s not as sharp, and the extra glass in the transmission path causes aberrations. Another thing that needed attention was the power supply. The seller mentioned that “sometimes it takes five minutes to warm up”. Sometimes it was more than five minutes, and the power supply would click click click away. So that needed fixing and it was easy enough to find a compatible new replacement, however it cost 200 Euros. Expensive! The third problem I noticed was that some of the scans were coming out stretched. Often about 10% too wide/long, sometimes more than that. My panoramics looked panoooooooramic. I did some research and someone suggested this might be a “buffering issue”, which I thought was nonsense. Doing some testing I heard slipping sounds when the scanner was pulling the film into the body. After more research I stumbled on a post that suggested the belts need replacing. I opened the scanner up, and sure enough: A ha! You can’t quite see that the one on the back is even worse. I replaced those with compatible belts: 535 synchroflex t 2.5/245. Problem solved. The fourth problem was that the film holders were old and/or had been mishandled. They were falling apart and held together with electrical tape or glue, which didn’t seem optimal. Replacements cost 350 Euros in total for the four I needed. They’re now available cheaper from China, since the patents have expired. Or, you know, China. They used to cost about 200 Euros each from Hasselblad. The fifth problem, which is a potential one and hasn’t manifested yet, is that the lamps may eventually need replacing. I picked up a couple for 25 Euros. That seemed like a reasonable thing to do while they’re still available. Success? Let’s add up the costs of acquiring this scanner and renovating it: Scanner: 1,750.- CHF Power Supply: 175.- CHF Belts: 25.- CHF Film Holders: 350.- CHF Lamps: 25.- CHF Total: 2,325.- CHF (c. 2,500 EUR) In the last year (since acquiring the scanner) I have scanned: c. 250 panoramics frames (~ 6,000 CHF) c. 2,500 medium format frames (~ 80,000 CHF) c. 200 large format frames (~ 9,000 CHF) The figures in parentheses are what it would have cost me to have that number of frames scanned by a third party. That is, er, quite a saving. Also quite a lucrative business model perhaps? I think I can argue the cost of the scanner was a very good investment, and I haven’t finished using it yet. Even if it were to stop working tomorrow, it has already paid for itself many times over. Could it stop working tomorrow? Yes, because of other issues that will be harder to solve. The Bigger Issue(s)? A Power Mac G4 (discontinued in 2004). This came with the scanner, the necessary hardware and software to drive it, and is almost certainly living on borrowed time. Spinning metal is never good in the long-term. I’ll maybe purchase a backup soon, as these can still be found for a couple of hundred Euros. The key thing though, is that this very expensive, very high quality scanner, will at some point be rendered useless by the upgrade treadmill because the software required to run it will be increasingly difficult to run. A scanner that is still used by businesses, educational institutions, and individuals like me. A scanner that originally cost tens of thousands of Euros less than a decade ago. The upgrade treadmill is constantly whirring away. This is from the top of the Seattle Space Needle. “Do not upgrade anything on computer”. Clearly that notice speaks of someone being bitten by an upgrade at some point. I wonder is anyone else feeling the fatigue? Security updates, sure I can understand. But feature creep and trivialities? No! What tangible benefits have the last ten, fifteen, or even twenty years of OS updates brought? Other than security, and compatibility with newer hardware? New hardware is great, really, but by association forced deprecation of older hardware. No! It feels like the upgrade treadmill gets faster and steeper every year. Add to that subscription lock-in and dead endpoints: “I couldn’t vacuum my house because an SSL cert had expired” is what someone told me earlier this year. Fortunately this person is a software engineer so ended up man-in-the-middling the network traffic to get the vacuum cleaner to work again (no SSL-pinning it seems). “GoPro is announcing the end of life of the GoPro Quik app for macOS, effective at the end of 2024”. They discontinued the former in favour of their mobile app, which requires an account, login, subscription, and so on. I just want to transfer the videos from the hardware, I don’t need any of this crap (I don’t need any of that crap, it turns out GoPro haven’t locked the device down enough to prevent using third party apps to access the files. Yet). And, of course, software has to be in everything. These days the scanner would/could have an embedded Raspberry PI? Just a keyboard and mouse input, monitor and USB output would reduce the surface area, connectivity issues, and software dependency. Or software is never done? Because: externalities. I guess software is “done” when it’s no longer supported? Marciano Planque has a good piece on this: When hardware products reach end-of-life (EOL), companies should be forced to open-source the software. I think that’s a fair thing to say. I suspect Hasselblad/Imacon never open-sourced the software due to licensing issues. Or they just lost the source. Or they just don’t care, I don’t know. Maybe some combination of the three. And, inevitably, discontinued hardware like this scanner. Or, that is to say, discontinued parts? What about regulation changes? The panoramics I shoot are with a camera that was discontinued in 2004 because EU regulation banned lead solder in circuit boards. The company decided redesigning the parts wasn’t worth it. Old hardware has new exciting ways to fail. As time goes on components will fail or loosen - components that were expected to last decades. Then that results in tribal knowledge, or worse link rot and QR code rot. A lot of this stuff is hidden in walled gardens. There’s a Facebook Imacon group, for example. Why in the ever-loving fuck is a group for technical people, by technical people, on Facebook? Then there’s misleading AI. “My flextight scans are coming out stretched, what might the problem be?” LLM’s have gobbled up all the right information, and all the wrong information. Or information that is massively out of date. Nowhere in the suggestions here does it mention the belts might need replacing, which, according to my own research, is the most common reason these days. Legacy Software A decade ago I wrote an essay that also hit the front page of HN: All Software is Legacy. I think it is still relevant today, some parts not so much given we are now in The Age of Prompt, but mostly it’s still true. Nicholas always said “legacy software is the ugly stuff that makes you money”, which I think is true. But now it’s the stuff that surrounds us, like when I want to withdraw cash (guess what software most cash machines are still running?). Or when I want to take a train - when I gave this talk in Germany I had to get from the airport to the city centre. The ticket machines were disabled with a sign saying “no longer in use, download the app”. Then register. Then buy the ticket. I just want to give you money. Or when I wanted to pay for parking while stopping off at some random town in the UK - the same situation as with the ticket machines. “Download the app, register, pay”. Fuck that, I went and parked somewhere else. I just want to park, I don’t want to fight with software. Or if I want to hire a bike (not pictured: the half dozen apps on my phone to hire a bike). And when I want to buy stuff from a shop… One of the self-checkouts crashed recently in the coop, rebooting into a version of SUSE Linux from well over a decade ago. We’re collectively creating more and more of this everyday, letting it out into the world where it becomes a future liability for someone or the death knell for something. A pile of bikes, an unplugged ticket machine, a top of the line but no longer driveable scanner. References Imacon Users Group (the non-Facebook group) The state of Hasselblad Flextight scanners (2019) 1951 USAF resolution test chart Vlads Test Target Printer Story Original Scanner Blog Responses to HN Camera Scanning All Software is Legacy Repair Cafe

a week ago • 2 votes
The Caretakers

If you haven’t been paying attention, which is likely given the small corner of the web I’m going to cover, you may have missed that BINGOS just hit 247 consecutive monthly uploads to CPAN1. This is noteworthy as it surpasses DROLSKY’s record set a while back. That’s an upload to CPAN, once a month, for over twenty years. Once a month minimum I should say, as many of the authors on the board are uploading more frequently than that, some even daily. My own monthly upload streak is a mere 149, or twelve and a half years, and really that’s only because I pay attention to the board. Once a month I will check it to remind myself to dig through my CPAN distributions for bugs, issues, and refactoring opportunities. This keeps their internals fresh in my mind for when reports from users or contributors come in. The other authors on the board maintain an order of magnitude more distributions than I do, so are more likely to get a steady stream of reports from users. A stream that is turning into a torrent given the current place we find ourselves in the world of open source software. And that’s where I start to worry a little. Here in the Perl world we’ve been quietly getting on with it since the first shouts of “nobody uses Perl anymore” going back almost as long as BINGOS’ streak. Quietly maintaining critical systems, billion dollar businesses, backwards compatibility, security patches, the occasional new feature, and yearly releases of the interpreter. But that torrent of reports and issues is going to become a flood, and it’s not just the Perl ecosystem that is at risk. CPAN being Perl’s software repository. The upload stats can be found at cpan.io ↩

8th Aug 2026 • 1 votes
A World of First Drafts

Five short essays on where we are currently at. A World of First Drafts Dreams and the Uncanny Valley Fiddlers in the Room The Last Hike Semantic Satiation of an Acronym A World of First Drafts I recently picked up a copy of “An Evening With Windham Hill”, a collection of early 1980s live performances by some of Windham Hill’s popular-at-the-time acoustic guitar players. I primarily bought the LP for the recording of “Turning: Turning Back” by Alex deGrassi, as it was the only way to get this on vinyl; the 1992 retrospective that also contains the track was never (officially) released on anything but CD. The album contains a more interesting track on it, that being the first (?) performance of a Michael Hedges composition. Introducing the performance Hedges says “This is a new piece for [that] started out for guitar and then, er, all of a sudden it needed piano and about a week ago it needed bass so… We need to play it tonight. It’s dedicated to Steve Reich, and it’s called Spare Change.” The performance starts out very Hedges like, with his (now) distinct playing style and tapping on his instrument, but then after about forty seconds the piano comes in and Hedges influence seems to be diluted. He pulls it back but seems to be fighting with the piano, and when the bass solo arrives at three minutes Hedges is then completely lost to Manring. The three instruments then battle for the remaining two minutes of the composition leaving us at an ending that feels unresolved. So unresolved it takes the audience several seconds to realise the performance is over. The composition and its performance was very much a first draft. Some interesting ideas in places, but ultimately lacking cohesion, unsatisfying, and even forgettable. Hedges’ voice (that of his guitar) is lost amongst the parts that aren’t his. When Hedges released the final version two years later, on the album “Aerial Boundaries”, he knew the piece needed work so he made some major changes. The first was to drop the key two semitones lower. I guess he removed the capo from his guitar. The second change was more substantial: Hedges decided to replace the piano and bass parts with his own guitar, spending over one hundred hours recording sounds, looping them, playing them backwards, splicing tapes, pulling them apart and sticking them together, experimenting, all to get the textures that fit. Over one hundred hours in the studio in 1983, and likely many hours after those first drafts in 1982, to create a five minute long piece of music. Hedges could have just released the original arrangement, but he knew it was mediocre so continued to refine it. The version realised on “Aerial Boundaries”, and the liner notes explicitly use the verb realised rather than recorded, is probably the most striking work in Hedges’ entire discography. Because there are other remarkable recordings on the album, especially the title track, and due to the near impossibility of performing the track live, “Spare Change” is often ignored. That doesn’t mean Hedges’ time and effort to refine it was a waste, as the track still stands out today. Along with the first live recording, the final version is a permanent record of an idea elevated to something interesting, influential, even epochal. Fingerstyle guitar was going through some major changes in the early eighties, and Hedges was one of several important composers and performers at the time. Had he just sat back, been satisfied with the first draft, not torn it apart to rebuild, it would have been forgotten. I wouldn’t be writing about it today, and it’s possible it may have been discarded and failed to make the final cut for the album. Perhaps the composition would have been dredged up some time in the future, when an artist’s career reaches that inevitable scraping of the barrel stage. Alas, Hedges was killed in an automobile accident in 1997, at the age of 43. The obvious place I’m going with all this is that I increasingly feel like we are moving into a world of first drafts. One where ideas are formed but then delegated to something else to refine them. The problem with this approach is there is nothing new to pull from the tombola, you still have a first draft at the end of that process and haven’t gone anywhere compelling. I’ve even had someone defend their approach with “the ideas are all mine”. But ideas are easy, execution and refinement are hard and that’s where your individuality comes through. Maybe your ideas are boring, and your arguments are weak? I’d still like to read them in your style. Writers have built entire careers on this. The reason I read your blog, or listen to your music, or watch your videos, or look at your photos is because I’m interested in how you see and react to and interpret the world. If your voice and idiom is lost to the machine’s then you are no longer interesting and I’m no longer interested. Dreams and the Uncanny Valley A few years after moving to the mountains I had an idea for a photo project that would involve shooting images of well known vistas and then subtly replacing some of the mountain ranges with different peaks. Or removing parts entirely. Or just messing with the horizon in some way. I played with this a little when I printed some postcards of the Matterhorn, flipping the mountain on its horizontal axis to result in a mirror image. I assumed people would notice, given that concrete chocolate mountain is in the top five of recognisable peaks. Nobody did, at least not for five years. Or nobody said anything as they weren’t quite sure if what they were looking at was wrong, a literal uncanny valley moment. Sometime later I had a dream, or a nightmare since I tend not to remember my dreams, in which everything was subtly wrong. Like something out of a half forgotten Philip K. Dick story or tired science fiction trope. All my friends faces were a little different, but I was the only one that could see this. All the music I knew was off in some way, a different tempo or transposed, or key lyrics were swapped with synonyms. Nobody else noticed and sang along like it had always been this way. They were oblivious to the changes. The books and films all had slightly different titles and plots, and the actors cast were not the same. The walk I took to work was a diversion, my keys were interchanged, the office was rearranged, my keyboard layout was swapped. At lunch the microwave controls were on the left, the food tasted weird. None of my colleagues had any qualms. Finally I realised that if that was the case externally, at the macroscopic level, then was it also the case at an atomic level? I rushed home to find my Roche Biochemical Pathways poster, unfolded it (the wrong way), and stared at the molecules trying to recall the chirality of the amino acids, nucleic acids, and sugars. But I couldn’t remember, it was too long since I had done any biochemistry. And then I woke up. Fiddlers in the Room There’s a certain type of software engineer, I’ve met several times in my ongoing career, that I like to term a “fiddler”. When the fiddler is tasked with solving a problem, instead of thinking “I will solve this problem” they instead think “I will solve this particular class of problem”. The non-fiddler will pick an existing solution, and if there isn’t one that fits they will code for the explicit problem at hand. The fiddler will build an entire system to cater for a hypothetical future and other people’s unknowns. The result is something that is not actually generic enough to solve the particular class of problem, being too brittle, and is too broad to solve the original specific problem in a maintainable way. An excess baggage of logic and abstractions means cognitive overload. Having to think about the generic class of problem when trying to maintain for the specific problem is always, well, a problem. Don’t get me wrong. Fiddlers are important, and have contributed to many critical parts of many ecosystems, and some fiddlers are very good at what they do; however, they are the minority and, more often than not, a one hit wonder. A key tell of some fiddlers is that they like to use the latest tools to facilitate their fiddling, and this is the most dangerous fiddler of all. New tools are more likely to change quickly, and churn is the enemy in software engineering. Now it seems that everyone can be a fiddler, so my term has become a tautology or changed to mean something else. Fiddling has been flipped on its head. Now it seems fiddlers can trivially generate for the explicit problem at hand, pulled from a vast corpus of existing classes of problems, likely previously generated by the previous generation of fiddlers. Those not-quite-generic-enough-and-a-little-too-brittle solutions. And if the fiddling is not enough? Have an entire orchestra. Keep screaming “COMPUTER, DO SOMETHING!” If and when it all goes wrong? Yes: “Switching over to manual control… good luck!” and of course: “This is the world’s smallest violin, and it’s playing a sad song for you.” The Last Hike The last hike took place on the 25th May 2026. Temperature was 20ºC, humidity 53%. Distance 13.00km, altitude change +/- 566m. 735kcal was expended in energy, with an average heart rate of 107bpm (max 153, min 70). Duration was 3h21m. The computer decided the effort was “moderate”. The last hike was somewhat more successful than the last bike ride, which took place on the 19th May 2026. Temperature was 10ºC, humidity 66%. Distance 11.56km, altitude change +/- 258m. 356kcal was expended in energy, with an average heart rate of 138bpm (max 161, min 69). Duration was 41:31. Max speed was 51km/h when a catastrophic puncture of the back tire took place, sending the rider veering off course into a 1m wide grass verge between the asphalt and a barbed wire fence. The rider was thrown over the handlebars of the bike, onto their right shoulder. Some moderate skin grazing occurred, but a six hour visit to the hospital was required to confirm no bones were broken. The computer, knowing nothing of this, decided the effort was “moderate”. I decided to stop tracking all this stuff because I don’t know what advantages it begets other than the obvious: exercise is beneficial. The disadvantages are not worth it, trying to beat personal bests, going faster, worrying about regressions, contributing to the world’s largest and longest ongoing clinical trial? My father used to run a lot in his youth. He would take me in the pram when he ran, and he ran so much he wore through two sets of wheels. There’s little evidence of that now, maybe a few photos in our attic, and his arthritic knees. It’s possible that his running has nothing to do with his knees, I know. None of this was tracked, so who knows? Even if it was tracked we wouldn’t know anyway. All of this data works at a population level, individually you need to ask yourself if you really need to know the minutia. Exercise is beneficial until the bones hit the asphalt. Even then, the net benefit is worth it. Just move a few times a week for more than a few minutes and you’ll feel better. Semantic Satiation of an Acronym There has to be an end state to all this, eventually? Surely? A point at which it just becomes the norm, at which it is everyday, normal, mundane? A point at which we can just all get on with our lives and benefit from the good, and be protected from the bad. A point where the thing producing so much noise is reduced to a background hum that can be ignored. I can't seem to escape this in any place. Technical discussions normally reserved for office work hours spill out into social settings. The pub, the dinner table, with strangers on a plane, in a queue, at a gig. I hear horror stories of wannabee software engineers making all the mistakes newbies make, but with none of the guardrails or mentoring to steer them in the safe direction. I end up talking shop with people I would never want to be my colleagues. I'm not trying to gatekeep here, I have no kingdom to protect, rather I am worried that the bar being so low will cause many people to trip over it. I do find this stuff useful, in the very specific areas it is currently limited to. It will become more useful in other areas over time and hopefully less dangerous, but that's an unknown for now. I didn't think semantic satiation was possible with an acronym. Can I please just go one day without hearing about this. Just one.

14th Jun 2026 • 1 votes
Print Sales, Costs, And Profit: 2025

Apparently the bottom is falling out of the art market, at least for those in the high-end segment1. That market of investors, speculators, and dealers who are working with mind-boggling prices. This isn’t surprising. The buyers in this market aren’t interested in art, they’re interested in the financial returns, and the last few years have seen more compelling investments in other areas. Cryptocurrencies, NFTs, and now AI. Why faff about with tangibles, which need shipping, storage2, maintenance, and insurance? Why bother with art that needs decades to see a return, when the alternatives can be turned around in months with a few clicks? At least when the bottom falls out of the art market you will still have something to hang on your wall, eh? Anyway, it’s three years since I bought a printer. I thought I would provide an update on running costs, print sales, and profit. I’m happy to be completely transparent about all this as I know it can be difficult to figure this out if you are thinking of creating your own printing business. Culan et Croix des Chaux, December 2024 Again, the reason why I do this from November to November is that the ski season runs December to April, and usually the most sales are during that period. Also I bought the printer in November 2022 so I’m just carrying on that way. You can find details from the previous year here. All the photos here are work shot in the last twelve months, I’ll cover this in a bit more detail below in the “Thoughts” section. Year Three Investment I needed to purchase more paper and ink this year, as expected. I also decided to buy a stand for the printer, as it was available at a significant discount. This gives me more freedom in my studio as I can now move the printer around easily. It also will allow me to add a second roll feeder in the future, should I decide I need that. Investment figures are included in the costs below. Income Income is made up of sales through a local gallery that I am part of, along with web sales: 11,409.- CHF - sales through the gallery 642.- CHF - sales through web shop = 12,051.- CHF total income Total income is up by about 20% compared to 2024 on both gallery sales and web shop sales. Costs I include the loan repayment here as that is an ongoing cost. I had to purchase consumables (paper and ink), at a total of 1,370.- CHF however that cost isn’t included in the total here as it will be amortised in future years under the “print consumables” category when they are used to make prints. First Tracks, January 2025 Rental cost of space in the gallery was reduced due to having overpayments of rent and my loan to the gallery, from the previous year, paid back to me. -4,058.- CHF - loan repayment -1,500.- CHF - rental cost of space in the gallery -672.- CHF - print consumables (paper, ink) -130.- CHF - investment (printer stand) -3,408.- CHF - framing/mounting/postage -2,323.- CHF - commission to the gallery -300.- CHF - web shop subscription -62.- CHF - stripe / paypal fees 0.- CHF - advertising (none this year) = -12,453.- CHF total costs Total costs are down by about 15% compared to 2024. Total Profit (Loss) 12,051.- CHF total income -12,453.- CHF total costs = -402.- CHF total loss Thoughts Holy crap, I only made a 402.- CHF loss this year. What happened? A combination of debt owed to me by the gallery and better sales meant I came close to breaking even. The original investment loan also accounted for 4,058.- CHF in costs, but this loan is now fully repaid and will not factor into future costs. I was lucky this year in that one of the local schools made a bulk order of small framed prints, which contributed to the boost in sales. If I were to net off the cost of the loan, but also this large order (since this is a rare occurrence), I would have a profit of about 2,000.- CHF. So still a good year, relatively speaking. However, that’s about 175.- CHF per month profit. You can’t live on this, and it’s an absolute grind if you want to start increasing that. Again, the reality remains that a physical space is absolutely essential for getting the work out there and sold - just look at the figures, approx twenty times the amount of income than from the web sales. Verbier the day after after record snowfall, April 2025 The Content Treadmill This year’s “trip to a ski resort in another country” was Chamonix, which I hadn’t been to in over a decade. I hired a guide for some off piste ventures, and we had a great time exploring the mountains. The snow was reasonable at the start of the week, but a bit naff at the end, but I did manage to shoot a couple of photos that have been added to the work for sale. We already have a trip planned for next year, and are considering the year after. This should result in more work. I should note that the work is a byproduct of my first interest - snowboarding. Really I’m not looking for locations or trips with a thought to shoot photos, they are purely an afterthought. If the places I go are suitable, and the conditions allow, then a photo might result. I have no interest in creating “content”, as that’s an absolute grind as well. I might shoot a couple of good photos a year, really. The work I shoot is very specific to a place, as I’ve talked about previously. If I wanted it to sell I would be looking to display it in galleries present in those locations. Sunset Over Le Col de la Croix, October 2025 And if I did want to boost the income from all of this I would have to jump on that content treadmill. Be thinking all the time about where to go next, what to shoot, and how to sell it. Be pushing out work all the time. That means producing mediocre work, that you’re not happy with, and that is a drag. Going out with the intent of creating content? Where’s the fun in that? The Storm Hits the Art Market ↩ Geneva Freeport ↩

1st Nov 2025 • 1 votes
A Survey of the Ticket (Re)Selling Landscape

TL;DR? It’s a quagmire, essentially. But let’s dive in. Possible strong language ahead, and copious footnotes with links to much more information below. scalper noun [ C ] US informal /ˈskæl.pər/ (UK: tout) someone who buys things, such as theatre tickets, at the usual prices and then sells them, when they are difficult to get, at much higher prices: “A scalper offered me a $20 ticket for the concert for $90.” We got tickets for Oasis, and it wasn’t easy. In fact, it was an absolute pain in the arse. There seems to have been a change in the ticket buying, selling, scalping, and reselling landscape of late, or the nature of this event has just brought it to my attention. Doing some research around this all suggests it’s always been bad, and yes it has always been bad. But this bad? I’m not so sure. I’d speculate it’s now so easy to put together something to automatically poll for (and purchase) tickets, that anyone who is after them can do it and anyone who wants to scalp can do it as well. The consequence of all this is that a ticket can be resold multiple times. Originally, which is then listed for sale ethically, then snapped up by a scalper, then reduced to sell nearer the event, then grabbed by another scalper, then sold once more to a genuine fan again. All the while the platforms are taking obscene amounts of fees every time that single ticket sells and sells again. Here’s some observations from the experience of trying to purchase Oasis tickets, and the platforms we ended up looking at to do this. Ticketmaster Part of me feels sorry for Ticketmaster (or Live Nation, it seems). I believe they have a shit tonne of technical debt that would make even the most seasoned senior nope out before getting beyond the first interview. This includes such madness as a custom operating system built on top of VAX that is holding up most of the backend. Given the age of these systems it’s all now layers of emulation. They’ve tried to rewrite it more than once but never get anything close enough to the existing implementation in terms of performance and reliability1. Part of me wonders if that might be an interesting challenge, but the other part wonders why they don’t just fix these with a sledgehammer. Radiohead recently announced shows and worked to reduce the inevitable stampede by way of a two stage lottery; you had to first register an interest, with a verified mobile number and email address and then, if you were lucky, you were sent a code that allowed you to join a queue on the day of the sale. A moderately high bar to those that automate scalping, with an aspect of a lottery to reduce scalping even more. That system brought the queue size down to manageable numbers and did cut the scalping markets. Not completely, but significantly given the number of tickets on the resale sites was, and remains, relatively low: Note - dozens, rather than hundreds, and being sold at relatively low prices. Also note the warning that “Resale of tickets is prohibited for this event. The ID of everyone entering the event will be checked to ensure it matches the name on their personal ticket.” … Resale of tickets is prohibited says the resale site. Eh? More on this below in the Stubhub section. The final part of me assumes Ticketmaster just doesn’t care. Their platform can handle the load, and they make gobs of money so why fix the non issues? Sure, the site can struggle sometimes, but that’s transient. The tickets will get sold eventually. Why worry? The experience with Oasis was about the same as most in demand events - join a queue and then hope there’s something left when you get to the end of it. The queues were absurd in size. Hundreds of thousands in line for a venue of a tenth of that capacity, with most in the queue likely buying multiple tickets. The double punch was that at the same time we were sat in the queue, some minutes then hours after the sale opened, we saw thousands of tickets being sold on resale sites. Scalped in seconds, then listed at hugely marked up prices before the official ticket sales had even sold out2. So we didn’t get tickets out of this system when it opened for sales over a year ago. We weren’t massively bothered, oh well we won’t see them. Then a work event coincided with the final gigs at Wembley in September 2025. Thus began the attempts to get some on the resale market. Stubhub A resale site with no limitations. Tickets sold out on the primary sales platform? Try a resale site where you can find them at ten times the price. Hmm, right… And I mean ten times the price. Oasis tickets for the general admission floor standing areas in Wembley were selling at, minimum, 600 GBP. Seated 1,000 GBP+. Absurd. What I find particularly irksome is that Stubhub will allow listings of tickets from the moment the official sales open. If this is supposed to be a fan resale site, why isn’t there at least some sort of delay? A week after the official sales open? A day? A few hours? Something? A genuine fan doesn’t buy four tickets and then immediately relist them at ten times the price. Stubhub is aiding the scalpers. I’ve had a script for a while that will watch Stubhub and alert me when tickets appear below a certain price3. This is to work around the lack of support for alerts in Stubhub’s platform. I set the price to a reasonable amount, about 250 GBP for floor standing tickets. Still high, but not absurd. The script merely alerts me, it doesn’t do anything to add tickets to a basket, purchase them, etc. Just a headsup that some tickets are listed. So I setup an alert for Oasis tickets, specifically for the general admission floor standing areas in Wembley. In other words the cheapest tickets. I actually did this not long after the first sales opened, in the hope that we might grab tickets and then we could plan a trip. It’s always been a useful script, but absolutely nothing came out of it this time. The script didn’t alert me once. Resale tickets for this event were all listed so high that it was never going to succeed. No luck on Stubhub. Twickets Twickets are one of the ethical resale sites, and they make a big point of this fact in their marketing spiel. The problem, which became clear with this event, is that scalpers have automated the purchasing of tickets through their site and that locks out the real fans. To summarise: real fans who can’t make the gig sell them at face value on Twickets. Scalpers automate purchasing these and then resell them on other sites at inflated prices. This completely fucks the site’s ethos. It’s infuriating trying to purchase through Twickets as the way it works is tickets are listed until their purchase is confirmed, and adding them to a basket holds a lock on them for ten minutes. The scalpers have automated this such that their bot will add to the basket, giving the scalper the option to get in front of the screen to complete the purchase. Sometimes the scalper can’t get to whatever device they need to in time. So what happens? The tickets are available to grab again and their own bot, or another bot, or one of many other bots, adds to their basket and holds the lock on the tickets again. Rinse and repeat. Tickets would be listed, but when you tried to purchase you were told “Looks like another fan grabbed this ticket just before you - when someone starts the checkout process, we temporarily hold the ticket so they have a fair chance to finish buying it. If they don’t complete their purchase, the ticket will be released back for others to buy.” The tickets would sit listed for over an hour, meaning they were being added to a basket, timing out, added, timing out, added, over and over again. The above tickets were an example of that, eventually selling almost two hours after they were listed. Who takes this many attempts to purchase? I can understand a couple of timeouts, maybe even twice that many, but an hour and a half? Eight timeouts before a successful purchase? Bots, nothing but bots or a very lucky actual fan that manages to slip in between the cracks and hits the purchase button at the right millisecond. 90 quid tickets, on their way to being resold for a much higher price. That’s likely 1,000 quid profit for the scalper. We saw this happen multiple times, I’d say we saw twenty tickets sold at face value going to scalpers on this ethical resale site, before we gave up. Twickets claim they have strong bot defence, but the evidence suggests otherwise and that they are failing in their mission statement4. There are technical solutions to this, and I’ll give them this suggestion for free: they don’t have alerts for in demand events, which is understandable as sending out hundreds of thousands of notifications, emails, etc, is costly and would lead to a stampede to their site causing more problems. Instead they should send out notifications in batches - don’t list the tickets, make them only available through an unlisted unique URL, which is provided via the notification, send out small batches of notifications to randomly picked users (say, one hundred at once) at ten minute intervals, and when the tickets are sold - stop. The key here is the URL in the notification should itself expire, meaning the bot race is eliminated. Don’t get to it in 10mins? Too late. If you chuck the ticket in the basket but don’t complete payment? Sorry, you lost your chance. I know the above suggestion is a bit “shit Hacker News says”, but come on Twickets - sort this out as you’re more Fuckwits than Twickets at the moment. TicketSwap Another ethical resale site that seems to be doing better than Twickets, but I’m not convinced. They do send notifications, and they might be using batching as I talk about above to reduce load5. So fair play. However we noticed that some sellers appear to be selling multiples of tickets via this platform, as distinct listings over different hours/days/weeks, and that in itself is fishy. Witness this seller (username removed) who we saw at least seven or eight times selling moderate value tickets via the app. At the end of the process, i.e. shortly before the gig, they had sold over twenty individual tickets. That seems off, and perhaps it was a scalper offloading. My suggestion here would be to not allow a seller to list tickets for the same event more than once, but that’s probably trivial to work around if you’re a scalper. Reddit Reddit is full of scammers, to the point that it’s just hilarious to read their attempts to make excuses around what they’re doing. The pattern is usually an account registered very recently, perhaps a day or two, or a week or two; alternatively they have an account of much longer standing with higher karma, but they have deleted/hidden the entire submission and comment history. Those are probably compromised accounts, and at least one I saw was playing the longer game: It started as an account registered a year ago, with the first submissions cleary karma farming in some of the larger subreddits. Get the account to some degree of trustworthiness. They pivoted to trying to resell tickets to in demand gigs about a month ago. These scammers were typically bait and switch when it came to proving they had the tickets, or the payment method. One even claimed they had a friend at Ticketmaster and therefore couldn’t sell on the resale sites as it would get their friend into trouble. Bollocks. I’m not saying there are no legitimate sellers on Reddit, but with in demand gigs your odds of being scammer are almost 100% so don’t risk it. I saved a couple of the threads as PDF printouts for posterity, so entertain yourself here and here if you want to. The second thread is particularly silly as I get into an argument with a poster who is advertising their own alerting service for Twickets, which they charge for, knowing full well that there is no way it can beat the other bots. They even had a page on their own site effectively admitting this fact (Mainbotpy section). Selling a service you know to be ineffective? Another scam. Viagogo Where we ended up purchasing tickets. The same observations apply as for Stubhub, which is not surprising as they are the same company with different frontends, but Viagogo are particularly egregious as they don’t list tickets with their fees so you are in for a shock when you hit the final payment page. Fees came in at almost 50% of the listing amount: This was more than I wanted to pay at the point of purchase, but likely the least I was going to pay. The fees are a piss take. I don’t think there’s much more to say about this one. However the interesting observation is making an incorrect assumption that prices would come down as the event got closer, with scalpers getting more desperate to offload their unsold tickets. That didn’t happen, and considering the profits involved probably never happens. I watched the listing values slowly decrease over the period of about a week, along with the number of available tickets, and purchased two on the Wednesday before the Saturday event. I’d figured the amount of time we’d spent on this was way beyond sunk costs territory. Even if the tickets continued to decrease in price we had already lost more than their value in time. The Wednesday turned out to be the lowest point they hit, and they started going up again reaching bonkers prices on the Friday and then even higher on the Saturday morning: You probably can’t see it there. Just know that there were 42 tickets left on the day, with the cheapest being £1,741. The thing is, these were selling, or at least appeared that way since the number of listings reduced to dozens a couple of hours later, and then was down to single figures a couple of hours before the event. I don’t think the bottom one here sold though: Just Write a Better Bot? We’re beyond that. It only takes a single scalper to ruin it for everyone, and if everyone decides to run bots to buy tickets… well, I don’t think I need to expand on why this is a terrible idea. So What Now? The Resale market needs killing. Dead. As in: no resale of tickets for events is ever possible and you just have to eat that fact. Fuck all the businesses that have built a model on gouging fans, you’re all gobshites. The lot of you. Line goes up? Fuck your line. If you buy a ticket it should be cryptographically stored in your account in a way that cannot be shared (time linked rotating qr/barcodes like most ticket apps use these days). And it cannot be sold. Period. You can’t make the gig, you change plans, you’re ill? Too bad. That would stop the scalping market instantly. And if we aren’t going to kill the resale market, how about using some recent buzzword tech to improve the situation? NF-ucking-Ts that would have some tangible use and allow artists to benefit from the resales, track how many times tickets are being resold, as well as giving information that allows identification of the scalpers so they can be named and shamed? Do it. Please? Some information here, here, and here that notes “It runs a proprietary operating system on a proprietary emulator. Tickets are actually the equivalent of blocks on a file system. It’s incredibly complex in the way that it’s able to represent the physical layout of venues, view obstructions, language stuff, and most famously the packing algorithm.” A much more comprehensive article can be found here, which does cover some of “The Host”. ↩ I posted about this on Reddit at the time (archive of the thread here) and one comment suggested that the resale sites themselves are scalping for high demand events. Unproven, but it’s a possibility. ↩ I gave a lightning talk about this last year. The script is pretty basic, and the most interesting thing is probably StubHub building large parts of their site through JSON so it really makes this a trivial thing to do. ↩ Twickets do have bot defence in place, but it’s the most naive and/or simple implementation. Their thresholds for some checks are also way too high - it’s possible to sit and refresh the page every second and not see a 429 (“too many requests”) response. Many of their endpoints are also not rate limited, allowing for easy automation of polling, and their “private” API endpoints are not as private as they should be, nor are they properly secured. ↩ I suspect the batching is effectively done through whichever provider they are using to send the notifications to mobile devices. Some of the events (as seen in the screenshot) show 30,000+ people waiting for tickets. That’s not a massive number, but enough to cause quite a stampede should notifications appear on all those devices at once. I do assume it’s also trivial to automate grabbing the tickets after receiving the notification, but the “not all at once” delivery is going to render that process ineffective. TicketSwap may suffer from the same issues as Twickets, but that’s unproven without further investigation (I only ever use TicketSwap via a mobile app). ↩

17th Oct 2025 • 1 votes

More in programming

Clip of me singing Despard in Ruddigore in 2013

A clip of me singing a funny song from Gilbert and Sullivan’s Ruddigore back in 2013

5 hours ago • 1 votes
How and Why fork() Uses Copy-on-Write

In this video, we look at why fork() needs copy-on-write, how it works inside the kernel, and a memory usage problem that Instagram encountered with Python.

12 hours ago • 1 votes
What we lost when we lost comments

Comments require commitment, but they’re worth it.

18 hours ago • 1 votes
Lighthouse map

Lovely global map with animated lights sweeping the waters

20 hours ago • 1 votes
Warming up the Puma master before it forks

Basecamp 5 runs on Puma in cluster mode: one master process with preload_app! and 63 single-threaded workers per host, deployed as a Docker container with Kamal. We serve Basecamp from several sites. Each site has its own web hosts and a read replica of the database, and writes go to a single primary database in one of them. On our busiest hosts, each deploy left up to 2,000 requests waiting while the new workers warmed up. We reduced those queues by running signed-in requests through the app in the Puma master, before it forked the workers. Why 63 single-threaded workers? Basecamp has always served web requests from processes rather than threads. It ran on Unicorn, which only does processes, until we moved to Puma in January 2025, and we kept the same setup: workers (Concurrent.physical_processor_count * 1.3).ceil threads 1, 1 preload_app! On a 48-core host that’s 63 workers, each handling one request at a time. We chose 1.3 after benchmarking HEY in 2023, when we moved our apps out of the cloud and onto our own hardware. We tested several combinations of workers and threads with a mix of GET and POST requests on a 32-vCPU VM. Every multithreaded configuration we tested was slower and handled fewer requests than single-threaded workers. Adding workers beyond about 1.2 to 1.3 per vCPU brought little benefit. The threaded workers spent a lot of their time waiting for Ruby’s global VM lock. That made single-threaded workers a good fit for this workload, and we use the same setup for Basecamp. An app that spends more time waiting on its database or other services may benefit from more threads, so benchmark your own app. The other reason is the app itself. Basecamp has class-level state in places and has never needed to be thread-safe. With one request per process, it still doesn’t. Processes do use more memory than threads, and preload_app! reduces the difference. The master loads the app once and the workers share its memory through copy-on-write until they write to it. Shopify’s comparison of Ruby execution models explains the trade-off well. In the HEY benchmark the best setup came to about 260 MB of PSS per core, where PSS counts each shared page once, split between the processes using it, and the gap to a threaded setup was smaller than we’d expected. What Puma does on each host when a container starts: one master, then 63 forked workers that share its memory until they write to it. Two things about this setup matter for the rest of the post. A worker that’s compiling or loading something is fully blocked — there’s no other thread to pick up the next request. And whatever the master has in memory before it forks, all 63 workers share. Whatever they build after the fork, they build 63 times. What happens when we deploy Kamal starts the new container alongside the old one, and kamal-proxy moves the host’s traffic across as soon as the health check passes. At that moment, the new workers have handled health checks but no customer requests. preload_app! means the master loads the app once and the workers inherit it through fork. That covers the code. It doesn’t cover anything Ruby and Rails set up on first use: YJIT compiled code. YJIT compiles a method once it’s been called a certain number of times. The master calls very little during boot, so every worker compiles the same methods again on its own first requests. Compiled templates. Action View turns each ERB template into a Ruby method the first time it’s rendered. The schema cache. Active Record reads each model’s columns from the database the first time that model is used. Inline caches and memoized values throughout Ruby, Rails and the app. All 63 workers did all of this at once, while serving the traffic the old container had been handling a second earlier. In the test environment with YJIT on, the first request to a project page on a cold process took 652 ms, 151 ms of it YJIT compiling. The same request to a warm process took 28 ms. In production, CPU time per request peaked at around 200 ms while kamal-proxy moved traffic to the new container, against about 30 ms once the workers had warmed up. A host with spare CPU absorbs this. Every one of our web hosts has 48 cores and 63 workers, but each Amsterdam host serves around 250 requests per second, against 25 to 60 at our other sites. In Amsterdam the slow first requests turned into a queue. At a peak-hour deploy, the Puma backlog on an Amsterdam host reached anywhere from 250 to 2,238 requests, and kamal-proxy’s p99 response time hit about 10 seconds. Eron, our Director of Operations, had been tracking this since June. Another server in Amsterdam would help, but it would take weeks to arrive, so we also wanted to make deploys cheaper on the hardware we already had. What didn’t work We tried a few things first. In June, Donal tested the first two on a single Amsterdam host, comparing it with its neighbors, and they ruled out two likely causes. Warming each worker’s database connections. Puma’s before_fork hook clears the master’s connections, and each worker opened its own on its first request. Opening them in before_worker_boot instead made no difference. Queries on a freshly booted production host were already under a millisecond, so connections weren’t the problem. A synthetic request in each worker. Next, each worker made a few requests in before_worker_boot to an internal controller that touched every model. That ran the middleware, routing and Active Record paths, but it ran them in 63 workers at once — exactly the CPU spike we were trying to avoid. And a request with no real data renders no real views, so most of the app stayed cold. Spreading YJIT compilation out. Delaying YJIT in each worker by a random interval spread the compiling out over a few minutes, but every worker still ran interpreted until its delay ended. The queue didn’t change. Reforking from a warm worker. This is what Shopify’s Pitchfork does: let one worker serve traffic until it’s warm, then fork the others from it. Puma has an experimental version called fork_worker, and on beta it worked — the reforked workers were warm after three to five requests, where fresh ones took up to 30 seconds. But with fork_worker the template is worker 0, and it keeps serving requests. If it exits, the workers waiting to be forked never start (puma/puma#3596). If it gets no traffic, the refork never happens, which is what we saw on beta. Instacart have a mold_worker patch that promotes a warm worker to a template that stops serving, but it isn’t in a Puma release. We have a branch of it, and we may come back to it. That last experiment did show us where the fix was, though. Everything a warm worker has that a cold one lacks is in its memory, and fork copies memory. The master already has the app loaded. It just never runs it. Run the requests in the master So now, before the master binds its socket and forks, it makes the app’s own requests, in-process, the way a signed-in user would. Rack has a hook for exactly this. Rack::Builder#warmup takes a block that’s called once with the built app, before the server starts. rails server builds the app from config.ru, so the change to boot is one line: require_relative "config/environment" warmup { WarmUp.configured.run } if ENV["WARM_UP"] run Rails.application With preload_app! this runs in the master, and the workers inherit whatever it did. Puma binds its socket after the app is built, so until the warm-up finishes the health check’s connection is refused and kamal-proxy keeps retrying. No request reaches a worker that hasn’t been warmed. The warm-up has three steps. After precompiling the views, it gives the page requests and schema loading a shared 20-second budget, checked before each page or model. 1. Precompile the views actionview_precompiler reads every template for its render calls and compiles each one with the locals it’s passed. For us that’s 1,394 templates in about two seconds. A first request to a project page then compiles 2 templates instead of 44. 2. Request the pages, signed in A small browser class makes the requests through Rack::MockRequest, with the two cookies a real sign-in sets, then goes back for each page’s lazy Turbo frames: class WarmUp::Browser def initialize(signed_in_as:) @client = Rack::MockRequest.new(Rails.application) @headers = { "HTTP_USER_AGENT" => "Basecamp warm-up", "HTTP_COOKIE" => cookie_for(signed_in_as), "bc3.warm_up" => true } end def visit(path) page = get(path) frames_in(page).each { |id, src| get(src, "HTTP_TURBO_FRAME" => id) } end private def get(path, headers = {}) @client.get("https://#{host}#{path}", @headers.merge(headers)) end def frames_in(page) Nokogiri::HTML5(page.body).css("turbo-frame[src]").map { |frame| [ frame["id"], frame["src"] ] } end end The requests are signed in. The user is a monitoring account we already use for automated checks, and the pages are its own project, Campfire, to-dos, documents and messages. Public pages weren’t enough: after warming up with signed-out pages only, the first signed-in request to the projects page still took 131 ms, because authentication, the signed-in controllers and their views had never run. With signed-in pages it took 40 ms. cookie_for writes the same signed cookie the sign-in controller does, using the app’s own cookie jar, so there’s no API token and no secret to store. The frames are followed. The busiest HTML requests in production aren’t pages at all but Turbo frames — the sidebar badge, the inbox, the navigation menus. The browser parses each page and requests its <turbo-frame src> URLs with the Turbo-Frame header, so those controllers and views get warmed too. Our first four pages turned into 60 requests. The requests are excluded from rate limiting. They are internal, so they do not count against the rate limits that apply to real visitors. 3. Load the rest of the schema The page requests load the schema for the models they touch. The last step loads the rest, from the read replica: ApplicationRecord.reading do models.lazy.take_while { time_left? }.each { |model| model.load_schema if model.table_exists? } end The step checks 261 models and loads any schema information still missing. Those database round trips add up when the primary is far away: outside a request, Active Record uses the writing role, and from a host a long way from the primary each round trip is tens of milliseconds. Reading from the local replica brings the step down from about 20 seconds to 3.5. The pages go first because they load most of the schema anyway. If the time budget runs out, the step stops, logs how many models it got through, and the workers load the rest on first use like they always did. Rails can also load the schema from a dumped cache file at boot (bin/rails db:schema:cache:dump), which would make this step unnecessary. We don’t ship one in our image yet, because the dump needs a database to read from at build time, and we have several databases to cover. It’s on the list. What to close before the fork Running requests in the master opens things the master never opened before, and every worker inherits them. Two processes writing to the same socket will corrupt each other’s traffic, so you need to know what’s open before you fork. The way to find out is to list the master’s open file descriptors — ls -l /proc/<pid>/fd — before and after a warm-up, in an environment set up like production. Development wasn’t enough for us: it stores files on disk, so our S3 connections only showed up in production. Then, for each thing that’s open, check how its library handles a fork. We found three kinds: Already handled. Plenty of libraries detect a fork on their own, either by recording the PID they connected from and reconnecting in the child, by opening per-process files, or by resetting their thread pools. Redis clients, metrics libraries and concurrency libraries tend to be in this group. Check, but you probably don’t need to do anything. Already closed. Database connections are the classic one, and most Puma configs already clear them in before_fork. Anything else that’s opened per process — we have a SQLite cache the workers open on boot — needs closing when the warm-up finishes. Needs a new step. HTTP clients with keep-alive connections are the ones to look for: cloud SDKs with connection pools, tracing exporters, error reporters. They usually have no fork handling at all. We empty the aws-sdk connection pools in before_fork, and we run the warm-up untraced so the OpenTelemetry exporter never opens its connection to Tempo in the first place. Once that’s done, before_fork finishes with Process.warmup, which Ruby 3.3 added for this purpose: a major GC, a heap compaction, and every surviving object promoted to the old generation, so the memory pages the workers share change as little as possible afterwards. Choosing the pages The first list was the four pages that ran the busiest requests on beta. Once the warm-up was live, production showed us which endpoints were still cold. For one deploy, we compared each endpoint’s mean duration in the six minutes after kamal-proxy moved traffic to the new container with the same endpoint an hour later, then multiplied the difference by the number of requests in those six minutes. That gives the extra time each endpoint cost us because it was cold: Endpoint Cold Warm Requests in 6 min Extra seconds Campfire 246 ms 70 ms 6,490 1,140 Projects (JSON API) 84 ms 50 ms 22,077 771 Docs & Files 262 ms 177 ms 4,996 421 To-dos tool 205 ms 113 ms 4,018 371 To-dos (JSON API) 33 ms 16 ms 18,738 320 The pages already in the warm-up showed what to expect: the project page kept a 36 ms gap after a deploy, and the to-do page 10 ms. We’ve proposed adding these five requests, and expect them to add about five to seven seconds to the page step. The two JSON endpoints were a surprise. The warm-up’s page list had no API requests in it, so nothing on the API path had run before the first real request: not the API controllers, and not the Jbuilder templates rendering real records. Precompiling the views covers JSON templates too, but it isn’t a substitute for running the request. Results The warm-up is on for all 68 web hosts. With the first four pages it took 12 to 16 seconds per host: about 2 seconds to precompile the views, 7 to 9 for the 60 requests, and 3.5 for the schema. Deploys take that much longer per host, and we raised the deploy timeout from 30 to 60 seconds to cover it. In Amsterdam, at a peak-hour deploy: During deploy Before After Peak Puma backlog per host 250–2,238 requests 19–223 requests Peak kamal-proxy p99 about 10 s 2.4–4.8 s Peak CPU time per request 201–214 ms 88–132 ms Peak database time per request 56–69 ms 39–47 ms The same eight hosts at three deploys on 1 October, an hour apart, as the warm-up went from one host to four to all eight. The deploy in the middle, with four hosts warmed and four not, shows why every host needed the warm-up. Each warmed host recovered faster on its own: mean request duration peaked at 130 to 173 ms, against 203 to 311 ms on the hosts that weren’t warmed. But the backlog on all eight was about the same, because they were all waiting on the same database. Mean request duration on each host at the 07:21 UTC deploy. Blue hosts warmed up in the master before forking, orange hosts did not. Memory came down too. The workers now share compiled templates, YJIT code and the schema with the master instead of each building their own copy. On beta, the view precompiler alone took a busy worker’s private memory from 174–202 MB to 119–135 MB. Thirty minutes after the deploy, the web containers used about 39 GB less memory than the previous day’s containers at the same age and traffic. Amsterdam served most of our traffic at the times we tested. In Amsterdam, each new container used about 2 GB less just after traffic moved to it, which lowers the peak while the old and new containers overlap. Working with Claude Claude Code helped throughout. It combed through the per-worker backlogs and per-endpoint timings in Prometheus and Loki after each deploy, worked out the cold-versus-warm cost of each endpoint, and prepared the changes and the pull request descriptions with the benchmarks in them. We decided what to try, deployed it and read the results. If you do this Warm the master before it forks. Compile common code and templates and load their schema in the master, so workers inherit that work. With preload_app!, Rack::Builder#warmup runs before the workers start accepting traffic. Use the app’s real requests. Public pages, internal endpoints and synthetic queries warm the paths they run and nothing else. Signed-in requests to real records, frames included, run what production runs. Measure the cold penalty per endpoint. The difference between an endpoint’s cold and warm duration, times its request count after a deploy, ranks the pages worth adding. Ours weren’t the ones we’d have guessed, and two of them were JSON. Check what the warm-up leaves open. List the master’s file descriptors after a warm-up and account for every one before the fork. Two of ours needed changes. Set a time budget. A warm-up that runs long on one slow host fails the deploy on that host. Ours gives the page requests and schema loading a shared 20-second budget, checked before each page or model, puts the most valuable pages first, and logs what it skipped. Reforking from a warm worker, as Pitchfork does, solves the same problem continuously rather than once at boot, and it would warm paths no fixed list of pages covers. We may still get there: our branch brings Instacart’s mold_worker up to date with Puma’s main branch and fixes the bugs we found in it. But warming the master works with the Puma we already run, took a few days to implement, and substantially reduced the queues after deployment.

yesterday • 1 votes
📚 BoredReading

You seem to be enjoying this.

Join free to unlock everything.

Create free account

Already have an account? Sign in