More from Lee A Johnson
If you haven’t been paying attention, which is likely given the small corner of the web I’m going to cover, you may have missed that BINGOS just hit 247 consecutive monthly uploads to CPAN1. This is noteworthy as it surpasses DROLSKY’s record set a while back. That’s an upload to CPAN, once a month, for over twenty years. Once a month minimum I should say, as many of the authors on the board are uploading more frequently than that, some even daily. My own monthly upload streak is a mere 149, or twelve and a half years, and really that’s only because I pay attention to the board. Once a month I will check it to remind myself to dig through my CPAN distributions for bugs, issues, and refactoring opportunities. This keeps their internals fresh in my mind for when reports from users or contributors come in. The other authors on the board maintain an order of magnitude more distributions than I do, so are more likely to get a steady stream of reports from users. A stream that is turning into a torrent given the current place we find ourselves in the world of open source software. And that’s where I start to worry a little. Here in the Perl world we’ve been quietly getting on with it since the first shouts of “nobody uses Perl anymore” going back almost as long as BINGOS’ streak. Quietly maintaining critical systems, billion dollar businesses, backwards compatibility, security patches, the occasional new feature, and yearly releases of the interpreter. But that torrent of reports and issues is going to become a flood, and it’s not just the Perl ecosystem that is at risk. CPAN being Perl’s software repository. The upload stats can be found at cpan.io ↩
Five short essays on where we are currently at. A World of First Drafts Dreams and the Uncanny Valley Fiddlers in the Room The Last Hike Semantic Satiation of an Acronym A World of First Drafts I recently picked up a copy of “An Evening With Windham Hill”, a collection of early 1980s live performances by some of Windham Hill’s popular-at-the-time acoustic guitar players. I primarily bought the LP for the recording of “Turning: Turning Back” by Alex deGrassi, as it was the only way to get this on vinyl; the 1992 retrospective that also contains the track was never (officially) released on anything but CD. The album contains a more interesting track on it, that being the first (?) performance of a Michael Hedges composition. Introducing the performance Hedges says “This is a new piece for [that] started out for guitar and then, er, all of a sudden it needed piano and about a week ago it needed bass so… We need to play it tonight. It’s dedicated to Steve Reich, and it’s called Spare Change.” The performance starts out very Hedges like, with his (now) distinct playing style and tapping on his instrument, but then after about forty seconds the piano comes in and Hedges influence seems to be diluted. He pulls it back but seems to be fighting with the piano, and when the bass solo arrives at three minutes Hedges is then completely lost to Manring. The three instruments then battle for the remaining two minutes of the composition leaving us at an ending that feels unresolved. So unresolved it takes the audience several seconds to realise the performance is over. The composition and its performance was very much a first draft. Some interesting ideas in places, but ultimately lacking cohesion, unsatisfying, and even forgettable. Hedges’ voice (that of his guitar) is lost amongst the parts that aren’t his. When Hedges released the final version two years later, on the album “Aerial Boundaries”, he knew the piece needed work so he made some major changes. The first was to drop the key two semitones lower. I guess he removed the capo from his guitar. The second change was more substantial: Hedges decided to replace the piano and bass parts with his own guitar, spending over one hundred hours recording sounds, looping them, playing them backwards, splicing tapes, pulling them apart and sticking them together, experimenting, all to get the textures that fit. Over one hundred hours in the studio in 1983, and likely many hours after those first drafts in 1982, to create a five minute long piece of music. Hedges could have just released the original arrangement, but he knew it was mediocre so continued to refine it. The version realised on “Aerial Boundaries”, and the liner notes explicitly use the verb realised rather than recorded, is probably the most striking work in Hedges’ entire discography. Because there are other remarkable recordings on the album, especially the title track, and due to the near impossibility of performing the track live, “Spare Change” is often ignored. That doesn’t mean Hedges’ time and effort to refine it was a waste, as the track still stands out today. Along with the first live recording, the final version is a permanent record of an idea elevated to something interesting, influential, even epochal. Fingerstyle guitar was going through some major changes in the early eighties, and Hedges was one of several important composers and performers at the time. Had he just sat back, been satisfied with the first draft, not torn it apart to rebuild, it would have been forgotten. I wouldn’t be writing about it today, and it’s possible it may have been discarded and failed to make the final cut for the album. Perhaps the composition would have been dredged up some time in the future, when an artist’s career reaches that inevitable scraping of the barrel stage. Alas, Hedges was killed in an automobile accident in 1997, at the age of 43. The obvious place I’m going with all this is that I increasingly feel like we are moving into a world of first drafts. One where ideas are formed but then delegated to something else to refine them. The problem with this approach is there is nothing new to pull from the tombola, you still have a first draft at the end of that process and haven’t gone anywhere compelling. I’ve even had someone defend their approach with “the ideas are all mine”. But ideas are easy, execution and refinement are hard and that’s where your individuality comes through. Maybe your ideas are boring, and your arguments are weak? I’d still like to read them in your style. Writers have built entire careers on this. The reason I read your blog, or listen to your music, or watch your videos, or look at your photos is because I’m interested in how you see and react to and interpret the world. If your voice and idiom is lost to the machine’s then you are no longer interesting and I’m no longer interested. Dreams and the Uncanny Valley A few years after moving to the mountains I had an idea for a photo project that would involve shooting images of well known vistas and then subtly replacing some of the mountain ranges with different peaks. Or removing parts entirely. Or just messing with the horizon in some way. I played with this a little when I printed some postcards of the Matterhorn, flipping the mountain on its horizontal axis to result in a mirror image. I assumed people would notice, given that concrete chocolate mountain is in the top five of recognisable peaks. Nobody did, at least not for five years. Or nobody said anything as they weren’t quite sure if what they were looking at was wrong, a literal uncanny valley moment. Sometime later I had a dream, or a nightmare since I tend not to remember my dreams, in which everything was subtly wrong. Like something out of a half forgotten Philip K. Dick story or tired science fiction trope. All my friends faces were a little different, but I was the only one that could see this. All the music I knew was off in some way, a different tempo or transposed, or key lyrics were swapped with synonyms. Nobody else noticed and sang along like it had always been this way. They were oblivious to the changes. The books and films all had slightly different titles and plots, and the actors cast were not the same. The walk I took to work was a diversion, my keys were interchanged, the office was rearranged, my keyboard layout was swapped. At lunch the microwave controls were on the left, the food tasted weird. None of my colleagues had any qualms. Finally I realised that if that was the case externally, at the macroscopic level, then was it also the case at an atomic level? I rushed home to find my Roche Biochemical Pathways poster, unfolded it (the wrong way), and stared at the molecules trying to recall the chirality of the amino acids, nucleic acids, and sugars. But I couldn’t remember, it was too long since I had done any biochemistry. And then I woke up. Fiddlers in the Room There’s a certain type of software engineer, I’ve met several times in my ongoing career, that I like to term a “fiddler”. When the fiddler is tasked with solving a problem, instead of thinking “I will solve this problem” they instead think “I will solve this particular class of problem”. The non-fiddler will pick an existing solution, and if there isn’t one that fits they will code for the explicit problem at hand. The fiddler will build an entire system to cater for a hypothetical future and other people’s unknowns. The result is something that is not actually generic enough to solve the particular class of problem, being too brittle, and is too broad to solve the original specific problem in a maintainable way. An excess baggage of logic and abstractions means cognitive overload. Having to think about the generic class of problem when trying to maintain for the specific problem is always, well, a problem. Don’t get me wrong. Fiddlers are important, and have contributed to many critical parts of many ecosystems, and some fiddlers are very good at what they do; however, they are the minority and, more often than not, a one hit wonder. A key tell of some fiddlers is that they like to use the latest tools to facilitate their fiddling, and this is the most dangerous fiddler of all. New tools are more likely to change quickly, and churn is the enemy in software engineering. Now it seems that everyone can be a fiddler, so my term has become a tautology or changed to mean something else. Fiddling has been flipped on its head. Now it seems fiddlers can trivially generate for the explicit problem at hand, pulled from a vast corpus of existing classes of problems, likely previously generated by the previous generation of fiddlers. Those not-quite-generic-enough-and-a-little-too-brittle solutions. And if the fiddling is not enough? Have an entire orchestra. Keep screaming “COMPUTER, DO SOMETHING!” If and when it all goes wrong? Yes: “Switching over to manual control… good luck!” and of course: “This is the world’s smallest violin, and it’s playing a sad song for you.” The Last Hike The last hike took place on the 25th May 2026. Temperature was 20ºC, humidity 53%. Distance 13.00km, altitude change +/- 566m. 735kcal was expended in energy, with an average heart rate of 107bpm (max 153, min 70). Duration was 3h21m. The computer decided the effort was “moderate”. The last hike was somewhat more successful than the last bike ride, which took place on the 19th May 2026. Temperature was 10ºC, humidity 66%. Distance 11.56km, altitude change +/- 258m. 356kcal was expended in energy, with an average heart rate of 138bpm (max 161, min 69). Duration was 41:31. Max speed was 51km/h when a catastrophic puncture of the back tire took place, sending the rider veering off course into a 1m wide grass verge between the asphalt and a barbed wire fence. The rider was thrown over the handlebars of the bike, onto their right shoulder. Some moderate skin grazing occurred, but a six hour visit to the hospital was required to confirm no bones were broken. The computer, knowing nothing of this, decided the effort was “moderate”. I decided to stop tracking all this stuff because I don’t know what advantages it begets other than the obvious: exercise is beneficial. The disadvantages are not worth it, trying to beat personal bests, going faster, worrying about regressions, contributing to the world’s largest and longest ongoing clinical trial? My father used to run a lot in his youth. He would take me in the pram when he ran, and he ran so much he wore through two sets of wheels. There’s little evidence of that now, maybe a few photos in our attic, and his arthritic knees. It’s possible that his running has nothing to do with his knees, I know. None of this was tracked, so who knows? Even if it was tracked we wouldn’t know anyway. All of this data works at a population level, individually you need to ask yourself if you really need to know the minutia. Exercise is beneficial until the bones hit the asphalt. Even then, the net benefit is worth it. Just move a few times a week for more than a few minutes and you’ll feel better. Semantic Satiation of an Acronym There has to be an end state to all this, eventually? Surely? A point at which it just becomes the norm, at which it is everyday, normal, mundane? A point at which we can just all get on with our lives and benefit from the good, and be protected from the bad. A point where the thing producing so much noise is reduced to a background hum that can be ignored. I can't seem to escape this in any place. Technical discussions normally reserved for office work hours spill out into social settings. The pub, the dinner table, with strangers on a plane, in a queue, at a gig. I hear horror stories of wannabee software engineers making all the mistakes newbies make, but with none of the guardrails or mentoring to steer them in the safe direction. I end up talking shop with people I would never want to be my colleagues. I'm not trying to gatekeep here, I have no kingdom to protect, rather I am worried that the bar being so low will cause many people to trip over it. I do find this stuff useful, in the very specific areas it is currently limited to. It will become more useful in other areas over time and hopefully less dangerous, but that's an unknown for now. I didn't think semantic satiation was possible with an acronym. Can I please just go one day without hearing about this. Just one.
Apparently the bottom is falling out of the art market, at least for those in the high-end segment1. That market of investors, speculators, and dealers who are working with mind-boggling prices. This isn’t surprising. The buyers in this market aren’t interested in art, they’re interested in the financial returns, and the last few years have seen more compelling investments in other areas. Cryptocurrencies, NFTs, and now AI. Why faff about with tangibles, which need shipping, storage2, maintenance, and insurance? Why bother with art that needs decades to see a return, when the alternatives can be turned around in months with a few clicks? At least when the bottom falls out of the art market you will still have something to hang on your wall, eh? Anyway, it’s three years since I bought a printer. I thought I would provide an update on running costs, print sales, and profit. I’m happy to be completely transparent about all this as I know it can be difficult to figure this out if you are thinking of creating your own printing business. Culan et Croix des Chaux, December 2024 Again, the reason why I do this from November to November is that the ski season runs December to April, and usually the most sales are during that period. Also I bought the printer in November 2022 so I’m just carrying on that way. You can find details from the previous year here. All the photos here are work shot in the last twelve months, I’ll cover this in a bit more detail below in the “Thoughts” section. Year Three Investment I needed to purchase more paper and ink this year, as expected. I also decided to buy a stand for the printer, as it was available at a significant discount. This gives me more freedom in my studio as I can now move the printer around easily. It also will allow me to add a second roll feeder in the future, should I decide I need that. Investment figures are included in the costs below. Income Income is made up of sales through a local gallery that I am part of, along with web sales: 11,409.- CHF - sales through the gallery 642.- CHF - sales through web shop = 12,051.- CHF total income Total income is up by about 20% compared to 2024 on both gallery sales and web shop sales. Costs I include the loan repayment here as that is an ongoing cost. I had to purchase consumables (paper and ink), at a total of 1,370.- CHF however that cost isn’t included in the total here as it will be amortised in future years under the “print consumables” category when they are used to make prints. First Tracks, January 2025 Rental cost of space in the gallery was reduced due to having overpayments of rent and my loan to the gallery, from the previous year, paid back to me. -4,058.- CHF - loan repayment -1,500.- CHF - rental cost of space in the gallery -672.- CHF - print consumables (paper, ink) -130.- CHF - investment (printer stand) -3,408.- CHF - framing/mounting/postage -2,323.- CHF - commission to the gallery -300.- CHF - web shop subscription -62.- CHF - stripe / paypal fees 0.- CHF - advertising (none this year) = -12,453.- CHF total costs Total costs are down by about 15% compared to 2024. Total Profit (Loss) 12,051.- CHF total income -12,453.- CHF total costs = -402.- CHF total loss Thoughts Holy crap, I only made a 402.- CHF loss this year. What happened? A combination of debt owed to me by the gallery and better sales meant I came close to breaking even. The original investment loan also accounted for 4,058.- CHF in costs, but this loan is now fully repaid and will not factor into future costs. I was lucky this year in that one of the local schools made a bulk order of small framed prints, which contributed to the boost in sales. If I were to net off the cost of the loan, but also this large order (since this is a rare occurrence), I would have a profit of about 2,000.- CHF. So still a good year, relatively speaking. However, that’s about 175.- CHF per month profit. You can’t live on this, and it’s an absolute grind if you want to start increasing that. Again, the reality remains that a physical space is absolutely essential for getting the work out there and sold - just look at the figures, approx twenty times the amount of income than from the web sales. Verbier the day after after record snowfall, April 2025 The Content Treadmill This year’s “trip to a ski resort in another country” was Chamonix, which I hadn’t been to in over a decade. I hired a guide for some off piste ventures, and we had a great time exploring the mountains. The snow was reasonable at the start of the week, but a bit naff at the end, but I did manage to shoot a couple of photos that have been added to the work for sale. We already have a trip planned for next year, and are considering the year after. This should result in more work. I should note that the work is a byproduct of my first interest - snowboarding. Really I’m not looking for locations or trips with a thought to shoot photos, they are purely an afterthought. If the places I go are suitable, and the conditions allow, then a photo might result. I have no interest in creating “content”, as that’s an absolute grind as well. I might shoot a couple of good photos a year, really. The work I shoot is very specific to a place, as I’ve talked about previously. If I wanted it to sell I would be looking to display it in galleries present in those locations. Sunset Over Le Col de la Croix, October 2025 And if I did want to boost the income from all of this I would have to jump on that content treadmill. Be thinking all the time about where to go next, what to shoot, and how to sell it. Be pushing out work all the time. That means producing mediocre work, that you’re not happy with, and that is a drag. Going out with the intent of creating content? Where’s the fun in that? The Storm Hits the Art Market ↩ Geneva Freeport ↩
I̸ ̸s̸t̸a̷r̶t̵e̷d̵ ̴w̵a̷t̶c̸h̶i̴n̸g̵ ̴t̵h̴e̶ ̸P̶y̴t̴h̸o̶n̴ ̵d̶o̵c̷u̵m̷e̴n̴t̷a̸r̶y̴,̶ ̸[̶r̷e̶c̶e̴n̷t̵l̷y̷ ̷a̶v̵a̸i̷l̸a̷b̷l̴e̷ ̸o̷n̸ ̵Y̴o̷u̶T̴u̴b̶e̸]̴(̵h̸t̵t̷p̴s̸:̴/̸/̴w̸w̴w̶.̵y̶o̶u̵t̵u̴b̶e̸.̴c̷o̷m̷/̵w̵a̴t̷c̴h̵?̷v̴=̵G̵f̶H̶4̵Q̷L̸4̸V̸q̴J̵0̴)̷,̴ ̴b̴u̵t̷ ̴I̴ ̸c̷o̸u̸l̶d̴n̸’̴t̷ ̷g̷e̷t̵ ̴i̸n̸t̶o̵ ̷i̷t̸ ̴b̵e̶c̷a̷u̴s̷e̷ ̵t̴h̴e̶r̵e̴ ̸w̷a̶s̵ ̶s̵o̵m̶e̶t̷h̵i̸n̴g̵ ̵i̶n̵ ̴t̵h̸e̵ ̴b̶a̶c̵k̵g̵r̵o̸u̷n̶d̷ ̶c̷a̷u̶s̷i̴n̵g̴ ̷a̶ ̷c̸o̴n̷s̸t̷a̸n̵t̸ ̴d̷i̷s̸t̷r̴a̵c̴t̶i̶o̵n̷.̴ Oh wait, sorry - let me try that again. I started watching the Python documentary, recently available on YouTube, but I couldn’t get into it because there was something in the background causing a constant distraction. Music. It was music. It’s a shame as I’m interested in watching the documentary, but an artistic decision has made it difficult for me to do so. If you can’t relate then imagine if I had written this entire blog post with Zalgo text, like the opening paragraph. That’s what the experience feels like to me. Because I really love music. So if I hear it, I listen to it. This is really hard for me to switch off as music commands my attention. It’s one of the reasons I never have music on while doing something else. If two things are fighting for my attention, one of them being music, the music always wins. Always. I find it a contradiction to edit together a bunch of interviews and then add background music. The interviewee’s words are the music, they should stand alone and not require background noise. Doing otherwise is to treat the audience as if they are not intelligent enough to understand the nuances of the interview. Oh minor key, me sad now… Fuck off, please. It’s likely I’m feeling particularly attuned to this as we watched a couple of films over the weekend. The first, “Highest 2 Lowest”, being a sub-par film reduced to mediocrity by feeling like one long iPhone advertisement, featuring copious unnecessary background music in moments of character dialogue. I pointed this out to my wife about 10mins into the film, and then she couldn’t ignore it. Contrast this with Les triplettes de Belleville, a largely dialogue free film, which I hadn’t seen since its release twenty years ago. A film which knows when to hold back on the music and when it is necessary. One that also plays with the musical themes - Glenn Gould’s piano playing, repeated in an improvisation using a bicycle wheel, then interpreted on jazz piano a little while later. Backgrounds are important, but they can be massively distracting. Skipping through the timeline of the documentary it seems that it is at least 80% talking heads, in other words something that’s not visually compelling, so I’ll likely just read the transcript. The problem in doing that is also the loss of nuance. Nuance is important, so that’s unfortunate, but the artistic decision to stomp all over the interviews with background music already destroyed any of it that existed in this documentary and replaced it with something entirely different. I’ll talk about this a bit in the next blog post.
More in programming
I owe a lot of my professional identity and success to CSS-Tricks. CSS-Tricks repeatedly gave me the opportunity to write for them. In doing so, they helped to both socialize and normalize accessibility as a mainstream frontend concern. I’m deeply thankful to them for this. The team was also a joy to work with, notably Geoff Graham. He’s a mensch, and one of the nicest people you can interact with in the frontend web space. If you have not been following the news about the site, Kevin Powell has a good video about the whole situation: Content skipped. I’m not speaking on behalf of Geoff, Chris, or others involved with running the current version of CSS-Tricks. I’ve got skin in the game as an author. This is my personal opinion, born of my feelings and beliefs. I think a lot of the web’s infrastructure should be co-ops, and CSS-Tricks is knowledge infrastructure. To that point, I should also point out that the website covers far more than just CSS. The corporate model of ownership can be a risk. If infrastructure is not part of a corporation’s core strategy, it is not a priority. As Kevin’s video touched on, it seems like promotion via owning the frontend content space isn’t part of Digital Ocean’s strategy anymore. It is not that CSS-Tricks does not have value. It is that Digital Ocean cannot see it. It is deeply, tragically ironic to me that Digital Ocean allowed this to transpire. This is because I know for a fact that the techniques and philosophies shared by CSS-Trick authors helped to shape iterations of their product’s UI. Some may be quick to point out that this knowledge now—illegally—exists inside of LLM training data, so the risk of the website going away is mitigated. To this, know that we should be striving to keep resources like CSS-Tricks going. Human creativity is the force that creates new techniques, strategies, and technologies. The web will calcify without voices sharing what they know, forever locking us into endless permutations of a fixed point in time. Unlike corporations, co-ops don’t have to be motivated by profit. By not needing to prioritize growth at all costs it means co-ops can instead prioritize and incentivise things like preservation and cultivation. It is also a successful model of operation, one that even already exists, and flourishes in the tech space. Collective ownership can also serve as checks and balances for, and protection against hierarchical decision-making. I only need to point to the chaotic and aberrant decisions many CEOs in the technology space have been making as of late to demonstrate the value of this approach. Paddy Srinivasan, if you somehow wind up reading this: Save some face and take a big swing. Give CSS-Tricks back to the people who love it.
How can something that “just works” be so annoying? situation We live in Cambridge off a little road down a drive in shared ownership between us and our neighbouring houses. All the utilities are buried under this drive, including the phone line. anticipation Over the last few years we have been canvassed repeatedly by CityFibre saying that they can deliver fibre all way to our house. I saw them digging trenches and leaving tails of purple fibre cladding along nearby roads, ready to hook up all the houses. I thought they would need to do something similar to deliver fibre to us. So when they turned up and knocked on our door, I talked to their salesbods and walked them up and down the drive and pointed out where the existing BT line goes. Then they gave up trying to sell to us. This happened about three times. disaffection We were not eager enough for an upgrade to deal with these impediments. notification A few months ago we were told that CityFibre would soon come and do the upgrade, since there’s a nationwide deadline for turning off the copper phone network at the end of the year. We expected that this would force them to actually plan some digging works, so we talked to our neighbours about it. We were all ready for some huge faff to follow the next visit by the CityFibre bods. installation CityFibre turned up on the promised morning bright and early. To our enormous surprise, a brown fibre housing was already poking out of the ground next to our copper phone line. It had been fed through 50 metres of 5cm duct without us being aware they were even working on the street. Within a couple of hours, the technicians had drilled through our wall, installed the ONT, blown fibre through the unexpected pipe, plugged in the CPE (superficially identical to the old one), and left telling us to anticipate that it might not work properly until tomorrow. activation Around lunch time, the copper phone line stopped working completely. Some faff ensued, switching all our devices over to the new WiFi network. For a while we thought this was the death of our land line, but in the course of debugging other issues, I realised that the router has a built-in VoIP adapter (I don’t think we were told it has a built-in VoIP adapter) so I plugged the phone in and it Just Worked: they had ported our phone number across and everything. Flawless. I was seriously impressed. rumination It has been a few weeks since the switchover, and apart from a couple of horrible Clown-afflicted IoT devices, it has been fairly smooth. What prompted me to write this up was realising that we delayed this upgrade for years because the sales people were not given enough technical information about how the installation process works: the fact that houses typically have a 5cm duct containing the copper lines (probably standard for the last 40 years) and the fact that fibre can be shoved through a few tens of metres without difficulty. And worse, the sales people didn’t have an esclation path for difficult cases: they just gave up instead. From a technical point of view, the installation was impeccable. (I guess the loose 24 hour window for the cutover time was because OpenReach and CityFibre don’t have tight requirements on ISP reconfiguration schedules.) From the sales point of view, it was crap. Maybe it would have gone faster if we offered to switch early without asking if the drive would be a problem? But I guess the difference between “yes!” and “yes, but will this be a problem?” is too much to expect from a minimum-wage door-to-door salesbod whose employer didn’t give them enough information or any escalation path.
I listen to a lot of podcasts, and I like how they fit around other tasks. I press play, lock my phone, and put it down. I’m free to wash the dishes, fold the laundry, or shop for groceries. Unfortunately, more and more information is only published as a video. Technical talks, conference sessions, video essays – they don’t work in an audio-only podcast app. I could convert these videos to MP3 files, but that breaks down the moment a video isn’t pure spoken word. If a speaker says, “Look at this slide” or holds up a diagram, an audio-only file leaves me stranded. I don’t want to give up the podcast player I like, nor stare at a screen for an hour – but I do want the information in these videos. To solve this, I’m abusing my podcast player’s chapter support. This gives me the best of both worlds: I can listen to a video as audio-first, and glance at my lock screen if I need a moment of visual context. The idea: Chapters every few seconds MP3 files can have ID3 metadata, and ID3 metadata can include chapters. A chapter covers a particular time range, and it can have an associated title, description, and cover art. My podcast app of choice is Overcast, which can’t play videos, but it does have robust chapter support. I can jump between chapters, navigate a table of contents, and see per-chapter cover art. To get videos into Overcast, I’m creating MP3 files with a new chapter every few seconds, and the per-chapter cover art is a corresponding frame from the video. As I play the file, I get a slow, stop-motion-like rendition of the original video. If my phone is locked, I can glance at my lock screen and see the current frame in the Now Playing screen. Overcast is developed by Marco Arment, and I got this idea from Forecast, his app for adding chapters to podcasts. In particular, I was struck by its ability to create chapters that don’t display in the chapter list – ideal if I don’t want a table of contents with hundreds of entries. As I was developing my script, I compared my output to the output from Forecast to ensure I was creating the chapters correctly. The code: FFmpeg and Mutagen There are three steps in this process: Convert a video file to an MP3 Extract images from the video at a fixed interval Insert the images as hidden chapters in the MP3 file Let’s go through each in turn. 1. Convert a video file to an MP3 Converting a video file to an MP3 is a single FFmpeg command: ffmpeg -i video.mp4 audio.mp3 This is consistently the slowest step of the process, and I do wonder if I could use different settings or an alternative encoder to make it go faster – but it’s not slow enough to be worth further investigation. 2. Extract images from the video at a fixed interval Extracting images from a video needs a more complicated FFmpeg command: ffmpeg -i video.mp4 \ -vf 'fps=1/5,scale=iw*sar:ih,scale=min(iw\,945):min(ih\,945):force_original_aspect_ratio=decrease' \ thumbnail_%04d.jpg This extracts an image every 5 seconds, downscales any image larger than 945 pixels square (while preserving the original aspect ratio), and saves the results as sequentially numbered JPEG images (thumbnail_0001.png, thumbnail_0002.png, and so on). The key is the -vf flag, which defines two FFmpeg filters: The fps filter selects one frame every 5 seconds (fps=1/5). The first scale filter scales the width based on the sample aspect ratio (scale=iw*sar:ih). Without this filter, frames can be stretched and distorted. The second scale filter scales the input video, preserving the original aspect ratio (force_original_aspect_ratio=decrease), and ensuring the output images fit within 945×945px or the size of the input video, whichever is smaller. My limit is 945 pixels because that’s the largest size that cover art is shown on my iPhone. This filter still isn’t completely correct – it sometimes creates images from portrait videos that are smaller than I’m expecting – but it’s good enough. These are only thumbnails for glancing at, and if I want to change it later, I can always do the image resizing outside FFmpeg. 3. Insert the images as hidden chapters in the MP3 file Inserting the chapters into the MP3 file is more complicated. Although FFmpeg has basic support for ID3 metadata, as far as I know, it can’t insert chapters with per-chapter artwork. Instead, I’m going to reach for Python and the Mutagen library. Here’s the code to add a chapter to an MP3 file: from mutagen.id3 import APIC, CHAP, ID3, PictureType audio = ID3("audio.mp3") with open("thumbnail_0001.jpg", "rb") as f: img_data = f.read() image_frame = APIC(mime="image/jpeg", type=PictureType.OTHER, data=img_data) chapter_frame = CHAP( element_id="chp1", start_time=0, end_time=5 * 1000, sub_frames=[image_frame] ) audio.add(chapter_frame) audio.save() This creates a single chapter that lasts the first 5 seconds (0 to 5000 milliseconds), and the per-chapter cover art is thumbnail_0001.jpg. If we ran this in a loop, we could add images for every 5 second slice of the original video. This code is inserting two frames into the ID3 metadata: The CHAP (chapter) frame contains the timing information, and it can have subframes for metadata like title, chapter art, or associated URL. The APIC (attached picture) subframe contains information about a picture, which can either be a blob of image data or a URL to an image on the web. Normally, you’d also insert a CTOC frame which defines a table of contents, but I don’t want a TOC with hundreds of 5-second chapters, so I’m deliberately not doing this here. This is allowed by the ID3 spec – you’re not required to insert a CTOC frame if you’re using chapters, and you can have chapters that aren’t listed in your table of contents. To work out which frames I needed, I used Forecast to create some chapters by hand, and I inspected their frames. In particular, loading an MP3 and calling Mutagen’s pprint() method shows a human-readable list of frames, and then I could drill into the individual fields: from mutagen.id3 import ID3 audio = ID3("audio.mp3") print(audio.pprint()) I wrapped all this code in a project called glancecast, which allows you to convert a video file with a single command, with optional flags to set the frame length and chapter art size: $ python3 glancecast.py interesting_talk.mp4 interesting_talk.mp3 The process takes a minute or so to complete, most of which is spent transcoding the video file to MP3. The resulting MP3s are usually 40 to 50 MB in size, which is very reasonable. The outcome: How it looks in practice Here’s what one of these “glanceable” podcasts looks like in Overcast and on my lock screen: Maggie Appleton presented this talk over two years ago and it’s been on my “talks to watch” list ever since. Once I put it in Overcast? I listened to it in less than a day. It’s not a lot of extra information, but enough that I can quickly glance down and get the gist of what a speaker is saying. Both views update with a new frame every few seconds, or I can put my phone in my pocket and ignore the screen. I’ve used this approach for half a dozen videos so far, and I’m happy with the results. I expect to keep using it, because I have a long queue of videos I’ve been meaning to watch. If you’d like to try this, check out glancecast for the full code and instructions. [If the formatting of this post looks odd in your feed reader, visit the original article]
Andrew Baker, the current Group CIO at Capitec Bank wrote an interesting piece on AI and open source, and how these tools that generate code according to one’s specification may replace the general reliance on open source implementations done by contributors around the world. I’d really recommend reading it. I have great admiration and respectContinue reading "AI Isn’t Replacing Open Source"