More from the jsomers.net blog
Now that we have actually good AI, I have this vision of a form of computing that doesn’t involve me using a computer so much. Imagine you had the day’s emails to go through. It would be nice if the ones that required a simple decision could be dispatched with a few pen-strokes: I could write down a date that would work for that meeting; check a box to accept that invitation; etc. If an email required me to review a draft, I'd love to mark up a print version on my couch, sans screen, and have those notes scanned and sent off as if I'd done the whole thing on Google Docs. The point is not to give up on virtuality, but just to save the end user from having to interact with it. It's great to be able to send information to anyone in the world instantly; but let me do it without the glaring screen and the thousand distractions. Is such a thing even possible? I'm not sure, but I just now uploaded a first draft of this post, handwritten in a small notebook, into ChatGPT. Its transcription was nearly perfect. A first draft of this blog post Here's another example: when I really want to play with re-arrangements of a complicated outline, or when I want to collaborate with someone else on one, I find myself laying physical note cards on a table. Having real objects to work with allows for more flexibility than software. If the problem demands it, you can stack cards, draw on them, cut them, tape things to them, etc.—and none of these improvised ways of organizing has to be coded up in advance. Space turns out to be a good way to organize information. I gather that in the paper days if you had a big project you were working on, say a book, it would spill out into the room you were working in: chapter outlines pinned up on the walls, stacks of books in meaningful piles on the floors, folders with drafts and clippings. Certain sections of the work in progress would become associated in your mind with certain parts of the room. Chapter 3 would be over there. I rarely use paper or physical space this way because it lacks critical conveniences. A huge wall-mounted paper calendar is maybe the best way to plan and visualize a large coordinated effort. (In "making-of" documentaries you find that movie shoots are often planned this way.) Everyone can see it and point to it because it's at human-scale; you can express many dimensions of information simultaneously using shape, color, position, size, and any other physical attribute. But I have a hard time understanding how I'd keep such a calendar up to date. In practice, like almost everyone else, I use a virtual calendar that automatically adds events as I'm invited to them, syncs with other people's calendars, reconciles time zones effortlessly, and interoperates with other programs like email. Could we get the best of both worlds? In other words, shouldn't one goal of rapid technical advancement be some melding of the physical and virtual worlds such that I can sit quietly in an easy chair with pen and pad; or lay cards out on a table to organize my thoughts; or turn a room into the embodiment of a project; and yet have the same flexibility, portability, persistence, and remixability as in the digital versions of these things? I spend a lot of my day on screens. There are many problems with these things, articulated well by Bret Victor in the context of his Dynamicland project. Screens are small, antisocial, and they have a tiny vocabulary of affordances compared to physical objects. Plus they have the problem that they make it difficult to just use your calendar, todo list, or map—or even just respond to a friend's message—without encountering something else along the way, like a social network, short-form video, Slack, the news, or some other notification. To state the obvious: your phone is the best place to keep your calendar and inbox and todo list because you always have it with you, but of course that makes it ripe for other intruders. Bundling makes your phone indispensable, but also a menace. If nothing else I'd like it if operating systems and web browsers helped me be less distracted and frenetic, instead of encouraging exactly that multi-tasking freneticism. When I opened my phone or computer, it'd be nice if it was constrained to operate in a mode purpose-built for whatever task I intended to use it for. If I want to look something up, for instance, my phone should be a look-up machine (ie no texts, no apps, no ads, just a place for my question and the answer). If I want to compose a word of the day entry, I should launch straight into a browser with the tabs I regularly use for that, including the CRM, Webster's dictionary, and the OED; if I want to work on an article, my computer should assume the form of a typewriter, word processor, or McPhee mode note-processor depending on what stage I'm at. In each of these modes everything else should disappear, inaccessible. At least then you could mimic in software that thing you get from physical objects—which is that they are usually built to do one, and only one, thing well. My alarm clock, for instance, is just an alarm clock; and that's what I like about it! Not too long ago I spoke to a roboticist for an article I was writing, who worked with large autonomous earth-moving machines—e.g. a retrofitted excavator that could lift a boulder, scan its every edge and dimple, then model how it would settle amongst other boulders in a retaining wall before placing it there. He imagined a future in which such machines enabled a return to natural materials in the built world. He talked about old stone walls he’d seen in New England. Those walls, made of loose rocks found in situ, are lovely and sturdy, and adaptive—constantly rebuilt as farmers go about their work and notice areas that need patching up, adding stones they find lying around. But this is the very reason such walls aren't really built anymore. They're too labor-intensive. We live in a prefab world because the scarce thing now is not material or money but "the works and days of hands." I am moved by the idea that our future could feel less futuristic than pastoral. High tech could save us from high tech. We'd go back to the old interfaces without giving up the conveniences of the new ones. Read, write, communicate, create—and hardly ever see or touch a screen. I'm not sure that'll happen or be what people want. But shouldn't we be thinking of ways to use the new magic to spend less time tapping and clicking?
When I first started writing professionally, for the Atlantic’s website, I taught myself “reporting” with a simple self-made curriculum unfolding over six or seven articles. The first two pieces I wrote from my head, with reference to things I already knew or to books I’d read. For the third, I layered in firsthand experience, when […]
On podcasts it's pretty common to hear something like this: So Alexander Hamilton has just finished law school, and he's trying to make a name for himself. He's only been in New York a few years. So he takes on this case... The problem with the past tense ("Hamilton had just finished law school, and […]
The game of Five'Em was invented by two friends of mine, Ben Gross and Rich Berger, to combat Hold'Em fatigue. The rules are simple: You're dealt five hole cards instead of two, and after each round of community cards comes out (starting with the flop), you discard one of these extras. After the river is […]
More in programming
I listen to a lot of podcasts, and I like how they fit around other tasks. I press play, lock my phone, and put it down. I’m free to wash the dishes, fold the laundry, or shop for groceries. Unfortunately, more and more information is only published as a video. Technical talks, conference sessions, video essays – they don’t work in an audio-only podcast app. I could convert these videos to MP3 files, but that breaks down the moment a video isn’t pure spoken word. If a speaker says, “Look at this slide” or holds up a diagram, an audio-only file leaves me stranded. I don’t want to give up the podcast player I like, nor stare at a screen for an hour – but I do want the information in these videos. To solve this, I’m abusing my podcast player’s chapter support. This gives me the best of both worlds: I can listen to a video as audio-first, and glance at my lock screen if I need a moment of visual context. The idea: Chapters every few seconds MP3 files can have ID3 metadata, and ID3 metadata can include chapters. A chapter covers a particular time range, and it can have an associated title, description, and cover art. My podcast app of choice is Overcast, which can’t play videos, but it does have robust chapter support. I can jump between chapters, navigate a table of contents, and see per-chapter cover art. To get videos into Overcast, I’m creating MP3 files with a new chapter every few seconds, and the per-chapter cover art is a corresponding frame from the video. As I play the file, I get a slow, stop-motion-like rendition of the original video. If my phone is locked, I can glance at my lock screen and see the current frame in the Now Playing screen. Overcast is developed by Marco Arment, and I got this idea from Forecast, his app for adding chapters to podcasts. In particular, I was struck by its ability to create chapters that don’t display in the chapter list – ideal if I don’t want a table of contents with hundreds of entries. As I was developing my script, I compared my output to the output from Forecast to ensure I was creating the chapters correctly. The code: FFmpeg and Mutagen There are three steps in this process: Convert a video file to an MP3 Extract images from the video at a fixed interval Insert the images as hidden chapters in the MP3 file Let’s go through each in turn. 1. Convert a video file to an MP3 Converting a video file to an MP3 is a single FFmpeg command: ffmpeg -i video.mp4 audio.mp3 This is consistently the slowest step of the process, and I do wonder if I could use different settings or an alternative encoder to make it go faster – but it’s not slow enough to be worth further investigation. 2. Extract images from the video at a fixed interval Extracting images from a video needs a more complicated FFmpeg command: ffmpeg -i video.mp4 \ -vf 'fps=1/5,scale=iw*sar:ih,scale=min(iw\,945):min(ih\,945):force_original_aspect_ratio=decrease' \ thumbnail_%04d.jpg This extracts an image every 5 seconds, downscales any image larger than 945 pixels square (while preserving the original aspect ratio), and saves the results as sequentially numbered JPEG images (thumbnail_0001.png, thumbnail_0002.png, and so on). The key is the -vf flag, which defines two FFmpeg filters: The fps filter selects one frame every 5 seconds (fps=1/5). The first scale filter scales the width based on the sample aspect ratio (scale=iw*sar:ih). Without this filter, frames can be stretched and distorted. The second scale filter scales the input video, preserving the original aspect ratio (force_original_aspect_ratio=decrease), and ensuring the output images fit within 945×945px or the size of the input video, whichever is smaller. My limit is 945 pixels because that’s the largest size that cover art is shown on my iPhone. This filter still isn’t completely correct – it sometimes creates images from portrait videos that are smaller than I’m expecting – but it’s good enough. These are only thumbnails for glancing at, and if I want to change it later, I can always do the image resizing outside FFmpeg. 3. Insert the images as hidden chapters in the MP3 file Inserting the chapters into the MP3 file is more complicated. Although FFmpeg has basic support for ID3 metadata, as far as I know, it can’t insert chapters with per-chapter artwork. Instead, I’m going to reach for Python and the Mutagen library. Here’s the code to add a chapter to an MP3 file: from mutagen.id3 import APIC, CHAP, ID3, PictureType audio = ID3("audio.mp3") with open("thumbnail_0001.jpg", "rb") as f: img_data = f.read() image_frame = APIC(mime="image/jpeg", type=PictureType.OTHER, data=img_data) chapter_frame = CHAP( element_id="chp1", start_time=0, end_time=5 * 1000, sub_frames=[image_frame] ) audio.add(chapter_frame) audio.save() This creates a single chapter that lasts the first 5 seconds (0 to 5000 milliseconds), and the per-chapter cover art is thumbnail_0001.jpg. If we ran this in a loop, we could add images for every 5 second slice of the original video. This code is inserting two frames into the ID3 metadata: The CHAP (chapter) frame contains the timing information, and it can have subframes for metadata like title, chapter art, or associated URL. The APIC (attached picture) subframe contains information about a picture, which can either be a blob of image data or a URL to an image on the web. Normally, you’d also insert a CTOC frame which defines a table of contents, but I don’t want a TOC with hundreds of 5-second chapters, so I’m deliberately not doing this here. This is allowed by the ID3 spec – you’re not required to insert a CTOC frame if you’re using chapters, and you can have chapters that aren’t listed in your table of contents. To work out which frames I needed, I used Forecast to create some chapters by hand, and I inspected their frames. In particular, loading an MP3 and calling Mutagen’s pprint() method shows a human-readable list of frames, and then I could drill into the individual fields: from mutagen.id3 import ID3 audio = ID3("audio.mp3") print(audio.pprint()) I wrapped all this code in a project called glancecast, which allows you to convert a video file with a single command, with optional flags to set the frame length and chapter art size: $ python3 glancecast.py interesting_talk.mp4 interesting_talk.mp3 The process takes a minute or so to complete, most of which is spent transcoding the video file to MP3. The resulting MP3s are usually 40 to 50 MB in size, which is very reasonable. The outcome: How it looks in practice Here’s what one of these “glanceable” podcasts looks like in Overcast and on my lock screen: Maggie Appleton presented this talk over two years ago and it’s been on my “talks to watch” list ever since. Once I put it in Overcast? I listened to it in less than a day. It’s not a lot of extra information, but enough that I can quickly glance down and get the gist of what a speaker is saying. Both views update with a new frame every few seconds, or I can put my phone in my pocket and ignore the screen. I’ve used this approach for half a dozen videos so far, and I’m happy with the results. I expect to keep using it, because I have a long queue of videos I’ve been meaning to watch. If you’d like to try this, check out glancecast for the full code and instructions. [If the formatting of this post looks odd in your feed reader, visit the original article]
Andrew Baker, the current Group CIO at Capitec Bank wrote an interesting piece on AI and open source, and how these tools that generate code according to one’s specification may replace the general reliance on open source implementations done by contributors around the world. I’d really recommend reading it. I have great admiration and respectContinue reading "AI Isn’t Replacing Open Source"
A framework for thinking about when AI involvement is additive or a violation
Why we need richer, thicker interfaces and better boundary objects for collaborative planning with agents
A look at 10 foundational pillars that enable agents to operate more competently and more efficiently in any codebase.