Full Width [alt+shift+f] Shortcuts [alt+shift+k]
Sign Up [alt+shift+s] Log In [alt+shift+l]
1

Fast Linear Gradient Fills with OffscreenCanvas

from SOS [alt+shift+b] in programming

TLDR: Demo is here, Code is here, App is here In Kidz Fun Art, the web app for tablets I’ve built for my kids and hopefully yours, I recently added a nice little feature where you can fill in any area with a linear gradient. It’s highly responsive to the user moving their pen/finger, and can quickly let them change the direction and spacing of the gradient at something like 60 frames per second. This post describes the technical details of how it is achieved. To see it in action, either try it out yourself at https://kidzfun.art or on the minimal demo page, or watch the video below Step 1 The user clicks inside some shape that they want to fill with a gradient, in the example below it’s the green diamond. At this point, the app sends a few things to the Worker Thread: An OffscreenCanvas. A Canvas is a 2D drawing element in a web browser. For performance reasons, you can transfer control of a Canvas to a Worker thread, so that as the user is moving their pen around in the single threaded user thread, any paint operations can happen simultaneously without blocking the user’s actions. A copy of the pixel data from the user’s Canvas, containing the green diamond you see above Some data about the point that the user clicked, and the colours to use in the gradient Step 2 Now the Worker thread performs a simple flood fill with a solid colour, starting from the point that the user clicked. The resulting pixels are set to black in the OffscreenCanvas (the actual colour doesn’t matter, it just has to be opaque). While doing this we record the bounding box around the pixels for use later. Step 3 Since we now know the bounding box for the filled in pixels, we can fill it with a linear gradient using the context.createLinearGradient function, as below const colours = ["#FF0000", "#FFFFFF"]; const gradient = context.createLinearGradient(x1, y1, x2, y2); colours.forEach((color, index) => { const stopPosition = index / (colours.length - 1); ...
27th May 2024

Stay updated

Get a weekly newsletter with the top 5 articles worth reading every week.

More from SOS

Clean your Mac & Windows/Linux PC with Disk Space

TLDR: Download the best disk cleanup app in existence from diskspace.io. A fun experiment in writing the same app three times (well kind of 4 times) with AI, and being as optimised and OS native as possible. My Mac recently filled its 1TB drive and my quest to find out where the space had gone was very frustrating. The existing tools for tracking down disk space were slow, clunky and it was impossible to use them until their 30 minute crawl of my hard drive completed. I decided to fix this by building a new app called Disk Space, which takes the great features of my favourite old time app Disk Inventory X, makes it blazingly fast and adds much improved file/folder deletion aesthetics so that you can clean up safely and quickly, as well as highlighting recently created files so you can find what changed more quickly. Let’s get nerdy The three versions, Mac, Windows and Linux, are all mostly independent, with just a little shared C++ code. My goal was to make each app as small, fast, as native to its environment as possible. This meant not using any of the more common cross platform libraries, and leaning on Claude to do the work. MacOS The MacOS version of Disk Space is written fully in Swift, with no dependencies on other libraries. This was the first version I built. It detects how many CPU cores your machine has and optimises itself to maximise the throughput. The size of the installed app is just under 1MB, with about 300KB of that being multiple sizes of the icon, so the app itself is just over 600KB in size. Not bad! Releasing desktop apps for the Mac is actually not too bad an experience. You can choose to put it in the App Store, but then it can be a pain to release a new version. Apple offer a free notarisation service that signs the built app with your Apple Developer credentials, and it works well. For this reason, you can simply download the MacOS version of Disk Space from the site. Windows The Windows version is a direct port of the Swift code to C++. Similar to the MacOS version it optimises its operation based on the number of cores, but it also checks if the drive being scanned is a spinning disk or a solid state drive. If it’s a spinning disk there’s no point running many threads against it, it’s physically incapable of responding, so it caps the number of threads. Since it it just pure C++ with no dependencies pulled in, the installer is just about 600KB, pretty cool. Releasing apps on Windows these days generally means you are forced to either release through the Microsoft Store or pay for quite expensive yearly fees to have your app notarized. Without this, the user will be shown very scary warning dialogs, making the app very hostile to use. Since this is a small free app, I went the Microsoft Store route. Linux The Linux version shares some of the C++ with the Windows version, especially the code that draws the multi-coloured tree map on the right. Similar to the Windows version, it checks the hardware of your storage to best optimize itself and otherwise builds the UI using Linux native code with almost not dependencies. The first version Claude recommended depended on the GTK libraries for the UI, which was convenient, but it meant that using the app on any system that didn’t include those libraries would force the user to download hundreds of megabytes just to get a 500KB app running. Luckily, within an hour Claude had completely rewritten the app to be almost fully self contained. This means that the Linux app, which is packaged as an AppImage file, is just about 600KB all in, and you can simply download it from the site. Epilogue This was a fun experiment in building an identical app for all three operating systems while keeping it as native and optimised as possible. The hardest part was the hardware setup required for testing. I now have on my desk: My MacBook Pro (my primary machine). I do most of my work on this, run Claude and do all testing of the MacOS app. A small but powerful Windows Desktop. This is pretty great, as not only can Claude build and test Windows apps on it, it can also build and test (to some degree) Linux apps too. I use this as the main machine for those two operating systems. An ancient, 2009 MacBook Pro 17″ that I installed Linux on just for building this app. I use this for testing the Linux version on a real machine, not just on Windows WSL. It works relatively well, but with just 4GB of memory I won’t be doing any development on it any time soon. Still, it’s great to make use of the old hardware instead of throwing it away – I knew there was a reason I hung on to it! … a lot of messy crap I need to tidy up. Any day now…. I can’t believe you read this far, thanks! Now go get Disk Space from diskspace.io, your hard drive will thank you

4 weeks ago • 1 votes
Huge under-the-hood upgrade to Animations in Kidz Fun Art

One of the absolutely coolest features of Kidz Fun Art is the ability to create Animations. This was initially inspired by watching my nieces creating an animation on another Android app, so I focused purely on the creation case. This worked well, where most animations were under 50 frames in length, as it takes time to create them. I later added the ability to import Gif images, since it was a relatively simple change – parse the Gif image into frames, save them and boom, you can edit and re-export it. However this exposed a problem: Gifs can be huge, and Kidz Fun Art didn’t work well with thousands of animation frames of data. I wasn’t sure where the bottlenecks were, but at some point the app would just crash if the Gif was big enough, of if the user created an animation over 100 frames or so. This is all now fixed, and the app comfortably imports multi-Megabyte Gif images and provides a better user experience when some operations (like deleting hundreds of frames) are not instant. Using AI to find performance issues & subtle bugs While I have always written the vast majority of the code for Kidz Fun Art by hand (using AI for more complex things like WebGL shaders and C++ based paint brush simulation), it’s been invaluable recently for reviewing and testing the code. I asked Claude Opus 4.8 (the current frontier model as of July 2026) to identify the bottlenecks and it did a great job. There were multiple places where I initially wrote a function to take an action on a single frame that would then save the full animation, but later reused this function in a loop over all the frames. This caused the full animation to be saved hundreds of times in a few seconds, crashing the app. The list of frame thumbnails at the bottom of the screen rendered all thumbnails up front. This is fine for 50 but not for 1000. When saving a Gif, the file was far too large. This is because it wrote each frame in its entirety to the Gif image. The Gif standard obviously supports just writing the pixels that changed from the previous frame, and I wasn’t doing that. Deleting a large animation would take multiple seconds, with not user feedback. This was fine for other media types as they are more or less instant, but in this case it let the user click around the app, then have unexpected things happen 5 seconds later. It made the app feel broken. Gifs that stored some frames with the option to simply restore the previous frame were not handled, making the import of some images be inaccurate. There were a number of places where race conditions between multiple asynchronous actions could cause bugs. The AI was very good at finding places where my code should have been waiting for one to complete before beginning the second. There were multiple places where memory leaks occurred, specifically with not cleaning up event listeners. It found them all. When leaving the app open for days or weeks at a time without a reload, this could have been a real problem. What is better now? Animations now scale up to very large sizes, with instant access to all frames. You should be able to import basically any reasonable Gif image, and it will be exported in a highly optimized manner, with perfect colour matching per pixel. When deleting a large animation, you are told it is being deleted immediately, so it’s obvious the app is doing what you asked it to do. When importing a large Gif, you are shown a dialog telling you that it is happening, and blocking you doing other work until that completes. Many subtle bugs fixed. The frame list is fully virtualized, and scales up to essentially any size of animation. We have tests now! Another great use of AI is for writing tests, in this case laboriously creating dozens of large integration tests. These were invaluable in both validating the deep changes being made and in finding more edge cases and race conditions. Kidz Fun Art now has full end-to-end tests covering animations, layers, comics, cards, drawing, colouring, handwriting, maths and puzzles. I hate writing these, but AI doesn’t get bored, and I look forward to adding more and more regression tests in the future to keep quality high for all the world’s young artists out there.

2nd Jul 2026 • 1 votes
Spirograph fun in Kidz Fun Art

For a long time I’ve wanted to add Spirographs to my (awesome ) drawing app for kids, Kidz Fun Art, and today it’s ready! There was quite a bit of fun mathematics in getting it to feel natural and work with all sizes of circles, but it seems to have worked out very well! You can move the Spirograph around, change the size of the outer and inner circles, and draw in any colours you like. Read more about it on the main blog post here, try it out on the web at https://kidzfun.art , get it for iPad here, or download for Microsoft Windows here.

12th Jun 2026 • 1 votes
Mazers – a WebOS app rises again on iOS & iPad

Way, waaayyy back in 2010, I built a fun little game for the Palm WebOS series of phones called Mazer. I was happy with it, loads of people downloaded and played it, and then WebOS died. I recently found the source code again, and with the help of Claude AI I rewrote it to run on iOS and iPad! Get it for free today from the iOS App Store. (Android version coming soon) There are four different game types You can find your way around a simple maze, or race a terrifying fiery ball to the finish. Over 120 hand crafted obstacle courses to get around with worm holes, force fields, evil fiery balls, and more. My personal favourite, a Pacman like maze where the four ghosts chase your little ball around as you try to open the portal and get outta there!

8th Apr 2026 • 1 votes
Analyse and run simulations on your energy usage

I’ve been using the Irish energy provider Energia for 5 years or so (as of writing, 2026) and they used to have a useful insights dashboard that let me analyse my power usage. Well, they seem to have removed it so I built a handy dashboard that anyone can use. It’s at https://energy.chofter.com/ , try it out! You simply download your power usage information as a CSV file (a spreadsheet) from their site, currently at https://energyonline.energia.ie/my-account/half-hourly-usage/ . Then drop the file into the web app and it will: Show a useful overview of your usage for the full time period You can configure your current home setup This includes specifying your current tariff, whether or not you have a home battery or a car Compare usage versus last year Shows a heatmap of your usage by every 30 minutes Simulate the change in cost if you change your home setup Try out what would happen if you kept your consumption the same but changed your tariff, or added a battery or a car. This one is particularly useful.

11th Mar 2026 • 1 votes

More in programming

Clip of me singing Despard in Ruddigore in 2013

A clip of me singing a funny song from Gilbert and Sullivan’s Ruddigore back in 2013

10 hours ago • 1 votes
How and Why fork() Uses Copy-on-Write

In this video, we look at why fork() needs copy-on-write, how it works inside the kernel, and a memory usage problem that Instagram encountered with Python.

17 hours ago • 1 votes
What we lost when we lost comments

Comments require commitment, but they’re worth it.

23 hours ago • 1 votes
Two-Stack Sliding-Window Aggregation

An aggregation is some kind of summary of a set of data. This can be the sum, length, minimum, etc. It is quite common to want to calculate such a summary repeatedly, e.g. “the maximum noise level in dB for the past 30 seconds” for a nuisance detector. In such a case we say there is a sliding window over our data, and we want to aggregate over our window. If our aggregation is a binary operator with an inverse, like integer sums, there is a very easy solution using a double-ended queue: from collections import deque class SlidingWindowSum: def __init__(self): self.sum = 0 self.elems = deque() def push(self, x): self.sum += x self.elems.append(x) def pop(self): self.sum -= self.elems.popleft() def eval(self): return self.sum But what if our operator has no inverse? This is actually the case for most interesting summaries such as minimum, quantile, approximate unique count (for example using HyperLogLog), etc. In fact, even something as simple as a floating-point sum suffers from the fact that floating-point addition is not invertible. For example, if you ever have a NaN in your input data with the above naive algorithm your sum will forever remain NaN, even long after the bad value has left your window. Six years ago I came up with an algorithm for maintaining just the minimum/maximum in a sliding window and posted it to cs.stackexchange. I now consider this algorithm pointless, because it turns out there is a simple and efficient algorithm that solves this problem for a very wide class of aggregations. I’m writing this blog post to spread the word, because I feel it should be more widely known. Folklore I came across this algorithm while reading a far more advanced paper, Low-Latency Sliding-Window Aggregation in Worst-Case Constant Time by Tangwongsan et al. Why is this paper titled low-latency? Because it does the same as what I’m about to describe, but in O(1) time for each step. However, in it they also described a “two-stack” algorithm, which does it in amortized O(1), and is far, far simpler. Amortized O(1) means that across many operations the total amount of work per element is constant, but an individual operation can take much longer. This is almost always fine, unless you absolutely need a low upper bound on latency. Funnily enough that paper attributes this algorithm to “adamax” from a 2011 Stack Overflow post. They in turn credit a 2001 lecture note by D. Sleator for the inspiration. However, this lecture note does not describe a sliding window aggregate, it describes the classical two-stack algorithm for implementing a FIFO queue and does amortized analysis on it. Ultimately I would not be surprised to find that this algorithm was already described in an obscure paper from the 1970s, seeing how simple and brilliant it is. Two stacks Like the authors of the paper, I will generalize the two-stack algorithm to arbitrary associative aggregation functions. By abstracting the aggregation as a set of functions, empty(), unit(x), combine(x, y) and finalize(x), you can describe many possible aggregations, for example a mean: empty = lambda: (0, 0) unit = lambda x: (x, 1) combine = lambda x, y: (x[0] + y[0], x[1] + y[1]) finalize = lambda x: x[0] / x[1] if x[1] else None I’d like to note here that these functions have the following signatures: fn empty() -> Agg; fn unit(x: Value) -> Agg; fn combine(x: Agg, y: Agg) -> Agg; fn finalize(x: Agg) -> Out; I’m making a distinction here between Value, Agg and Out because while they seem superficially similar for something like an integer sum, for an approximate unique count on strings you would have (Value, Agg, Out) = (String, HyperLogLogSketch, u64), three wildly different types. Without further ado, the algorithm: class TwoStackAgg: def __init__(self): self.values = [] self.values_agg = empty() self.cum_aggs = [] def push(self, x): self.values.append(x) self.values_agg = combine(self.values_agg, unit(x)) def pop(self): if not self.cum_aggs: cum_agg = empty() while self.values: cum_agg = combine(unit(self.values.pop()), cum_agg) self.cum_aggs.append(cum_agg) self.values_agg = empty() self.cum_aggs.pop() def eval(self): return finalize( combine(self.cum_aggs[-1], self.values_agg) if self.cum_aggs else self.values_agg ) That’s it, the entire algorithm. There’s two stacks (values and cum_aggs) and one more aggregate, values_agg. At any point in time values_agg holds the aggregate of values, and cum_aggs contains the cumulative aggregates of all values in our window that aren’t in values, in reverse order. From this we can get the aggregate over our entire window in constant time by by combining the last value of cum_aggs with values_agg. The neat part is that (assuming w is our window size) every wth operation we drain all of values and maintain a running aggregate while pushing the partial cumulative aggregates onto cum_aggs. This is what makes it amortized O(1), doing O(w) internal operations every wth pop bounds the total amount of work per element to O(1), even though a singular operation might not be constant time. I think this is best visualized. Suppose we sum [1, 2, ..., 10] with a fixed-size sliding window of four elements, then the state on each eval() call would look like this (values_agg not shown as it is simply the aggregate of the values): cum_aggs values out [] [] = 0 [] [1] = 1 [] [1, 2] = 1 + 2 [] [1, 2, 3] = 1 + 2 + 3 [] [1, 2, 3, 4] = 1 + 2 + 3 + 4 [4, 3 + 4, 2 + 3 + 4] [5] = 2 + 3 + 4 + 5 [4, 3 + 4] [5, 6] = 3 + 4 + 5 + 6 [4] [5, 6, 7] = 4 + 5 + 6 + 7 [] [5, 6, 7, 8] = 5 + 6 + 7 + 8 [8, 7 + 8, 6 + 7 + 8] [9] = 6 + 7 + 8 + 9 [8, 7 + 8] [9, 10] = 7 + 8 + 9 + 10 [8] [9, 10] = 8 + 9 + 10 [] [9, 10] = 9 + 10 [10] [] = 10 [] [] = 0 In total the memory usage is O(w), where w is your maximum window size. Note that for simplicity of analysis and the example I assumed a fixed-size window w, but there is nothing about the two-stack algorithm that requires this. You can call push(x) and pop() as many times as you’d like between each eval(), growing and shrinking the window size as needed. Floating-point non-associativity Note that we required above that our aggregate combine is associative, meaning: combine(combine(x, y), z) = combine(x, combine(y, z)) Technically speaking, floating-point addition doesn’t respect this. Nevertheless, the above algorithm is still very useful because the results closely match the expected outcome, even more so if you use a compensated summation algorithm like Kahan summation. Another neat thing about the two-stack algorithm is that it doesn’t require commutativity, if you follow the above implementation precisely. The order of operands is maintained, which can matter for things like string concatenation. However, there is a second very useful property of the above algorithm. Each aggregate is strictly a combination of the elements in the window, and none outside the window. This means if your window contains a NaN or infinity (or some other outlier), that value only poisons the windows that contain it rather than the rest of your computation. But even without NaN or infinity it is useful, due to not propagating errors endlessly. E.g. if your sliding window starts with [1e20, 1], this is what would happen with a naive rolling sum: >>> 1e20 + 1 - 1e20 - 1 -1.0 Compensated summation will reduce these effects, but not making your result depend on values outside of the window will eliminate long-term error accumulation entirely.

23 hours ago • 1 votes
Lighthouse map

Lovely global map with animated lights sweeping the waters

yesterday • 1 votes
📚 BoredReading

You seem to be enjoying this.

Join free to unlock everything.

Create free account

Already have an account? Sign in