More from The personal website of Matthew Ström
Album art didn’t always exist. In the early 1900s, recorded music was still a novelty, overshadowed by sales of sheet music. Early vinyl records were vastly different from what we think of today: discs were sold individually and could only hold up to four minutes of music per side. Sometimes, only one side of the record was used. One of the most popular records of 1910, for example, was “Come, Josephine, in My Flying Machine”: it clocked in at two minutes and 39 seconds. via Wikipedia The packaging of these records was strictly utilitarian: a brown paper sleeve to protect the record from dust, printed with the name of the record label or the retailer. Rarely did the packaging include any information on the disc inside; the label on the center of the disc was all there was to differentiate one record from another. But as record sales started to show signs of life, music publishers took note. Columbia Records, one of the first companies to sell music on discs, was especially successful. They pioneered the sale of songs in bundles: the individual discs were bound together in packages resembling photo albums, partly to protect the delicate shellac that the records were made of, partly to increase their sales. They resembled photo albums, so Columbia called them “record albums.” There were many more technological breakthroughs that made it possible to mass-manufacture and distribute music throughout the world at affordable prices. The five-minute-long 78 rpm discs were replaced by 20-minute discs that ran at 33 ⅓ rpm, which were replaced by the hour-long 12″ LP we know today. Delicate shellac was replaced by the more resilient (and cheaper) vinyl. Both recording technology and consumer electronics were always evolving, allowing more dynamic music to fit into smaller packages and be played on smaller, higher-fidelity stereos. The invention of album art can get lost in the story of technological mastery. But among all the factors that contributed to the rise of recorded music, it stands as one of the few that was wholly driven by creators themselves. Album art — first as marketing material, then as pure creative expression — turned an audio-only medium into a multi-sensory experience. This is the story of the people who made music visible. The prophet: Alex Steinweiss Alex Steinweiss was born in 1917, the son of eastern European immigrants. Growing up in Brooklyn, New York, Steinweiss took an early interest in art and earned a scholarship to Parsons School of Design. On graduating, he worked for Austrian designer Joseph Binder, whose bold, graphic posters had influenced design for the first decades of the 1900s. The Most Important Wheels in America, Association of American Railroads (1952) via Moma Joseph Binder, Österreichs Wiederaufbau Ausstellung Salzburg (1933) via Moma Joseph Binder, Air Corps U.S. Army (Winning entry for the MoMA National Defense Poster Competition [Army Air Corps Recruiting]) via Moma After his work with Binder, Steinweiss was hired by Columbia Records to produce promotional displays and ads, but the job didn’t stick. At the outbreak of World War II, he went to work for the Navy’s Training and Development Center in New York City, designing teaching material and cautionary posters. When the war ended, Steinweiss went back to freelancing for Columbia. At a lunch meeting in 1948, company president Ted Wallerstein mentioned that Columbia would soon introduce a new kind of record that, spinning at a slower speed of 33 ⅓ rpm, could hold more music than the older 78 rpm discs. But there was a problem: the smaller, more intricate grooves on the discs were being damaged by the heavy paper sleeves used for the 78s. After the lunch, Steinweiss went to work to create a new, safer jacket for the records. But his vision for the new packaging went beyond just its construction. “The way records were sold was ridiculous,” Steinweiss said. “The covers were brown, tan or green paper. They were not attractive, and lacked sales appeal.” He suggested that Columbia should spend more money on packaging, convinced that eye-catching designs would help sell records.1 His first chance to prove his case was a 1940 compilation by the songwriters Rodgers and Hart — one of the first releases on the new microgroove 33 ⅓ records. For it, he asked the Imperial Theater (located one block west of Times Square) to change the lettering on their marquee to read “SMASH SONG HITS BY RODGERS & HART." Steinweiss had a photographer take a photo, and back in his studio, superimposed “COLUMBIA RECORDS’’ on the image to match the perspective and style of the signage. The last touch, a nod to the graphic abstraction of his mentor Joseph Binder, were orange lines arcing around the marquee in the exact size of the record underneath. Album art was born. Smash Song Hits by Rodgers & Hart via RateYourMusic Steinweiss would go on to design hundreds of covers for Columbia from 1940 to 1945. His methodology was rigorous; the covers went beyond nice pictures to be visual representations of the music itself. Before most people owned a TV set, Steinweiss’s album covers were affordable multi-sensory entertainment. Looking at the album cover and listening to the music created an experience that was more than the sum of its parts. “I tried to get into the subject,” he explains, "either through the music, or the life and times of the composer. For example, with a Bartók piano concerto, I took the elements of the piano—the hammers, keys, and strings—and composed them in a contemporary setting using appropriate color and rendering. Since Bartók is Hungarian, I also put in the suggestion of a peasant figure.” via RateYourMusic Steinweiss was prophetic: His colorful compositions sold records. Newsweek reported that sales of Bruno Walter’s recording of Beethoven’s “Eroica” symphony increased 895% with its new Steinweiss cover.” 2 Eroica The challenger: Reid Miles From 1940 to 1950, Columbia Records was the dominant force in music sales. Buoyed by Steinweiss’s initial successes, Columbia hired more artists and designers to produce album art. Jim Flora led the charge from 1947–1950 with irreverent illustrations and more daring explorations of typography, and like Steinweiss, his work mirrored the music on the records. During the era, Columbia began to focus much more on popular music. Flora’s campy compositions screamed “this isn’t your parent’s music.” Gene Krupa and His Orchestra via JimFlora.com Jim Flora's cover for Bix and Tram via JimFlora.com Jim Flora's cover for Kid Ory and His Creole Jazz Band via JimFlora.com But while Columbia was focusing on making it into the hit parade, an upstart label was honing in on a sound that would come to define the era; Blue Note Records, founded in 1939, was fixated on the jazz underground. From its founding and throughout the 1950s, Blue Note focused on “hot jazz,” a mutant strain of jazz descending from the big band swing era, often including twangy banjoes, wailing clarinets, and rambunctious New Orleans second-line-style drumming. Founder Alfred Lion wrote the label’s manifesto: Blue Note Records are designed simply to serve the uncompromising expressions of hot jazz or swing, in general. Any particular style of playing which represents an authentic way of musical feeling is genuine expression. … Blue Note records are concerned with identifying its impulse, not its sensational and commercial adornments.3 One way Blue Note stood out from labels like Columbia was their dedication to their artists. Many of the working musicians of the ’50s lived like vampires, waking up after dusk and playing gigs into the early hours of the morning, then rehearsing until dawn. Blue Note would record their artists in the pre-dawn hours, giving musicians time to rest up before their next night’s gigs started. Art Blakey, Thelonius Monk, Charlie Parker, Dizzy Gillespie, and John Coltrane are household names now; but then, because of their drinking, drug use, and frenetic schedules, labels wouldn’t work with them. Blue Note embraced them, feeding their fires of creative innovation and creating an updraft for the insurgency of jazz to come. Album art was one more revolutionary way for Blue Note to explore “genuine expression.” Just as they fostered talented musicians, they’d give young designers a chance to shine. Alfred Lion’s childhood friend Francis Wolff had joined the label as a producer and photographer; he’d shoot candid portraits of the musicians as they worked. Then, designers like Paul Bacon, Gil Mellé (himself a musician), and John Hermansader would pair Wolff’s black-and-white photos with a single, bright color, then juxtapose them with stark, sans-serif type. Genius of Modern Music Vol. 1 via Deep Groove Mono Gil Mellé's cover featuring Francis Wolff's photography for his band's New Faces — New Sounds via Deep Groove Mono John Hermansader's cover featuring Francis Wolff's photography for George Wallington's Showcase via Deep Groove Mono As the 1960s approached, the musicians Blue Note worked so hard to cultivate were forging new styles, leaving behind the swing-era pretense of jazz as dance music. Charlie Parker and Bud Powell kept speeding up the tempo and stuffing more chords into progressions. Max Roach started playing the drums like a boxer, bobbing and weaving around the beat with skittering cymbals, waiting for the right moment to land a single monumental “thud” of a kick drum. Without the drums keeping a steady rhythm, bass players like Milt Hinton and Gene Ramey had to furiously mark out time with eighth notes, traversing chords by plucking up and down the scale. This was bebop, and it was musicians’ music. Blue Note’s ethos of artistic integrity was the perfect Petri dish for virtuosic musicians to develop innovative sounds — they worked in small ensembles, often just five players, constantly scrambling and re-arranging instrumentation, playing harder and faster and louder. Then, around 1955, just as Blue Note was hitting its stride, Wolff met a 28-year-old designer named Reid Miles. Miles had recently moved to New York and had been working for John Hermansader at Esquire magazine. He was a big fan of classical music but wasn’t so interested in jazz. Wolff convinced Miles to start designing covers for Blue Note all the same and kicked off one of the most influential partnerships in modern design. The first cover Miles created was for vibraphone player Milt Jackson; it picked up from the established art style, with Wolff’s photos and a single bright hue. But the type was even more exaggerated, and the photo took up more than half the cover. White dots overlayed on Jackson’s mallets were the perfect abstraction of the staccato tones of the vibraphone. It’s a great cover, but it was just a hint of what was to come. via Ariel Salminen A common theme of Miles’ covers was the emphasis on Wolff’s photography. We’re familiar with these iconic images today, but at the time they were revolutionary; before, black musicians like Louis Armstrong and Ella Fitzgerald were portrayed in tuxedos and evening gowns, posed smiling genially or laughing, rendered so as to not offend the largely white listening audience. Wolff’s portraits were candid, realistic, showing black musicians at work. For example, the cover for Art Blakey’s The Freedom Rider shows Blakey lost in a moment, almost entirely obscured by a cymbal. The drummer is smoking a cigarette, but it’s barely hanging onto the corner of his lip — his mouth is half-open, his brows clenched in a moment of agony or ecstasy. Miles would let the photo fill up the entire cover, cramming the name of the record into whatever empty space was available. The Freedom Rider via London Jazz Collector Miles sometimes reversed this relationship, pioneering the use of typography to convey the spirit of the music. His cover for Jackie McLean’s It’s Time! is composed of an edge-to-edge grid of 243 exclamation marks; a postage stamp picture of McLean graces the upper corner, almost a punchline. Lee Morgan’s The Rumproller is another type-only cover, this time with the title smeared out from corner to corner, like it was left on a hot dashboard for the day. Larry Young’s Unity has no photo at all; the four members of the quartet become orange dots resting in (or bubbling out of) the bowl of the U. It's Time via Ariel Salminen Reid Miles' cover for Lee Morgan's The Rumproller via Fonts in Use Reid Miles' cover for Larry Young's Unity Miles fulfilled the Blue Note manifesto. His album covers pushed the envelope of graphic design just as the artists on the records inside continued to break new ground in jazz. With the partnership of Miles and Wolff, alongside Alfred Lion’s commitment to artistic integrity, Blue Note became the standard-bearer for jazz. Columbia Records couldn’t help but notice. Even though Blue Note wasn’t nearly as commercially successful as Columbia, their willingness to take risks had established them as a much more sophisticated, innovative, and creative label; to compete for the best talent, Columbia would need to find a way to win the attention of both artists and listeners. The master: S. Neil Fujita Sadamitsu Fujita was born in 1921 in Waimea, Hawaii. He was assigned the name Neil in boarding school — leading up to World War II, anti-Japanese sentiment was rampant, especially in Hawaii. Fujita moved to LA to attend art school, but his studies were cut short in 1942 when Franklin Roosevelt signed executive order 9066, allowing the imprisonment of Japanese Americans living on the west coast. Fujita was sent to Wyoming, where he enlisted in the 442nd Regimental Combat Team. Before the war was over, he’d see combat in Italy, France, and the Pacific theater. After the war, Fujita finished his studies in LA. He quickly made a name for himself in the advertising world; his résumé landed on the desk of Bill Golden, the art director for CBS, which owned Columbia Records. Alex Steinweiss, the first album artist and Columbia’s ace in the hole, had moved on to RCA. Columbia needed a new direction. Golden called Fujita and asked him to run the art department. Fujita would be building a whole new team, replacing the relationships that Columbia had built with art studios for hire. This wasn’t going to be the hardest part of Fujita’s work; when offering him the job, Golden warned him that he’d experience a lot of racist attitudes still simmering in the wake of World War II.4 Still, Fujita agreed to take the job. Fujita’s first covers fit in with the work that Reid Miles was doing at Blue Note: single-color accents set against black-and-white photography. The Jazz Messengers via Discogs Fujita's cover for Miles Davis' 'Round About Midnight via Discogs In 1959, jazz was leaving the stratosphere. Ornette Coleman was performing what he called “free jazz,” frenetic, inscrutable compositions that drew backlash and praise in equal parts. John Coltrane recorded Giant Steps with a level of virtuosity that even his own bandmates struggled to keep up with. Miles Davis recorded Kind of Blue, which would go on to be regarded as one of the best recordings of all time. Fujita was also breaking ground at Columbia. He was one of the first directors to hire both men and women in a racially integrated office.5 He delegated work, tapping painters, illustrators, and photographers to contribute to covers. Fujita himself trained to be a painter before starting his career in design; he started looking for ways to incorporate his own original paintings into the covers: “We thought about what the picture was saying about the music,” Fujita recalled, “and how we could use that to sell the record. And abstract art was getting popular so we used a lot more abstraction in the designs—with jazz records especially.” He got the perfect opportunity to make his mark with two albums released in 1959: Charles Mingus’s Mingus Ah Um and Dave Brubeck’s Time Out. Mingus Ah Um Fujita's cover for Dave Brubeck's Time Out Fujita’s abstract paintings reflected the pure exuberance of Mingus’ and Brubeck’s music. In the case of Mingus Ah Um, the divisions and intersections spanning the cover read like a beam of light passing through exotic lenses, magnifiers, refractors, and prisms; through his music, Mingus was reflecting on the transition of jazz from popular entertainment to mind-expanding creative exercise. For Time Out, the wheels and rollers spooling out across the page echo the way that Brubeck’s quartet was experimenting with how time signatures could be interlocked, multiplied, and divided to create completely new textures and musical patterns. Fujita’s covers made it plain: Jazz was art. ’59 turned out to be a watershed for both jazz and album art. Brubeck’s Time Out went to #2 on the pop charts in 1961, and was the first jazz LP to sell more than a million copies; “Take Five,” the album’s standout hit, would also become the first jazz single to sell a million copies. For a unique moment in time, the music and art worlds were being propelled forward by a commercially successful record. Fujita’s paintings were making their way into millions of homes, driving sales of records by the vanguards of jazz. Fujita left Columbia records shortly after these major successes. “I wanted to be something other than just a record designer,” he said, “so I left to go on my own.” He’d go on to design the book covers for Truman Capote’s In Cold Blood and Mario Puzo’s The Godfather — when the latter was turned into Francis Coppola’s breakthrough film, Fujita’s design was used for its title and promotional art. But he’d continue to design album covers, creating paintings for each one. Far Out, Near In Fujita's cover for Dony Byrd and Gigi Gryce's Modern Jazz Perspective Fujita's cover for Columbia's recording of Glenn Gould performing Berg, Křenek, and Schoenberg. The next generation As jazz continued to evolve throughout the ’60s and ’70s, melding with rock ’n roll to produce punk, electronic, R&B, and rap, album art evolved alongside. Packaging became more sophisticated: multi-disc albums came in folding cases called gatefolds, accompanied by booklets of photography and art. New printing techniques allowed for brighter colors, shiny foil stamps, and textured finishes. Budgets for production grew larger and larger. The Beatles’ Sgt. Pepper’s Lonely Hearts Club Band featured an elaborate photo of the band members, 57 life-sized photograph cutouts, and nine wax sculptures. For the first time for a rock EP, the lyrics to the songs were printed on the back of the cover. In another first, the paper sleeve inside was not white but a colorful abstract pattern instead. Also inside was a sheet of cardboard cutouts, including a postcard portrait of Sgt. Pepper, a fake mustache, sergeant stripes, lapel badges, and a stand-up cutout of the Beatles themselves. The zany campiness of Sgt. Pepper’s could only be matched by an absurd gift box full of toys and games. The stark loneliness of the Beatles’ next album would be paired with a plain white cover, without even ink to fill in the impression of the words “The Beatles” on the front. Sgt. Pepper's Lonely Hearts Club Band, designed by Jann Haworth and Peter Blake and photographed by Michael Cooper The cover of The Beatles, designed by Richard Hamilton via Reddit The most famous artists and designers of each generation would try their hand at album art. Salvador Dali, Andy Warhol, Saul Bass, Keith Haring, Annie Leibovitz, Jeff Koons, Shepard Fairey, and Banksy would all create work for albums. Some of those pieces would become the most recognizable ones in an artist’s catalog. Greatest Hits by The Modern Jazz Quartet Andy Warhol's cover for The Velvet Underground & Nico via Leo Reynolds Saul Bass's cover for Frank Sinatra Conducts Tone Poems of Color via Moma Keith Haring's cover for David Bowie's Without You Annie Leibovitz and Andrea Klein's cover for Bruce Springsteen's Born In The USA Jeff Koons' cover for Lady Gaga's Artpop Shepard Fairey's cover for The Smashing Pumpkins' Zeitgeist Banksy's cover for Blur's Think Tank None of this would have been possible without the contributions of Alex Steinweiss, Jim Flora, Paul Bacon, Gil Mellé, John Hermansader, Reid Miles, S. Neil Fujita, and others. If not for the arms race between Columbia Records and Blue Note for the best art and the best artists of the ’50s, many artists would never have found their career. And in some cases, an album like The Rolling Stones’ Sticky Fingers would be remembered more for its art than for its music. When music was first pressed into discs, design was less than an afterthought. Today, album art is an extension of music itself. Footnotes & References https://www.nytimes.com/2011/07/20/business/media/alex-steinweiss-originator-of-artistic-album-covers-dies-at-94.html ↩︎ https://web.archive.org/web/20120412033422/http://www.adcglobal.org/archive/hof/1998/?id=318 ↩︎ https://web.archive.org/web/20080503055603/https://www.bluenote.com/History.aspx ↩︎ https://www.hellerbooks.com/pdfs/voice_s_neil_fujita.pdf ↩︎ https://www.nationalww2museum.org/war/articles/s-neil-fujita ↩︎
It used to be easy to pick colors for design systems. Years ago, you could pick a handful of colors to match your brand’s ethos, or start with an off-the-shelf palette (remember flatuicolors.com?). Each hue and shade served a purpose, and usually had a quirky name like “idea yellow” or “innovation blue”. This hands-on approach allowed for control and creativity, resulting in color schemes that could convey any mood or style. But as design systems have grown to keep up with ever-expanding software needs, the demands on color palette have grown exponentially too. Modern software needs accessibility, adaptability, and consistency across dozens of devices, themes, and contexts. Picking colors by hand is practically impossible. This is a familiar problem to the Stripe design team. In “Designing accessible color systems,” Daryl Koopersmith and Wilson Miner presented Stripe’s approach: using perceptually uniform color spaces to create aesthetically pleasing and accessible systems. Their method offered a new approach to selection to enhance beauty and usability, grounded in scientific understanding of human vision. In the four years since that post, Stripe has stretched those colors to the limit. The design system’s resilience through massive growth is a testament to the team’s original approach, but last year we started to see the need for a more flexible, scalable, and inclusive color system. This meant both an expansion of our color palette and a rethinking of how we generate and apply these colors to accommodate our still-growing products. This essay will take you through my attempts to solve these problems. Through this process, I’ve created a tool for generating expressive, functional, and accessible color systems for any design system. I’ll share the full code of my solution at the end of the essay; it represents not just a technical solution but a philosophical shift in how we think about color in design systems, emphasizing the balance between creativity and inclusivity. Why don’t the existing tools work? So what makes a good color palette? Through the looking glass: perceptual uniformity Picking the right color space Using OKHsl First steps with generated scales Scaling up Making scales expressive: Leveraging hue and saturation Hue Saturation and Chroma In practice: Crafting colors with functions Pick base hues Add functions for hue, saturation, and lightness Calculate the colors for each scale number Making scales adaptive: Using background color as an input Making scales accessible: Building in the WCAG contrast calculation Step 1: Calculate a target contrast ratio based on scale step Step 2: Calculate lightness based on a target contrast ratio Step 3: Translate from XYZ Y to OKHsl L Putting it all together: All the code you need What does it look like in practice? What we’ve learned and where we’re going Why don’t the existing tools work? In the past few years, I’ve come across dozens of tools that promise to generate color palettes for design systems. Some are simple, like Adobe Color, which generates color palettes based on a single input color, or even an image. Others are more complex, like Colorbox, which generates color scales based on a long list of parameters, easing curves, and input hues. But I’ve found that each of these tools has critical limitations. Complex tools like Colorbox or color x color allow for a high degree of customization, but they require a lot of manual input and don’t provide guidelines for accessibility. Simple tools like Adobe’s Color and Leonardo provide more constraints and accessibility features, but they does so at the expense of flexibility and extensibility. None of the tools I’ve found can integrate tightly with an existing design system; all are simply apps that generate an initial set of colors. None can respond to the unique constraints of your design system, or adapt as you add more themes, modes, or components. That’s why I ended up going back to first principles, and decided to build up a framework that can be adapted to any codebase, design tool, or end user interface. So what makes a good color palette? To build palettes from first principles, we need a strong conceptual foundation. A great color palette is like a Swiss Army knife, built to address a wide array of needs. But that same flexibility can make the system unwieldy and clunky. Through years of working on design systems, two principles have emerged as a constant benchmark for quality color palettes: utility and consistency. A color palette with high utility is vital for a robust design system, encompassing both adaptability and functionality. It should offer a wide array of shades and hues to cater to diverse use cases, such as status changes—reds for errors, greens for successes, and yellows for warnings—and interaction states like hovering, disabled, or active selections. It’s also essential for signifying actionable items like links and buttons. Beyond functionality, an adaptable palette enables smooth transitions between light, dark, and high contrast modes, supporting the evolution of your product and differing brand expressions. This ensures that your user interfaces remain consistent and recognizable across various platforms and usage contexts. Moreover, an adaptable palette underscores a commitment to accessibility—it should provide accessible contrast ratios across all components, accommodating users with visual impairments, and offer high-contrast modes that enhance visibility and functionality without sacrificing style. Consistency is another crucial aspect of a well-designed color palette. Despite the diverse range of components and their variants, a consistent palette maintains a coherent visual language throughout the system. This coherence ensures that elements like badges retain a consistent visual weight, avoiding misleading emphasis, and the relative contrast of components remains balanced between dark and light modes. This consistency helps preserve clarity and hierarchy, further enhancing the user experience and the overall aesthetics of the design system. As you’ll see, even simple questions about these goals reveals a deep rabbit hole of possible solutions. Through the looking glass: perceptual uniformity The principles of utility and consistency make selecting a color palette more complex. There’s a question at the heart of both constraints: what makes two colors look different? We have an intuitive sense that yellow and orange are more similar than green and blue, but can we prove it objectively? Scientists and artists have spent the last decade puzzling this out, and their answer is the concept of perceptual uniformity. Perceptual uniformity is rooted in how our eyes work. Humans see colors because of the interaction between wavelengths of light and cells in our eyes. In 1850, before we could look at cells under a microscope, scientist Hermann von Helmholtz theorized that there were three color vision cells (now known as cones) for blue, green, and red light. Thomas Young and Hermann von Helmholtz assumed that the eye’s retina consists of three different kinds of light receptors for red, green and blue. Public Domain via Wikipedia Most modern screens depend on this century-old theory, mixing red, green, and blue light to produce colors. Every combination of these colors produces a distinct one; 10% red, 50% green, and 25% blue light create the cartoon green of the Simpson’s yard. 75% red, 85% green, and 95% blue is the blindingly pale blue of the snow in Game of Thrones. Von Helmholtz was amazingly close to the truth, but until 1983, we didn’t have a full understanding of the exact way that each cell in our eyes responds to light. While it’s true that we have three kinds of color vision cells, and that each responds strongly to either red, green, or blue light, the full mechanism of color vision is much more nuanced. So, while it’s technologically simple to mix red, green, and blue lights to reproduce color, the red, green, and blue coordinate system — the RGB color space — isn’t perceptually uniform. Picking the right color space Despite not being perceptually uniform, many design systems still use RGB color space (and its derivative, HSL space) for picking colors. But over the past century, scientists and artists have invented more useful ways to map the landscape of color. Whether it’s capturing skin tones accurately in photographs or creating smooth gradients for data visualization, these different color spaces give us perceptually uniform paths through a gamut. Lab is an example of a perceptually uniform color space. Developed by the International Commission on Illumination, or CIE, the Lab color space is designed to be device-independent, encompassing all perceivable colors. Its three dimensions depict lightness (L), and color opponents (a and b) — the latter two varying between green-red and blue-yellow axes respectively. This makes it useful for measuring the differences between colors. However, it’s not very intuitive; for example, unless you’ve spent a lot of time working with the lab color space, it’s probably hard to imagine what a pair of (a, b) values like (70, -15), represents.1 LCh (Luminosity, Chroma, hue) is more ergonomic, but still perceptually uniform color space. It’s a cylindrical color space, which means that along the hue axis, colors change from red to blue to green, and then back to red — like traveling on a roundabout. Along the way, each color appears equally bright and colorful. Moving along the luminosity axis, a color appears brighter or dimmer but equally colorful, like adjusting a flashlight’s distance from a painted wall. Along the chroma axis, a color stays equally bright but looks more or less colorful, like it’s being mixed with different amounts of gray paint. The LCh color space. Note the uneven peaks of chroma at different hues. via Hueplot LCh trades off some of lab’s functionality for being more intuitive. But LCh can be clunky, too, because the C (chroma) axis starts at 0 and don’t have a strict upper limit. Chroma is meant to be a relative measure of a color’s “colorfulness”. Some colors are brighter and more colorful than others: is a light aqua blue as colorful as a neon lime green? How does a brick red compare to a grape soda purple? The chroma scale is meant to make these comparisons possible. But try for a moment to imagine a sea green as rich and deep as an ultraviolet blue. Lab and LCh both let you specify these “impossible” colors that don’t have a real-world representation. In technical parlance, they’re called “out of gamut,” since they can’t be produced by screens, or seen by human eyes. The existence of out-of-gamut colors makes it hard to reliably build a color system in LCh or lab color space. Finding colors with consistent properties is a manual process; when Stripe was building its previous color system using lab, the team made a specialized tool for visualizing the boundaries of possible colors, allowing designers to tweak each shade to maximize its saturation. This isn’t a tenable solution for most teams; what if there was a color space that combined the simplicity of RGB and HSL with the perceptual uniformity of lab and LCh? Björn Ottosson, creator of the OKLab color space, did just that in his blog post “OKHsv and OKHsl — two new color spaces for color picking.” OKHsl is similar to Lch in that it has three components, one for hue, one for colorfulness, and one for lightness. Like LCh, the hue axis is a circle with 360 degrees. The lightness axis is similar to Lch’s luminosity, going from 0 to 1 for every hue. In place of Lch’s chroma channel, though, OKHsl uses an absolute saturation axis that goes from 0 to 1 for every hue, at every lightness. 0 represents the least saturated color (grays ranging from white to black), and 1 represents the most saturated color available in the sRGB gamut. The OKHsl color space. It’s a cylinder, which makes it much better for generating color palettes. via Hueplot Practically, OKHsl allows for easier color selection and manipulation. It bypasses the issues found in LCh or lab, creating an intuitive, straightforward, and user-friendly system that can produce the desired colors without worrying about out-of-gamut colors. That’s why it’s the best space for generating color pallettes for design systems. Using OKHsl Practically speaking, to use OKHsl, you need to be able to convert colors to and from sRGB. This is a fairly straightforward calculation, but it’s not built into most design tools. Bjorn Ottosson linked the javascript code to do this conversion in his blog post, and the library colorjs.io will soon have support for OKHsl. Going forward, I’ll assume you have a way to convert colors to and from OKHsl. If you don’t, you can use the code I’ve written to generate a color palette in OKHsl, and then convert it to sRGB for use in your design system. First steps with generated scales To get started generating our color scales, we need a few values: The hue of the color we want to generate The saturation of the color we want to generate A list of lightness values we want to generate For example, we can generate a cool neutral color scale by choosing these values: Hue: 250 Saturation: 5 Lightness values: Light: 85 Medium: 50 Dark: 15 Using those values to pick colors in the OKHsl color space, we get the following palette: Neutral OKHsl sRGB Hex 250, 5, 85 #d2d5d8 250, 5, 50 #73787c 250, 5, 15 #212325 We can do the same thing for all our colors, picking numbers to build out the entire system. OKHsl sRGB Hex OKHsl sRGB Hex 250, 5, 85 #d2d5d8 250, 90, 85 #b6d9fd 250, 5, 50 #73787c 250, 90, 50 #1a7acb 250, 5, 15 #252628 250, 90, 15 #022342 OKHsl sRGB Hex OKHsl sRGB Hex OKHsl sRGB Hex 145, 90, 85 #6af778 20, 90, 85 #fec3ca 100, 90, 85 #eed63d 145, 90, 50 #388b3f 20, 90, 50 #d32d43 100, 90, 50 #877814 145, 90, 15 #0c2a0e 20, 90, 15 #45060f 100, 90, 15 #282302 Scaling up For bigger projects, you’ll often need more than just three shades per color. Choosing the right number can be tricky: too few shades limit your options, but too many can cause confusion. This can seem daunting, particularly in the early stages of your design system. But there’s a method to simplify this: use a consistent numbering system, ensuring your color choices remain versatile no matter how your system evolves. This system is often referred to as ‘magic numbers.’ If you’re familiar with Tailwind CSS or Material Design, you’ve seen this in action. Instead of naming shades like ‘light’ or ‘dark,’ each shade gets a number. For instance, in Tailwind, the scale goes from 0 to 1,000, and in Material Design, it’s 0 to 100. The extremes often correspond to near-white or near-black, with middle numbers denoting pure hues. The beauty of this system is its flexibility. If you initially use shades named ‘red 500’ and ‘red 700’, and later need something in between, you can simply introduce ‘red 600’. This keeps your design adaptable and intuitive. Another bonus of magic numbers is that we can often plug the number directly into a color picker to scale the lightness of the shade. That’s why, for the rest of this essay, I’ll call these scale numbers. For example, if we wanted to create a more extensive color scale for our blues, we could use the following values in the OKHsl color space: Blue Scale Number OKHsl sRGB Hex 0 250, 90, 100 #ffffff 10 250, 90, 90 #cfe5fe 20 250, 90, 80 #9dccfd 30 250, 90, 70 #68b1f9 40 250, 90, 60 #3395ed 50 250, 90, 50 #1b7acb 60 250, 90, 40 #0f60a3 70 250, 90, 30 #08477c 80 250, 90, 20 #032f55 90 250, 90, 10 #01172e 100 250, 90, 0 #000000 We’ve turned the scale number into the lightness value with the function $$L(n) = 1-n$$. In this formula, n is a normalized value — one that goes from 0 to 1 — that represents our scale number, and L(n) is the lightness value in OKHsl. It turns out that using functions and formulas in combination with scale numbers is a powerful way to create expressive color scales that can power beautiful design systems. Making scales expressive: Leveraging hue and saturation One advantage of scale numbers is the ability to plug them directly into a color picker to dictate the lightness of a shade. But scale numbers really show their usefulness and versatility when you leverage them across every component of your color system. That means venturing beyond lightness to explore hue and saturation, too. Hue When using scale numbers to control lightness, it’s easy to assume hue and saturation will behave consistently across the lightness range. However, our perception of color doesn’t work that simply. Hue can appear to shift dramatically between light and dark shades of the same color due to a phenomenon called the Bezold–Brücke effect: colors tend to look more purplish in shadows and more yellowish in highlights. So if we want to maintain consistent hue perception, we can use scale numbers for adapting the hues of our color scales. As lightness decreases, blues and reds should shift slightly towards more violet/purple tones to counteract the Bezold–Brücke effect. Likewise, as lightness increases, yellows, oranges, and reds should shift towards more yellowish hues. 23 Purple without (top) and with (bottom) accounting for the Bezold–Brücke shift Red without (top) and with (bottom) accounting for the Bezold–Brücke shift In both examples above, we’ve used the scale number to shift the hue slightly as the lightness increases. This looks like the following formula: $$H(n) = H_{base} + 5*(1 - n)$$. H(n) is the hue at a given normalized scale value; Hbase is the “base hue” of the color. The 5*(1 - n) term means the hue will change by 5 degrees as the scale number goes from one end to the other. If you’re using this formula, you should tweak the numbers to your liking. By making hue a function of lightness, with the scale number adjusting hue accordingly, hues look more consistent and harmonious across the entire scale. The shifts don’t need to be large – even subtle hue variations of a few percentage points can perceptually compensate for natural hue shifts with changing brightness. Saturation and Chroma From our understanding of the CIE LCh color space and its sibling, the OKHsl color space, we know that colors generally attain their peak chroma around the middle of the lightness scale.4 In design, this presents a fantastic opportunity. By designing our color scales such that the midpoint is the most chromatically rich, we can make sure that our colors are the most vibrant and saturated where it matters most. Conversely, as we veer towards the lightness extremes, we can have chroma values that taper off, ensuring that our brightest and darkest shades remain subtle and balanced. OKHsl gives us a saturation component that goes from 0% to 100% of the possible chroma at a given hue and lightness value. We can take advantage of this by using the normalized scale number as an input to a function that goes from a minimum saturation to a maximum and back again. Green with constant saturation (top) and varying saturation (bottom) In practice, the formula for achieving this looks like this: $$S(n) = -4n^2 + 4n$$, where S(n) is the saturation at a given (normalized, as before) scale value n. The formula is an upside-down parabola, which starts at 0% and peaks at 100% when the scale value is 0.5. You can add a few terms to adjust the minimum and maximum saturation if you’d like to adjust the scale further: Neutrals, for example, don’t need a high maximum saturation. But most colors do well moving between 0% and 100% saturation. In practice: Crafting colors with functions Let’s put this into practice and generate an extensive color scale with only a handful of functions. Functions allow us to build a flexible framework that is resilient to change and can be easily adapted to new requirements; if we need to add colors, tweak hues, or adjust saturation, we can do so without rewriting the entire system. Pick base hues First, let’s pick a handful of base hues. At the very least, you’ll need a blue for interactive elements like links and buttons, and green, red, and yellow for statuses. Your neutrals need a hue, too; though it won’t show up much, a cool neutral and a warm neutral have very different effects on the overall system. Neutral Blue Green Red Yellow Base hue (Hbase) 250 250 145 20 70 Add functions for hue, saturation, and lightness Next, let’s use the functions we came up with earlier to indicate how the colors should change depending on scale numbers. Neutral Blue Green Red Yellow Base hue (Hbase) 250 250 145 20 70 Hue function $$H(n) = 250$$ $$H(n) = H_{base} + 5*(1 - n)$$ Saturation function $$S(n) = -0.8n^2 + 0.8n$$ $$S(n) = -4n^2 + 4n$$ Lightness function $$L(n) = 1-n$$ The hue function is a constant for neutrals, and for the colors we use the function that accounts for the Bezold–Brücke shift. As for saturation, the neutral colors have a maximum saturation of 20% instead of the full 100%; the rest of the colors use the function that goes from 0% to 100% and back. The lightness function is the same for all colors. Calculate the colors for each scale number Now let’s let the math work its magic. For each scale number, for every color, we have all the information we need from the base hue, hue function, saturation function, and lightness function. sRGB Hex Scale Number Neutral Blue Green Red Yellow 0 #ffffff #ffffff #ffffff #ffffff #ffffff 10 #e0e3e6 #dae4f0 #d8e8d4 #f9d1d6 #f7e9a3 20 #bfc8d1 #aacaf1 #9adb90 #f1b5b7 #ebbe83 30 #9fadbd #73aff6 #67c55b #f7838c #e09c34 40 #8193a6 #2e92f9 #39ac30 #fa405e #c3810a 50 #67798c #0077d8 #009100 #dd0042 #a26900 60 #506070 #065faa #227021 #ae0f33 #815304 70 #3c4752 #0e477c #255125 #7e1a28 #5f3e0b 80 #292f35 #12304d #1c351c #4e1b1e #3e290f 90 #141619 #0d1722 #101910 #221111 #1d150b 100 #000000 #000000 #000000 #000000 #000000 Of course, this palette is fairly basic and might not be optimal for your needs. But using formulas and functions to calculate colors from scale numbers has a powerful advantage over manually picking each color; you can make tweaks to the formulas themselves and instantly see the entire palette adapt. Making scales adaptive: Using background color as an input Today, color modes like dark mode and high-contrast accessibility mode are table stakes in design systems. So, if you’re picking colors manually, you have to pick an additional 50 colors for each mode, carefully balancing the unique perception of color against each different background. However, with the functions-and-formulas approach to picking colors, we can abstract a color palette to respond to any background color we might want to apply. Let’s go back to the lightness formula we used in the previous palettes: Using this formula, the lightness will decrease as the scale number increases. In dark mode, we want the opposite: lightness should increase as the scale number increases. We can use a more detailed formula to switch the direction of our scale if the lightness of a background color is less than a specific value: 0.18 \\n, & \text{if } Y_{b} The “Yb” in this equation is the background color’s Y value in the XYZ color space. As I explained at the beginning of this essay, color spaces are different ways of mapping all the colors in a gamut; XYZ is an extremely precise and comprehensive color space. While the X and Z components don’t map neatly to phenomenological aspects of a color (like a and b in the LAB color space), the Y component represents the luminosity of a color. You may be wondering why we’re using another color space (in addition to OKHsl) to dictate lightness. This is because the WCAG (Web Content Accessibility Guidelines) color contrast algorithm compares Y values in XYZ space, which will be more relevant in the next section. A color with a Y value of 0.18 will have the particular quality of passing WCAG contrast level AA5 on both pure white ( #ffffff) and pure black ( #000000). That makes it a good test to see if a color is a light background (Yb > 0.18) or a dark background (Yb < 0.18). Using this equation for our color system, we can now get both dark mode and light mode colors, calculated automatically based on the background color we choose. The color palette calculated with a background color of #000000 (Yb = 0) sRGB Hex Scale Number Neutral Blue Green Red Yellow 0 #000000 #000000 #000000 #000000 #000000 10 #141619 #0e1722 #10190f #221112 #1c150b 20 #292f35 #132f4f #1e351a #4e1a20 #3d2a0f 30 #3c4752 #10467f #275122 #7e192b #5e3e0b 40 #506070 #075eac #25701e #ae0e36 #805304 50 #67798c #0077d8 #009100 #dd0042 #a26900 60 #8193a6 #2993f8 #35ac35 #fa405c #c4810a 70 #9fadbd #6fb0f6 #61c660 #f78489 #e29b35 80 #bfc8d1 #a7caf1 #96db94 #f1b5b5 #ecbd86 90 #e0e3e6 #d9e4f0 #d6e9d6 #f0dedd #eee0d1 100 #ffffff #ffffff #ffffff #ffffff #ffffff Making scales accessible: Building in the WCAG contrast calculation One of the most helpful aspects of scale numbers is that they can simplify accessibility substantially. The first time I saw this feature was with the US Web Design System’s (USWDS) design tokens.The USWDS color tokens have scale numbers from 0–100; using any tokens that have a scale number of 50 or more guarantees that those colors will meet the WCAG color contrast criteria at AA level. This makes designing accessible interfaces much easier. Instead of manually running each color pairing through a color contrast check, you can compare the scale numbers of the design tokens and instantly know if the combination meets accessibility criteria. When I first set out to build out a system of functions for Stripe’s color palette, this was the most daunting part of the challenge. Going in, I wasn’t even sure if it was possible to systematically target contrast ratios across all hues. However, after seeing the technique used in Adobe’s Leonardo, I had some degree of hope that such a function existed. After many false starts and dead ends, I found the right set of operations. Step 1: Calculate a target contrast ratio based on scale step Stripe’s color scales follow the lead of the USWDS; when scale numbers differ by 500 or greater, those two colors conform to the AA-level contrast ratio of 4.5:1. This means that when neutral.500 is used on top of neutral.0 (or vice versa), the color combination should be accessible. To accomplish this with calculated colors, it’s important to understand how WCAG’s contrast ratio is measured. A contrast ratio like 4.5:1 is the output ® of the following formula, which the WCAG calls “relative luminance”:6 In this equation, L1 is the luminance (i.e., the Y value of the color in XYZ color space) of the lighter color, and L2 is the luminance of the darker color. So how do we use this knowledge to transform scale steps into contrast ratios? Well, we know step 0 and 500 need to have a ratio of 4.5. Step 100 and step 600 also need to have a ratio of 4.5, and so on, up the scale. This is a feature of exponential equations; equally-spaced points along the function have consistent ratios. Exponential equations also model the growth of a population, or the spread of a virus. It happens that luminosity is also an exponential function of scale step, which shouldn’t be surprising if you know a bit of calculus. Exponential functions take the form $$f(x) = e^{kx}$$, where k is some constant. In our case, we’ll call the function r(x) (for contrast ratio), where x is a number between 0 and 1 that represents our scale step; we need to solve for k to find the exact constant that produces the correct contrast ratios. Since r(0.5) should be 4.5 — that is, scale step 500 has a contrast ratio of 4.5:1 with step 0 — we start with $$4.5 = e^{k * 0.5}$$. Solving for k yields $$k = ln(20.25)$$. To make this a little easier to work with, we can use a close approximation of this value, 3.009. And if browsers were perfect pieces of software, that would be that. But color in web browsers is a tricky technical problem. Specifically, when you convert an RGB color like rgb(129.3, 129.3, 129.3) to a hex color, it’s rounded off; the result is #818181, which is exactly rgb(129, 129, 129). The formula we derived, $$r(x) = e^{3.008x}$$ is exact, so if you round a color’s values at all after calculating it, you may end up with inaccessible colors. Therefore, in testing this function, I’ve found that adding a little extra contrast to the overall system helps guard against rounding errors. The final formula I used to calculate the contrast ratio from a scale step is as follows: Where r(x) is the target contrast ratio and x is a number from 0 to 1 that represents the scale number. If your scale numbers (like Stripe’s) goes from 0 to 1,000, then a scale number of 500 correlates to x=0.5. Step 2: Calculate lightness based on a target contrast ratio Now that we have a function to calculate a contrast ratio based on our scale number, let’s return to the relative luminance equation: If we solve this equation for L2, we get an equation for the luminosity of a color with the desired contrast ratio with a given color. This is true as long as L1 is greater than L2. Put another way, this covers cases where we’re generating a darker color than our given (background) color. For the opposite case, we can use the same formula, solved for L1 instead of L2. This gets us the following piecewise equation: 0.18, \\ R (Y_b + 0.05) - 0.05 & \text{if } Y_b As explained earlier, the 0.18 in this equation represents the luminosity of “middle gray,” a color equally contrasting with #000000 and #ffffff; Each case depends on whether the background color is dark or light.7 So, for example, if I want a foreground to have a 4.5:1 contrast ratio with the background color, I can calculate the luminosity of that color by inputting the luminosity of the background as Yb and the contrast ratio as 4.5. If the background is #ffffff, which has a luminosity of 1, Yf comes out to 0.183. We can substitute in our function for r(x) to get the following: 0.18, \\ e^{3.04x} (Y_b + 0.05) - 0.05 & \text{if } Y_b This is a function that takes: A number from 0 to 1 that represents a scale number, and The Y value of a background color, and provides the Y value (i.e., luminance) of a color at the given scale number. Step 3: Translate from XYZ Y to OKHsl L Despite its scientific accuracy, XYZ is not a great colorspace to work in for generating color scales for design systems — while we can step through the Y values in a fairly straightforward way, calculating X and Z values of a given color requires matrix multiplication. Instead, we can translate XYZ’s Y value into OKHsl’s l value with the following two-step process: First, we can use the following formula to convert the Y value to the lightness value in lab:8 0.0088564516 \end{cases}$$ Then, OKHsl uses a “toe” function to map the lab lightness value to a perceptually accurate lightness value. Essentially it adds a little space to the dark end of the spectrum. This function is a little complicated: The math gets a lot more manageable if we put it all into a javascript function: const YtoL = Y => { if (Y <= 0.0088564516) { return Y * 903.2962962; } else { return 116 * Math.pow(Y, 1/3) - 16; } } const toe = l => { const k_1 = 0.206 const k_2 = 0.03 const k_3 = (1+k_1)/(1+k_2) return 0.5*(k_3*l - k_1 + Math.sqrt((k_3*l - k_1)*(k_3*l - k_1) + 4*k_2*k_3*l)) } const computeScaleLightness = (scaleValue, backgroundY) => { let foregroundY; if (backgroundY > 0.18) { foregroundY = (backgroundY + 0.05) / Math.exp(3.04 * scaleValue) - 0.05; } else { foregroundY = Math.exp(3.04 * scaleValue) * (backgroundY + 0.05) - 0.05; } return toe(YtoL(foregroundY)); } The function computeScaleLightness takes two values, the normalized scale value and the Y value of your background color, and returns an OKHsl L (lightness) value for the color at that scale step. With this, we have all the pieces we need to generate a complete accessible color palette for any design system. Putting it all together: All the code you need Now we have all the components to write a complete color generation library. // utility functions const YtoL = (Y) => { if (Y <= 0.0088564516) { return Y * 903.2962962; } else { return 116 * Math.pow(Y, 1 / 3) - 16; } }; const toe = (l) => { const k_1 = 0.206; const k_2 = 0.03; const k_3 = (1 + k_1) / (1 + k_2); return ( 0.5 * (k_3 * l - k_1 + Math.sqrt((k_3 * l - k_1) * (k_3 * l - k_1) + 4 * k_2 * k_3 * l)) ); }; const normalizeScaleNumber = (scaleNumber, maxScaleNumber) => scaleNumber / maxScaleNumber; // hue, chroma, and lightness functions const computeScaleHue = (scaleValue, baseHue) => baseHue + 5 * (1 - scaleValue); const computeScaleChroma = (scaleValue, minChroma, maxChroma) => { const chromaDifference = maxChroma - minChroma; return ( -4 * chromaDifference * Math.pow(scaleValue, 2) + 4 * chromaDifference * scaleValue + minChroma ); }; const computeScaleLightness = (scaleValue, backgroundY) => { let foregroundY; if (backgroundY > 0.18) { foregroundY = (backgroundY + 0.05) / Math.exp(3.04 * scaleValue) - 0.05; } else { foregroundY = Math.exp(3.04 * scaleValue) * (backgroundY + 0.05) - 0.05; } return toe(YtoL(foregroundY)); }; // color generator function const computeColorAtScaleNumber = ( scaleNumber, maxScaleNumber, baseHue, minChroma, maxChroma, backgroundY, ) => { // create an OKHsl color object; this might look different depending on what library you use const okhslColor = {}; // normalize scale number const scaleValue = normalizeScaleNumber(scaleNumber, maxScaleNumber); // compute color values okhslColor.h = computeScaleHue(scaleValue, baseHue); okhslColor.s = computeScaleChroma(scaleValue, minChroma, maxChroma); okhslColor.l = computeScaleLightness(scaleValue, backgroundY); // convert OKHsl to sRGB hex; this will look different depending on what library you use return convertToHex(okhslColor); }; For this code to work, you’ll need a library to convert from OKHsl to sRGB hex. The upcoming version of colorjs.io supports this, as does culori. I’ve marked where that matters, in case you’d like to use a different color conversion utility. What does it look like in practice? Here are some examples of the same design in a number of themes, with different background colors: Three generated color palettes By adjusting the hue, chroma, and saturation when we generate our colors, we can get a broad and expressive range of hues, while ensuring each shade is accessible when used in the same context. What we’ve learned and where we’re going At Stripe, we’ve implemented this approach to generating color palettes. It’s now the foundation of the colors in our design system, Sail. The color generation function is also available to the users of our design system; this means that teams can offer theming features to end users, which is especially useful when Stripe’s merchants embed our UI in their own applications. One important lesson I learned while on this journey is the importance of token APIs. This is a bit of an esoteric topic and might be worthy of its own essay. The short version is: Using color aliases (like color.button.background referring to color.action.500 referring to color.base.blue.500) allows theming to happen “behind the scenes,” and ensures that components don’t need to update their code when switching themes. So where do we go from here? There are two features that I’d like to explore in the future to make this approach to color even more robust. First, I’d like to develop an alternative color lightness scale for APCA. The APCA color contrast function is an alternative to the current WCAG contrast ratio function. It purports to more accurately reflect contrast between colors, taking into account the “polarity” of the colors (e.g., dark-on-light or light-on-dark) and the font size of any text. The math behind the APCA contrast function is a bit more complicated than the WCAG function, and my early experiments weren’t very successful. Second, I’d like to extend this approach to work in wide-gamut color spaces like display P3. Currently, OKHsl only covers the sRGB gamut; more and more screens are capable of displaying colors beyond the sRGB gamut, offering even more possibilities for accessible color palettes. Calculating a P3 version of OKHsl should be possible, but it’s definitely outside the scope of my current ability/comprehension. Ultimately, however, the approach outlined in this essay should be a solid basis for generating colors for any design system. No matter how many hues you need, how expressive you’d like to be, how many shades your system consists of, or what kinds of themes you design, the set of functions I’ve covered will provide accessible color combinations. Special thanks to Dmitry Belyaev for providing feedback on a draft of this essay. Footnotes & References (70, -15) is the coordinate for pink in lab colors space. ↩︎ R. W. Pridmore, “Bezold–Brücke Hue-Shift as Functions of Luminance Level, Luminance Ratio, Interstimulus Interval and Adapting White for Aperture and Object Colors,” Vision Research 39, no. 19 (1999): 3873-3891. ↩︎ Jesús Lillo et al., “Lightness and Hue Perception: The Bezold-Brücke Effect and Colour Basic Categories,” Psicológica 25, no. 1 (2004): 23-43. ↩︎ However, it’s important to note that this peak can vary slightly depending on the specific hue in question. ↩︎ AA is generally accepted as the standard for accessibility. A and AAA ratings exist, but are much more lax and more more strict, respectively. You can read more about conformance levels on the W3C website. ↩︎ https://www.w3.org/WAI/GL/wiki/Contrast_ratio ↩︎ This isn’t extremely rigorous; you might want a “light theme” that starts from a dark gray background and gets darker as the scale number increases. I’ll leave that as an exercise to the reader. This formula will cover the typical dark and light mode calculations. ↩︎ If you’re like me and get suspicious when you see oddly specific numbers like 903.2962962 in equations like these, a quick explanation: unlike in the RGB color space, the XYZ color space has no “true white.” Because our eyes can perceive true white differently according to what light source is used, to transfer colors in and out of XYZ color space we often need to also define true white. The most common values are defined by something cryptically called the “CIE standard illuminant D65”, which corresponds roughly to what white looks like on a clear day in northern Europe. I am not making this up. ↩︎
There’s a lot of fear in the air. As AI gets better at design, it’s natural for designers to be worried about their jobs. But I think the question — will AI replace designers? — is a waste of time. Humans have always invented technology to do their work for them and will continue to do so as long as we exist. Let’s use our curiosity and creativity to imagine how technology will help us be better, more efficient, and more impactful. So, in that spirit, I’d like to share a metaphor that I think paints a picture of how the job of design will change in the next decade. Flying by wire Commercial airplanes are some of the most complicated machines humans have ever built. It took all of Wilbur Wright’s skill to fly the first powered airplane for a minute, just 10 feet off the ground. That plane, the Wright Flyer, could carry one person; it weighed 745 pounds with fuel and could reach a height of 30 feet at a maximum speed of 30 miles per hour.1 The Airbus A380, currently the world’s largest commercial airliner, weighs over a million pounds when fully loaded. It can carry up to 853 people, flying up to 43,000 feet at a cruising speed of 561 mph — 85% the speed of sound. Just two people pilot the A380.2 The cockpit of an Airbus A380. Photo by Steve Jurvetson, CC BY 2.0 The A380, and all modern commercial airplanes, wouldn’t exist without something called “fly-by-wire.” Fly-by-wire is a system that translates a pilot’s inputs — changing the throttle to speed up or slow down, controlling pitch and roll with the yoke, turning knobs and dials in the cockpit — into coordinated movements of the airplane’s engines and control surfaces. The first fly-by-wire systems were a veritable nervous system of electric relays and motors; today, they’re sophisticated computers in the belly of the plane. Originally, fly-by-wire had nothing to do with automation. As airplanes got larger, the cables, rods, and hydraulic links connecting the cockpit to the rest of the plane became a monumental design challenge. By replacing those complex, bulky components with electrical wires and switches, airplanes would be lighter and easier to maintain, with more room for passengers and cargo. The first commercial airplane with a fly-by-wire system was the supersonic Concorde jet. At speeds of over Mach 1, it would be almost impossible for a pilot to move the control surfaces of the airplane through sheer mechanical force; fly-by-wire allowed pilots to smoothly operate the plane at any speed. And because sudden changes at top speed could be catastrophic, the fly-by-wire system could use analog circuitry to smooth out a pilot’s inputs. An experimental fly by wire system in the Vought F-8 Crusader using data-processing equipment adapted from the Apollo Guidance Computer As fly-by-wire systems became more common, they went from faithfully transferring pilot’s inputs to interpreting and adjusting them. The Airbus A320, introduced in 1988, featured the first digital (computerized) fly-by-wire system; it included “flight envelope protection,” a system that prevents pilots from taking any action that would cause damage to the airplane. Depending on the speed, altitude, and phase of flight, the fly-by-wire system will ignore certain pilot inputs altogether. Fly-by-wire has been the focus of both scrutiny and praise since its introduction. On one hand, it has saved lives: When US Airways Flight 1549 (an Airbus A320) flew through a flock of birds on takeoff, it lost all power. The pilots had to make an emergency landing in the Hudson river, flying the airplane unusually low and slow, risking putting the plane into an uncontrollable stall. The fly-by-wire system, with its flight envelope protection, ensured the plane could maneuver at the very edge of its capability, leading to a controlled landing with only a few serious injuries for those aboard. On the other hand, fly-by-wire has been criticized for replacing parts of pilots’ expertise. In 2009, Air France Flight 447 (another Airbus A320), crashed in the Atlantic Ocean, killing all 228 passengers and crew. An investigation into the cause of the crash concluded that the autopilot and fly-by-wire protections started to malfunction when ice crystals interfered with the aircraft’s sensors; the pilots, used to flying with the safety of flight envelope protection, couldn’t correct for the errors, stalled the plane, and crashed into the ocean. Whether you think fly-by-wire is a crucial innovation or a crutch, its effect on the airline industry is easy to demonstrate. Bigger planes that fly farther can carry more passengers to more destinations. From 1970 to 2019, the number of airline passengers worldwide has grown over 1,400%, from 310 million to 4.4 billion.3 In the same time period, the number of commercial pilots — pilots holding “commercial” or “airline transport” licenses — has increased 27%, from 208,027 in 1969 to 265,810 in 2019. Pilots’ salaries have stayed consistently high: an average airline captain made about $51,750 a year in 19754, the equivalent to $287,770 in 20235. An airline captain with six years of experience can expect to make $285,460 today.6 Designing by wire Just as fly-by-wire systems have made pilots more efficient (not redundant), AI and automation will make designers more effective. Imagine a design-by-wire system. The job of the designer is to indicate what they want the desired outcome to be. Like a pilot pushing the throttle to make the airplane accelerate, a designer could assemble a wireframe or configure a screen to enable a user to accomplish a task. The design-by-wire system could then interpret the designer’s instructions. The system could change aspects of the design to use the latest design systems components in the correct way. The system could optimize the design to make implementation cheaper, faster, or less prone to bugs. The system could automatically fix accessibility issues, or add information to address accessibility concerns like keyboard shortcuts, screen reader labels, or high-contrast and reduced-motion variations. A designer could list out hypotheses about the design (like “will this convert the most users to paid plans?”), and the system could provide designs for multivariate testing, along with test or research plans. The system could automate QA by testing designs against simulated user behavior, adjusting the designs to cover the wide and unpredictable cases of real user interaction. You can already see these kinds of systems taking shape. Noya promises to take wireframes and turn them into production code using an existing design system. Galileo AI claims to be able to create fully-editable designs from a single text description. Diagram’s Genius aims to provide contextual suggestions in Figma, filling out designs with the click of a button. These are just early tech previews, but they paint a picture of AI becoming a core component of our design tools. At the end of the day, design-by-wire systems are centered around the designer. Like the pilot of an Airbus A380, the designer becomes the operator of a fantastically complicated machine. There is a risk in designing by wire. If designers don’t understand how the system works, they risk losing control, becoming less effective than they were before. That’s why it’s important that we become experts in AI; we don’t have to be able to write the code that drives these tools, but we need to understand the way the systems work. To use the machine to its full potential, the designer has to understand the intricacies of its operation: Pilots, by analogy, train for years in simulators before they step foot in the cockpit of a real airliner. Here’s a few resources you can use to learn more about GPTs, the systems driving the current boom in AI: Stephen Wolfram’s “What is ChatGPT Doing … and Why Does It Work?” is an incredibly in-depth exploration of the technology and concepts, with interactive code. Ted Chiang’s “ChatGPT Is a Blurry JPEG of the Web” uses analogies to explain the strengths and weaknesses of GPT technology. 3Blue1Brown has a 5-part video series explaining the basics of neural networks, including how they are trained. The Coding Train has 26 videos on neural networks, in which Daniel Schiffman builds and explains various components and variations of neural nets. There’s also an accompanying chapter in Schiffman’s book The Nature of Code. All four of the above authors are amazing teachers who mix mathematical depth with intuitive analogies and mental models. Conclusion AI will change our jobs in ways we can’t imagine. This has been happening to airline pilots since the advent of fly-by-wire systems. In the transition to newer, faster, larger, and more efficient airplanes, pilots have needed more and more technical understanding and skill in interacting with the computers that fly their planes. But pilots are still needed. Likewise, designers won’t be replaced; they’ll become operators of increasingly complicated AI-powered machines. New tools will enable designers to be more productive, designing applications and interfaces that can be implemented faster and with less bugs. These tools will expand our brains, helping us cover accessibility and usability concerns that previously took hours of effort from UX specialists and QA engineers. It’ll take years of training to become an expert at designing with these new AI-powered systems. Start now, and you’ll stay ahead of the curve. Wait, and the challenge won’t come from the AI itself; it’ll be other designers — ones who are skilled at AI-powered design — who will come for your job. Footnotes & References “Wright Flyer.” In Wikipedia, February 19, 2023. https://en.wikipedia.org/w/index.php?title=Wright_Flyer&oldid=1140234161. ↩︎ “Airbus A380.” In Wikipedia, February 10, 2023. https://en.wikipedia.org/w/index.php?title=Airbus_A380&oldid=1138648822. ↩︎ “Top 15 Countries with Departures by Air Transport - 1970/2020 -.” Accessed February 20, 2023. https://statisticsanddata.org/data/top-15-countries-with-departures-by-air-transport-1970-2020/. ↩︎ “Industry Wage Survey. Scheduled Airlines.” Scheduled Airlines, Bulletin / Bureau of Labor Statistics, 1977 1972, 2 v. https://catalog.hathitrust.org/Record/009881222. ↩︎ Calculated with https://www.in2013dollars.com/us/inflation/1975?amount=51750 ↩︎ “Major Airline Pilot Salary: First Officer and Captain Pay in 2023 / ATP Flight School.” Accessed February 18, 2023. https://atpflightschool.com/become-a-pilot/airline-career/major-airline-pilot-salary.html. ↩︎
After four months of parental leave, I came back to work and noticed something different. Many of the words, phrases, acronyms, and figures of speech I took for granted no longer held the same meanings. Specifically, the word “platform:” before my time off, I would breeze over it in any presentation or document without a second thought. Now, I find it totally devoid of any meaning. I decided to start sketching out some possible defintions, and quickly uncovered something exciting: platform design isn’t just another flavor of UX or product design. There are challenges, mental models, and requisite skills that set platform design apart. I lead the platform design team;the work that we can do as platform designers gives us extreme leverage in solving user problems. We have the opportunity to make an outsize difference in the experience of our users. What follows is the result of my exploration of what sets platform design apart. If you are a platform designer, too, I hope it gives you an idea of the potential of your work. Interfaces Application designers fit components, modules, containers, and data together to create the best user experience possible. Platform designers additionally focus on the layers between all those elements — the interfaces — to ensure that everything that can be possibly built on the platform has the same high bar of quality. Developers use the concept of an API to plan the way two things fit together (“interface” is the I in API). Good APIs can change the world: TCP/IP enables every computer in the world to communicate with every other. Stripe’s famous “7 lines of code” could connect any piece of software to the complex global banking system. Jeff Bezos’ insistence on API design at Amazon allowed the company to turn its own infrastructure into its most profitable product (AWS). Platform designers have to understand and plan experience APIs: interfaces both in space (how elements appear beside each other, in front of or behind each other, inside or surrounding each other) and in time (how elements or entire screens appear before or after each other, how to communicate causal relationships). With the right interfaces, any application built on the platform will have a great user experience. Incentives Designing applications requires understanding the balance between motivation and friction. It’s almost mathematical: if a user’s motivation is greater than the friction they experience, the user will complete a task. If users aren’t completing a task, you can increase the motivation through marketing and guidance, or decrease the friction through usability improvements or automation. Designing a platform involves the same balancing act, but with an additional variable: whether we intend to or not, we influence the incentives for both application builders and end users. Platform incentives are complex. Take Spotify, for example: they create incentives for artists in the form of paying for each stream of a song. So, in 2014, funk band Vulfpeck released an album on Spotify called Sleepify consisting entirely of 30-second-long silent “songs”. Vulfpeck encouraged fans to play the album on repeat while they slept. After it garnered 5 million streams, Spotify removed the album, but paid the band accordingly; Vulfpeck used the $20,000 it earned to go on tour.1 Incentives are a powerful tool unique to platform design. To create a successful platform, we need to think deeply about all the ways in which incentives can lead to desired outcomes, or ways in which they motivate bad — or, in Vulfpeck’s case, merely mischievous — actors. Emergence Application users usually behave in predictable ways. We can study their tendencies and preferences, then create experiences to fit our observations. Data like task success, retention rate, user sentiment, and engagement tells us how our predictions matched the real world. Over time, by researching and iterating, we come closer to meeting business goals. We can’t always predict platform users’ behavior. The people building applications on the platform and the people using those applications form a feedback loop; both groups develop creative and unexpected ways to use (or abuse) the platform. Think of Twitter users inventing hashtags and @-mentions, or early bulletin board users using punctuation to create emoticons. These ideas spread far beyond their original application, becoming deeply-embedded cross-platform features. This is called emergence. Emergence presents a unique opportunity in design. When behavior is predictable, we design tightly-tuned experiences (“happy paths”) to realize the best outcomes for users. When behavior is emergent, users’ creativity becomes a multiplier on top of our own, exponentially increasing the best outcomes for both users and business. Second-order thinking Application design is all about first-order thinking. A user interacts with the application, and something happens as a result — cause and effect. Causes and effects can be separated by space and time, but their constant tick and tock drives the crucial engagement loop of every successful product. Platform design requires second-order thinking, where first-order effects are causes, too. A great example of this is attributed to Warren Buffett: imagine a crowd watching a parade. A few people stand on their tiptoes — that’s a first-order cause. Now, they can see better — the first-order effect. What happens next? All the people behind them have to stand on their tiptoes, too — that’s the second-order effect. In the end, everyone is worse off, and nobody can see any better.2 Second-order thinking requires creativity. Platform designers have to ask: how will interfaces and incentives create emergent behavior? How will those behaviors change the incentives? What can we build to channel these feedback loops towards our goals? Though these thought exercises will never fully predict the outcomes, without them, a platform is doomed. Case study: Lego Let’s put all these pieces together.3 Lego is one of the most successful toy companies of all time due to their rigorous approach to platform design. Lego bricks are clearly fun to play with, but Lego’s multi-generational success story goes deeper than that — to interfaces, incentives, emergence, and second-order thinking . Interfaces Lego didn’t invent interlocking plastic bricks; Hilary Page, owner of Kidicraft, secured a patent for injection-molded building blocks in 1940. But Lego succeeded, and Kidicraft didn’t. It wasn’t the bricks themselves that mattered. It’s was the way they fit together — their interfaces. The level of precision of a Lego brick is mesmerizing. Manufacturing tolerances (exact to 2 microns, less than the width of a human hair) are so rigorously maintained that every lego brick ever manufactured attaches to every other brick. Pieces fit together and lock tightly, but can be pulled apart easily by a child. Instructions for lego sets also keep children in mind, only using pictures to communicate.4 But the simplicity of the design is misleading: if you have 6 standard lego bricks, you can put them together in 915,103,765 different ways.5 The ingenious design of standard interfaces, together with an ironclad commitment to precise adherence to those standards, is why Lego succeeds. Incentives Lego incentivizes play. Specifically, they encourage builders to (literally) think outside the box, combining pieces from widely different sets to produce new and inventive designs. They do this in both direct and indirect ways, and through both positive and negative interventions: Direct Indirect Positive Sponsored events like Lego Build Day, where builders are encouraged to explore ideas like “take a space ship and rebuild it into a panda bear hammock.” Allowing third parties to sell and resell individual Lego pieces, enabling builders in any part of the world to find any piece they want. Negative Setting clear rules for first-party designers of Lego sets that set the tone for builders: for example, sets can’t require bricks to be assembled in ways that can’t be disassembled later Using trade protections to keep low-quality third-party parts from entering the market, keeping the guarantee of quality for lego bricks high. We can imagine a Lego that allows for parts with unique pieces that don’t fit with others, incentivizing consumers to constantly buy new kits, and locking third parties out of the market. But it’s hard to see how this version of Lego would be as successful. Emergence There are many examples of how Lego users have adapted the system to do things the original designers never intended. In some cases, Lego has even brought those innovations back into the core system. In 1985, a group of computer scientists had a big idea. What if they connected a programming language to physical machines, like Lego creations? Partnering with Lego through the MIT Media Lab, the researchers created Lego Logo, a beginner-friendly robotics programming language that was optimized for creativity. Since then, programmable Legos — “Mindstorms” and “Technics,” offshoots of the original Lego Logo products — have been used to build everything from a Rubik’s Cube-solving robot to customizable prosthetics for children. Lego also encourages emergence in passive ways. Take Brickit, an app created by Leonid Aleksandrov. Snap a photo of your Lego pieces and it’ll show you what you can build, with 3D instructions. You can even scan finished creations and Brickit will reverse-engineer the instructions so you can share your ideas with others. Surprisingly, Lego hasn’t sued Aleksandrov – they’ve allowed Brickit and its community of creators to flourish. Second-order thinking Lego is no stranger to second-order thinking. As a toy company, they are aware of the impact their decisions have on the lives of children and the planet. Specifically, Lego has avoided selling sets that include military equipment or vehicles. No tanks, no bomber planes, no soldiers with guns. The ban extends beyond finished designs — for a long time, the company didn’t even sell gray-colored bricks, the first choice for making military machines. This self-imposed restriction is especially impressive in the light of how profitable war-like toys can be; think of Nerf guns, green army men, and first-person shooter video games.6 Another example of second-order thinking lies in Lego’s plans for the future. Lego bricks are made of ABS, a petroleum-derived plastic. ABS can’t be recycled, meaning Lego produces more than 100,000 metric tons of single-use plastic every year. To reduce their environmental impact, Lego has invested hundreds of millions of dollars in sustainability, resulting in plant-derived plastic flora, and fully recycled paper packaging. And in 2021, after 3 years of research and 250 material tests, a fully recycled brick — made of PET from used water bottles — was ready for mass production.7 tl;dr Interfaces, incentives, emergence, and second-order thinking constitute the biggest differences between platform and application design. Interfaces are the points of contact between elements, where simplicity and flexibility can lead to efficiency at scale. Incentives drive the motivation of both platform- and end- users. By designing incentives, we can re-invest users’ energy, amplifying desired outcomes and preventing undesired results. Emergence is the open-ended feedback loop that platforms can create and maintain. By designing for emergence, not against it, we enable users to discover applications we never imagined. Second-order thinking lets us plan for, and potentially tame, the complexities that threaten to turn platforms into dead ends — or worse. By putting these concepts together, designers can take advantage of the unique leverage platform design provides, efficiently solving thorny problems at scale. Footnotes & References “How Funk Band Vulfpeck Took on Spotify.” CNBC, 2 Apr. 2018, www.cnbc.com/video/2018/04/02/how-funk-band-vulfpeck-took-on-spotify.html. ↩︎ “Second-Order Problem.” Farnam Street, April 7, 2019. https://fs.blog/second-order/. ↩︎ Pun fully intended. ↩︎ The instructions for the largest lego set (The Millennium Falcon) run over 450 pages, all without a single word. ↩︎ Eilers, Søren. “A Lego Counting Problem.” University of Copenhagen, April 7, 2005. https://web.math.ku.dk/~eilers/lego.html. ↩︎ Lendon, Brad. “Lego Won’t Make Modern War Machines, but Others Are Picking up the Pieces.” CNN, December 13, 2020. https://www.cnn.com/style/article/lego-military-sets-intl-hnk-dst/index.html. ↩︎ White, Jeremy. “How Lego Perfected the Recycled Plastic Brick.” Wired, July 11, 2021. https://www.wired.com/story/lego-recycled-plastic-brick/. ↩︎
More in design
For many years, I have maintained a text file called “A Rubric for Website Design Critique.” It is relatively short, but used nonetheless; I’ve returned to it, off and on, for most of my career, referring back, adding things, removing things, adjusting. Its purpose is to standardize how I challenge the work I do and the work I am shown, and even a standard needs maintenance. The documented rubric has five components. Of information architecture, it asks, Is the priority apparent? Does it make sense? Is it actionable? Of layout: Do the visual elements support the architecture? Is the page as scannable as a high-fidelity asset as it was a wireframe? Of accessibility: Is there adequate contrast? Has text been hidden in images? Can a screen reader properly navigate? And of visual language, Is there coherence? Is it consistent? I emphasized documented earlier because it was never complete. The fifth component is art direction, and after that heading in my document is nothing. The file just ends. It isn’t like me to leave something unfinished. I don’t like ragged edges, even when I know they’re natural and sometimes essential. And time and again over the years, I’ve had a chance to wonder at this empty space. Why is it there? Why is it difficult to fill? Perhaps I’m just not the person to fill it. Perhaps that’s where my expertise ends. However, looking again at this unfinished document recently, I have come to a different conclusion. The first four sections — Information Architecture, Layout, Accessibility, Visual Language — are inspection routines. Each one asks a question that has an answer, and the answer can — should — be able to be found by someone who is not me. Art direction is not like that. And that’s why every time I attempted to fill out structured guidance I produced a list of things I did not actually believe… and then deleted them. And so, the section remained empty, which is its own kind of answer, and not a very useful one. Here is a better attempt. Order Is the Floor The first four sections are about order. They ask whether a page is arranged so that it can be seen, perceived, and understood. That is the floor, and a great deal of professional work never gets off it. In fact, the majority of my career has been focused on getting interaction design off the floor. On my team we have run a periodic competition called The Tidiest Designer, where each person submits a composition file for inspection. We look for order, consistency, clarity, and utility. We do not do this because order alone makes design good. We do it because order is what allows good design to happen. Many beautiful, client-applauded comps have been chaotic disasters underneath their presentation modes, and not surprisingly, conflict-inducing when actually produced. Order facilitates that the promise of design becomes its function. But deeper than that, when order is our foundation, we can spend more of our critical energy on the responsible rendering of taste. Section five is that rendering. It is the point at which intent stops being organized and starts being expressed. What follows is not a set of criteria, then, but five places to stand while you look, in the order I tend to look, with a test attached to each that someone else can run. Where a test comes back “I don’t know,” that is a finding. Most designers are intuitive in their creation, which is not a bad thing. But without cross-examining, reinforcing, studying, enriching, and systematizing what begins with our intuition, we end up with something that is meaningful to us and arbitrary to everyone else. This — arbitrariness — is the most common condition of professional design work, and it is nearly invisible from the inside. The Key Every good piece of design has at least one detail that unlocks how the whole thing works. Good designers notice it immediately. Everyone else responds to it without knowing they have. It might be a rule, a crop, a single color used once, a piece of type set deliberately against the grid. Whatever it is, the rest of the composition should be clearing a path for it. A designer on my team once brought me a set of ads for a maker of high-end audio equipment, built around the idea of choice. Two arrows ran in parallel and then diverged, one rendered in color veering off to the left, the other in white, passing it before turning right. The white arrow was the key. It overpowered the bolder colored one simply by pushing further into the space, and its arc carried the eye down to the copy and the call to action. Then I noticed that its curve radius quietly echoed the skewed, rotated “o” in the client’s logotype, and that those two arrows were the only shapes in the entire ad other than text. That last part is the lesson. The key was doing three jobs at once, and everything else had gotten out of its way. The Key Test. Name the key in one sentence. Then say what the composition does to protect it. If nothing on the page is deferring to anything else, there is no key, only assembly. If you can name three, there is also no key, because three keys is zero keys. The Structure Underneath Structure does more work than content while convincing its audience of the opposite. This is the oldest secret in graphic design and painters have known it longest. Mondrian said that every true artist has been inspired more by the beauty of lines and colors and the relationships between them than by the concrete subject of the picture. I adore that because it explains why I can find inspiration in a page of text before I have read a single word. A page held up by its photography is not designed. It is dressed. It is also why I stay in wireframe far longer than most people would think reasonable, finalizing layout with grey boxes and grey lines even when the real material is sitting right there. If it is beautiful on the merits of its structure, it will hold almost any image and almost any text. The Structure Tests. The first one is a classic for graphic designers: Squint until the type turns to grey and the images turn to shapes, and see whether the hierarchy still reads. The other takes a bit more work but, I think is better: Put a grey box where the hero image is and a line of Latin where the headline is. If the design dies, the image was doing the design’s job, and the next round of content will expose it. The Point of View This is the one most design work fails, and it fails in hiding, because nothing is obviously, visually wrong. When we constantly reference existing solutions, our work gravitates toward the mean. We solve for expectations rather than needs. We optimize for recognition rather than revelation. The result is competent and anonymous, and it passes every inspection above. Restraint, on the other hand, is the visible evidence that somebody was directing. It shows up as absence, which makes it hard to credit and easy to skip. The Point of View Test. Put your design beside three others in its category and swap the logos or identifying marks. If this doesn’t break or seriously undermine your work — if your work is that interchangeable — then it has no direction. It is conventional in the truest sense. The harder version of this test is a question you really must ask at various stages of your work: What did I deliberately not do? or What is this not doing? If you cannot answer, then nothing was decided. Such a thing will age at exactly the rate of its category. And because it followed the category’s lead, it will always be behind. What It Is Saying Imagery and type say something before anyone reads a word, and what they say is frequently not what the business does. A few years ago I ran an informal study on a client’s homepage to prove a hunch. They sell technology and expertise to wineries, and they wanted to connect the heritage and craft their customers care about to the stability their technology provides. So they leaned hard on old-style typefaces and historical imagery, to make prospects feel at home. It looked really nice, but I was worried that’s all it did. Traffic was being paid for, and not enough was converting. So, I ran a transient attention test. Participants had eight seconds with the homepage, scrolling but not clicking, and then the page was closed and they were asked what stood out and what the page was for. The page said “commerce technology” and “wine brands” in scannable, plain text. And yet, every participant recalled the imagery instead — an ancient Greco-Roman tapestry — and volunteered words like “history” and “archaeology.” Not one person mentioned wine. Not one mentioned technology. The page was well written. But for its viewers, it was about the wrong thing. The Imagery Test. Give someone outside the project eight seconds to view your design. Afterward, ask what the thing they just saw was — what does the company do? what was the page for? Do not accept a paraphrase of the headline. Ask what the pictures told them. The gap between their answer and the actual business is the size of the art direction problem. Durability Good design is evergreen. The reactions I trust are the ones that survive a week, and the ones I distrust tend to arrive fastest. Anything resting on a technique currently in fashion has a short window before a browser, a platform, or simply everyone else’s adoption closes it. Both of the tests here buy the same thing at different scales: distance. A week of it shows you what belongs to this year. An hour of it shows you what belongs to the last hour of your own looking. The Dated Test. Leave the composition open in a tab and come back to it after a week, even if it has already progressed through reviews, as most things will in that time. Then, name what on it is dated to this year, and ask of each whether it is carrying an idea or just carrying a date. A composition can survive one or two decisions that belong to its moment. It does not survive being made of them. The Interval Test. This one goes after a different fragility, one that lives in your read of the work rather than in the work itself. Clutter accumulates precisely because the eye that added it has stopped seeing it. I have always found that coming back to a finished but unshared design after even just a few hours away, sometimes minutes, has resulted in needed editorial moves. What you have been staring at is porous to every other thing held on your screen or in your recent memory, and your working brain is an unwitting cheat. Breaks expose that immediately. Take enough of them and the work stops absorbing its surroundings. Preference and Judgment Taste is that combination of preference, personality, and perceived novelty that lets an observer tell your work from someone else’s. It belongs in the work. It does not belong in the verdict. I have sat in too many reviews where a real critique was offered, understood, and then dissolved by “well, we like it.” That is nice. But who cares if you like it? Does it do what it is supposed to do? Or is it possible that the things you like about it get in the way? The way through is not to suppress the reaction but to keep going after it. Name what you are responding to, then say what it is doing for the work. If it is doing nothing for the work, you have found a preference. If it is doing something, you have found a judgment, and now you have to justify it, which is the only part of design that has ever been hard. To make it somewhat easier, do not defend it. Sell it. Don’t believe the lie that “good design just works” as if it will be self-evident in the eye of the beholder and embraced without question. Nothing could be further from the truth. Good design often requires advocacy. Every rubric wants to become an inspection. In art school you always knew a critique was going nowhere when someone would ummm and ahhh, approach the piece, back away from it, approach it again, and finally ask, “is this, ummm, is this balsa wood?” They just had to say something, and what a thing is made of was the best they could do. The digital equivalent is talking about the canvas, the type foundry, the plugins, or opening the inspector. None of those are relevant to assessing a design’s quality. Sections one through four can be inspected. Section five has to be seen — by you first, and yet, outside of yourself — which takes time and, more importantly, conviction. I do not think that section five will ever be as short as the others, or as portable. It takes longer to run than all four of them combined. For years I read that as a defect in my system. But lately I have started to suspect it is the only part of the rubric that will still be worth anything in a few years, because production is becoming generative and design is not. Which leaves me somewhere I have not settled. The first four sections are the ones a machine can already run. The fifth is the one it cannot, so the fifth is where the work is going. But the fifth is also the one nobody has ever managed to teach quickly. I do not yet know whether that is a problem to solve or a fact to accept. Better yet, maybe it’s a distant horizon to embrace, because it means we have somewhere left to go. P.S. I have left creative direction out of this entirely, which is a cheat. In my own notes it sits above art direction, closer to the conceptual end of the spectrum that runs down through graphic design to the mechanics of a build. That is a different piece, and I suspect a harder one.
Four years ago, I wrote “How to pick the least wrong colors.” The gist is: picking a categorical color palette is an optimization problem. There’s no such thing as the right colors. But if you use the right cost function, and the right kind of hill climbing, you can at least get the least wrong ones. Since the original post I’ve been slowly picking away at improvements and new approaches. Now that we’re past the singularity, I’ve put a few coding robots on the job. It’s reassuring that many of my assumptions were good ones! The robots have been able to improve the code, bridging some of the gaps in my own knowledge. Today, I’m publishing an updated version of the algorithm as an npm package, along with a fancy GUI version. While there’s still more to do, I’m proud of how far I’ve been able to take it. What’s new New evaluators More controls The public API and a CLI What’s improved The annealing algorithm Configurable color space and distance metric The results One more thing Acknowledgements What’s new New evaluators Almost as soon as I published the first version, I realized that the cost function lends itself really well to modularity. Beyond my initial evaluation functions, I could design new ones, and provide a framework for anyone to plug in their own. As a recap, my original criteria for good categorical colors, mapped to evaluation functions: Similarity — a way of measuring the similarity of one palette to another, useful for providing art direction and getting brand alignment Energy — the colors should be different from each other so they aren’t liable to be confused from one another Range — the differences between the colors should be consistent so unintended groupings don’t appear Color vision deficiency — simulating the colors under different types of color blindness (red-green, blue-yellow, partial to full tritanopia) Here’s the new evaluators: JND — strongly reject palettes that have two or more colors that are too similar Avoid — the mirror image of the similarity evaluation, push colors away from a user-defined set Contrast — compares colors, keeping them above the WCAG AA color contrast floor. Can be used with a background color to maintain contrast on a chart’s background Saliency — uses color naming study data to prefer colors that are easy to name Name difference — the mirror image of saliency, avoiding colors that share names Each of these evaluators can be weighted, indicating the kinds of tradeoffs and priorities you’d like for your color palette. Additionally, the whole evaluator system is pluggable: you can define your own evaluators and have them drive the optimizer! More controls Colors can now be fixed in place, or pinned to a particular order, making it easier to load in existing palettes and optimize all or just some of the colors. Individual channels of each color can be locked, too, meaning you can keep the saturation or hue of a color fixed while optimizing its lightness. This works in any color space. The public API and a CLI The whole package is now a proper library, with a public API. This means: 1. the whole thing is now distributable through npm, with proper versioning, 2. there’s a CLI, making it much more ergonomic for both humans and agents. The API allows for full configuration of the algorithm, as well as loading in colors to optimize. Output can be in raw color values, CSS properties, or DTCG JSON. There’s also a new reportJndIssues endpoint that allows you to evaluate palettes without optimizing them, which is useful to compare a generated palette to commonly-used ones (like Observable, d3, IBM Carbon, and more). What’s improved The annealing algorithm When I wrote the initial algorithm in 2022, I had just learned about simulated annealing. I’ll be honest: I don’t know much more today than I did then. But with AI-assisted research, I was able to solve some questions I had about the initial implementation. Now, the algorithm picks the correct starting temperature based on some random initial samples. Mutation also happens in a scaled manner, so colors change less towards the end of the optimization schedule. Iterations can be capped to prevent very long runs, and the whole thing is much, much more performant. Configurable color space and distance metric The first version of the algorithm worked in RGB space. Now, it defaults to okhsl, but even this is configurable. Individual channels can be constrained to dial in the palette’s boundaries. Also, you can choose which color distance metric you’d like to use (but the library uses CIEDE2000 by default). This flexibility is powered largely by a move from chroma.js to culori. I’ve learned a ton about color spaces since 2022, so being able to mix and match color spaces with distance metrics has been extremely useful. The results The category-colors library reliably produces better results than other palette-generating tools and industry-standard color palettes. Compared to other palette-generating tools, category-colors has more control. Palettailor, for example, optimizes for pure color difference, without accounting for color vision deficiency. QualPal brings some of the optimization parameters, but doesn’t allow for steering towards or away from arbitrary colors. Scores at 8 colors ΔEMinimum ΔEworst of CVD Name differenceMinimum Uniformitylower is better category-colors 22.6 ±1.6 13.7 ±2.0 0.35 ±0.14 best in column 0.30 ±0.02 best in column QualPal 1.1.0 24.7 21.8 best in column 0.10 0.44 Palettailor 26.6 ±2.4 best in column 4.5 ±1.7 0.34 ±0.16 0.34 ±0.04 Colorgorical 15.8 ±3.3 4.1 ±1.6 0.09 ±0.06 0.42 ±0.03 All numbers are at 8 colors. Rows with ± are mean ± standard deviation over 10 palettes; rows without are deterministic and produce one palette. category-colors and Palettailor are 10 independent runs on the same seeds; Colorgorical’s row is 10 palettes from its authors’ own sampling script at equal criterion weights. QualPal was run with CVD on, matched bounds, and takes no seed. Name difference is Heer & Stone’s 1 − cosine; Colorgorical’s own interface reports a Hellinger distance instead. Shaded cells are the best value in their column. Compared to industry-standard palettes, category-colors can produce more optimal palettes, especially at high cardinality. Scores at 8 colors ΔEMinimum ΔEworst of CVD Name differenceMinimum Uniformitylower is better category-colors 22.6 ±1.6 best in column 13.7 ±2.0 best in column 0.35 ±0.14 0.30 ±0.02 best in column Okabe–Ito 21.3 8.8 0.06 0.34 Observable 10 18.4 0.6 0.40 0.34 Tableau 10 18.1 3.2 0.24 0.32 d3 category10 16.2 1.6 0.84 best in column 0.40 ColorBrewer Set3 13.7 1.9 0.16 0.32 IBM Carbon 12.8 5.0 0.11 0.34 Same run: 8 colors, 10 trials. Reference palettes are deterministic, so they're single values. Shaded cells are the best value in their column. One more thing I’ve built a UI that consumes the package and makes it easy to generate and optimize palettes. This has been the biggest request since I published the initial essay, so it’s the thing I’m excited to share. It’s ridiculously overengineered, but hey, what else are personal projects for? Acknowledgements Many measurements come from published research: Gaurav Sharma, Wencheng Wu and Edul Dalal for CIEDE2000; Gustavo Machado, Manuel Oliveira and Leandro Fernandes for the color vision deficiency simulation; Maureen Stone, Danielle Albers Szafir and Vidya Setlur for the size-dependent just-noticeable-difference result; Jeffrey Heer and Maureen Stone, whose color naming models and the c3 data from the Stanford Visualization Group power both the saliency and name-difference evaluators. Existing palettes: Masataka Okabe and Kei Ito’s Color Universal Design set; Matthew Petroff’s sequences; and Mark Harrower and Cynthia Brewer’s ColorBrewer. Other generators laid a lot of the groundwork: Kecheng Lu and colleagues (Palettailor), Connor Gramazio, David Laidlaw and Karen Schloss (Colorgorical), Johan Larsson (QualPal), and Chin Tseng, Arran Zeyu Wang, Ghulam Jilani Quadri and Danielle Albers Szafir (CatPAW). Andrew McNutt, Maureen Stone and Jeffrey Heer’s color-buddy has also been indispensable. Finally, Dan Burzo’s culori made it easy to make this library colorspace-agnostic.
This is part of a new experiment I started in an effort to document the process of making Niche design.
Weekly curated resources for designers — thinkers and makers.
Users parse a layout before they read its labels. Whitespace, borders, alignment, color, and motion determine what belongs together. When these cues fight the content, users attach the label, price, warning, status, or action to the wrong object. Proximity, similarity, enclosure, and the other Gestalt cues guide the eye, snapping visual chaos into clarity.