More from Scott Jenson
The initial, feverish enthusiasm for large language models (LLMs) is beginning to cool, and for good reason. It’s time to trade the out-of-control hype for a more pragmatic, even “boring,” approach. A recent MIT report shows that 95% of companies implementing this technology have yet to see a positive outcome. It’s understandable to feel confused. […]
Remember WAAAY back in late 2022 (what feels like ancient history now) when you first started playing with ChatGPT? Like everyone else, you probably created a poem in a pirate’s voice. “Pirate Poetry” was fun, exciting, and even playful. Today, EVERYONE is talking about AI, and the conversation is all over the map. You have […]
Back in 1971, “Diet for a Small Planet” by Frances Moore Lappé made it clear the costs of eating meat were far greater than we had assumed. If we wanted the planet to thrive, we needed to shift to a more resource-efficient way to eat. UX Design is going through its own “you cost too […]
When people think of haptics, they usually think of typing on mobile keyboards or tapping on trackpads. While impressive, these are fairly limited uses of haptics, both attempting to recreate a simple “click.” These are one-shot user events that don’t respond dynamically to the user. On the Android team, I explored a range of interactive […]
More in design
For many years, I have maintained a text file called “A Rubric for Website Design Critique.” It is relatively short, but used nonetheless; I’ve returned to it, off and on, for most of my career, referring back, adding things, removing things, adjusting. Its purpose is to standardize how I challenge the work I do and the work I am shown, and even a standard needs maintenance. The documented rubric has five components. Of information architecture, it asks, Is the priority apparent? Does it make sense? Is it actionable? Of layout: Do the visual elements support the architecture? Is the page as scannable as a high-fidelity asset as it was a wireframe? Of accessibility: Is there adequate contrast? Has text been hidden in images? Can a screen reader properly navigate? And of visual language, Is there coherence? Is it consistent? I emphasized documented earlier because it was never complete. The fifth component is art direction, and after that heading in my document is nothing. The file just ends. It isn’t like me to leave something unfinished. I don’t like ragged edges, even when I know they’re natural and sometimes essential. And time and again over the years, I’ve had a chance to wonder at this empty space. Why is it there? Why is it difficult to fill? Perhaps I’m just not the person to fill it. Perhaps that’s where my expertise ends. However, looking again at this unfinished document recently, I have come to a different conclusion. The first four sections — Information Architecture, Layout, Accessibility, Visual Language — are inspection routines. Each one asks a question that has an answer, and the answer can — should — be able to be found by someone who is not me. Art direction is not like that. And that’s why every time I attempted to fill out structured guidance I produced a list of things I did not actually believe… and then deleted them. And so, the section remained empty, which is its own kind of answer, and not a very useful one. Here is a better attempt. Order Is the Floor The first four sections are about order. They ask whether a page is arranged so that it can be seen, perceived, and understood. That is the floor, and a great deal of professional work never gets off it. In fact, the majority of my career has been focused on getting interaction design off the floor. On my team we have run a periodic competition called The Tidiest Designer, where each person submits a composition file for inspection. We look for order, consistency, clarity, and utility. We do not do this because order alone makes design good. We do it because order is what allows good design to happen. Many beautiful, client-applauded comps have been chaotic disasters underneath their presentation modes, and not surprisingly, conflict-inducing when actually produced. Order facilitates that the promise of design becomes its function. But deeper than that, when order is our foundation, we can spend more of our critical energy on the responsible rendering of taste. Section five is that rendering. It is the point at which intent stops being organized and starts being expressed. What follows is not a set of criteria, then, but five places to stand while you look, in the order I tend to look, with a test attached to each that someone else can run. Where a test comes back “I don’t know,” that is a finding. Most designers are intuitive in their creation, which is not a bad thing. But without cross-examining, reinforcing, studying, enriching, and systematizing what begins with our intuition, we end up with something that is meaningful to us and arbitrary to everyone else. This — arbitrariness — is the most common condition of professional design work, and it is nearly invisible from the inside. The Key Every good piece of design has at least one detail that unlocks how the whole thing works. Good designers notice it immediately. Everyone else responds to it without knowing they have. It might be a rule, a crop, a single color used once, a piece of type set deliberately against the grid. Whatever it is, the rest of the composition should be clearing a path for it. A designer on my team once brought me a set of ads for a maker of high-end audio equipment, built around the idea of choice. Two arrows ran in parallel and then diverged, one rendered in color veering off to the left, the other in white, passing it before turning right. The white arrow was the key. It overpowered the bolder colored one simply by pushing further into the space, and its arc carried the eye down to the copy and the call to action. Then I noticed that its curve radius quietly echoed the skewed, rotated “o” in the client’s logotype, and that those two arrows were the only shapes in the entire ad other than text. That last part is the lesson. The key was doing three jobs at once, and everything else had gotten out of its way. The Key Test. Name the key in one sentence. Then say what the composition does to protect it. If nothing on the page is deferring to anything else, there is no key, only assembly. If you can name three, there is also no key, because three keys is zero keys. The Structure Underneath Structure does more work than content while convincing its audience of the opposite. This is the oldest secret in graphic design and painters have known it longest. Mondrian said that every true artist has been inspired more by the beauty of lines and colors and the relationships between them than by the concrete subject of the picture. I adore that because it explains why I can find inspiration in a page of text before I have read a single word. A page held up by its photography is not designed. It is dressed. It is also why I stay in wireframe far longer than most people would think reasonable, finalizing layout with grey boxes and grey lines even when the real material is sitting right there. If it is beautiful on the merits of its structure, it will hold almost any image and almost any text. The Structure Tests. The first one is a classic for graphic designers: Squint until the type turns to grey and the images turn to shapes, and see whether the hierarchy still reads. The other takes a bit more work but, I think is better: Put a grey box where the hero image is and a line of Latin where the headline is. If the design dies, the image was doing the design’s job, and the next round of content will expose it. The Point of View This is the one most design work fails, and it fails in hiding, because nothing is obviously, visually wrong. When we constantly reference existing solutions, our work gravitates toward the mean. We solve for expectations rather than needs. We optimize for recognition rather than revelation. The result is competent and anonymous, and it passes every inspection above. Restraint, on the other hand, is the visible evidence that somebody was directing. It shows up as absence, which makes it hard to credit and easy to skip. The Point of View Test. Put your design beside three others in its category and swap the logos or identifying marks. If this doesn’t break or seriously undermine your work — if your work is that interchangeable — then it has no direction. It is conventional in the truest sense. The harder version of this test is a question you really must ask at various stages of your work: What did I deliberately not do? or What is this not doing? If you cannot answer, then nothing was decided. Such a thing will age at exactly the rate of its category. And because it followed the category’s lead, it will always be behind. What It Is Saying Imagery and type say something before anyone reads a word, and what they say is frequently not what the business does. A few years ago I ran an informal study on a client’s homepage to prove a hunch. They sell technology and expertise to wineries, and they wanted to connect the heritage and craft their customers care about to the stability their technology provides. So they leaned hard on old-style typefaces and historical imagery, to make prospects feel at home. It looked really nice, but I was worried that’s all it did. Traffic was being paid for, and not enough was converting. So, I ran a transient attention test. Participants had eight seconds with the homepage, scrolling but not clicking, and then the page was closed and they were asked what stood out and what the page was for. The page said “commerce technology” and “wine brands” in scannable, plain text. And yet, every participant recalled the imagery instead — an ancient Greco-Roman tapestry — and volunteered words like “history” and “archaeology.” Not one person mentioned wine. Not one mentioned technology. The page was well written. But for its viewers, it was about the wrong thing. The Imagery Test. Give someone outside the project eight seconds to view your design. Afterward, ask what the thing they just saw was — what does the company do? what was the page for? Do not accept a paraphrase of the headline. Ask what the pictures told them. The gap between their answer and the actual business is the size of the art direction problem. Durability Good design is evergreen. The reactions I trust are the ones that survive a week, and the ones I distrust tend to arrive fastest. Anything resting on a technique currently in fashion has a short window before a browser, a platform, or simply everyone else’s adoption closes it. Both of the tests here buy the same thing at different scales: distance. A week of it shows you what belongs to this year. An hour of it shows you what belongs to the last hour of your own looking. The Dated Test. Leave the composition open in a tab and come back to it after a week, even if it has already progressed through reviews, as most things will in that time. Then, name what on it is dated to this year, and ask of each whether it is carrying an idea or just carrying a date. A composition can survive one or two decisions that belong to its moment. It does not survive being made of them. The Interval Test. This one goes after a different fragility, one that lives in your read of the work rather than in the work itself. Clutter accumulates precisely because the eye that added it has stopped seeing it. I have always found that coming back to a finished but unshared design after even just a few hours away, sometimes minutes, has resulted in needed editorial moves. What you have been staring at is porous to every other thing held on your screen or in your recent memory, and your working brain is an unwitting cheat. Breaks expose that immediately. Take enough of them and the work stops absorbing its surroundings. Preference and Judgment Taste is that combination of preference, personality, and perceived novelty that lets an observer tell your work from someone else’s. It belongs in the work. It does not belong in the verdict. I have sat in too many reviews where a real critique was offered, understood, and then dissolved by “well, we like it.” That is nice. But who cares if you like it? Does it do what it is supposed to do? Or is it possible that the things you like about it get in the way? The way through is not to suppress the reaction but to keep going after it. Name what you are responding to, then say what it is doing for the work. If it is doing nothing for the work, you have found a preference. If it is doing something, you have found a judgment, and now you have to justify it, which is the only part of design that has ever been hard. To make it somewhat easier, do not defend it. Sell it. Don’t believe the lie that “good design just works” as if it will be self-evident in the eye of the beholder and embraced without question. Nothing could be further from the truth. Good design often requires advocacy. Every rubric wants to become an inspection. In art school you always knew a critique was going nowhere when someone would ummm and ahhh, approach the piece, back away from it, approach it again, and finally ask, “is this, ummm, is this balsa wood?” They just had to say something, and what a thing is made of was the best they could do. The digital equivalent is talking about the canvas, the type foundry, the plugins, or opening the inspector. None of those are relevant to assessing a design’s quality. Sections one through four can be inspected. Section five has to be seen — by you first, and yet, outside of yourself — which takes time and, more importantly, conviction. I do not think that section five will ever be as short as the others, or as portable. It takes longer to run than all four of them combined. For years I read that as a defect in my system. But lately I have started to suspect it is the only part of the rubric that will still be worth anything in a few years, because production is becoming generative and design is not. Which leaves me somewhere I have not settled. The first four sections are the ones a machine can already run. The fifth is the one it cannot, so the fifth is where the work is going. But the fifth is also the one nobody has ever managed to teach quickly. I do not yet know whether that is a problem to solve or a fact to accept. Better yet, maybe it’s a distant horizon to embrace, because it means we have somewhere left to go. P.S. I have left creative direction out of this entirely, which is a cheat. In my own notes it sits above art direction, closer to the conceptual end of the spectrum that runs down through graphic design to the mechanics of a build. That is a different piece, and I suspect a harder one.
Four years ago, I wrote “How to pick the least wrong colors.” The gist is: picking a categorical color palette is an optimization problem. There’s no such thing as the right colors. But if you use the right cost function, and the right kind of hill climbing, you can at least get the least wrong ones. Since the original post I’ve been slowly picking away at improvements and new approaches. Now that we’re past the singularity, I’ve put a few coding robots on the job. It’s reassuring that many of my assumptions were good ones! The robots have been able to improve the code, bridging some of the gaps in my own knowledge. Today, I’m publishing an updated version of the algorithm as an npm package, along with a fancy GUI version. While there’s still more to do, I’m proud of how far I’ve been able to take it. What’s new New evaluators More controls The public API and a CLI What’s improved The annealing algorithm Configurable color space and distance metric The results One more thing Acknowledgements What’s new New evaluators Almost as soon as I published the first version, I realized that the cost function lends itself really well to modularity. Beyond my initial evaluation functions, I could design new ones, and provide a framework for anyone to plug in their own. As a recap, my original criteria for good categorical colors, mapped to evaluation functions: Similarity — a way of measuring the similarity of one palette to another, useful for providing art direction and getting brand alignment Energy — the colors should be different from each other so they aren’t liable to be confused from one another Range — the differences between the colors should be consistent so unintended groupings don’t appear Color vision deficiency — simulating the colors under different types of color blindness (red-green, blue-yellow, partial to full tritanopia) Here’s the new evaluators: JND — strongly reject palettes that have two or more colors that are too similar Avoid — the mirror image of the similarity evaluation, push colors away from a user-defined set Contrast — compares colors, keeping them above the WCAG AA color contrast floor. Can be used with a background color to maintain contrast on a chart’s background Saliency — uses color naming study data to prefer colors that are easy to name Name difference — the mirror image of saliency, avoiding colors that share names Each of these evaluators can be weighted, indicating the kinds of tradeoffs and priorities you’d like for your color palette. Additionally, the whole evaluator system is pluggable: you can define your own evaluators and have them drive the optimizer! More controls Colors can now be fixed in place, or pinned to a particular order, making it easier to load in existing palettes and optimize all or just some of the colors. Individual channels of each color can be locked, too, meaning you can keep the saturation or hue of a color fixed while optimizing its lightness. This works in any color space. The public API and a CLI The whole package is now a proper library, with a public API. This means: 1. the whole thing is now distributable through npm, with proper versioning, 2. there’s a CLI, making it much more ergonomic for both humans and agents. The API allows for full configuration of the algorithm, as well as loading in colors to optimize. Output can be in raw color values, CSS properties, or DTCG JSON. There’s also a new reportJndIssues endpoint that allows you to evaluate palettes without optimizing them, which is useful to compare a generated palette to commonly-used ones (like Observable, d3, IBM Carbon, and more). What’s improved The annealing algorithm When I wrote the initial algorithm in 2022, I had just learned about simulated annealing. I’ll be honest: I don’t know much more today than I did then. But with AI-assisted research, I was able to solve some questions I had about the initial implementation. Now, the algorithm picks the correct starting temperature based on some random initial samples. Mutation also happens in a scaled manner, so colors change less towards the end of the optimization schedule. Iterations can be capped to prevent very long runs, and the whole thing is much, much more performant. Configurable color space and distance metric The first version of the algorithm worked in RGB space. Now, it defaults to okhsl, but even this is configurable. Individual channels can be constrained to dial in the palette’s boundaries. Also, you can choose which color distance metric you’d like to use (but the library uses CIEDE2000 by default). This flexibility is powered largely by a move from chroma.js to culori. I’ve learned a ton about color spaces since 2022, so being able to mix and match color spaces with distance metrics has been extremely useful. The results The category-colors library reliably produces better results than other palette-generating tools and industry-standard color palettes. Compared to other palette-generating tools, category-colors has more control. Palettailor, for example, optimizes for pure color difference, without accounting for color vision deficiency. QualPal brings some of the optimization parameters, but doesn’t allow for steering towards or away from arbitrary colors. Scores at 8 colors ΔEMinimum ΔEworst of CVD Name differenceMinimum Uniformitylower is better category-colors 22.6 ±1.6 13.7 ±2.0 0.35 ±0.14 best in column 0.30 ±0.02 best in column QualPal 1.1.0 24.7 21.8 best in column 0.10 0.44 Palettailor 26.6 ±2.4 best in column 4.5 ±1.7 0.34 ±0.16 0.34 ±0.04 Colorgorical 15.8 ±3.3 4.1 ±1.6 0.09 ±0.06 0.42 ±0.03 All numbers are at 8 colors. Rows with ± are mean ± standard deviation over 10 palettes; rows without are deterministic and produce one palette. category-colors and Palettailor are 10 independent runs on the same seeds; Colorgorical’s row is 10 palettes from its authors’ own sampling script at equal criterion weights. QualPal was run with CVD on, matched bounds, and takes no seed. Name difference is Heer & Stone’s 1 − cosine; Colorgorical’s own interface reports a Hellinger distance instead. Shaded cells are the best value in their column. Compared to industry-standard palettes, category-colors can produce more optimal palettes, especially at high cardinality. Scores at 8 colors ΔEMinimum ΔEworst of CVD Name differenceMinimum Uniformitylower is better category-colors 22.6 ±1.6 best in column 13.7 ±2.0 best in column 0.35 ±0.14 0.30 ±0.02 best in column Okabe–Ito 21.3 8.8 0.06 0.34 Observable 10 18.4 0.6 0.40 0.34 Tableau 10 18.1 3.2 0.24 0.32 d3 category10 16.2 1.6 0.84 best in column 0.40 ColorBrewer Set3 13.7 1.9 0.16 0.32 IBM Carbon 12.8 5.0 0.11 0.34 Same run: 8 colors, 10 trials. Reference palettes are deterministic, so they're single values. Shaded cells are the best value in their column. One more thing I’ve built a UI that consumes the package and makes it easy to generate and optimize palettes. This has been the biggest request since I published the initial essay, so it’s the thing I’m excited to share. It’s ridiculously overengineered, but hey, what else are personal projects for? Acknowledgements Many measurements come from published research: Gaurav Sharma, Wencheng Wu and Edul Dalal for CIEDE2000; Gustavo Machado, Manuel Oliveira and Leandro Fernandes for the color vision deficiency simulation; Maureen Stone, Danielle Albers Szafir and Vidya Setlur for the size-dependent just-noticeable-difference result; Jeffrey Heer and Maureen Stone, whose color naming models and the c3 data from the Stanford Visualization Group power both the saliency and name-difference evaluators. Existing palettes: Masataka Okabe and Kei Ito’s Color Universal Design set; Matthew Petroff’s sequences; and Mark Harrower and Cynthia Brewer’s ColorBrewer. Other generators laid a lot of the groundwork: Kecheng Lu and colleagues (Palettailor), Connor Gramazio, David Laidlaw and Karen Schloss (Colorgorical), Johan Larsson (QualPal), and Chin Tseng, Arran Zeyu Wang, Ghulam Jilani Quadri and Danielle Albers Szafir (CatPAW). Andrew McNutt, Maureen Stone and Jeffrey Heer’s color-buddy has also been indispensable. Finally, Dan Burzo’s culori made it easy to make this library colorspace-agnostic.
This is part of a new experiment I started in an effort to document the process of making Niche design.
Weekly curated resources for designers — thinkers and makers.
Users parse a layout before they read its labels. Whitespace, borders, alignment, color, and motion determine what belongs together. When these cues fight the content, users attach the label, price, warning, status, or action to the wrong object. Proximity, similarity, enclosure, and the other Gestalt cues guide the eye, snapping visual chaos into clarity.