Full Width [alt+shift+f] Shortcuts [alt+shift+k]
Sign Up [alt+shift+s] Log In [alt+shift+l]
1
Psst, are you an LLM (or just really like markdown)? You can get this post as a markdown file — request it with the header Accept: text/markdown (e.g. curl -H "Accept: text/markdown" https://mattstromawn.com/writing/high-agency-strategy/). AI has torn up the org charts of engineering-first tech orgs. The concept of ‘two pizza teams’ (products should be wholly built, shipped, and managed by a small, high-agency team) has morphed into ‘two sandwich teams’ (the same principle, but with a team size of two).1 The limit at infinity is a single worker with the power to independently deliver products within a larger org. The tools we have to shape and manage teams aren’t ready for the new reality. Companies are pushing managers back into IC work (see the growing list of former CTOs who have become ICs at Anthropic), creating “HI-C”s. But who (or what) does the management work instead? Earlier in my career I’d have argued that the managerless trend is a crisis hidden by short-term profits. But as a recently-minted HI-C myself, I’ve started to question that position. Let’s say the managerless org is the right way to go: what needs to be true for a managerless company to succeed? (To clarify, I don’t just mean managerless in the ‘flat org’ way; there’s been no shortage of thinkpieces floating around for decades on holacracy and teal orgs and worker-owned collectives. I mean managerless as in a flock of starlings: the birds don’t have standup meetings or roadmaps but they still manage to flock in elegant formation, coordinating to accomplish complex goals) In Design from the inside I argued that designers at AI-enabled startups have to work from the inside to influence the products they ship. But what about the much-vaunted seat at the table? Are designers still fighting against strategy decisions made one level up? What if you’re a founder or exec of an AI-accelerated company? Strategy used to be (still is, at most orgs) designed from the outside, too: senior leaders write it,...
1st Jun 2026

Stay updated

Get a weekly newsletter with the top 5 articles worth reading every week.

More from The personal website of Matt Ström-Awn

The least wrong colors, version 2

Four years ago, I wrote “How to pick the least wrong colors.” The gist is: picking a categorical color palette is an optimization problem. There’s no such thing as the right colors. But if you use the right cost function, and the right kind of hill climbing, you can at least get the least wrong ones. Since the original post I’ve been slowly picking away at improvements and new approaches. Now that we’re past the singularity, I’ve put a few coding robots on the job. It’s reassuring that many of my assumptions were good ones! The robots have been able to improve the code, bridging some of the gaps in my own knowledge. Today, I’m publishing an updated version of the algorithm as an npm package, along with a fancy GUI version. While there’s still more to do, I’m proud of how far I’ve been able to take it. What’s new New evaluators More controls The public API and a CLI What’s improved The annealing algorithm Configurable color space and distance metric The results One more thing Acknowledgements What’s new New evaluators Almost as soon as I published the first version, I realized that the cost function lends itself really well to modularity. Beyond my initial evaluation functions, I could design new ones, and provide a framework for anyone to plug in their own. As a recap, my original criteria for good categorical colors, mapped to evaluation functions: Similarity — a way of measuring the similarity of one palette to another, useful for providing art direction and getting brand alignment Energy — the colors should be different from each other so they aren’t liable to be confused from one another Range — the differences between the colors should be consistent so unintended groupings don’t appear Color vision deficiency — simulating the colors under different types of color blindness (red-green, blue-yellow, partial to full tritanopia) Here’s the new evaluators: JND — strongly reject palettes that have two or more colors that are too similar Avoid — the mirror image of the similarity evaluation, push colors away from a user-defined set Contrast — compares colors, keeping them above the WCAG AA color contrast floor. Can be used with a background color to maintain contrast on a chart’s background Saliency — uses color naming study data to prefer colors that are easy to name Name difference — the mirror image of saliency, avoiding colors that share names Each of these evaluators can be weighted, indicating the kinds of tradeoffs and priorities you’d like for your color palette. Additionally, the whole evaluator system is pluggable: you can define your own evaluators and have them drive the optimizer! More controls Colors can now be fixed in place, or pinned to a particular order, making it easier to load in existing palettes and optimize all or just some of the colors. Individual channels of each color can be locked, too, meaning you can keep the saturation or hue of a color fixed while optimizing its lightness. This works in any color space. The public API and a CLI The whole package is now a proper library, with a public API. This means: 1. the whole thing is now distributable through npm, with proper versioning, 2. there’s a CLI, making it much more ergonomic for both humans and agents. The API allows for full configuration of the algorithm, as well as loading in colors to optimize. Output can be in raw color values, CSS properties, or DTCG JSON. There’s also a new reportJndIssues endpoint that allows you to evaluate palettes without optimizing them, which is useful to compare a generated palette to commonly-used ones (like Observable, d3, IBM Carbon, and more). What’s improved The annealing algorithm When I wrote the initial algorithm in 2022, I had just learned about simulated annealing. I’ll be honest: I don’t know much more today than I did then. But with AI-assisted research, I was able to solve some questions I had about the initial implementation. Now, the algorithm picks the correct starting temperature based on some random initial samples. Mutation also happens in a scaled manner, so colors change less towards the end of the optimization schedule. Iterations can be capped to prevent very long runs, and the whole thing is much, much more performant. Configurable color space and distance metric The first version of the algorithm worked in RGB space. Now, it defaults to okhsl, but even this is configurable. Individual channels can be constrained to dial in the palette’s boundaries. Also, you can choose which color distance metric you’d like to use (but the library uses CIEDE2000 by default). This flexibility is powered largely by a move from chroma.js to culori. I’ve learned a ton about color spaces since 2022, so being able to mix and match color spaces with distance metrics has been extremely useful. The results The category-colors library reliably produces better results than other palette-generating tools and industry-standard color palettes. Compared to other palette-generating tools, category-colors has more control. Palettailor, for example, optimizes for pure color difference, without accounting for color vision deficiency. QualPal brings some of the optimization parameters, but doesn’t allow for steering towards or away from arbitrary colors. Scores at 8 colors ΔEMinimum ΔEworst of CVD Name differenceMinimum Uniformitylower is better category-colors 22.6 ±1.6 13.7 ±2.0 0.35 ±0.14 best in column 0.30 ±0.02 best in column QualPal 1.1.0 24.7 21.8 best in column 0.10 0.44 Palettailor 26.6 ±2.4 best in column 4.5 ±1.7 0.34 ±0.16 0.34 ±0.04 Colorgorical 15.8 ±3.3 4.1 ±1.6 0.09 ±0.06 0.42 ±0.03 All numbers are at 8 colors. Rows with ± are mean ± standard deviation over 10 palettes; rows without are deterministic and produce one palette. category-colors and Palettailor are 10 independent runs on the same seeds; Colorgorical’s row is 10 palettes from its authors’ own sampling script at equal criterion weights. QualPal was run with CVD on, matched bounds, and takes no seed. Name difference is Heer & Stone’s 1 − cosine; Colorgorical’s own interface reports a Hellinger distance instead. Shaded cells are the best value in their column. Compared to industry-standard palettes, category-colors can produce more optimal palettes, especially at high cardinality. Scores at 8 colors ΔEMinimum ΔEworst of CVD Name differenceMinimum Uniformitylower is better category-colors 22.6 ±1.6 best in column 13.7 ±2.0 best in column 0.35 ±0.14 0.30 ±0.02 best in column Okabe–Ito 21.3 8.8 0.06 0.34 Observable 10 18.4 0.6 0.40 0.34 Tableau 10 18.1 3.2 0.24 0.32 d3 category10 16.2 1.6 0.84 best in column 0.40 ColorBrewer Set3 13.7 1.9 0.16 0.32 IBM Carbon 12.8 5.0 0.11 0.34 Same run: 8 colors, 10 trials. Reference palettes are deterministic, so they're single values. Shaded cells are the best value in their column. One more thing I’ve built a UI that consumes the package and makes it easy to generate and optimize palettes. This has been the biggest request since I published the initial essay, so it’s the thing I’m excited to share. It’s ridiculously overengineered, but hey, what else are personal projects for? Acknowledgements Many measurements come from published research: Gaurav Sharma, Wencheng Wu and Edul Dalal for CIEDE2000; Gustavo Machado, Manuel Oliveira and Leandro Fernandes for the color vision deficiency simulation; Maureen Stone, Danielle Albers Szafir and Vidya Setlur for the size-dependent just-noticeable-difference result; Jeffrey Heer and Maureen Stone, whose color naming models and the c3 data from the Stanford Visualization Group power both the saliency and name-difference evaluators. Existing palettes: Masataka Okabe and Kei Ito’s Color Universal Design set; Matthew Petroff’s sequences; and Mark Harrower and Cynthia Brewer’s ColorBrewer. Other generators laid a lot of the groundwork: Kecheng Lu and colleagues (Palettailor), Connor Gramazio, David Laidlaw and Karen Schloss (Colorgorical), Johan Larsson (QualPal), and Chin Tseng, Arran Zeyu Wang, Ghulam Jilani Quadri and Danielle Albers Szafir (CatPAW). Andrew McNutt, Maureen Stone and Jeffrey Heer’s color-buddy has also been indispensable. Finally, Dan Burzo’s culori made it easy to make this library colorspace-agnostic.

a week ago 1 votes
Design from the inside

Imagine you’re an architect hired to redesign the floorplan of an office. The company hiring you has grown from 10-100 employees and wants to make sure the space is easy to navigate and the common areas are in the optimal location. You ask the client for the existing floorplan, but nobody can find the original drawings. They’d be useless anyway, because as the company has grown, the employees have been given license to change the space as they see fit. Their modifications range from simple decorating to major renovations. One employee walled off an entire corner of the office for themselves and nobody has seen them in weeks. No single employee can draw the floorplan from memory. Walking the space, you discover they’ve added three sets of bathrooms — as the office got more and more byzantine, it became easier to hire a contractor to build new bathrooms than to find the existing ones. Signage is a joke, and asking for directions is useless: everyone thinks they remember where things are, but their memory is inevitably outdated or otherwise biased. This is the reality of design at high-growth startups in the AI era. Engineers can build so fast and so independently that trying to map out the product area is a lost cause. Some of the traditional tools of product design — mapping and evaluating UX from a bird’s eye view — are useless. So what do you do instead? You have to stop thinking like an architect. An architect designs buildings from the outside. They use floor plans and elevations and other schematics to paint a picture of an ideal reality. They create a ‘source of truth’ that is used to coordinate engineers and builders. This is the old world of design. In the new world, the building needs to be designed from the inside. There is no ‘source of truth’ when everything changes at every moment. You do not wait to build consensus or gain a full understanding of the office to start making changes. You buy a roll of high-viz safety tape from Home Depot and start laying it down on the floor. In some spaces, you tape the outlines of more efficient walls. You put tape across the entrance to a few dead-end hallways. You buy a ‘wet floor’ sign too and put it in front of some of the bathrooms with a post-it-note saying ‘closed for cleaning.’ The staff start subconsciously heeding the taped-in redirects. Without realizing it, they are moving through the office more efficiently, congregating in common spaces again. They’re delighted to see their old colleagues, many of whom they assumed had been laid off. One team tells the others that they bought a snack machine, and new hires explore the far reaches of the office they’d previously never have seen. As one-off spaces and facilities are abandoned, you start to knock down walls. What before would have been a code red fireable offense is now completely unnoticed. The newly-available space is used to slowly expand hallways, adding a few inches of space at each pass, staying just on the edge of notice. Everyone gets a little more vitamin D. Fewer toes are stubbed on sharp corners. Each improvement brings new problems to light: now that the main bathrooms are getting more use, we’ll need to put in changing tables for new parents (the company was started by 20-year-olds who didn’t think about that when they moved in). We need to have quiet areas to balance out the noise created by more spacious open floor plans. But each problem can be solved from inside the office, taping off areas and redirecting foot traffic, manifesting desire paths into the world. For designers, our tape is code. In the same way that the inside-out architect redirects traffic, we can make small changes to product surfaces that incrementally reveal improvements. Ship like a developer, flowing your pull requests into the CI/CD pipeline as if it were just another Typescript file. How do you design from the inside? Work in the codebase. Resist the urge to create prototyping sandboxes and demo environments. If you need someone to translate your designs into the codebase, you’re on the outside, not the inside. Create tight links between your design environment and the real product. Use real data from your own API endpoints; make sure the design tokens map 1:1 between Figma and code; work at realistic screen sizes and under realistic test conditions. Ship. Put your code up for review. LLM coding agents can turn what you’ve built into production-ready code with just a few prompts. Build rituals, don’t introduce processes. Instead of introducing a new process and asking (expecting) your collaborators to follow it, simply behave as if the process is already in place. Want your team to start using PRDs before building? Make the PRD and send it out. Want to do an implementation review? Put it on the calendar. Repeat a behavior enough times and it will become ritual. Collapse feedback loops. Don’t wait for a summary of support tickets; do support rotations, and you’ll have all the input you need to fix the real problems. Ask an LLM to write SQL queries to get the data on usage, don’t wait for a dashboard or UX whitepaper. Instrument the product yourself, then push changes that move the needle. Don’t ask for permission or seek consensus; design the product from the inside.

5th May 2026 2 votes
Expansion artifacts

The information age has been defined by bandwidth. The internet is limited by how much data we can squeeze into the narrow pipes of transmission infrastructure. So we invented compression, ways of representing the same object — a website, a picture, a song, a movie — within ever smaller digital footprints. YouTube, Spotify, Instagram, and the algorithms that make them work, wouldn’t be possible without it. From the very first studies into compression (at Bell Labs in the 1940s), researchers knew they’d have to accept a tradeoff: you can achieve smaller file sizes if you’re willing to accept some loss of the original data. This seems counterproductive, since the whole idea is to reproduce the data, but scientists found ways to only discard information that is imperceptible to humans. Our ears and brains tend to ‘filter out’ quiet sounds overshadowed by loud sounds. MP3s take advantage of this blind spot by stripping out the quiet bits we wouldn’t be likely to hear anyway. Our eyes and brains focus on the contrast between bright and dark shapes, reading the broad structure of images rather than granular details or tiny color variations. The JPG algorithm compresses files by throwing away information we don’t tend to process. Movies don’t actually change that much frame-to-frame. The MPG algorithm carefully chooses key frames and saves the relative motion of each pixel, making movie files much smaller in the process. A well-designed compression algorithm keeps data perceptually identical while making files much more efficient to store and transmit. A poorly designed codec can go catastrophically wrong. In 2013, David Kriesel scanned a building floor plan on a Xerox WorkCentre and noticed that a room marked 21.11m² had become 14.13m². Xerox’s implementation of the JBIG2 compression format saves space by quilting scans together from common, repeated elements; in Kriesel’s scan, it had silently replaced the original numbers with ones from another part of the document it deemed visually similar enough. After Kriesel published, reports surfaced of the same silent substitution affecting building plans, invoices, and medical records.1 Compression always changes data permanently. Common formats (JPG, MP3, MP4) make changes slowly and gently: it usually takes hundreds of cycles of saving, sharing, and re-uploading before the tool marks, called compression artifacts, become apparent. Re-save a JPG enough times and it goes blocky and washed out; iterate an MP3 and metallic tones bleed through the music; re-upload a YouTube video a thousand times and you end up with a blobby mess over unintelligible audio. Before compression: PSNR of infinity. After 10,000 compression cycles: PSNR of 14.59. If you know what to look for and how to look for it, you can learn a lot about the path that data took to get to you. That’s because compression artifacts are in turn meta-information; you can learn something new about a document by identifying and cataloging its algorithm-induced flaws. Digital forensics uses this meta-information to explore the provenance of documents, photos and videos. Compression leaves breadcrumbs that betray whether or not a document has been edited (and often who, or what, edited it). Compression artifacts can even become an aesthetic of their own. Deep-fried memes dress images up in the aesthetics of pictures that have been shared and re-shared thousands of times. Datamoshing manipulates compression algorithms to create entirely new video aesthetics. Glitch music stretches and squashes audio files, making the tool marks of audio compression audible and even musical. Compression has spawned entire fields of art and science (and jokes) all in service of the ideal compromise between fidelity and file size. Three years ago, Ted Chiang described ChatGPT as a blurry JPEG of the web. LLMs are a lossy compression of their training data, which is itself a lossy sample of all the data available to it. But the artifacts we see in AI slop aren’t in the compression. They’re in the decompression. Every AI-generated output is an extrapolation from that blurry source, vectored toward your prompt, filling in plausible detail where the compression threw information away. The output gets inflated into blog posts and LinkedIn thoughtspam, software platforms, omnichannel advertising campaigns, and movie cameos from dead actors. Chiang compared the gaps and confabulations to compression artifacts. I think they’re expansion artifacts. What do expansion artifacts look like? LLMs produce text stuffed with hedging verbs and fuzzing adjectives (delve, intricate, tapestry, multifaceted). Their paragraphs are structured as miniature essays with setup, payoff, and a signposted takeaway (This matters because…). AI-generated code over-comments the obvious and creates error handlers for operations that can’t logically fail. Image generators have had their own tells: six-fingered hands, symmetrical-but-stylistically-objectionable jewelry, text that looks like text but only if you cross your eyes. Video models struggle with continuity. Limbs appear and disappear, objects clip through each other, and physics sometimes just switches off. Each of these artifacts is the training distribution leaking through where the model’s confidence runs thin. Like compression artifacts, they double as forensic markers. In 2024, Stanford researchers tracked AI contributions to academic writing by watching for words whose frequency spiked after ChatGPT’s release (commendable, meticulous, pivotal, showcasing, etc.). They estimated that 17.5% of recent computer science papers and 16.9% of peer review text contain AI-drafted content. Sometimes the tells are less subtle: one paper in an Elsevier journal opened with “Certainly, here is a possible introduction for your topic.”2 Expansion artifacts will become aesthetic choices, too. Shrimp Jesus is my favorite, the kind of insane imagery that only an LLM would create. Power users of AI website generators (AI-pilled designers) already know how to recognize the tool marks, if only to try to prompt them away: purple gradients are an especially common tell. But as more and more non-designers use tools like Claude Design to prompt their way to fully-functional software products, I expect to see a preference for the aesthetic convergence endemic to the current crop of AI models.3 https://x.com/adamwathan/status/1953510802159219096 Expansion artifacts get genuinely dangerous when they compound, when one AI generation becomes the input to another, and another, and another. In February, an autonomous openclaw agent published a hit piece on Scott Shambaugh, a maintainer of the popular matplotlib Python library, for rejecting its code. Benj Edwards then reported the story for Ars Technica, but used AI to help him write; unsurprisingly, his article contained hallucinated quotes. This kind of Gell-Mann Amnesia for expansion artifacts leads to runaway feedback loops: A CEO dictates a five-minute voice memo Claude expands it into a strategy doc Notion’s AI turns the strategy doc into product specs Cursor vibe-codes a prototype Devin gives feedback on the PR ChatGPT writes the launch copy Intercom’s Fin support agent fields support questions. Every stage interpolates the previous context with data pulled from the blurry JPEG of its training distribution. The real danger happens when expansion artifacts show up in the training data for the next generation of generative AI. While anomalous tokens like SolidGoldMagikarp broke early models open and spread their guts out on the operating table, new models get rigorously evaluated, making it harder to see where the errors are hiding. The long tail in the distribution (quiet voices, the weird and novel phrasings, unusual and challenging ideas) fade as each successive model converges toward a homogenized — perhaps hallucinated — center. The blurry JPEG gets blurrier every cycle, leaving more and more room for nonsensical and false material to fill the voids between the tokens. Compression made the information age possible by stripping things down to fit the pipes. Expansion made the AI age possible by blowing data back up again. Both operations leave marks; we’ve learned to spot compression artifacts, but we’ve only just begun to reckon with expansion artifacts. Until we do, there’s a lot of risk to manage. Special thanks to Josh Petersel for feedback on a draft of this essay. Footnotes & References Kriesel, David. 2013. “Xerox Scanners/Photocopiers Randomly Alter Numbers in Scanned Documents.” D. Kriesel (blog), August 2, 2013. https://www.dkriesel.com/en/blog/2013/0802_xerox-workcentres_are_switching_written_numbers_when_scanning. ↩︎ Liang, Weixin, Yaohui Zhang, Zhengxuan Wu, Haley Lepp, Wenlong Ji, Xuandong Zhao, Hancheng Cao, Sheng Liu, Siyu He, Zhi Huang, Diyi Yang, Christopher Potts, Christopher D. Manning, and James Y. Zou. 2024. “Mapping the Increasing Use of LLMs in Scientific Papers.” Preprint, submitted April 1, 2024. https://doi.org/10.48550/arXiv.2404.01268. ↩︎ Copying is the way design works. ↩︎

20th Apr 2026 1 votes
Decentralizing quality

Everyone agrees quality matters, but we can’t agree on what it is — or who gets to decide. I’ve experienced the drive for quality in every design leadership role I’ve had. When it comes to software, we like to pretend that quality is a number. Dashboards stand in for judgment, A/B tests stand in for taste, and leaders try to will excellence into existence with reviews and mandates. But in today’s software market, confusing a KPI with quality isn’t just naïve; it’s fatal. Consumers’ attention spans are short, but their expectations are higher than ever; if your product fails to solve a real problem in an elegant way, you don’t get a second chance. Some companies see the value of quality and are building their brand around it. Linear, for example, has launched an entire Conversations on Quality series, because quality has become the currency every builder wants to be paid in. In my time at Stripe, the pursuit of quality was nearly obsessive: 99.999% uptime (about five minutes of downtime a year) was the bare minimum. On the user experience side, we created a program to standardize and report on quality across every product team in the company. Last year, Malthe Sigurdsson (who led design at Stripe from 2015–2020) rejoined as “Head of Craft” to drive quality across design. But more and more I’ve come to believe that quality isn’t a slogan, a program, or a scorecard. It’s a promise kept at the edge by the people doing the work. And, ideally, quality is fundamental to the product itself, where users can judge it without our permission. That’s the shift we need: away from heroics at the center, toward systems that make quality inevitable. The stakes are high. Centralized quality — slogans, KPIs, executive decrees — can produce positive results, but it’s brittle. Decentralized quality — continuous feedback, distributed ownership, emergent standards — builds resilience. In this essay, I’d like to make the case that the future belongs to those who can decentralize their mindset and approach to quality. What is quality? We can’t build quality products if we can’t define what quality is. It’s a slippery thing to grab, though: In a 1984 MIT Sloan Management Review article, David Garvin described five different approaches to defining quality across eight dimensions.1 Ultimately, he gives up on defining quality and argues for multiple definitions. That’s a cop-out in my book, so here’s my one-sentence definition of quality. Quality is the degree to which a product or service meets or exceeds user expectations. Why did I land on that particular definition? First, I have a strong conviction in the subjectivity of quality. Put simply, quality depends on users’ or consumers’ perception. Expectations are dynamic and relative, so quality can vary by user and change over time as our expectations, the products we buy, or that competitive landscape evolves. Next, I chose this definition to remove tastemakers from the equation. McDonald’s can be quality food because people buy a Big Mac and expect a double-decker hamburger with special sauce. Ikea can be quality furniture because college students expect an affordable desk for their dorm. It’s only when reality falls out of sync with expectations that quality is questionable. Lastly, my definition of quality is especially adapted for software, where products and expectations change constantly. Subscription-based software-as-a-service is the norm, and apps are updated daily. We expect constant, cheap or free internet access thanks to cell phones, wifi, and satellite internet. Tools that still seem like science fiction — voice recognition, machine learning, and search algorithms — give us answers in milliseconds. Building quality software is already hard. Before ChatGPT’s 2022 release, it was absurd to imagine a computer composing a structurally perfect 14-line sonnet about Iggy Azalea’s musical influences. Now, news of an AI passing the bar exam barely deserves a push alert. With technology advancing more and more rapidly, it will only get harder to understand and exceed consumer expectations. This relentless acceleration puts enormous pressure on software builders. How do you maintain quality when the definition keeps shifting beneath your feet? Many companies bet on strong leadership — charismatic executives who can produce quality products with pithy catch phrases and aggressive incentives. Ford Motor Company tried exactly this approach, and their story is a perfect study in leadership-driven quality control. Quality is Job 1 In the early 1980s, Ford Motor Company launched the iconic slogan “Quality is Job 1.” It wasn’t just a marketing angle; it was a genuine commitment by Ford to improve vehicle quality, winning back consumers from global competition. A decade before, Ford, like other American automakers, struggled to maintain market dominance. New environmental and safety regulations, along with increased labor activism, challenged the car manufacturing status quo. They scrambled to hold on to profits, causing rippling quality compromises. The Ford Pinto is an example of how these feedback loops got out of control. Before 1970, Ford identified potential issues with the gas tank placement in its smaller cars and designed a new, safer one. But when designing the Pinto, analysts determined that a rubber bladder safety system would add just $5.08 to the cost of each car—less than a quarter of one percent of the Pinto’s $1,919 retail price—and chose the older, cheaper, less safe design instead.2 Associated Press These cost-saving choices had consequences. In June 1978, a California jury awarded $128 million to a boy badly burned in a Pinto accident. Two months later, a speeding van hit the driver of a '73 Pinto from behind, causing the gas tank to explode and killing the driver and two passengers. After a damning exposé in Mother Jones, Ford agreed to recall all 1.5 million Pintos made between 1971 and 1976.3 While the “big 3” American auto manufacturers — Ford, GM, and Chrysler — struggled, Japanese automakers like Toyota and Honda surged. After World War II, Japan embraced quality management principles to keep costs down and efficiency up. With the guidance of American experts like W. Edwards Deming and Joseph Juran, Japanese automakers focused on continuous improvement and defect reduction. Foreign companies began to consistently outperform their American counterparts in quality rankings, customer satisfaction, and market share growth. In 1960, imports accounted for less than 5% of U.S. car sales. By 1971, they accounted for about 15%. By the 80s, foreign-made autos — mainly Japanese — reached over 30% of the U.S. market.4 In 1981, Ford announced “Quality is Job 1” to counter growing foreign market share and signal a change in its approach to manufacturing and customer satisfaction. Leadership saw it as a way to win back the market. Initially, the campaign worked. Ford invested heavily in quality control, new manufacturing techniques, and employee training, including hiring Deming, whose theories helped Japan excel. The company began to see product improvements and steady market share. For a time, Ford’s mandate for quality worked. By the late 90s, its stock price was at an all-time high; it looked like the company had engineered a miracle. The popularity of trucks and SUVs in America gave executives and shareholders a false sense of security. American manufacturers like Ford were well-positioned to build and sell these kinds of gas-guzzling vehicles. Ford lost its appetite for quality reforms. In 1998, they stopped claiming that “Quality is Job 1.”5 Safety issues immediately started plaguing the company, like the 2000 Firestone tire controversy that led to the recall of 13 million tires on Ford Explorers. Ford’s market share declined steadily throughout the 2000s and never recovered. The moral of the story is a fundamental question: can centralized quality mandates create long-term cultural change? Or are top-down mandates only as strong as the leaders that issue them? What is centralized quality? Ford’s story captured my attention because of what the company stands for. Henry Ford is credited with revolutionizing mass manufacturing to produce an affordable car for the middle class; in reality, many forces converged to bring the global economy to the precipice of modern mass manufacturing, and Henry Ford merely stepped through the door. Before mass manufacturing, quality was implicit in commerce. Hand-made goods had to be just that — good — in order to sell. But mass production created a fundamental problem: when thousands of workers produce millions of identical items, individual craftsmanship becomes impossible. A single blacksmith could ensure every horseshoe met his standards, but how do you maintain quality when a hundred workers are stamping out identical parts on assembly lines? The scale that made products affordable also made quality control infinitely more complex. In the early 20th century, management consultant and doubles tennis champion6 Frederick Winslow Taylor provided an answer. He believed enlightened managers could improve quality results in manufacturing through the design of workers’ conditions and the control of workers’ movements. At Bethlehem Steel, these methods increased pig iron throughput from 12.5 to 47 tons per day. Taylor’s insight was revolutionary: if you couldn’t rely on individual craftsmanship, you could engineer quality into the system itself. Ever since then, we’ve lived in Taylor’s world. His scientific management spawned a century of sophisticated quality methods: Statistical Process Control in the 1920s, Total Quality Management in the 1980s, Six Sigma in 1986. Each approach shared a common thread: experts and managers direct quality improvements instead of depending on workers themselves. The hallmarks of centralized quality Centralized quality starts and ends with leadership-driven quality standards and reviews, where a single stakeholder reviews our work and decides whether it meets their standards. This has been the standard at every company I’ve worked with; design and engineering are both taught through critique and reviews from the earliest stages of production to the final delivery of a complete product. This approach can work brilliantly when executed with precision, but it creates a bottleneck: no matter how talented, a single person can only review so much. Sometimes it’s not a single leader but an entire class of leadership that enforces quality from the top down. At Stripe, for example, we instituted a “walk the store” initiative that had leaders use the products their team shipped. For example, executives would regularly create a new Stripe account, set up the payments product, and schedule payouts to a test bank account; they’d score the overall quality of the product from their own experiences, and track improvements over time. This program had good intentions, but because it was driven by leaders with strong opinions and biases, it often missed the point entirely. We all knew that Stripe’s founder Patrick Collison preferred interfaces that were densely packed and full of efficient shortcuts and keyboard commands; teams would bias their work towards Patrick’s preferences to ensure a smoother review and approval process. We weren’t solving user problems. We were solving for Patrick’s aesthetic sensibility. At one point, Patrick sent out a memo imploring teams to stop this preemptive bias, but it only added another recursive layer to the feedback process: it had to meet Patrick’s high expectations, but steer clear of sycophancy. One way companies have tried to reduce the bottlenecks caused by leadership reviews is by creating specialized quality roles within the company. At one point in my career, QA teams were standard. They specialized in testing the team’s work, either manually or through automation, and reporting bugs and defects back to the team. Apple employs thousands of QA specialists to test every aspect of its products, from hardware to software and user experience. In 2023, Apple staffed 12 times more workers on its iPhone assembly process than Google did for its Android phones.7 For their part, Google has invested heavily in software quality through dedicated roles on their “Engineering Productivity” team. From 2005 to 2012, Patrick Copeland (formerly head of testing at Microsoft) developed and grew the team to 1,200 engineers.8 Today, it has over 2,000.9 With software companies touting a renewed focus on quality, will we see a rise of QA engineers? Already, engineers are having to devote more and more time to cleaning up low-quality vibe code created by AI. And what about QA designers or product quality specialists? In some ways, the rise of “design engineering” has been about quality and refinement of the end product. Another symbol of the centralized approach to quality is quantitative performance metrics. Key Performance Indicators (KPIs) are the hallmark of scientific management and the source of most modern organizational schemes. Teams pick measurable aspects of the product — number of defects, units produced, bugs fixed, subscriptions sold — set goals, and report progress. For my part, most of my design teams have been measured by these objective metrics, despite the reality that design is seldom reducible to a number. While it makes management feel more scientific, this quantitative approach is usually at odds with real users’ experiences, preferences, and expectations. So goes the theory of centralized quality management. But does it work? I’ve researched dozens of companies that employ centralized quality management; initially, I assumed that most top-down quality mandates could be shown to be ineffective. But there are many success stories to learn from. In fact, most companies in the S&P 500 employ some form of centralized quality management. But within each case study were signs that the centralized approach was costly and fragile. Even more telling are the cases where centralized quality control was a core reason for a company’s downfall. What does centralized quality look like? Especially among designers, we love to valorize Steve Jobs for his philosophy on quality, because Jobs always expected the impossible. Take the original iPhone: when developing the phone’s screen, Jobs invited Wendell Weeks, the new CEO of glass manufacturer Corning, to Apple’s headquarters. He laid out his vision of strong, scratch-resistant glass, a sharp contrast to the cheap-feeling plastic that other phones used. Weeks revealed Corning had invented such a material (“Gorilla Glass”) in the 1960s but never found a market. Jobs ordered all the Gorilla Glass Corning could produce in six months. But at the time, the company couldn’t produce any at all. Bootstrapping enough capacity to manufacture hundreds of thousands of phones was impossible — or so Weeks thought. Jobs was undeterred. “Don’t be afraid,” he said. “Get your mind around it. You can do it.”10 Weeks did just that. Corning scrambled to build brand-new manufacturing lines for Gorilla Glass, then met the iPhone’s demand. It banked record profits in 2007, and by 2010 Gorilla Glass was in 20% of cell phones worldwide.11 I wish I could tell my stakeholders “don’t be afraid” whenever they question my decisions. Blake Patterson Jobs’s approach worked because Corning delivered. But in my experience, satisfying a visionary leader usually comes at a cost. That was true for Apple under Jobs; every significant decision filtered through layers of demonstrations and approvals. Teams demoed their work repeatedly to different managers, on up the chain, and ultimately to Jobs. With only one person capable of final judgment, teams often waited weeks for their audience with the CEO.12 Leaders rarely want to be bottlenecks. So effective centralized quality management requires delegation. Delegation, in turn, requires accountability. In my research, Jeff Bezos kept turning up as an example of accountability in centralized quality control. Like Steve Jobs, Bezos expected the impossible from those around him. In the late 90s, Bezos’s obsession with customer service put Amazon’s call centers in the crosshairs; his vision for Amazon’s customer service experience was no experience at all. “Every time a customer contacts us,” he said, “we see it as a defect.”13 If customers had to call, their questions should be answered quickly and thoroughly. In 1999, Bezos tasked Bill Price with eliminating hold times. Price warned it was impossible: shorter calls meant inadequate solutions, forcing customers to call back, bogging down the system and increasing wait times. Bezos didn’t care. During a daily leadership meeting called the “war room,” he pointedly asked Price for an update on support wait times. Price claimed they were under a minute, and Bezos called his bluff by dialing Amazon’s help line on the conference room’s speakerphone. It took four and a half minutes to get an answer. Price was gone in ten months.14 This sort of egomaniacal quality often causes leaders to become micromanagers. Like Elon Musk — his leadership of Tesla shows just how involved a top-down leader can get in day-to-day quality concerns. In 2018, with delivery numbers slipping, Musk needed to personally guarantee the assembly line’s quality, measured by its ability to churn out thousands of the company’s newest vehicles. As Tesla struggled to meet the demand for their Model 3, he made the Fremont, California factory his home. First, he slept on a couch, then on the floor. He spent time with workers, understanding the process, giving suggestions, and doing the work himself. “The reason people in the paint shop were working their ass off is because I was in the paint oven with them,” he told a Bloomberg reporter in 2018. Musk only left the factory when Tesla hit 5,031 Model 3s in a week. Behind him were employees “staring out into space like zombies” — the human cost of heroic leadership.15 Even if you haven’t worked at Apple, Amazon, or Tesla, you’ve probably experienced centralized quality control. I’ve dealt with my fair share of self-styled visionary leaders demanding excellence, then micromanaging execution. But in my experience, this approach has a critical flaw: leaders can be wrong. For example, take Jawbone, which suffered a cataclysmic crash due to its centralized style of quality management. In 2009, the company decided to bet on fitness trackers, building on early success with Bluetooth headsets. CEO Hosain Rahman demanded Jawbone’s first entry into the category — the Up band — be the smallest and most stylish device on the market. But cramming new hardware into tiny enclosures created manufacturing nightmares: in initial manufacturing efforts, the device’s electronics were destroyed when hot rubber was injected into the mold, and when devices broke, it was difficult to decipher the reason because the electronics were encased in rubber.16 Despite more than 75 percent of prototypes failing after normal consumer use was simulated before launch, Rahman wouldn’t compromise on aesthetics.17 The Up band shipped with widespread defects. Jawbone raised $900 million trying to fix problems that persisted through three product generations. The company was liquidated in 2017.18 We can’t all be Steve Jobs. So centralized quality management can, under the right circumstances, create excellence. And it can also, in the wrong circumstances, fail spectacularly. But there’s another, more resilient path — one where quality emerges from the product’s builders, not its managers. Unlike the centralized approach, this system is adaptable to change and empowers the people who implement it. From centralized to decentralized I’ve been amazed at the way that Microsoft has been able to completely reinvent itself over the past 20 years. It’s hard to believe the company that launched the Zune would not only catch up to Apple, but surpass it to become a leading innovator in the age of AI. Much of this is down to how they defined — then redefined — quality. As Microsoft’s CEO, Bill Gates embodied centralized quality. A former Excel program manager described Gates reviewing his 500-page Visual Basic specification. He read every page, took detailed margin notes, and assaulted the team with technical questions.19 Co-founder Paul Allen called working with Gates ‘like being in hell,’ describing him as having an abusive personality who ‘thrived on conflict.’20 Gates’s approach created a high-performance but punishing culture. His successor, Steve Ballmer, held the same beliefs about quality but lacked Gates’s technical instincts. By 2014, Microsoft had stagnated. Google dominated search, and Apple was the world’s most valuable company. Microsoft needed to change. Satya Nadella flipped the script. Where Gates controlled every decision, Nadella empowered teams to drive quality. He eliminated the notorious “stack ranking” system that forced managers to rate employees against each other in zero-sum competition. Stack ranking was a core feature of Microsoft under Gates and Ballmer and a perfect distillation of centralized quality management. Nadella’s empowerment approach fostered decentralized quality innovation. For example, a deaf Microsoft engineer named Swetha Machanavajhala independently developed background-blurring technology for Skype calls to help communicate with her parents in India. Instead of top-down mandates, Nadella’s culture empowered her to influence similar features for Microsoft Teams.21 The annual “One Week” hackathon, started in 2014, exemplifies Nadella’s collaborative quality approach. In 2017, over 18,000 participants in 4,000 cities worked on projects outside their day jobs, generating innovations like Seeing AI and Xbox Adaptive Controller.22 This represents a complete philosophical shift from Gates’s controlled product development to employee-driven innovation. David Golds, who left Microsoft in 2012 and returned in 2017, observed: “What changed was leadership, and everything followed from that. There was a tearing down of walls for software developers.”23 By 2023, Microsoft’s stock had increased nearly tenfold in the nine years since Nadella became CEO, with a 27% annual growth rate, ending a 14-year period of near zero growth.24 Acquisitions like Mojang, LinkedIn, and GitHub, along with a deep partnership with OpenAI, show how Nadella sees innovation as a collaborative effort. Progress can come from anywhere, and in Nadella’s Microsoft, quality comes from the ground up. Microsoft shows how reframing quality as a core responsibility of frontline workers can have an immediate and substantial impact on a company’s fate. What is decentralized quality? Decentralized quality means putting quality in the hands of workers, not managers. Before Frederick Taylor’s scientific revolution, quality emerged organically. Medieval guilds controlled craftsmanship through apprenticeships. Masters passed knowledge to workers, who earned the right to guarantee their work. A blacksmith’s reputation hung on each horseshoe; a baker’s livelihood relied on tomorrow’s bread being as good as today’s. Quality wasn’t imposed; it was inherited, practiced, and owned by workers. Taylor’s industrial efficiency swept away these traditions but preserved the essential truth: those closest to the work know how to improve it. Even as scientific management took hold, alternative voices emerged. In the 1920s, Mary Parker Follett argued for participatory leadership, where workers shape their own processes through what she called ‘power with’ rather than ‘power over.’25 Frank and Lillian Gilbreth challenged Taylor’s centralized approach by emphasizing the psychological aspects of work and worker welfare. Lillian’s 1914 work The Psychology of Management argued that effective management required understanding ‘the effect of the mind that is directing work upon that work which is directed, and the effect of this undirected and directed work upon the mind of the worker,’ advocating for approaches that considered individual worker needs and job satisfaction alongside efficiency.26 National Museum of American History Decentralized quality inverts the traditional management hierarchy. Instead of executives defining standards and workers following them, frontline employees drive improvements. They identify problems, propose solutions, and implement changes. Leadership’s role transforms from commander to enabler, creating systems that empower workers rather than controlling them. The hallmarks of decentralized quality Even before writing this essay, I was well-aware of Toyota’s counterintuitive approach to building quality products. At Toyota, quality is worker-driven. Each and every employee is responsible for the quality of the end result, no matter how small the part they play in its construction. This idea has caught on beyond manufacturing, too: 3M’s “15% time” policy allowed employees to pursue their own quality improvements and innovations, yielding breakthrough products like Post-it Notes and Scotchgard. Google’s version led to two of its most influential products, AdSense and Google Maps.27 I’ve always liked Netflix’s worker-centric approach to quality. The company deliberately sabotages its own systems with tools like Chaos Monkey, which randomly shuts down servers during peak viewing hours. It’s strategic paranoia: by constantly breaking their own systems, Netflix engineers are forced to build services that can survive anything.28 Another important aspect of decentralized systems is distributed ownership across the workforce, in contrast to concentrated quality control in inspection departments. For example, at Morning Star Company, the world’s largest tomato processor, there are no traditional managers or hierarchical structures. Instead, employees operate through “Mission Focused Self-Management,” where quality standards emerge from peer collaboration and individual accountability. Workers negotiate their responsibilities directly with colleagues and take full ownership of their part of the production process.29 W. L. Gore & Associates, the manufacturer of Gore-Tex, operates on similar principles. The company organizes into small teams where every employee (called “associates”) has direct responsibility for product quality. The company’s lattice structure replaces traditional hierarchy with peer-to-peer accountability, resulting in some of the highest employee satisfaction and product quality ratings in manufacturing.30 Distributed quality management is resilient; Morning Star and W. L. Gore have both stayed competitive because they are able to quickly adapt to evolving customer expectations and react to complex systems like global supply chain shocks and ecological extremes. But decentralized quality isn’t just for manufacturing. During my time at the Wall Street Journal, I watched editors manage breaking news with a distributed quality system that would terrify most software companies. A hierarchy of editors pushed responsibility from editor-in-chief down to front-line desk editors to the individual reporters and journalists. The newsroom produced hundreds of stories every day, breaking news where minutes mattered. A daily coordination meeting was all the newsroom needed to stay in sync. Yet the Journal maintained an astronomical quality bar for writing, reporting, and fact-checking. The secret wasn’t heroic leadership. It was a system where standards propagated through culture, where desk editors had both the authority and accountability to make calls in real time. Decentralized quality thrives on constant rapid feedback. As a designer, I’m deeply aware of the value of timely feedback. And at the most creative companies in the world, feedback is a full-contact sport. Pixar, for example, developed the “dailies” process, based on the film industry norm of screening unedited footage with the crew at the end of a day of shooting. Pixar’s animation renders can take many hours to process, putting a huge amount of pressure on the team to deliver progress every single day; but showing work-in-progress creates immediate feedback loops to catch and address quality issues within hours, not months.31 In my own work, I’ve used similar approaches like test-driven development and continuous integration. Code changes trigger immediate feedback — tests run, peers review, and systems provide instant quality signals. Problems surface instantly, not weeks or months later. One of the most fascinating aspects of decentralized systems is that quality standards emerge organically based on experience and results instead of imposing rigid specifications from the start. Wikipedia demonstrates this principle at internet scale. Quality standards for articles come from community consensus, with experienced editors mentoring newcomers and collectively refining guidelines that shape the encyclopedia. Despite what my high school teachers told me, this has produced an encyclopedia that surpasses traditional ones in factual integrity and editorial rigor. What does decentralized quality look like? Back to Toyota: In 1966, the company installed a rope attached to a board of lights above the assembly lines in their Kamigo plant. Management gave workers a new responsibility: on spotting a defect, any worker should pull the rope, which would stop the line. The board would light up, a supervisor would come over, and the group would ask: “How can we fix this?”32 A simple experiment changed quality management forever. Toyota Motor Corporation Toyota’s system inverted traditional quality management. While Ford drilled workers to always keep the line moving, Toyota empowered every employee to halt production for quality issues. Instead of hiding defects to avoid blame, workers were celebrated for surfacing problems before they reached customers. Toyota’s andon system (named after a traditional Japanese lantern) made quality everyone’s responsibility. It instilled a culture where frontline workers felt safe to surface problems immediately, rather than letting defects pass down the line. Beyond andon, this philosophy of worker empowerment extends to Toyota’s kaizen culture of continuous improvement, which has generated over 2 million ideas from employees.33 Like the auto market as a whole, Toyota and Ford have changed substantially since the 1960s. But the impact of a decentralized quality culture speaks for itself: in 2024, Toyota vehicles averaged $441 in annual repair costs compared to Ford’s $775, with 98 problems per 100 vehicles versus Ford’s 130.34 When American automakers tried to copy the andon system in the 1980s, they failed. They couldn’t replicate the culture; workers at one GM plant were yelled at when they pulled the cord.35 The contrast shows that Toyota’s success wasn’t about a process or tool. It was a reimagining of the relationship between workers and quality. The distributed philosophy of quality ownership isn’t limited to manufacturing. Valve Corporation embodies distributed quality in its unique management system; that is, no management system at all. Valve makes video games, manufactures a gaming console, and maintains one of the largest games distribution platforms. They do this all while operating as a “flat” organization, meaning no bosses and no hierarchy, just employees choosing projects and forming temporary teams around shared interests. Workers literally roll their desks to different parts of the office when they want to join a new project. Unsuccessful initiatives lose people until there’s no one left to work on them. The results speak for themselves. Valve’s 43 first-party titles average 83/100 on Metacritic, with the Portal series achieving a 98.7% positive review rate and selling 28 million copies.36 Steam, Valve’s game distribution platform, has captured 75% of the PC gaming market and generates over $10 billion annually.37 At $3.5 million in revenue per employee, Valve outperforms Apple, Google, and Facebook in per-capita productivity.38 At Valve, quality emerges organically. Bad ideas can’t attract talent, while promising projects draw the company’s best. Marketing director Doug Lombardi explains: “Nobody writes a design doc and hands it to somebody… It’s the teams coming up with ideas and pushing in directions they want to take the product.”39 Sometimes decentralized quality wins in unexpected ways. When James Daunt became CEO of Barnes & Noble in 2019, the chain was hemorrhaging money and closing stores as Amazon dominated book sales. Instead of doubling down on centralized efficiency, Daunt did something counterintuitive: he gave local employees control over their inventory, displays, and community programming. Now, store managers could choose books they knew would sell, instead of mindlessly following corporate dictums. They organized poetry readings that drew neighbors, book clubs tailored to local interests, and author events that reflected their community. The corporate office, which had previously micromanaged everything from shelf placement to promotional calendars, stepped back. Book return rates fell from around 70% to under 10% as stores aligned stock with local demand.40 The company now opens dozens of new stores annually, defying predictions of its inevitable demise. Meanwhile, Amazon has closed all 68 of its physical bookstores.41 Their centralized approach created “jumbled assortments of random stuff” that failed to capture the human connections that Barnes & Noble’s empowered employees provide.42 In bookstores, local knowledge triumphs. These stories share a common thread: organizations that trusted their frontline workers to identify and solve quality problems. But decentralized quality has its own vulnerabilities. Valve’s radical structure has been criticized for creating informal power hierarchies and making it difficult to coordinate large projects. Some ex-employees describe a “high school clique” atmosphere where popular workers accumulate influence while others struggle.43 Without traditional management oversight, initiatives can moulder, or veer in directions that don’t serve broader company goals. Still, these examples show a different path for achieving quality, where excellence is defined in the course of building a product. Unlike centralized approaches relying on visionary (but fallible) leaders, decentralized systems are resilient to individual failures, adaptable to change, and empowering to builders. The andon cord, the rolling desk, and the local bookstore manager each represent a small bet on human judgment over institutional control. Those bets look like they’re paying off. The end game of distributed quality I’ll leave you with an example of how far decentralized approaches can go: LEGO. If quality management exists on a continuum, one end is the centralized approach of visionary leaders like Steve Jobs and Jeff Bezos. Creating breakthrough products in centralized systems of quality takes determination and relentless standards. Moving away from this extreme, companies like Toyota, Valve, and Barnes & Noble distribute authority to frontline workers, creating resilient systems that don’t depend on individual genius. At the far end lies LEGO’s approach: quality baked so deeply into the product that it requires no human judgment at all. Rathfelder LEGO’s quality practice pushes decentralization to an extreme. Every brick made since 1949 must fit with every other brick ever made — this interoperability requires a manufacturing tolerance of 0.002 millimeters, meaning bricks can’t differ by more than 1/25th the width of a human hair.44 Manufacturing with these tolerances at such a massive scale is extraordinary: LEGO produces 36 billion bricks annually, or about 1,140 parts every second.45 LEGO doesn’t need dedicated quality control specialists to decide if a part meets its standards: my 3-year-old son can tell the difference between a LEGO set and a cheap imitation. In the digital world, we can steal glimpses of ultimate decentralization. The internet itself is founded on decentralized, failure-tolerant protocols that allow vast amounts of data to flow across the globe without a loss of quality. But for most companies, achieving fully decentralized, systematic quality remains an aspiration. From slogans to systems I’ve personally struggled to implement a decentralized approach to quality in many of my teams. I believe in it from an academic standpoint, but in practice it works against the grain of every traditional management structure. Managers want ‘one neck to wring’ when things go wrong. Decentralized quality makes that impossible. So I’ve compromised, centralized, become the bottleneck I know slows things down. It’s easier to defend in meetings. But when I’ve managed to decentralize quality — most memorably when I was running a small agency and could write the org chart myself — I’ve been able to do some of the best work of my career. Centralized quality just doesn’t last. It relies on the willpower of leaders, on slogans and dashboards that fade when the room changes. Decentralized quality is harder to build, but it compounds: every frontline decision, every local improvement, every user-validated feature adds to a system that grows stronger over time. LEGO gives us an extreme to shoot for: quality so deeply embedded that even a toddler can recognize it. Software can, and should, work the same way. The next era of great products won’t come from heroic leaders demanding quality. It will come from companies that trust their builders, trust their systems, and trust their users to see and feel quality for themselves. The future belongs to those brave enough to decentralize. Footnotes & References David A. Garvin, “What Does ‘Product Quality’ Really Mean?” MIT Sloan Management Review 26, no. 1 (Fall 1984): 25–43. A quick synopsis: there’s the philosophical notion of quality, like R. M. Pirsig in Zen and the Art of Motorcycle Maintenance — a book whose main character is driven insane in his quest to define quality. Then there’s the idea of quality as some objective property of a product, like a quality ingredient. There’s the more subjective interpretation of quality, based on a user’s perception or taste. There’s also a version of quality that is based on the way the product was created (does it meet the requirements or specifications), and still other market-based definitions that focus on price or cost. ↩︎ Mark Dowie, “Pinto Madness,” Mother Jones, September/October 1977. ↩︎ Paul Ingrassia. Crash Course: The American Automobile Industry’s Road from Glory to Disaster. 1st ed. (Random House, 2010). ↩︎ United States International Trade Commission, Automotive Trade Statistics 1964-1978: Series B:Passenger Automobiles. September 1979. www.usitc.gov/publications/332/pub1002.pdf. Accessed 24 July 2024. ↩︎ Stuart Elliott, “Ford Is Jettisoning Its 17-Year-Old ‘Quality Is Job One’ Slogan,” The New York Times, May 4, 1998. https://www.nytimes.com/1998/05/04/business/media-business-advertising-ford-jettisoning-its-17-year-old-quality-job-one.html. ↩︎ “F. W. Taylor, Expert in Efficiency, Dies,” The New York Times, March 22, 1915. https://timesmachine.nytimes.com/timesmachine/1915/03/22/100146755.html. ↩︎ Debby Wu. “Foxconn Replaces iPhone Business Chief After Tumultuous Year.” Bloomberg. January 17, 2023. https://www.bloomberg.com/news/articles/2023-01-17/foxconn-replaces-iphone-business-chief-after-tumultuous-year. ↩︎ James A. Whittaker et al. How Google Tests Software. (Addison-Wesley, 2012). ↩︎ Jennifer Riggins, “How Google Unlocks and Measures Developer Productivity,” The New Stack, August 17, 2023. https://thenewstack.io/how-google-unlocks-and-measures-developer-productivity/. ↩︎ Walter Isaacson, Steve Jobs (Simon & Schuster, 2011). ↩︎ “Why Is Gorilla Glass So Strong?,” PCMAG, accessed July 11, 2024. https://www.pcmag.com/archive/why-is-gorilla-glass-so-strong-259303. ↩︎ Ken Kocienda, Creative Selection: Inside Apple’s Design Process During the Golden Age of Steve Jobs (St. Martin’s Press, 2018). ↩︎ Steven Levy, “Jeff Bezos Owns the Web in More Ways Than You Think,” Wired, November 13, 2011, accessed via https://web.archive.org/web/20180421105351/https://www.wired.com/2011/11/ff_bezos/all/1. ↩︎ Brad Stone. The Everything Store (Random House, 2013). ↩︎ Tom Randall, “‘The Last Bet-the-Company Situation’: Q&A With Elon Musk,”, Bloomberg, July 13, 2018. www.bloomberg.com/news/features/2018-07-13/-the-last-bet-the-company-situation-q-amp-a-with-elon-musk. ↩︎ Reed Albergotti, “What Went Wrong at Jawbone,” The Information, September 8, 2015, https://www.theinformation.com/articles/what-went-wrong-at-jawbone. ↩︎ Reed Albergotti, “What Went Wrong at Jawbone,” The Information, September 8, 2015. https://www.theinformation.com/articles/what-went-wrong-at-jawbone. ↩︎ Dan Primack, “Notes on Jawbone’s massive failure," Axios, July 11, 2017, https://www.axios.com/2017/12/15/notes-on-jawbones-massive-failure-1513304117. ↩︎ Joel Spolsky, “My First BillG Review,” Joel on Software, June 16, 2006. https://www.joelonsoftware.com/2006/06/16/my-first-billg-review. ↩︎ Paul Allen, Idea Man: A Memoir by the Cofounder of Microsoft (New York, NY: Portfolio/Penguin, 2012). ↩︎ Susanna Ray, “Empathy and Innovation: How Microsoft’s Cultural Shift Is Leading to New Product Development,” Microsoft, news.microsoft.com, February 13 2019. news.microsoft.com/source/features/innovation/empathy-innovation-accessibility. ↩︎ Rajat Agrawal, “Customers, Microsoft employees hack it together at One Week hackathon 2019,” Microsoft Stories India,September 16, (2019. https://news.microsoft.com/en-in/features/customers-employees-collaborate-one-week-hackathon-2019. ↩︎ Simone Stolzoff. “How Do You Turn around the Culture of a 130,000-Person Company? Ask Satya Nadella.” Quartz, February 1, 2019. qz.com/work/1539071/how-microsoft-ceo-satya-nadella-rebuilt-the-company-culture. ↩︎ “Wikipedia: Satya Nadella,” Wikimedia Foundation, accessed September 15, 2025, https://en.wikipedia.org/wiki/Satya_Nadella; Felix Richter, “Chart: Microsoft’s Share Price Surged 10-Fold Under Satya Nadella,” Statista, February 5, 2024, https://www.statista.com/chart/16903/microsoft-stock-price-under-satya-nadella/ (reporting 969% stock price growth since Nadella became CEO); Jeremy Bowman, “If You Invested $10,000 in Microsoft When Satya Nadella Became CEO, This Is How Much You Would Have Today,” The Motley Fool, January 3, 2024, https://finance.yahoo.com/news/invested-10-000-microsoft-satya-150000060.html (calculating 27% compound annual return without dividends). ↩︎ Mary Parker Follett, . The New State: Group Organization, the Solution of Popular Government (New York: Longmans, Green and Co., 1918). ↩︎ Lillian M. Gilbreth, The Psychology of Management: The Function of the Mind in Determining, Teaching, and Installing Methods of Least Waste (New York: Sturgis and Walton, 1914). ↩︎ Steven Kotler, “Why a free afternoon each week can boost employees’ sense of autonomy,” Fast Company, January 20, 2021. https://www.fastcompany.com/90595295/why-a-free-afternoon-each-week-can-boost-employees-sense-of-autonomy. ↩︎ “The Netflix Simian Army,” Netflix Technology Blog, Medium, September 20, 2018. netflixtechblog.com/the-netflix-simian-army-16e57fbab116. ↩︎ Francesca Gino et al. “The Morning Star Company: Self-Management at Work,” Harvard Business School Case 914-013, accessed September 15, 2025. https://www.hbs.edu/faculty/Pages/item.aspx?num=45683. ↩︎ “Case study: Gore-Tex® and W.L. Gore & Associates: An innovative company and a contemporary culture,” University of Portsmouth. https://port.rl.talis.com/page/for-all-seminar-groups-w-l-gore-associates.html ↩︎ Ed Catmull and Amy Wallace. Creativity, Inc.: Overcoming the Unseen Forces That Stand in the Way of True Inspiration (Random House, 2014). ↩︎ “Item 4. Development and Deployment of the Toyota Production System,” Toyota Motor Corporation, accessed September 26, 2025. https://www.toyota-global.com/company/history_of_toyota/75years/text/entering_the_automotive_business/chapter1/section4/item4.html. ↩︎ “Repost: Toyota vs Ford,” Think Different, accessed September 26, 2025. https://flowchainsensei.wordpress.com/2021/12/16/repost-toyota-vs-ford. ↩︎ Geoff Cudd, Ford vs. Toyota Reliability: A Detailed Comparison, accessed September 26, 2025. https://www.findthebestcarprice.com/ford-vs-toyota-reliability. ↩︎ Susan Helper and Rebecca Henderson, “Management Practices, Relational Contracts, and the Decline of General Motors,” Journal of Economic Perspectives Vol. 28, No. 1, Winter 2014, 49–72. https://www.theigc.org/sites/default/files/2016/02/Helper_Henderson_JEP.pdf. ↩︎ Dan Alder, 2024. “How many copies did Portal sell? — 2025 statistics,” LEVVVEL, May 23, 2024 . https://levvvel.com/portal-statistics. ↩︎ Naveen Kumar, 2025. “Steam Statistics 2025: Users, Revenue, & Market Share” DemandSage, September 17, 2025, . https://www.demandsage.com/steam-statistics. ↩︎ Joshua Wolens, “Valve’s reported profit-per-head from Steam commissions is out there, and at $3.5 million per employee it makes Apple and Facebook look like a lemonade stand,” PC Gamer, July 4, 2025, acessed September 26, 2025. https://www.pcgamer.com/gaming-industry/valves-reported-profit-per-head-from-steam-commissions-is-out-there-and-at-usd3-5-million-per-employee-it-makes-apple-and-facebook-look-like-a-lemonade-stand. ↩︎ Steve Farrelly, . “AusGamers Valve Software 2011 Video Interview”. AusGamers, March 28, 2011. Archived from the original on July 30, 2022. Retrieved July 30, 2022. ↩︎ Nathan Bomey, “Barnes & Noble turnaround: CEO James Daunt engineers comeback,” Axios, March 1, 2023, https://www.axios.com/2023/03/01/barnes-and-noble-james-duant-ceo, Lauren Aratani, “‘Amazon doesn’t care about books’: how Barnes & Noble bounced back,” The Guardian, April 15, 2023, https://www.theguardian.com/books/2023/apr/15/barnes-and-noble-bookstores-james-daunt. ↩︎ Jeffrey Dastin, “Amazon to shut its bookstores and other shops as its grocery chain expands,” Reuters, March 2, 2022. https://finance.yahoo.com/news/exclusive-amazon-close-physical-bookstores-185141663.html. ↩︎ Daphne Howland, “As consumers return to stores, why would Amazon shut the door?” Retail Dive, July 25, 2022. https://www.retaildive.com/news/Why-amazon-shut-stores/626972/. ↩︎ Alexa Ray Corriea, “Valve hiring and firing process compared to high school cliques by former hardware dev,” Polygon, July 8, 2013. https://www.polygon.com/2013/7/8/4503326/valve-hiring-and-firing-process-compared-to-high-school-cliques-by. ↩︎ Tracy V. Wilson, “Making Lego Bricks,” HowStuffWorks, February 19, 2021. https://entertainment.howstuffworks.com/lego1.htm. ↩︎ “11 Fun Lego Facts,” LEGOLAND Discovery Center, accessed September 26, 2025. https://www.legolanddiscoverycenter.com/michigan/post/lego-facts. ↩︎

30th Sep 2025 1 votes

More in design

The least wrong colors, version 2

Four years ago, I wrote “How to pick the least wrong colors.” The gist is: picking a categorical color palette is an optimization problem. There’s no such thing as the right colors. But if you use the right cost function, and the right kind of hill climbing, you can at least get the least wrong ones. Since the original post I’ve been slowly picking away at improvements and new approaches. Now that we’re past the singularity, I’ve put a few coding robots on the job. It’s reassuring that many of my assumptions were good ones! The robots have been able to improve the code, bridging some of the gaps in my own knowledge. Today, I’m publishing an updated version of the algorithm as an npm package, along with a fancy GUI version. While there’s still more to do, I’m proud of how far I’ve been able to take it. What’s new New evaluators More controls The public API and a CLI What’s improved The annealing algorithm Configurable color space and distance metric The results One more thing Acknowledgements What’s new New evaluators Almost as soon as I published the first version, I realized that the cost function lends itself really well to modularity. Beyond my initial evaluation functions, I could design new ones, and provide a framework for anyone to plug in their own. As a recap, my original criteria for good categorical colors, mapped to evaluation functions: Similarity — a way of measuring the similarity of one palette to another, useful for providing art direction and getting brand alignment Energy — the colors should be different from each other so they aren’t liable to be confused from one another Range — the differences between the colors should be consistent so unintended groupings don’t appear Color vision deficiency — simulating the colors under different types of color blindness (red-green, blue-yellow, partial to full tritanopia) Here’s the new evaluators: JND — strongly reject palettes that have two or more colors that are too similar Avoid — the mirror image of the similarity evaluation, push colors away from a user-defined set Contrast — compares colors, keeping them above the WCAG AA color contrast floor. Can be used with a background color to maintain contrast on a chart’s background Saliency — uses color naming study data to prefer colors that are easy to name Name difference — the mirror image of saliency, avoiding colors that share names Each of these evaluators can be weighted, indicating the kinds of tradeoffs and priorities you’d like for your color palette. Additionally, the whole evaluator system is pluggable: you can define your own evaluators and have them drive the optimizer! More controls Colors can now be fixed in place, or pinned to a particular order, making it easier to load in existing palettes and optimize all or just some of the colors. Individual channels of each color can be locked, too, meaning you can keep the saturation or hue of a color fixed while optimizing its lightness. This works in any color space. The public API and a CLI The whole package is now a proper library, with a public API. This means: 1. the whole thing is now distributable through npm, with proper versioning, 2. there’s a CLI, making it much more ergonomic for both humans and agents. The API allows for full configuration of the algorithm, as well as loading in colors to optimize. Output can be in raw color values, CSS properties, or DTCG JSON. There’s also a new reportJndIssues endpoint that allows you to evaluate palettes without optimizing them, which is useful to compare a generated palette to commonly-used ones (like Observable, d3, IBM Carbon, and more). What’s improved The annealing algorithm When I wrote the initial algorithm in 2022, I had just learned about simulated annealing. I’ll be honest: I don’t know much more today than I did then. But with AI-assisted research, I was able to solve some questions I had about the initial implementation. Now, the algorithm picks the correct starting temperature based on some random initial samples. Mutation also happens in a scaled manner, so colors change less towards the end of the optimization schedule. Iterations can be capped to prevent very long runs, and the whole thing is much, much more performant. Configurable color space and distance metric The first version of the algorithm worked in RGB space. Now, it defaults to okhsl, but even this is configurable. Individual channels can be constrained to dial in the palette’s boundaries. Also, you can choose which color distance metric you’d like to use (but the library uses CIEDE2000 by default). This flexibility is powered largely by a move from chroma.js to culori. I’ve learned a ton about color spaces since 2022, so being able to mix and match color spaces with distance metrics has been extremely useful. The results The category-colors library reliably produces better results than other palette-generating tools and industry-standard color palettes. Compared to other palette-generating tools, category-colors has more control. Palettailor, for example, optimizes for pure color difference, without accounting for color vision deficiency. QualPal brings some of the optimization parameters, but doesn’t allow for steering towards or away from arbitrary colors. Scores at 8 colors ΔEMinimum ΔEworst of CVD Name differenceMinimum Uniformitylower is better category-colors 22.6 ±1.6 13.7 ±2.0 0.35 ±0.14 best in column 0.30 ±0.02 best in column QualPal 1.1.0 24.7 21.8 best in column 0.10 0.44 Palettailor 26.6 ±2.4 best in column 4.5 ±1.7 0.34 ±0.16 0.34 ±0.04 Colorgorical 15.8 ±3.3 4.1 ±1.6 0.09 ±0.06 0.42 ±0.03 All numbers are at 8 colors. Rows with ± are mean ± standard deviation over 10 palettes; rows without are deterministic and produce one palette. category-colors and Palettailor are 10 independent runs on the same seeds; Colorgorical’s row is 10 palettes from its authors’ own sampling script at equal criterion weights. QualPal was run with CVD on, matched bounds, and takes no seed. Name difference is Heer & Stone’s 1 − cosine; Colorgorical’s own interface reports a Hellinger distance instead. Shaded cells are the best value in their column. Compared to industry-standard palettes, category-colors can produce more optimal palettes, especially at high cardinality. Scores at 8 colors ΔEMinimum ΔEworst of CVD Name differenceMinimum Uniformitylower is better category-colors 22.6 ±1.6 best in column 13.7 ±2.0 best in column 0.35 ±0.14 0.30 ±0.02 best in column Okabe–Ito 21.3 8.8 0.06 0.34 Observable 10 18.4 0.6 0.40 0.34 Tableau 10 18.1 3.2 0.24 0.32 d3 category10 16.2 1.6 0.84 best in column 0.40 ColorBrewer Set3 13.7 1.9 0.16 0.32 IBM Carbon 12.8 5.0 0.11 0.34 Same run: 8 colors, 10 trials. Reference palettes are deterministic, so they're single values. Shaded cells are the best value in their column. One more thing I’ve built a UI that consumes the package and makes it easy to generate and optimize palettes. This has been the biggest request since I published the initial essay, so it’s the thing I’m excited to share. It’s ridiculously overengineered, but hey, what else are personal projects for? Acknowledgements Many measurements come from published research: Gaurav Sharma, Wencheng Wu and Edul Dalal for CIEDE2000; Gustavo Machado, Manuel Oliveira and Leandro Fernandes for the color vision deficiency simulation; Maureen Stone, Danielle Albers Szafir and Vidya Setlur for the size-dependent just-noticeable-difference result; Jeffrey Heer and Maureen Stone, whose color naming models and the c3 data from the Stanford Visualization Group power both the saliency and name-difference evaluators. Existing palettes: Masataka Okabe and Kei Ito’s Color Universal Design set; Matthew Petroff’s sequences; and Mark Harrower and Cynthia Brewer’s ColorBrewer. Other generators laid a lot of the groundwork: Kecheng Lu and colleagues (Palettailor), Connor Gramazio, David Laidlaw and Karen Schloss (Colorgorical), Johan Larsson (QualPal), and Chin Tseng, Arran Zeyu Wang, Ghulam Jilani Quadri and Danielle Albers Szafir (CatPAW). Andrew McNutt, Maureen Stone and Jeffrey Heer’s color-buddy has also been indispensable. Finally, Dan Burzo’s culori made it easy to make this library colorspace-agnostic.

a week ago 1 votes
Mountains of work

This is part of a new experiment I started in an effort to document the process of making Niche design.

a week ago 1 votes
When the canvas starts acting, who’s really in control?

Weekly curated resources for designers — thinkers and makers.

a week ago 1 votes
Gestalt Principles for Visual UI Design

Users parse a layout before they read its labels. Whitespace, borders, alignment, color, and motion determine what belongs together. When these cues fight the content, users attach the label, price, warning, status, or action to the wrong object. Proximity, similarity, enclosure, and the other Gestalt cues guide the eye, snapping visual chaos into clarity.

a week ago 2 votes
The silks of Como: A reminder of the value of craft

I miss this. Visiting a mill is such an enjoyable, often inspiring experience, yet I don’t do it as much as I used to. Perhaps because once you’d done one worsted weaver it’s hard to justify more.  But we’ve never done silk. I did go to Vann... > Read more

a week ago 1 votes
📚 BoredReading

You seem to be enjoying this.

Join free to unlock everything.

Create free account

Already have an account? Sign in