Full Width [alt+shift+f] Shortcuts [alt+shift+k]
Sign Up [alt+shift+s] Log In [alt+shift+l]
1
Small companies are, by default, very transparent. When there are 4 people working in a room, you have a direct line of sight on what everybody else is doing, and why. Your docs, Slack channels, and repositories are open to everybody. When the CEO has an epiphany that changes everything, you all know right away – probably because you were at lunch together when it happened. Thus, startup founders will often get religion about transparency. “Our culture,” they’ll declare, “is to be radically transparent! Everything defaults to open. We hire adults, expect them to do great work, and give them the context they need.” Yay transparency! And this works pretty well. Transparent orgs tend to delegate more effectively, have higher accountability, less politics, faster trust, and just plain ship more. Transparency helps bigger orgs adapt more quickly to the ground truth, responding to customer signals that execs might not be directly exposed to. But, at a certain scale, radical transparency strains. Some idle musing by the CEO sends a team off on an unimportant side quest. A well-justified compensation anomaly upsets a group who is missing background information. A 450-message Slack thread about bike shed paint color choices devolves into factions, hashtags, and philosophical arguments about the morality of taupe. #nevertaupe And if you talk to people at a large yet highly transparent company, you’ll hear about the hazards of the relentless firehose. A thousand shared Slack channels, to start. But also a glut of docs – some critical, most unmaintained. Then there’s the meeting notes, meeting recordings, and meeting invites. Plus proposals, requests for comment, and requests to comment on your proposals’ comments’ resolutions. “So, you like information, eh? Well, have all the information in the world!” How do you make sense of all this? While some people are tenaciously able to find, within this chaos, the important info they need to do great work, a lot of otherwise-capable...
31st Mar 2026

Stay updated

Get a weekly newsletter with the top 5 articles worth reading every week.

More from Allen Pike

Test Coverage Won't Save You

Forestwalk’s CTO Jenn Cooper shares what she’s been learning about tests, after a couple years of increasingly coding with agents: Most discussions about AI-native development jump from this problem – agents’ tendency to accumulate tech debt – directly to tests. … Tests verify that code does what it did before. Whether what it did was even the right way to do it is a separate question. She argues that while agents make it easy to have rigorous traditional test coverage, at best unit tests maintain local code cohesion. At worst, they can actually make it harder to improve what agents are worst at: the wider coherence of the entire codebase. So far I’ve been impressed with how effective the broader automated checks she describes can be to guard against agentic nonsense.

8th Jun 2026 1 votes
Building for Voice In, Visuals Out

Recently, Andrej Karpathy argued that the ideal interaction pattern for AI models is voice in, visuals out: Audio is the human-preferred input to AIs, but vision is the preferred output from them. Around a ~third of our brains are a massively parallel processor dedicated to vision; it is the 10-lane superhighway of information into brain. The claim is that while “text in, markdown out” is the mode most people use LLMs today, what we should be building toward is a Jarvis-like mode where we primarily speak to AI – and it primarily responds with UI, video, or other visuals. Let’s check in on where we’re at for both halves of this claim: visuals as output, and voice as input. Visuals Out Humans love looking at things! While it can be convenient to be able to listen to our computers speak, waiting through a voice response feels kinda… ugh. You can increase the speaking rate, but fundamentally, the fastest way for a computer to give humans information is to display it. We’re faster at reading text than we are at listening, but that’s just the start. There’s a good reason computers long ago evolved past text-only terminals: richer interfaces are often faster, clearer, nicer, and more useful. The power of human vision has facilitated a rich history of computers showing people stuff. At first, LLMs weren’t great at producing visuals, often spending many tokens to produce half-baked results. However, Anthropic’s Thariq Shihipar recently wrote how HTML is increasingly a viable output format to supplant Markdown, for certain model responses. This is great, since HTML is a powerful way to show visuals. Going beyond text can give us dynamic: Hierarchy (sidebars, columns, navigation) Exploration (drill ins, filters, expansion) Direct manipulation (scrolling, dragging) Data visualizations (graphs, charts, dashboards) Mockups and prototypes (show, not tell) Illustrative images and video (pelicans, bicycles) Thus the DOS era of AI begins to end. While it will be a while before general-purpose agents consistently return compelling HTML in response to arbitrary requests, visual responses are already practical for vertical agents – it helps to do one thing well. Recent months have seen a noticeable uptick in AI features producing useful diagrams, charts, sliders, and so on. So, yep. Visual output is a natural fit for AI, and we’re already going beyond plain text. Voice in On the other hand, most people are ambivalent about the idea of talking to AI. We were promised the Star Trek computer, or Jarvis, but so far we’ve gotten Siri and automated spam calls. There’s merit to the skepticism. Fundamentally, voice is never going to be the only input mode for computers. Just as we sometimes need voice because our hands are occupied, other times it’s impractical to speak aloud for social or confidentiality reasons. And even when we can speak, voice alone isn’t enough – effective computer use will always require more precise inputs, such as mouse clicks and drags. However, voice is a deeply human and useful input mode. For example, it’s excellent for getting out our not-yet-organized thoughts and observations. While ChatGPT voice mode is substantially dumber than its text mode, it can still be useful for organizing your thoughts – advanced rubber-ducking. Compared to text, speech also contains additional nuance and detail. Voice is not just words – it’s intonation, timing, tone, pitch, energy, and emphasis. Where a transcript would only see okay, how you voice the “okay” might convey “Sounds good!”, “Tell me more”, “I kind of doubt that.” or “Get the hell out of my office.” This is why we call somebody if we need to have an emotional conversation, rather than sending misinterpretable text messages. We speak faster than we type in terms of WPM, so together with the additional details in our voice, we simply put out more information per second via voice than from a keyboard. The Tyranny of Latency So, great. Talking to AI and having it respond with visuals are both natural and highly useful. Why aren’t we doing this all the time? If you’ve actually used AI voice systems, you’ve probably noticed that they’re usually slow, dumb, or both. In order to feel fast, we’ve known since the 60s that computers should respond within about 100ms, and that in order to keep users’ sense of flow, they need to respond within about 1000ms (1 second). Even before networks and giant neural nets, it could be a challenge to hit these bars. But voice AI adds a substantial new hurdle. Humans are more sensitive to lagged voice than we are to lagged visuals. For a fully fluid voice conversation with interruptions going both ways, the latency bar is about 200ms. More than that, and interruptions feel janky and annoying. You’ve experienced this on voice calls with other humans: if there’s a noticeable lag and you’re stepping on one another’s words, you back off into a more stilted turn-taking conversation style. At best, this is what we get with common AI applications today: slow, single-duplex turn-takers. They listen until it seems like you’ve stopped, generate a response, then stream until it sounds like you’ve started saying something, at which point they abruptly stop. While 200ms is a long time in traditional computing terms – a smooth animation frame needs to render in just 16ms – you’ll find 200ms is not a long time to do the complex work of sending a user’s voice over the network, making sense of it, generating a voice response, and sending it back. In order to achieve the required latency, applications generally do voice inference with rather small models. The most advanced voice model most people have tried, ChatGPT’s rather outdated voice mode, is profoundly dumb compared to GPT 5.5 or Claude Opus 4.8. Even if you understand why this is the case, it’s fun to watch that guy who awkwardly gets it to misadvise him1. But there is hope. Earlier this month Thinking Machines gave a preview of their approach for realtime voice models, which they call Interaction Models. These are full-duplex systems, which means we’re finally getting simultaneous perception and generation. Rather than switching between generation and listening, these streaming models slice time into 200ms chunks, interleaved continuously. While 200ms isn’t enough to generate a very smart response, that fast streaming model can call slower, smarter models to do things like lookups, reasoning, and generating artifacts – then return the results in 200ms chunks when they’re ready. Now, this is all very exciting, and I’m excited to see where it goes. But despite the claim “The model instantly reacts to visual cues”, even their demo videos show a noticeable and sometimes awkward lag between stimuli and voice responses. This is partly because it’s early – Thinking Machines was only founded last year. But it’s partly because humans are just that sensitive to voice delays. It’s a fundamentally difficult problem. However. Humans are less sensitive to laggy visuals. Since visuals are less intrusive than a voice response, you get the more permissive 1000ms response budget that we’re used to when building computer programs. This is convenient, since voice → visuals is a great interaction mode. Voice In, Visuals Out The good news is that you don’t need to wait for Thinking Machines or any other model advances to build useful voice in, visuals out experiences today. Here’s a quick example of what voice in, visuals out can feel like: not a chat, but a live visual representation of what you’re working on. The Cedarloop voice agent can help outline notes, file bugs, and do other in-meeting work. Here are a few latency approaches to keep in mind if you’re working on voice-in, visuals-out agents: The underlying model needs to be very fast. Any slower than p50 latency of 700ms and p95 of 1200ms will feel janky. Meanwhile, it’s common to see small requests on “fast” models that have over 5000ms of p95 latency 🫠 You need to send uncomfortably short time slices for inference. Err on the side of sending incomplete text rather than waiting for two-second pauses, and use context engineering to have the model heal any errors. Keep your context prefixes stable, so they can be well-cached. 90%+ of our input tokens are cached, and thus far faster (and cheaper) than if we were sending fresh context every request. Tokens are slow, and HTML is token-heavy. Realtime visuals-out needs to use efficient formats out of the LLM, which can then be displayed in a rich web or native view. Get it dialled in right, and you can build delightful-feeling experiences. If you’re working on these kinds of realtime apps, I’d love to chat – happy to share what we’ve been learning, and hear what others have been finding. GPT-Realtime-2 recently launched in the API with “GPT-5-class reasoning,” but is not in ChatGPT yet. And so far, Claude has no realtime multimodal model at all. ↩

31st May 2026 1 votes
We Can Do Hard Things

Years ago, back when I was leading a mobile dev team, my friend had an idea for a business. You see, back then the most frustrating thing about mobile dev was the final step: getting your app on actual phones. Builds, provisioning, and code signing made for a harrowing trial, festooned with obtuse errors and other sharp spikes. So, Dennis had a pitch for me. “What if,” he asked, “we did all your apps’ builds and provisioning and signing for you, in the cloud?” I raised an eyebrow. “Well, obviously that would be great. In theory. But it would be too annoying to build that. Apple drops Xcode versions and switches submission requirements with no warning. And you’d need to make sure that…” He stopped me with a wave. “Right, but: if we did it, and it worked. Would you use it?” “Well, of course we would. But I don’t think you want to run this.” My attempt to discourage him didn’t work. Perversely, the idea that this was a hard problem got him more excited. He immediately dove in. Three years later, Buddybuild was acquired with fanfare. They’d accomplished what they set out to do, made a tidy profit, and they were even able to keep theirgreat team here in Vancouver. Wisely they ignored me, and chose to do the hard thing. The Nice Thing About Hard Things Doing something hard yet pointless is foolish. But doing something hard yet valuable has a lot of benefits. It’s easier to recruit a great team to tackle hard, worthwhile problems. It leads to less competition, due to schlep blindness. It’s a great way to hone your ambition and discipline – over time, working on hard things feels less hard. Consider that. If you have a great team, less competition, but more ambition and discipline, then you’re set up to do well. These days are well suited to attempting hard things. Our tools are improving so fast that a project which seemed straightforward last year might be trivial next year. Better to dial up the ambition a bit. Of course, there are a few pitfalls to trying hard things. You’re more likely to burn out, for one – it’s very important to sleep, exercise, and manage your own energy when your work is kicking your ass. And it can sometimes be difficult to tell when the “hard and purposeful” parts end, and when the “overcomplicating things” or “naive folly” begins. I highly recommend having a co-founder that finds hard and purposeful problems motivating, yet takes a dim view of overcomplication. Doing hard things is best not attempted alone. But, all in all, it’s a good default. We can do hard things. So, let’s.

30th Apr 2026 2 votes
Maggie Appleton on Gas Town and Coding Agent Orchestration

Maggie was already perhaps the best writer on the intersection of engineering and design, but now that she’s joined Github Next, she’s also extremely keyed in to where tools for coding are going. Her piece on Gas Town and orchestrating coding agents is sharp and worth reading in full. As the pace of software development speeds up, we’ll feel the pressure intensify in other parts of the pipeline: thoughtful design, critical thinking, user research, planning and coordination within teams, deciding what to build, and whether it’s been built well. The most valuable tools in this new world won’t be the ones that generate the most code fastest. They’ll be the ones that help us think more clearly, plan more carefully, and keep the quality bar high while everything accelerates around us. We’ve known for a couple years now that faster coding will mean non-coding work will increasingly be a bottleneck, and now it’s happening. Deciding what to build – and whether it’s been built well – was already one of the most important tasks on a software team. But in the face of tools that can add anything to your product, desirable or not, this judgement becomes the core of the work.

13th Feb 2026 1 votes

More in startups

Divorced from reality

Fundamentals will matter

6 days ago 2 votes
Tuesday

My breathing was a constant problem.

6 days ago 2 votes
Industry roundup #13

When "buy, don't build" stops applying to dev tools

a week ago 1 votes
AI, tools and transformation

It’s very tempting to imagine that AI turns everyone into a tool-builder - now everyone can just ask the model to make the software they need, and apps as we know them are dead. I think that misunderstands how most people think and where software actually comes from, and more importantly, it isn’t a path to change how companies actually work.

a week ago 2 votes
Paralyzed by the Perfect Resume

Those with the most to lose are often the least likely to take big swings

a week ago 2 votes
📚 BoredReading

You seem to be enjoying this.

Join free to unlock everything.

Create free account

Already have an account? Sign in