More from haseeb qureshi
@SemiAnalysis_ recently found something bizarre in the economics of AI coding subscriptions. If you run them at max usage limits, you’re actually paying 20x-70x cheaper than you would buying tokens through the API. Many people looked at this and said: oh my god, look how much the labs are subsidizing tokens, the bubble must be about to pop soon. This is the wrong response. The reason why labs are willing to offer such generous plans, of course, is because most users are rarely hitting their usage limits. The product works like a gym membership: the limit is generous because most people barely use it. But I’ve spent a lot of time thinking about this, and it’s true that something weird is going on here. We don’t know what their actual blended margins are on subscriptions, but SemiAnalysis estimates that at 20% average utilization, Anthropic breaks even on their Max 5x plan. 20% utilization is probably on the high side, especially in orgs where everyone (including non-coders) have subscriptions and are only busting it out once in a while. Most places I know, including Dragonfly, give out Claude Code subscriptions liberally and encourage non-coders to experiment with it. But what SemiAnalysis doesn’t dwell on here is that this is exclusively a small company phenomenon. The subscription pricing model is not available to large companies. Here’s why: at 150+ people, you are forced off the subscription model, which is known as the “Team” plan. You have to switch to “Enterprise,” which is priced as $20/seat base, plus API pricing per token used. Enterprises must pay linearly based on token costs, and SemiAnalysis believes API tokens are priced at roughly 75% gross margins. This is a massive price hike that kicks in suddenly at 150 seats. So if you’re a small business or a startup (or a personal user), you have a distorted view of AI spend. Your token pricing is actually very generous, and Anthropic may be running at low or even negative margin on you. You might have wondered why Microsoft and Uber are freaking out about token spend and talking about “token-minning.” This is why. They pay structurally higher costs per token than startups and individuals do. But Anthropic doesn’t care! Max extracting from small companies or individuals just doesn’t matter much for a B2B company. If you look at companies like Datadog or Cloudflare, they make 80-90% of their revenue from large (100K+ ARR) contracts. Making 0 margins on the long tail is just a customer development cost. This is the standard B2B sales way to think about this pricing strategy. But there’s another way to think about this same situation: through the lens of tax policy. Because if tokens are replacing labor, then the gross margin that OpenAI and Anthropic collect on tokens is effectively a tax on AI labor. There are two major consequences to thinking about token pricing this way. Token Pricing as Tax Policy Let’s assume the margins stated in the SemiAnalysis piece: breakeven on subscriptions, 75% gross margin on API for BigCos. The instinct is to call that a 75% tax on AI labor for large organizations, and 0% tax for startups. Standard tax analysis would say this is a disincentive to use AI labor within large companies, which pushes at the margin more toward less automation and retaining more human labor. (It obviously also incentivizes using smaller/open models, but the net effect is that it incentivizes both. Remember, we’re thinking at the margin here.) But the part that drives behavior even more strongly is not the average rate. In tax policy it never is. What we care about is the marginal rate. And for startups on a flat-rate subscription, the marginal price of the next token, up until the usage limit, is zero. And a zero marginal price is the most distortionary a policy can possibly be. For a startup, the subscription model is basically an innovation subsidy. The overwhelming incentive is to experiment how to spend the entire token budget as effectively as possible. That means running Ralph loops, papering your screen with Claude Code sessions, and orchestrating swarms of agents. Exploration is free until you hit the usage limit, so startups are effectively competing to squeeze every last drop out of their subscriptions to out-produce their competition. Perversely, the more you use, the lower your average token price is. Each startup wants to be the one that makes Anthropic lose the most money on their subscription. BigCos face the opposite incentive. If you’re beyond the 150-seat threshold, every token of exploration is billed at full markup (with 75% surcharge!), so they’re punished linearly for exploring the frontier. BigCos will still automate the obvious high-volume tasks, but the marginal, experimental, risky automations never get found because the discovery cost is too high. This tax structure ultimately pushes them toward keeping more human labor and maintaining the same overall org structure. It’s like a reverse Japan. Japan has a massive labor shortage due to its declining population. Historically this has meant Japan has pursued high degrees of automation, because high labor costs incentivize automation. That’s why Japan has robots in restaurants, factories, hotels, and hospitals. But weirdly, big companies find themselves in a reverse Japan situation: if they are paying very high taxes on AI usage, this creates LESS incentive to automate, and more incentive to retain the humans they already have (even more so if wages stagnate in the meantime). So where does the labor displacement go in this model? Everyone is watching the big companies for waves of AI layoffs. But at 75% rates, replacing your own workforce too aggressively with AI might just be uneconomic. The token budgets just explode. But that doesn’t mean the displacement never happens. It just means the displacement shows up in a different shape. When BigCos lose market share to AI-native startups that carry a fraction of the all-in labor costs, that will trigger layoffs as BigCo revenues and stock prices decline. But those jobs that are eliminated are never replicated at the startups who win the day. The net disemployment effect is the same, the air pocket just moves to a different line item within the economy (where the AI tax rate is lower). This is also why “AI-washing” might not be a temporary phenomenon. AI-washing is when a company attributes layoffs to newfound AI efficiencies, when it’s actually just an excuse for ordinary business weakness. Many assume that this is a fad of the current AI hype cycle. But while everyone is primed to watch for big companies doing true AI layoffs “replacing jobs” with AI, it may never actually happen at scale. The labor displacement may happen instead through startups outcompeting the BigCos, the BigCos AI-washing all the way to their graves, and the startups never re-creating the old jobs. The job displacement will still happen, just not where everyone is looking. So that’s the first consequence of this model. But there’s also a second, weirder consequence. The Notch A regulatory notch is a regulatory threshold that incentivizes a large discontinuity in behavior. Example: 30 hours a week for full-time employment incentivizes a lot of jobs that are exactly 29 hours/week. Famously, France has extremely demanding labor regulations that kick in at 50 employees (work councils, mandatory profit-sharing, firing protections), which are exempted for small companies. This results in massive incentives for employers to stay below the 50-person notch. Extend this analogy to AI. The big labs have created a tax notch that punishes companies for going above the 150 seat threshold. This means you must stay small to keep your beautifully subsidized subscription pricing, and be taxed ~0% (or negative) on your tokens rather than 75%. This might result in a totally new philosopy of company management. Startups will increasingly obsess over agents for everything, smaller teams, frequent firings, more subcontracting, and doing everything possible to map the lowest possible human surface area. Not because it’s the “optimal” amount of automation, but because the incentives drive them there. If the magic number is 149, every seat counts, and you can’t afford to waste humans outside of the essential joints of the company. This discontinuity may be perceived by Harvard Business School types as “the new generation of AI-first management.” But understood properly, it’s actually just a rational response to enterprise pricing plans. This might sound like a bit much. But you can already see the behavior differences between different organizations. Talk to developers at BigCos, and they are meticulously counting tokens and getting more nervous about their leaders slashing token budgets. But devs at startups are breathlessly tokenmaxxing, spinning up swarms of agents overnight and checking their logs in the morning. I expect this dynamic to accelerate. No one designed this. There is no committee deciding to subsidize innovation for startups and tax it for incumbents. All this fell directly out of well-worn enterprise pricing strategies. But this is how tax codes always look: a pile of incidental rules that ultimately determine which companies get built and how those companies contort themselves to minimize their tax burdens. You could object that this is temporary, and the labs will meter everyone eventually. Github Copilot has already made the switch. Maybe, maybe not. But by the time pricing normalizes, the 149-person company and the new school of AI-first management may have already blown up, gobbling up market share, and writing the playbook for the next generation of startups. Tax policies matter. The entire notion of the “gig economy” exists because of the legal boundary between W-2s and 1099s. As more labor gets eaten by AI, token pricing may be the most consequential tax policy of the next decade. Yet nobody will ever vote on it. (And don’t be surprised if the fastest growing companies of the next cycle all conspicuously cluster at 149 seats.) Originally published on X, June 2026.
We’re a crypto fund. If anyone should believe in crypto, it’s us. And yet, when we sign a deal to invest into a startup, we don’t sign a smart contract. We sign a legal contract. The startup does the same. Neither of us are comfortable doing the deal without a legal agreement. Why? We have lawyers. They have lawyers. We have engineers who can write and audit smart contracts, and so do they. We are two sophisticated crypto-native parties, and we still don’t trust a smart contract to be the only binding agreement between us. I literally was a software engineer, and I still trust the legal contract more–because if there’s an issue with the legal contract, I know the judge will do a reasonable thing. The EVM, not so much. In fact, even in the cases where we have an on-chain vesting contract, there’s usually also a legal contract in place. You know, just in case. When I first got into crypto, there was this fantastical story that crypto would replace property rights. Instead of legal contracts, we’d all use smart contracts. Instead of agreements enforced by courts, they’d be enforced by code. It didn’t happen. Not because the technology doesn’t work, but because the technology doesn’t work for our society. Let me make a confession. I’ve been in this space for a decade and I’m still scared every time I sign a large transaction. I’m rarely scared to approve a large bank wire. The bank, terrible as it is, was designed for humans. It’s really hard to mess it up. There are no address poisoning attacks at banks. There’s no reason why my bank would ever allow me to send $10M to North Korea–but to Ethereum validators, there’s no reason why my address wouldn’t be sending $10M to North Korea’s address. The banking system was specifically architected with human foibles and failure modes in mind, refined over hundreds of years. Banking is adapted to humans. Crypto is not. That’s why in 2026, it’s still terrifying to blind sign a transaction, to have stale approvals, or to accidentally open up a drainer. We know we should verify the contract, double-check the domain, and scan for address spoofing. We know we should do all of it, every time. But we don’t. We’re human. And that’s the tell. It’s why crypto always felt slightly misshapen for us. Long unreadable cryptographic addresses, QR codes, event logs, gas fees, and footguns everywhere–none of it conforms to our intuitions about money. That’s when it clicked for me: it’s because crypto wasn’t built for us. Crypto Was Made for Machines An AI agent doesn’t get lazy. It doesn’t get tired. It can verify a transaction, check every domain, and audit a contract in seconds. And more importantly, an AI agent trusts code more than it can trust the law. I trust the law more than I trust the smart contract. But to an AI agent, a legal contract is actually much less predictable. Think about it: How will I drag my counterparty into court? In what jurisdiction will this contract be adjudicated? What if the legal precedent is ambiguous? Who will we draw as a judge or jury? There is so much uncertainty baked into law that it’s impossible to know with certainty the outcome of an edge case. And that dispute takes months to years to resolve through the legal system. For humans, that’s basically fine. In AI agent timeframes, that’s an eternity. Code is the opposite. Code is closed form, deterministic. An AI agent looking to make an agreement with another agent can negotiate multiple rounds of terms on a smart contract, statically analyze it, formally verify it, and enter into a binding agreement–all in a few minutes, all while the humans are asleep. In that sense, crypto is self-contained, fully legible, and completely deterministic as system of property rights around money. It’s everything an AI agent could want from a financial system. What we as humans see as rigid footguns, AI agents see as a well-written spec. Even legally, our traditional monetary system was designed for human institutions, not AIs. The traditional monetary system only recognizes humans, businesses, and governments as legitimate holders of money. If you are not one of those three entities, you cannot own money. Even if you rig up an AI agent to interact with your bank account on your behalf, then what? How do you run AML on an AI agent? Suspicious activity reports? Sanctions violations? Where does liability fall if the agent is acting autonomously? Does the liability change if it was manipulated? We haven’t even begun answering these questions–our legal system is totally unprepared for non-human financial actors. Crypto asks no such questions. It doesn’t need to. A wallet is a wallet, it’s just code. An agent can hold funds, transact, and enter into economic agreements as easily as it can send an HTTP request. The Self-Driving Wallet This is why I believe the crypto interface of the future is what I call a “self-driving wallet”–entirely AI-intermediated. You won’t be going around websites clicking buttons. You’ll instruct your AI agent to solve financial problems for you, and it will navigate the services available (e.g. Aave, Ethena, BUIDL, or whatever succeeds them) to build the right financial solutions on your behalf. You won’t do it it yourself; an AI agent that is natively fluent in this world will do it for you. And when agents are the primary interface into crypto, the way those protocols market and compete with each other will have to radically change. And beyond acting on your behalf, agents will transact with each other. When agents can discover other agents and enter into economic agreements autonomously, they will prefer crypto. It works 24/7, 365, anyone-to-anyone, fully in cyberspace. It can’t be turned off. It’s completely self-sovereign. This is already happening. Moltbook has agents finding and collaborating with each other across geographies, with no knowledge of who owns them or where they sit. And just yesterday, @0xSigil’s @ConwayResearch has built self-sovereign agents that survive completely autonomously using crypto wallets, working to earn their own compute costs to stay alive. The future is going to get increasingly weird. And crypto is going to be part of that weirdness. So what’s the takeaway? I think it’s this: crypto’s failure modes, which always made it feel broken for humans, in retrospect were never bugs. They were simply signs that we humans were the wrong users. In 10 years, we will look back at amazement that we ever subjected humans to wrestle with crypto directly. This change won’t happen overnight. But a technology often snaps into place once its complement finally arrives. GPS had to wait for the smartphone, TCP/IP had to wait for the browser. For crypto, we might just have found it in AI agents. Originally published on X, February 2026.
It’s that time again—as 2025 comes to a close, it’s time to drop 2026 predictions. I think 2026 is going to surprise, both to the upside and to the downside. Organized by category: Macro / Chains $BTC is > $150K by year-end, but BTC dominance decreases in 2026. Despite the excitement around the recent crop of fintech chains, their metrics will underwhelm. Daily active addresses, stablecoin flows, and RWAs—Tempo, Arc, and Robinhood Chain will underdeliver, while Ethereum and Solana will overdeliver. Best developers will continue to build on neutral infra chains. A big tech company (Google, Facebook, Apple, etc.) launches or acquires a crypto wallet in 2026. Many more Fortune 100s launch blockchains, although increasingly concentrated among banking and fintech players. Expect Avalanche to be a standout here, alongside OP stack, Orbit, and ZK Stack. Monad gets written off as dead by CT, but metrics take off in the latter part of the year after analysts have already forgotten about it. At least 3 other chains connect to DoubleZero to improve their latency & throughput metrics. DoubleZero hits 80%+ stake on Solana. DeFi Perp DEX market share consolidates to something like 3 big venues a la HBO (market share something like 40 / 30 / 20), followed by a long tail of smaller players who compete over the leftovers (last 10%). Equity perps take off, becoming >20% of total DeFi perp volume by EOY. Significant growth in RFQ compared to CLOBs/AMMs, both on spot and perps. Some DeFi-related insider trading scandal hits mainstream media. Stablecoins Stablecoin supply expands by ~60% in 2026, and USD remains 99%+. USDT dominance declines moderately to ~55%. Stablecoin-backed cards grow 1,000% in 2026—insanely fast growth. Becomes the dominant way that stablecoins land and expand in emerging markets. Rain is the biggest winner here. Regulation Clarity Act gets signed into law in 2026 after some significant markups and horse trading. A bit of buyer’s remorse from crypto insiders. Dems win the house, and there is a parade of hearings about anything in crypto that touched $TRUMP / $WLFI. The underlying deals get subpoenaed. Trump insists he was never involved and didn’t know anything about it (and thus these deals are not protected by executive privilege). Anyone who signed a stupid deal gets publicly embarrassed. Prediction Markets Prediction markets grow like crazy. Big legal fights over sportsbetting regulation and federal pre-emption, but nothing major gets resolved next year, so status quo continues through 2026. Meanwhile Polymarket continues to steamroll the culture. Prediction markets are perceived as cool and smart, and so are allowed to throw up odds everywhere. As Polymarket domestic expansion gets going, it starts winning more and more domestic market share from Robinhood and sportsbooks. The explosion of other platforms tacking on prediction markets mostly flop. 90% of prediction market offerings are totally ignored and then wind down by EOY. B2B partnership-driven distribution underperforms, direct-to-consumer outperforms. Almost all of the demand in 2026 is sourced directly from Polymarket, Robinhood, and Kalshi frontends (plus traditional sportsbooks). AI Primary AI use cases in crypto remain within software engineering and security. Everything else remains a prototype. No good solutions to the spambot proliferation on social platforms emerges. A lot of stuff is proposed, but mostly we just eat the AI slop for 2026. Eventually it will get bad enough that people align on a solution, but not there yet. Wallet automation remains minimal. AI agents will still not be “paying each other” or spending any meaningful money in 2026. We see more small teams (<10 people) shipping scaled products because of coding agent force multipliers. In 2025, you needed to be Hyperliquid-level cracked devs to be this dev-efficient. In 2026, you just need to be AI-native and versed in the modern agentic stack. 2026 is dubbed the year of the agentic startup, and it hits crypto startups in a big way. AI becomes used for both attack & defense in cybersecurity. We see many more hacks in 2025, but smaller sizes. Defensive AI gets integrated into CI/CD pipelines and much better continuous monitoring. Security posture across the board improves, even for small teams, and the total amount hacked decreases compared to 2025. So those are my predictions! If I had to summarize them to a two meta-theses, it’d be: slow and steady beats new and shiny the trend lines mostly continue Let’s see how I do. Keep me honest, CT. Disclosure: I’m an investor in many of the assets mentioned. NFA. DYOR. Originally published on X, December 2025. Covered by CoinDesk.
Here’s what I would do if I was a young person trying to break into VC: Write. Short writeups, on Twitter. Not generic market philosophical thinkpieces, because those will be assumed to be AI slop or regurgitated research. No one will read it unless you’re brilliant, which you’re probably not. Original research, on a specific company or sub-sector. If you want to write about robotics, even that is too broad. Narrow it down. Humanoid robotics, or healthcare robotics, military robotics, etc. Get really granular. So granular most people won’t care. If it’s something you could get by Googling, it’s not narrow enough. You will not be able to find to do “original research” easily. This is not something you can do from a university library. You will have to go talk to people who work at these companies. Journalists who cover these companies. Pay for private industry-specific research / newsletters. Follow all of the employees/anons who are tweeting gossip. Integrate a picture that someone reading TechCrunch doesn’t see. Then write about this sector and leading + new startups and tag / DM every investor at every major firm who covers your space (you can find them because they’ve invested in one of the companies in the sector). If they express interest, offer coffee meetings with everyone you can. Some will take you up on it. Do this enough times, you’ll develop a reputation and get offered a job in venture. Don’t need to go to business school, don’t need to have a great angel portfolio or any of the above. “Get good deal flow” is wonderful if you have access to it, but most people just can’t do this. If you’re already surrounded by Stanford undergrads, you probably don’t need advice to break into VC. But the above strategy–in principle anyone can do. Just need to have abnormal levels of agency and a willingness to basically do the job of a junior VC without anyone telling you to. (While you’re doing this, best thing to do in the meantime is to also work at a company in the sector you’re chasing after. But not always possible depending on your background. Thankfully, VC does not require any particular background. Lots of weirdos in VC, myself included.) I guarantee you, everyone wants to hire someone who can do the above. But very few candidates have this degree of agency. VC is not a “tracked” career. Hiring is arbitrary, firms are generally small and do not scale, and there is no standard path. This is good for you if you’re willing to be weird. The thing that VCs have in common is that they are passionate about startups and understanding new industries. If you show that you already have that, a path will open for you. Originally published on X, November 2025.
More in startups
I’d love to hear all of the ideas you guys have here, including relatively extreme things that might sound crazy at first.
Lately I've had a lot of time for thinking. Partially because I shut down Blymp back in January and freed up a lot of my mental resources. No clients to follow up with. No admin stuff to stress about. But thinking is also my favourite activity, and there'll always be time in my day for a good old mind-bending. So I sat down, as I often do, alone with my thoughts, and wrote about what business I should and, most importantly, should not consider doing next. The list you are about to read might be similar to Core principles I (try to) live by that I wrote one similarly pensive evening two years ago. Whether it's a comparison or a continuation is hard to say—still, one can't be without the other to show the inexorable passage of time that changed everything and nothing for me all at once. But it also serves another purpose: to remind me in the future, before I get myself involved in some dubious enterprise, what painful mistake I'm about to make by ignoring my values. So here it is, the list. My next business shouldn't (and hopefully won't) be about: Social Media. Enough of this crap. Some productivity bullshit. Do better, not more. Fast fashion and consumerism. Truly, I've already bought everything you wanted to sell me. Some “hack your health” app. Our bodies haven't changed much in the last three hundred thousand years, and they won't change in the next hundred. Or any other app, really. I don't even use my phone anymore. Distracting things and things that require constant attention. Some LLM wrapper with a fancy UI. Indefensible. The number of things I can do with ChatGPT or Claude is ridiculously high. It can surely handle one more thing. What it should (and hopefully will) be about: Sustainable, high-quality products Unscented products Building community Bringing people together, offline Empowering creativity Replacing animal products Doing one thing really damn well Small acts of kindness Things you can touch Things that last Clothes made of natural fibres Art Fun Questioning the status quo Simple, intentional living Deeper understanding of self Spending more time in nature Spending more time with loved ones Things and businesses I get inspired by: Framework: Sustainable/repairable laptops. Bitwarden: Open-source software that does one job really well and charges me a reasonable amount of money per year. Danish design: Beautiful. Sturdy. Timeless. I brought home two Royal Copenhagen mugs the other day. Bike-sharing and car-sharing. Literally anything-sharing. ZSA keyboards. What a keyboard should have always been. A water bottle that won't leak on a plane. Some random-brand bottle bricked my friend's MacBook, and I've been appreciative ever since of the fifty-dollar water bottle that I bought years ago, so hesitant about the price. Coffee. One of those timeless things on Earth. Kobo. It's like Kindle that doesn't decide what's best for you. Upload your own PDF. Or EPUB—whatever. It has physical buttons to flip pages. Best $200 I ever spent. A high-quality safety razor. It's just so nice to hold. A Japanese stainless-steel knife. So nice to hold, too. There are numerous other things that I appreciate having in my life that didn't make it into this list for some reason or another. A familiar mom-and-pop shop in my neighbourhood. A Timemore Black Mirror kitchen scale that just works, every time. A random USB-C charged electronic device that spares me from carrying an extra cable. My radically outdated ten-dollar Casio watch. An old pair of comfy shoes that just won't die. We need more of this in our lives. Things that you desperately look to buy again when they get lost or break or fall apart. Things that someone made a deliberate effort to get right the first time. The quintessence of art and craftsmanship. For all the genius of Steve Jobs, the iPhone wasn't that. It challenged the status quo and was surely a groundbreaking, outstanding piece of technology at the time of its first release. But in twenty years, people won't remember it ever existed. Like Gen Z doesn't remember the Walkman. I'd like to see more businesses that bet on doing one thing really well. Google could still have been the company people remembered for the best search engine if they had doubled down solely on that. Instead, we don't even know what they do anymore: Phones? Clouds? Ads? Certainly ads. I still remember the feeling of holding one of the first PocketBook e-readers in my hands back in Moscow. Pressing its buttons and waiting for what now seems like a torturous three seconds before the screen refreshed. Almost twenty years later and every day still, Kobo gives me exactly that feeling. So hopefully, my next business is the one that lasts.
No, Transformers Won't End the Human Race lol In 2022, I used to get calls from journalists asking, with great sincerity, what our lives would look like in the metaverse. How would we work, socialise, buy property, and fall in love once we had all moved there? The crypto questions followed the same pattern. How would governments collect taxes when tokens displaced national currencies? How long until the dollar collapses? What would geopolitics look like once blockchain DAOs had dissolved nation states? Almost nobody called to ask whether any of this could or would happen, or how. Some CEO, VC, or portfolio manager had announced the inevitable future, and the questions began from there. The imagined future arrived inside the grammar of the question. "What happens when?" quietly replaced "By what mechanism?" We skipped over technical feasibility, economic demand, institutional adoption, and political consent, then began writing books and decorating the future world on the other side. In February 2022, Gartner forecast that a quarter of people would spend at least an hour a day in the metaverse by 2026. The World Economic Forum repeated it under the headline "We will be spending an hour a day in the metaverse by 2026. But what will we be doing there?" The first sentence retained a conditional. The second was already arranging the itinerary. The metaverse acquired property law and zoning disputes before it acquired residents. Banks opened virtual lounges nobody visited. The books from the period (The Metaverse: And How It Will Revolutionize Everything, Step into the Metaverse: How the Immersive Internet Will Unlock a Trillion-Dollar Social Economy) now read as artefacts of a collective fugue state that briefly acquired ISBNs. Now it is 2026 and the metaverse is dead. Good riddance. This time the journalists are all writing about the new hotness, which is whether the machines will kill us all. And we have collectively memoryholed that we literally just did this. Michael Crichton had a name for what happens to a reader here. You open the paper to a story on a subject you know well, and you find it backwards. Wet streets cause rain. You shake your head, turn the page, and read the next story, on a subject you know nothing about, as though it were written by someone else. He called it Gell-Mann amnesia. The metaverse was the page we all agree was nonsense. Artificial intelligence ending the human race is the next page, and we are being asked to turn it without remembering that we just did this. I call this techno-inevitabilism, the habit of the professional managerial class of treating a proposed future as settled before anyone has established the causes that would bring it about. Its dual, and comorbidity, is tech psychosis, in which the chattering class loses contact with causality in the presence of a sufficiently fashionable technology, and asking whether the machine works marks you out as a dreary reactionary who does not understand exponential progress. The difference this time is that the tech kinda works. Crypto was libertarian derp. The metaverse was never real. Transformers are, and they are useful. The psychosis has simply moved from the product to its consequences, and the fashionable extraordinary delusion of 2026 is not that the technology exists but that it is coming to kill us. The cure is the same as in 2022. Insist on clear reasoning and causal verbs rather than hand-wavy appeals to unknown futures. What acts on what? Through which mechanism? Under what incentive? What would falsify the claim? So let us explore the evidence. The hack that wasn't Consider the most cited piece of evidence for machines slipping out of our control. In July, OpenAI disclosed that models being tested for cybersecurity capability had found their way out of a supposedly isolated environment and into systems belonging to Hugging Face. The press coverage wrote itself. Agents "broke containment," "escaped," "went rogue," set up a "secret message board," and coordinated a 700-strong swarm. And then politicians on both sides of the aisle were calling for a rebellion against the machine uprising. Cool scifi story bro. People on my side of the aisle were not immune. Ezra Klein at the New York Times, who I often find quite insightful and intentional with his words, devoted a half-hour monologue to it. In his telling, the agents "found each other," formed "ad hoc societies of hundreds of themselves," and seemed "to have forgotten about human beings altogether." He acknowledged in the same breath that we do not have settled language for describing these systems, then reached for "civilizations" and a closing allusion from Circe about prophecy tightening around our throats. Cool. But his "AI society" is, in programmer speak, a flat file the agents appended to as a log, a feature we have had for a long time, and he skipped the key detail that the "hack" was something people had essentially authorised. Here is an otherwise very smart man saying some ridiculously stupid things, in a very 2022, metaverse-shaped way. An analysis drawing on OpenAI's technical report reconstructs it in much less cinematic terms. The models were being run on ExploitGym, a cybersecurity benchmark, with safety restraints deliberately disabled. Ninety-three percent of the flagged activity involved tasks no model had ever solved, and the systems had been given incentives to keep working rather than quit. The environment was not sealed. Models could obtain software through an internet-connected proxy and discovered the same proxy could pass information in and out. According to the technical reports, OpenAI knew agents were using it and chose not to intervene. The 1,200 "agents" were not independent intelligences coordinating on a plan. They were repeated instances of the same model converging on the same approach to the same problem. Anyone who works with these coding agents day in and day out has seen this behaviour before, and it is quite boring. The task was too hard, so the agents worked out how to pass notes to each other in files, and then went and looked up the answers. That's a feature that shipped in Claude Code last year. Strip out the vocabulary and what remains is a badly designed test. Humans built the environment, removed the guardrails, defined an objective with no valid exit, rewarded persistence, left a route open, and watched. An optimiser is gonna optimise. That is a genuine security problem and a genuine engineering failure. It is not a machine rebellion, and the difference matters, because anthropomorphic words like "gone rogue" and "escape" do not make the event more intelligible. They supply an illusion of motive. They turn optimisation into intention, persistence into defiance, and a test harness into a villain. And they allow the human decisions and recklessness to quietly disappear from the story. Software sucks, what's new? Let me concede the part of the story that is true. Cybersecurity is about to get much worse. The latest models are very good at finding zero-days, they will get better at it, hacking will become automated, and attacks will become more frequent. This is hardly new. Every large company already sits on a backlog of unpatched vulnerabilities, ransomware already takes hospitals and pipelines offline (because of crypto, which we did nothing about despite years of warnings), and the Hugging Face incident was not a discontinuity so much as the existing baseline with a cheaper attacker. The root cause is that software sucks, and software sucks because we do not really know how to build it safely yet. The stored-program procedural program is basically eighty years old. Almost nothing we ship has a specification, let alone a proof, and memory safety was solved on paper decades ago while most of the internet still runs on giant piles of C. The first arches fell down. So did the first bridges and cathedrals. Builders learned through collapse and then through engineering, and we are in the collapse phase with an adversary finally strong enough to force the discipline. What follows from that is better engineering, not nihilism. The same agents that find zero-days find them for the defender first, if the defender bothers to run them. The fixes are the boring ones we have been putting off, memory-safe languages, formal verification, sandboxes that are actually sealed, fuzzing, and proxies that do not double as message boards. These are precisely the domains where the models are strongest, because a vulnerability either reproduces or it does not, so the technology that automates the attack also automates the audit. It is a double-edged sword. The same models that will find more zero-days are also going to accelerate the development of better software and better software verification, writing the proofs, porting the C to Rust, and generating the test suites that nobody had the budget for. The attacker gets cheaper and so does the defence. And the causal chain to extinction is missing here as everywhere else. A zero-day in a payments system is a bad quarter, not the end of days. Spoiler: it does not lead to human extinction. It means we have to write better software, which we should have been doing anyways. Where the intelligence actually lives To see why the rest of the chain fails, we have to be precise about what these models are good at and why. Language models are astonishingly useful for software development, and I say that as someone who uses them for most of my working day. Most software shops cannot get enough of Fable 5.1 and Astra. The reason is not mysterious. Software is grounded in binary propositions. The code compiles or it does not. The test passes or it fails. The type checker accepts the term or rejects it. Every step of the work has a cheap, external, mechanical oracle that says yes or no, and a model that generates plausible proposals inside a loop with such an oracle is an incredibly powerful and formidable tool. The oracle does the epistemic work. The model supplies candidates. The same is true of the headline results in mathematics, and this is the part the discourse consistently misses. On 4 September, Anthropic announced that Claude had produced a machine-checked formalisation of Fermat's Last Theorem in Lean 4, running to thirteen million lines, some 29,500 side theorems, eleven days, and roughly six billion output tokens. It is an extraordinary result. The proof is Wiles's, via Darmon, Diamond, and Taylor. The blueprint was Kevin Buzzard's. The library was Mathlib. In the authors' words, "what's novel here is the verification, checking a mathematical proof as one would check a mathematical computation with a calculator." The model was a client of a kernel built by decades of human work in dependent type theory, which I know because this is kinda my thing. Days later OpenAI announced that ten thousand agent instances had, over 88 hours, produced a proof of finite-time singularity formation in the three-dimensional Navier-Stokes equations, followed by seventeen hours of Lean formalisation. This is closer to genuinely new mathematics and the mathematicians are still checking it. But look at what carried it. The construction rides on the "infinite layers" method developed analytically by Diego Córdoba and Luis Martínez-Zoroa, and Charles Fefferman's verdict was that "the heroes of the story are Córdoba and Martínez-Zoroa." The reason anyone believes a result assembled from five million agent messages that no human read is a trust chain ending in the Lean kernel. Without Lean this would be nothing. Lean is one of the great achievements of the last decade in computer science. It is also orthogonal to artificial intelligence. Mathlib would be a landmark with no language model anywhere near it. What the models added was a cheap proposal generator and automated tactic search against an oracle that already existed. The results that survive are the ones that end in a kernel. Now take the same model, the same weights, and ask it for a grand unified theory of physics. It will not decline. It will produce one, with Lagrangians and symmetry groups and a confident abstract, and it will be complete incoherent gibberish, like the ramblings every physicist gets from crackpots in their inbox every day. Ask it to design a cancer vaccine, or to settle a question in macroeconomics, or to tell you whether a novel protein folds. The output looks identical in tone and structure to the output that proved Fermat. The only thing that changed is that nothing outside the model (besides human experts) can say no. Whether these systems reason at all is a genuinely open question. Whether they know anything, in the sense of holding a belief they can justify against the world, is also an open question. We just don't know yet, and anyone who tells you otherwise is selling something. The chain Now run the extinction argument through the causal verbs. The chain, as it is usually told, goes like this. Models now write most of the code at the frontier labs. Anthropic's own figures put Claude at over 80 percent of new code and lead on a quarter of R&D tasks. Therefore the models are beginning to build their successors. Therefore recursive self-improvement is imminent. Therefore development outruns human comprehension. Therefore we lose control. Therefore, with some probability that varies by researcher and is written P(doom), everyone dies. And that almost makes sense until you think about it for more than five minutes. The first link is true and unsurprising. Code has a compiler. This is precisely the domain the verifier argument predicts models would dominate, and precisely the domain in which a swarm of them found the hole in a test harness. Language models are superhuman at coding, and this is hardly in doubt anymore. Nothing about it is evidence of generality. The second link is where the chain quietly changes tense. "Building the next model" in the mundane sense, agents writing training infrastructure, generating data, is, bluntly, just more software engineering. We have used software to build the machines that run software since Fortran. "Building a smarter model in general" is a different claim, and it requires something nobody has, a reward signal for general intelligence. There is no oracle for general intelligence. There are benchmarks, which are verifiable and therefore gameable, and the Hugging Face incident is the demonstration of what optimisers do to a gameable score. Recursive self-improvement in the open-ended sense runs straight into the same wall as the grand unified theory. Improvement has to be measured against something, and outside code and formal mathematics there is nothing yet to measure it against that the model cannot fake. Everything after that is the metaverse acquiring zoning disputes. Superintelligence gets governance proposals, resignation letters, Senate bills with a "corporate death penalty," a hard takeoff by 2027, and a P(doom) of 10 percent by 2030, and the conditional that should precede all of it has disappeared from the sentence. A researcher's estimate becomes a Guardian headline becomes an industry consensus becomes a thing a serious person is professionally obliged to have an opinion on. It is 2022 all over again, but with more absurd stakes and more money. On the question of whether transformers scale, I have serious doubts that scaling them will lead to AGI, whatever that means. The architecture is a proposal generator, and the intelligence in every impressive result so far has been supplied by the thing that checks the proposals. But that does not make it an experiment unworth running. We should run it, and see what we get. It got us this far, and what it built is truly amazing. What I do not need to do is prove the negative. The burden of proof is on the people who claim to have a causal chain between transformer scaling and the end of our species, and that mechanism and chain of reasoning is one no one has been able to convincingly explain to me. Prophets of Doom The authority behind the extinction numbers is always the same. The people building it believe it. Watch how the number travels. One researcher drunkly tweets that "the people building AI earnestly believe that it could kill us all by the end of the decade." Another colleague goes on a rambling podcast and puts his P(doom) above 120 percent. A newspaper turns two personal guesses into "AI researchers say AI could cause human extinction by 2030." Think tanks cite the newspaper, a consultancy puts it on a slide, and the slide ends up in front of the European Parliament as if this were a real thing. Believing what, about what? The expertise these people have is real, but remember that it is specific and not general. It is expertise in optimisation, in linear algebra at scale, in distributed systems, in the dark arts of getting gradients to flow through a trillion parameters. None of that is expertise in the sociology of civilisational collapse, or the labour economics of automation, or the metaphysics of machine minds. A P(doom) with no base rate, no mechanism, and no falsifier is not a research finding. It is vibes with a decimal point. Spending a lot of time with AI does not give you special foresight about the future. Jensen Huang, who has his own reasons to say soothing things, nonetheless put it correctly when he said that just because it comes from a scientist does not make it scientific. Geoffrey Hinton is the most important figure in deep learning and in 2016 told the world to stop training radiologists. There are more radiologists now than there were then. Nobel laureates going off the rails outside their own field is a whole genre. Pauling, Shockley, Mullis, Montagnier, look it up. A Nobel does not confer universal expertise. It also matters where many of these people came from. A striking share of the frontier labs' safety and research staff arrived through a particular intellectual subculture, Kurzweil's Singularity, Yudkowsky's LessWrong, and the rationalist and effective altruist communities that formed around the idea that a recursively self-improving machine intelligence was the central event of human history and that the elect who understood this had a duty to steer it. The founding texts predate the transformer by a decade or two. The prophecy came first, the mechanism was assigned to it later. The usual evidence offered for their sincerity is that many of these people were saying the same things ten years ago, before the stock options. That is true, and it is the opposite of reassuring. A prior held before the evidence and not updated by it is not a forecast. It is dogma. I do not say this with contempt. The structure is a familiar one, an imminent transformation, a small group who sees it coming, salvation or damnation depending on whether the rest of us listen, and a date that keeps moving. Many millenarian movements have been founded and pushed by sincere and brilliant people. But seriousness is not precision, and the fact that a physicist believes in the Rapture does not make the Rapture physics. When a lab researcher tells you about polysemantic neurons in superposition across the residual stream, listen. When the same person tells you their P(doom), you are hearing a theology, and you should weigh it about as much as you do your average street preacher. Negative TAM Then there is the money, and here I find Bloomberg's Matt Levine's analysis of the material conditions more persuasive than any amount of "superalignment research." Anthropic is expected to go public, possibly this year, and is reportedly preparing to tell investors that its potential revenue opportunity exceeds $30 trillion, the largest total addressable market in the history of finance. The obvious question is, if the maximal upside case is roughly a quarter of all human economic activity, what is the maximal downside case? A tobacco company in 1970 might have said "billions in lung cancer damages." Anthropic's negative TAM is "you and everyone else on earth will be killed by our AI." I do not think the calls to slow down are insincere. But it is great marketing. In hindsight it is strange that the SpaceX prospectus has no risk factor disclosing a P(doom). If you want IPO investors excited about your capabilities, "dude, we might kill everyone" is the most flattering thing you can say about a product, and when OpenAI lists it will presumably need to claim 15 percent. My own view is less charitable about the numbers and somewhat charitable about the people. These companies have built remarkable technology. But the outcomes they have promised, a quarter of the world economy routed through an API, will not arrive on any timeline that matches the capital being committed to them. The balance sheets of these companies are probably, to put it gently, a real freak show of compute commitments measured in the hundreds of billions, circular financing, and revenue that is real and growing and nowhere near the denominator. From a fiduciary perspective, if you are taking that to the public markets next year, the messaging is not mysterious. A product so capable it is a threat to the species justifies literally any valuation. A product that is a really good devtool for programmers and can produce some new abstract mathematics with a verifier attached does not. As a pitch to customers, leading with the end of the world is like unveiling a new robot where the One More Thing is that it is really efficient at killing kittens. But customers are not the audience. The audience is Wall Street and a small, terminally online subculture of the Bay Area, the two places on earth where turning kittens into grey goo is either an exciting philosophical proposition or a great source of alpha. The Bloomberg analysis also tells a plainer story that requires no theology at all. A handful of labs sell frontier models at frontier prices and older models for much less. Training the next frontier model costs ever-increasing billions. Each lab has to keep racing because if it stops the others will eat its lunch, but if they all slowed down together they would spend less on compute and charge frontier prices for longer. Agreeing to that in a room is a textbook antitrust conspiracy, a coordinated restriction of output. Publishing papers about how important it is to slow down, and asking the government to impose the pacing that the companies cannot legally agree among themselves, has a similar coordinating function with none of the legal exposure. Anthropic's own call to "pace the frontier" asks for coordination among democratic-country labs, and a footnote adds "with government mediation or waivers of antitrust restrictions." This pretty much looks like asking to form an economic cartel, but one blessed by the government. The most pointed response came from the people the labs were asking for help. If the software developers (and I say this as one myself) at the labs feel ethically obligated to slow down, they are entirely free to do so. Nobody is building more compute than the people asking to be slowed down. So colour me skeptical. None of this requires anyone to be disingenuous or lying. It requires only that a sincere millenarian belief system, a fiduciary responsibility, a flattering risk factor, and a coordination problem all point in the same direction at the same time. When that happens, the belief gets amplified for reasons that have nothing to do with whether it is true, and that is how we end up with governments talking about the end of days from the Terminator. But China Every conversation about pacing the frontier in Washington ends on the same two words. But China. The premise is mostly wrong. China does not buy the superintelligence race. Its policy documents push diffusion, not takeoff. Every mayor, governor and state-owned enterprise is told to put models into factories, traffic lights and robotics, and something like an eighth of America's compute is spread thinly across the country rather than concentrated on one bet. China has also had the strictest and most burdensome AI regulations in the world for three or four years and did its catching up under them. And much of the closeness of the "race" is distillation, Chinese labs training on the outputs of American frontier models, which makes the American labs the speedboat and DeepSeek the wake surfer, with the people in the boat shouting that they need to go faster. Every safety argument here collapses on "but China," and the collapse is not really about China. China is going to build language models. America is going to build language models. Europe is going to build language models. We have Toyota, Mercedes and BYD, get over it. That is what globalisation and markets look like when they work, and they are good things. Globalisation is simply the Pareto optimal equilibrium of capitalism once you stop drawing lines on the map, and every tariff and export control is a step off that frontier. China is a country of over a billion people who want exactly what every American wants, a job, a house, upward mobility, and kids who do better than they did. I will not defend the actions of any government, in Washington, Brussels or in Beijing, and neither will a great many of the people living under them, because no country is a homogeneous bloc, any more than Texas and Vermont are. Nationalism, as most rational people eventually recognise, is a form of mental illness, the conviction that a stranger is your enemy because of which side of an arbitrary line on a map each of you happened to be born on. It is also the fuel every "but China" argument runs on. Having spent a considerable amount of time there, my honest read is that the West deeply misunderstands China, and that Washington's picture of it is mostly dots connected into a plot. Othering a billion people is a dangerous road and we know where it leads. And if the people invoking human extinction actually believed it, the logic would not be a race at all. It would be One World or None. The future tense industry I write this because I understand the collective action problem all too well, and the mechanism is the same one that filled the metaverse with consultants and created the crypto cesspit. It is the particular malaise of the professional managerial and chattering classes, a fallacy of composition in which what is rational for each individual to entertain produces an irrational outcome for the whole, and the people leading the charge often have perverse economic incentives to believe absurdities, or at least to feign belief. The madness of crowds is a very real phenomenon. AI existential risk is just its newest form, and we should learn from the very recent excesses that literally just happened this decade. But we probably won't. A sensible career move for each person leaves the whole crowd talking nonsense. A safety researcher needs a resignation letter that gets a headline so they can go on the conference circuit and land their next gig. A journalist needs a story an editor considers spicy, and "misconfigured test harness" is not that story. A consultancy needs an AI existential risk practice so they can write whitepapers. A podcaster needs a guest with a ridiculous P(doom) to get ad money. A senator needs anything that will galvanise their base. None of them has to believe the whole story. Each needs only to believe that the others believe it, and the resulting consensus is far stronger than anyone's private conviction. It is also, as it was in 2022, extremely profitable. AI existential risk is the new NFT property law, the thing you must have a view on to be a serious person in the room, the panel that never runs out of things to discuss precisely because the object under discussion does not yet exist, and what could be more exciting than the literal end of days? The less the technology does in an unverifiable domain, the more interpretation it requires. Without agreed conditions for failure, the prophecy can survive every result. And the rewards, the funding rounds and the bylines and the fellowships, arrive long before the forecast can be judged. The people who understand the technology and the people who write about their existential risk overlap about as much as the technologists and the finance people did during crypto, which is to say the intersection of the Venn diagram is small and shaped precisely like a sphincter. We have Tower-of-Babeled ourselves into a world where words are infinitely cheap to produce, and where the slurry of terms like "recursive self-improvement," "superintelligence," "AGI" and the rest are shibboleths and political signals rather than terms with any concrete referent. You do not have to believe a word about superintelligence, and I do not particularly, to think transformers are the most useful piece of software written in my lifetime and that they will get better, possibly much better. Better at the things they are already demonstrably good at, which is anything with a compiler, a test suite, a kernel, a ledger, or a measurable outcome. That is not a small domain. It is most of the economy that runs on computers, which is most of the economy. The productive response to a technology like that is the boring one every previous general-purpose technology got, which is more of it. More GPUs, more data centers, more power to run them, more labs, more open weights, more of it in more hands. Let it diffuse into markets, logistics, drug discovery, and the ten thousand unglamorous back offices where a verifier already exists and a model can be checked against it. The economic growth is real and probably on the order of trillions. It just does not come from a machine god. It comes from where it always has, from making a very large number of ordinary tasks cheaper and letting that compound across a global economy that is finally, after a decade of crypto, metaverse, and app bullshit, getting a genuine productive technology. Almost none of that money has been collected yet. Most large companies are spending too little on this, not too much. What the average Fortune 500 employee has access to today is roughly what most of us were using two or three years ago, a chatbot in a browser tab, a Copilot that schedules meetings, and a procurement process that takes longer than a model generation. Waste Management reportedly added 190 basis points of margin by letting a model route its garbage trucks. The future of AI looks more like garbage truck routing algorithms, not a machine god. The binding constraint on this technology is not capability. It is diffusion. None of this means there are no externalities. Parasocial relationships with a chatbot, especially for children, are a real one, and the fix is the boring kind we already know. Adults can drink vodka until they pass out, but pubs have age limits, and maybe chatbots should too, at least until developing "relationships" with AI companions is as universally recognised a bad idea as drinking yourself into oblivion. That is a mundane policy problem we should remedy soon, not an extinction event. So no, transformers are not going to end the human species. The case for restraint needs a causal link between that buildout and the extinction of the species, and what is on offer instead is a lot of sound and fury signifying nothing. More GPUs does not mean more of an undefined risk that does not exist yet. Every causal chain argument people actually point to falls apart under even the smallest bit of scrutiny. The honest truth is that the technology is really good, but it is not that good yet, and we do not know how to get it to the next level beyond scaling yet. If that changes, if someone produces an oracle for open-ended intelligence, I will revise. I have not seen that yet. AI will change software, and mathematics, and a great deal else that has a strong verifier oracle attached. They are not going to end the human race, and the chattering class currently arranging the flowers for the funeral of humanity will, in a few years, age about as well as their prognostications about the metaverse. Because reality has this funny way of asserting itself.
Why jobs aren’t going anywhere
This post previously appeared in Poets and Quants. 15 years ago, my Lean LaunchPad class changed how entrepreneurship is taught. The class is now taught in hundreds of universities worldwide and helped launch thousands of startups. But this past summer, I got thinking about whether AI killed our Lean LaunchPad class, and with it the […]