Full Width [alt+shift+f] Shortcuts [alt+shift+k]
Sign Up [alt+shift+s] Log In [alt+shift+l]
1

We are going to kill “unalive”

from Anil Dash [alt+shift+b] in AI

I talk to a lot of old people, those who were born in the 20th century, and if I ask them what the word “unalive” means, they usually have no idea what I’m talking about, except for some of them who have kids or who study Internet culture. This will, of course, probably seem very weird to most people who were born in or grew up in the 21st century. Just to recap for the olds: “unalive” is the word you use to represent concepts like dying, or death, or killing or being killed, on digital platforms where saying those words accurately will cause the algorithm to punish or censor you. Or, maybe, where the perception is that using those words will result in being censored by the algorithm, and no one is actually willing to find out what happens if you use the forbidden words. This sort of attack on people’s expression started on platforms like TikTok, where nearly all content is distributed through an algorithmic feed, but has since become ubiquitous in nearly all digital media that we see. In fact, these tics are now so prevalent that it’s routine to hear people using this kind of language in everyday life, even though there’s not yet an algorithm to appease in the physical world. I’ve heard people say, out loud, “he unalived himself”, in reference to someone dying by suicide. And all of this has become even more visible in recent days as online conversation has turned to discussion of the horrific lack of accountability around the tragic rape case at Cornell University. Across the Internet, people are routinely referring to the central crime in the case as r*pe or “grape” or even using the 🍇 emoji, without a second thought for what it means that the very word can’t be said online anymore. Or, at least, the assumption is that it can’t be said. To be clear, I am very much in favor of people using content warnings or sensitivity markers for content, and fine with people using abbreviations like “SA” for references to disturbing or triggering topics like sexual assault;...
14 hours ago

Stay updated

Get a weekly newsletter with the top 5 articles worth reading every week.

More from Anil Dash

VC isn’t VC anymore — understanding the rise of Cancer Capital

We really, really need to talk about venture capital. Because it’s not “venture capital” anymore. There’s a huge disconnect between what most people think of VC, where an investor has a big fund and cuts checks to help a founder build a company, and the current reality, where a handful of billionaire extremists use the cover of “VC” to advance an outrageous agenda where they’re accountable to no one. I’m gonna explain this from a standpoint that almost never gets articulated: I’ve personally raised tens of millions of dollars in venture capital funding as CEO of startups, and been directly involved as a board member or advisor in raising hundreds of millions more. I’ve sat in board rooms, across the table from the people I’m talking about here, or been at the industry events that they frequent. So this isn’t sour grapes because these VCs wouldn’t cut me a check, or some chip on my shoulder about these investors due to a business deal. This is what I know about these bad actors because I’m part of the community of creators and inventors who build the things that they used to invest in — back when they still cared about innovation. Many of the trends in society and politics that people are most angry about, from data centers being forced down everyone’s throats, to all of our favorite apps and services being enshittified, to politicians being paid to ignore the will of the people, are all being supercharged by these cancer capitalists. They have warped the structure of venture capital into a form of oligarchy that answers to no market, no regulators, and no voters. So it’s worth understanding exactly how they did it. How Venture Capital became Cancer Capital I’ll be breaking these points down in further detail, but just to begin framing the concept, I’ll lay out the core idea here in some bullet points (so that you’re not tempted to run this whole thing through an LLM): Venture capital was only supposed to be a tiny segment of the overall capital market, but it has expanded to become the primary form of funding that new companies consider — it was never the only way, and it didn’t used to be the default one VC was meant to be a small percentage of overall investment because it represents the high-risk, high-reward part of a portfolio; to be healthy, most of a portfolio — or most of an economy — needs to focus on assets that are more stable and predictable. But a cancer grows from a cell that a body needs in small, healthy amounts, and that turns deadly when it grows without limit until it harms, or even kills, its host. As regulations have gotten looser in recent years, a handful of venture capital firms have become “do everything” funds that combine private equity with their existing VC businesses, and expand to manage massive stockpiles of tens of billions of dollars The 1% of VC firms that get this big stop being exposed to the risk in their own investments at all — when you collect 2% a year to manage $50 billion, that’s a billion dollars landing in your pocket annually whether any company you funded lives or dies. Those firms also stop legally even being venture capital firms, making them unaccountable to markets, founders, or the law — and that’s how they become “Cancer Capital” Meanwhile, the 99% of “normal” VCs don’t have the power or funding of the Cancer Capital firms, but are forced to play on the field that those firms define, even if they don’t like the way they do business Since the Cancer Capital firms have become so powerful, the overall balance of power between founders and VCs has flipped; instead of founders having a company that VCs would try to fund, now VCs publish extremist political manifestos, and “founders” are just the people who are selected to carry out parts of those plans The rest of the world doesn’t know: New founders and workers entering the tech industry are unaware that Cancer Capital has taken over, so many are still trying to play by the old rules, and can’t figure out why their ideas are being pushed into serving the goals of the Cancer Capital firms Politicians and media still look at VC as if it works like it did 10 or 20 years ago, and cheer them on like they’re funding job creation or enabling new companies to grow, when their primary goal is concentrating power and wealth into the hands of the Cancer Capital tycoons. They keep getting fooled by this, over and over. These days, venture firms are increasingly getting their funds from pension funds and retail retirement accounts, meaning the public (you!) are increasingly holding the bag for the parts of their portfolios that actually have some risk, even if you never intentionally made that choice The shift away from IPOs in the tech industry has also encouraged these Cancer Capital firms to find ways to cash out long before companies ever go public, meaning they can make a massive return off of companies that never make a penny of profit, even if regular investors get screwed by the stock of a company once it actually gets listed on the public stock market. Part of why this has gotten so corrupt is the way the Cancer Capital firms have transformed themselves into their post-VC forms. Because they’re not legally VC firms anymore, they’re free to buy shares directly from founders, or hold unlimited amounts of publicly-traded stock — exactly what they couldn’t do as regular VCs. They can even sell their investment in a company as an asset to another one of their own funds, and then book the increase in value as a profit, all without the company ever having made a penny. Another racket: a company that’s raised a bunch of cash in a funding round can buy out its early investors if they’re one of these post-VCs, so they can get paid off even if their portfolio company has never made a penny in profits or revenues. All of this self-dealing, and the way that they’re isolated from any accountability, has made these firms become more and more shameless in their behavior. Former Andreessen Horowitz partner John O’Farrell publicly called out the firm (a rarity — the company is notoriously vindictive towards those who it decides are disloyal) for what he called its “political infiltration” of AI policy. Marc Andreessen, Ben Horowitz and their firm have put $115.3 million into this midterm cycle — nearly double their $63 million in 2024, and more than any other billionaire donor in the country, even including Elon Musk. Molly White, whose Tech Influence Watch tracks this money against FEC filings, shows that a16z alone accounts for more than 20% of all political contributions from the entire cohort of crypto and AI companies it follows. And they’re funneling these funds to candidates in both parties. This is a huge escalation from the tentative baby steps that folks like Zuckerberg were making in the Obama era, working on benign issues like trying to help immigrants. And of course, it gets a lot worse than just their lobbying. As I have frequently noted, Andreessen Horowitz hired a man as a partner at their firm despite his having no background or qualifications in tech, finance, or startups whatsoever. His only discernible qualification was that he had choked my unarmed neighbor Jordan Neely to death on a subway car. This is how brazen, how toxic and destructive, we’ve allowed the industry formerly known as venture capital to become. We must understand that it is no longer a financial machine that is used to fund startups, but a political and social machine focused on dismantling democracy and civil society. And it’s time to act accordingly. Up next: we’ll dive into the specifics of many of the points laid out above, to understand more about how we got here.

2nd Sep 2026 • 2 votes
Introducing Dashboard Touch, a build-your-own version of Touch ID

For years, I’ve wanted to have a standalone version of Apple’s Touch ID authentication feature for my Mac, but without having to use an Apple keyboard. (I generally like their keyboards, but my daily driver keyboard these days is a big clicky mechanical beast.) I’d gone down various dead ends of trying to find substitutes, and even checked out the efforts where people had ripped apart expensive Apple keyboards just to scavenge the Touch ID sensors out of them. None of them quite solved the problem. So today, I’m sharing an open source project called Dashboard Touch, which lets you make your own Touch ID-style sensor for your Mac, using low-cost off-the-shelf part. It’s based on an extensive refactoring of the excellent tinyTouch project by Zimeng Xiong, who recently cracked the code on how to make a useful fingerprint scanner system that’s also reasonably secure for regular Mac users. (You should definitely check out his project and support his new hardware build if you’re interested in this stuff.) I took my own approach to this work because I wanted to focus a lot on having a friendly web interface for configuring exactly how the fingerprint sensor system works on your computer. When you get Dashboard Touch set up, it presents you with a nice web interface that runs right on your own Mac, letting you do things like set the color of the ring light on the fingerprint sensor, or capture your fingerprints so they’re recorded in the system. Behind the scenes, the way the system works couldn’t be simpler. You buy a little fingerprint sensor, and a small microcontroller, wire them together (it was actually fun to get back to soldering stuff!), and then plug them into your computer with a regular USB cable. After you run the setup script, you just go to the web interface and add your finger(s) to the system. Once you’re running, your Mac runs as normal, except any time the system prompts you to type in your password, you can just swipe your fingertip on the sensor and Dashboard Touch will type your password in for you. At a technical level, the device is actually literally pretending to be a keyboard. (You can look over the code on GitHub and get a feel for the approach pretty quickly.) After it’s installed, it’s basically a set-it-and-forget-it kind of thing. You don’t need to do anything else for it to Just Work. Your password is only ever stored securely on your Mac, and your fingerprints are only ever stored securely on your sensor. You can erase them at any time and nothing talks to the internet at all except the one manual update checker, where you can see if there’s a new version of Dashboard Touch, but only when you intentionally click the button to request it to do so. Overall, this isn’t the kind of system you should use if you’re protecting a bank vault, but if your computer is physically secure and nobody is going to have extended unsupervised access to your Dashboard Touch setup without your permission, you should be fine. Dash, not bored Personally, it’s been really fun to get back to making things. As you can see in my introductory video, I ended up creating an enclosure for my fingerprint sensor in my woodshop, so that it would match my desk that I recently built. Even creating and editing the intro video was a fun project that took me out of my usual comfort zone. Nearly every task in this project has had me stretching to do things that I’m pretty bad at, from firmware coding to security review to detailed carpentry to video editing. But I just love the idea of putting things out there again for people to hack on, and there's also something really satisfying about being able to be super-opinionated about the design and user interface of something after so many years of working on teams where I had to collaborate. (Even though I always got to collaborate with brilliant people, it's different when you get to pick every pixel!) I just also have been missing the era of the web when most of what I saw online was weird and fun things that regular people were building, and I realized that I can't mourn the absence of those kinds of projects unless I invest my own energy into building some of those kinds of things myself. So, here's one! Let me know what you think, and if you've got any ideas for how to make this thing better. Or, of course, if you find any bugs that I should fix. I hope you have fun touching the blinking lights!

20th Aug 2026 • 1 votes
Becoming Skilled at Making Documents

The vast majority of the documents people use to do business are really quite poor. Presentations that make your eyes glaze over, memos that are inscrutable or unclear, and all kinds of artifacts that say more about how they were created than whatever message they were ostensibly trying to communicate. It's been one of my great frustrations for years, and a big part of why I wrote Make Better Documents a while ago. That post captured a list of the suggestions I've been giving people for years on how to make better, more effective documents that can actually do work for you, instead of fighting at cross purposes to your larger goals. To my great surprise, that list of suggestions on how to make better documents got a pretty huge response, and a lot of people told me they found it really helpful. So now, I've created a Better Documents skills.md file for people who use LLM tools like Claude to help assist them in creating business documents, to prompt their AI tools to make better documents by default. If you're not familiar, agent skills are simple text files that describe new capabilities or processes that LLMs can take advantage of when carrying out tasks. (They're Markdown files — more proof of how Markdown is taking over the world!) The way this skill works is that it's distilled the broad principles I outlined in that post into a series of 5 tests, covering areas like whether you've properly considered your target audience, whether the overall structure is correct, if you've overdone things with your formatting, and if things are named clearly, and then either generates a new file that follows those rules, or reviews an existing document to make sure it is obeying best practices. It's nothing too fancy, but I've been using it for a while, and shared it with a few friends, and people have told me they found it handy and it's improved some of their routine documents. I'm especially glad that people have found it useful even if they're the kind of folks who would never let an LLM generate a document on their behalf, but do think software tools are useful for things like spell check or grammar check. I see this as being a tool in that kind of category. If you're familiar with skills, the install process is really simple and works just like any other skill. This skill is totally free and open source (if you have improvements, send along a pull request on GitHub, or if you're not a coder type, just email me or hit me up on social media and let me know what fixes/suggestions you've got), so there are no encumbrances or restrictions on its use. I am curious if it's useful to people, especially if you find it handy to use more broadly at a company, so don't be shy to get in touch if you find it valuable. Here's to us all enduring fewer terrible presentations!

24th Jul 2026 • 2 votes
Maybe it's time for lots of little indie AIs to take over

“[T]here can be alternatives. What we can imagine is, rather than the ChatGPT killer, a lot of different little AIs from little responsible players.” That’s me, in The Guardian a few days ago, trying to distill a message that I’ve been trying to get out as broadly as possible for quite a while now. It's sort of like hoping a comet will take out the major AI players and a bunch of smaller new players will be the smarter, better-adapted mammals that take their place instead. We’re in another one of those big inflection points for AI. Trump administration policymakers for AI suspended access to Anthropic’s newest product. All of these policymakers have a web of investments in competing players — including SpaceX, which is about to IPO — and the corruption and grift of this cohort are so extensive that it’s impossible to judge what the actual risks and reality are around any of these platforms or technologies, since no one involved is an honest broker. It’s a shame More broadly, there’s been the widespread pushback against AI culturally, one that is undeniably strongest amongst those who were born in this century. But the adoption patterns and usage data show that even younger people are using some AI tools. And that’s a pattern that we’ve seen before, with social media. We have a significant group of people knowing that a technology contradicts some of their values, preferences, or beliefs, but using it anyway. Sometimes it’s due to the coercive nature of the platforms themselves, and how they insinuate themselves into our lives, to the point where we don’t even realize we’re using them. Sometimes they are forced upon us by the creators of the platforms, since they have so much power over the devices we use, and the tools that we rely on for things like doing our jobs, or communicating with our loved ones or our communities. There are millions of people who don’t like that they’re using LLMs provided by the Big AI companies, but end up using them anyway. Just like there are hundreds of millions of people who don’t like that they’re on the giant social networking platforms like Facebook, but end up using them anyway. The feelings that people walk away from those experiences with are often guilt, or shame, or embarrassment, or resentment — all some of the most negative and destructive emotions that humans can experience. Actual alternatives But if people want to get the benefits of some of these technologies, without either the shame of supporting the harms of Big AI, or the unpredictability of being beholden to corrupt billionaires bickering with one another, there are finally starting to be other options. As I mentioned in (One) Good AI Is Here, it’s possible for creators working in their own communities to now make AI tools that serve their specific needs, without causing all the harms that make people object to Big AI. This feels like the true alternative to the narrative of “inevitability” that so much of the hyper-funded AI industry is trying to push, while also not forcing people into a quiet life of AI guilt if they still find some utility in some aspects of these tools. Right now, those who (rightly!) object to Big AI due to their platforms’ impact on the environment, or labor, or their extractive use of content without consent, or its many other potential harms, are generally not aware of, or often open to, the idea of there being small, human-scale tools created by and for communities that are accountable for those tools over time. But my suspicion is that it is not only possible to make these tools, there may in fact already be many of these tools in existence, and we’re just not as familiar with them because they’ve been quietly serving their specific niches without having multi-billion-dollar campaigns promoting them. What I'm unabashedly hoping to do (and I think the Guardian story reflects some momentum in that regard), is shift the narrative from focusing on running away from the bad thing in AI, to finding the good thing that we're running toward. There are alternatives that we could be affirmatively choosing, ones that look at questions like the one I asked more than a year ago, "[https://www.anildash.com/2025/05/01/what-would-good-ai-look-like/](What Would "Good" AI Look Like?")", and offer answers that might give us hope instead of just the righteous rage and anger we feel when we let our imaginations be constrained by the limits of what Big AI offers.

15th Jun 2026 • 1 votes

More in AI

What is Codemode

More than a year ago I wrote a few posts here that recommended people not to load custom tools into their context (or MCP servers) but to just use more scripts. Most importantly I wrote that Code Is All You Need and I wrote about that MCP needs code. With Pi 1.0 we now added MCP support via Codemode which in some ways is a long time coming, but then also maybe somewhat surprising to some. So I want to share some updated thoughts on this blog on what this all means. What Are Tools When a harness like Pi provides tools for an LLM to call, it does so by supplying some tool definitions which then translate into some token structure on the server side. Whether a model is encouraged to call a tool is the result of the reinforcement learning process. Something I wrote about before if you want to learn more. One of the reasons we strongly lean towards CLI and bash is because it allows easy composition of calls, and because the model also learns how the file system works when it’s trained. So when it invokes a tool like echo foo > /tmp/test.txt the model also learns that after that tool call, there is now a file called test.txt in /tmp. However bash has one fundamental limitation which is that it can only compose programs that run. And there are some things, which are not programs, but native tools to the LLM and they sort of have to be. The most obvious example here is read or view_image. If a multimodal model needs to read an image, it cannot use cat for that because the harness needs to inject the actual image payload into the protocol of the LLM. Another quite vivid example are sub agents. In order to spawn and orchestrate sub agents, it’s tricky to avoid tools that are provided by the harness. While in theory the agent could provide a CLI tool that talks to the outer harness via environment variables and Unix sockets, it’s a rather crude process. It however has another issue, and that is where the code runs. Brains vs Hands To better understand that, it’s important to think a bit more about where all the bits and pieces run. There really usually are two different systems involved. The first is the brain, the harness: it runs on one machine. It’s trusted. The second is often the same machine, but it’s really where the tools are executing: the hands. In Pi we now call this the execution environment, but you can think of it as the target of all the operations. Crucially what is important for us, is that there is a dividing line between the harness brain and the target environment that runs bash and executes the tools. And splitting this in half has some really important consequences. For a start it means that they are running on different file systems and they have different levels of trust. If you for instance use a sandboxing solution like Gondolin your bash stuff will be sandboxed just fine, but the harness itself will not be. Orchestrating The Harness Which brings us to what Codemode really does: it’s a way for the LLM to express and orchestrate complex operations on the harness side, but not the execution environment side. Codemode runs in the harness, in its own sandbox. In case of Pi it’s running in QuickJS within a WASM runtime with intentional limitations: no network, no file system, no timers, limited RAM. The only way is to call more tools. You could also imagine that Codemode could run Scheme or some other language as well. If you are not familiar with Codemode, it’s basically just a way to issue tool calls from within some language, in our case JavaScript. That allows you to compose those calls without necessarily going through the LLM’s context. Credit for naming goes to our friends at Cloudflare who coined it. For instance if you issue a bash call as a regular tool call in the LLM, then we only throw the trailing 2000 lines into the context and if the agent wants more, it needs to look at the overflow file itself. If however the agent issues that invocation via Codemode, then the Codemode side gets larger outputs sent structurally. Most importantly, because Codemode is JavaScript the agent can express concurrent operations and basic workflows. A common way in which you see agents now use this, is to first probe at 5-10 items from some tool response to see what it looks like, and to then write a Codemode script that processes the next n items. Codemode also allows you to throw state into the transcript! That means that one Codemode invocation can stash away data, that the next call in the session can load again. And remember: this is on the harness host, not the sandbox. In case of Pi, Codemode also allows you to issue calls that naturally do not make any sense in Pi’s traditional interface. For instance if you want to generate images with an image model or you want to classify some text with a one shot classifier model, those Pi APIs are exposed via Codemode, but not via regular tools where they would just waste context. What It Looks Like So now that we talked a bunch about it, it’s probably worth being a bit more explicit about it. Let’s walk ourselves through some invocations of Codemode of recent Pi sessions of mine. Note that none of this code is human written. It’s from real sessions of Pi, just re-indented for your viewing pleasure. The agent starts using Codemode automatically either because it’s a task where the model already naturally picks up that tool, or because a user asked it to. Note that Codemode is by default only enabled in Pi when MCP is enabled, but you can turn it on with "defaultTools": ["+codemode"] in the settings. Just ask Pi to enable it for you. Generating Images Let’s start simple with image generation. Image generation is a feature that Pi supports in the AI SDK core, but it’s not a tool that the agent can use. In the past the only way to use image models has been to write a bespoke extension or to have the agent run node itself and use the internal image APIs. However because we expose quite a few of the internal model APIs within Codemode, it means that the agent can use it: const [painter] = await models.getAvailableOfType("image"); const result = await models.generateImages(painter, { input: [{ type: "text", text: "A cute little puppy sitting on a grassy " + "lawn, soft natural light, photorealistic" }], }); if (result.stopReason !== "stop") return result.errorMessage; for (const block of result.output) { if (block.type === "image") image(block); else text(block.text); } Note that the call to image() sends the image back as image content to the LLM. On the harness side it feeds it directly into both the agent, as well as onto disk as a temporary artifact in case the agent wants to be able to pass that image back to bash. Classifying Things Similar things apply to classifier models such as Jev. They also do not fit well into the workflows of an agent through the typical tools. But rather than making a bespoke tool available, Codemode just allows the agent to reach into the AI SDK and invoke those directly. Here you can see how Jev is used to mass process GitHub issues for a quick sentiment analysis: const jev = await models.getModelOfType("classifier", "typesafe", "jev-latest"); const r = await tools.bash({ command: "gh issue list --state open --limit 100 " + "--json number,title,body,comments", }); const issues = JSON.parse(r.output); const results = await Promise.all(issues.map(async (issue) => { const res = await models.classify(jev, { state: { title: issue.title, body: (issue.body || "").slice(0, 4000), comments: issue.comments.slice(-5).map(c => c.body.slice(0, 800)), }, questions: { sentiment: { type: "choice", instructions: "What is the overall sentiment of the author towards pi?", criteria: { positive: "Appreciative, happy, constructive praise", neutral: "Matter-of-fact report or request without emotion", negative: "Frustrated, annoyed, upset, or angry", }, }, frustration: { type: "score", instructions: "How frustrated is the reporter?", criteria: ["not at all", "mildly", "clearly frustrated", "very angry"], }, kind: { type: "choice", instructions: "What kind of issue is this?", criteria: { bug: "Bug report or regression", feature: "Feature request or enhancement", question: "Question or support request", other: "Docs, discussion, meta, spam", }, }, }, }); if (res.stopReason !== "stop") { return { n: issue.number, title: issue.title, error: res.errorMessage }; } return { n: issue.number, title: issue.title, ...res.answers }; })); store("sentiment_results", results); return results .filter(r => !r.error) .sort((a, b) => b.frustration.score - a.frustration.score) .slice(0, 12) .map(r => `#${r.n} ${r.frustration.score.toFixed(2)} [${r.kind.choice}] ${r.title}`); Note how in that above example we also call store() which dumps the result of that execution into the session transcript. A future invocation of Codemode can thus read back that result if it wants to. The Promise.all here is fine, because Pi limits the total number of concurrent tool executions itself to four and maintains a queue for the rest. A more adventurous example is to use Jev to drive a game engine for debugging purposes: Codemode with Jev for Game Debugging Here it knows about my tankctl command and it built itself quickly a minimal harness around it to drive a game loop to assist a user with debugging a problem. Note how it built a 30 step loop in which each step goes back to both the game engine to get a text dump of what’s going on, and then to Jev to determine what to do next: const jev = await models.getModelOfType("classifier", "typesafe", "jev-latest"); const tank = async (cmd) => (await tools.bash({ command: `tools/tankctl "${cmd}"` })).output; await tank("start --map assets/maps/night_arena.map"); const questions = { action: { type: "choice", instructions: "You control the tank '@' in a top-down tank game. " + "Choose the best next action.", criteria: { attack: "an enemy has line of sight to you and you can fire at it", approach: "no enemy has line of sight; drive toward the nearest enemy", dodge: "an enemy shot is heading at you and will hit soon", powerup: "a powerup is close and no enemy threatens you", }, }, }; function commandFor(choice, st) { const p = st.player; const enemy = st.enemies.filter(e => !e.dead) .sort((a, b) => (b.los - a.los) || (a.dist - b.dist))[0]; if (choice === "attack" && enemy) { return `fire_at tank ${enemy.id}; frames 30 until clear,damage,kill`; } if (choice === "dodge") { // move perpendicular to the closest incoming shot const s = st.projectiles.filter(s => !s.yours) .sort((a, b) => a.eta - b.eta)[0]; const dir = s && Math.abs(s.vel[0]) > Math.abs(s.vel[1]) ? (p.pos[1] > s.pos[1] ? "+down" : "+up") : (p.pos[0] > (s ? s.pos[0] : 0) ? "+right" : "+left"); return `input ${dir}; frames 20 until damage; input stop`; } const powerup = st.powerups.filter(u => u.available) .sort((a, b) => a.dist - b.dist)[0]; if (choice === "powerup" && powerup) { return `goto ${powerup.pos[0]} ${powerup.pos[1]} 180`; } return enemy ? `goto ${enemy.pos[0]} ${enemy.pos[1]} 90` : null; } const log = []; for (let step = 0; step < 30; step++) { const st = JSON.parse(await tank("state")); if (st.state !== "playing") break; const threats = st.projectiles .filter(s => !s.yours && s.miss_dist < 1.5 && s.eta < 1.5) .map(s => `incoming shot dist ${s.dist} eta ${s.eta}s`) .join("\n") || "no incoming shots"; const r = await models.classify(jev, { state: { map: await tank("view 8"), threats, hp: st.player.hp }, questions, }); if (r.stopReason !== "stop") { log.push(`#${step} classifier error: ${r.errorMessage}`); break; } const choice = r.answers.action.choice; const cmd = commandFor(choice, st); if (!cmd) break; log.push(`#${step} hp=${st.player.hp} ${choice} -> ${await tank(cmd)}`); } return log.join("\n"); Calling MCP Servers Lastly, Codemode obviously is great for calling MCP servers. And because we do not actually expose any of the MCP tools to the LLM, the agent first uses provided APIs to issue a tool search within Codemode to discover what it might be able to do with the connected servers. This form of progressive discovery makes the whole MCP business work well enough for a lot of use cases today. Here for instance you can see the agent reach for the Sentry MCP straight away, even without discovering the tools, presumably because it has learned during the RL process already about what the Sentry MCP looks like. But it learns from what we inject into the system prompt, that the Sentry server is available to begin with. It’s not completely guessing here. const orgs = await tools.mcp__sentry__find_organizations({}); const { organizations } = orgs.structuredContent; const results = await Promise.allSettled(organizations.map(org => tools.mcp__sentry__find_projects({ organizationSlug: org.slug, regionUrl: org.regionUrl, }) )); return organizations.map((org, i) => { const r = results[i]; if (r.status !== "fulfilled") return { org: org.slug, error: String(r.reason) }; if (r.value.isError) return { org: org.slug, error: r.value.content }; return { org: org.slug, projects: r.value.structuredContent.projects.map(p => p.slug), }; }); Modern MCP Is A Fight I really don’t want to talk too much about MCP here, but MCP is in fact a protocol that greatly benefits from Codemode. The problem in parts is that MCP in practice often targets harnesses that do not (yet?) use Codemode. But the tide is shifting. In the meantime, a temporary crutch has been to do what Cloudflare did, and do Codemode within the MCP server. But now we have Codemode in Codemode which is pretty bad. It means double JSON escaping, easy for smaller models to get confused by and the inner code cannot call the outer tools. So if you for instance use the Cloudflare MCP servers in Pi, the agent needs to write JavaScript and funnel it through more JavaScript. This is really not optimal, but it’s also understandable that this is happening: const accRes = await tools.mcp__cloudflare__execute({ code: `async () => { const r = await cloudflare.request({ method: "GET", path: "/accounts" }); return r.result.map(a => ({ id: a.id, name: a.name })); }`, }); const accounts = JSON.parse(accRes.content.map(c => c.text).join("")); const out = []; for (const account of accounts) { const r = await tools.mcp__cloudflare__execute({ account_id: account.id, code: `async () => { const r = await cloudflare.request({ method: "GET", path: \`/accounts/\${accountId}/workers/scripts\`, }); return r.result.map(s => ({ id: s.id, modified: s.modified_on })); }`, }); out.push({ account: account.name, workers: r.content.map(c => c.text).join("") }); } return out; MCP Desires So to end things off: how well does Codemode work with MCP today? Well … not amazingly well. That’s because MCP servers are not really targeting harnesses that use Codemode yet (though at this point I think most harnesses support it). For this to work well some recommendations: Structured content: Codemode wants calls to return some nicely formatted JSON. So that needs to come back from the server, and many don’t do that yet. The outputSchema system in MCP is great for that. Consistent results: an interesting failure case is when an MCP server does not return consistent data. For instance because it tries to token optimize things depending on how many items are in the result set. This can cause an initial probe with 5 items to succeed, but then fail when the server returns the maximum batch size. Large binary data: today MCP does not yet support large binary data so quite a few use cases that are really interesting do not work well at all yet. You end up with all kinds of weird workarounds such as pre-signed URLs to allow file uploads then to happen through non MCP channels. Composable tool search: the MCP server might know better than the MCP client which tool is appropriate for a task. But there is no good mechanism today that allows a harness to fan out tool searches across multiple MCP servers. It’s all emergent behavior and it does not scale well to multiple active servers. Future of Codemode So where does this leave us? Is this a reversal of what I wrote a year ago where I encouraged CLIs? I don’t think so. In fact, the MCP ecosystem from my perspective picked up on exactly what we pointed out a year ago works: code. But Codemode goes beyond MCP in that it can act as a capable mechanism within the harness to express more freedom for the agent. There are however also some things that we still need to figure out. For one, durability with Codemode is trickier. We might have to adopt some ideas from durable workflow engines here to snapshot invocations. Or maybe, something like Starlark is a better composition language than JavaScript given its deterministic nature. Images, binary data and just the inability of this pattern to work with smaller models is also something that needs to be fleshed out. So it’s for sure not a perfect solution yet, but it’s quite a useful pattern that I expect us to leverage more.

14 hours ago • 1 votes
Coming soon: New York City’s hearing on AI risks

New York tries to take care of its own

yesterday • 1 votes
Dyson CameraJet

So when I saw Dyson had a $500 toothbrush, I was excited. Finally, advertising that targets me! I love brushing my teeth, and I have more money than I know how to spend. Not because I’m particularly rich, but because most stuff doesn’t really appeal to me. Like if I owned a helicopter it would just be a headache, because like imagine one day I get a call from the hangar saying the hangar is flooding and the water is rising and you need to move your helicopter. I’m thousands of miles away and need a helicopter pilot in the next 30 minutes, a new place to store it, was the maintenance even done will we even be able to take off on short notice and really I just am upset with myself because I made the poor decision to purchase a helicopter, and once I come back to reality I feel relieved that I don’t own a helicopter and this scenario will never happen to me. I do however, by means of my birthday, own a Dyson CameraJet (pictured above). It broke within 30 seconds of the first brushing. None of the LEDs turn on anymore. I spent an hour investigating, finally opening the user removable battery compartment to find the Spearmint Dyson Low-foaming mouth rinse had leaked inside. And by how the toothbrush is designed, it’s clear the entire electronics compartment was flooded with the stuff. Here’s the top comment on Reddit about this toothbrush. Apparently this is happening to everyone, “a potential for water seepage” they say. Dyson wants me to find the receipt and return it through some obtuse process that probably doesn’t work, dude it was a gift I just want my $500 toothbrush to work. They claim they worked on it for 6 years, but it’s clear their QA Process doesn’t include putting any liquid in the device. It clearly should, ideally for all devices but at least for spot checks on some. It’s sad to see this. At comma, we put every comma four in a highly stressful environment for 16 hours, a superset of the state it’s in driving, while testing all peripherals: the camera, IMU, GPS, screen, etc… We have gotten the failure rate super low by doing this, and for the few that do fail it’s usually after a while. There’s no excuse for a mature consumer electronics company to not design a procedure to fully test the functionality of each device before shipping. This shows some serious dysfunction at the company, and they should take this as a wake up call to fix their processes and issue a recall for the toothbrush. Dyson, if you see this post, e-mail me when I can drop by the Dyson store in ifc mall Hong Kong and swap it for a new one. I don’t want a stupid process, I want a real technical explanation of the issue and a working fancy toothbrush.

4 days ago • 1 votes
📚 BoredReading

You seem to be enjoying this.

Join free to unlock everything.

Create free account

Already have an account? Sign in