Full Width [alt+shift+f] Shortcuts [alt+shift+k]
Sign Up [alt+shift+s] Log In [alt+shift+l]
1

Fridays With Bob

from IEEE Spectrum [alt+shift+b] in AI

When I started at Spectrum 25 years ago, a senior editor suggested that I find a “rabbi,” by which he meant someone who could mentor me in how EEs approach problems and evaluate potential solutions. I didn’t find one right away. Then in 2005 we decided to do a special report, focusing on the challenges of enterprise software development. I suggested we invite IEEE Life Senior Member Robert N. Charette, a self-described risk ecologist, prolific book author, and leading authority on risk management and software engineering, to explore in our pages the myriad reasons software projects fail. His seminal article “Why Software Fails” is still read in university engineering classes today. IEEE Life Senior Member Robert N. Charette is one of IEEE Spectrum’s most prolific authors.Robert N. Charette It was, as they say, the beginning of a beautiful friendship. I had found my rabbi, one who shared my love of writing. We settled into a rhythm that would last more than 20 years, talking on Friday mornings about a range of topics including the growing ubiquity of software in our lives. So when I became Spectrum’s website editor in 2007, he was the first contributor I tapped to start a regular blog (remember those?). The Risk Factor was born and over the course of more than 10 years and 1,750 posts, Bob chronicled hundreds of software debacles, culminating in “Lessons From a Decade of IT Failures,” which won a Jesse H. Neal Award for Best Infographics in 2016. Ironically, yet predictably, those infographics were created in a software package that is no longer supported and thus are lost to the bits of time. “I like the expression on the fish just before it’s going to be swallowed by the heron.”Robert N. Charette Bob, however, was not a one-trick pony. In between his full-time job running his two management consultancies and raising a future biochemist and a future civil engineer, his daughters Maura and Megan, he also wrote many deeply reported and insightful articles....
1st Aug 2026

Stay updated

Get a weekly newsletter with the top 5 articles worth reading every week.

More from IEEE Spectrum

Electricity Theft Is Rampant, but Delhi Found a Fix

It’s disheartening how much power gets generated and then promptly lost as it travels through grid networks. This leaking of electricity happens when it vanishes as heat as well as when it is pilfered by thieves and nonpaying customers. More than half of the countries that track these metrics lost at least 10 percent of their electricity in 2023, according to the World Bank. Losses topped 20 percent for 24 of those nations. Two countries lost more than half of what they generated. With numbers this high, cutting down on losses makes sense. Electricity demand is rising beyond what many grid operators can supply; reducing waste would help meet some of that demand without having to build new power plants. Plus, when the power comes from fossil fuels, any loss means emitting even more greenhouse gases into the atmosphere. And electric losses hit the bottom lines of power providers, which ultimately pass those costs on to everyone else. The trouble is, reducing electricity losses is a hard and expensive process that takes a long time. Typically, the less maintained the grid infrastructure, the more electricity that’s lost. And the more fragile the region’s law enforcement and government, the more prevalent the power theft. Natural disasters and war make things worse. Delhi’s Power Grid Comeback Fixing a power grid requires a systemic approach across many sectors. There’s no one technology that will solve the problem. At the outset, the obstacles to success may feel insurmountable. Equipment across entire grid networks must be updated. Multiple arms of government must agree to reforms and coordinate to ensure power providers are set up to succeed. Regulations must be written or revised, investments made, cultures changed. The city of Delhi did all those things. Over the past 25 years, it cut its electricity losses from about 50 to 5 percent. How the city pulled off that impressive feat is the focus of “The Epic Comeback of Delhi’s Power Grid” by Mini Shaji Thomas, an electrical engineer at the university Jamia Millia Islamia who has lived in Delhi since the 1990s. She gives us a view from the inside—as a resident and a power systems expert. Delhi is a shining example, but some other regions have significantly lowered electricity losses over the last quarter century too. The country of Georgia went from losses of over 16 percent in 2002 to about 8 percent in 2023. In Singapore, losses dropped from 6.6 percent to a nearly nonexistent 0.2 percent over the same time period. Global Electricity Theft Crisis But there are many parts of the world where electricity losses remain a problem or have gotten worse. In Jamaica, where power theft is rampant, losses have hovered between 21 and 28 percent for years. Argentina’s losses nearly doubled between 2015 and 2023, going from an all-time low of about 12 percent to an all-time high of nearly 24 percent. The main problem: Transmission and distribution companies lacked the capital to maintain and upgrade their networks, which left equipment operating under stress. A delay in the installation of smart meters has allowed thieves to more easily siphon power and tamper with meters. Thomas says she hopes her account of Delhi’s grid comeback will serve as a blueprint for others. It’s possible to replicate the sweeping changes Delhi made, she says. But it “requires a concerted effort from all stakeholders, customers, the utility, the government, and their employees.”

6 days ago • 1 votes
Cash in on the AI Boom by Renting Out Your Spare Compute

If you own an at-home server, a gaming computer, or just a laptop that doesn’t get much love, listen up. You can now put that spare computing power to use and earn some passive income in the process. AI companies are hungry for more compute to run AI inference—the process of using a pre-trained model to respond to queries—and they’re willing to pay you for it. “Imagine Uber or Airbnb, but for AI inference computing tasks,” says Ilman Shazhaev, founder and CEO of Far Labs, based in Abu Dhabi. The AI boom has spurred on construction of massive data centers, often damaging local communities by raising electricity prices, straining local water resources, causing environmental damage and noise, and being just plain ugly. Huge data centers are likely not going anywhere—training new frontier models and running AI models from leading companies will likely still be the purview of these behemoths. But now, several companies are providing AI inference on smaller, mostly open-source models. They are running inference on pre-existing computing power spread throughout homes and small businesses, and compensating owners. “Everyone thinks the only way to do it is data centers. And data centers are extractive for the communities in which they’re built, and they don’t return services or taxes or much of anything to the people there. So why not just turn this whole thing on its head?” says John Federico, founder and CEO of Evolving Edge, in Austin, Texas. “The compute power is out there. If you can orchestrate it, then you’re actually adding value to those communities directly.” The idea isn’t entirely new: From 1999 to 2020, a volunteer-based project called SETI@Home used spare computers to search for signs of extraterrestrial life in radio telescope data, for instance. But now, commercial companies are eager to use the same strategy. Shazhaev’s Far Labs is launching its platform Far AI in the coming weeks, while Federico’s Evolving Edge is currently in open beta. Other companies, like Bless Network, Salad, and Gradient have started to provide similar platforms over the last year. Connecting to the network Federico has been a computer hobbyist since youth, and he has amassed a whole server in his basement to run his projects. “It just hit me one day, there’s all this talk about not having enough compute, and I just thought, well, 92 percent of the country has broadband, and you have people like me who have mini data centers in a closet,” he says. Federico sees the potential hosts as people much like himself who have already invested in home servers, and he aims to make the process of selling spare compute as seamless for them as possible. “Sign up for the program, install an application,” Federico says. “All we want to do is run jobs on your machine when you tell us we’re allowed to. The only thing we do is monitor the resource usage. And of course, you can give us a schedule.” With a large enough network of devices, the platform would have compute available whenever it’s needed. Privacy and security are primary concerns for such hosts. To reassure the users that their local data is secure, and that no malware will be downloaded to their devices, the team open-sourced their scheduling software. “The node software is open source, so anyone can look at it, see what it does. All we want to do is run jobs on your machine when you tell us we’re allowed to,” Federico says. Far Labs’ Shazhaev explains that the company’s software is designed around a principle known as “least privilege”: granting both the host and the user the least access possible to accomplish the task. Inference runs as an isolated workload with authenticated, encrypted communication and explicit limits on the GPU, CPU, memory, storage, and network resources it may use. Customers do not receive arbitrary access to the host machine, and providers can inspect resource use, pause the node, revoke access, and remove the software at any time. The protection also works in the other direction. Workloads are segmented and only the minimum required information is exposed to an individual node. Sensitive enterprise workloads can be restricted to controlled hardware rather than routed through consumer devices. Divide and conquer Massive data centers still have advantages from the user perspective: top of the line GPUs and CPUs, high speed networking, thick cables, and sophisticated cooling. User devices are usually less powerful, more varied, and less reliably connected to one another. “This is quite a difficult issue from the science angle,” Shazhaev says. “You want to do a similar level of tasks that are happening in those high infrastructure data centers, and run them on the user device with limited capacity.” Evolving Edge’s Federico says this is an issue for the largest, state-of-the art AI models. But those are not always needed and are often not even preferred. “There are numerous companies, once they reach a certain scale, suddenly paying for tokens on a state-of-the-art frontier model [that] no longer makes sense for their needs,” he says. “Instead, they are fine-tuning open-source models for specific tasks that they have in their business. These models don’t require anywhere near the resources that some of the state-of-the-art models do. It’s just using the right tool for the job.” Smaller, open-source models can often fit on a single user device. But if that fails, there are tools to split a single inference task over multiple GPUs or CPUs. Evolving Edge is using an open-source tool called Ray to perform this splitting, while Far Labs has developed its own proprietary software that not only splits the workload, but wraps the splitting in a layer of security and reliability-providing software. “One thing we have done is we shared the model,” Shazhaev says. “We take the model, we cut it into many pieces, then these pieces will be distributed through different devices. And we have an orchestrator and a load balancer which manage the task flow, so each device processes a part of the task. Then we combine the answers in the main brain, the orchestrator.” Through a combination of using smaller, more task-specific models, and splitting larger models between disparate devices, the teams claim they can perform inference much cheaper than a traditional data center “because we don’t have capital expenditure,” Shazhaev says. Gaming PCs are a common source of spare computational power in the home. Dizzaract The distributed advantage Not only is it cheaper to run inference this way, it is also more reliable, Shazhaev claims. The companies have access to a distributed network of computing resources, rather than one giant device that can experience outages. Shazhaev compares this to cryptocurrencies, and their resilience through decentralization. “Today, to shut down Bitcoin, you need to nuke the whole planet. Here, we have the same concept,” Shazhaev says. Federico explains that this resiliency would be beneficial not just for AI inference, but for all kinds of applications, including smart cities, environmental sensors, autonomous vehicles, and more. During an Amazon Web Services outage in 2026, for example, smart beds were stuck in their upright positions and their users couldn’t adjust them. Federico says that a distributed network where everything doesn’t need to be routed through a single data center, say, in Ashburn, Va., would make those kinds of outages much less impactful. “We could lose 100 nodes in a network of 250,000 and it wouldn’t matter,” he says. If the network of user devices is substantial enough, every job can be routed to a nearby device, decreasing the latency. Far Labs claims a latency of 100 milliseconds or less on its platform. The lower cost and lower latency of this approach may even enable new use cases, such as in-game AI video generation, which is currently prohibitively slow and expensive. “OpenAI last year had $30 billion in revenue, but they closed the financial year at an $8 billion loss. Why? The official reason is due to the high cost of inference,” Shazhaev says. “And those are mostly text models. For gameplay, you have audio, video, animations: It’s heavy data, and you need real-time responses. So, we’ve been trying to solve this issue.” All of these companies are trying to tap into an untapped resource of local compute, and hoping it’ll benefit the device hosts and users alike. “All these big guys are running around building data centers,” Shazhaev says, “but I believe there is enough compute power that already exists in the world.”

1st Sep 2026 • 1 votes
The Orbital Data Center Hype Machine Is Already in Orbit

“The lowest-cost place to put AI will be in space, and that will be true within two years, maybe three at the latest,” SpaceX founder Elon Musk told the World Economic Forum in Davos this past January, as his company was preparing to go public. Later that month, SpaceX filed an application with the Federal Communications Commission for an orbital data center constellation of up to 1 million satellites in low Earth orbit, 500 to 2,000 kilometers above Earth. And just three days before the IPO, he discussed some initial design specifications for a new AI-1 satellite data center in a video interview. Musk is prone to hyperbole when it comes to timelines. Full self-driving cars by 2017. First human mission to Mars in 2024. Ten thousand Optimus humanoid robots by the end of 2025. Et cetera. For orbital data centers, which he says will be a cost-effective alternative to terrestrial data centers within three years, the math won’t make sense for several years, if ever. Consider this: There are roughly 14,500 active satellites in orbit. Musk’s Starlink constellation accounts for about two thirds of those. Both the launch cadences and satellite-manufacturing capacity would have to scale up astronomically to deploy a million orbital data center satellites. For context, there have been roughly 7,000 orbital launches in all of human history. To loft 1 million satellites into low Earth orbit on SpaceX’s Starship, which is designed to carry up to 60 satellites per vehicle, would require 16,666 launches exclusively devoted to satellite deployments. Considering that SpaceX launched a record 165 orbital missions in 2025, even at 10 times that cadence, it would take a decade. And how long would it take to build 1 million satellites, given Starlink’s current pace of around 4,000 per year and a generous tenfold increase in capacity? Short of a manufacturing revolution, try 25 years. The reality is that the vision of massive constellations of orbital data centers is nowhere close to being realized. As this month’s cover story, “Why Orbital Data Centers Are So Hard” by Andrew Cavalier of ABI Research, makes clear, the reality is that the vision of massive constellations of orbital data centers is nowhere close to being realized. Dina Genkina, IEEE Spectrum’s computing and hardware editor, put the idea into perspective: “Starcloud (a startup that has applied to the FCC for an 88,000 orbital data center satellite constellation) sent one Nvidia H100 GPU in space so far. Their radiator was too weak to let the chip run at full power.” As Cavalier shows, cooling even a single Nvidia H100 GPU in space is difficult: It draws 700 watts, which will require 1.4 square meters of radiator at 60 °C. A 40-kilowatt rack of servers will need an 80-m² radiator; a 100-megawatt data center will require 2,500 of those radiators. Some astronomers are understandably concerned that a million satellites with giant radiative wings would blot out the stars. So if the economics doesn’t make sense, if the chips are at the mercy of the radiative ravages of space, and if humanity will lose its view of the stars, not to mention increasing the risk of triggering the Kessler syndrome, why are the hyperscalers hyping orbital data centers? Genkina offered the obvious answer: sweet, sweet moolah. “The Elon Musk part of it is honestly genius because he’s got xAI building the data centers, SpaceX sending them to space, and Tesla building solar panels,” Genkina says. “It’s almost like he’s paying himself.” Two Analyst’s Views of SpaceX’s Proposed AI1 Data Center Satellite Michael Pierce, Principal at Technology Strategy Partners Musk’s timelines are notoriously overly ambitious, but I think SpaceX’s orbital data centers might reach cost parity with terrestrial data centers in 5 to 10 years. The Starlink laser-link network already exists as the communication backbone for any SpaceX compute constellation, and that infrastructure is what no new entrant can replicate quickly. The chip-agnostic payload design probably reflects their disclosed difficulty securing AI silicon as much as any modularity philosophy. My view is that the only realistic near-term application is a SpaceX mega-constellation for inference. Training workloads likely cannot tolerate the synchronization and latency constraints of a distributed orbital system. Our report analyzed the market from the integrator’s vantage point, but AI1 is what it looks like when one player has assembled all the necessary advantages simultaneously. The question is whether the terrestrial data center industrial base will degrade or improve on economics. I don’t have insight into SpaceX’s internal costs, as opposed to public pricing, on all their components, so it’s hard to say if they’ll completely dominate or not. Even if they are not cost competitive with terrestrial data centers for another 5 to 10 years, it may simply be faster to get new compute that just happens to be in space. Matt Hasan, AI strategist and independent consultant My initial view is that AI1 does not fundamentally change the rationale for space-based data centers as much as it changes the timeline and scale. The underlying drivers remain the same: escalating AI compute demand, growing power constraints on terrestrial grids, and the desire to colocate energy generation with computation. What AI1 does signal is that the concept is beginning to move from theoretical discussion toward engineering and capital allocation decisions. The announcement adds credibility to the idea that hyperscale computing infrastructure may eventually expand beyond terrestrial constraints rather than simply competing for increasingly scarce grid capacity on Earth. That said, significant economic and technical questions remain. Launch costs, maintenance, hardware replacement cycles, thermal management, latency-sensitive workloads, and overall system economics will ultimately determine whether space-based data centers become a mainstream extension of AI infrastructure or remain a niche capability for specialized applications. The key development is not that these questions have been resolved, but that major industry players now appear willing to invest resources toward answering them.

1st Jul 2026 • 1 votes
IEEE President’s Note: Designing a Safer Digital World for Kids

Children born after 2013 are the first generation to grow up fully immersed in digital systems, which weren’t designed with them in mind. One‑third of the world’s Internet users are younger than 18, according to UNICEF, yet these systems shaping their daily lives were built for adults. They were optimized for engagement and designed long before people understood how profoundly digital environments influence children. For engineers and technical professionals, online safety is not an abstract policy debate. It is a design challenge that demands rigor, systems thinking, and ethical foresight. Governments around the world are also beginning to recognize the problem. Policymakers from across Australia, Brazil, the European Union, Indonesia, and the United States are responding to risks engineers have long understood: Addictive features, inappropriate content, opaque data practices, and algorithmic systems shape user behavior in ways that their creators did not fully predict. For years, technology moved faster than governance. Now governance is trying to catch up. Global Shift Toward Design Reform Supporting National Digital Ambitions In Athens this year I met with senior leaders of Greek government agencies and key national research institutions. Greece is moving quickly on digital transformation and responsible technology governance, and our discussions reinforced IEEE’s role as a trusted, neutral collaborator. We focused on supporting Greece’s ambitions in digital modernization and public‑sector innovation. We also discussed responsible AI and age-appropriate digital design in Europe and elsewhere. These engagements, grounded in shared values and long‑term commitment, strengthened IEEE’s presence within the European ecosystem and opened new pathways for collaboration on trustworthy AI and child‑focused digital well‑being. The European Union and the United Kingdom have been among the first to act, embedding age‑appropriate digital design into their broader children’s rights agenda. Drawing on IEEE expertise and global best practices, Indonesia is the first country in Asia, and Brazil is the first country in Latin America, to adopt age-appropriate design regulation. Australia is aiming to limit access to harmful content and addictive design features through age restrictions on certain platforms. And in the United States, in addition to federal efforts, states including California, New York, and Utah are enacting approaches including age-appropriate design principles. Across these efforts, a shared realization is emerging. Protecting children online is not simply about filtering content or adding parental controls. It requires rethinking the architecture of digital systems regarding how data is collected, how algorithms make decisions, how interfaces influence attention, and how AI interacts with the developing minds of young users. Engineers and technical professionals understand that design choices are never neutral. They encode values, incentives, and assumptions. When the user is a child, those choices carry greater weight. This is where IEEE’s work becomes more essential. Protecting Children Online For more than a decade, IEEE has been building technical and ethical foundations for safer digital experiences. The first IEEE standard on age-appropriate design in 2021 marked a turning point. It offers a structured, principled approach to designing with children’s rights in mind. The Institute’s 2022 article “Use a New IEEE Standard to Design a Safer Digital World for Kids” highlights how the standard helps translate those principles into engineering practice. Today the IEEE Standards Association’s (SA) Trustworthy Digital Experiences portfolio provides a practical, technically grounded framework for governments and industry. Spanning ethical design, data governance, algorithmic transparency, and child‑focused digital well‑being, it has already initiated discussions with government stakeholders around the world. This work helps bridge the gap between engineering realities and policy ambitions. No single country can solve these challenges alone. Many policymakers lack access to the combined expertise in technology, governance, and children’s rights needed to act quickly and effectively. This collaborative effort helps close that gap. The stakes are high. Without coordinated action, public policy will continue to lag behind technology, leaving children exposed to risks that could have been mitigated through thoughtful design. But with the right frameworks, governments can ensure digital systems respect children’s rights, support healthy development, and promote well‑being. IEEE’s emerging standards and collaborative technology policy work offer a path forward. By grounding national efforts in evidence‑based, rights-aligned design principles, IEEE is helping governments move from reactive regulation to proactive, coherent, and globally informed strategies for protecting children online. Safeguarding childhood in the digital age is both a moral imperative and an engineering challenge. And IEEE is helping to lead the way. —Mary Ellen Randall IEEE president and CEO Please share your thoughts with me: [email protected]. This article appears in the June 2026 print issue.

1st Jun 2026 • 1 votes

More in AI

Agent governance looks less like moral education and more like institutional economics

how agents turn ruthless on paper first

6 hours ago • 1 votes
We are going to kill “unalive”

I talk to a lot of old people, those who were born in the 20th century, and if I ask them what the word “unalive” means, they usually have no idea what I’m talking about, except for some of them who have kids or who study Internet culture. This will, of course, probably seem very weird to most people who were born in or grew up in the 21st century. Just to recap for the olds: “unalive” is the word you use to represent concepts like dying, or death, or killing or being killed, on digital platforms where saying those words accurately will cause the algorithm to punish or censor you. Or, maybe, where the perception is that using those words will result in being censored by the algorithm, and no one is actually willing to find out what happens if you use the forbidden words. This sort of attack on people’s expression started on platforms like TikTok, where nearly all content is distributed through an algorithmic feed, but has since become ubiquitous in nearly all digital media that we see. In fact, these tics are now so prevalent that it’s routine to hear people using this kind of language in everyday life, even though there’s not yet an algorithm to appease in the physical world. I’ve heard people say, out loud, “he unalived himself”, in reference to someone dying by suicide. And all of this has become even more visible in recent days as online conversation has turned to discussion of the horrific lack of accountability around the tragic rape case at Cornell University. Across the Internet, people are routinely referring to the central crime in the case as r*pe or “grape” or even using the 🍇 emoji, without a second thought for what it means that the very word can’t be said online anymore. Or, at least, the assumption is that it can’t be said. To be clear, I am very much in favor of people using content warnings or sensitivity markers for content, and fine with people using abbreviations like “SA” for references to disturbing or triggering topics like sexual assault; we should provide people with as much context and control as possible when choosing what information they want to consume and when. I also know that sometimes, people use lesser terms for stressful subjects like death or assault to create a bit of ironic distance from painful or upsetting topics. But most of the different variations of wording and emojis are coming from trying to appease the platforms, and there’s a heavy cost for those who are worried about being mindful: If someone is using a tool to filter out content, it will no longer be effective because everyone is using misspellings and euphemisms and imagery to get around the algorithm. The spread of censored and mangled syntax is happening because people believe, or have experienced, that platforms will silence them for accurately describing the world in plain language. This shit is terrible, and it has to stop. You Were Not Born With These Constraints One of the things that’s most concerning to me is that an entire generation has grown up not realizing how extreme it is that their very language is being chosen for them by platforms run by people who hate that generation’s ability to express itself, and who hate the things it has to say. From their youngest days, this generation grew up watching people make stupid faces at them for YouTube thumbnails and never had a chance to reflect on the fact that those creators didn’t want to be humiliating themselves by making those expressions — the demands of the algorithms of Big Tech forced them to do that. The rituals of feeding the algorithm are so built into people’s everyday habits that they’re invisible to people who weren’t alive before today’s platforms took over. Every parent of my cohort remembers the first time they heard their toddler finish doing something cute in their living room, and then turn around and say, “please like and subscribe!” afterwards. It’s a ghastly, sickening feeling to confront the fact that our little kids were being brainwashed into thinking that every adorable thing they did should be followed by a prompt to provide data to Google. Over on Instagram, where people originally signed up thinking they were going to see someone’s vacation pictures, or shots of their cousin’s kids, you’re now stuck watching people beg for everyone to reply with cultish phrases in the comments, which will then earn them an obviously AI-generated response in return, all in service of “showing activity” to the algorithm, like it’s an angry god that needs a sacrifice. They’re just not sure exactly what the angry god wants. Your free speech was taken away from you, and the people who did it are the same ones who spent years pretending to care about “free expression”. They contrived examples of lack of free speech on college campuses while squashing protests, and cried crocodile tears about “cancel culture” while getting people fired for political criticism. Now they have no problem with billionaires deciding exactly what words everyone is allowed to say. Larry Ellison is not content with his family owning all of the movies and TV shows — his family has to control what words people are allowed to speak on TikTok, too. Elon Musk isn’t content to merely generate and distribute child sexual abuse material for profit — he wants to silence the messages of the few decent people who are foolish enough to remain on Twitter/X, too. (That’s why I wrote you a guide on how to get your organization off of that cursed platform.) Now that an entire generation has grown up using these Orwellian euphemisms, and all of the Big AI products are trained on the Internet that was created under this regime, do you think today’s AI tools even know that the real, uncensored world exists? If you can’t say “genocide” on any of the major platforms, yet those are the ones all of the Big AI tools used as their training data... well, then the AI tools sure aren’t very likely to know much about genocide, are they? Fuck the Algorithm Our creativity can be constrained by the language we use — our imaginations are limited by what we can think to say. If we’re trained to limit the words we speak just by habit, and those limits are put in place by people whose social, political, cultural and moral goals are the opposite of what we value, then our work is unalive before it is even born. The answer to this is simple: say what you mean. This will take, to some degree, courage. It may even take, I hesitate to say, some sacrifice. When I suggest this course of action to people, they inevitably say, “But it will cost me audience!” or “But what if I lose followers!” or “What if they demonetize me!” Okay, what if they do? What if they do. Are you willing to push on this? To make a point about it? To move to platforms where you can actually say what you mean? Or to remember that you already have platforms where you can say what you mean? On an email newsletter or podcast you can say whatever the hell you want and nobody can stop you. On my blog right here, I can even curse in a headline and it won’t affect anything about how my site operates. (And a reminder: Substack is not an email newsletter, and a Spotify show is not a podcast — they’ll be unaliving your distribution any day now.) If you are a 20th century relic like me, it is incumbent upon you to remind the generation that grew up inside the algorithm that another world is possible, and that we know this because we lived it. We were able to style a MySpace page in any way that we wanted; the code for LiveJournal was entirely open source so there was no part of the algorithm that was unknowable. A blog like the one you’re reading right now could be made by anyone, and put up for pennies, and nobody could stop it from being read by millions of people. (And that last one? It’s still possible.) If you are from this century, forget all that rambling bullshit about ancient history: all that matters is you getting what you deserve, because you’ve been fucked over by the same billionaires who’ve poisoned your planet and infested your world with slop. The best artists around you are invisible to you, and the most important statements by activists that you care about are being silenced. It’s not a conspiracy, it’s a system working as designed. And the proof is as obvious as the fact that the angriest activists you know can’t even talk about systemic abuses or state violence without having to put it in algorithmically approved speech or censoring their captions like they’re going to be read by 5-year-olds. It should make you furious. It is time to kill “unalive”. The response is simple: For every message you put out, start by saying what you mean. Don’t work backwards from what the algorithm wants or what a platform permits. Build a presence on every platform you can, even the ones where you have fewer followers or where you’re harder to find. Tell your audience that your speech being free matters more than corporate convenience. Keep saying it, and keep it positive: independence gets them better art, better information, and more connected communities. Build alliances with other artists, activists and people who share your values, and let them know you’re going to start sharing your work uncensored. Start releasing your work uncensored and see where the platforms push back. (You may be surprised: sometimes you were censoring yourself in anticipation of limits that weren’t even there.) If a platform does try to limit your reach or expression, make a LOT OF NOISE about it. Tell the press, rally your alliance, and spread the word on your other platforms, using the moment to build audience and raise support there. Get others to amplify the parts of your work that don’t violate platform policy, so the controversy drives people to the rest. Find the others pushing back on algorithmic control of expression, and raise and praise their work when they do the same. If we keep accepting the words that are forced upon us by TikTok and Meta and Google and all the rest, while platforms like Twitter/X allow the most hateful and harmful content in the world to be distributed completely unfettered, we’ll only see authoritarianism rise, and the harms against the vulnerable accelerate. But what breaks my heart almost as much is that we’ll see so many brilliant artists and activists and thinkers whose genius will be muted or silenced by mindless, heartless algorithms that capriciously decide who gets to say exactly what words, in what ways. I get angry every time I think about it. The tech tycoons get ever more brazen in what they’re willing to say publicly, boasting about how they’re going to cause the end of the world, or calling for ethnic cleansing, all while putting tighter and tighter reins on the speech and expression of ordinary people. It’s time for “unalive” to die.

yesterday • 1 votes
What is Codemode

More than a year ago I wrote a few posts here that recommended people not to load custom tools into their context (or MCP servers) but to just use more scripts. Most importantly I wrote that Code Is All You Need and I wrote about that MCP needs code. With Pi 1.0 we now added MCP support via Codemode which in some ways is a long time coming, but then also maybe somewhat surprising to some. So I want to share some updated thoughts on this blog on what this all means. What Are Tools When a harness like Pi provides tools for an LLM to call, it does so by supplying some tool definitions which then translate into some token structure on the server side. Whether a model is encouraged to call a tool is the result of the reinforcement learning process. Something I wrote about before if you want to learn more. One of the reasons we strongly lean towards CLI and bash is because it allows easy composition of calls, and because the model also learns how the file system works when it’s trained. So when it invokes a tool like echo foo > /tmp/test.txt the model also learns that after that tool call, there is now a file called test.txt in /tmp. However bash has one fundamental limitation which is that it can only compose programs that run. And there are some things, which are not programs, but native tools to the LLM and they sort of have to be. The most obvious example here is read or view_image. If a multimodal model needs to read an image, it cannot use cat for that because the harness needs to inject the actual image payload into the protocol of the LLM. Another quite vivid example are sub agents. In order to spawn and orchestrate sub agents, it’s tricky to avoid tools that are provided by the harness. While in theory the agent could provide a CLI tool that talks to the outer harness via environment variables and Unix sockets, it’s a rather crude process. It however has another issue, and that is where the code runs. Brains vs Hands To better understand that, it’s important to think a bit more about where all the bits and pieces run. There really usually are two different systems involved. The first is the brain, the harness: it runs on one machine. It’s trusted. The second is often the same machine, but it’s really where the tools are executing: the hands. In Pi we now call this the execution environment, but you can think of it as the target of all the operations. Crucially what is important for us, is that there is a dividing line between the harness brain and the target environment that runs bash and executes the tools. And splitting this in half has some really important consequences. For a start it means that they are running on different file systems and they have different levels of trust. If you for instance use a sandboxing solution like Gondolin your bash stuff will be sandboxed just fine, but the harness itself will not be. Orchestrating The Harness Which brings us to what Codemode really does: it’s a way for the LLM to express and orchestrate complex operations on the harness side, but not the execution environment side. Codemode runs in the harness, in its own sandbox. In case of Pi it’s running in QuickJS within a WASM runtime with intentional limitations: no network, no file system, no timers, limited RAM. The only way is to call more tools. You could also imagine that Codemode could run Scheme or some other language as well. If you are not familiar with Codemode, it’s basically just a way to issue tool calls from within some language, in our case JavaScript. That allows you to compose those calls without necessarily going through the LLM’s context. Credit for naming goes to our friends at Cloudflare who coined it. For instance if you issue a bash call as a regular tool call in the LLM, then we only throw the trailing 2000 lines into the context and if the agent wants more, it needs to look at the overflow file itself. If however the agent issues that invocation via Codemode, then the Codemode side gets larger outputs sent structurally. Most importantly, because Codemode is JavaScript the agent can express concurrent operations and basic workflows. A common way in which you see agents now use this, is to first probe at 5-10 items from some tool response to see what it looks like, and to then write a Codemode script that processes the next n items. Codemode also allows you to throw state into the transcript! That means that one Codemode invocation can stash away data, that the next call in the session can load again. And remember: this is on the harness host, not the sandbox. In case of Pi, Codemode also allows you to issue calls that naturally do not make any sense in Pi’s traditional interface. For instance if you want to generate images with an image model or you want to classify some text with a one shot classifier model, those Pi APIs are exposed via Codemode, but not via regular tools where they would just waste context. What It Looks Like So now that we talked a bunch about it, it’s probably worth being a bit more explicit about it. Let’s walk ourselves through some invocations of Codemode of recent Pi sessions of mine. Note that none of this code is human written. It’s from real sessions of Pi, just re-indented for your viewing pleasure. The agent starts using Codemode automatically either because it’s a task where the model already naturally picks up that tool, or because a user asked it to. Note that Codemode is by default only enabled in Pi when MCP is enabled, but you can turn it on with "defaultTools": ["+codemode"] in the settings. Just ask Pi to enable it for you. Generating Images Let’s start simple with image generation. Image generation is a feature that Pi supports in the AI SDK core, but it’s not a tool that the agent can use. In the past the only way to use image models has been to write a bespoke extension or to have the agent run node itself and use the internal image APIs. However because we expose quite a few of the internal model APIs within Codemode, it means that the agent can use it: const [painter] = await models.getAvailableOfType("image"); const result = await models.generateImages(painter, { input: [{ type: "text", text: "A cute little puppy sitting on a grassy " + "lawn, soft natural light, photorealistic" }], }); if (result.stopReason !== "stop") return result.errorMessage; for (const block of result.output) { if (block.type === "image") image(block); else text(block.text); } Note that the call to image() sends the image back as image content to the LLM. On the harness side it feeds it directly into both the agent, as well as onto disk as a temporary artifact in case the agent wants to be able to pass that image back to bash. Classifying Things Similar things apply to classifier models such as Jev. They also do not fit well into the workflows of an agent through the typical tools. But rather than making a bespoke tool available, Codemode just allows the agent to reach into the AI SDK and invoke those directly. Here you can see how Jev is used to mass process GitHub issues for a quick sentiment analysis: const jev = await models.getModelOfType("classifier", "typesafe", "jev-latest"); const r = await tools.bash({ command: "gh issue list --state open --limit 100 " + "--json number,title,body,comments", }); const issues = JSON.parse(r.output); const results = await Promise.all(issues.map(async (issue) => { const res = await models.classify(jev, { state: { title: issue.title, body: (issue.body || "").slice(0, 4000), comments: issue.comments.slice(-5).map(c => c.body.slice(0, 800)), }, questions: { sentiment: { type: "choice", instructions: "What is the overall sentiment of the author towards pi?", criteria: { positive: "Appreciative, happy, constructive praise", neutral: "Matter-of-fact report or request without emotion", negative: "Frustrated, annoyed, upset, or angry", }, }, frustration: { type: "score", instructions: "How frustrated is the reporter?", criteria: ["not at all", "mildly", "clearly frustrated", "very angry"], }, kind: { type: "choice", instructions: "What kind of issue is this?", criteria: { bug: "Bug report or regression", feature: "Feature request or enhancement", question: "Question or support request", other: "Docs, discussion, meta, spam", }, }, }, }); if (res.stopReason !== "stop") { return { n: issue.number, title: issue.title, error: res.errorMessage }; } return { n: issue.number, title: issue.title, ...res.answers }; })); store("sentiment_results", results); return results .filter(r => !r.error) .sort((a, b) => b.frustration.score - a.frustration.score) .slice(0, 12) .map(r => `#${r.n} ${r.frustration.score.toFixed(2)} [${r.kind.choice}] ${r.title}`); Note how in that above example we also call store() which dumps the result of that execution into the session transcript. A future invocation of Codemode can thus read back that result if it wants to. The Promise.all here is fine, because Pi limits the total number of concurrent tool executions itself to four and maintains a queue for the rest. A more adventurous example is to use Jev to drive a game engine for debugging purposes: Codemode with Jev for Game Debugging Here it knows about my tankctl command and it built itself quickly a minimal harness around it to drive a game loop to assist a user with debugging a problem. Note how it built a 30 step loop in which each step goes back to both the game engine to get a text dump of what’s going on, and then to Jev to determine what to do next: const jev = await models.getModelOfType("classifier", "typesafe", "jev-latest"); const tank = async (cmd) => (await tools.bash({ command: `tools/tankctl "${cmd}"` })).output; await tank("start --map assets/maps/night_arena.map"); const questions = { action: { type: "choice", instructions: "You control the tank '@' in a top-down tank game. " + "Choose the best next action.", criteria: { attack: "an enemy has line of sight to you and you can fire at it", approach: "no enemy has line of sight; drive toward the nearest enemy", dodge: "an enemy shot is heading at you and will hit soon", powerup: "a powerup is close and no enemy threatens you", }, }, }; function commandFor(choice, st) { const p = st.player; const enemy = st.enemies.filter(e => !e.dead) .sort((a, b) => (b.los - a.los) || (a.dist - b.dist))[0]; if (choice === "attack" && enemy) { return `fire_at tank ${enemy.id}; frames 30 until clear,damage,kill`; } if (choice === "dodge") { // move perpendicular to the closest incoming shot const s = st.projectiles.filter(s => !s.yours) .sort((a, b) => a.eta - b.eta)[0]; const dir = s && Math.abs(s.vel[0]) > Math.abs(s.vel[1]) ? (p.pos[1] > s.pos[1] ? "+down" : "+up") : (p.pos[0] > (s ? s.pos[0] : 0) ? "+right" : "+left"); return `input ${dir}; frames 20 until damage; input stop`; } const powerup = st.powerups.filter(u => u.available) .sort((a, b) => a.dist - b.dist)[0]; if (choice === "powerup" && powerup) { return `goto ${powerup.pos[0]} ${powerup.pos[1]} 180`; } return enemy ? `goto ${enemy.pos[0]} ${enemy.pos[1]} 90` : null; } const log = []; for (let step = 0; step < 30; step++) { const st = JSON.parse(await tank("state")); if (st.state !== "playing") break; const threats = st.projectiles .filter(s => !s.yours && s.miss_dist < 1.5 && s.eta < 1.5) .map(s => `incoming shot dist ${s.dist} eta ${s.eta}s`) .join("\n") || "no incoming shots"; const r = await models.classify(jev, { state: { map: await tank("view 8"), threats, hp: st.player.hp }, questions, }); if (r.stopReason !== "stop") { log.push(`#${step} classifier error: ${r.errorMessage}`); break; } const choice = r.answers.action.choice; const cmd = commandFor(choice, st); if (!cmd) break; log.push(`#${step} hp=${st.player.hp} ${choice} -> ${await tank(cmd)}`); } return log.join("\n"); Calling MCP Servers Lastly, Codemode obviously is great for calling MCP servers. And because we do not actually expose any of the MCP tools to the LLM, the agent first uses provided APIs to issue a tool search within Codemode to discover what it might be able to do with the connected servers. This form of progressive discovery makes the whole MCP business work well enough for a lot of use cases today. Here for instance you can see the agent reach for the Sentry MCP straight away, even without discovering the tools, presumably because it has learned during the RL process already about what the Sentry MCP looks like. But it learns from what we inject into the system prompt, that the Sentry server is available to begin with. It’s not completely guessing here. const orgs = await tools.mcp__sentry__find_organizations({}); const { organizations } = orgs.structuredContent; const results = await Promise.allSettled(organizations.map(org => tools.mcp__sentry__find_projects({ organizationSlug: org.slug, regionUrl: org.regionUrl, }) )); return organizations.map((org, i) => { const r = results[i]; if (r.status !== "fulfilled") return { org: org.slug, error: String(r.reason) }; if (r.value.isError) return { org: org.slug, error: r.value.content }; return { org: org.slug, projects: r.value.structuredContent.projects.map(p => p.slug), }; }); Modern MCP Is A Fight I really don’t want to talk too much about MCP here, but MCP is in fact a protocol that greatly benefits from Codemode. The problem in parts is that MCP in practice often targets harnesses that do not (yet?) use Codemode. But the tide is shifting. In the meantime, a temporary crutch has been to do what Cloudflare did, and do Codemode within the MCP server. But now we have Codemode in Codemode which is pretty bad. It means double JSON escaping, easy for smaller models to get confused by and the inner code cannot call the outer tools. So if you for instance use the Cloudflare MCP servers in Pi, the agent needs to write JavaScript and funnel it through more JavaScript. This is really not optimal, but it’s also understandable that this is happening: const accRes = await tools.mcp__cloudflare__execute({ code: `async () => { const r = await cloudflare.request({ method: "GET", path: "/accounts" }); return r.result.map(a => ({ id: a.id, name: a.name })); }`, }); const accounts = JSON.parse(accRes.content.map(c => c.text).join("")); const out = []; for (const account of accounts) { const r = await tools.mcp__cloudflare__execute({ account_id: account.id, code: `async () => { const r = await cloudflare.request({ method: "GET", path: \`/accounts/\${accountId}/workers/scripts\`, }); return r.result.map(s => ({ id: s.id, modified: s.modified_on })); }`, }); out.push({ account: account.name, workers: r.content.map(c => c.text).join("") }); } return out; MCP Desires So to end things off: how well does Codemode work with MCP today? Well … not amazingly well. That’s because MCP servers are not really targeting harnesses that use Codemode yet (though at this point I think most harnesses support it). For this to work well some recommendations: Structured content: Codemode wants calls to return some nicely formatted JSON. So that needs to come back from the server, and many don’t do that yet. The outputSchema system in MCP is great for that. Consistent results: an interesting failure case is when an MCP server does not return consistent data. For instance because it tries to token optimize things depending on how many items are in the result set. This can cause an initial probe with 5 items to succeed, but then fail when the server returns the maximum batch size. Large binary data: today MCP does not yet support large binary data so quite a few use cases that are really interesting do not work well at all yet. You end up with all kinds of weird workarounds such as pre-signed URLs to allow file uploads then to happen through non MCP channels. Composable tool search: the MCP server might know better than the MCP client which tool is appropriate for a task. But there is no good mechanism today that allows a harness to fan out tool searches across multiple MCP servers. It’s all emergent behavior and it does not scale well to multiple active servers. Future of Codemode So where does this leave us? Is this a reversal of what I wrote a year ago where I encouraged CLIs? I don’t think so. In fact, the MCP ecosystem from my perspective picked up on exactly what we pointed out a year ago works: code. But Codemode goes beyond MCP in that it can act as a capable mechanism within the harness to express more freedom for the agent. There are however also some things that we still need to figure out. For one, durability with Codemode is trickier. We might have to adopt some ideas from durable workflow engines here to snapshot invocations. Or maybe, something like Starlark is a better composition language than JavaScript given its deterministic nature. Images, binary data and just the inability of this pattern to work with smaller models is also something that needs to be fleshed out. So it’s for sure not a perfect solution yet, but it’s quite a useful pattern that I expect us to leverage more.

yesterday • 1 votes
Do you actually like hard problems?

I spent years becoming the engineer people reached for. Now they still reach for me, but for different reasons.

yesterday • 1 votes
Coming soon: New York City’s hearing on AI risks

New York tries to take care of its own

2 days ago • 1 votes
📚 BoredReading

You seem to be enjoying this.

Join free to unlock everything.

Create free account

Already have an account? Sign in