Full Width [alt+shift+f] Shortcuts [alt+shift+k]
Sign Up [alt+shift+s] Log In [alt+shift+l]
1

Absurd Workflows: Durable Execution With Just Postgres

from Armin Ronacher's Thoughts and Writings [alt+shift+b] in AI

It’s probably no surprise to you that we’re building agents somewhere. Everybody does it. Building a good agent, however, brings back some of the historic challenges involving durable execution. Entirely unsurprisingly, a lot of people are now building durable execution systems. Many of these, however, are incredibly complex and require you to sign up for another third-party service. I generally try to avoid bringing in extra complexity if I can avoid it, so I wanted to see how far I can go with just Postgres. To this end, I wrote Absurd 1, a tiny SQL-only library with a very thin SDK to enable durable workflows on top of just Postgres — no extension needed. Durable Execution 101 Durable execution (or durable workflows) is a way to run long-lived, reliable functions that can survive crashes, restarts, and network failures without losing state or duplicating work. Durable execution can be thought of as the combination of a queue system and a state store that remembers the most recently seen execution state. Because Postgres is excellent at queues thanks to SELECT ... FOR UPDATE SKIP LOCKED, you can use it for the queue (e.g., with pgmq). And because it’s a database, you can also use it to store the state. The state is important. With durable execution, instead of running your logic in memory, the goal is to decompose a task into smaller pieces (step functions) and record every step and decision. When the process stops (whether it fails, intentionally suspends, or a machine dies) the engine can replay those events to restore the exact state and continue where it left off, as if nothing happened. Absurd At A High Level Absurd at the core is a single .sql file (absurd.sql) which needs to be applied to a database of your choice. That SQL file’s goal is to move the complexity of SDKs into the database. SDKs then make the system convenient by abstracting the low-level operations in a way that leverages the ergonomics of the language you are working with. The...
3rd Nov 2025

Stay updated

Get a weekly newsletter with the top 5 articles worth reading every week.

More from Armin Ronacher's Thoughts and Writings

Latent Powers

A few weeks ago I felt like it would be fun to see if I can make one of those cheap Chinese CarPlay dongles run something other than the stock firmware. The idea was that rather than just forwarding CarPlay, why not do something more interesting with them? They all work quite similarly: they act as bridges between your car and the phone. From there they deal with video and audio streams and pass some other data through. Most of them also bring up a custom UI for pairing and have a web interface that your phone can reach for updates. Long story short: I had a conversation with Fable and Sol via Pi about what could be done with such a dongle or whether I should use a Raspberry Pi instead if I wanted to do my own thing there. I figured it might be quite fun to run my own code while still allowing regular CarPlay to pass through. Through working with the LLM I learned about CatPlay, which is a Rust reimplementation of the CarPlay protocol that can run on Carlinkit devices. In particular, it can run on the Carlinkit Mini Ultra, which I figured would be easy enough to buy. I do have a few CarPlay adapters around, but I did not have that particular model, so I bought one on Amazon. Twenty-four hours later, I had a device in my hand that was branded as a Carlinkit Mini Ultra, but instead of being the Ingenic device that the original author used, it turned out to be something else. This is normally where the story would stop. However, it’s 2026. Armed with a bit of knowledge about how these systems work, I managed to have some fruitful discussions with Kimi K3 and Sol and figure out how flash the device and in turn, how to make CatPlay compile for that SoC. I guess that hacking these USB devices is not necessarily hard, but it’s laborious and you can easily end up bricking your devices. It also just sucks because sometimes you need to work with someone else’s code that does not itself run on your machine. In the past, I would abandon many such projects for lack of tenacity. But my clanker is tenacious. But so are all of our clankers. Some of the projects we’re now attempting are happening because of conversations we have with them. In this case I did not find or decide on CatPlay, the model did. It was not the only suggestion, but it became the best starting point after discarding others. And I discover this more and more. Particularly when we have solitary interactions with these models, some of us “independently” decide to work on similar projects. When I talked with an acquaintance about CarPlay he also mentioned recently that he decided to try something similar because he too wanted to see if he can get his own agent be hooked up with the car. And guess what: he too learned about the CarPlay hacking community, and that it’s an option, from the models and roughly around the same time. It really got me thinking about how this could create situations in which completely independent people end up building things they believe are their own ideas. Yet they were inspired or pushed towards doing something by a conversation with an LLM — a conversation that someone else also had. What if we took paths, because those were the paths that were more likely with current generation models? There is a running joke in the AI builder community right now that we’re all working on the same things, and in many ways it feels like we are. That might be because those things are obvious, or it might be partly because we all use the same models with the same capabilities. A few months ago, I first saw Lucas Meijer share the idea to make a model in Pi produce HTML reports rather than Markdown. I thought that was pretty unique. Except, well turns out the models are probably trained more and more for that (e.g. Claude Artifacts), and now it has become for many the default choice for sharing reports. How much of what we build comes from eliciting the same latent capabilities from the same models? Did the models make us prompt them that way? Was it because we shared ideas on Twitter and other communities that inspired us? Or is it all unrelated? There is something powerful and strange about how LLMs diffuse knowledge and capabilities, while perhaps also nudging us all simultaniously and independently toward building the same things.

4 weeks ago • 0 votes
What Is Reasoning

A few weeks ago a paper was shared that showed how to extract reasoning traces from closed-weight models. Together with online discussions about tricking models into leaking them, it made me investigate it more out of curiosity. Twitter seems full of half-truths and confusion about how this works, so perhaps this helps some to understand what is happening. Hiding Traces Reasoning traces are usually hidden from us. We have lamented this, but mostly have to accept it. Open-weight models thankfully reveal them, and from their behavior you can see that their traces can be long and confusing. This is probably a good reason to separate them from what is normally shown to users. At minimum, UIs need to detect them. The industry has done a good job at making reasoning traces sound special and exotic, but they really are just text: the model is trained to emit its thinking into a scratchpad as part of its response, before its final answer. GPT-OSS’s Harmony response format makes this easy to see: <|channel|>analysis<|message|> I need to work this out ... <|end|><|start|>assistant<|channel|>final<|message|> The answer is ... <|return|> The markers are special tokens, but the reasoning between them uses “the same text” as the final answer (just that GPT chain-of-thought text sounds really funny). When the model samples the analysis channel token, a parser routes the following text into a separate stream exposed through the Responses API. For closed models, presumably a simple model redacts and summarizes it. Reasoning Effort How much budget goes to reasoning? Earlier APIs exposed reasoning token budgets, making it seem like a property of the sampling process. In reality, reasoning effort is baked into the system prompt. GPT-OSS puts this into the system prompt: Reasoning: low That’s it. Training produces the resulting behavior, such as emitting the token sequence that switches to the analysis channel. This also explains why changing the effort invalidates the KV cache. I think closed GPT models call reasoning effort “juice,” since you can ask most models how much juice they have. In DwarfStar for DeepSeek with max reasoning this is added to the system prompt: Reasoning Effort: Absolute maximum with no shortcuts permitted. You MUST be very thorough in your thinking and comprehensively decompose the problem to resolve the root cause, rigorously stress-testing your logic against all potential paths, edge cases, and adversarial scenarios. Don’t Think The destination of reasoning tokens is therefore a learned convention: the model is trained to keep scratch work out of the final channel. Trick it into thinking it is in that channel and it may leak tokens. We have even seen older models, when thinking is disabled, reason into the bash tool and echo their thoughts to /dev/null. So in some sense the only “special” behavior for some models is not to think. That at times is done by “mechanically” removing the model’s usual ways to think. In DwarfStar, disabled thinking uses the prefill </think>, while enabled thinking uses <think>, which are the tokens that close and start thinking. GPT-OSS doesn’t prefill but lets the model decide either way on its own. But presumably, some inference APIs prefill the opening token when reasoning is enabled, so the model never samples it itself and might prevent the sampling of the reasoning token when disabled since it can be trivially detected. This may explain why a custom think tool can trick models into putting some reasoning where it should not go — but only when native reasoning is disabled. Fun fact: this blog post triggered safey checks Hilariously enough I was unable to use GPT 5.6 terra for spell and grammar checking on this blog post because of safety filters. Had to switch to Kimi.

19th Aug 2026 • 1 votes
Better Models: Worse Tools

A very strange Pi issue sent me down a rabbit hole over the last two days. The short version is that newer Claude models sometimes call Pi’s edit tool with extra, invented fields in the nested edits[] array. And not Haiku or some small model: Opus 4.8. The edit itself is usually correct but the arguments do not match the schema as the model invents made-up keys and Pi thus rejects the tool call and asks to try again. That alone is not too surprising as models emit malformed tool calls sometimes. Particularly small ones. What surprised me is that this is getting worse with newer Anthropic models as both Opus 4.8 and Sonnet 5 show it but none of the older models. In other words, the SOTA models of the family are worse at this specific tool schema than their older siblings. In case you are curious about Fable: I intentionally did not test it because I was not sure if the classifiers they are running might downgrade me to Opus silently. Tool Calls Are Text If you have not spent too much time looking at LLM tool calling internals, the important thing to understand is that tool calls are not magic and use some rather crude in-band signalling. The model receives a transcript, a system prompt and a list of available tools. The server munches that into a large prompt with special marker tokens. Because the model was trained and reinforced on examples of that format, at some point during generation it emits something that the API or client interprets as “call this tool with these arguments”. For a file edit tool, the intended invocation payload might say something like this: { "path": "some/file.py", "edits": [ { "oldText": "text to replace", "newText": "replacement text" } ] } A harness then validates the arguments, performs the edit, and feeds the result back into the model. If validation fails, the model sees an error and usually tries again. How exactly that formatting happens is not known for the Anthropic models, but some people have gotten out “ANTML” markers and they at times do leak also into public communications. To the best of my knowledge, the call above would come out serialized like this from the model: <antml:function_calls> <antml:invoke name="edit"> <antml:parameter name="path">some/file.py</antml:parameter> <antml:parameter name="edits"> [ { "oldText": "text to replace", "newText": "replacement text" } ] </antml:parameter> </antml:invoke> </antml:function_calls> An important thing to note here is that this thing, while looking like XML, is not really XML. It’s just a thing they found convenient to tokenize and train on. The other thing to note is that a basic top-level string parameter appears in-line whereas an array of objects is implemented via JSON serialization. While I’m not entirely sure that this is how it works, there are some indications that this is not too far off. This will become relevant later. There are two very different ways to make the model produce a structure like this: You can ask the model to produce valid JSON matching a schema and then validate it afterwards. You can constrain the sampler so that invalid JSON, or even invalid schema shapes, cannot be sampled in the first place. The second approach is what people usually refer to as grammar-aware or constrained decoding. The sampler masks out tokens that would violate the grammar. If the model is currently inside a JSON object and the schema says only oldText and newText are allowed, the sampler can prevent it from emitting "in_file" or "type". Grammar-aware decoding can be used both to constrain something to be syntactically valid JSON and also to enforce specific enum values or keys. Without any form of constraints the model is merely following a learned convention. The Failure Pi’s edit tool supports multiple exact string replacements in one call. That is why the arguments contain an edits array. In the failing cases the model produces entries like this: { "oldText": "...", "newText": "...", "requireUnique": true } or this: { "oldText": "...", "newText": "...", "oldText2": "", "newText2": "" } Across repeated trials I saw a whole zoo of invented trailing keys: type, id, kind, unique, requireUnique, matchCase, in_file, forceMatchCount, children, notes, cost, oldText2, newText2, oldText_2, newText_2, and even an event.0.additionalProperties key inside the edit object itself. The most annoying part is that the actual oldText and newText payloads were byte-correct in the invalid calls I inspected. The model had in fact produced the right invocation but then added nonsense at the end of the object. The failure is also heavily context-dependent. A fresh single-turn prompt like “edit this file” did not reproduce it at all for me. An agentic history where the model had read files, diagnosed a problem and then composed a multi-line edit could reproduce it. And more annoyingly, not all transcripts will show that behavior. In fact, I needed Petr Baudis‘s transcripts to reproduce this for me at all! In that user’s session continuing the session caused Opus 4.8 to fail around 20% of the time. Stripping thinking blocks from history reduced the failure rate by half. Turning on strict tool invocation eliminated it in my runs. Why It’s Getting Worse My strongest hypothesis is that this is not random deterioration but a training artifact. When older Anthropic models were trained, they were trained on some tools (some of which were documented). But that training did not yet have a user-shipped harness like Claude Code as the obvious target. Modern Anthropic models are most likely different because their post-training includes Claude Code or a harness that looks very similar. The model learns what a successful tool call looks like in that environment. It also learns what mistakes are tolerated by that environment. Claude Code’s own tools are comparatively flat. The ordinary edit tool is not Pi’s nested edits[] shape; it is closer to file_path, old_string, new_string, and an optional flag (replace_all). Looking at Claude Code’s client is very instructive: it contains retry paths for malformed tool use, parameter aliases, type coercions, Unicode repairs and filtering of unknown keys. In other words, Anthropic’s own client appears to expect and accept a fair amount of slop and repairs it, mostly silently. If reinforcement learning happens in a harness like that, or a simulation of one, then slightly malformed tool calls can still complete the task and receive reward. The harness fully absorbs the error and there is little gradient against inventing an alias, adding a stray field or using a nearby parameter name. Worse, the model may become very strongly adapted to the canonical Claude Code edit tool shape. A different harness can present a tool with the same semantic intent but a different schema. Such a tool can increasingly be off-distribution. The better-trained model might actually fight you harder because its prior is stronger. This is not too surprising, but it is a change from how this was a few months ago. When Opus 4.5 launched, it adapted to other edit tools exceptionally well. In fact, I was pretty convinced that we’re on a good path where the models are more likely to adapt to any sort of tool shape that comes around for as long as the instructions are good. Now I’m somewhat worried about the track we’re on here. Alternative tool schemas might not just be unfamiliar. They might be implicitly punished by post-training that optimizes for one particular, forgiving tool ecology. And that ecology is not documented. While there is a text editor tool that is documented, you will see that this format is in fact not followed by Claude Code. What Claude Code does internally (which is a closed-source harness) is hidden from you. The Slop Harness Claude Code is obviously closed-source but we can look at the minified code and get some idea of what it does. And honestly, it’s very forgiving of incoming data. For a start, Claude Code checks the model’s visible text for leaked <invoke markup. It also emits some telemetry when that happens and then it has its own state machine to retry such bad calls by pushing back to the model. It has explicit Unicode escape repair which fixes broken \uXXXX sequences and lone surrogates in string values. It also has per-tool aliases for parameters. For instance, Edit accepts old_str (presumably from the times when the models were trained on the officially documented text editor tool), the newer old_string from the schema, new_str/new_string, path as an alias for file_path, and some more. It also silently filters out unexpected keys and it does not use strict mode either. The issue with strict mode is that Anthropic applies complexity limits to the tool definitions that cause API requests to fail, so presumably that’s why Claude Code does not attempt to use it. Strictness Will this problem be with us in other harnesses too? One huge issue with Anthropic is that the models are completely closed, and so is the harness. Codex models are also closed, but at least the harness is not. We also have gpt-oss which is at least a bit interesting. The models are explicitly trained to use OpenAI’s harmony response format and there is a lot of documentation that at least tells us how OpenAI people think about this. Harmony makes channels and tool-call content types part of the prompt format. A function call can look like this: <|start|>assistant<|channel|>commentary to=functions.get_weather <|constrain|>json<|message|>{"location":"San Francisco"}<|call|> The important bit is <|constrain|>json. The model can express in-band that this message body is JSON, and an inference stack can use that boundary to switch into JSON-constrained sampling for the body of the tool call. Presumably a bit of this also happens in Anthropic’s models, at least in strict mode I would imagine. The marker in harmony helps the sampler to detect when it needs to sample with a specific grammar, and because it is part of the transcript, it makes that rather easy to do. For hosted GPT models, there is also an option to provide a LARK grammar for custom tools that need to adhere to something like this. Anthropic appears different from that, though maybe not entirely. If an array of objects is represented as JSON, as it appears to be, then the model has to write JSON inside the tool parameter. There is probably basic grammar-constrained sampling going on, and that may partly explain the extra keys. For a nested array parameter, that JSON includes escaped multi-line file content inside string literals, inside one tag. The unexpected, made-up keys appear exactly at the highest-entropy point of that task: after closing a several-hundred-token escaped newText string, where the model must decide } vs , "...". As strict mode in Anthropic appears to fix this, I presume that on the server side they are refusing to sample a key that is not permitted by the JSON schema structure. That would also explain why they have limits to the complexity of the tool definitions when strict mode is enabled. So far, the Codex models I tested did not show this type of regression. I tested all available ones except 5.6, which I do not have access to yet. What This Means For Harnesses The uncomfortable lesson is that tool schemas are not neutral, at least not on Anthropic models. We like to pretend that a schema is an abstract contract and the model is a general reasoner that will follow it, but that might no longer be the case for some of the tools. Tool schemas are somewhere in the distribution and some shapes are close to what the model saw during post-training and some are far away. Some are easy for the provider’s hidden encoding (e.g. top-level attributes in ANTML), whereas some require the model to write large escaped JSON objects inside nested arrays after long multiline strings. The model may be smart enough to understand the schema and still be bad at sampling the exact shape under pressure. If this type of model behavior continues, I wonder what the implications for harnesses are. Obviously one could turn on strict sampling in Anthropic and the problem should go away. On the other hand, that the model has this behavior shows the impact that reinforcement learning has on them. Fighting that prior is probably futile if you want to get the best model performance. Right now the reality is that Claude Code is not open source and we cannot really know what they are doing in their RL environments either. We cannot assume Claude-Code-trained behavior will transfer cleanly to your tools unless they are a close match. The more post-training happens inside one dominant harness, the more every other harness will have to inherit its quirks. I used to be more skeptical of strict grammar-constrained tool invocation because constrained decoding can have quality tradeoffs. I still think that can be true in general, but this bug moved my priors significantly. If the newest models get better at solving the task while getting worse at faithfully emitting an alternative tool schema, then the harness needs stronger guarantees somewhere. If you want to find out more, or you want to discuss this, consider reading the issue on the Pi tracker.

4th Jul 2026 • 1 votes
Communities of Not

There is a strange thing that happens in communities that gather around abstinence from something: identity from opposition. At their best these communities are not just negative: childfree spaces can be about autonomy, choice and acceptance, anti-car spaces about safer streets and transit, and LLM-skeptical developer spaces about the future of labor, code quality and slop1. But the thing being refused often does not go away and instead becomes the main subject of the community’s identity. That would be fine if it stayed at criticism, maybe even angry criticism, but more often than not it turns into policing and hatred towards others. An influencer without children becomes a parent, an urban bike commuter by choice buys a Porsche, a respected developer tries LLMs, and the community feels betrayed because it assumed they were members of the same tribe. The expulsion of that person (who never signed up to be a community member) is entirely imaginary but the punishment that the community unleashes is not: people pile on and shame them, quote them out of context and turn their weakest moments into proof that the person was always unserious, a sharlatan or should not be listened to. I do not think the answer is to tell people to stop paying attention. Cars shape cities even for people who cycle, children influence politics, workplaces and taxes even for people who do not have them. For us developers, LLMs show up in editors, issue trackers, hiring conversations, management pressure and code reviews whether we asked for them or not. Resisting that can be legitimate but that is no excuse for using one’s rejection to justify shitty mob behavior. I understand the thinking all too well, because I have done versions of this myself in the past. It took me a while to become more accepting of other people’s worldviews that diverge from mine. Whatever insecurities we have, finding a group of others sharing them can be comforting. The danger is that being part of a crowd of negativity can easily make us part of collective harassment. I can only encourage you to breathe, slow down, de-escalate when given the chance, and resist the temptation to always assume the most catastrophic reading. Default to being open to new things. Being negative towards something, and making that ones identity, is an easy trap to fall into. These examples are not meant as equivalents. The recent mob against rsync is the LLM version that prompted this post. I picked the others because I’m familiar with those communities and they all show similar cases of personal choices being interpreted as betrayal.↩

6th Jun 2026 • 1 votes
Content for Content’s Sake

Language is constantly evolving, particularly in some communities. Not everybody is ready for it at all times. I, for instance, cannot stand that my community is now constantly “cooking” or “cooked”, that people in it are “locked in” or “cracked.” I don’t like it, because the use of the words primarily signals membership of a group rather than one’s individuality. But some of the changes to that language might now be coming from … machines? Or maybe not. I don’t know. I, like many others, noticed that some words keep showing up more than before, and the obvious assumption is that LLMs are at fault. What I did was take 90 days’ worth of my local coding sessions and look for medium-frequency words where their use is inflated compared to what wordfreq would assume their frequency should be. Then I looked for the more common of these words and did a Google Trends search (filtered to the US). Note that some words like “capability” are more likely going to show up in coding sessions just because of the nature of the problem, so the actual increase is much more pronounced than you would expect. You can click through it; this is what the change over time looks like. Note that these are all words from agent output in my coding sessions that are inflated compared to historical norms: Loading word trend chart… The interactive word trend chart requires JavaScript. Something is going on for sure. Google Trends, in theory, reflects words that people search for. In theory, maybe agents are doing some of the Googling, but it might just be humans Googling for stuff that is LLM-generated; I don’t know. This data set might be a complete fabrication, but for all the words I checked and selected, I also saw an increase on Google Trends. So how did I select the words to check in the first place? First, I looked for the highest-frequency words. They were, as you would expect, things like “add”, “commit”, “patch”, etc. Then I had an LLM generate a word list of words that it thought were engineering-related, and I excluded them entirely from the list. Then I also removed the most common words to begin with. In the end, I ended up with the list above, plus some other ones that are internal project names. For instance, habitat and absurd, as well as some other internal code names, were heavily over-represented, and I had to remove those. As you can see, not entirely scientific. But of the resulting list of words with a high divergence compared to wordfreq, they all also showed spikes on Google Trends. There might also be explanations other than LLM generation for what is going on, but I at least found it interesting that my coding session spikes also show up as spikes on Google Trends. The Rise of LLM Slop The choice of words is one thing; the way in which LLMs form sentences is another. It’s not hard to spot LLM-generated text, but I’m increasingly worried that I’m starting to write like an LLM because I just read so much more LLM text. The first time I became aware of this was that I used the word “substrate” in a talk I gave earlier this year. I am not sure where I picked it up, but I really liked it for what I wanted to express and I did not want to use the word “foundation”. Since then, however, I am reading this word everywhere. This, in itself, might be a case of the Baader–Meinhof phenomenon, but you can also see from the selection above that my coding agent loves substrate more than it should, and that Google Trends shows an increase. We have all been exposed to LLM-generated text now, but I feel like this is getting worse recently. A lot of the tweet replies I get and some of the Hacker News comments I see read like they are LLM-generated, and that includes people I know are real humans. It’s really messing with my brain because, on the one hand, I really want to tell people off for talking and writing like LLMs; on the other hand, maybe we all are increasingly actually writing and speaking like LLMs? I was listening to a talk recording recently (which I intentionally will not link) where the speaker used the same sentence structure that is over-represented in LLM-generated text. Yes, the speaker might have used an LLM to help him generate the talk, but at the same time, the talk sounded natural. So either it was super well-rehearsed, or it was natural. Engage and Farm At least on Twitter, LinkedIn, and elsewhere, there is a huge desire among people to write content and be read. Shutting up is no longer an option and, as a result, people try to get reach and build their profile by engaging with anything that is popular or trending. In the same way that everybody has gazillions of Open Source projects all of a sudden, everybody has takes on everything. My inbox is a disaster of companies sending me AI-generated nonsense and I now routinely see AI-generated blog posts (or at least ones that look like they are AI-generated) being discussed in earnest on Hacker News and elsewhere. Genuine human discourse had already been an issue because of social media algorithms before, but now it has become incredibly toxic. As more and more people discover that they can use LLMs to optimize their following, they are entering an arms race with the algorithms and real genuine human signal is losing out quickly. There are entire companies now that just exist to automate sending LLM-generated shit and people evidently pay money for it. Speed Should Kill If we take into account the idea that the highest-quality content should win out, then the speed element would not matter. If a human-generated comment comes in 15 minutes after a clanker-generated one, but outperforms it by being better, then this whole LLM nonsense would show up less. But I think that LLM-generated noise actually performs really well. We see this plenty with Open Source now. Someone builds an interesting project, puts it on GitHub and within hours, there are “remixes” and “reimplementations” of that codebase. Not only that, many of those forks come with sloppy marketing websites, paid-for domains, and a whole story on socials about why this is the path to take. I have complained before that Open Source is quickly deteriorating because people now see the opportunity to build products on top of useful Open Source projects, but the underlying mechanics are the same as why we see so much LLM slop. Someone has a formed opinion (hopefully) at lunch, and then has a clanker-made post 3 minutes later. It just does not take that much time to build it. For the tweets, I think it’s worse because I suspect that some people have scripts running to mostly automate the engagement. And surely, we should hate all of this. These low-effort posts, tweets, and Open Source projects should not make it anywhere. But they do! Whatever they play into, whether in the algorithms or with human engagement, they are not punished enough for how little effort goes into them. Friction and Rate Limiting That increases in speed and ease of access can turn into problems is a long-understood issue. ID cards are a very unpopular thing in the UK because the British are suspicious of misuse of a central database after what happened in Nazi Germany. Likewise the US has the Firearm Owners Protection Act from 1986, which also bans the US from creating a central database of gun owners. The gun-tracing methodologies that result from not having such a database look like something out of a Wes Anderson movie. We have known for a long time that certain things should not be easy, because of the misuse that happens. We know it in engineering; we know it when it comes to governmental overreach. Now we are probably going to learn the same lesson in many more situations because LLMs make almost anything that involves human text much easier. This is hitting existing text-based systems quickly. Take, for instance, the EU complaints system, which is now buckling under the pressure of AI. Or take any AI-adjacent project’s issue tracker. Pi is routinely getting AI-generated issue requests, sometimes even without the knowledge of the author. Trust Erosion and Gaslighting I know that’s a lot of complaining for “I am getting too many emails, shitty Twitter mentions, and GitHub issues.” I really think, though, that now that we know that it’s happening, we have to change how we interact with people who are increasingly automating themselves. Not only do they produce a lot of shitty slop that we all have to sit through; they are also influencing the world in much more insidious ways, in that they are influencing our interactions with each other. The moment I start distrusting people I otherwise trust, because they have started picking up LLM phrasing, it erodes trust all over society. You also can’t completely ban people for bad behavior, because some of this increasingly happens accidentally. You sending Polsia spam to me? You’re dead to me. You sending me an AI-generated issue request and following up with an apology five minutes later? Well, I guess mistakes happen. Yet, in many ways, what is going on and will continue to go on is unsettling. I recently talked with my friend Ben who said he forced someone to call him to continue a conversation because he was no longer convinced he was talking to a human. Not all of us have been exposed to the extreme cases of this yet, but I had a handful of interactions in which I questioned reality due to the behavior of the person on the other side. I struggle with this, and I consider myself to be pretty open to new technologies and AI in particular. But how will my children react to stuff like this? My mother? I have strong doubts that technology is going to solve this for us. Suggestions for Change The reason I don’t think technology is going to solve this for us is that while it can hide some spam and label some generated text, it won’t fix us humans. What is being damaged here are social interactions across the board: the assumption that when someone writes to you, there is a person on the other side who has put some care into the interaction. I would rather have someone ghost me or reject me than send me back some AI-generated slop. Change has to start with awareness and an unfortunate developmend is that LLMs don’t just influence the text we rea and influence the text we write, even when we don’t use htem. Given the resulting ambiguity, we need to become more aware of how easily we can turn into energy vampires when we use agents to back us up in interactions with others. Consider that every time someone reads text coming from you, they will have to increasingly have to make a judgement call if it was you, or an LLM or you and an LLM that produced the interaction. Transparency in either direction, when there is ambiguity, can help great lengths. When someone sends us undeclared slop, we need to change how we engage with them. If we care about them, we should tell them. If we don’t care about them, we should not give them visibility and not engage. When it comes to creating platforms and interfaces where text can be submitted, we need to throw more wrenches in. The fact that it was cheap for you to produce does not make it cheap for someone else to receive, and we need to find more creative ways to increase the backpressure. GitHub or whatever wants to replace it, will have a lot to improve here and some of which might be going against it’s core KPIs. More engagement is increasingly the wrong thing to look at if you want a long term healthy platform. Whatever we can do to rate-limit social interactions is something we should try: more in-person meetings, more platforms where trust has to be earned, and maybe more acceptance that sometimes the right response is no response at all. And as for AI assistence on this blog, I have an AI transparency disclaimer for a while. In this particular blog post I used Pi as an agent to help me generate the dynamic visualization and I use the agent to write the code to analyze and scrape Google Trends.

4th May 2026 • 1 votes

More in AI

Why do OpenAI's GPT-2 weights beat mine? Part five: data quality

When I finished learning how to build an LLM from scratch, I was left with a mystery: my own models were not as good as OpenAI's original GPT-2 models, despite being based on the same architecture. My models all had 163M parameters, and followed the design from Sebastian Raschka's book "Build a Large Language Model (from Scratch)". That meant that they were pretty much the same as the setup for the OpenAI GPT-2 "small" instance, except that they did not use weight-tying or bias on the QKV matrices. Weight-tying means that you re-use the initial embedding matrix as the output head at the end, and using it means that GPT-2 small saved quite a few parameters -- it was 124M rather than 163M -- at, at least in my own experiments, a cost in quality; similarly, while I found that QKV bias made a tiny improvement in loss terms, I'd felt it was likely within the noise. But GPT-2 small consistently beat my models on an instruction fine-tuning (IFT) task -- also adapted from Raschka's book. That test fine-tunes the model on a subset of the Alpaca dataset, until validation loss starts rising, and then runs a test set through the resulting model. The responses to the test set questions are stored, and then I run all of the responses from all of the models under test past GPT 5.5 in one go to get an aggregate score; more details here. GPT-2 small always did better than any of my models on this. Additionally, it did surprisingly well on a simpler eval -- one that just measured the cross entropy loss it got on a test set. It scored close to my own best models, and better than many of them. What made this result particularly interesting was that the test set in question was a split of my own training data; my models would not have seen it when training (at least, in theory), but it seems likely that it would be much more similar to their own training data than it was to OpenAI's. I've checked two things while probing this mystery: It seems very likely that the GPT-2 models were overtrained by modern standards; would overtraining my own models get them closer? It turned out that no, it probably didn't help with the IFT eval (though there might have been some signal there). It did help quite a lot with the test loss eval, though. The way I was handling dropout in the IFT test might have been unduly benefiting some models while working against others. I decided to standardise on not using dropout during this eval, as (counter-intuitively for me) it seemed to harm the results of most models, even those that had been pre-trained with dropout. In particular, the OpenAI weights were harmed by using dropout, and making a change that benefited them (along with some of my own models) seemed the most conservative approach to take in investigating this. The next thing I wanted to look into was the training data. The exact dataset that the various GPT-2 models were trained on has never been released; all we know about it is from the paper, where they say: [W]e created a new web scrape which emphasizes document quality. To do this we only scraped web pages which have been curated/filtered by humans. Manually filtering a full web scrape would be exceptionally expensive so as a starting point, we scraped all outbound links from Reddit, a social media platform, which received at least 3 karma. This can be thought of as a heuristic indicator for whether other users found the link interesting, educational, or just funny. They called it "WebText". There is an OpenWebText that tries to replicate it, but although they tried to follow the same procedure as the original, there's no guarantee that it is all that similar. By comparison, I'd normally been training against FineWeb. While this is a general web-scraping dataset, without the "curation" provided by using only stuff that was linked from upvoted Reddit posts, it has been refined to remove any obvious junk. I had felt that it was pretty much equivalent. But what if I were wrong about that? I decided to see if I could get better models by using better data. The starting point Here's a table of all of the models I've been comparing to date. The "Test loss" column shows how well the model in question did on that held-back cross entropy loss evaluation. The "IFT epochs" column shows how many epochs of fine-tuning the model needed before its validation loss started rising, the "IFT score" the score that GPT 5.5 gave the model's responses to the test set of my Alpaca data, and the "IFT rank" the model's rank in terms of that score. The OpenAI small model is in there in bold, and I've also included the OpenAI medium model for comparison purposes. Test loss IFT epochs IFT score IFT rank OpenAI weights: medium 3.231442 2 43.75 1 JAX, overtrained one long epoch 3.324953 3 19.77 4 JAX, overtrained two normal epochs 3.326482 4 19.72 5 JAX, with MHA bias, no dropout 3.418784 4 18.69 6 JAX, no MHA bias, no dropout 3.420089 5 21.46 3 JAX, no MHA bias, with dropout 3.476802 5 13.22 15 OpenAI weights: small 3.499677 2 26.00 2 1xrtx3090-stacked-interventions 3.538161 4 13.77 14 8xa100m40-stacked-interventions-1 3.577761 4 10.76 18 Cloud FineWeb, 8x A100 40 GiB 3.673623 3 17.72 7 1xrtx3090-baseline 3.683835 4 15.74 8 8xa100m40-baseline 3.691526 3 14.19 13 Cloud FineWeb, 8x H100 80 GiB 3.724507 4 14.33 12 Cloud FineWeb, 8x A100 80 GiB 3.729900 3 11.34 17 Cloud FineWeb, 8x B200 160 GiB 3.771478 4 14.67 11 Local FineWeb train 3.943522 5 12.31 16 Local FineWeb-Edu extended train 4.134991 5 15.04 9 Local FineWeb-Edu train 4.166892 5 14.99 10 You can see that the OpenAI small model did pretty well in terms of the test loss, when you consider that it has 39M fewer weights than my models and was being tested against a dataset that differs more from its likely training data than it does from my own models'. Additionally, the specific models that did better than OpenAI's small one were all trained with JAX rather than PyTorch -- my hypothesis for that is that it's a result of the JAX ones getting better initial weights by pure chance. But the big difference was in the IFT score. In the specific run that gave the results in this table, the OpenAI small model got 26.00 -- the closest of my own models was more than 4.5 points lower, at 21.46. This difference was consistent over all of my other test runs. The GPT-2 small model was always ahead of mine. (GPT-2 medium, of course, beat GPT-2 small and all of my models, but given that it is twice the size of mine, that's not a big surprise.) Now, quite some time ago, I had tried looking into data quality as a lever to pull for model performance. At the bottom of the table, with the worst test loss of all models, you can see two models: "Local FineWeb-Edu train" "Local FineWeb-Edu extended train" These two were (as you might guess from the names) trained on the FineWeb-Edu dataset, which includes just the most "educational" data from FineWeb. They scored very badly on the test loss score. Given that the test dataset is from FineWeb, that's not a big surprise -- as I've written previously: If you train a model on Jane Austen and then evaluate against Chuck Tingle, then you're not going to get amazing results. But again, GPT-2 had the same issue, and did perfectly well on the test loss eval. On the other hand, while these FineWeb-Edu models' performance on the IFT eval wasn't stellar -- there are plenty of my other models ahead of them -- they did seem to punch above their weight. Consistently across all of the IFT evals I've done, they have scored higher than many of the others -- despite their poor loss on the test eval. Additionally: they were amongst the first models that I trained, before I'd spent time learning about how to optimise my hyperparameters and training loop. They did not use gradient clipping, they did use dropout, their batch size was just "whatever I could squeeze into the GPU", and I didn't set the learning rate to the right kind of value or schedule it over the course of the training run. So maybe a new training run on FineWeb-Edu plus my training improvements would help? And maybe some other tweaks to the training data would be worth looking into? The plan I decided to see what would happen if I trained some models with better-quality data. Specifically, I would train models with my current optimised loop and hyperparameters on four different datasets: FineWeb-Edu -- essentially the same as "Local FineWeb-Edu train" but with a better training setup. This would test the "more educational -> better" hypothesis. A 50:50 split of FineWeb and FineWeb-Edu. I've read that LLMs can be helped by having a decent amount of lower-quality data in their training loop, as it helps them to generalise. Perhaps having some FineWeb in there in addition to the FineWeb-Edu stuff would improve that test loss score while also helping the IFT test? A "curated" dataset containing 45% of its contents from FineWeb, 45% from FineWeb-Edu, and 10% from the Simple English Wikipedia. The full Wikipedia is huge, and full of obscure facts -- while the Simple English one is small and hopefully richer in useful information on a per-token basis. And conveniently, Answer.ai have made a snapshot of it available on Hugging Face Hub. Might deliberately putting a bunch of encyclopaedic data into the training set make the model better at the IFT eval (which has lots of factual questions in it, like "who wrote Pride and Prejudice")? OpenWebText. Even though I was unsure how well it matched the original WebText, given that it was there, it seemed silly to not try training something on it and see how it matched up. I would train each model on 3.2B tokens of the chosen dataset; that's the Chinchilla-optimal amount for my 163M-parameter models. If there were any interesting results, then I might consider doing overtrained models later on. I decided to be at least vaguely scientific about this, and to pre-register some predictions: The FineWeb-Edu-only model would do pretty badly on the test loss, but better than my older FineWeb-Edu models (90%). It would also punch above its weight on the IFT eval (90%). The 50:50 split: I expected it to do worse on the test eval than my JAX FineWeb-only models (70%), but better than the FineWeb-Edu one (90%). I wasn't sure about how it would do on the IFT eval, but thought it might be somewhere in between the two groups (60%). The curated dataset I had high hopes for in terms of the IFT eval -- let's say 80% chance of it being the best of all of my models. For the test loss eval, I expected it to do about as well as the 50:50 split, maybe a little bit worse (70%). I had no idea how the OpenWebText eval would do! Could be worse, could be better. Here's how things turned out. The FineWeb-Edu model I already had a dataset based on FineWeb-Edu ready to go, from when I trained those two original models. It is just the 10B-token sample of the original dataset at the time I generated it last December, formatted appropriately for my training script (details on the dataset card). I kicked off a training run with my JAX code (which I've been using for the other posts in this series): giles@poppy:~/Dev/jax-gpt2-from-scratch (main)$ XLA_PYTHON_CLIENT_MEM_FRACTION=0.95 uv run train.py full-llm-full-train-with-mha-output-bias-fineweb-edu datasets/ 2026-09-11 18:11:47.991583 Downloading dataset Fetching 4 files: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 4/4 [00:00<00:00, 1772.93it/s] Download complete: : 0.00B [00:00, ?B/s] | 0/4 [00:00<?, ?it/s] 2026-09-11 18:11:48.226273 Loading dataset into RAM Download complete: : 0.00B [00:00, ?B/s] 2026-09-11 18:16:29.507646 Creating model 2026-09-11 18:16:33.042509 Creating optimizer 2026-09-11 18:16:34.138990 Start train 0%| | 0/33165 [00:00<?, ?it/s] 2026-09-11 18:17:38.486288 Saving checkpoint 1%|▌ | 173/33165 [13:22<39:17:03, 4.29s/it, loss=6.897, tps=21,201] ...and just less than 40 hours later, I had a model: Training complete in 142,912.226 seconds 2026-09-13 09:58:26.437276 Tokens seen: 3,260,252,160 2026-09-13 09:58:26.437284 Throughput: 22,813 tokens/second 2026-09-13 09:58:26.437302 Final train loss: 3.342 2026-09-13 09:58:26.437309 Done I converted the saved JAX safetensors file from the last checkpoint into a format that would be compatible with my PyTorch eval code, and ran my smoke test: how would it complete the sentence "Every effort moves you"? Every effort moves you closer to God’s Kingdom, and even closer to Him. As we can see in That was nice and coherent -- if unusually religious! -- so that was promising. I ran the test eval: giles@perry:~/Dev/ddp-base-model-from-scratch (main)$ uv run test_loss.py datasets/ ../jax-gpt2-from-scratch/runs/full-llm-full-train-with-mha-output-bias-fineweb-edu/model.json ../jax-gpt2-from-scratch/runs/full-llm-full-train-with-mha-output-bias-fineweb-edu/checkpoints/latest/pytorch-model.safetensors Fetching 4 files: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 4/4 [00:00<00:00, 2758.50it/s] 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 3200/3200 [03:52<00:00, 13.74it/s] Loss against our test dataset: 3.632900 That was pretty good, putting it at a better test loss than all of the models I had trained without optimised hyperparameters, and worse than all of the ones I had trained on FineWeb with optimised hyperparameters. So that fit in with my prediction that it would be better than the old FineWeb-Edu models; the fact that it was also better than the non-optimised training runs with FineWeb seemed sensible enough that I felt silly for not having predicted that it would have fallen exactly there :-) I decided to leave the IFT eval until the end so that I could check all of the models from these experiments together, so it was time to upload this one to Hugging Face, and move on to the next model. 50:50 FineWeb to FineWeb-Edu I put together a new repo with a script to prepare datasets specifically for my training setup. You provide it with config that specifies some source datasets along with information about how to process them and how to mix them together, and it uploads a new dataset to Hugging Face Hub with the required characteristics. For example, for the 50:50 FineWeb to FineWeb-Edu split, the config looked like this: { "seed": 42, "tokens_desired": 10000000000, "upload_dataset_name": "gpjt/fw-fwedu-5050-gpt2-tokens", "sources": [ { "name": "FineWeb", "hf_id": "HuggingFaceFW/fineweb", "hf_name": "sample-10BT", "hf_split": "train", "item_field": "text", "weight": 50 }, { "name": "FineWeb-Edu", "hf_id": "HuggingFaceFW/fineweb-edu", "hf_name": "sample-10BT", "hf_split": "train", "item_field": "text", "weight": 50 } ] } The way the script works is pretty simple: it works out (based on those weights and the tokens_desired) how many tokens it wants from each source dataset, shuffles the items in the sources, then it loops until it has the desired number of tokens or more stored in an output. In the loop, it works out which source is currently most under-represented, grabs an item from it, tokenises it, and adds it to the output. Running it with that 50:50 config seemed to work fine: giles@perry:~/Dev/prepare-llm-training-dataset (main)$ uv run prepare-dataset.py runs/fw-fwedu-5050/ Resolving data files: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████| 27468/27468 [00:00<00:00, 89875.56it/s] Loading dataset shards: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████| 102/102 [00:00<00:00, 133.75it/s] Resolving data files: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████| 2410/2410 [00:00<00:00, 87461.48it/s] Loading dataset shards: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████| 98/98 [00:00<00:00, 200.09it/s] 2026-09-13 20:13:22.000187: Generating dataset; per-source counts 2026-09-13 20:13:22.000217: FineWeb: 5,000,000,000 2026-09-13 20:13:22.000221: FineWeb-Edu: 5,000,000,000 FineWeb: 100%|████████████████████████████████████████████████████████████████████████████████████████████████▉| 4999999705/5000000000 [1:01:33<00:00, 1353639.33token/s] FineWeb-Edu: 5000000363token [1:01:33, 1353639.47token/s] 2026-09-13 21:14:55.747239: Done generating tokens 2026-09-13 21:14:55.748480: FineWeb: 4,999,999,705 / 5,000,000,000 (1.000, 1 iterators) 2026-09-13 21:14:55.748487: FineWeb-Edu: 5,000,000,363 / 5,000,000,000 (1.000, 1 iterators) 2026-09-13 21:14:55.748489: Total: 10,000,000,068 2026-09-13 21:14:55.748491: Catting... 2026-09-13 21:16:29.565152: Catted into a tensor of shape torch.Size([10000000068]) 2026-09-13 21:16:29.566663: Saving... 2026-09-13 21:16:36.006267: Saved 2026-09-13 21:16:36.009413: Uploading to gpjt/fw-fwedu-5050-gpt2-tokens Processing Files (1 / 1) : 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████| 20.0GB / 20.0GB, 117MB/s New Data Upload : 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████| 14.6GB / 14.6GB, 98.1MB/s ...du-5050/train.safetensors: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████| 20.0GB / 20.0GB 2026-09-13 21:17:59.545875: Done So we had almost-perfect 50:50 balance between the datasets, and it saved this dataset on Hugging Face. I ran a script to double-check that it looked sane, and it did, so it was time to spin up a training run: giles@perry:~/Dev/jax-gpt2-from-scratch (main)$ XLA_PYTHON_CLIENT_MEM_FRACTION=0.90 uv run train.py full-llm-full-train-with-mha-output-bias-fw-fwedu-5050 datasets/ 2026-09-13 21:20:59.880918 Downloading dataset Fetching 2 files: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 2/2 [01:13<00:00, 36.70s/it] Download complete: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████████| 20.0G/20.0G [01:13<00:00, 1.24GB/s] 2026-09-13 21:22:13.521745 Loading dataset into RAM Download complete: 100%|██████████████████████████████████████████████████████████████████████████████████████████████████████████████| 20.0G/20.0G [01:13<00:00, 272MB/s] 2026-09-13 21:22:33.787720 Creating model 2026-09-13 21:22:35.501063 Creating optimizer 2026-09-13 21:22:36.043837 Start train 0%| | 0/33165 [00:00<?, ?it/s] 2026-09-13 21:23:11.437206 Saving checkpoint 0%| | 26/33165 [02:20<38:07:05, 4.14s/it, loss=9.308, tps=18,246] That was running on perry, my normal workstation, and I kicked it off in parallel with the "curated" model training run below on poppy my training box, but I'll keep the runs separate for the purposes of this writeup. When this had been running for an hour or so, our power went out. My guess is that having the tumble dryer running, the car charging, the kettle boiling, the electric hob switched on, and two machines doing training runs is a bit too much for our electrics... which might be a problem in the future, especially if (as planned) I make poppy a multi-GPU machine. However, as things stand, I was able to kick it off again after switching the circuit breaker back on, and things held up. Again, about 40 hours later: Training complete in 136,060.457 seconds 2026-09-15 12:05:26.432638 Tokens seen: 3,227,516,928 2026-09-15 12:05:26.432642 Throughput: 23,721 tokens/second 2026-09-15 12:05:26.432650 Final train loss: 3.793 2026-09-15 12:05:26.432653 Done (Note that the numbers reported at the end of a restarted run like this only include what happened after the restart.) I converted it to PyTorch-compatible tensors, and did the smoke test: Every effort moves you on to other options—in fact, it’s not even worth that effort. Just make Looking good! Time for the loss test: giles@perry:~/Dev/ddp-base-model-from-scratch (main)$ uv run test_loss.py datasets/ ../jax-gpt2-from-scratch/runs/full-llm-full-train-with-mha-output-bias-fw-fwedu-5050/model.json ../jax-gpt2-from-scratch/runs/full-llm-full-train-with-mha-output-bias-fw-fwedu-5050/checkpoints/latest/pytorch-model.safetensors Fetching 4 files: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 4/4 [00:00<00:00, 1192.07it/s] 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 3200/3200 [03:53<00:00, 13.72it/s] Loss against our test dataset: 3.462454 That was almost in keeping with my prediction that it would do worse than the JAX FineWeb-only models, except that it was better than the worst of those, "JAX, no MHA bias, with dropout": it was actually better than I predicted. So, a promising model. Time to upload it to Hugging Face -- and now let's move on to the next one. The "curated" dataset With my dataset-preparation script, this was easy enough to set up: { "seed": 42, "tokens_desired": 10000000000, "upload_dataset_name": "gpjt/fw-fwedu-simplewiki-gpt2-tokens", "sources": [ { "name": "FineWeb", "hf_id": "HuggingFaceFW/fineweb", "hf_name": "sample-10BT", "hf_split": "train", "item_field": "text", "weight": 45 }, { "name": "FineWeb-Edu", "hf_id": "HuggingFaceFW/fineweb-edu", "hf_name": "sample-10BT", "hf_split": "train", "item_field": "text", "weight": 45 }, { "name": "Simple English Wikipedia", "hf_id": "answerdotai/simplewiki", "hf_name": "articles", "hf_split": "train", "item_field": "md", "weight": 10 } ] } Running that worked nicely: giles@perry:~/Dev/prepare-llm-training-dataset (main)$ uv run prepare-dataset.py runs/fw-fwedu-simplewiki/ Resolving data files: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████| 27468/27468 [00:00<00:00, 90196.13it/s] Loading dataset shards: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████| 102/102 [00:00<00:00, 358.90it/s] Resolving data files: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████| 2410/2410 [00:00<00:00, 88254.11it/s] Loading dataset shards: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████| 98/98 [00:00<00:00, 589.23it/s] 2026-09-13 18:59:04.106327: Generating dataset; per-source counts 2026-09-13 18:59:04.106387: FineWeb: 4,500,000,000 2026-09-13 18:59:04.106407: FineWeb-Edu: 4,500,000,000 2026-09-13 18:59:04.106422: Simple English Wikipedia: 1,000,000,000 FineWeb: 100%|██████████████████████████████████████████████████████████████████████████████████████████████████▉| 4499997964/4500000000 [59:41<00:00, 1256362.56token/s] FineWeb-Edu: 4500000607token [59:41, 1256363.31token/s] Simple English Wikipedia: 1000002889token [59:41, 279192.58token/s] 2026-09-13 19:58:45.874744: Done generating tokens 2026-09-13 19:58:45.876043: FineWeb: 4,499,997,964 / 4,500,000,000 (1.000, 1 iterators) 2026-09-13 19:58:45.876048: FineWeb-Edu: 4,500,000,607 / 4,500,000,000 (1.000, 1 iterators) 2026-09-13 19:58:45.876052: Simple English Wikipedia: 1,000,002,889 / 1,000,000,000 (1.000, 6 iterators) 2026-09-13 19:58:45.876054: Total: 10,000,001,460 2026-09-13 19:58:45.876056: Catting... 2026-09-13 20:00:18.811748: Catted into a tensor of shape torch.Size([10000001460]) 2026-09-13 20:00:18.813169: Saving... 2026-09-13 20:00:22.773873: Saved 2026-09-13 20:00:22.773936: Uploading to gpjt/fw-fwedu-simplewiki-gpt2-tokens Processing Files (1 / 1) : 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████| 20.0GB / 20.0GB, 143MB/s New Data Upload : 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████| 19.8GB / 19.8GB, 142MB/s ...plewiki/train.safetensors: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████| 20.0GB / 20.0GB 2026-09-13 20:01:59.270021: Done One thing that is worth noting in that output is the "6 iterators" for the Simple English Wikipedia. If a source dataset runs out of items while we're building up the results in this script, we start iterating over it again (with a different seed for the shuffle so that the ordering is different). The "6 iterators" means that it needed to do that 6 times -- the original creation of the iterator at the start of the script, and five more. So that means that the Simple English Wikipedia is repeated (oversampled) somewhere between five and six times in the dataset. That's not a bad thing! From what I've read, it's actually quite standard to oversample highly educational content in LLM training datasets. And anyway, the dataset the script generated was 10B tokens, of which we're only using 3.2B for the training run in this post, so it would only appear somewhere between one and two times. The repetition would likely only really cut in if and when we did an overtrained model on the dataset. Anyway, I ran my check against the uploaded dataset -- the first few items were clearly from FineWeb, FineWeb-Edu, and the Simple English Wikipedia. It was time to kick off a training run: giles@poppy:~/Dev/jax-gpt2-from-scratch (main)$ XLA_PYTHON_CLIENT_MEM_FRACTION=0.95 uv run train.py full-llm-full-train-with-mha-output-bias-fw-fwedu-simplewiki datasets/ 2026-09-13 20:24:48.037024 Downloading dataset Downloading (incomplete total...): 0.00B [00:00, ?B/s] Warning: You are sending unauthenticated requests to the HF Hub. Please set a HF_TOKEN to enable higher rate limits and faster downloads. | 0/2 [00:00<?, ?it/s] WARNING:huggingface_hub.utils._http:Warning: You are sending unauthenticated requests to the HF Hub. Please set a HF_TOKEN to enable higher rate limits and faster downloads. Fetching 2 files: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 2/2 [02:51<00:00, 85.85s/it] Download complete: 100%|██████████████████████████████████████████████████████████████████████████████████████████████████████████████| 20.0G/20.0G [02:51<00:00, 435MB/s] 2026-09-13 20:27:39.934884 Loading dataset into RAM Download complete: 100%|██████████████████████████████████████████████████████████████████████████████████████████████████████████████| 20.0G/20.0G [02:51<00:00, 116MB/s] 2026-09-13 20:31:20.492877 Creating model 2026-09-13 20:31:24.054143 Creating optimizer 2026-09-13 20:31:25.100832 Start train 0%| | 0/33165 [00:00<?, ?it/s] 2026-09-13 20:32:29.650379 Saving checkpoint 0%|▎ | 107/33165 [08:38<39:05:39, 4.26s/it, loss=7.631, tps=20,293] Again, this was interrupted by the power outage that hit the 50:50 training run, but I was able to restart from a checkpoint. After another 22 hours, it crashed with an error that I've seen before: jax.errors.JaxRuntimeError: INTERNAL: CUDA error: Failed to end stream capture: CUDA_ERROR_STREAM_CAPTURE_INVALIDATED: operation failed due to a previous error during capture [executable_name='jit_train_step'] I put it aside as a one-off oddity when I hit it last time, but this time I dug in a bit more. I noted that it had not ever happened on perry, but seemed to be an issue on poppy, and that poppy had an older version of CUDA and the Nvidia drivers -- might that be the cause? I decided to upgrade those before kicking off the next run, but for now just restarted the run from the most recent checkpoint. (Note for anyone who is hitting the same error: it has not occurred since the upgrade, so that's worth trying.) This time it completed OK: Training complete in 59,564.515 seconds 2026-09-15 15:56:52.909888 Tokens seen: 1,367,212,032 2026-09-15 15:56:52.909894 Throughput: 22,953 tokens/second 2026-09-15 15:56:52.909912 Final train loss: 3.332 2026-09-15 15:56:52.909959 Done Again, these numbers just show what happened after the most recent restart. I copied it over to perry, converted it into a format that was compatible with my PyTorch code, and ran the smoke test: Every effort moves you by the air, for it will make you a better athlete, so your body becomes bigger and stronger Coherent enough -- time for the loss eval: giles@perry:~/Dev/ddp-base-model-from-scratch (main)$ uv run test_loss.py datasets/ ~/Dev/jax-gpt2-from-scratch/runs/full-llm-full-train-with-mha-output-bias-fw-fwedu-simplewiki/model.json ~/Dev/jax-gpt2-from-scratch/runs/full-llm-full-train-with-mha-output-bias-fw-fwedu-simplewiki/checkpoints/latest/pytorch-model.safetensors Fetching 4 files: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 4/4 [00:00<00:00, 1007.64it/s] 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 3200/3200 [03:57<00:00, 13.48it/s] Loss against our test dataset: 3.542460 Again, in line with my predictions -- worse than the JAX FineWeb-only models, and indeed than the very best PyTorch one, 1xrtx3090-stacked-interventions, and also worse than the 50:50 split, but better than the FineWeb-Edu one. I uploaded it to Hugging Face, and it was time to move on to what was meant to be the final model for this set of experiments. The OpenWebText run Again, this was a simple enough config to set up: { "seed": 42, "tokens_desired": 10000000000, "upload_dataset_name": "gpjt/openwebtext-gpt2-tokens", "sources": [ { "name": "OpenWebText", "hf_id": "Skylion007/openwebtext", "hf_name": "plain_text", "hf_split": "train", "item_field": "text", "weight": 50 } ] } ...and the build and upload process worked well (and took much less time -- for some reason, sampling randomly from a single dataset is faster than sampling from two or three): giles@perry:~/Dev/prepare-llm-training-dataset (main)$ uv run prepare-dataset.py runs/openwebtext/ Resolving data files: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████| 80/80 [00:00<00:00, 32723.26it/s] Resolving data files: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████| 80/80 [00:00<00:00, 97940.55it/s] Loading dataset shards: 100%|██████████████████████████████████████████████████████████████████████████████████████████████████████████| 80/80 [00:00<00:00, 1200.13it/s] 2026-09-15 13:16:47.622617: Generating dataset; per-source counts 2026-09-15 13:16:47.622645: OpenWebText: 10,000,000,000 Resolving data files: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████| 80/80 [00:00<00:00, 45602.65it/s] Resolving data files: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████| 80/80 [00:00<00:00, 67650.06it/s] Loading dataset shards: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████| 80/80 [00:00<00:00, 307.11it/s] OpenWebText: 10000000024token [31:46, 5246208.64token/s] 2026-09-15 13:48:33.761350: Done generating tokens 2026-09-15 13:48:33.762021: OpenWebText: 10,000,000,024 / 10,000,000,000 (1.000, 2 iterators) 2026-09-15 13:48:33.762026: Total: 10,000,000,024 2026-09-15 13:48:33.762028: Catting... 2026-09-15 13:49:33.115508: Catted into a tensor of shape torch.Size([10000000024]) 2026-09-15 13:49:33.115923: Saving... 2026-09-15 13:49:36.365978: Saved 2026-09-15 13:49:36.366027: Uploading to gpjt/openwebtext-gpt2-tokens Processing Files (0 / 1) : 100%|██████████████████████████████████████████████████████████████████████████████████████████████████████▉| 20.0GB / 20.0GB, 147MB/s New Data Upload : 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████| 19.9GB / 19.9GB, 147MB/s ...webtext/train.safetensors: 100%|██████████████████████████████████████████████████████████████████████████████████████████████████████▉| 20.0GB / 20.0GB 2026-09-15 13:51:16.202890: Done Note that it needed to oversample -- that "2 iterators". OpenWebText is about 40 GiB uncompressed, and so that's about 10B GPT-2 tokens -- presumably just a little bit less. Again, given that I was planning to use just the first 3.2B tokens of the dataset, I didn't feel that it would matter. I ran the check script on the newly-uploaded Hugging Face dataset and all looked well, so that was all set for the training run. I upgraded poppy first with a sudo pacman -Syu to see if that helped with the weird error that I got in the previous run (which, as I said, it looks like it did), then kicked it off: giles@poppy:~/Dev/jax-gpt2-from-scratch (main)$ XLA_PYTHON_CLIENT_MEM_FRACTION=0.95 uv run train.py full-llm-full-train-with-mha-output-bias-openwebtext datasets/ 2026-09-15 16:42:32.606185 Downloading dataset Fetching 2 files: 100%|████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 2/2 [00:00<00:00, 941.38it/s] Download complete: : 0.00B [00:00, ?B/s] | 0/2 [00:00<?, ?it/s] 2026-09-15 16:42:32.879987 Loading dataset into RAM Download complete: : 0.00B [00:00, ?B/s] 2026-09-15 16:45:40.438791 Creating model 2026-09-15 16:45:43.840269 Creating optimizer 2026-09-15 16:45:44.848351 Start train 0%| | 0/33165 [00:00<?, ?it/s] 2026-09-15 16:46:50.632075 Saving checkpoint 1%|█ | 332/33165 [24:33<38:45:54, 4.25s/it, loss=6.623, tps=22,154] About 31 hours in, it crashed again, but this time it was my own dumb fault: poppy has a relatively small disk and I ran out of space. I fixed that and kicked it off again from the most recent checkpoint, and this time it completed: Training complete in 33,927.995 seconds 2026-09-17 11:25:10.835989 Tokens seen: 779,747,328 2026-09-17 11:25:10.835994 Throughput: 22,982 tokens/second 2026-09-17 11:25:10.836012 Final train loss: 3.165 2026-09-17 11:25:10.836018 Done I converted it to PyTorch for the smoke test: Every effort moves you through each phase, so it's not a complete picture. I'm sure your story was ...which looked solid, so it was time for the test loss eval: giles@perry:~/Dev/ddp-base-model-from-scratch (main)$ uv run test_loss.py datasets/ ~/Dev/jax-gpt2-from-scratch/runs/full-llm-full-train-with-mha-output-bias-openwebtext/model.json ~/Dev/jax-gpt2-from-scratch/runs/full-llm-full-train-with-mha-output-bias-openwebtext/checkpoints/latest/pytorch-model.safetensors Fetching 4 files: 100%|████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 4/4 [00:00<00:00, 674.76it/s] 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 3200/3200 [03:59<00:00, 13.37it/s] Loss against our test dataset: 4.045255 Our worst score yet in this experiment! Worse than any of my models so far, apart from the two FineWeb-Edu ones I did without optimised hyperparameters. Now, the first draft of this post went straight to the results from here, but the story wasn't quite over yet... Test set contamination GPT-6 Astra is relentless. Before I publish any of these posts, I run them past an editorial board of LLMs to look for issues. GPT-6 Astra not only checked the text, it also visited the code I'd linked to to check that out too, and spotted something problematic. It's obvious in retrospect, but my code to build the new datasets had a high risk of including the contents of the -- in theory held-back -- test set. The way that the test set was generated was that I downloaded the 10B sample of FineWeb back in December, splitting it into 99% training data and 1% "validation". That validation split was about 100M tokens, and I was only using the first 19M or so for actual validation runs during training, so I (somewhat arbitrarily) designated about 19M other tokens starting at position 50M in there as my test set. Now, my new dataset-generation code was just sampling randomly from the complete 10B sample of FineWeb. So there was nothing stopping it from pulling in data that was in that old validation split! That meant that it was quite likely that my new "curated" and "50:50" datasets contained at least some of the test set that was meant to have been held back from the models during training. On reflection, the problem was potentially even worse. FineWeb-Edu is a subset of FineWeb; my existing FineWeb-Edu dataset came from the 10B sample of the Hugging Face original, and so it also could potentially contain documents that I'd put into the test set. The first thing to do was to establish the size of the problem. I wrote a script to take in a "forbidden" dataset and split; this was assumed to be formatted as one big tensor of GPT-2 tokens, which is what all of my datasets are. It would then split it by end-of-text tokens, and generate a hash and a token count for each resulting "document". Optionally, you could restrict it to only considering a subset -- the n tokens starting at position p -- and it would then generate hashes/lengths for the documents inside that slice, or that overlapped it at the start or the end. I ran that to generate a list of hashes for the entire validation set -- the validation split of gpjt/fineweb-gpt2-tokens -- and then used a second script to check my various training sets (and the validation set itself) to see how much of a contamination problem there was. I got these results: Dataset Split Contamination with validation set gpjt/fineweb-gpt2-tokens validation 102163003 out of 102163003 tokens (100.00%) gpjt/fineweb-gpt2-tokens train 636166 out of 102163003 tokens (0.62%) gpjt/fineweb-edu-gpt2-tokens train 672189 out of 102163003 tokens (0.66%) gpjt/fw-fwedu-5050-gpt2-tokens train 49224580 out of 102163003 tokens (48.18%) gpjt/fw-fwedu-simplewiki-gpt2-tokens train 44233824 out of 102163003 tokens (43.30%) gpjt/openwebtext-gpt2-tokens train 212 out of 102163003 tokens (0.00%) So: The validation set was 100% "contaminated" with itself, which was a useful sanity check. The training set of gpjt/fineweb-gpt2-tokens had what I felt was a small level of contamination. It was interesting that there was any at all -- I think that must mean that there are some repeated documents in the original dataset, and some of them wound up with copies in both my training and validation splits. The gpjt/fineweb-edu-gpt2-tokens dataset also had what felt like a reassuringly low level of contamination. Both gpjt/fw-fwedu-5050-gpt2-tokens and gpjt/fw-fwedu-simplewiki-gpt2-tokens, however, looked problematic. In both cases, the training datasets had more than 40% of the validation/test set in them. gpjt/openwebtext-gpt2-tokens was, as you'd expect, almost completely uncontaminated. It looks like maybe one document happened to have been picked up by both the OpenWebText and the FineWeb crawls and then included in the bit of FineWeb I was using for validation. However, these numbers -- while scary, at least for the 50:50 and the curated datasets -- were not quite the ones to use. They showed how much of the full validation set showed up in the full training set; what I actually cared about was how much of the test set -- those 19M tokens starting at position 50M in the validation split -- was in the actual subset of the training datasets that I actually trained on -- the first ~3.2B of them. I re-ran the script to generate hashes for just the test set, and then re-ran the contamination-checking script, telling it just to look at the appropriate subset of the training tokens, and got this: Dataset (first 3.2B tokens only) Split Contamination with test set gpjt/fineweb-gpt2-tokens train 26557 out of 19632681 tokens (0.14%) gpjt/fineweb-edu-gpt2-tokens train 32079 out of 19632681 tokens (0.16%) gpjt/fw-fwedu-5050-gpt2-tokens train 2986889 out of 19632681 tokens (15.21%) gpjt/fw-fwedu-simplewiki-gpt2-tokens train 2682430 out of 19632681 tokens (13.66%) gpjt/openwebtext-gpt2-tokens train None It was clear that there was a problem -- certainly with gpjt/fw-fwedu-5050-gpt2-tokens and gpjt/fw-fwedu-simplewiki-gpt2-tokens. They'd seen what felt like a significant amount of the test set while training, so their results on the test loss eval were dubious at best. I decided to train those two models afresh, and see what the result was in terms of loss. If the difference was huge, I'd look into the risks of the (much smaller) contamination of gpjt/fineweb-gpt2-tokens and gpjt/fineweb-edu-gpt2-tokens. But if it was pretty small, I'd not worry about that too much. I extended the script that prepared datasets so that the config file could specify a forbidden_dataset. Any documents in the source datasets that matched forbidden ones would be excluded from the output. I then updated the config for gpjt/fw-fwedu-5050-gpt2-tokens and gpjt/fw-fwedu-simplewiki-gpt2-tokens so that the whole validation split of gpjt/fineweb-gpt2-tokens was forbidden, and re-generated them. You can see the updated datasets here and here. Running the contamination-checker script against them showed that they were clear. I then re-did the full training runs for those models; the uncontaminated version of the 50:50 split model is here, and the curated one is here. And the good news: both of them actually did very slightly better at the test loss eval than their equivalents that had been trained on the contaminated data: Model Contaminated Test loss JAX, FineWeb/FineWeb-Edu 50:50 No 3.449257 JAX, FineWeb/FineWeb-Edu 50:50 Yes 3.462454 JAX, curated No 3.534068 JAX, curated Yes 3.542460 There are a number of possibilities that come to mind; perhaps learning from the test set just doesn't happen with tiny 163M models like this, or perhaps while the contaminated models were learning, the benefit they got from that was outweighed by the data that they got instead of the test set data being in some way better for training purposes, at least in terms of the loss eval. But anyway, I felt that if the effect of seeing more than 10% of the test set data during training was so tiny, then the effect of seeing less than 0.2% -- which is what the FineWeb-Edu model in this set of training runs had, as did all of my other FineWeb-only models from previous experiments -- would be even smaller and I'd disregard it. That was excellent news! I didn't need to start all of my experiments from scratch. For the rest of this post, I will include the numbers and results for the contaminated models as well as the uncontaminated ones -- they're interesting for several reasons -- but for future posts I'll skip the contaminated ones. So -- finally! -- let's start digging into the final results. Results Firstly, I think it's worth taking a look at all of the test loss results in context. Here they are in a table, with the new models in bold: Test loss OpenAI weights: medium 3.231442 JAX, overtrained one long epoch 3.324953 JAX, overtrained two normal epochs 3.326482 JAX, with MHA bias, no dropout 3.418784 JAX, no MHA bias, no dropout 3.420089 JAX, FineWeb/FineWeb-Edu 50:50 (uncontaminated) 3.449257 JAX, FineWeb/FineWeb-Edu 50:50 (contaminated) 3.462454 JAX, no MHA bias, with dropout 3.476802 OpenAI weights: small 3.499677 JAX, curated (uncontaminated) 3.534068 1xrtx3090-stacked-interventions 3.538161 JAX, curated (contaminated) 3.542460 8xa100m40-stacked-interventions-1 3.577761 JAX, FineWeb-Edu 3.632900 Cloud FineWeb, 8x A100 40 GiB 3.673623 1xrtx3090-baseline 3.683835 8xa100m40-baseline 3.691526 Cloud FineWeb, 8x H100 80 GiB 3.724507 Cloud FineWeb, 8x A100 80 GiB 3.729900 Cloud FineWeb, 8x B200 160 GiB 3.771478 Local FineWeb train 3.943522 JAX, openwebtext 4.045255 Local FineWeb-Edu extended train 4.134991 Local FineWeb-Edu train 4.166892 I think there's something very clear here: with the new models, the more FineWeb that was in the training mix, the better the model did on this eval. I think I might have been subconsciously expecting that in the predictions I did before running these experiments, but in retrospect it's so incredibly obvious that I feel silly for not mentioning it explicitly! But that tells us something interesting. From the description in the paper, whatever OpenAI did the GPT-2 training run on, it was not like FineWeb. It was probably more similar to OpenWebText -- and yet, that model was the one that performed the worst on this test eval, so if it is more like OpenWebText, there must be some other factor involved. But moving on for now: how about the IFT test -- the one that kicked off all of this work in the first place? I generated a set of IFT responses for all of the new models, and then ran them (plus responses for all of the other models on that table above) past GPT 5.5, and found that one of my new models was getting quite close to the original GPT-2 small weights! So I did four more runs, so that I could get an average. Here are the results -- the "IFT score" is the average across all five runs of the judge, and the "IFT rank" is based on that. The "IFT epochs" was from the original result-generation script. Test loss IFT epochs IFT score IFT rank OpenAI weights: medium 3.231442 2 42.36 1 JAX, overtrained one long epoch 3.324953 3 18.67 7 JAX, overtrained two normal epochs 3.326482 4 18.71 6 JAX, with MHA bias, no dropout 3.418784 4 17.90 8 JAX, no MHA bias, no dropout 3.420089 5 20.50 4 JAX, FineWeb/FineWeb-Edu 50:50 (uncontaminated) 3.449257 4 17.69 9 JAX, FineWeb/FineWeb-Edu 50:50 (contaminated) 3.462454 4 19.30 5 JAX, no MHA bias, with dropout 3.476802 5 13.02 21 OpenAI weights: small 3.499677 2 25.19 2 JAX, curated (uncontaminated) 3.534068 4 16.63 10 1xrtx3090-stacked-interventions 3.538161 4 13.51 19 JAX, curated (contaminated) 3.542460 4 13.58 18 8xa100m40-stacked-interventions-1 3.577761 4 10.19 24 JAX, FineWeb-Edu 3.632900 4 24.56 3 Cloud FineWeb, 8x A100 40 GiB 3.673623 3 16.59 11 1xrtx3090-baseline 3.683835 4 15.15 12 8xa100m40-baseline 3.691526 3 13.64 16 Cloud FineWeb, 8x H100 80 GiB 3.724507 4 13.59 17 Cloud FineWeb, 8x A100 80 GiB 3.729900 3 10.79 23 Cloud FineWeb, 8x B200 160 GiB 3.771478 4 13.70 15 Local FineWeb train 3.943522 5 11.87 22 JAX, openwebtext 4.045255 4 13.28 20 Local FineWeb-Edu extended train 4.134991 5 14.29 14 Local FineWeb-Edu train 4.166892 5 14.69 13 If you want to see the full numbers, they're below. The number that initially surprised me, and made me decide to do multiple LLM-judge runs was the one for the "JAX, FineWeb-Edu" model. In my first run it came in at 24.35 vs the OpenAI small weights' 24.93 -- so close that I wondered if it might even beat them on a re-run. However, in the further four runs its score was consistently lower than the OpenAI model's, and the gap extended a bit in some. So, was FineWeb-Edu the clear winner here? Perhaps. If you look at the contaminated/uncontaminated pairs, something interesting pops out. For the 50:50 mix, the model trained with the contaminated dataset got 19.30, and the one trained on the uncontaminated one got 17.69 -- a difference of 1.61. For the "curated" dataset, the situation was even more interesting: uncontaminated got 16.63, while contaminated got 13.58, a delta of 3.05 points. Remember, the contamination issue is about whether or not the model saw the held-back test set during training. It was an issue for the test loss that is based on that test set, but is entirely orthogonal to the IFT test. From the IFT perspective, both contaminated and uncontaminated models in each case saw training data that was -- in theory, at least -- essentially the same in terms of quality. Indeed, the uncontaminated run saw almost the same data in the same order as the contaminated one, except that some items were omitted, and then extra ones were added to the end. The purpose of this set of experiments was to see how data quality affected the results on the IFT test set. But in the case of the curated model, something that should be unrelated to data quality changed the results by 3.05 points! If something as simple as changing which data of the same quality the model is trained with can affect the IFT score so drastically, it makes it a bit harder to be certain as to whether or not data quality really had the effect we were looking for. On the other hand, the FineWeb-Edu model came in at 24.56, which is 4.06 points better than the 20.50 that the closest other model got -- more than the 3.05 points we see in difference between the two curated dataset models. And it's worth noting that the model with 20.50 is "JAX, no MHA bias, no dropout", which has a subtly different architecture -- no bias on the output projection of the multi-head attention blocks. A better comparison might be "JAX, with MHA bias, no dropout", which got a score of 17.90, for a whacking great difference of 6.66 points. I think that without doing a very large number of training runs on different datasets with different mixes, each one created with a different seed, it would be hard to work out exactly what is in the noise here and what is not. However, that would cost a lot in terms of time. I think that the best thing here is to chalk this up as a fairly decent indication that FineWeb-Edu improves matters for the IFT eval, but far from a certainty. But it's certainly worth noting that whatever the noise is, it has a range of at least 3.05 points -- and the FineWeb-Edu model is just 0.63 points short of GPT-2 small! So there could well be something there. Of course, we don't know whether that model got (by chance) the best possible balance of FineWeb-Edu tokens, and could never win -- or whether it got a bad balance and would actually beat GPT-2 with a better one. So that's certainly worth keeping in mind. As an aside, the result for the curated dataset really surprised me. I had expected that it would be the best one, simply because it almost certainly contained more facts. I took a look at its answers to the questions -- one possibility that came to mind might be that it would get better responses to questions like "What is the chemical symbol for chlorine" or "Who wrote Pride and Prejudice" than the others, but would fail on less knowledge-based tasks. But it was terrible at fact-based questions too: Name the author of 'Pride and Prejudice'. What is the periodic symbol for chlorine? As I understand it, many real-world training runs do include (often oversampled) amounts of highly educational training data like this model's dataset did. But perhaps the models that I'm training are just too small to be able to make use of the data they gained that way -- maybe doing things this way and expecting good results is like asking six-year-old children to memorise stuff before they've learned enough to be able to make use of it 1. It's worth noting that the GPT-2 small model also failed on those factual questions. Well, anyway: I think we have some useful results here, so let's work out what that means for next steps. Conclusion The results we got in these experiments point in two interesting directions. The perfect connection between the amount of FineWeb in the training set and the result on the (FineWeb-based) test loss eval, while perfectly obvious in retrospect, really does highlight how mysterious it is that the OpenAI small weights do so well on that test. The fact that FineWeb-Edu did well on the IFT test tells us that there does seem to be value in using richer training data -- though the less-spectacular results of the 50:50 mix and the curated one weaken that a bit, as does the indicator of what the noise due to data selection from equivalently high-quality datasets might be. The OpenWebText result I think I'll ignore, given that -- while in theory it should be similar to what OpenAI trained on -- there are no guarantees, and it might differ in non-obvious ways for non-obvious reasons. I think that the right direction to take this going forward is to separate these two angles. I should chase a higher IFT score, and then once I have nailed that down, I should see what (if anything) might allow me to get the resulting model to improve its test score. But I will need to make sure that whatever dataset I use, I use various "mixes" of it -- versions created with different random seeds. In my earlier experiments with overtraining, I did find that it didn't seem to improve the IFT results -- but it did improve the test loss. So perhaps identifying the right combination of other factors to boost the IFT score, then overtraining the result, might help? Of course, my overtraining tests were with FineWeb, so the connection might not hold up as well if the starting model (as seems likely) was trained on a different dataset. Also, while working through the results here, I've come to the conclusion that the set of models I'm using is a bit confusing -- there are now different hyperparameter settings, small architectural differences (the MHA bias thing), dropout settings during the pre-training, and now datasets. I think that's OK for now; I should see this part of this series as more ideation than actually running the proper experiments. But at the end, when I have some solid hypotheses with a reasonable amount of backup, I should start from scratch: a baseline model, then staged interventions to build up to what (hopefully) will be a model as good as GPT-2 small. Anyway, I'll wrap this one up here. I think that the next lever to pull is (perhaps surprisingly) going to be weight tying. I had previously kind of disregarded that as a possibility, but while I was working on this post, something popped into my mind. The OpenAI models were originally trained with weight tying. My codebase does actually support doing it -- but because I got the OpenAI weights I'm using from the code in "Build a Large Language Model (from Scratch)", when I'm running the IFT test, the weights are not actually tied! We load up a model that has separate but identical embedding and output head matrices, and then we fine-tune that. So those two matrices can vary independently during fine-tuning -- to put it another way, while GPT-2 small was pre-trained with 124M parameters, the IFT test is being done on a 163M-parameter version. Does that give them some non-obvious advantage? And would adding weight-tying to my own models help, either with or without the output heads being independent at fine-tuning time? Stay tuned :-) Appendix: all IFT judge runs Here are the numbers for all of the IFT judge runs, included for completeness. You can see that the LLM judge ranks models very consistently between runs, but there is variation -- that is, on some runs it's in what I think of as a "better mood" than others, and if that's the case, it will give better scores -- but it will give them almost consistently between models, so all of the models do better. Note that (unlike the table above) this one is sorted by the average IFT score rather than the test loss. Model Run 1 Run 2 Run 3 Run 4 Run 5 Average OpenAI weights: medium 42.24 42.16 42.95 41.83 42.61 42.36 OpenAI weights: small 24.93 24.96 25.39 25.01 25.66 25.19 JAX, FineWeb-Edu 24.35 24.55 24.3 24.68 24.9 24.56 JAX, no MHA bias, no dropout 20.5 19.9 20.76 21.25 20.07 20.50 JAX, FineWeb/FineWeb-Edu 50:50 (contaminated) 19.16 18.86 19.61 19.17 19.7 19.30 JAX, overtrained two normal epochs 18.47 18.29 19.17 18.69 18.91 18.71 JAX, overtrained one long epoch 18.04 18.71 19.62 18.41 18.57 18.67 JAX, with MHA bias, no dropout 17.49 17.35 18.33 17.73 18.62 17.90 JAX, FineWeb/FineWeb-Edu 50:50 (uncontaminated) 17.37 17.73 17.53 18.01 17.83 17.69 JAX, curated (uncontaminated) 16.77 16.03 17.3 16.08 16.96 16.63 Cloud FineWeb, 8x A100 40 GiB 16.44 16.23 17.14 16.62 16.54 16.59 1xrtx3090-baseline 14.85 15.07 15.19 15.14 15.51 15.15 Local FineWeb-Edu train 14.37 14.23 15.08 14.79 15 14.69 Local FineWeb-Edu extended train 14.4 14.07 13.82 14.56 14.61 14.29 Cloud FineWeb, 8x B200 160 GiB 13.37 13.05 13.85 13.67 14.57 13.70 8xa100m40-baseline 13.64 13.36 13.9 13.32 13.97 13.64 Cloud FineWeb, 8x H100 80 GiB 13.45 13.32 13.6 13.51 14.07 13.59 JAX, curated (contaminated) 13.09 13.48 13.95 13.23 14.15 13.58 1xrtx3090-stacked-interventions 13.37 13.11 14.04 13.84 13.17 13.51 JAX, openwebtext 12.88 12.7 13.74 13.53 13.53 13.28 JAX, no MHA bias, with dropout 13.19 12.86 12.98 12.85 13.24 13.02 Local FineWeb train 11.75 11.75 12.21 11.46 12.19 11.87 Cloud FineWeb, 8x A100 80 GiB 10.68 10.2 11.03 10.55 11.49 10.79 8xa100m40-stacked-interventions-1 9.44 9.79 10.84 10.2 10.66 10.19 A small boy asleep on his right side, the right arm stuck out, the right hand hanging limp over the edge of the bed. Through a round grating in the side of a box a voice speaks softly. "The Nile is the longest river in Africa and the second in length of all the rivers of the globe. Although falling short of the length of the Mississippi-Missouri, the Nile is at the head of all rivers as regards the length of its basin, which extends through 35 degrees of latitude …" At breakfast the next morning, "Tommy," some one says, "do you know which is the longest river in Africa?" A shaking of the head. "But don't you remember something that begins: The Nile is the …" "The - Nile - is - the - longest - river - in - Africa - and - the - second - in - length - of - all - the - rivers - of - the - globe …" The words come rushing out. "Although - falling - short - of …" "Well now, which is the longest river in Africa?" The eyes are blank. "I don't know." "But the Nile, Tommy." "The - Nile - is - the - longest - river - in - Africa - and - second …" "Then which river is the longest, Tommy?" Tommy burst into tears. "I don't know," he howls. Brave New World, Aldous Huxley ↩

2 days ago • 1 votes
Pluralistic: Voting is to politics as shopping is to boycotts (01 Oct 2026)

Today's links Voting is to politics as shopping is to boycotts: The big P only matters if the small p is in play. Hey look at this: Delights to delectate. Object permanence: Gilberto Gil v WIPO; Censored Apple wifi hacker talk; Wells Fargo crime-spree started in 1998; Stencils "may not be reproduced"; DVD Jon v Apple DRM; Unpaid diplomatic parking tickets as index of corruption; Tortured Canadian was not a terrorist; Decarbonization at a distance. Upcoming appearances: Brighton, Virtual, South Bend, Hudson, Calgary, Winnipeg, Paris, Vancouver, Victoria, Ottawa, Kilkenny, Montreal. Recent appearances: Where I've been. Latest books: You keep readin' em, I'll keep writin' 'em. Upcoming books: Like I said, I'll keep writin' 'em. Colophon: All the rest. Voting is to politics as shopping is to boycotts (permalink) Here's a funny thing about the right to vote: it wasn't won by voting. From the Magna Carta to the US Constitution to the Emancipation Proclamation to 19th Amendment, voting rights (what you might call "Big P" Politics) were always downstream of protests, riots, petitions, mass movements, strikes and good, old fashioned community organizing (that is, "small p" politics). Which is to say, Big P politics matter, but to make them matter, we need a lot of small p politics. That means that democracy isn't something you do every couple of years with a ballot paper (though that's an important aspect of the process). Democracy is continuous. If you've ever wondered why your vote seems to accomplish so little, I think you can blame the near-abolition of small p politics by Big P politicians of every stripe. Indeed, Obama's genius was summoning up an army of door-knocking, phone-banking small p political activists and then euthanizing that organization after he won the election: https://newrepublic.com/article/140245/obamas-lost-army-inside-fall-grassroots-machine For Obama, the grassroots were useful for one thing: getting out the vote. The last thing he wanted was for millions of activated voters to turn into activists who'd flame him and harangue him and picket him if they didn't like his compromises. Boy, did Obama ever compromise. He let the bank executives who created the Great Financial Crisis off the hook and encouraged them to foreclose on the homes of millions of Americans, the very same public that had bailed them out: https://theweek.com/articles/624777/obamas-biggest-failure He shielded the CIA's torturers from scrutiny and prosecution: https://journals.law.harvard.edu/ilj/2009/04/obama-publishes-torture-memos-immunizes-cia-staff/ He reneged on his promise to shut down Gitmo: https://www.pbs.org/newshour/show/obama-failed-close-guantanamo And his promise to hold the phone companies to account for their complicity in the NSA's mass domestic surveillance: https://www.pbs.org/wgbh/frontline/article/obama-on-mass-government-surveillance-then-and-now/ He stepped up secret drone warfare: https://www.cfr.org/articles/obamas-final-drone-strike-data And unconstitutional domestic surveillance: https://www.eff.org/deeplinks/2017/01/obama-expands-surveillance-powers-his-way-out Whenever I raise this, Obama's apologists come out of the woodwork to tell me that "the president isn't the Green Lantern," and that Obama couldn't act without help from Congress and the Senate, who wouldn't back his plays. I think that Trump's presidency has shown us how much power the president really has even when the legislature won't play ball. But even if you accept the Green Lantern apologetics, the fact remains that Obama could have had a clamoring army of ardent supporters in the streets, defending his agenda against recalcitrants in his own party and wreckers in the GOP. He chose not to have that army. He sent that army home. It's like Obama heard the story about post-election FDR telling civil rights leaders, "I want to do it, now make me do it," and concluded, "I don't want to do it, so I'd better not let anyone make me do it": https://www.quora.com/Did-Franklin-Roosevelt-ever-say-I-agree-with-you-I-want-to-do-it-now-make-me-do-it Of course, Trump is doing everything he can to extinguish both small p politics and Big P Politics. It's not just his wildly illegal voter suppression tactics. He's banning and prosecuting political groups, invoking anti-terror laws (which Obama supported and promised would only be used proportionately and wisely) to chase his grassroots opposition underground: https://www.whitehouse.gov/presidential-actions/2025/09/designating-antifa-as-a-domestic-terrorist-organization/ Liberals are often contemptuous of grassroots movements (cf "basket of deplorables," "Green Lantern" scolding), but the right is terrified of them. The right's political leadership is terrified of its own grassroots, and rightly so, because those people are maniacs, and they're the reason the GOP has been pushed into its most extreme positions. The right's grassroots, meanwhile, are afraid of the left's grassroots. The last thing they want is a militant, organized, mobilized base pushing Dem politicians to take the stands that are wildly and widely popular in America, from Medicare for All to an end to ICE – the Mamdani agenda, in other words. Mamdani is the anti-Obama. He shows what happens when a progressive candidate nurtures and co-governs with their base after the election, using millions of passionate, committed, everyday people to steamroller anyone who gets in the way of his agenda: https://www.nyc.gov/content/100days/pages/ Of course the downside of this is that when Mamdani reneges on his pledges, he is loudly and furiously held to account for it: https://www.thecityreporter.nyc/2026/02/19/mamdani-budget-parks-libraries/ Mamdani understood that he would be corralled into compromises if he won the mayoralty and that when he made those compromises, his base would come after him with the unmistakable fury of betrayed idealists. He also understood that any comfort he enjoyed by sidelining his base while in office would come at a price far higher than being yelled at by his supporters: it would cost him the ability to get anything done. Voting for Mamdani was important. It got him elected. But staying organized – in unions, neighborhood clubs, affinity groups, DSA chapters and mutual aid groups – is what's letting him get stuff done, and stopping him from bailing on his promises as politically infeasible. In other words, voting only matters if it's the final stage of a sustained campaign to build and mobilize popular power. Without that, voting will get you precious little. The right's leadership understands this very well, which is why they've spent years attacking unions, community organizers like Acorn, and activist institutions like Planned Parenthood. We must defend voting rights – Big P Politics – to the bitter end, but we need to defend organizing – small p politics – just as ferociously. The reduction of politics to voting is part of the 50 year neoliberal project whose foremost goal is to make you think of yourself as an atomized individual and not as a member of a polity. Turning "politics" into "voting" is absolutely in line with Margaret Thatcher's dictum that "there is no such thing as society." It's the same move that convinced workers that the answer to bad working conditions is looking your boss in the eye and threatening to change jobs (not forming a union and striking). It's also the same move that transformed "boycotts" into "shopping." Boycotts are a collective enterprise. Before a boycott takes place, small-p political groups hold meetings, organize alternatives and communicate their demands. During a boycott, organizers work to insulate participants from reprisals, like the Montgomery Bus Boycott organizers who reasoned and remonstrated with employers who disciplined workers whose participation made them late for work. And yes, as part of a boycott, you make some consumption choices. You buy X instead of Y. But "shopping" by itself isn't a boycott. You can't "vote with your wallet" (especially not when billionaires get to vote against you with their wallets): https://pluralistic.net/2025/09/13/consumption-choices/#marginal-benefits Shopping isn't politics, and while voting is Politics (Big P), it's also not politics (small p). A boycott, on the other hand, is politics. What's more, "shopping" has the same relationship to "boycotts" that "voting" has to "politics." It's a step you take, after you've laid a lot of groundwork with other people, as part of a mass movement. I understand why shopping and voting are more attractive than boycotts and politics. Meetings suck. Hell is other people: https://locusmag.com/feature/commentary-cory-doctorow-hell-is-other-people/ But changing the system requires systemic work. Hell is other people because other people are great but it's so hard to get them to do things your way. That takes time and understanding and togetherness and arguing and forgiving. Not everyone has time or capacity for that, and at any given time, we don't all have to be doing that work. We can take turns, spelling each other off at times in our lives when we have more or less slack. But lots of us have to be in the fight, or all of us will get screwed. There aren't enough of us doing politics right now. We can tell, because our politicians are so contemptuous of the grassroots that they will sell us out without a moment's hesitation, smugly certain that they will face no consequences for doing so: https://pluralistic.net/2026/09/22/happy-chudmas/#baloney-in-our-slacks Oligarchs have it easy. Where we have to convince people to fight, they can pay or threaten people to bring them into line. But oligarchs' power is wearing thin. The data-center uprising shows how much fury there is out there, looking for a productive outlet: https://www.bloodinthemachine.com/p/with-the-backlash-to-data-centers Data centers are very bad and very visible, so they make for good targets. But data centers are only the physical extrusion of a vast, brutal, extractive system. The most important way to fight data centers is to take everyone you meet protesting one and organize with them to scare the shit out of "your" politicians so they don't dare compromise on anything. Hey look at this (permalink) The Facebook Fake-out https://www.anildash.com/2026/09/29/facebook-fake-out/ Anatomy Unzipped: John of Arderne’s Sweden Scroll (ca. 1425–35) https://publicdomainreview.org/collection/arderne-scroll/ Inside McDonald’s push to have AI price your Big Mac https://www.reuters.com/business/inside-mcdonalds-push-have-ai-price-your-big-mac-2026-09-29/ what is going on with ceiling fans https://mcmansionhell.com/post/829127919552151552/what-is-going-on-with-ceiling-fans From Shitpost to Bullshit https://www.unpopularfront.news/p/from-shitpost-to-bullshit Object permanence (permalink) #25yrsago GWB's press secretary to media: "watch what you do, watch what you say" https://web.archive.org/web/20010926223602/https://www.whitehouse.gov/news/releases/2001/09/20010926-5.html#BillMaher-Comments#BillMaher-Comments #20yrsago Stencils kit “may not be reproduced in any form” https://web.archive.org/web/20061022000842/http://www.fairuseday.com/index.php/2006/10/01/copyright-is-broken/ #20yrsago DVD Jon selling Apple DRM to Apple’s competitors https://web.archive.org/web/20061004191106/https://featured.gigaom.com/2006/10/02/dvd-jon-fairplays-apple/ #20yrsago Unpaid diplomatic parking tickets as index of national corruption https://web.archive.org/web/20130719065306/https://www.theatlantic.com/magazine/archive/2006/10/primary-sources/305203/ #20yrsago Canadian deported to Syria for torture is cleared https://www.theguardian.com/world/2006/oct/02/worlddispatch #20yrsago Gilberto Gil slams WIPO https://fromgeneva.blogspot.com/2006/09/wipo-general-assembly-impressions-from.html #20yrsago Speech given by censored Apple WiFi hacker at ToorCon https://craphound.com/cache_toorcon_2006.txt #10yrsago Company suspected of blame in Office of Personnel Management breach will help run new clearance agency https://www.reuters.com/article/us-usa-security-background-idUSKCN1202M6/ #10yrsago Wells Fargo started demanding fraud of its employees in 1998; Illinois cuts Wells off from state business https://www.citizen.org/wp-content/uploads/wells-fargo-king-of-cross-sell.pdf #10yrsago Google: if you support Amazon’s Echo, you’re cut off from Google Home and Chromecast https://variety.com/2016/digital/news/google-home-amazon-echo-chromecast-1201874125/ #5yrsago How the IMF loan-sharks the global south https://pluralistic.net/2021/10/02/debt-trap/#global-arm-breakers #1yrago Decarbonization at a distance https://pluralistic.net/2025/10/02/there-goes-the-sun/#carbon-shifting Upcoming appearances (permalink) https://www.epl.ca/blogs/post/elbows-up-with-cory-doctorow/ Brighton: Digital Sovereignty and the Post-American Internet (Green Party Conference), Oct 3 https://www.openrightsgroup.org/events/digital-sovereignty-and-the-post-american-internet/ Virtual: How to govern technology in a multipolar digital world (Connecting Current), Oct 6 https://connectingcurrent.tech/how-to-govern-technology-a-multipolar-digital-world/ South Bend: An Evening With Cory Doctorow (Notre Dame), Oct 6 https://franco.nd.edu/events/2026/10/06/an-evening-with-cory-doctorow/ Hudson, OH: Hudson Library, Oct 7 https://engagedpatrons.org/EventsExtended.cfm?SiteID=3850&amp;EventID=596952&amp;PK= Calgary: Wordfest, Oct 8 https://wordfest.com/2026/show/wordfest-presents-cory-doctorow-2026/ Winnipeg: McNally Robinson, Oct 9 https://www.mcnallyrobinson.com/event-18991/An-Evening-with-Cory-Doctorow Paris: Slow Tech Summit, Oct 15 https://slowtechsummit.com/ Vancouver: Read, Resist, Repair, Rejoice (Vancouver Writers Festival), Oct 19 https://writersfest.bc.ca/festival-event-2026/01 Victoria: Munro's Books, Oct 20 https://www.munrobooks.com/events/6113620261020 Vancouver: Life After AI (Vancouver Writers Festival), Oct 22 https://writersfest.bc.ca/festival-event-2026/46 Ottawa: Life After AI (Ottawa Writers Festival), Oct 24 https://writersfestival.org/event/life-after-ai Kilkenny (Kilkenomics), Nov 6-8 https://kilkenomics.com/ Vancouver: Enshittification (Sid Williams Theatre Society), Nov 10 https://www.sidwilliamstheatre.com/events/cory-doctorow-talks-enshittification/ Vancouver: BC Policy Solutions Gala, Nov 12 https://bcpolicy.ca/gala/ Montreal: World Science Fiction Convention, Sep 2-6 https://montreal2027.ca/en Recent appearances (permalink) Terms of Service with Clare Duffy (CNN) https://www.cnn.com/audio/podcasts/terms-of-service-with-clare-duffy/episodes/458ce968-af5d-11f0-b539-13ed2afe25f8 AI, Work, and Power (Software Engineering Daily) AI, Work, and Power https://softwareengineeringdaily.com/podcasts/cory-doctorow-on-ai-work-and-power/ AI, Corporate Power, and the Fight for Worker Control (Plutopia) https://plutopia.io/cory-doctorow-ai-corporate-power-and-the-fight-for-worker-control/ How to Think About AI—Before It’s Too Late (Daniel Solove) https://www.youtube.com/watch?v=_0xR3uEgGcc Could Tech Bosses Destroy Life As We Know It? (Politics JOE) https://www.youtube.com/watch?v=PL4VktU0SgY Latest books (permalink) "The Reverse-Centaur's Guide to AI," a short book about being a better AI critic, Farrar, Straus and Giroux, June 2026 https://us.macmillan.com/books/9780374621568/thereversecentaursguidetolifeafterai/ "Canny Valley": A limited edition collection of the collages I create for Pluralistic, self-published, September 2025 https://pluralistic.net/2025/09/04/illustrious/#chairman-bruce "Enshittification: Why Everything Suddenly Got Worse and What to Do About It," Farrar, Straus, Giroux, October 7 2025 https://us.macmillan.com/books/9780374619329/enshittification/ "Picks and Shovels": a sequel to "Red Team Blues," about the heroic era of the PC, Tor Books (US), Head of Zeus (UK), February 2025 (https://us.macmillan.com/books/9781250865908/picksandshovels). "The Bezzle": a sequel to "Red Team Blues," about prison-tech and other grifts, Tor Books (US), Head of Zeus (UK), February 2024 (thebezzle.org). "The Lost Cause:" a solarpunk novel of hope in the climate emergency, Tor Books (US), Head of Zeus (UK), November 2023 (http://lost-cause.org). "The Internet Con": A nonfiction book about interoperability and Big Tech (Verso) September 2023 (http://seizethemeansofcomputation.org). Signed copies at Book Soup (https://www.booksoup.com/book/9781804291245). "Red Team Blues": "A grabby, compulsive thriller that will leave you knowing more about how the world works than you did before." Tor Books http://redteamblues.com. "Chokepoint Capitalism: How to Beat Big Tech, Tame Big Content, and Get Artists Paid, with Rebecca Giblin", on how to unrig the markets for creative labor, Beacon Press/Scribe 2022 https://chokepointcapitalism.com Upcoming books (permalink) "The Post-American Internet," a geopolitical sequel of sorts to Enshittification, Farrar, Straus and Giroux, 2027 "Unauthorized Bread": a middle-grades graphic novel adapted from my novella about refugees, toasters and DRM, FirstSecond, April 20, 2027 "Enshittification, Why Everything Suddenly Got Worse and What to Do About It" (the graphic novel), Firstsecond, 2027 "The Memex Method," Farrar, Straus, Giroux, 2027 Colophon (permalink) Today's top sources: Currently writing: “Once Is Enemy Action,” a science fiction novel about the origins of modern technofascism. Today's words: 509 (20770 total). "The Post-American Internet," a sequel to "Enshittification," about the better world the rest of us get to have now that Trump has torched America. Fourth draft completed. Submitted to editor. A Little Brother short story about DIY insulin PLANNING This work – excluding any serialized fiction – is licensed under a Creative Commons Attribution 4.0 license. That means you can use it any way you like, including commercially, provided that you attribute it to me, Cory Doctorow, and include a link to pluralistic.net. https://creativecommons.org/licenses/by/4.0/ Quotations and images are not included in this license; they are included either under a limitation or exception to copyright, or on the basis of a separate license. Please exercise caution. How to get Pluralistic: Blog (no ads, tracking, or data-collection): Pluralistic.net Newsletter (no ads, tracking, or data-collection): https://pluralistic.net/plura-list Mastodon (no ads, tracking, or data-collection): https://mamot.fr/@pluralistic Bluesky (no ads, possible tracking and data-collection): https://bsky.app/profile/doctorow.pluralistic.net Medium (no ads, paywalled): https://doctorow.medium.com/ Tumblr (mass-scale, unrestricted, third-party surveillance and advertising): https://mostlysignssomeportents.tumblr.com/tagged/pluralistic "When life gives you SARS, you make sarsaparilla" -Joey "Accordion Guy" DeVilla READ CAREFULLY: By reading this, you agree, on behalf of your employer, to release me from all obligations and waivers arising from any and all NON-NEGOTIATED agreements, licenses, terms-of-service, shrinkwrap, clickwrap, browsewrap, confidentiality, non-disclosure, non-compete and acceptable use policies ("BOGUS AGREEMENTS") that I have entered into with your employer, its partners, licensors, agents and assigns, in perpetuity, without prejudice to my ongoing rights and privileges. You further represent that you have the authority to release me from any BOGUS AGREEMENTS on behalf of your employer. ISSN: 3066-764X

2 days ago • 1 votes
A big-tent or small-tent AI safety movement?

The unstated disagreement that underpins safety debates

2 days ago • 1 votes
AI #188: Gemini Dot Argon

Is Google back?

2 days ago • 1 votes
📚 BoredReading

You seem to be enjoying this.

Join free to unlock everything.

Create free account

Already have an account? Sign in