More from Sam Altman
Here is a photo of my family. I love them more than anything. Images have power, I hope. Normally we try to be pretty private, but in this case I am sharing a photo in the hopes that it might dissuade the next person from throwing a Molotov cocktail at our house, no matter what they think about me. The first person did it last night, at 3:45 am in the morning. Thankfully it bounced off the house and no one got hurt. Words have power too. There was an incendiary article about me a few days ago. Someone said to me yesterday they thought it was coming at a time of great anxiety about AI and that it made things more dangerous for me. I brushed it aside. Now I am awake in the middle of the night and pissed, and thinking that I have underestimated the power of words and narratives. This seems like as good of a time as any to address a few things. First, what I believe. *Working towards prosperity for everyone, empowering all people, and advancing science and technology are moral obligations for me. *AI will be the most powerful tool for expanding human capability and potential that anyone has ever seen. Demand for this tool will be essentially uncapped, and people will do incredible things with it. The world deserves huge amounts of AI and we must figure out how to make it happen. *It will not all go well. The fear and anxiety about AI is justified; we are in the process of witnessing the largest change to society in a long time, and perhaps ever. We have to get safety right, which is not just about aligning a model—we urgently need a society-wide response to be resilient to new threats. This includes things like new policy to help navigate through a difficult economic transition in order to get to a much better future. *AI has to be democratized; power cannot be too concentrated. Control of the future belongs to all people and their institutions. AI needs to empower people individually, and we need to make decisions about our future and the new rules collectively. I do not think it is right that a few AI labs would make the most consequential decisions about the shape of our future. *Adaptability is critical. We are all learning about something new very quickly; some of our beliefs will be right and some will be wrong, and sometimes we will need to change our mind quickly as the technology develops and society evolves. No one understands the impacts of superintelligence yet, but they will be immense. Second, some personal reflections. As I reflect on my own work in the first decade of OpenAI, I can point to a lot of things I’m proud of and a bunch of mistakes. I was thinking about our upcoming trial with Elon and remembering how much I held the line on not being willing to agree to the unilateral control he wanted over OpenAI. I’m proud of that, and the narrow path we navigated then to allow the continued existence of OpenAI, and all the achievements that followed. I am not proud of being conflict-averse, which has caused great pain for me and OpenAI. I am not proud of handling myself badly in a conflict with our previous board that led to a huge mess for the company. I have made many other mistakes throughout the insane trajectory of OpenAI; I am a flawed person in the center of an exceptionally complex situation, trying to get a little better each year, always working for the mission. We knew going into this how huge the stakes of AI were, and that the personal disagreements between well-meaning people I cared about would be amplified greatly. But it’s another thing to live through these bitter conflicts and often to have to arbitrate them, and the costs have been serious. I am sorry to people I’ve hurt and wish I had learned more faster. I am also very aware that OpenAI is now a major platform, not a scrappy startup, and we need to operate in a more predictable way now. It has been an extremely intense, chaotic, and high-pressure few years. Mostly though, I am extremely proud that we are delivering on our mission, which seemed incredibly unlikely when we started. Against all odds, we figured out how to build very powerful AI, figured out how to amass enough capital to build the infrastructure to deliver it, figured out how to build a product company and business, figured out how to deliver reasonably safe and robust services at a massive scale, and much more. A lot of companies say they are going to change the world; we actually did. Third, some thoughts about the industry. My personal takeaway from the last several years, and take on why there has been so much Shakespearean drama between the companies in our field, comes down to this: “Once you see AGI you can’t unsee it.” It has a real "ring of power” dynamic to it, and makes people do crazy things. I don’t mean that AGI is the ring itself, but instead the totalizing philosophy of “being the one to control AGI”. The only solution I can come up with is to orient towards sharing the technology with people broadly, and for no one to have the ring. The two obvious ways to do this are individual empowerment and making sure democratic system stays in control. It is important that the democratic process remains more powerful than companies. Laws and norms are going to change, but we have to work within the democratic process, even though it will be messy and slower than we’d like. We want to be a voice and a stakeholder, but not to have all the power. A lot of the criticism of our industry comes from sincere concern about the incredibly high stakes of this technology. This is quite valid, and we welcome good-faith criticism and debate. I empathize with anti-technology sentiments and clearly technology isn’t always good for everyone. But overall, I believe technological progress can make the future unbelievably good, for your family and mine. While we have that debate, we should de-escalate the rhetoric and tactics and try to have fewer explosions in fewer homes, figuratively and literally.
AI has gotten remarkably better in recent years; ChatGPT can do amazing things that we take for granted. This is as it should be, and is the story of human progress. But behind the blinking circle, nicely abstracted away, is the greatest story of human ingenuity I have ever seen. A lot of people have worked unbelievably hard to discover how to build something that most experts thought was impossible on this timeframe, and to build a company to deliver products at massive scale to let people benefit from it. Most people who use ChatGPT will never think about the people that put so much work into it, which is totally ok, but just to take a minute of your time… There are two people I'd like to mention that OpenAI would not be OpenAI without: Jakub Pachocki and Szymon Sidor. Time and again, they combine research and engineering to solve impossible problems. They have not gotten enough public credit, but they decided to scale up RL as a baseline to see where it broke when the conventional wisdom was that it didn't scale which led to our Dota result, built much of the infrastructure that enabled a lot of our scientific discoveries, led GPT-4 pretraining, drove together with Ilya and Lukasz the initial ideas that led to the reasoning breakthrough, and have made significant progress exploring new paradigms. Jakub is our chief scientist. He once described Szymon as “indefatigable”, which is as perfect of a use of that word as I have ever heard. OpenAI has not yet thrown a problem at them they have not been able to solve; I have heard about partnerships like there is research labs of the past where two people are able to complement each other so well, but it is very special to get to watch it unfold over the years.
We are past the event horizon; the takeoff has started. Humanity is close to building digital superintelligence, and at least so far it’s much less weird than it seems like it should be. Robots are not yet walking the streets, nor are most of us talking to AI all day. People still die of disease, we still can’t easily go to space, and there is a lot about the universe we don’t understand. And yet, we have recently built systems that are smarter than people in many ways, and are able to significantly amplify the output of people using them. The least-likely part of the work is behind us; the scientific insights that got us to systems like GPT-4 and o3 were hard-won, but will take us very far. AI will contribute to the world in many ways, but the gains to quality of life from AI driving faster scientific progress and increased productivity will be enormous; the future can be vastly better than the present. Scientific progress is the biggest driver of overall progress; it’s hugely exciting to think about how much more we could have. In some big sense, ChatGPT is already more powerful than any human who has ever lived. Hundreds of millions of people rely on it every day and for increasingly important tasks; a small new capability can create a hugely positive impact; a small misalignment multiplied by hundreds of millions of people can cause a great deal of negative impact. 2025 has seen the arrival of agents that can do real cognitive work; writing computer code will never be the same. 2026 will likely see the arrival of systems that can figure out novel insights. 2027 may see the arrival of robots that can do tasks in the real world. A lot more people will be able to create software, and art. But the world wants a lot more of both, and experts will probably still be much better than novices, as long as they embrace the new tools. Generally speaking, the ability for one person to get much more done in 2030 than they could in 2020 will be a striking change, and one many people will figure out how to benefit from. In the most important ways, the 2030s may not be wildly different. People will still love their families, express their creativity, play games, and swim in lakes. But in still-very-important-ways, the 2030s are likely going to be wildly different from any time that has come before. We do not know how far beyond human-level intelligence we can go, but we are about to find out. In the 2030s, intelligence and energy—ideas, and the ability to make ideas happen—are going to become wildly abundant. These two have been the fundamental limiters on human progress for a long time; with abundant intelligence and energy (and good governance), we can theoretically have anything else. Already we live with incredible digital intelligence, and after some initial shock, most of us are pretty used to it. Very quickly we go from being amazed that AI can generate a beautifully-written paragraph to wondering when it can generate a beautifully-written novel; or from being amazed that it can make live-saving medical diagnoses to wondering when it can develop the cures; or from being amazed it can create a small computer program to wondering when it can create an entire new company. This is how the singularity goes: wonders become routine, and then table stakes. We already hear from scientists that they are two or three times more productive than they were before AI. Advanced AI is interesting for many reasons, but perhaps nothing is quite as significant as the fact that we can use it to do faster AI research. We may be able to discover new computing substrates, better algorithms, and who knows what else. If we can do a decade’s worth of research in a year, or a month, then the rate of progress will obviously be quite different. From here on, the tools we have already built will help us find further scientific insights and aid us in creating better AI systems. Of course this isn’t the same thing as an AI system completely autonomously updating its own code, but nevertheless this is a larval version of recursive self-improvement. There are other self-reinforcing loops at play. The economic value creation has started a flywheel of compounding infrastructure buildout to run these increasingly-powerful AI systems. And robots that can build other robots (and in some sense, datacenters that can build other datacenters) aren’t that far off. If we have to make the first million humanoid robots the old-fashioned way, but then they can operate the entire supply chain—digging and refining minerals, driving trucks, running factories, etc.—to build more robots, which can build more chip fabrication facilities, data centers, etc, then the rate of progress will obviously be quite different. As datacenter production gets automated, the cost of intelligence should eventually converge to near the cost of electricity. (People are often curious about how much energy a ChatGPT query uses; the average query uses about 0.34 watt-hours, about what an oven would use in a little over one second, or a high-efficiency lightbulb would use in a couple of minutes. It also uses about 0.000085 gallons of water; roughly one fifteenth of a teaspoon.) The rate of technological progress will keep accelerating, and it will continue to be the case that people are capable of adapting to almost anything. There will be very hard parts like whole classes of jobs going away, but on the other hand the world will be getting so much richer so quickly that we’ll be able to seriously entertain new policy ideas we never could before. We probably won’t adopt a new social contract all at once, but when we look back in a few decades, the gradual changes will have amounted to something big. If history is any guide, we will figure out new things to do and new things to want, and assimilate new tools quickly (job change after the industrial revolution is a good recent example). Expectations will go up, but capabilities will go up equally quickly, and we’ll all get better stuff. We will build ever-more-wonderful things for each other. People have a long-term important and curious advantage over AI: we are hard-wired to care about other people and what they think and do, and we don’t care very much about machines. A subsistence farmer from a thousand years ago would look at what many of us do and say we have fake jobs, and think that we are just playing games to entertain ourselves since we have plenty of food and unimaginable luxuries. I hope we will look at the jobs a thousand years in the future and think they are very fake jobs, and I have no doubt they will feel incredibly important and satisfying to the people doing them. The rate of new wonders being achieved will be immense. It’s hard to even imagine today what we will have discovered by 2035; maybe we will go from solving high-energy physics one year to beginning space colonization the next year; or from a major materials science breakthrough one year to true high-bandwidth brain-computer interfaces the next year. Many people will choose to live their lives in much the same way, but at least some people will probably decide to “plug in”. Looking forward, this sounds hard to wrap our heads around. But probably living through it will feel impressive but manageable. From a relativistic perspective, the singularity happens bit by bit, and the merge happens slowly. We are climbing the long arc of exponential technological progress; it always looks vertical looking forward and flat going backwards, but it’s one smooth curve. (Think back to 2020, and what it would have sounded like to have something close to AGI by 2025, versus what the last 5 years have actually been like.) There are serious challenges to confront along with the huge upsides. We do need to solve the safety issues, technically and societally, but then it’s critically important to widely distribute access to superintelligence given the economic implications. The best path forward might be something like: Solve the alignment problem, meaning that we can robustly guarantee that we get AI systems to learn and act towards what we collectively really want over the long-term (social media feeds are an example of misaligned AI; the algorithms that power those are incredible at getting you to keep scrolling and clearly understand your short-term preferences, but they do so by exploiting something in your brain that overrides your long-term preference). Then focus on making superintelligence cheap, widely available, and not too concentrated with any person, company, or country. Society is resilient, creative, and adapts quickly. If we can harness the collective will and wisdom of people, then although we’ll make plenty of mistakes and some things will go really wrong, we will learn and adapt quickly and be able to use this technology to get maximum upside and minimal downside. Giving users a lot of freedom, within broad bounds society has to decide on, seems very important. The sooner the world can start a conversation about what these broad bounds are and how we define collective alignment, the better. We (the whole industry, not just OpenAI) are building a brain for the world. It will be extremely personalized and easy for everyone to use; we will be limited by good ideas. For a long time, technical people in the startup industry have made fun of “the idea guys”; people who had an idea and were looking for a team to build it. It now looks to me like they are about to have their day in the sun. OpenAI is a lot of things now, but before anything else, we are a superintelligence research company. We have a lot of work in front of us, but most of the path in front of us is now lit, and the dark areas are receding fast. We feel extraordinarily grateful to get to do what we do. Intelligence too cheap to meter is well within grasp. This may sound crazy to say, but if we told you back in 2020 we were going to be where we are today, it probably sounded more crazy than our current predictions about 2030. May we scale smoothly, exponentially and uneventfully through superintelligence.
Our mission is to ensure that AGI (Artificial General Intelligence) benefits all of humanity. Systems that start to point to AGI* are coming into view, and so we think it’s important to understand the moment we are in. AGI is a weakly defined term, but generally speaking we mean it to be a system that can tackle increasingly complex problems, at human level, in many fields. People are tool-builders with an inherent drive to understand and create, which leads to the world getting better for all of us. Each new generation builds upon the discoveries of the generations before to create even more capable tools—electricity, the transistor, the computer, the internet, and soon AGI. Over time, in fits and starts, the steady march of human innovation has brought previously unimaginable levels of prosperity and improvements to almost every aspect of people’s lives. In some sense, AGI is just another tool in this ever-taller scaffolding of human progress we are building together. In another sense, it is the beginning of something for which it’s hard not to say “this time it’s different”; the economic growth in front of us looks astonishing, and we can now imagine a world where we cure all diseases, have much more time to enjoy with our families, and can fully realize our creative potential. In a decade, perhaps everyone on earth will be capable of accomplishing more than the most impactful person can today. We continue to see rapid progress with AI development. Here are three observations about the economics of AI: 1. The intelligence of an AI model roughly equals the log of the resources used to train and run it. These resources are chiefly training compute, data, and inference compute. It appears that you can spend arbitrary amounts of money and get continuous and predictable gains; the scaling laws that predict this are accurate over many orders of magnitude. 2. The cost to use a given level of AI falls about 10x every 12 months, and lower prices lead to much more use. You can see this in the token cost from GPT-4 in early 2023 to GPT-4o in mid-2024, where the price per token dropped about 150x in that time period. Moore’s law changed the world at 2x every 18 months; this is unbelievably stronger. 3. The socioeconomic value of linearly increasing intelligence is super-exponential in nature. A consequence of this is that we see no reason for exponentially increasing investment to stop in the near future. If these three observations continue to hold true, the impacts on society will be significant. We are now starting to roll out AI agents, which will eventually feel like virtual co-workers. Let’s imagine the case of a software engineering agent, which is an agent that we expect to be particularly important. Imagine that this agent will eventually be capable of doing most things a software engineer at a top company with a few years of experience could do, for tasks up to a couple of days long. It will not have the biggest new ideas, it will require lots of human supervision and direction, and it will be great at some things but surprisingly bad at others. Still, imagine it as a real-but-relatively-junior virtual coworker. Now imagine 1,000 of them. Or 1 million of them. Now imagine such agents in every field of knowledge work. In some ways, AI may turn out to be like the transistor economically—a big scientific discovery that scales well and that seeps into almost every corner of the economy. We don’t think much about transistors, or transistor companies, and the gains are very widely distributed. But we do expect our computers, TVs, cars, toys, and more to perform miracles. The world will not change all at once; it never does. Life will go on mostly the same in the short run, and people in 2025 will mostly spend their time in the same way they did in 2024. We will still fall in love, create families, get in fights online, hike in nature, etc. But the future will be coming at us in a way that is impossible to ignore, and the long-term changes to our society and economy will be huge. We will find new things to do, new ways to be useful to each other, and new ways to compete, but they may not look very much like the jobs of today. Agency, willfulness, and determination will likely be extremely valuable. Correctly deciding what to do and figuring out how to navigate an ever-changing world will have huge value; resilience and adaptability will be helpful skills to cultivate. AGI will be the biggest lever ever on human willfulness, and enable individual people to have more impact than ever before, not less. We expect the impact of AGI to be uneven. Although some industries will change very little, scientific progress will likely be much faster than it is today; this impact of AGI may surpass everything else. The price of many goods will eventually fall dramatically (right now, the cost of intelligence and the cost of energy constrain a lot of things), and the price of luxury goods and a few inherently limited resources like land may rise even more dramatically. Technically speaking, the road in front of us looks fairly clear. But public policy and collective opinion on how we should integrate AGI into society matter a lot; one of our reasons for launching products early and often is to give society and the technology time to co-evolve. AI will seep into all areas of the economy and society; we will expect everything to be smart. Many of us expect to need to give people more control over the technology than we have historically, including open-sourcing more, and accept that there is a balance between safety and individual empowerment that will require trade-offs. While we never want to be reckless and there will likely be some major decisions and limitations related to AGI safety that will be unpopular, directionally, as we get closer to achieving AGI, we believe that trending more towards individual empowerment is important; the other likely path we can see is AI being used by authoritarian governments to control their population through mass surveillance and loss of autonomy. Ensuring that the benefits of AGI are broadly distributed is critical. The historical impact of technological progress suggests that most of the metrics we care about (health outcomes, economic prosperity, etc.) get better on average and over the long-term, but increasing equality does not seem technologically determined and getting this right may require new ideas. In particular, it does seem like the balance of power between capital and labor could easily get messed up, and this may require early intervention. We are open to strange-sounding ideas like giving some “compute budget” to enable everyone on Earth to use a lot of AI, but we can also see a lot of ways where just relentlessly driving the cost of intelligence as low as possible has the desired effect. Anyone in 2035 should be able to marshall the intellectual capacity equivalent to everyone in 2025; everyone should have access to unlimited genius to direct however they can imagine. There is a great deal of talent right now without the resources to fully express itself, and if we change that, the resulting creative output of the world will lead to tremendous benefits for us all. Thanks especially to Josh Achiam, Boaz Barak and Aleksander Madry for reviewing drafts of this. *By using the term AGI here, we aim to communicate clearly, and we do not intend to alter or interpret the definitions and processes that define our relationship with Microsoft. We fully expect to be partnered with Microsoft for the long term. This footnote seems silly, but on the other hand we know some journalists will try to get clicks by writing something silly so here we are pre-empting the silliness…
The second birthday of ChatGPT was only a little over a month ago, and now we have transitioned into the next paradigm of models that can do complex reasoning. New years get people in a reflective mood, and I wanted to share some personal thoughts about how it has gone so far, and some of the things I’ve learned along the way. As we get closer to AGI, it feels like an important time to look at the progress of our company. There is still so much to understand, still so much we don’t know, and it’s still so early. But we know a lot more than we did when we started. We started OpenAI almost nine years ago because we believed that AGI was possible, and that it could be the most impactful technology in human history. We wanted to figure out how to build it and make it broadly beneficial; we were excited to try to make our mark on history. Our ambitions were extraordinarily high and so was our belief that the work might benefit society in an equally extraordinary way. At the time, very few people cared, and if they did, it was mostly because they thought we had no chance of success. In 2022, OpenAI was a quiet research lab working on something temporarily called “Chat With GPT-3.5”. (We are much better at research than we are at naming things.) We had been watching people use the playground feature of our API and knew that developers were really enjoying talking to the model. We thought building a demo around that experience would show people something important about the future and help us make our models better and safer. We ended up mercifully calling it ChatGPT instead, and launched it on November 30th of 2022. We always knew, abstractly, that at some point we would hit a tipping point and the AI revolution would get kicked off. But we didn’t know what the moment would be. To our surprise, it turned out to be this. The launch of ChatGPT kicked off a growth curve like nothing we have ever seen—in our company, our industry, and the world broadly. We are finally seeing some of the massive upside we have always hoped for from AI, and we can see how much more will come soon. It hasn’t been easy. The road hasn’t been smooth and the right choices haven’t been obvious. In the last two years, we had to build an entire company, almost from scratch, around this new technology. There is no way to train people for this except by doing it, and when the technology category is completely new, there is no one at all who can tell you exactly how it should be done. Building up a company at such high velocity with so little training is a messy process. It’s often two steps forward, one step back (and sometimes, one step forward and two steps back). Mistakes get corrected as you go along, but there aren’t really any handbooks or guideposts when you’re doing original work. Moving at speed in uncharted waters is an incredible experience, but it is also immensely stressful for all the players. Conflicts and misunderstanding abound. These years have been the most rewarding, fun, best, interesting, exhausting, stressful, and—particularly for the last two—unpleasant years of my life so far. The overwhelming feeling is gratitude; I know that someday I’ll be retired at our ranch watching the plants grow, a little bored, and will think back at how cool it was that I got to do the work I dreamed of since I was a little kid. I try to remember that on any given Friday, when seven things go badly wrong by 1 pm. A little over a year ago, on one particular Friday, the main thing that had gone wrong that day was that I got fired by surprise on a video call, and then right after we hung up the board published a blog post about it. I was in a hotel room in Las Vegas. It felt, to a degree that is almost impossible to explain, like a dream gone wrong. Getting fired in public with no warning kicked off a really crazy few hours, and a pretty crazy few days. The “fog of war” was the strangest part. None of us were able to get satisfactory answers about what had happened, or why. The whole event was, in my opinion, a big failure of governance by well-meaning people, myself included. Looking back, I certainly wish I had done things differently, and I’d like to believe I’m a better, more thoughtful leader today than I was a year ago. I also learned the importance of a board with diverse viewpoints and broad experience in managing a complex set of challenges. Good governance requires a lot of trust and credibility. I appreciate the way so many people worked together to build a stronger system of governance for OpenAI that enables us to pursue our mission of ensuring that AGI benefits all of humanity. My biggest takeaway is how much I have to be thankful for and how many people I owe gratitude towards: to everyone who works at OpenAI and has chosen to spend their time and effort going after this dream, to friends who helped us get through the crisis moments, to our partners and customers who supported us and entrusted us to enable their success, and to the people in my life who showed me how much they cared. [1] We all got back to the work in a more cohesive and positive way and I’m very proud of our focus since then. We have done what is easily some of our best research ever. We grew from about 100 million weekly active users to more than 300 million. Most of all, we have continued to put technology out into the world that people genuinely seem to love and that solves real problems. Nine years ago, we really had no idea what we were eventually going to become; even now, we only sort of know. AI development has taken many twists and turns and we expect more in the future. Some of the twists have been joyful; some have been hard. It’s been fun watching a steady stream of research miracles occur, and a lot of naysayers have become true believers. We’ve also seen some colleagues split off and become competitors. Teams tend to turn over as they scale, and OpenAI scales really fast. I think some of this is unavoidable—startups usually see a lot of turnover at each new major level of scale, and at OpenAI numbers go up by orders of magnitude every few months. The last two years have been like a decade at a normal company. When any company grows and evolves so fast, interests naturally diverge. And when any company in an important industry is in the lead, lots of people attack it for all sorts of reasons, especially when they are trying to compete with it. Our vision won’t change; our tactics will continue to evolve. For example, when we started we had no idea we would have to build a product company; we thought we were just going to do great research. We also had no idea we would need such a crazy amount of capital. There are new things we have to go build now that we didn’t understand a few years ago, and there will be new things in the future we can barely imagine now. We are proud of our track-record on research and deployment so far, and are committed to continuing to advance our thinking on safety and benefits sharing. We continue to believe that the best way to make an AI system safe is by iteratively and gradually releasing it into the world, giving society time to adapt and co-evolve with the technology, learning from experience, and continuing to make the technology safer. We believe in the importance of being world leaders on safety and alignment research, and in guiding that research with feedback from real world applications. We are now confident we know how to build AGI as we have traditionally understood it. We believe that, in 2025, we may see the first AI agents “join the workforce” and materially change the output of companies. We continue to believe that iteratively putting great tools in the hands of people leads to great, broadly-distributed outcomes. We are beginning to turn our aim beyond that, to superintelligence in the true sense of the word. We love our current products, but we are here for the glorious future. With superintelligence, we can do anything else. Superintelligent tools could massively accelerate scientific discovery and innovation well beyond what we are capable of doing on our own, and in turn massively increase abundance and prosperity. This sounds like science fiction right now, and somewhat crazy to even talk about it. That’s alright—we’ve been there before and we’re OK with being there again. We’re pretty confident that in the next few years, everyone will see what we see, and that the need to act with great care, while still maximizing broad benefit and empowerment, is so important. Given the possibilities of our work, OpenAI cannot be a normal company. How lucky and humbling it is to be able to play a role in this work. (Thanks to Josh Tyrangiel for sort of prompting this. I wish we had had a lot more time.) [1] There were a lot of people who did incredible and gigantic amounts of work to help OpenAI, and me personally, during those few days, but two people stood out from all others. Ron Conway and Brian Chesky went so far above and beyond the call of duty that I’m not even sure how to describe it. I’ve of course heard stories about Ron’s ability and tenaciousness for years and I’ve spent a lot of time with Brian over the past couple of years getting a huge amount of help and advice. But there’s nothing quite like being in the foxhole with people to see what they can really do. I am reasonably confident OpenAI would have fallen apart without their help; they worked around the clock for days until things were done. Although they worked unbelievably hard, they stayed calm and had clear strategic thought and great advice throughout. They stopped me from making several mistakes and made none themselves. They used their vast networks for everything needed and were able to navigate many complex situations. And I’m sure they did a lot of things I don’t know about. What I will remember most, though, is their care, compassion, and support. I thought I knew what it looked like to support a founder and a company, and in some small sense I did. But I have never before seen, or even heard of, anything like what these guys did, and now I get more fully why they have the legendary status they do. They are different and both fully deserve their genuinely unique reputations, but they are similar in their remarkable ability to move mountains and help, and in their unwavering commitment in times of need. The tech industry is far better off for having both of them in it. There are others like them; it is an amazingly special thing about our industry and does much more to make it all work than people realize. I look forward to paying it forward. On a more personal note, thanks especially to Ollie for his support that weekend and always; he is incredible in every way and no one could ask for a better partner.
More in AI
Anthropic now treats the Pope as a competitor in a battle for souls
When I finished learning how to build an LLM from scratch, I was left with a mystery: my own models were not as good as OpenAI's original GPT-2 models, despite being based on the same architecture. My models all had 163M parameters, and followed the design from Sebastian Raschka's book "Build a Large Language Model (from Scratch)". That meant that they were pretty much the same as the setup for the OpenAI GPT-2 "small" instance, except that they did not use weight-tying or bias on the QKV matrices. Weight-tying means that you re-use the initial embedding matrix as the output head at the end, and using it means that GPT-2 small saved quite a few parameters -- it was 124M rather than 163M -- at, at least in my own experiments, a cost in quality; similarly, while I found that QKV bias made a tiny improvement in loss terms, I'd felt it was likely within the noise. But GPT-2 small consistently beat my models on an instruction fine-tuning (IFT) task -- also adapted from Raschka's book. That test fine-tunes the model on a subset of the Alpaca dataset, until validation loss starts rising, and then runs a test set through the resulting model. The responses to the test set questions are stored, and then I run all of the responses from all of the models under test past GPT 5.5 in one go to get an aggregate score; more details here. GPT-2 small always did better than any of my models on this. Additionally, it did surprisingly well on a simpler eval -- one that just measured the cross entropy loss it got on a test set. It scored close to my own best models, and better than many of them. What made this result particularly interesting was that the test set in question was a split of my own training data; my models would not have seen it when training (at least, in theory), but it seems likely that it would be much more similar to their own training data than it was to OpenAI's. I've checked two things while probing this mystery: It seems very likely that the GPT-2 models were overtrained by modern standards; would overtraining my own models get them closer? It turned out that no, it probably didn't help with the IFT eval (though there might have been some signal there). It did help quite a lot with the test loss eval, though. The way I was handling dropout in the IFT test might have been unduly benefiting some models while working against others. I decided to standardise on not using dropout during this eval, as (counter-intuitively for me) it seemed to harm the results of most models, even those that had been pre-trained with dropout. In particular, the OpenAI weights were harmed by using dropout, and making a change that benefited them (along with some of my own models) seemed the most conservative approach to take in investigating this. The next thing I wanted to look into was the training data. The exact dataset that the various GPT-2 models were trained on has never been released; all we know about it is from the paper, where they say: [W]e created a new web scrape which emphasizes document quality. To do this we only scraped web pages which have been curated/filtered by humans. Manually filtering a full web scrape would be exceptionally expensive so as a starting point, we scraped all outbound links from Reddit, a social media platform, which received at least 3 karma. This can be thought of as a heuristic indicator for whether other users found the link interesting, educational, or just funny. They called it "WebText". There is an OpenWebText that tries to replicate it, but although they tried to follow the same procedure as the original, there's no guarantee that it is all that similar. By comparison, I'd normally been training against FineWeb. While this is a general web-scraping dataset, without the "curation" provided by using only stuff that was linked from upvoted Reddit posts, it has been refined to remove any obvious junk. I had felt that it was pretty much equivalent. But what if I were wrong about that? I decided to see if I could get better models by using better data. The starting point Here's a table of all of the models I've been comparing to date. The "Test loss" column shows how well the model in question did on that held-back cross entropy loss evaluation. The "IFT epochs" column shows how many epochs of fine-tuning the model needed before its validation loss started rising, the "IFT score" the score that GPT 5.5 gave the model's responses to the test set of my Alpaca data, and the "IFT rank" the model's rank in terms of that score. The OpenAI small model is in there in bold, and I've also included the OpenAI medium model for comparison purposes. Test loss IFT epochs IFT score IFT rank OpenAI weights: medium 3.231442 2 43.75 1 JAX, overtrained one long epoch 3.324953 3 19.77 4 JAX, overtrained two normal epochs 3.326482 4 19.72 5 JAX, with MHA bias, no dropout 3.418784 4 18.69 6 JAX, no MHA bias, no dropout 3.420089 5 21.46 3 JAX, no MHA bias, with dropout 3.476802 5 13.22 15 OpenAI weights: small 3.499677 2 26.00 2 1xrtx3090-stacked-interventions 3.538161 4 13.77 14 8xa100m40-stacked-interventions-1 3.577761 4 10.76 18 Cloud FineWeb, 8x A100 40 GiB 3.673623 3 17.72 7 1xrtx3090-baseline 3.683835 4 15.74 8 8xa100m40-baseline 3.691526 3 14.19 13 Cloud FineWeb, 8x H100 80 GiB 3.724507 4 14.33 12 Cloud FineWeb, 8x A100 80 GiB 3.729900 3 11.34 17 Cloud FineWeb, 8x B200 160 GiB 3.771478 4 14.67 11 Local FineWeb train 3.943522 5 12.31 16 Local FineWeb-Edu extended train 4.134991 5 15.04 9 Local FineWeb-Edu train 4.166892 5 14.99 10 You can see that the OpenAI small model did pretty well in terms of the test loss, when you consider that it has 39M fewer weights than my models and was being tested against a dataset that differs more from its likely training data than it does from my own models'. Additionally, the specific models that did better than OpenAI's small one were all trained with JAX rather than PyTorch -- my hypothesis for that is that it's a result of the JAX ones getting better initial weights by pure chance. But the big difference was in the IFT score. In the specific run that gave the results in this table, the OpenAI small model got 26.00 -- the closest of my own models was more than 4.5 points lower, at 21.46. This difference was consistent over all of my other test runs. The GPT-2 small model was always ahead of mine. (GPT-2 medium, of course, beat GPT-2 small and all of my models, but given that it is twice the size of mine, that's not a big surprise.) Now, quite some time ago, I had tried looking into data quality as a lever to pull for model performance. At the bottom of the table, with the worst test loss of all models, you can see two models: "Local FineWeb-Edu train" "Local FineWeb-Edu extended train" These two were (as you might guess from the names) trained on the FineWeb-Edu dataset, which includes just the most "educational" data from FineWeb. They scored very badly on the test loss score. Given that the test dataset is from FineWeb, that's not a big surprise -- as I've written previously: If you train a model on Jane Austen and then evaluate against Chuck Tingle, then you're not going to get amazing results. But again, GPT-2 had the same issue, and did perfectly well on the test loss eval. On the other hand, while these FineWeb-Edu models' performance on the IFT eval wasn't stellar -- there are plenty of my other models ahead of them -- they did seem to punch above their weight. Consistently across all of the IFT evals I've done, they have scored higher than many of the others -- despite their poor loss on the test eval. Additionally: they were amongst the first models that I trained, before I'd spent time learning about how to optimise my hyperparameters and training loop. They did not use gradient clipping, they did use dropout, their batch size was just "whatever I could squeeze into the GPU", and I didn't set the learning rate to the right kind of value or schedule it over the course of the training run. So maybe a new training run on FineWeb-Edu plus my training improvements would help? And maybe some other tweaks to the training data would be worth looking into? The plan I decided to see what would happen if I trained some models with better-quality data. Specifically, I would train models with my current optimised loop and hyperparameters on four different datasets: FineWeb-Edu -- essentially the same as "Local FineWeb-Edu train" but with a better training setup. This would test the "more educational -> better" hypothesis. A 50:50 split of FineWeb and FineWeb-Edu. I've read that LLMs can be helped by having a decent amount of lower-quality data in their training loop, as it helps them to generalise. Perhaps having some FineWeb in there in addition to the FineWeb-Edu stuff would improve that test loss score while also helping the IFT test? A "curated" dataset containing 45% of its contents from FineWeb, 45% from FineWeb-Edu, and 10% from the Simple English Wikipedia. The full Wikipedia is huge, and full of obscure facts -- while the Simple English one is small and hopefully richer in useful information on a per-token basis. And conveniently, Answer.ai have made a snapshot of it available on Hugging Face Hub. Might deliberately putting a bunch of encyclopaedic data into the training set make the model better at the IFT eval (which has lots of factual questions in it, like "who wrote Pride and Prejudice")? OpenWebText. Even though I was unsure how well it matched the original WebText, given that it was there, it seemed silly to not try training something on it and see how it matched up. I would train each model on 3.2B tokens of the chosen dataset; that's the Chinchilla-optimal amount for my 163M-parameter models. If there were any interesting results, then I might consider doing overtrained models later on. I decided to be at least vaguely scientific about this, and to pre-register some predictions: The FineWeb-Edu-only model would do pretty badly on the test loss, but better than my older FineWeb-Edu models (90%). It would also punch above its weight on the IFT eval (90%). The 50:50 split: I expected it to do worse on the test eval than my JAX FineWeb-only models (70%), but better than the FineWeb-Edu one (90%). I wasn't sure about how it would do on the IFT eval, but thought it might be somewhere in between the two groups (60%). The curated dataset I had high hopes for in terms of the IFT eval -- let's say 80% chance of it being the best of all of my models. For the test loss eval, I expected it to do about as well as the 50:50 split, maybe a little bit worse (70%). I had no idea how the OpenWebText eval would do! Could be worse, could be better. Here's how things turned out. The FineWeb-Edu model I already had a dataset based on FineWeb-Edu ready to go, from when I trained those two original models. It is just the 10B-token sample of the original dataset at the time I generated it last December, formatted appropriately for my training script (details on the dataset card). I kicked off a training run with my JAX code (which I've been using for the other posts in this series): giles@poppy:~/Dev/jax-gpt2-from-scratch (main)$ XLA_PYTHON_CLIENT_MEM_FRACTION=0.95 uv run train.py full-llm-full-train-with-mha-output-bias-fineweb-edu datasets/ 2026-09-11 18:11:47.991583 Downloading dataset Fetching 4 files: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 4/4 [00:00<00:00, 1772.93it/s] Download complete: : 0.00B [00:00, ?B/s] | 0/4 [00:00<?, ?it/s] 2026-09-11 18:11:48.226273 Loading dataset into RAM Download complete: : 0.00B [00:00, ?B/s] 2026-09-11 18:16:29.507646 Creating model 2026-09-11 18:16:33.042509 Creating optimizer 2026-09-11 18:16:34.138990 Start train 0%| | 0/33165 [00:00<?, ?it/s] 2026-09-11 18:17:38.486288 Saving checkpoint 1%|▌ | 173/33165 [13:22<39:17:03, 4.29s/it, loss=6.897, tps=21,201] ...and just less than 40 hours later, I had a model: Training complete in 142,912.226 seconds 2026-09-13 09:58:26.437276 Tokens seen: 3,260,252,160 2026-09-13 09:58:26.437284 Throughput: 22,813 tokens/second 2026-09-13 09:58:26.437302 Final train loss: 3.342 2026-09-13 09:58:26.437309 Done I converted the saved JAX safetensors file from the last checkpoint into a format that would be compatible with my PyTorch eval code, and ran my smoke test: how would it complete the sentence "Every effort moves you"? Every effort moves you closer to God’s Kingdom, and even closer to Him. As we can see in That was nice and coherent -- if unusually religious! -- so that was promising. I ran the test eval: giles@perry:~/Dev/ddp-base-model-from-scratch (main)$ uv run test_loss.py datasets/ ../jax-gpt2-from-scratch/runs/full-llm-full-train-with-mha-output-bias-fineweb-edu/model.json ../jax-gpt2-from-scratch/runs/full-llm-full-train-with-mha-output-bias-fineweb-edu/checkpoints/latest/pytorch-model.safetensors Fetching 4 files: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 4/4 [00:00<00:00, 2758.50it/s] 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 3200/3200 [03:52<00:00, 13.74it/s] Loss against our test dataset: 3.632900 That was pretty good, putting it at a better test loss than all of the models I had trained without optimised hyperparameters, and worse than all of the ones I had trained on FineWeb with optimised hyperparameters. So that fit in with my prediction that it would be better than the old FineWeb-Edu models; the fact that it was also better than the non-optimised training runs with FineWeb seemed sensible enough that I felt silly for not having predicted that it would have fallen exactly there :-) I decided to leave the IFT eval until the end so that I could check all of the models from these experiments together, so it was time to upload this one to Hugging Face, and move on to the next model. 50:50 FineWeb to FineWeb-Edu I put together a new repo with a script to prepare datasets specifically for my training setup. You provide it with config that specifies some source datasets along with information about how to process them and how to mix them together, and it uploads a new dataset to Hugging Face Hub with the required characteristics. For example, for the 50:50 FineWeb to FineWeb-Edu split, the config looked like this: { "seed": 42, "tokens_desired": 10000000000, "upload_dataset_name": "gpjt/fw-fwedu-5050-gpt2-tokens", "sources": [ { "name": "FineWeb", "hf_id": "HuggingFaceFW/fineweb", "hf_name": "sample-10BT", "hf_split": "train", "item_field": "text", "weight": 50 }, { "name": "FineWeb-Edu", "hf_id": "HuggingFaceFW/fineweb-edu", "hf_name": "sample-10BT", "hf_split": "train", "item_field": "text", "weight": 50 } ] } The way the script works is pretty simple: it works out (based on those weights and the tokens_desired) how many tokens it wants from each source dataset, shuffles the items in the sources, then it loops until it has the desired number of tokens or more stored in an output. In the loop, it works out which source is currently most under-represented, grabs an item from it, tokenises it, and adds it to the output. Running it with that 50:50 config seemed to work fine: giles@perry:~/Dev/prepare-llm-training-dataset (main)$ uv run prepare-dataset.py runs/fw-fwedu-5050/ Resolving data files: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████| 27468/27468 [00:00<00:00, 89875.56it/s] Loading dataset shards: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████| 102/102 [00:00<00:00, 133.75it/s] Resolving data files: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████| 2410/2410 [00:00<00:00, 87461.48it/s] Loading dataset shards: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████| 98/98 [00:00<00:00, 200.09it/s] 2026-09-13 20:13:22.000187: Generating dataset; per-source counts 2026-09-13 20:13:22.000217: FineWeb: 5,000,000,000 2026-09-13 20:13:22.000221: FineWeb-Edu: 5,000,000,000 FineWeb: 100%|████████████████████████████████████████████████████████████████████████████████████████████████▉| 4999999705/5000000000 [1:01:33<00:00, 1353639.33token/s] FineWeb-Edu: 5000000363token [1:01:33, 1353639.47token/s] 2026-09-13 21:14:55.747239: Done generating tokens 2026-09-13 21:14:55.748480: FineWeb: 4,999,999,705 / 5,000,000,000 (1.000, 1 iterators) 2026-09-13 21:14:55.748487: FineWeb-Edu: 5,000,000,363 / 5,000,000,000 (1.000, 1 iterators) 2026-09-13 21:14:55.748489: Total: 10,000,000,068 2026-09-13 21:14:55.748491: Catting... 2026-09-13 21:16:29.565152: Catted into a tensor of shape torch.Size([10000000068]) 2026-09-13 21:16:29.566663: Saving... 2026-09-13 21:16:36.006267: Saved 2026-09-13 21:16:36.009413: Uploading to gpjt/fw-fwedu-5050-gpt2-tokens Processing Files (1 / 1) : 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████| 20.0GB / 20.0GB, 117MB/s New Data Upload : 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████| 14.6GB / 14.6GB, 98.1MB/s ...du-5050/train.safetensors: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████| 20.0GB / 20.0GB 2026-09-13 21:17:59.545875: Done So we had almost-perfect 50:50 balance between the datasets, and it saved this dataset on Hugging Face. I ran a script to double-check that it looked sane, and it did, so it was time to spin up a training run: giles@perry:~/Dev/jax-gpt2-from-scratch (main)$ XLA_PYTHON_CLIENT_MEM_FRACTION=0.90 uv run train.py full-llm-full-train-with-mha-output-bias-fw-fwedu-5050 datasets/ 2026-09-13 21:20:59.880918 Downloading dataset Fetching 2 files: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 2/2 [01:13<00:00, 36.70s/it] Download complete: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████████| 20.0G/20.0G [01:13<00:00, 1.24GB/s] 2026-09-13 21:22:13.521745 Loading dataset into RAM Download complete: 100%|██████████████████████████████████████████████████████████████████████████████████████████████████████████████| 20.0G/20.0G [01:13<00:00, 272MB/s] 2026-09-13 21:22:33.787720 Creating model 2026-09-13 21:22:35.501063 Creating optimizer 2026-09-13 21:22:36.043837 Start train 0%| | 0/33165 [00:00<?, ?it/s] 2026-09-13 21:23:11.437206 Saving checkpoint 0%| | 26/33165 [02:20<38:07:05, 4.14s/it, loss=9.308, tps=18,246] That was running on perry, my normal workstation, and I kicked it off in parallel with the "curated" model training run below on poppy my training box, but I'll keep the runs separate for the purposes of this writeup. When this had been running for an hour or so, our power went out. My guess is that having the tumble dryer running, the car charging, the kettle boiling, the electric hob switched on, and two machines doing training runs is a bit too much for our electrics... which might be a problem in the future, especially if (as planned) I make poppy a multi-GPU machine. However, as things stand, I was able to kick it off again after switching the circuit breaker back on, and things held up. Again, about 40 hours later: Training complete in 136,060.457 seconds 2026-09-15 12:05:26.432638 Tokens seen: 3,227,516,928 2026-09-15 12:05:26.432642 Throughput: 23,721 tokens/second 2026-09-15 12:05:26.432650 Final train loss: 3.793 2026-09-15 12:05:26.432653 Done (Note that the numbers reported at the end of a restarted run like this only include what happened after the restart.) I converted it to PyTorch-compatible tensors, and did the smoke test: Every effort moves you on to other options—in fact, it’s not even worth that effort. Just make Looking good! Time for the loss test: giles@perry:~/Dev/ddp-base-model-from-scratch (main)$ uv run test_loss.py datasets/ ../jax-gpt2-from-scratch/runs/full-llm-full-train-with-mha-output-bias-fw-fwedu-5050/model.json ../jax-gpt2-from-scratch/runs/full-llm-full-train-with-mha-output-bias-fw-fwedu-5050/checkpoints/latest/pytorch-model.safetensors Fetching 4 files: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 4/4 [00:00<00:00, 1192.07it/s] 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 3200/3200 [03:53<00:00, 13.72it/s] Loss against our test dataset: 3.462454 That was almost in keeping with my prediction that it would do worse than the JAX FineWeb-only models, except that it was better than the worst of those, "JAX, no MHA bias, with dropout": it was actually better than I predicted. So, a promising model. Time to upload it to Hugging Face -- and now let's move on to the next one. The "curated" dataset With my dataset-preparation script, this was easy enough to set up: { "seed": 42, "tokens_desired": 10000000000, "upload_dataset_name": "gpjt/fw-fwedu-simplewiki-gpt2-tokens", "sources": [ { "name": "FineWeb", "hf_id": "HuggingFaceFW/fineweb", "hf_name": "sample-10BT", "hf_split": "train", "item_field": "text", "weight": 45 }, { "name": "FineWeb-Edu", "hf_id": "HuggingFaceFW/fineweb-edu", "hf_name": "sample-10BT", "hf_split": "train", "item_field": "text", "weight": 45 }, { "name": "Simple English Wikipedia", "hf_id": "answerdotai/simplewiki", "hf_name": "articles", "hf_split": "train", "item_field": "md", "weight": 10 } ] } Running that worked nicely: giles@perry:~/Dev/prepare-llm-training-dataset (main)$ uv run prepare-dataset.py runs/fw-fwedu-simplewiki/ Resolving data files: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████| 27468/27468 [00:00<00:00, 90196.13it/s] Loading dataset shards: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████| 102/102 [00:00<00:00, 358.90it/s] Resolving data files: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████| 2410/2410 [00:00<00:00, 88254.11it/s] Loading dataset shards: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████| 98/98 [00:00<00:00, 589.23it/s] 2026-09-13 18:59:04.106327: Generating dataset; per-source counts 2026-09-13 18:59:04.106387: FineWeb: 4,500,000,000 2026-09-13 18:59:04.106407: FineWeb-Edu: 4,500,000,000 2026-09-13 18:59:04.106422: Simple English Wikipedia: 1,000,000,000 FineWeb: 100%|██████████████████████████████████████████████████████████████████████████████████████████████████▉| 4499997964/4500000000 [59:41<00:00, 1256362.56token/s] FineWeb-Edu: 4500000607token [59:41, 1256363.31token/s] Simple English Wikipedia: 1000002889token [59:41, 279192.58token/s] 2026-09-13 19:58:45.874744: Done generating tokens 2026-09-13 19:58:45.876043: FineWeb: 4,499,997,964 / 4,500,000,000 (1.000, 1 iterators) 2026-09-13 19:58:45.876048: FineWeb-Edu: 4,500,000,607 / 4,500,000,000 (1.000, 1 iterators) 2026-09-13 19:58:45.876052: Simple English Wikipedia: 1,000,002,889 / 1,000,000,000 (1.000, 6 iterators) 2026-09-13 19:58:45.876054: Total: 10,000,001,460 2026-09-13 19:58:45.876056: Catting... 2026-09-13 20:00:18.811748: Catted into a tensor of shape torch.Size([10000001460]) 2026-09-13 20:00:18.813169: Saving... 2026-09-13 20:00:22.773873: Saved 2026-09-13 20:00:22.773936: Uploading to gpjt/fw-fwedu-simplewiki-gpt2-tokens Processing Files (1 / 1) : 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████| 20.0GB / 20.0GB, 143MB/s New Data Upload : 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████| 19.8GB / 19.8GB, 142MB/s ...plewiki/train.safetensors: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████| 20.0GB / 20.0GB 2026-09-13 20:01:59.270021: Done One thing that is worth noting in that output is the "6 iterators" for the Simple English Wikipedia. If a source dataset runs out of items while we're building up the results in this script, we start iterating over it again (with a different seed for the shuffle so that the ordering is different). The "6 iterators" means that it needed to do that 6 times -- the original creation of the iterator at the start of the script, and five more. So that means that the Simple English Wikipedia is repeated (oversampled) somewhere between five and six times in the dataset. That's not a bad thing! From what I've read, it's actually quite standard to oversample highly educational content in LLM training datasets. And anyway, the dataset the script generated was 10B tokens, of which we're only using 3.2B for the training run in this post, so it would only appear somewhere between one and two times. The repetition would likely only really cut in if and when we did an overtrained model on the dataset. Anyway, I ran my check against the uploaded dataset -- the first few items were clearly from FineWeb, FineWeb-Edu, and the Simple English Wikipedia. It was time to kick off a training run: giles@poppy:~/Dev/jax-gpt2-from-scratch (main)$ XLA_PYTHON_CLIENT_MEM_FRACTION=0.95 uv run train.py full-llm-full-train-with-mha-output-bias-fw-fwedu-simplewiki datasets/ 2026-09-13 20:24:48.037024 Downloading dataset Downloading (incomplete total...): 0.00B [00:00, ?B/s] Warning: You are sending unauthenticated requests to the HF Hub. Please set a HF_TOKEN to enable higher rate limits and faster downloads. | 0/2 [00:00<?, ?it/s] WARNING:huggingface_hub.utils._http:Warning: You are sending unauthenticated requests to the HF Hub. Please set a HF_TOKEN to enable higher rate limits and faster downloads. Fetching 2 files: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 2/2 [02:51<00:00, 85.85s/it] Download complete: 100%|██████████████████████████████████████████████████████████████████████████████████████████████████████████████| 20.0G/20.0G [02:51<00:00, 435MB/s] 2026-09-13 20:27:39.934884 Loading dataset into RAM Download complete: 100%|██████████████████████████████████████████████████████████████████████████████████████████████████████████████| 20.0G/20.0G [02:51<00:00, 116MB/s] 2026-09-13 20:31:20.492877 Creating model 2026-09-13 20:31:24.054143 Creating optimizer 2026-09-13 20:31:25.100832 Start train 0%| | 0/33165 [00:00<?, ?it/s] 2026-09-13 20:32:29.650379 Saving checkpoint 0%|▎ | 107/33165 [08:38<39:05:39, 4.26s/it, loss=7.631, tps=20,293] Again, this was interrupted by the power outage that hit the 50:50 training run, but I was able to restart from a checkpoint. After another 22 hours, it crashed with an error that I've seen before: jax.errors.JaxRuntimeError: INTERNAL: CUDA error: Failed to end stream capture: CUDA_ERROR_STREAM_CAPTURE_INVALIDATED: operation failed due to a previous error during capture [executable_name='jit_train_step'] I put it aside as a one-off oddity when I hit it last time, but this time I dug in a bit more. I noted that it had not ever happened on perry, but seemed to be an issue on poppy, and that poppy had an older version of CUDA and the Nvidia drivers -- might that be the cause? I decided to upgrade those before kicking off the next run, but for now just restarted the run from the most recent checkpoint. (Note for anyone who is hitting the same error: it has not occurred since the upgrade, so that's worth trying.) This time it completed OK: Training complete in 59,564.515 seconds 2026-09-15 15:56:52.909888 Tokens seen: 1,367,212,032 2026-09-15 15:56:52.909894 Throughput: 22,953 tokens/second 2026-09-15 15:56:52.909912 Final train loss: 3.332 2026-09-15 15:56:52.909959 Done Again, these numbers just show what happened after the most recent restart. I copied it over to perry, converted it into a format that was compatible with my PyTorch code, and ran the smoke test: Every effort moves you by the air, for it will make you a better athlete, so your body becomes bigger and stronger Coherent enough -- time for the loss eval: giles@perry:~/Dev/ddp-base-model-from-scratch (main)$ uv run test_loss.py datasets/ ~/Dev/jax-gpt2-from-scratch/runs/full-llm-full-train-with-mha-output-bias-fw-fwedu-simplewiki/model.json ~/Dev/jax-gpt2-from-scratch/runs/full-llm-full-train-with-mha-output-bias-fw-fwedu-simplewiki/checkpoints/latest/pytorch-model.safetensors Fetching 4 files: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 4/4 [00:00<00:00, 1007.64it/s] 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 3200/3200 [03:57<00:00, 13.48it/s] Loss against our test dataset: 3.542460 Again, in line with my predictions -- worse than the JAX FineWeb-only models, and indeed than the very best PyTorch one, 1xrtx3090-stacked-interventions, and also worse than the 50:50 split, but better than the FineWeb-Edu one. I uploaded it to Hugging Face, and it was time to move on to what was meant to be the final model for this set of experiments. The OpenWebText run Again, this was a simple enough config to set up: { "seed": 42, "tokens_desired": 10000000000, "upload_dataset_name": "gpjt/openwebtext-gpt2-tokens", "sources": [ { "name": "OpenWebText", "hf_id": "Skylion007/openwebtext", "hf_name": "plain_text", "hf_split": "train", "item_field": "text", "weight": 50 } ] } ...and the build and upload process worked well (and took much less time -- for some reason, sampling randomly from a single dataset is faster than sampling from two or three): giles@perry:~/Dev/prepare-llm-training-dataset (main)$ uv run prepare-dataset.py runs/openwebtext/ Resolving data files: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████| 80/80 [00:00<00:00, 32723.26it/s] Resolving data files: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████| 80/80 [00:00<00:00, 97940.55it/s] Loading dataset shards: 100%|██████████████████████████████████████████████████████████████████████████████████████████████████████████| 80/80 [00:00<00:00, 1200.13it/s] 2026-09-15 13:16:47.622617: Generating dataset; per-source counts 2026-09-15 13:16:47.622645: OpenWebText: 10,000,000,000 Resolving data files: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████| 80/80 [00:00<00:00, 45602.65it/s] Resolving data files: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████| 80/80 [00:00<00:00, 67650.06it/s] Loading dataset shards: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████| 80/80 [00:00<00:00, 307.11it/s] OpenWebText: 10000000024token [31:46, 5246208.64token/s] 2026-09-15 13:48:33.761350: Done generating tokens 2026-09-15 13:48:33.762021: OpenWebText: 10,000,000,024 / 10,000,000,000 (1.000, 2 iterators) 2026-09-15 13:48:33.762026: Total: 10,000,000,024 2026-09-15 13:48:33.762028: Catting... 2026-09-15 13:49:33.115508: Catted into a tensor of shape torch.Size([10000000024]) 2026-09-15 13:49:33.115923: Saving... 2026-09-15 13:49:36.365978: Saved 2026-09-15 13:49:36.366027: Uploading to gpjt/openwebtext-gpt2-tokens Processing Files (0 / 1) : 100%|██████████████████████████████████████████████████████████████████████████████████████████████████████▉| 20.0GB / 20.0GB, 147MB/s New Data Upload : 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████| 19.9GB / 19.9GB, 147MB/s ...webtext/train.safetensors: 100%|██████████████████████████████████████████████████████████████████████████████████████████████████████▉| 20.0GB / 20.0GB 2026-09-15 13:51:16.202890: Done Note that it needed to oversample -- that "2 iterators". OpenWebText is about 40 GiB uncompressed, and so that's about 10B GPT-2 tokens -- presumably just a little bit less. Again, given that I was planning to use just the first 3.2B tokens of the dataset, I didn't feel that it would matter. I ran the check script on the newly-uploaded Hugging Face dataset and all looked well, so that was all set for the training run. I upgraded poppy first with a sudo pacman -Syu to see if that helped with the weird error that I got in the previous run (which, as I said, it looks like it did), then kicked it off: giles@poppy:~/Dev/jax-gpt2-from-scratch (main)$ XLA_PYTHON_CLIENT_MEM_FRACTION=0.95 uv run train.py full-llm-full-train-with-mha-output-bias-openwebtext datasets/ 2026-09-15 16:42:32.606185 Downloading dataset Fetching 2 files: 100%|████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 2/2 [00:00<00:00, 941.38it/s] Download complete: : 0.00B [00:00, ?B/s] | 0/2 [00:00<?, ?it/s] 2026-09-15 16:42:32.879987 Loading dataset into RAM Download complete: : 0.00B [00:00, ?B/s] 2026-09-15 16:45:40.438791 Creating model 2026-09-15 16:45:43.840269 Creating optimizer 2026-09-15 16:45:44.848351 Start train 0%| | 0/33165 [00:00<?, ?it/s] 2026-09-15 16:46:50.632075 Saving checkpoint 1%|█ | 332/33165 [24:33<38:45:54, 4.25s/it, loss=6.623, tps=22,154] About 31 hours in, it crashed again, but this time it was my own dumb fault: poppy has a relatively small disk and I ran out of space. I fixed that and kicked it off again from the most recent checkpoint, and this time it completed: Training complete in 33,927.995 seconds 2026-09-17 11:25:10.835989 Tokens seen: 779,747,328 2026-09-17 11:25:10.835994 Throughput: 22,982 tokens/second 2026-09-17 11:25:10.836012 Final train loss: 3.165 2026-09-17 11:25:10.836018 Done I converted it to PyTorch for the smoke test: Every effort moves you through each phase, so it's not a complete picture. I'm sure your story was ...which looked solid, so it was time for the test loss eval: giles@perry:~/Dev/ddp-base-model-from-scratch (main)$ uv run test_loss.py datasets/ ~/Dev/jax-gpt2-from-scratch/runs/full-llm-full-train-with-mha-output-bias-openwebtext/model.json ~/Dev/jax-gpt2-from-scratch/runs/full-llm-full-train-with-mha-output-bias-openwebtext/checkpoints/latest/pytorch-model.safetensors Fetching 4 files: 100%|████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 4/4 [00:00<00:00, 674.76it/s] 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 3200/3200 [03:59<00:00, 13.37it/s] Loss against our test dataset: 4.045255 Our worst score yet in this experiment! Worse than any of my models so far, apart from the two FineWeb-Edu ones I did without optimised hyperparameters. Now, the first draft of this post went straight to the results from here, but the story wasn't quite over yet... Test set contamination GPT-6 Astra is relentless. Before I publish any of these posts, I run them past an editorial board of LLMs to look for issues. GPT-6 Astra not only checked the text, it also visited the code I'd linked to to check that out too, and spotted something problematic. It's obvious in retrospect, but my code to build the new datasets had a high risk of including the contents of the -- in theory held-back -- test set. The way that the test set was generated was that I downloaded the 10B sample of FineWeb back in December, splitting it into 99% training data and 1% "validation". That validation split was about 100M tokens, and I was only using the first 19M or so for actual validation runs during training, so I (somewhat arbitrarily) designated about 19M other tokens starting at position 50M in there as my test set. Now, my new dataset-generation code was just sampling randomly from the complete 10B sample of FineWeb. So there was nothing stopping it from pulling in data that was in that old validation split! That meant that it was quite likely that my new "curated" and "50:50" datasets contained at least some of the test set that was meant to have been held back from the models during training. On reflection, the problem was potentially even worse. FineWeb-Edu is a subset of FineWeb; my existing FineWeb-Edu dataset came from the 10B sample of the Hugging Face original, and so it also could potentially contain documents that I'd put into the test set. The first thing to do was to establish the size of the problem. I wrote a script to take in a "forbidden" dataset and split; this was assumed to be formatted as one big tensor of GPT-2 tokens, which is what all of my datasets are. It would then split it by end-of-text tokens, and generate a hash and a token count for each resulting "document". Optionally, you could restrict it to only considering a subset -- the n tokens starting at position p -- and it would then generate hashes/lengths for the documents inside that slice, or that overlapped it at the start or the end. I ran that to generate a list of hashes for the entire validation set -- the validation split of gpjt/fineweb-gpt2-tokens -- and then used a second script to check my various training sets (and the validation set itself) to see how much of a contamination problem there was. I got these results: Dataset Split Contamination with validation set gpjt/fineweb-gpt2-tokens validation 102163003 out of 102163003 tokens (100.00%) gpjt/fineweb-gpt2-tokens train 636166 out of 102163003 tokens (0.62%) gpjt/fineweb-edu-gpt2-tokens train 672189 out of 102163003 tokens (0.66%) gpjt/fw-fwedu-5050-gpt2-tokens train 49224580 out of 102163003 tokens (48.18%) gpjt/fw-fwedu-simplewiki-gpt2-tokens train 44233824 out of 102163003 tokens (43.30%) gpjt/openwebtext-gpt2-tokens train 212 out of 102163003 tokens (0.00%) So: The validation set was 100% "contaminated" with itself, which was a useful sanity check. The training set of gpjt/fineweb-gpt2-tokens had what I felt was a small level of contamination. It was interesting that there was any at all -- I think that must mean that there are some repeated documents in the original dataset, and some of them wound up with copies in both my training and validation splits. The gpjt/fineweb-edu-gpt2-tokens dataset also had what felt like a reassuringly low level of contamination. Both gpjt/fw-fwedu-5050-gpt2-tokens and gpjt/fw-fwedu-simplewiki-gpt2-tokens, however, looked problematic. In both cases, the training datasets had more than 40% of the validation/test set in them. gpjt/openwebtext-gpt2-tokens was, as you'd expect, almost completely uncontaminated. It looks like maybe one document happened to have been picked up by both the OpenWebText and the FineWeb crawls and then included in the bit of FineWeb I was using for validation. However, these numbers -- while scary, at least for the 50:50 and the curated datasets -- were not quite the ones to use. They showed how much of the full validation set showed up in the full training set; what I actually cared about was how much of the test set -- those 19M tokens starting at position 50M in the validation split -- was in the actual subset of the training datasets that I actually trained on -- the first ~3.2B of them. I re-ran the script to generate hashes for just the test set, and then re-ran the contamination-checking script, telling it just to look at the appropriate subset of the training tokens, and got this: Dataset (first 3.2B tokens only) Split Contamination with test set gpjt/fineweb-gpt2-tokens train 26557 out of 19632681 tokens (0.14%) gpjt/fineweb-edu-gpt2-tokens train 32079 out of 19632681 tokens (0.16%) gpjt/fw-fwedu-5050-gpt2-tokens train 2986889 out of 19632681 tokens (15.21%) gpjt/fw-fwedu-simplewiki-gpt2-tokens train 2682430 out of 19632681 tokens (13.66%) gpjt/openwebtext-gpt2-tokens train None It was clear that there was a problem -- certainly with gpjt/fw-fwedu-5050-gpt2-tokens and gpjt/fw-fwedu-simplewiki-gpt2-tokens. They'd seen what felt like a significant amount of the test set while training, so their results on the test loss eval were dubious at best. I decided to train those two models afresh, and see what the result was in terms of loss. If the difference was huge, I'd look into the risks of the (much smaller) contamination of gpjt/fineweb-gpt2-tokens and gpjt/fineweb-edu-gpt2-tokens. But if it was pretty small, I'd not worry about that too much. I extended the script that prepared datasets so that the config file could specify a forbidden_dataset. Any documents in the source datasets that matched forbidden ones would be excluded from the output. I then updated the config for gpjt/fw-fwedu-5050-gpt2-tokens and gpjt/fw-fwedu-simplewiki-gpt2-tokens so that the whole validation split of gpjt/fineweb-gpt2-tokens was forbidden, and re-generated them. You can see the updated datasets here and here. Running the contamination-checker script against them showed that they were clear. I then re-did the full training runs for those models; the uncontaminated version of the 50:50 split model is here, and the curated one is here. And the good news: both of them actually did very slightly better at the test loss eval than their equivalents that had been trained on the contaminated data: Model Contaminated Test loss JAX, FineWeb/FineWeb-Edu 50:50 No 3.449257 JAX, FineWeb/FineWeb-Edu 50:50 Yes 3.462454 JAX, curated No 3.534068 JAX, curated Yes 3.542460 There are a number of possibilities that come to mind; perhaps learning from the test set just doesn't happen with tiny 163M models like this, or perhaps while the contaminated models were learning, the benefit they got from that was outweighed by the data that they got instead of the test set data being in some way better for training purposes, at least in terms of the loss eval. But anyway, I felt that if the effect of seeing more than 10% of the test set data during training was so tiny, then the effect of seeing less than 0.2% -- which is what the FineWeb-Edu model in this set of training runs had, as did all of my other FineWeb-only models from previous experiments -- would be even smaller and I'd disregard it. That was excellent news! I didn't need to start all of my experiments from scratch. For the rest of this post, I will include the numbers and results for the contaminated models as well as the uncontaminated ones -- they're interesting for several reasons -- but for future posts I'll skip the contaminated ones. So -- finally! -- let's start digging into the final results. Results Firstly, I think it's worth taking a look at all of the test loss results in context. Here they are in a table, with the new models in bold: Test loss OpenAI weights: medium 3.231442 JAX, overtrained one long epoch 3.324953 JAX, overtrained two normal epochs 3.326482 JAX, with MHA bias, no dropout 3.418784 JAX, no MHA bias, no dropout 3.420089 JAX, FineWeb/FineWeb-Edu 50:50 (uncontaminated) 3.449257 JAX, FineWeb/FineWeb-Edu 50:50 (contaminated) 3.462454 JAX, no MHA bias, with dropout 3.476802 OpenAI weights: small 3.499677 JAX, curated (uncontaminated) 3.534068 1xrtx3090-stacked-interventions 3.538161 JAX, curated (contaminated) 3.542460 8xa100m40-stacked-interventions-1 3.577761 JAX, FineWeb-Edu 3.632900 Cloud FineWeb, 8x A100 40 GiB 3.673623 1xrtx3090-baseline 3.683835 8xa100m40-baseline 3.691526 Cloud FineWeb, 8x H100 80 GiB 3.724507 Cloud FineWeb, 8x A100 80 GiB 3.729900 Cloud FineWeb, 8x B200 160 GiB 3.771478 Local FineWeb train 3.943522 JAX, openwebtext 4.045255 Local FineWeb-Edu extended train 4.134991 Local FineWeb-Edu train 4.166892 I think there's something very clear here: with the new models, the more FineWeb that was in the training mix, the better the model did on this eval. I think I might have been subconsciously expecting that in the predictions I did before running these experiments, but in retrospect it's so incredibly obvious that I feel silly for not mentioning it explicitly! But that tells us something interesting. From the description in the paper, whatever OpenAI did the GPT-2 training run on, it was not like FineWeb. It was probably more similar to OpenWebText -- and yet, that model was the one that performed the worst on this test eval, so if it is more like OpenWebText, there must be some other factor involved. But moving on for now: how about the IFT test -- the one that kicked off all of this work in the first place? I generated a set of IFT responses for all of the new models, and then ran them (plus responses for all of the other models on that table above) past GPT 5.5, and found that one of my new models was getting quite close to the original GPT-2 small weights! So I did four more runs, so that I could get an average. Here are the results -- the "IFT score" is the average across all five runs of the judge, and the "IFT rank" is based on that. The "IFT epochs" was from the original result-generation script. Test loss IFT epochs IFT score IFT rank OpenAI weights: medium 3.231442 2 42.36 1 JAX, overtrained one long epoch 3.324953 3 18.67 7 JAX, overtrained two normal epochs 3.326482 4 18.71 6 JAX, with MHA bias, no dropout 3.418784 4 17.90 8 JAX, no MHA bias, no dropout 3.420089 5 20.50 4 JAX, FineWeb/FineWeb-Edu 50:50 (uncontaminated) 3.449257 4 17.69 9 JAX, FineWeb/FineWeb-Edu 50:50 (contaminated) 3.462454 4 19.30 5 JAX, no MHA bias, with dropout 3.476802 5 13.02 21 OpenAI weights: small 3.499677 2 25.19 2 JAX, curated (uncontaminated) 3.534068 4 16.63 10 1xrtx3090-stacked-interventions 3.538161 4 13.51 19 JAX, curated (contaminated) 3.542460 4 13.58 18 8xa100m40-stacked-interventions-1 3.577761 4 10.19 24 JAX, FineWeb-Edu 3.632900 4 24.56 3 Cloud FineWeb, 8x A100 40 GiB 3.673623 3 16.59 11 1xrtx3090-baseline 3.683835 4 15.15 12 8xa100m40-baseline 3.691526 3 13.64 16 Cloud FineWeb, 8x H100 80 GiB 3.724507 4 13.59 17 Cloud FineWeb, 8x A100 80 GiB 3.729900 3 10.79 23 Cloud FineWeb, 8x B200 160 GiB 3.771478 4 13.70 15 Local FineWeb train 3.943522 5 11.87 22 JAX, openwebtext 4.045255 4 13.28 20 Local FineWeb-Edu extended train 4.134991 5 14.29 14 Local FineWeb-Edu train 4.166892 5 14.69 13 If you want to see the full numbers, they're below. The number that initially surprised me, and made me decide to do multiple LLM-judge runs was the one for the "JAX, FineWeb-Edu" model. In my first run it came in at 24.35 vs the OpenAI small weights' 24.93 -- so close that I wondered if it might even beat them on a re-run. However, in the further four runs its score was consistently lower than the OpenAI model's, and the gap extended a bit in some. So, was FineWeb-Edu the clear winner here? Perhaps. If you look at the contaminated/uncontaminated pairs, something interesting pops out. For the 50:50 mix, the model trained with the contaminated dataset got 19.30, and the one trained on the uncontaminated one got 17.69 -- a difference of 1.61. For the "curated" dataset, the situation was even more interesting: uncontaminated got 16.63, while contaminated got 13.58, a delta of 3.05 points. Remember, the contamination issue is about whether or not the model saw the held-back test set during training. It was an issue for the test loss that is based on that test set, but is entirely orthogonal to the IFT test. From the IFT perspective, both contaminated and uncontaminated models in each case saw training data that was -- in theory, at least -- essentially the same in terms of quality. Indeed, the uncontaminated run saw almost the same data in the same order as the contaminated one, except that some items were omitted, and then extra ones were added to the end. The purpose of this set of experiments was to see how data quality affected the results on the IFT test set. But in the case of the curated model, something that should be unrelated to data quality changed the results by 3.05 points! If something as simple as changing which data of the same quality the model is trained with can affect the IFT score so drastically, it makes it a bit harder to be certain as to whether or not data quality really had the effect we were looking for. On the other hand, the FineWeb-Edu model came in at 24.56, which is 4.06 points better than the 20.50 that the closest other model got -- more than the 3.05 points we see in difference between the two curated dataset models. And it's worth noting that the model with 20.50 is "JAX, no MHA bias, no dropout", which has a subtly different architecture -- no bias on the output projection of the multi-head attention blocks. A better comparison might be "JAX, with MHA bias, no dropout", which got a score of 17.90, for a whacking great difference of 6.66 points. I think that without doing a very large number of training runs on different datasets with different mixes, each one created with a different seed, it would be hard to work out exactly what is in the noise here and what is not. However, that would cost a lot in terms of time. I think that the best thing here is to chalk this up as a fairly decent indication that FineWeb-Edu improves matters for the IFT eval, but far from a certainty. But it's certainly worth noting that whatever the noise is, it has a range of at least 3.05 points -- and the FineWeb-Edu model is just 0.63 points short of GPT-2 small! So there could well be something there. Of course, we don't know whether that model got (by chance) the best possible balance of FineWeb-Edu tokens, and could never win -- or whether it got a bad balance and would actually beat GPT-2 with a better one. So that's certainly worth keeping in mind. As an aside, the result for the curated dataset really surprised me. I had expected that it would be the best one, simply because it almost certainly contained more facts. I took a look at its answers to the questions -- one possibility that came to mind might be that it would get better responses to questions like "What is the chemical symbol for chlorine" or "Who wrote Pride and Prejudice" than the others, but would fail on less knowledge-based tasks. But it was terrible at fact-based questions too: Name the author of 'Pride and Prejudice'. What is the periodic symbol for chlorine? As I understand it, many real-world training runs do include (often oversampled) amounts of highly educational training data like this model's dataset did. But perhaps the models that I'm training are just too small to be able to make use of the data they gained that way -- maybe doing things this way and expecting good results is like asking six-year-old children to memorise stuff before they've learned enough to be able to make use of it 1. It's worth noting that the GPT-2 small model also failed on those factual questions. Well, anyway: I think we have some useful results here, so let's work out what that means for next steps. Conclusion The results we got in these experiments point in two interesting directions. The perfect connection between the amount of FineWeb in the training set and the result on the (FineWeb-based) test loss eval, while perfectly obvious in retrospect, really does highlight how mysterious it is that the OpenAI small weights do so well on that test. The fact that FineWeb-Edu did well on the IFT test tells us that there does seem to be value in using richer training data -- though the less-spectacular results of the 50:50 mix and the curated one weaken that a bit, as does the indicator of what the noise due to data selection from equivalently high-quality datasets might be. The OpenWebText result I think I'll ignore, given that -- while in theory it should be similar to what OpenAI trained on -- there are no guarantees, and it might differ in non-obvious ways for non-obvious reasons. I think that the right direction to take this going forward is to separate these two angles. I should chase a higher IFT score, and then once I have nailed that down, I should see what (if anything) might allow me to get the resulting model to improve its test score. But I will need to make sure that whatever dataset I use, I use various "mixes" of it -- versions created with different random seeds. In my earlier experiments with overtraining, I did find that it didn't seem to improve the IFT results -- but it did improve the test loss. So perhaps identifying the right combination of other factors to boost the IFT score, then overtraining the result, might help? Of course, my overtraining tests were with FineWeb, so the connection might not hold up as well if the starting model (as seems likely) was trained on a different dataset. Also, while working through the results here, I've come to the conclusion that the set of models I'm using is a bit confusing -- there are now different hyperparameter settings, small architectural differences (the MHA bias thing), dropout settings during the pre-training, and now datasets. I think that's OK for now; I should see this part of this series as more ideation than actually running the proper experiments. But at the end, when I have some solid hypotheses with a reasonable amount of backup, I should start from scratch: a baseline model, then staged interventions to build up to what (hopefully) will be a model as good as GPT-2 small. Anyway, I'll wrap this one up here. I think that the next lever to pull is (perhaps surprisingly) going to be weight tying. I had previously kind of disregarded that as a possibility, but while I was working on this post, something popped into my mind. The OpenAI models were originally trained with weight tying. My codebase does actually support doing it -- but because I got the OpenAI weights I'm using from the code in "Build a Large Language Model (from Scratch)", when I'm running the IFT test, the weights are not actually tied! We load up a model that has separate but identical embedding and output head matrices, and then we fine-tune that. So those two matrices can vary independently during fine-tuning -- to put it another way, while GPT-2 small was pre-trained with 124M parameters, the IFT test is being done on a 163M-parameter version. Does that give them some non-obvious advantage? And would adding weight-tying to my own models help, either with or without the output heads being independent at fine-tuning time? Stay tuned :-) Appendix: all IFT judge runs Here are the numbers for all of the IFT judge runs, included for completeness. You can see that the LLM judge ranks models very consistently between runs, but there is variation -- that is, on some runs it's in what I think of as a "better mood" than others, and if that's the case, it will give better scores -- but it will give them almost consistently between models, so all of the models do better. Note that (unlike the table above) this one is sorted by the average IFT score rather than the test loss. Model Run 1 Run 2 Run 3 Run 4 Run 5 Average OpenAI weights: medium 42.24 42.16 42.95 41.83 42.61 42.36 OpenAI weights: small 24.93 24.96 25.39 25.01 25.66 25.19 JAX, FineWeb-Edu 24.35 24.55 24.3 24.68 24.9 24.56 JAX, no MHA bias, no dropout 20.5 19.9 20.76 21.25 20.07 20.50 JAX, FineWeb/FineWeb-Edu 50:50 (contaminated) 19.16 18.86 19.61 19.17 19.7 19.30 JAX, overtrained two normal epochs 18.47 18.29 19.17 18.69 18.91 18.71 JAX, overtrained one long epoch 18.04 18.71 19.62 18.41 18.57 18.67 JAX, with MHA bias, no dropout 17.49 17.35 18.33 17.73 18.62 17.90 JAX, FineWeb/FineWeb-Edu 50:50 (uncontaminated) 17.37 17.73 17.53 18.01 17.83 17.69 JAX, curated (uncontaminated) 16.77 16.03 17.3 16.08 16.96 16.63 Cloud FineWeb, 8x A100 40 GiB 16.44 16.23 17.14 16.62 16.54 16.59 1xrtx3090-baseline 14.85 15.07 15.19 15.14 15.51 15.15 Local FineWeb-Edu train 14.37 14.23 15.08 14.79 15 14.69 Local FineWeb-Edu extended train 14.4 14.07 13.82 14.56 14.61 14.29 Cloud FineWeb, 8x B200 160 GiB 13.37 13.05 13.85 13.67 14.57 13.70 8xa100m40-baseline 13.64 13.36 13.9 13.32 13.97 13.64 Cloud FineWeb, 8x H100 80 GiB 13.45 13.32 13.6 13.51 14.07 13.59 JAX, curated (contaminated) 13.09 13.48 13.95 13.23 14.15 13.58 1xrtx3090-stacked-interventions 13.37 13.11 14.04 13.84 13.17 13.51 JAX, openwebtext 12.88 12.7 13.74 13.53 13.53 13.28 JAX, no MHA bias, with dropout 13.19 12.86 12.98 12.85 13.24 13.02 Local FineWeb train 11.75 11.75 12.21 11.46 12.19 11.87 Cloud FineWeb, 8x A100 80 GiB 10.68 10.2 11.03 10.55 11.49 10.79 8xa100m40-stacked-interventions-1 9.44 9.79 10.84 10.2 10.66 10.19 A small boy asleep on his right side, the right arm stuck out, the right hand hanging limp over the edge of the bed. Through a round grating in the side of a box a voice speaks softly. "The Nile is the longest river in Africa and the second in length of all the rivers of the globe. Although falling short of the length of the Mississippi-Missouri, the Nile is at the head of all rivers as regards the length of its basin, which extends through 35 degrees of latitude …" At breakfast the next morning, "Tommy," some one says, "do you know which is the longest river in Africa?" A shaking of the head. "But don't you remember something that begins: The Nile is the …" "The - Nile - is - the - longest - river - in - Africa - and - the - second - in - length - of - all - the - rivers - of - the - globe …" The words come rushing out. "Although - falling - short - of …" "Well now, which is the longest river in Africa?" The eyes are blank. "I don't know." "But the Nile, Tommy." "The - Nile - is - the - longest - river - in - Africa - and - second …" "Then which river is the longest, Tommy?" Tommy burst into tears. "I don't know," he howls. Brave New World, Aldous Huxley ↩
Today's links Voting is to politics as shopping is to boycotts: The big P only matters if the small p is in play. Hey look at this: Delights to delectate. Object permanence: Gilberto Gil v WIPO; Censored Apple wifi hacker talk; Wells Fargo crime-spree started in 1998; Stencils "may not be reproduced"; DVD Jon v Apple DRM; Unpaid diplomatic parking tickets as index of corruption; Tortured Canadian was not a terrorist; Decarbonization at a distance. Upcoming appearances: Brighton, Virtual, South Bend, Hudson, Calgary, Winnipeg, Paris, Vancouver, Victoria, Ottawa, Kilkenny, Montreal. Recent appearances: Where I've been. Latest books: You keep readin' em, I'll keep writin' 'em. Upcoming books: Like I said, I'll keep writin' 'em. Colophon: All the rest. Voting is to politics as shopping is to boycotts (permalink) Here's a funny thing about the right to vote: it wasn't won by voting. From the Magna Carta to the US Constitution to the Emancipation Proclamation to 19th Amendment, voting rights (what you might call "Big P" Politics) were always downstream of protests, riots, petitions, mass movements, strikes and good, old fashioned community organizing (that is, "small p" politics). Which is to say, Big P politics matter, but to make them matter, we need a lot of small p politics. That means that democracy isn't something you do every couple of years with a ballot paper (though that's an important aspect of the process). Democracy is continuous. If you've ever wondered why your vote seems to accomplish so little, I think you can blame the near-abolition of small p politics by Big P politicians of every stripe. Indeed, Obama's genius was summoning up an army of door-knocking, phone-banking small p political activists and then euthanizing that organization after he won the election: https://newrepublic.com/article/140245/obamas-lost-army-inside-fall-grassroots-machine For Obama, the grassroots were useful for one thing: getting out the vote. The last thing he wanted was for millions of activated voters to turn into activists who'd flame him and harangue him and picket him if they didn't like his compromises. Boy, did Obama ever compromise. He let the bank executives who created the Great Financial Crisis off the hook and encouraged them to foreclose on the homes of millions of Americans, the very same public that had bailed them out: https://theweek.com/articles/624777/obamas-biggest-failure He shielded the CIA's torturers from scrutiny and prosecution: https://journals.law.harvard.edu/ilj/2009/04/obama-publishes-torture-memos-immunizes-cia-staff/ He reneged on his promise to shut down Gitmo: https://www.pbs.org/newshour/show/obama-failed-close-guantanamo And his promise to hold the phone companies to account for their complicity in the NSA's mass domestic surveillance: https://www.pbs.org/wgbh/frontline/article/obama-on-mass-government-surveillance-then-and-now/ He stepped up secret drone warfare: https://www.cfr.org/articles/obamas-final-drone-strike-data And unconstitutional domestic surveillance: https://www.eff.org/deeplinks/2017/01/obama-expands-surveillance-powers-his-way-out Whenever I raise this, Obama's apologists come out of the woodwork to tell me that "the president isn't the Green Lantern," and that Obama couldn't act without help from Congress and the Senate, who wouldn't back his plays. I think that Trump's presidency has shown us how much power the president really has even when the legislature won't play ball. But even if you accept the Green Lantern apologetics, the fact remains that Obama could have had a clamoring army of ardent supporters in the streets, defending his agenda against recalcitrants in his own party and wreckers in the GOP. He chose not to have that army. He sent that army home. It's like Obama heard the story about post-election FDR telling civil rights leaders, "I want to do it, now make me do it," and concluded, "I don't want to do it, so I'd better not let anyone make me do it": https://www.quora.com/Did-Franklin-Roosevelt-ever-say-I-agree-with-you-I-want-to-do-it-now-make-me-do-it Of course, Trump is doing everything he can to extinguish both small p politics and Big P Politics. It's not just his wildly illegal voter suppression tactics. He's banning and prosecuting political groups, invoking anti-terror laws (which Obama supported and promised would only be used proportionately and wisely) to chase his grassroots opposition underground: https://www.whitehouse.gov/presidential-actions/2025/09/designating-antifa-as-a-domestic-terrorist-organization/ Liberals are often contemptuous of grassroots movements (cf "basket of deplorables," "Green Lantern" scolding), but the right is terrified of them. The right's political leadership is terrified of its own grassroots, and rightly so, because those people are maniacs, and they're the reason the GOP has been pushed into its most extreme positions. The right's grassroots, meanwhile, are afraid of the left's grassroots. The last thing they want is a militant, organized, mobilized base pushing Dem politicians to take the stands that are wildly and widely popular in America, from Medicare for All to an end to ICE – the Mamdani agenda, in other words. Mamdani is the anti-Obama. He shows what happens when a progressive candidate nurtures and co-governs with their base after the election, using millions of passionate, committed, everyday people to steamroller anyone who gets in the way of his agenda: https://www.nyc.gov/content/100days/pages/ Of course the downside of this is that when Mamdani reneges on his pledges, he is loudly and furiously held to account for it: https://www.thecityreporter.nyc/2026/02/19/mamdani-budget-parks-libraries/ Mamdani understood that he would be corralled into compromises if he won the mayoralty and that when he made those compromises, his base would come after him with the unmistakable fury of betrayed idealists. He also understood that any comfort he enjoyed by sidelining his base while in office would come at a price far higher than being yelled at by his supporters: it would cost him the ability to get anything done. Voting for Mamdani was important. It got him elected. But staying organized – in unions, neighborhood clubs, affinity groups, DSA chapters and mutual aid groups – is what's letting him get stuff done, and stopping him from bailing on his promises as politically infeasible. In other words, voting only matters if it's the final stage of a sustained campaign to build and mobilize popular power. Without that, voting will get you precious little. The right's leadership understands this very well, which is why they've spent years attacking unions, community organizers like Acorn, and activist institutions like Planned Parenthood. We must defend voting rights – Big P Politics – to the bitter end, but we need to defend organizing – small p politics – just as ferociously. The reduction of politics to voting is part of the 50 year neoliberal project whose foremost goal is to make you think of yourself as an atomized individual and not as a member of a polity. Turning "politics" into "voting" is absolutely in line with Margaret Thatcher's dictum that "there is no such thing as society." It's the same move that convinced workers that the answer to bad working conditions is looking your boss in the eye and threatening to change jobs (not forming a union and striking). It's also the same move that transformed "boycotts" into "shopping." Boycotts are a collective enterprise. Before a boycott takes place, small-p political groups hold meetings, organize alternatives and communicate their demands. During a boycott, organizers work to insulate participants from reprisals, like the Montgomery Bus Boycott organizers who reasoned and remonstrated with employers who disciplined workers whose participation made them late for work. And yes, as part of a boycott, you make some consumption choices. You buy X instead of Y. But "shopping" by itself isn't a boycott. You can't "vote with your wallet" (especially not when billionaires get to vote against you with their wallets): https://pluralistic.net/2025/09/13/consumption-choices/#marginal-benefits Shopping isn't politics, and while voting is Politics (Big P), it's also not politics (small p). A boycott, on the other hand, is politics. What's more, "shopping" has the same relationship to "boycotts" that "voting" has to "politics." It's a step you take, after you've laid a lot of groundwork with other people, as part of a mass movement. I understand why shopping and voting are more attractive than boycotts and politics. Meetings suck. Hell is other people: https://locusmag.com/feature/commentary-cory-doctorow-hell-is-other-people/ But changing the system requires systemic work. Hell is other people because other people are great but it's so hard to get them to do things your way. That takes time and understanding and togetherness and arguing and forgiving. Not everyone has time or capacity for that, and at any given time, we don't all have to be doing that work. We can take turns, spelling each other off at times in our lives when we have more or less slack. But lots of us have to be in the fight, or all of us will get screwed. There aren't enough of us doing politics right now. We can tell, because our politicians are so contemptuous of the grassroots that they will sell us out without a moment's hesitation, smugly certain that they will face no consequences for doing so: https://pluralistic.net/2026/09/22/happy-chudmas/#baloney-in-our-slacks Oligarchs have it easy. Where we have to convince people to fight, they can pay or threaten people to bring them into line. But oligarchs' power is wearing thin. The data-center uprising shows how much fury there is out there, looking for a productive outlet: https://www.bloodinthemachine.com/p/with-the-backlash-to-data-centers Data centers are very bad and very visible, so they make for good targets. But data centers are only the physical extrusion of a vast, brutal, extractive system. The most important way to fight data centers is to take everyone you meet protesting one and organize with them to scare the shit out of "your" politicians so they don't dare compromise on anything. Hey look at this (permalink) The Facebook Fake-out https://www.anildash.com/2026/09/29/facebook-fake-out/ Anatomy Unzipped: John of Arderne’s Sweden Scroll (ca. 1425–35) https://publicdomainreview.org/collection/arderne-scroll/ Inside McDonald’s push to have AI price your Big Mac https://www.reuters.com/business/inside-mcdonalds-push-have-ai-price-your-big-mac-2026-09-29/ what is going on with ceiling fans https://mcmansionhell.com/post/829127919552151552/what-is-going-on-with-ceiling-fans From Shitpost to Bullshit https://www.unpopularfront.news/p/from-shitpost-to-bullshit Object permanence (permalink) #25yrsago GWB's press secretary to media: "watch what you do, watch what you say" https://web.archive.org/web/20010926223602/https://www.whitehouse.gov/news/releases/2001/09/20010926-5.html#BillMaher-Comments#BillMaher-Comments #20yrsago Stencils kit “may not be reproduced in any form” https://web.archive.org/web/20061022000842/http://www.fairuseday.com/index.php/2006/10/01/copyright-is-broken/ #20yrsago DVD Jon selling Apple DRM to Apple’s competitors https://web.archive.org/web/20061004191106/https://featured.gigaom.com/2006/10/02/dvd-jon-fairplays-apple/ #20yrsago Unpaid diplomatic parking tickets as index of national corruption https://web.archive.org/web/20130719065306/https://www.theatlantic.com/magazine/archive/2006/10/primary-sources/305203/ #20yrsago Canadian deported to Syria for torture is cleared https://www.theguardian.com/world/2006/oct/02/worlddispatch #20yrsago Gilberto Gil slams WIPO https://fromgeneva.blogspot.com/2006/09/wipo-general-assembly-impressions-from.html #20yrsago Speech given by censored Apple WiFi hacker at ToorCon https://craphound.com/cache_toorcon_2006.txt #10yrsago Company suspected of blame in Office of Personnel Management breach will help run new clearance agency https://www.reuters.com/article/us-usa-security-background-idUSKCN1202M6/ #10yrsago Wells Fargo started demanding fraud of its employees in 1998; Illinois cuts Wells off from state business https://www.citizen.org/wp-content/uploads/wells-fargo-king-of-cross-sell.pdf #10yrsago Google: if you support Amazon’s Echo, you’re cut off from Google Home and Chromecast https://variety.com/2016/digital/news/google-home-amazon-echo-chromecast-1201874125/ #5yrsago How the IMF loan-sharks the global south https://pluralistic.net/2021/10/02/debt-trap/#global-arm-breakers #1yrago Decarbonization at a distance https://pluralistic.net/2025/10/02/there-goes-the-sun/#carbon-shifting Upcoming appearances (permalink) https://www.epl.ca/blogs/post/elbows-up-with-cory-doctorow/ Brighton: Digital Sovereignty and the Post-American Internet (Green Party Conference), Oct 3 https://www.openrightsgroup.org/events/digital-sovereignty-and-the-post-american-internet/ Virtual: How to govern technology in a multipolar digital world (Connecting Current), Oct 6 https://connectingcurrent.tech/how-to-govern-technology-a-multipolar-digital-world/ South Bend: An Evening With Cory Doctorow (Notre Dame), Oct 6 https://franco.nd.edu/events/2026/10/06/an-evening-with-cory-doctorow/ Hudson, OH: Hudson Library, Oct 7 https://engagedpatrons.org/EventsExtended.cfm?SiteID=3850&EventID=596952&PK= Calgary: Wordfest, Oct 8 https://wordfest.com/2026/show/wordfest-presents-cory-doctorow-2026/ Winnipeg: McNally Robinson, Oct 9 https://www.mcnallyrobinson.com/event-18991/An-Evening-with-Cory-Doctorow Paris: Slow Tech Summit, Oct 15 https://slowtechsummit.com/ Vancouver: Read, Resist, Repair, Rejoice (Vancouver Writers Festival), Oct 19 https://writersfest.bc.ca/festival-event-2026/01 Victoria: Munro's Books, Oct 20 https://www.munrobooks.com/events/6113620261020 Vancouver: Life After AI (Vancouver Writers Festival), Oct 22 https://writersfest.bc.ca/festival-event-2026/46 Ottawa: Life After AI (Ottawa Writers Festival), Oct 24 https://writersfestival.org/event/life-after-ai Kilkenny (Kilkenomics), Nov 6-8 https://kilkenomics.com/ Vancouver: Enshittification (Sid Williams Theatre Society), Nov 10 https://www.sidwilliamstheatre.com/events/cory-doctorow-talks-enshittification/ Vancouver: BC Policy Solutions Gala, Nov 12 https://bcpolicy.ca/gala/ Montreal: World Science Fiction Convention, Sep 2-6 https://montreal2027.ca/en Recent appearances (permalink) Terms of Service with Clare Duffy (CNN) https://www.cnn.com/audio/podcasts/terms-of-service-with-clare-duffy/episodes/458ce968-af5d-11f0-b539-13ed2afe25f8 AI, Work, and Power (Software Engineering Daily) AI, Work, and Power https://softwareengineeringdaily.com/podcasts/cory-doctorow-on-ai-work-and-power/ AI, Corporate Power, and the Fight for Worker Control (Plutopia) https://plutopia.io/cory-doctorow-ai-corporate-power-and-the-fight-for-worker-control/ How to Think About AI—Before It’s Too Late (Daniel Solove) https://www.youtube.com/watch?v=_0xR3uEgGcc Could Tech Bosses Destroy Life As We Know It? (Politics JOE) https://www.youtube.com/watch?v=PL4VktU0SgY Latest books (permalink) "The Reverse-Centaur's Guide to AI," a short book about being a better AI critic, Farrar, Straus and Giroux, June 2026 https://us.macmillan.com/books/9780374621568/thereversecentaursguidetolifeafterai/ "Canny Valley": A limited edition collection of the collages I create for Pluralistic, self-published, September 2025 https://pluralistic.net/2025/09/04/illustrious/#chairman-bruce "Enshittification: Why Everything Suddenly Got Worse and What to Do About It," Farrar, Straus, Giroux, October 7 2025 https://us.macmillan.com/books/9780374619329/enshittification/ "Picks and Shovels": a sequel to "Red Team Blues," about the heroic era of the PC, Tor Books (US), Head of Zeus (UK), February 2025 (https://us.macmillan.com/books/9781250865908/picksandshovels). "The Bezzle": a sequel to "Red Team Blues," about prison-tech and other grifts, Tor Books (US), Head of Zeus (UK), February 2024 (thebezzle.org). "The Lost Cause:" a solarpunk novel of hope in the climate emergency, Tor Books (US), Head of Zeus (UK), November 2023 (http://lost-cause.org). "The Internet Con": A nonfiction book about interoperability and Big Tech (Verso) September 2023 (http://seizethemeansofcomputation.org). Signed copies at Book Soup (https://www.booksoup.com/book/9781804291245). "Red Team Blues": "A grabby, compulsive thriller that will leave you knowing more about how the world works than you did before." Tor Books http://redteamblues.com. "Chokepoint Capitalism: How to Beat Big Tech, Tame Big Content, and Get Artists Paid, with Rebecca Giblin", on how to unrig the markets for creative labor, Beacon Press/Scribe 2022 https://chokepointcapitalism.com Upcoming books (permalink) "The Post-American Internet," a geopolitical sequel of sorts to Enshittification, Farrar, Straus and Giroux, 2027 "Unauthorized Bread": a middle-grades graphic novel adapted from my novella about refugees, toasters and DRM, FirstSecond, April 20, 2027 "Enshittification, Why Everything Suddenly Got Worse and What to Do About It" (the graphic novel), Firstsecond, 2027 "The Memex Method," Farrar, Straus, Giroux, 2027 Colophon (permalink) Today's top sources: Currently writing: “Once Is Enemy Action,” a science fiction novel about the origins of modern technofascism. Today's words: 509 (20770 total). "The Post-American Internet," a sequel to "Enshittification," about the better world the rest of us get to have now that Trump has torched America. Fourth draft completed. Submitted to editor. A Little Brother short story about DIY insulin PLANNING This work – excluding any serialized fiction – is licensed under a Creative Commons Attribution 4.0 license. That means you can use it any way you like, including commercially, provided that you attribute it to me, Cory Doctorow, and include a link to pluralistic.net. https://creativecommons.org/licenses/by/4.0/ Quotations and images are not included in this license; they are included either under a limitation or exception to copyright, or on the basis of a separate license. Please exercise caution. How to get Pluralistic: Blog (no ads, tracking, or data-collection): Pluralistic.net Newsletter (no ads, tracking, or data-collection): https://pluralistic.net/plura-list Mastodon (no ads, tracking, or data-collection): https://mamot.fr/@pluralistic Bluesky (no ads, possible tracking and data-collection): https://bsky.app/profile/doctorow.pluralistic.net Medium (no ads, paywalled): https://doctorow.medium.com/ Tumblr (mass-scale, unrestricted, third-party surveillance and advertising): https://mostlysignssomeportents.tumblr.com/tagged/pluralistic "When life gives you SARS, you make sarsaparilla" -Joey "Accordion Guy" DeVilla READ CAREFULLY: By reading this, you agree, on behalf of your employer, to release me from all obligations and waivers arising from any and all NON-NEGOTIATED agreements, licenses, terms-of-service, shrinkwrap, clickwrap, browsewrap, confidentiality, non-disclosure, non-compete and acceptable use policies ("BOGUS AGREEMENTS") that I have entered into with your employer, its partners, licensors, agents and assigns, in perpetuity, without prejudice to my ongoing rights and privileges. You further represent that you have the authority to release me from any BOGUS AGREEMENTS on behalf of your employer. ISSN: 3066-764X
The unstated disagreement that underpins safety debates