Full Width [alt+shift+f] Shortcuts [alt+shift+k]
Sign Up [alt+shift+s] Log In [alt+shift+l]

Giles' blog

Sort By

Recent [alt+2]
Putting my JAX-trained models on the Hugging Face Hub I hadn't uploaded the models that I trained using JAX to the Hugging Face Hub because Transformers...
3 weeks ago
1

Putting my JAX-trained models on the Hugging Face Hub

from Giles' blog [alt+shift+b] in AI

3 weeks ago
I hadn't uploaded the models that I trained using JAX to the Hugging Face Hub because Transformers has been PyTorch-only since version 5 (though they say they're working to add interoperability with JAX in the future), so it would have been tough to get them working natively with...
A quick(ish) Chinchilla check I recently overtrained a couple of GPT-2 style models, training them both on 40 tokens per parameter...
7th Aug 2026
2

A quick(ish) Chinchilla check

from Giles' blog [alt+shift+b] in AI

7th Aug 2026
I recently overtrained a couple of GPT-2 style models, training them both on 40 tokens per parameter rather than the 20 per parameter that is generally regarded as "Chinchilla-optimal". The normal heuristic is that instead of doing that, you should scale up the number of tokens...
Writing an LLM from scratch, part 34b -- from bigrams to GPT-2, one component at a time (in JAX) This post is the capstone of the most long-running series on my blog. In December 2024 (!), I...
8th Jul 2026
2
8th Jul 2026
This post is the capstone of the most long-running series on my blog. In December 2024 (!), I started reading Sebastian Raschka's book "Build a Large Language Model (from Scratch)", and worked through it carefully. Being who I am, despite trying to apply a strict "no side...
Using Safetensors with Flax I'm porting my PyTorch LLM code to JAX, using Flax as the neural network layer. For various reasons...
4th Jun 2026
1

Using Safetensors with Flax

from Giles' blog [alt+shift+b] in AI

4th Jun 2026
I'm porting my PyTorch LLM code to JAX, using Flax as the neural network layer. For various reasons I wanted to use Safetensors to store checkpoints of the model. It took a little while to get it working; here's the trick I learned. If you look at the Safetensors docs, you'll...
10Gb/s Ethernet: using mini-heatsinks with a 10GBASE-T SFP+ module In my last post I showed the somewhat-scary temperatures I was getting on the MikroTik 10GBASE-T...
18th May 2026
1
18th May 2026
In my last post I showed the somewhat-scary temperatures I was getting on the MikroTik 10GBASE-T SFP+ module I have plugged into nigel, the 10Gb/s switch I have in my study. As I mentioned then, the plan was to try using some of the mini-heatsinks that people use on Raspberry...
Automating starting Lambda Labs instances I've been trying to get an 8x A100 instance on Lambda Labs to do a training run for my LLM from...
2nd Apr 2026
1

Automating starting Lambda Labs instances

from Giles' blog [alt+shift+b] in AI

2nd Apr 2026
I've been trying to get an 8x A100 instance on Lambda Labs to do a training run for my LLM from scratch series, but they're really busy at the moment, and it's rare to see anything. Thanks to the wonders of agentic coding, I spent an hour today getting something up and running to...
Writing an LLM from scratch, part 32e -- Interventions: the learning rate I'm still working on improving the test loss for a from-scratch GPT-2 small base model, trained on...
10th Mar 2026
1
10th Mar 2026
I'm still working on improving the test loss for a from-scratch GPT-2 small base model, trained on code based on Sebastian Raschka's book "Build a Large Language Model (from Scratch)". In my training code, I have this code to create the optimiser: optimizer =...
Writing an LLM from scratch, part 32a -- Interventions: training a baseline model I'm rounding out my series of posts on Sebastian Raschka's book "Build a Large Language Model (from...
4th Feb 2026
1
4th Feb 2026
I'm rounding out my series of posts on Sebastian Raschka's book "Build a Large Language Model (from Scratch)" by seeing how I could train the best base model I can from scratch on my own hardware. I started by training one in two days on my RTX 3090, and found that while it was a...
Writing an LLM from scratch, part 29 -- using DistributedDataParallel to train a base model from... I'm carrying on with my "extra credit" projects after finishing the main body of Sebastian Raschka's...
7th Jan 2026
1
7th Jan 2026
I'm carrying on with my "extra credit" projects after finishing the main body of Sebastian Raschka's book "Build a Large Language Model (from Scratch)". Having proven that I could train a GPT-2 small scale base model from scratch on my RTX 3090 in 48 hours, I wanted to try...
Writing an LLM from scratch, part 28 -- training a base model from scratch on an RTX 3090 Having worked through the main body of Sebastian Raschka's book "Build a Large Language Model (from...
2nd Dec 2025
1
2nd Dec 2025
Having worked through the main body of Sebastian Raschka's book "Build a Large Language Model (from Scratch)", I wanted to try an experiment: is it possible to train a base model of my own, on my own hardware? The book shows you how to train your LLM, does a basic training run on...
📚 BoredReading

You seem to be enjoying this.

Join free to unlock everything.

Create free account

Already have an account? Sign in