Full Width [alt+shift+f] Shortcuts [alt+shift+k]
Sign Up [alt+shift+s] Log In [alt+shift+l]
1

Two worlds: idea people and developers

from Tiny struggles [alt+shift+b] in technology

I recently attended a startup weekend and was surprised by the backgrounds of the other attendees. As a software engineer myself, I expected to see mostly engineers, but that wasn’t the case. Instead, I found myself surrounded by people with a diverse range of backgrounds and experiences, all of whom shared one thing in common: they had an idea. At the same time, couple months ago started running an Indie Hackers Meetup in Dublin, a software startup meetup that mostly attracts developers. And from other breaking news, as a result of that startup weekend I’m now deep into a new project. For the first time joining forces with a cofounder. The project is still in pretty early stages and I don’t know if it will become a serious business yet. Fingers crossed. This all got me thinking about the two distinct groups of people in the startup world - those with the skills to execute an idea but lacking a good business idea that would solve a real life problem and had financial potential, and those with great ideas but lacking the technical expertise to bring them to life. How can they meet and should they join forces as cofounders? Two groups of people Many developers dream of starting their own software businesses, while countless non-developers struggle to find the right development team to build their ideas. Do they need to join forces to build a company together? Well, no. There are many other ways. A developer can still get hired to work on another person’s idea and the dreamer with an idea can seek funding through an angel investment, family and friends or VC. But the truth is, these approaches have some serious drawbacks. For developers, it often means giving up control over the direction of the company they’re working for, in exchange for a relatively small amount of equity. And for entrepreneurs, it can mean giving up a significant chunk of their ownership in order to secure funding. And even then, they may struggle to find the right talent to build the technology...
29th Apr 2023

Stay updated

Get a weekly newsletter with the top 5 articles worth reading every week.

More from Tiny struggles

Implementing SARM on your VLA dataset in practice

1. Motivation: use big video dataset optimally for VLA training In the first part of this series, I explained what is SARM: Stage-Aware Reward Modeling for Long Horizon Robot Manipulation and how it can be used with a challenge such as Stanford Behavior Challenge (1200h of demonstrations over 50 diverse long horizon tasks). To sum up: Our main model for the robot (VLA) is trained on short windows (chunks) of data for which it predicts actions. Our SARM model is used to estimate progress within an episode and evaluates windows. We want to prioritize training on trajectory segments where the robot made meaningful progress toward task completion. In this post I will explain how I actually implemented this in practice. The code is now open on github. The core of the implementation follows closely the original paper. We will cover the following key areas: The design of the model Sequential Multimodal Architecture that utilizes a Global Anchor Frame. Data input shape and preparation Using the model for scoring the episodes Visual Validation of the predicted progress against our Stage-Aware Ground Truth. 2. The SARM Model Implementation See the source here. The model tackles a dual prediction problem: determining which stage of a task is being performed (classification) and how much progress has been made within that stage (regression). SARM provides a principled approach to stage-aware reward modeling by: Leveraging pretrained vision models (CLIP) for robust visual understanding Fusing multimodal information through Transformers Making hierarchical predictions (stage → progress within stage) Handling variable-length sequences efficiently Architecture Overview This architecture is particularly well-suited for tasks that have clear sequential structure and require fine-grained progress estimation within each stage. The SARM model follows a three-part design: Encoders - Process multimodal inputs (visual and proprioceptive) Shared Backbone - A Transformer that fuses information across time and modalities Dual Heads - Separate outputs for stage classification and progress regression Where the symbols are: B - batch size N - sequence size (multiple frames of images/data - more on that later) ┌────────────────────────────────────────────────────────────────────┐│ INPUT LAYER ││ ││ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ ││ │ Image Frames │ │ Joint States │ │ Task Index │ ││ │ (B,N,3, │ │ (B,N,256) │ │ (B,) │ ││ │ 224,224) │ │ │ │ │ ││ └──────┬───────┘ └───────┬──────┘ └────────┬─────┘ │└─────────┼─────────────────-┼──────────────────┼────────────────────┘ │ │ │ ▼ ▼ ▼┌─────────────────────────────────────────────────────────────────────┐│ ENCODER LAYER ││ ││ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ ││ │ CLIP (ViT) │ │ LayerNorm │ │ Embedding │ ││ │ [Frozen] │ │ + │ │ Layer │ ││ │ ↓ │ │ Linear │ │ │ ││ │ Linear │ │ │ │ │ ││ │ Projection │ │ Projection │ │ │ ││ │ │ │ │ │ │ ││ │ (512→768) │ │ (256→768) │ │ (50→768) │ ││ └──────┬───────┘ └──────┬───────┘ └───────┬──────┘ ││ │ │ │ ││ │ Visual │ State │ Task ││ │ Embeddings │ Embeddings │ Embedding ││ │ (B,N,768) │ (B,N,768) │ (B,1,768) ││ └─────────┬───────┴──────────────────┘ │└───────────────────┼─────────────────────────────────────────────────┘ │ ▼ ┌─────────────────┐ │ Element-wise │ │ Sum │ │ │ │ Visual + State │ │ + Task │ └────────┬────────┘ │ ▼ ┌─────────────────┐ │ Add Positional │ │ Bias to Frame 0│ └────────┬────────┘ │ │ Combined Embeddings │ (B,N,768) ▼┌────────────────────────────────────────────────────────────────────┐│ TRANSFORMER BACKBONE ││ ││ ┌───────────────────────────────────────────────────────────────┐ ││ │ Transformer Encoder (8 layers) │ ││ │ │ ││ │ ┌─────────────────────────────────────────────────────────┐ │ ││ │ │ Multi-Head Self-Attention (12 heads) │ │ ││ │ │ d_model = 768, │ │ ││ │ │ Dropout = 0.1 │ │ ││ │ └─────────────────────────────────────────────────────────┘ │ ││ │ ×8 │ ││ └───────────────────────────────────────────────────────────────┘ ││ ││ (with padding mask support) │└──────────────────────────────┬─────────────────────────────────────┘ │ │ Aggregated Features │ (B,N,768) ▼┌─────────────────────────────────────────────────────────────────────┐│ OUTPUT HEADS ││ ││ ┌──────────────────────┴──────────────-────────┐ ││ │ │ ││ ▼ ▼ ││ ┌─────────────────┐ ┌─────────────────┐ ││ │ Stage Head │ │ Subtask Head │ ││ │ (Classifier) │ │ (Regressor) │ ││ │ │ │ │ ││ │ Linear(768→512)│ ┌────────────────┤ Concat: │ ││ │ ReLU │ │ │ - Features(768)│ ││ │ Dropout(0.1) │ │ │ - Logits(100) │ ││ │ Linear(512→100)│─────────┘ │ │ ││ │ │ │ Linear(868→512) │ ││ │ Stage Logits │ │ ReLU │ ││ │ (B,N,100) │ │ Dropout(0.1) │ ││ └─────────────────┘ │ Linear(512→1) │ ││ │ Sigmoid │ ││ │ │ ││ │ Scalar Progress │ ││ │ (B,N) │ ││ └─────────────────┘ │└─────────────────────────────────────────────────────────────────────┘ Data Flow Explained 1. Input Processing The model accepts three types of inputs for each sequence: Image Frames (B, N, 3, 224, 224): A batch of N RGB images per sequence Joint States (B, N, D_state): Robot proprioceptive information (joint angles, positions, etc.) Task Index (B,): An integer identifying which task is being performed Where B is the batch size and N is the maximum sequence length, the actual data can be shorter and then we pad it. For every prediction, our model processes a sequence of frames, deliberately structured to provide maximum temporal context: Global Anchor Frame: The first frame of the episode is included in every sequence. This is a crucial engineering choice for long-horizon tasks, as it gives the Transformer a global, unchanging reference point for the task’s initial state. Subsampled Context Frames: Several preceding frames are included to capture recent history. Current Frame: The frame for which the progress prediction is required. 2. Encoding Stage Each modality is processed through its own encoder: Visual Encoding: Images are flattened from (B, N, 3, 224, 224) to (B*N, 3, 224, 224) Passed through a frozen CLIP ViT-B/32 model to extract visual features CLIP outputs 512-dimensional features per image Features are projected to the model dimension (768) via a linear layer Reshaped back to (B, N, 768) State Encoding: Joint states are normalized using LayerNorm Projected from dimension 256 to 768 via a linear layer Output: (B, N, 768) Task Encoding: Task index is converted to a learned embedding vector The embedding is replicated across the sequence: (B,) → (B, 1, 768) This embedding is broadcast and added to all timesteps 3. Multimodal Fusion The three encoded representations are combined: input_embeddings = visual_embeddings + state_embeddings + task_embedding Additionally, a learned positional bias is added only to the first frame: input_embeddings[:, 0, :] += positional_bias This creates a unified representation (B, N, 768) that contains information from all modalities. 4. Transformer Backbone The combined embeddings are processed through an 8-layer Transformer encoder: Architecture: Standard Transformer encoder with 12 attention heads Dimensions: 768-dimensional hidden states, 3072-dimensional feedforward layers Padding Support: The model accepts an optional padding mask (B, N) where True indicates padded positions Output: Aggregated features (B, N, 768) that capture temporal and multimodal dependencies 5. Dual Output Heads The model produces two types of predictions: Stage Head (Classification): Takes the aggregated features (B, N, 768) Passes through: Linear(768→512) → ReLU → Dropout → Linear(512→100) Outputs stage logits (B, N, 100) representing 100 possible task stages: 100 is a maximum number of task stages supported, in practice the stages will be task dependent Trained with cross-entropy loss Subtask Head (Regression): Concatenates aggregated features with stage logits: [features, stage_logits] → (B, N, 868) This conditioning allows progress estimation to be stage-aware Passes through: Linear(868→512) → ReLU → Dropout → Linear(512→1) → Sigmoid Outputs scalar progress (B, N) in the range [0, 1] Trained with MSE loss Loss Computation The SARMWithLoss wrapper handles training: Masking: Only non-padded positions are included in loss calculation Stage Loss: Cross-entropy between predicted logits and ground truth stage labels Progress Loss: MSE between predicted progress and ground truth progress values Total Loss: Weighted sum of both losses (default weights: 1.0 each) total_loss = (stage_loss_weight × stage_loss) + (progress_loss_weight × progress_loss) Including the loss calculation within the model wrapper made the training code simpler and more standard. Key Design Decisions These decisions follow the original SARM paper: Why freeze CLIP?CLIP is pretrained on massive image-text datasets and provides robust visual features. Freezing it prevents overfitting on smaller robotics datasets and reduces computational cost. Why condition subtask head on stage predictions?We estimate progress within a stage, not the whole episode, so it’s stage dependent. Why add positional bias only to the first frame?The first frame often contains important context about the initial state. The positional bias helps the model distinguish the starting point from subsequent frames. Supposedly such anchoring is very effective for video models. Why use variable-length sequences with padding? We use Rewind Augmentation when we sometimes generate longer sequences that ‘mess up’ progress on purpose by replaying older frames in the reverse order. Because of that we need to handle sequences of varied length. This augmentation is critical for the model to learn how undoing progress looks like. 3. Complex data preparation The data for SARM has to be prepared in a very particular way. The core of the sampling is implemented in the custom dataloaders here. Temporal Sampling Strategy SARM doesn’t sample frames uniformly. Instead, it uses a sophisticated sampling strategy designed to provide temporal context: def prepare_indices(ep_first_frame_idx, idx, skip_count=30, default_length=8, rewind_prob=0.05): # Sample backwards in time with skip_count intervals indices = [idx - i * skip_count for i in range(default_length)] # Always include the first frame of the episode indices.append(ep_first_frame_idx) # Reverse so time flows forward indices = list(reversed(indices)) # 5% chance: add "rewound" frames for temporal augmentation if random() < rewind_prob: num_extra = random.integers(2, 5) indices += [indices[-1 - i] for i in range(1, num_extra + 1)] return indices This creates sequences with the following structure: Sequence construction (skip_count=30, ~1 second at 30 FPS):┌───────┬───────┬───────┬───────┬───────┬───────┬───────┬───────┬───────┐│Frame 0│ t-7s │ t-6s │ t-5s │ t-4s │ t-3s │ t-2s │ t-1s │ t ││(start)│ │ │ │ │ │ │ │(curr) │└───────┴───────┴───────┴───────┴───────┴───────┴───────┴───────┴───────┘With 5% probability, add rewind frames:┌───────┬───────┬─────────────┬───────┬───────┬───────┐│... │ t │ t-1s (again)│ t-2s │ t-3s │(curr) │└───────┴───────┴─────────────┴───────┴───────┴───────┘ Why this design? Always anchor to episode start: Frame 0 provides consistent context about initial conditions Uniform temporal spacing: 1-second intervals capture motion patterns without redundancy Rewind augmentation: Teaches the model temporal reversibility and robustness Future context avoided: Model only sees past and present, not future frames Delta Timestamps Pattern The sampling strategy is complemented by a clever timestamping scheme: DELTA_TIMESTAMPS = [HIGH_NEGATIVE_TIMEDELTA] + [-7 + i for i in range(8)]# Results in: [1e6, -7, -6, -5, -4, -3, -2, -1, 0] When applied to current timestamp t: Frame 0: Gets timestamp ≈ -∞ (approximated as episode start) Frames 1-7: Get timestamps [t-7, t-6, …, t-1] Frame 8: Gets timestamp t (current) This ensures consistent temporal windows regardless of where you are in the episode.My custom dataset SARMDataset uses a dataset provided by the BEHAVIOR codebase under the hood that allows specifying delta timestamps for more efficient sampling. Variable-Length Sequence Handling Real episodes have variable lengths, and sequences can have different numbers of frames. SARM handles this with: def collate_fn(batch): # Find max sequence length in batch max_length = max(sample["sequence_length"] for sample in batch) # Pad all sequences to max_length batched_images = torch.zeros(batch_size, max_length, C, H, W) batched_padding_mask = torch.ones(batch_size, max_length, dtype=torch.bool) for i, sample in enumerate(batch): seq_len = sample["sequence_length"] batched_images[i, :seq_len] = sample["images"] batched_padding_mask[i, :seq_len] = False # False = valid, True = padding The padding mask is then passed to the Transformer to ensure padded positions don’t contribute to attention or loss: Example batch with lengths [9, 11, 13, 9]:Padded to max_length=13:┌─────────────┬─────────────┬─────────────┬─────────────┐│ Seq 1 (9) │ Seq 2 (11) │ Seq 3 (13) │ Seq 4 (9) │├─────────────┼─────────────┼─────────────┼─────────────┤│ [V][V]...[V]│ [V][V]...[V]│ [V][V]...[V]│ [V][V]...[V]││ [P][P][P][P]│ [P][P] │ │ [P][P][P][P]│└─────────────┴─────────────┴─────────────┴─────────────┘ V = Valid token P = Padding token (masked out) Inference Sampling Strategy To use SARM for VLA training, we need to run our original video dataset through SARM. But that dataset was huge to begin with! But we don’t need to evaluate every frame (with its proceeding sequence). With 5-second sampling at 30 FPS, I evaluate only 1 out of every 150 frames (5s × 30 FPS), reducing computational cost by 150×. I implemented a following dataloader (source). Why jitter? Jitter prevents the model from overfitting to fixed timestamps and produces more robust progress estimates by sampling at slightly varied intervals rather than exact multiples of 5 seconds. See INFERENCE_README for more details on the inference. Translating the progress to VLA training weights Additionally I implement the progress mapping to the weights following the SARM paper (equations 8-9):- Computes progress deltas r̂ᵢ = φ(t+Δ) - φ(t)- Uses running statistics (μ, σ) to normalize- Applies linear ramp between (μ - 2σ) and (μ + 2σ)- Optionally uses threshold κ for decisive weighting See the code in weight utils. 4. Training & Results This implementation of SARM was multi-task, however since there was so much data, I decided that it would be easier to evaluate and visualize it on a single task first. Having 200 episodes for each task, I divided the data into the following sets: “train_episodes”: 1-90, “val_episodes”: 91-105, “test_episodes”: 106-200 In general performance on the validation set wasn’t the best indicator of actual model performance when I analyzed it on the test dataset. Training for more steps was helpful. For training details see the config and the training script. The 10k-step snapshot has been trained on a single RTX5090 over several hours. The model is also compatible with training on MPS. The key bottleneck was the dataset access and video processing. Visualizations & analysis I performed detailed analysis on how well the models were predicting the progress here. Ground Truth First, it’s important how the ‘ground truth’ data looks like, here is a visualization: Based on the data annotations and the stage statistics I was able to generate our ‘Ground truth’ of progress. It was also a useful sanity check if the ground truth data looks right, e.g. having negative progress in the ground truth data would mean that there were bugs. We were only adding ’negative progress’ through the Rewind augmentation later on. Comparing models vs ground truth and each other Model checkpoint comparison vs ‘ground truth’ on a sample of episodes: The first 3 episodes were in the training data and the 2 last ones weren’t present.You can see here that the yellow (10k steps) model is better fitted to the data in the training set. Understanding the bias of the model So the model wasn’t perfect, what type of mistakes was it making? Overall, the model was leaning towards underestimating the progress. And the key problem was from predicting wrong stage number. Applicability and Limitations The caveat here that the task 8 was multimodal, the stages could be done in variable order, breaking the fundamental assumption of the fix stage order in SARM. Visualizations for a task fitting SARM assumptions would look better. The SARM Assumption: Stage-Aware modeling assumes a generally linear path through semantic checkpoints (Stage 1 $\rightarrow$ Stage 2 $\rightarrow$ Stage 3…). The Failure Case: Multimodal Progress: If a task allows for subtasks to be completed in an arbitrary order (e.g., “Tidy up the room”), the progress estimation becomes inherently multimodal, and our regression model, forced to average these possibilities, loses accuracy. Handling different sequences of stages in demonstrations Removing outliers If the majority of demonstrations are done in a consistent way, then we can remove the outlier demonstrations that create confusion. (Annotations to generate ground truth are enough to blacklist such episodes). Subtask splitting The SARM model implemented here can handle multiple tasks. Therefore, if a task can be done using different sequences of stages, we can transform it into set of related tasks with different demonstration variations. If there are multiple different ways represented in similar proportions, e.g. ‘pick up toy 1’ then ‘pick up toy 2’, and the reverse, we can change into two tasks pick_up_toys_1_2, pick_up_toys_2_1. Equal sampling We can also decide not use SARM for such tasks. SARM in Behavior challenge For the BEHAVIOR challenge specifically, we were very time constrained, and we ended up not having enough time to apply the model for the final checkpoint training, additionally, only about 30% of tasks fulfilled the fixed stage ordering for SARM. We performed quick fine tuning with weighted sampling earlier on a subset of data (not based on SARM), but it was difficult to see if it was actually helpful (eval in general was pretty challenging). Despite these challenges, SARM remains a promising approach for datasets with proper stage annotations and sequential task structure. Our analysis on the 30% of tasks with fixed stage ordering showed the model could accurately track progress. For the remaining tasks, the subtask splitting approach outlined above could make SARM applicable, potentially enabling more efficient VLA training through intelligent data selection across the full dataset.

14th Dec 2025 • 1 votes
Robotics Hackathon in Bimanual Manipulation in Munich

How a LinkedIn Post Led Me to a Munich Basement with Millions of Euros Worth of Robotics Equipment My LinkedIn feed has become a stream of robotics content over the past few months. As someone diving deep into AI robotics after years in ML/AI/RL, I’ve been deliberately connecting with people pushing the boundaries of the field. So when Nicolas Keller’s post about Munich being “the world’s best place to build robots” appeared in my feed, it immediately got my attention. A bimanual manipulation hackathon in Munich, organized in just three weeks. How cool is that? Here’s what they promised: Hands-on with dual-cobot humanoid upper-body setups, 2 Franka Emika Pandas, 2 depth cameras, 1 rgb, RTX5090 in each station Teleoperation with Meta Quest VR headsets Contributing to the MINGA research paper (with co-author potential) Collaborating in Munich’s growing robotics ecosystem I had three weeks to apply, get accepted, and arrange travel during what turned out to be Oktoberfest (the Lederhosen on the robot in the picture should have been a hint! It just made the stay extra expensive). I have recently missed the LeRobot’s worldwide hackathon and wanted to jump on the opportunity. The prospect of 30+ dual-arm robot setups, high-end GPUs, industry mentors and meeting other people passionate about robotic manipulation made it a no-brainer decision for me. The hackathon experience Most hackathons are short and they only involve software. Hardware Hackathons are much more rare, especially where hardware is provided and high end. The organizers promised a lot - and I have to say that they over-delivered, even though they operated on a very short timeline! 3 weeks! The hackathon wasn’t perfect; we hit some technical issues with the provided codebase and the robot controllers.The lab space (KI Fabrik in the Deutsches Museum) was full of amazing robots, powerful workstations, and a 3D workshop, but it also had its downsides - it was hot, humid and you could get trapped there due to the limited number of keys! The schedule was focused on building: after one day of setup and tutorials, it was essentially “09:00–open ended — Building time” for six straight days. The main communication happened on Discord. The setup was industrial-grade: 30+ Franka Emika Panda dual-arm configurations for about 40 participants. Each setup came with Meta Quest headsets running custom teleoperation software. There was a full workshop with 3D printing capabilities for custom grippers. The compute power was serious—workstations with the latest hardware that most of us don’t have access to. The initial goal of the organizers was to attract local students (mostly from TUM), but the hackathon was just too attractive. The organizers ran the selection process based on a Typeform where you had to justify your presence (CV, motivation, experience) and the final mix of people contained: PhD researchers, startup founders, industry engineers, and ambitious students. There was a significant number of people who traveled from other countries. Everyone wanted to be there. Industry engagement and realistic use cases The real differentiator was the industry backing. BMW and Siemens provided realistic challenges to be solved, explained the details, provided physical materials and sponsored prizes. Additionally, there were helpful lectures and mentoring from Nvidia, Hugging Face LeRobot and KIT (a big German university). When Sunday’s final presentations came, BMW and Siemens employes showed up to judge the results personally. This wasn’t academic theory - teams were working on problems that companies actually need solved, with the decision-makers accessible throughout and present for the final outcomes. The main organizers were TUM and Poke & Wiggle. I was very impressed with both. TUM showed great support for students and entrepreneurship. Poke&Wiggle people pulled everything together from the technical side. They were staying late and even hosted some participants coming from abroad in their homes! They were testing their own software stack while building what they claimed would become the largest public bimanual manipulation dataset. My strategy & Learning Probably my favorite thing about this hackathon was that it was extremely collaborative. Yes, people tried to win, but I was able to learn both from my team as well as from others. Team formation & collaboration I came to the event without knowing anyone and needed to form a team. It was actually a pretty common experience, many people didn’t know who to pair with. I have been in setups like this before and that experience helped. To form a good team you want the best people, but also it’s really hard to assess people very quickly and people who seem great at the start might not be able to give the best performance. E.g. they might not be fully available, you might not get along very well, etc. So my strategy was to talk to most people, see how they think and how experienced they are. I was selecting for getting along, enthusiasm and general intelligence. It was less important to me that someone was inexperienced as long as they were energetic and open minded. We weren’t the most effective or best organized, but we really enjoyed our time together, made good progress and learned a lot. Our team also shifted a bit during the 7 days (one person got sick and we adopted another one). We used a WhatsApp group to share resources, set up a GitHub repo for shared scripts and shared some notes on a Google doc. Our strategy was as follows: get familiar with the teleoperation trying out all available tasks and assess the task feasibility for the human operators train and deploy the models ASAP to test the pipeline (yes, we detected bugs and further limitations) pursue the most promising tasks, refine the dataset collection and experiment with the models for good performance Learnings We didn’t manage to win any categories. I think that my group was more focused on learning and experimentation, instead of purely competing to win. Some takeaways: If the task can’t be done by the human, the robot won’t be able to do it Training loss during model training is not a good predictor of the real life performance Robot safety mechanisms were critical to avoid breaking the robot Evaluation in real life is risky and some simulation setup would be helpful End-effector control with inverse kinematics was often causing the robots to get stuck due to joint limits; it required special care during teleoperation for the demos so that the actual policy wouldn’t block the robot Two arms are much harder than one Dataset quality matters a lot (recovery examples, non-Markovian states are confusing, noise/operators in the setup) What we tested 3 different task setups with hundred+ demonstrations each variations in model training (steps, parameters, models, action space) and dataset selection (recovery episodes ratio, bad episodes) different control modes (delta and absolute) we could only try actions in the EE space, the joint space controllers weren’t working correctly SMOLVLA and ACT models from Lerobot libraries performed experiments for generalization and resilience (e.g. messing with cameras was making robots much less effective) we also visualized attention maps for the ACT Models Here is an example of an attention map: (Kudos to physical AI Interpretability Repo - the author was there during the event and helped us a bit with the setup). Misc I got a pretty good feel for teleoperation in VR (it’s hard though!) I managed to get the arms to crash with each other and got my robots stuck countless times I spent some time setting up a simulation environment with Panda in MuJoCo and playing with Isaac Sim (approach abandoned in the end) I read several papers recommended by other participants I wasn’t able to try Pi0/Pi0.5 as the ready snapshots are for the same type of robot (Panda), but in a very different action space and we didn’t have time/resources for fine-tuning from scratch (80GB+ GPU memory required) What I wished I could do: try out Groot / Pi0.5 simulation, sim-to-real, and RL fine-tuning for the trained VLA (a simple VLA would be perfect!) Posts describing the experiences of some of my teammates: @Artur and @Andrea. We Need More of This After seven days of intense collaboration in a basement, we didn’t revolutionize the future of robotics, but we all learned a lot and everyone came back home more experienced and inspired. It was a great event! And I want to see more events like this in Europe, because the future of robotics doesn’t have to happen in SFO (or China) Europe has the ingredients for world-class robotics innovation. We have strong engineering talent and strong reasons to invest (aging population)! I live in Poland now, a place that produces some of the world’s best software engineers, but the innovation is lacking. I talked to two universities in Warsaw and they don’t really innovate or even follow the current state of the art yet for embodied intelligence. And it’s a shame, because with libraries such as LeRobot, open hardware, and open-source simulation engines, the space is now much more accessible. I was very impressed with TUM and many of the students. I would like to support the ecosystem in Poland and Europe. Munich proved it’s possible. Let me know if you would like to help! Next steps I left the event pretty drained, but also very excited! JI am still following up on the various threads I started during the hackathon. I also started to look at another exciting challenge that is focused on household tasks in the simulation. Currently I’m especially interested in the approaches combining VLAs with RL, like the ones outlined in the SimpleVLA paper and would love to participate in more hardware hackathons, ideally combining simulation and real world learning.

5th Oct 2025 • 1 votes
Creating a Robotics Experimentation Environment: My Experience and Practical Lessons

Part 1 of a series on practical robotics experimentation As part of my journey into robotics, I found myself facing a classic problem: I wanted to experiment with different learning approaches for robotic manipulation, but I needed a flexible playground where I could quickly test ideas without being locked into any single framework or workflow. The result is gym-so100-c, a simulation environment built around the Standard Open Arm SO101 that bridges multiple machine learning libraries—Stable-Baselines3, the imitation library, and Hugging Face’s LeRobot—all in one cohesive environment. The “Sim-First” Decision Even though I had access to physical robots, I chose a sim-first approach for a simple reason: I’m more of a software person who enjoys the comfort of my home office and the flexibility to keep working while traveling, rather than spending long days in a lab. This decision shaped everything about the project. I needed a simulation that was: Physically realistic enough to eventually transfer to hardware Fast enough for thousands of training episodes Flexible enough to work with different learning paradigms Simple enough that I could understand and modify every component Why Build Another Gym Environment? You might wonder: why not just use an existing simulation? I wanted something both flexible and reflecting my hardware platform well in simulation. The robotics world offers many compelling options. The Simulation Landscape I Considered: Isaac Sim was tempting—NVIDIA’s powerhouse with photorealistic rendering and advanced physics. But it requires a proper GPU setup, and I wanted to start experimenting immediately rather than waiting for hardware upgrades. ManiSkill is an exciting newer option with great task diversity and modern ML integration. I’m actually quite excited to try this next—it seems to hit the sweet spot of realism and ease of use. Gazebo/ROS represents the traditional robotics stack: mature, well-supported, with endless plugins. But the learning curve felt steep for someone coming from a pure ML background, and I wanted to focus on learning algorithms rather than robotics middleware. PyBullet similar to MuJoCo. Also popular in RL enviornments. Why I Chose MuJoCo + gym-aloha:The answer came down to immediate productivity. I could adapt gym-aloha and start experimenting within days, not weeks. It’s proven in the papers I was trying to replicate, has excellent contact modeling for manipulation, and enjoys a large ecosystem of compatible tools. The Gym Interface Standard:OpenAI Gym (now Gymnasium) defines a standard interface that every reinforcement learning environment implements: obs, info = env.reset() # Start a new episodeobs, reward, terminated, truncated, info = env.step(action) # Take an action This simple interface is incredibly powerful because it means the same environment can work with: RL libraries like Stable-Baselines3 (SAC, PPO, HER…) Imitation learning libraries like imitation Custom training loops or evaluation pipelines Any future framework that follows the standard The Evolution Plan:This environment is just the beginning. As I move toward more realistic scenarios, I’ll likely migrate to Isaac Sim or ManiSkill. But for rapid prototyping and algorithm comparison, this MuJoCo setup has been perfect. Key insight: There are many good options. If your requirements are not very specific, look for something popular and that has something similar to what you need that you can quickly adapt. This flows naturally from the question and sets up the technical details that follow. Standing on the Shoulders of ALOHA Rather than building from scratch, I adapted the gym-aloha project, which implements the dual-arm ALOHA platform used in several influential imitation learning papers. The ALOHA Foundation:ALOHA (A Low-cost Open Hardware Arm) proved that effective manipulation learning was possible with relatively simple hardware. The gym-aloha implementation provided: MuJoCo physics foundation Well-designed observation and action spaces My Adaptation:I modified gym-aloha for a single SO101 arm (5-DOF + gripper) to match my hardware target: # Simple registration exampleregister( id="gym_so100/SO100CubeToBin-v0", entry_point="gym_so100.env:SO100Env", max_episode_steps=700, nondeterministic=True, kwargs={"obs_type": "so100_pixels_agent_pos", "task": "so100_cube_to_bin"},) Once registered, creating the environment is straightforward: import gym_so100 # triggers env registrationimport gymnasium as gymenv = gym.make("gym_so100/SO100CubeToBin-v0")obs, info = env.reset() The Task: Cube-to-Bin I focused on one fundamental manipulation task: bin-a-cube. A red cube starts at a random position on the table, and the goal is to place it inside a fixed gray bin. This task is deceptively simple but covers the core challenges of manipulation: Perception: Locating the cube and understanding spatial relationships Planning: Approaching the cube from a graspable angle Control: Executing smooth, coordinated motion Manipulation: Grasping, lifting, and precise placement The MuJoCo scene includes: Robot: SO101 single arm with position-controlled actuators Workspace: Table, free-moving cube, and goal bin Sensors: Joint positions, gripper state, and camera views Sites: Tracking points for reward computation and success detection Control Paradigms: Joint vs. End-Effector Space Following gym-aloha’s design, I implemented joint-space control where actions directly specify target joint positions: Aspect Joint-space control Action Target joint positions → data.ctrl Control loop Actuators drive joints toward commanded positions Learning Policy learns in robot’s natural DOF Transfer Direct mapping to real hardware I chose joint-space over end-effector control for cleaner transfer to my target hardware. While end-effector control (where you specify gripper poses and let MuJoCo’s constraints solve for joint angles) can be more intuitive, it adds complexity that I wanted to avoid initially. What’s Inside the Environment As of August 2025, gym-so100-c includes: Simulation Environment & Tasks: MuJoCo-based SO101 simulation environment Cube-to-bin task with configurable reward shaping Training Integration Scripts: Reinforcement learning with Stable-Baselines3 (SAC) Imitation learning with the imitation library Training integration with LeRobot (ACT, Diffusion, VLAs) Control & Data Collection: Teleoperation via keyboard or gamepad Episode recording for demonstration datasets Dataset conversion to LeRobot format Early Design Lessons Physics Engine Choice: MuJoCo was the obvious choice for its speed, stability, and excellent contact modeling. It works well on CPU, it was perfect on MacBook. I am excited about IsaacSIM, but I don’t have a powerful GPU yet. I am open to trying other engines in the future. Observation Space Design: I experimented with different observation combinations: pixels_agent_pos: Camera images + joint positions agent_pos: Joint positions only (for faster training) pixels: Camera images only (for vision-based policies) The mixed approach (pixels_agent_pos) worked best, giving policies both rich visual information and precise proprioceptive feedback.This is also what ACT paper and most SOTA approaches are using. Reward Engineering: In Reinforce Learning an agent collects rewards interacting with the environment and tries to maximize them. The trick is to make such rewards that the agent does what you want from it and also that it can learn pretty well. Shaping the reward structure (designing the incentives) is called Reward Engineering and it’s suprisingly tricky to get right! For example if you reward getting close to the block too much, the agent might chose to hover over it and never grab it. If you don’t give any rewards until the task is complete and if the task is too hard, there is no feedback to learn from. I implemented both sparse rewards (success/failure only) and dense rewards with approach shaping. Dense rewards proved much more reliable for learning, though they required more careful tuning. I also implememented HER (hindsight experience replay) that automates the reward shaping by trying to teach the robot behaviors from previous trajectories, but the results were mixed in my setup. The Integration Challenge The real value of this environment isn’t just the simulation — it’s the integration layer that lets me seamlessly move between different learning approaches. In the next post, I’ll dive into what I learned from training experiments with SAC, behavior cloning, and modern imitation learning methods. But the foundation was crucial: having a single environment that could work with multiple learning paradigms, generate consistent datasets, and provide reliable evaluation metrics. Sometimes the unglamorous infrastructure work is what makes everything else possible. Next up: Training experiments and what I learned about the practical differences between reinforcement learning and imitation learning approaches. What’s your experience with simulation environments for robotics? Have you found certain design decisions that made experimentation much easier or harder? I’d love to hear about it.

13th Aug 2025 • 1 votes
Generating Hundreds of Consistent Illustrations with Gemini Image Generation

A deep dive into building an automated illustration pipeline for storytelling applications Introduction AI image generation has revolutionized creative workflows, but there’s a significant difference between generating a single stunning image and producing hundreds of consistent illustrations for a complete project. When building storylearner.app, we faced the challenge of generating book illustrations that maintained visual consistency while telling compelling stories through imagery. We’ve used powerful Gemini multimodal models, mostly because they are fast and because the experimental ones are available for free. This article explores the technical and creative challenges of large-scale AI illustration generation, showcasing techniques for achieving visual consistency and building robust pipelines that can handle the complexity of full book illustration projects. Below are images created for a chapter of a book adapted for the storylearner platform: The article has a companion colab notebook, so you can play with the examples yourself. The Consistency Challenge Understanding Visual Consistency in AI AI image generation models operate somewhat like a company of talented artists, each with amnesia. Every generation is essentially a fresh start unless you provide explicit context. This creates unique challenges when you need to maintain character consistency, style coherence, and narrative flow across hundreds of images. Consider this simple experiment: generating three images of “the same pig with wings and a top hat flying over a futuristic city” in different weather conditions. Even with identical prompts and seeds, subtle inconsistencies emerge—eye colors change, proportions shift, and the overall character can feel different. The Impact of Seeds and Prompts Seeds act like selecting a specific artist from your AI company. Using the same seed with identical prompts yields consistent results, but even small prompt variations can dramatically alter the output. We discovered that: Same model + same prompt + same seed = identical results Same model + same prompt + different seed = completely different character Same model + slightly altered prompt + same seed = often produces a different character entirely This sensitivity means that scaling up requires careful orchestration of all these variables. Building Consistency Through Reference Images The Reference Image Approach The most reliable method we found for maintaining character consistency involves using reference images. Here’s how it works: Generate an initial character/scene using carefully crafted prompts Upload this image as a reference for subsequent generations Include explicit instructions like “Use the supplied image as a reference for how the pig should look like” This approach significantly improves consistency, though it can also cause new issues in some very specific cases. With the following reference image: You can get this different, but consistent one! Here is an example of triggering a specific case: This specific case can happen when a prompt is pretty similar to the one that generated the initial image and the seed is identical.So, when using the reference image and similar prompt, I would actually recommend to change the seed or not set it. See the reference colab for details. Practical Implementation # Upload reference imagefiles = [client.files.upload(file="reference_character.png")]# Create parts with reference and promptparts = [ types.Part.from_uri( file_uri=files[0].uri, mime_type=files[0].mime_type, ), types.Part.from_text( text=prompt + "\nUse the supplied image as a reference for character appearance" ),]# Generate with referenceresponse = client.models.generate_content( model=IMAGE_MODEL, contents=parts, config=types.GenerateContentConfig( response_modalities=['Text', 'Image'], # No seed set. )) The Storylearner.app Illustration Pipeline High-Level Architecture Our production pipeline consists of three main stages: Idea Generation: Story text + guidelines → 3 illustration concepts per scene Idea Selection: Multiple concepts → best ideas chosen for the complete set Image Generation: Selected ideas + style guidelines + reference images → final illustrations Stage 1: Brainstorming Illustration Ideas Rather than feeding story text directly to image generation (which often produces poor results), we separate conceptualization from execution: class IdeaGenerator: def generate_illustration_ideas(self, text: str, context: str): prompt = """ You are a visual scene designer. Based on the story below, describe 3 different highly detailed and imaginative illustration ideas. Do not include any people or humanoid figures. Focus on setting, atmosphere, lighting, symbolic objects, and environmental storytelling. Story: {text} Context: {context} """ # Returns 3 detailed scene descriptions per text excerpt This approach generates rich, detailed scene descriptions that serve as blueprints for image generation. Stage 2: Intelligent Selection To avoid repetitive illustrations (like “a ship, a ship, a ship” in a sea voyage story), we use an AI selector to choose the best combination of ideas: class IllustrationSelector: def select_illustrations(self, chapter): prompt = """ Choose the best idea for each illustration considering that: - The set should be diverse - Illustrations shouldn't contain people - Prefer illustrations matching the provided titles """ # Returns optimal selection indices Stage 3: Consistent Generation The final generation stage uses: Style guidelines (detailed visual specifications) Reference images for style consistency Persistent chat sessions for maintaining context Retry mechanisms for handling API limitations Visual Guidelines and Style Consistency Crafting Effective Style Guidelines We developed comprehensive style guidelines that go beyond simple style names: Style: watercolorTechnique: Combine soft watercolor washes with fine ink line work for contrast and detail.Brushwork: Embrace visible brush strokes, blooming, and natural texture.Ink Lines: Use varied line weights for depth; apply cross-hatching or stippling for texture.Color Palette: Limit to a few harmonious hues with gentle gradations.Forms: Use simplified, geometric shapes; focus on essence over detail.White Space: Treat negative space as part of the composition.Texture: Highlight watercolor paper's natural texture and color variation.Atmosphere: Create light, airy scenes with openness and subtle contrast.Aesthetic: Preserve a hand-drawn look—embrace imperfections and human touch. Chat-Based Generation for Context Continuity Using persistent chat sessions helps maintain consistency within illustration sets: def generate_set_of_illustrations(ideas_with_file_paths, pass_image=True): chat = client.chats.create( model=IMAGE_MODEL, config=types.GenerateContentConfig(response_modalities=["Text", "Image"]), ) # Initialize with style guidelines and reference image initial_prompt = f""" You are a creative artist helping on an illustration project. Create {len(ideas_with_file_paths)} beautiful illustrations. VISUAL GUIDELINES: {VISUAL_GUIDELINES} """ # Generate each illustration within the same chat context for idea in ideas_with_file_paths: response = chat.send_message(format_illustration_prompt(idea)) # Process and save generated image Avoiding Common Pitfalls Critical Design Decisions Through extensive experimentation, we identified several key strategies: Avoid Human Close-ups: Character face consistency is extremely challenging. Focus on environmental storytelling instead. No Violence or Gore: Keep illustrations family-friendly and avoid content that might trigger safety filters. Diversify Scene Types: The selection stage prevents repetitive imagery across the complete set. Decouple Ideation from Generation: Separating concept creation from image generation improves both quality and debuggability. Real-World Example: Illustrating The Three Musketeers Let’s walk through illustrating a chapter from The Three Musketeers: Context and Settings First, we establish the story context: Setting: France, primarily Meung and Paris, early 17th centuryHistorical Context: Political tensions between French monarchy and Cardinal RichelieuMain Characters: D'Artagnan, Athos, Porthos, Aramis, Cardinal Richelieu, Milady de Winter Generated Ideas For the chapter opening, our system generated these concepts: “The Jolly Miller Inn Chaos”: Exterior scene with a yellow pony, scattered debris, and dramatic lighting hinting at recent altercation “Broken Sword”: Close-up of shattered steel on cobblestones, symbolizing lost honor and broken dreams “Inn Kitchen Aftermath”: Dimly lit interior with earthenware, bandages, and flickering candlelight Selection and Generation The selector chose the most diverse and narratively appropriate ideas, which were then generated using our reference image and style guidelines, producing illustrations that maintain visual consistency while telling the story effectively. Key Insights and Best Practices Consistency vs. Perfection: Perfect consistency isn’t always necessary—visual coherence in style and mood often matters more than exact character matching. The Artist Analogy: Think of AI models as artists with amnesia. You need to provide context, references, and clear instructions for each interaction. Pipeline Modularization: Breaking the process into idea generation, selection, and execution improves quality and maintainability. Style Guidelines Matter: Detailed, specific style descriptions work better than simple style names. Reference Images Are Crucial: Upload and reference style examples for best consistency results. Long Sessions are Fragile: It often works until a point, and at some point it fails poorly, e.g. inserting objects from a previous illustration into the following ones. Conclusion There is still a huge gap between a carefully handcrafted demo on an AI company blog and practical usage of the technology at scale. Things don’t work out well straight out of the box, but you can make these amazing tools work for you with help of systematic thinking about consistency, quality control, and robust engineering practices. While challenges remain, I hope you’ll enjoy the techniques we’ve developed at storylearner.app to power your own projects!

20th Jul 2025 • 1 votes
Escaping LLM piping mess with nifty engineering

In this post I’ll walk through how I upgraded a set of tangled Python notebooks—responsible for thousands of LLM calls—into a robust content-adaptation studio powered by: an async FastAPI pipeline, a disk-first Next.js frontend, and a small suite of custom CLI tools. Re-engineering the stack was essential for my own sanity: the notebooks were fragile, slow to iterate on, and far too labor-intensive to babysit. I was also facing content quality challenges that were pretty much impossible to address in the old code base, that re-engineering unlocked. My hope is that the story also nudges you to build (or level-up) your own tooling instead of settling for one-off notebooks. We’ll cover: Engineering constraints – huge text volumes, strict meaning preservation, multiple target languages. The original notebook setup – what worked and where it hurt. Pain points – why small hacks no longer cut it. The new architecture – key design choices, novel elements (with screenshots), and how they solve the earlier pain. Outcomes & takeaways – higher quality, less toil, faster experiments, and patterns you can reuse in your own LLM workflows. Shape of the problem Shape of the problem: large quantities of text that need to be processed in a very specific way: retain meaning of the original (no text disappearing or altered significantly) consistent between parts translated & simplified resulting content needs to consistently have good quality automatic quality assessment and quality repair (revisions) A similar challenge would be relevant in translating legal documents, healthcare documents, etc. Story Learner Book Adaptation needs In my project StoryLearner I offer adapted books for language learning at a specific level. For example, “Las Aventuras de Sherlock Holmes, in A2, Spanish”. We use LLMs for both language/level adaptation as well as for illustrations.It’s a lot of LLM calls (easily thousands for a single book), because books are long and there are many elements for a single adaptation. A book has chapters, chapters have parts for easier reading. Each book/chapter/page has an custom illustration. Additionally, there are titles and descriptions to be adapted and transcribed. Books are long and we want the adapted text to retain the meaning, while making the language simple (aligned with the target level), natural sounding and correct. Trouble with a flaky, slow pipeline and hard to assess output The pipeline was was quite a feat! It was a lot of LLM calls, built mostly in colab/jupyter notebooks. RAW BOOK | v[book_stripping] | v[chapter_extraction] | v[chapter_simplification] (English) | v[chapter_partification] (English) | \ | \ | ---> [illustration_generation] | v[adapt] ──▶ [Lang 1] │ [Lang 2] │ [Lang 3] │ [Lang 4] │ [Lang 5] │ [Lang 6] | Illustration generation was on its own a pretty interesting pipeline (more about it in a separate post!). Here is one of the adaptation notebooks: As you can see, it has its own table of contents on the side. It’s easily thousands of lines of code and prompts. And the hundreds of outputs (text and images) could make it very, very, long. To the point that it would have rendering issues. On top of it there was: Separate notebook for book_narration Another notebook for upload to storylearner (via API). The “pipeline” worked. It serialized partial outputs in a way that partial redos/continuations were possible, it offered decent visualization. It was adjustable (just add/tweak a notebook cell!). However it was fragile and assessing quality/redoing content was painful and slow. And it was very frustrating to me, especially since the quality was important to my partners (language schools). Main pain points LLM reliability issues (no resources, surprise safety controls kicking in, running out of quota or LLMs not following instructions) broke downstream steps. The adaptation process was slow, it wasn’t taking advantage of paralellization well The pipeline was already so brittle enough that meaningful experiments were nearly impossible Colab/notebooks encouraged slapping things together instead of proper engineering with encapsulation and tests Python notebooks having rendering bugs because they were so long and had so many outputs (large text/many images) Low confidence in language level, name/format consistency, and preserved meaning without having a strict review/repair process and having some examples of problems with quality. Reviewing 60+ chapters across six languages was slow and manual - was infeasible for me. A copy of an adaptation notebook per book (for visualization/auditability) was duplicating the code and making it harder to maintain New Content Adaptation Tools I haven’t built everything at once. I started with a frontend using the existing disk format of book adaptations, then as it was easier to see what was going on, I progressively built more and more backend migrating specific functionalities. Fast API gives a nice UI out of the bat to call APIs, which was nice for trying things out. But the workflows were simply to long to drive them by hand, so that is how CLI tools came to be. CLIs/Frontend were in big part written by copilot coding agent. I also extensively discussed the component prompts with LLMs 😀. Backend – bookadaptation (FastAPI) all functionality behind endpoints, all async / internally parallel where safe (28 separate endpoints) All IO and LLM calls done async Pipeline stages are classes; 11 adapter subclasses (Gemini Flash, Gemini Pro, GPT-4o, etc.). skip and use_cached flags run a no-op or reuse artefacts while the file tree stays unchanged. STAGE_DEPENDENCIES mapping declares primary & secondary inputs and outputs for every stage. Hierarchical on-disk structure; files get a _{revision_number} suffix for multi-round outputs. Heavy use of controlled generation (using schemas) SQL instrumentation: All prompts, schemas, settings, outputs, latency, and errors logged to SQLite. Frontend – ContentTools (Next.js Server Components) Specialised views: Book language adaptations Illustrate all artefacts of the adaptation pipeline (inputs, outpus) Easy debugging of what happened during the QA Easy comparisons between experimental implementations “Final for publish” view. Illustrations: All book illustrations as grids (chapters, chapter parts) Illustration deep dives - ideas and the best ideas Reads JSON & WebP directly from disk—no extra HTTP hop. CLI tools (all async friendly python) adaptation CLI Adapts and revises every chapter until it’s good enough Drives the pipeline stages implemented in the service (many endpoints right) Gracefully retries LLM driven decisions: Uses output from the overall review to finish the chapter or go for more rounds or partially skip stages narration CLI illustration CLI publishing CLI Everything runs locally, but could as well run on a server. Yes, I ended up building a lightweight custom pipeline orchestration… 💀 Flow of a book adaptation with automatic QA (multiple revisions) RAW BOOK | v[book_stripping] | v[chapter_extraction] | v[chapter_simplification] (English) | v[chapter_partification] (English) | \ | \ | ---> [illustration_generation] | v[adapt] ──▶ [Lang 1] │ [Lang 2] │ [Lang 3] │ [Lang 4] │ [Lang 5] │ [Lang 6] | v[review_chapter] ←──────────────┐ | │ v │[revise_chapter] │ | │ v │[review_consistency] │ | │ v │[revise_consistency] │ | │ v │[review_meaning_cohesion] │ | │ v │[revise_meaning_cohesion] │ | │ v │[review_titles] │ | │ v │[revise_titles] │ | │ v │[review_chapter_title] │ | │ v │[revise_chapter_title] │ | │ v │[review_overall] ──────────────┘ | v [promote_content] | | +--> [book_narration] (from adapted text) | +--> [upload / publish] QA rounds are controlled by the output of the review_overall stage. Examples of frontend enabling fast QA Example of problematic images that can be ’easily spotted’ by a trained eye. Chapter level issues overview for an adaptation workflow: Part level debugging/review of what was suggested/applied in the QA pipeline: Review and Revision—kept deliberately apart One key design choice was to decouple “reviewer” from “reviser”. The reviewer node reads the necessary context and produces a structured list of issues,while the reviser node sees only the affected slice plus those suggestions. Why keep them separate? Audit clarity The reviewer’s JSON lives as its own artefact, so you can diff, grep, or hand-editthe feedback without touching the text itself. It’s also easier to audit automatic revisions and spot ‘additional helpfulness’. Smaller prompts, cheaper calls A reviser that operates on just the target part + suggestions uses far fewer tokensthan one that re-ingests the whole chapter. Less collateral damage Narrow context means the reviser can’t “helpfully” rewrite good paragraphs in other sections or even just completely forget them. True parallelism Parts are context-isolated, so multiple revisions can run concurrently—no giant chapter-wide lock. Targeted rollbacks If a revision introduces a new issue, it’s easy to rollback. What the flow looks like review_meaning_cohesion scans consistency_revised_stories and writes meaning_cohesion_review_1.json: { "issues": [ {{ "part_number": 1, "suggestion": "Change sentence '...sentence...' to '...corrected sentence...' to ensure consistency with previous parts.", "reason": "Inconsistent character name across parts.", "severity": "high" }}, ... ], } During the revise_meaning_cohesion stage, revisions to specific parts are applied concurrently, e.g. reviser for part 8 only sees the text of part 8 and the suggestions for part 8. Other novel elements Stable DAG via no-op nodes – skipping or caching never changes filenames or dependencies, so UI and controller logic stay simple. Declarative STAGE_DEPENDENCIES – each endpoint validates its own inputs and fails fast if artefacts are missing. Pluggable adapter subclasses – swapping models or prompt strategies is a config change, not a refactor. Direct-disk reading Server Side React Components – Suprisingly trivial frontend code. SQLite error forensics – a single query surfaces “prohibited-content” or other LLM failures. Hot-reload mid-run – tweak prompts or error handling while a 60-chapter fan-out is running; retries pick up the change without restart. WebP illustration storage – generated art compresses very well; files are roughly 10× smaller than raw outputs. LLM-driven controller decisions – the CLI uses the review_overall output to decide whether to launch another revision round and which stages to skip. Personal wins The biggest win is defending my personal sanity. I no longer have to babysit a set of fragile notebook based pipelines based on fallible and untrustworthy LLMs. Higher confidence in quality – every chapter passes level, consistency, meaning, and overall reviews—automatic multi-round fixes if needed. Far less toil – no more babysitting fragile notebooks; the pipeline self-checks and fails early. Fast, parallel experimentation – new adapters or prompts run side-by-side with hot-reload; iteration is “super fast and fun.” Cost flexibility – total spend is higher (as there are more LLM calls), but the modular design lets me fall back to cheaper or local models whenever I choose. And the cool thing is that building this tooling was heavily accelerated by a coding assistant/agent, so it was significantly faster and more fun than I would have expected from the scope of the reengineering.

25th Jun 2025 • 1 votes

More in technology

Inside a 1980s filter chip that uses switched capacitors

Sometimes it's easier to identify an IC with a microscope. While sorting a box of old ICs, CuriousMarc came across some Harris ICs labeled "F1-10-5", a mysterious part number that didn't show up in any databooks. Since unidentifiable ICs are useless, he gave me one to analyze. Conveniently, it was in a ceramic package, so I could open it up with a quick tap from a chisel. Under the microscope, the chip's most striking feature was a grid of square capacitors. With all those capacitors, I guessed that it was a switched-capacitor filter. The die provided another clue: the part number HF-10. With this information, we quickly found that the chip was Harris's version of the standard MF10 switched-capacitor filter chip.1 The Harris integrated circuit, labeled F1-10-5 (or maybe FI-10-5), with a 1985 date code. Photo courtesy of CuriousMarc. Switched-capacitor filters were a popular way to implement analog filters in the 1980s. Rapidly switching capacitors in and out of a circuit enabled the construction of single-chip filters that were easy to use and performed well. The MF10, introduced by National Semiconductor in 1981, provides two flexible filters on a chip; each filter acts as a low-pass filter, band-pass filter, or a high-pass filter. The filter's characteristics are simple to control with a few external resistors. The Harris HF-10 die under the microscope with the main functional blocks labeled. (Click for a larger image.) Since I had the chip under the microscope, I took the opportunity to analyze it more closely. The white lines are the metal wiring that connects the chip's circuitry. Under the metal layer are two layers of polysilicon (reddish) and the underlying silicon (gray). The top and bottom halves of the chip are mostly mirror images, corresponding to the chip's two filters. The distinctive reddish squares in the middle of the chip are 72 tiny capacitors, constructed from polysilicon. Above the capacitors, CMOS switches turn on and off at the clock frequency, switching capacitors in and out of the circuit. Each filter uses three operational amplifiers (op amps), outlined in red. At the right are the three outputs from the three op amps: high pass, band pass, and low pass. The control circuitry is on the left: clock level shifting, clock shaping, frequency ratio handling, startup circuitry, and current sinks to provide fixed currents to other parts of the chip. Around the edges of the silicon die, 20 hair-thin bond wires connect the die to its 20 external pins. The die has some interesting chip art: a Harris logo and an outline of Florida; Harris was headquartered in Melbourne, Florida. The initials on the die are presumably the engineers who designed the chip. Some interesting images from the die. Switched capacitor circuits The filter is based on switched-capacitor circuits. A switched capacitor can replace a resistor in certain circuits, as shown below. The switches are controlled by a clock signal; the switches alternately close in clock phase 1 and phase 2 (ϕ1 and ϕ2). In phase 1, the capacitor is charged to the input voltage. In phase 2, the capacitor passes charge to the output. By rapidly toggling the switches, charge is (almost) steadily passed to the output. The larger the capacitance, the more charge that is passed through. Likewise, a higher frequency passes more charge. It can be shown that the circuit matches a resistor with resistance of 1/(fC): a higher capacitance and frequency correspond to lower resistance. A switched capacitor can replace a resistor. Why would you replace a simple resistor with this complicated switching circuit? In an integrated circuit, resistors are inaccurate and inconveniently large, especially high-value resistors. Replacing a large resistor with a small capacitor saves space on the die. Moreover, it is easy to generate an extremely accurate clock frequency with an inexpensive quartz crystal, making the filter's frequency highly accurate. Finally, the equivalent resistance can be changed simply by changing the clock frequency, making it easy to tune or sweep the filter. On-chip capacitors are fairly inaccurate, with the capacitance typically varying by 20% from chip to chip due to variations in manufacturing conditions. However, this isn't a problem in the MF10 because the circuitry was designed to depend on the ratio between capacitances, which is stable. Specifically, the MF10 uses 72 identical square capacitors, which will have almost identical capacitances. Careful examination shows that some of the capacitors are separate, while others are connected in groups of 8 to form larger capacitors.2 This yields a highly accurate ratio of 8:1 between the grouped capacitors and the individual capacitors, even though the absolute capacitance will vary from chip to chip. Each capacitor is constructed from two layers of polysilicon,3 forming the plates of the capacitor, separated by a thin layer of insulating oxide that acts as the dielectric. I estimate that each capacitor square is 5 picofarads. The grid of capacitors in the MF10. I've added yellow lines to show how the capacitors are grouped. The switches are above and below the capacitors. This chip uses one more trick with switched capacitors: it inverts the voltage while acting as a resistor. In the switched-capacitor circuit below, there are four switches. The capacitor charges to the input voltage during phase 1, the same as before. But duing phase 2, note that the top plate of the capacitor is grounded, while the output comes from the bottom plate. If the capacitor was charged to, say, 1 volt, the top plate is 1 volt above the bottom plate. So if the top plate is grounded, then the bottom plate must be at -1 V. (This is the same idea as a charge pump.) This circuit turns out to yield a more accurate filter because some parasitic capacitances cancel out. By using four switches, the switched capacitor can invert the voltage. The op-amp integrator The heart of most analog circuits is the operational amplifier, or op-amp. An op-amp takes two inputs and amplifies the difference by many orders of magnitude. Normally, an op-amp is configured with negative feedback, which forces the two inputs to be essentially the same. Op-amps are useful not only for amplification, but for filtering, buffering, summing, and other tasks. A basic op-amp integrator. The filter chip uses op-amps as integrators, to integrate an input voltage over time. The circuit above shows a simple op-amp integrator. The input voltage produces a current that flows through the resistor and charges the capacitor, so the capacitor holds the integral of the input voltage over time. You might expect that the left side of the capacitor would become positive as it charges. However, the op-amp's feedback forces both inputs to ground, so instead the right side of the capacitor becomes negative. Thus, the output is the negative integral.4 The MF10 chip uses the circuit above, except the resistor is replaced with a switched capacitor. The capacitor across the op-amp is not switched, but consists of either 8 or 16 capacitors from the capacitor grid. The CMOS switches The CMOS switch is the technology that makes the switched-capacitor filter possible. A CMOS switch has a fairly low resistance (maybe tens of ohms) when closed and an enormously high resistance (hundreds of megohms) when open. This high resistance ensures that the charge doesn't leak out of the capacitors. A CMOS switch is constructed by combining an NMOS transistor and a PMOS transistor. The NMOS transistor and PMOS transistor are opposites. An NMOS transistor is good at pulling the output low, while a PMOS transistor is good at pulling the output high, so in combination they provide an effective switch. An NMOS transistor is turned on by a high voltage on the gate, while a PMOS transistor is turned on by a low voltage on the gate. Thus, a CMOS switch requires two control signals of opposite polarity, which is a minor inconvenience. A CMOS switch. The diagram above shows how a switch is implemented with an NMOS transistor and a PMOS transistor in parallel. When the control line is high, and the inverted control line is low, both transistors turn on, providing a path through the switch circuit. When the control line is low (and the inverted line high), the transistors turn off, opening the switch. The chip uses CMOS switches in pairs, with one switch on and the other off. This forms the equivalent of a toggle switch that connects either A or B to the output. This circuit is simply two CMOS switches, with separate control lines for each switch, as shown below. In the MF10, the switch toggles at the clock frequency. During one clock phase, the switch is connected to A, while the switch is connected to B during the other clock phase. The schematic on the right, below, is the same circuit, but reorganized to match the layout on the die. A double-throw CMOS switch. The photo below shows a CMOS switch on the die, constructed from two PMOS transistors and two NMOS transistors. The four control lines run horizontally in polysilicon, forming a transistor gate where they cross doped silicon. The upper PMOS and NMOS transistors are driven by the clock phase 1 (Φ1) signals, while the lower transistors are driven by the phase 2 signals. CMOS switches on the die. The metal layer was removed to show the transistors. One problem with switched-capacitor filters is that the clock can generate switching noise that appears in the chip's outputs. The MF10 uses several techniques to reduce clock noise. Each set of transistors is surrounded by two isolation rings: one positive and one negative. These block noise from traveling through the silicon substrate. Note that the rings have opposite polarity for the NMOS transistors and the PMOS transistors. The light tan region in the photo above is a second layer of polysilicon. This polysilicon is connected to ground, providing a shield layer over the switching circuits. For the photo above, I removed the metal layer with acid5 to make the transistors more visible. The photo below shows the original die, with the metal layer connecting the transistors. The small black circles are connections between the metal layer and silicon or polysilicon. The same CMOS switches, showing the metal layer. Putting it together: the state variable filter There are many ways of creating a filter. The MF10 chip uses a technique called the state variable filter, invented in 1967. This circuit acts as three filters, with high-pass, band-pass, and low-pass outputs. Moreover, the circuit is flexible since the frequency, the gain, and the filter quality (Q) can be varied independently. It uses three op-amps: one to sum signals and two for integration. By changing how the values are summed, the characteristics of the filters can be changed. The diagram below shows a simplified representation of a state variable filter. The mathematics behind a state variable filter is complicated, so I won't get into it. In short, the signal, the integral, and the double integral form the three state variables that define the state of the system. Simplified diagram of a state variable filter, with two integrators. Inspired by North Coast Synthesis. The block diagram below shows how the filter is represented in the MF10 datasheet.6 The diagram is similar to the diagram above, with three op-amps. However, the summing circuitry has been separated out. Moreover, the feedback paths are not shown explictly. Instead, resistors are connected between the chip's external pins (squares) to configure the filter as desired. The mode switch at the top allows the low-pass feedback to be controlled by an external pin (SA/B). Block diagram of one of the filter sections. Adapted from the datasheet. The schematic below is my reverse-engineered schematic of the filter, as implemented on the chip. It closely matches the block diagram, but fills in the details. In the block diagram, the summing circuit (circle) adds one signal and subtracts two signals. This summing circuit is implemented with the three switched capacitors on the left, which act as summing resistors. Note that one switch is grounded during phase 1, while the others are grounded during phase 2; switching the polarity implements addition versus subtraction. The top sum input is either feedback from the low-pass output or ground, selected by an input pin. A CMOS switch is used here, but the switch is static, not clocked, so it doesn't use protection rings and shielding like the other switches. My reverse-engineered schematic of one of the filters. Click this image (or any other) for a larger version. The integrators have switched capacitors on the inputs, acting as resistors. The integration capacitor is either 8 or 16 "squares" of capacitance, selected by a ratio selection pin. This controls the ratio between the clock frequency and the filter frequency, either 50:1 or 100:1.7 Although the integration capacitors are attached to a CMOS switch, the switch is static, so the capacitors act as regular capacitors, not switched capacitors. The op-amps The op-amps are fairly standard CMOS op-amps, built from about 35 transistors. (You might get a lower count if you try counting the transistors below, since some of the blocks are multiple transistors.) The op-amp transistors are much larger than the CMOS switch transistors (very bottom, center). On the die, each op-amp is split into two parts: the differential amplifier on the left and an additional amplification stage on the right. A large capacitor (pinkish) sits between the halves. My first thought was that this was the integration capacitor, but it is just a frequency compensation capacitor, common in many op-amps to stabilize the output. The op-amps also have large transistors next to the output pins; these transistors are functionally part of the op-amps, but located next to the pins to minimize resistance. One of the chip's op-amps. I removed the metal layer to make the transistors visible. One unusual feature of the op-amps is a low-power mode. Pulling a particular IC pin low causes the chip to stop filtering and enter a low-power mode, reducing power consumption by 70%. This is implemented by shutting down the "current mirror" circuits that provide fixed currents to the op-amps and other parts of the chip. The non-overlapping clock generator The MF10 chip is driven by external clock signals, one for each filter, with the frequency of the filter proportional to the clock frequency. The photo of the CMOS switches earlier showed that the clock drives four control lines for the switches. You might think that two control lines would be sufficient: the clock and the inverted clock. The problem is that it is very important to avoid having both switches closed at the same time, even for a moment, as that will short the inputs and corrupt the signals. Instead, the two switches have separate control lines that enforce a small gap between when one switch opens and the other one closes. This is implemented with the circuit below that takes an input clock signal and produces the four outputs that drive the switches. The circuit to generate non-overlapping clock signals. There is a delay between when gate A or B turns on and when the corresponding output changes. The idea behind the circuit is that a phase is blocked from going high until after the other phase goes low, with a pair of inverters providing additional delay. In more detail, suppose the input clock drops from high to low. Gate A will turn off, causing the phase 1 output (ϕ1) to drop after a few gate delays (A delay). Gate B can't turn on until ϕ1 goes low. After additional gate delays, ϕ2 goes high. The behavior is similar when the input clock goes high. Gate B turns off, causing ϕ2 to go low after a delay. This allows gate A to turn on, turning on ϕ1 after more delay. To summarize, after a phase is turned off, there is a delay before the other phase turns on, so the two phases never overlap. The clock-shaping circuitry is implemented with CMOS logic gates. The photo above shows this circuitry under the microscope, with the metal layer removed. The rectangular blocks are doped silicon that forms transistors. The darker regions on the left are NMOS transistors and the lighter regions on the right are PMOS transistors. A CMOS gate consists of NMOS and PMOS transistors working together. The PMOS transistors are larger because PMOS transistors are slightly less efficient than NMOS transistors. The dark circles are contacts between the silicon and the metal layer on top. The copper-colored lines are not metal but a special type of silicon called polysilicon. When a polysilicon line crosses doped silicon, it forms the gate of a transistor. The pinks and greens are due to thin-film interference from a thin layer of oxide that didn't completely dissolve; the silicon is actually gray. The ternary input A weird feature of the chip is the input pin that selects the ratio between the input clock and the filter frequency. In effect, this is a digital input with three values. Tying the pin to the high supply voltage selects a 50:1 ratio. Tying the pin to the midpoint between the supply voltages selects a 100:1 ratio. Pulling the pin to the low supply voltage stops the filter and puts the chip into a low-power mode.8 To handle the three-level input, the input goes through two separate buffers, one that transitions at a lower voltage and one that transitions at a higher voltage. Thus, the two buffers separate the middle signal level. Each buffer consists of a special inverter feeding into a regular inverter. Before explaining the special inverters, I'll review how a regular CMOS inverter works. A CMOS inverter is constructed from a PMOS transistor and an NMOS transistor. When the input is high, the NMOS transistor turns on and pulls the output to ground. When the input is low, the PMOS transistor turns on and pulls the output high. Thus, the input signal is inverted. A CMOS inverter is constructed from a PMOS transistor and an NMOS transistor. In the die photo, you can see the four PMOS transistors (light gray) and four NMOS transistors (darker), forming four inverters. When a polysilicon line (copper-colored) crosses a doped silicon region, it forms the gate of a transistor. For this picture, I dissolved the metal layer in acid so the transistors are visible. The metal layer connected the transistors to complete the wiring of the inverters: it connects the two "out1" contacts to "in2" and connects the two "out2" contacts to the rest of the chip. For the second buffer, "out3" connects to "in4" and so forth. The four inverters that handle the ternary input. I flipped the image to make the orientation better. In this circuit, the length of the transistor gates is varied to make the inverters activate at different voltage levels. Six of the transistor gates are normal (orange arrows); the PMOS gates are wider (in the vertical direction) than the NMOS gates because PMOS transistors are inherently weaker. However, two of the transistor gates are unusually long (horizontal direction, red), making the transistors weak since the current must travel a longer distance. The inverter on the left has a weak PMOS transistor. If the input is high or low, the inverter will operate normally. But if the input is in the middle, both transistors will partially turn on. Since the PMOS transistor is very weak, the NMOS transistor will "win", pulling the output low. Thus, the leftmost inverter treats a medium-level input as a 1, outputting a 0. The third inverter is the opposite; the NMOS transistor has a long, winding gate, so it is weak. In this case, a medium-level input will partially turn on both transistors, but the PMOS transistor will "win", pulling the output high. To summarize, the two inverters have opposite behavior for a middle-level signal, allowing the three input levels to be distinguished. Since the output from a special inverter may be weak, the output goes to a normal inverter to amplify the signal. Conclusions Like most semiconductor companies, Harris has a complicated history. Harris started way back in 1895 as a printing press company. Harris moved into high technology in the 1950s and 1960s, acquiring various radio and electronics companies. In particular, Harris entered the IC business in 1967, when it acquired Radiation, Inc., renaming it Harris Semiconductor a few years later. (We've encountered some Radiation modules in Apollo systems, but I haven't written about them yet.) Harris got out of the semiconductor business in 1999, spinning off Intersil, which was later acquired by the Japanese semiconductor firm Renesas. In 2019, Harris merged with L3 Technologies to become L3Harris, the eighth-largest defense contractor in the US. As for switched-capacitor filters, they have lost popularity as filtering is now more easily done in the digital domain. Texas Instruments acquired National Semiconductor (and the MF10) in 2011; TI's website shows the MF10 as active but expensive and out of stock, so it's probably no longer being manufactured. State variable filters are still used in the synthesizer world both because of their flexibility and because they provide low-pass, band-pass, and high-pass filters in one unit. For more, follow me on Bluesky (@righto.com), Mastodon (@[email protected]), or RSS. Thanks to CuriousMarc for providing the IC. AI statement: Despite the presence of the em dash, no AI was used in the writing of this article (details). Notes and references Once we found the "HF-10" part number, a search turned up a National Semiconductor databook that confirmed that the Harris HF-10 was a direct replacement for the National Semiconductor MF10. It remains a mystery why the Harris chip is externally labeled "F1-10-5" rather than "HF-10". This format doesn't resemble other Harris part numbers. I would suspect a military part number, but it is completely different from the military formats that I've seen on other chips, such as JM38510 numbers or NSN numbers. ↩ You might wonder why the larger capacitors are formed by connecting eight smaller capacitor squares, rather than making one capacitor that is eight times as big. The reason is to get better matching between the two capacitor sizes. A capacitor that is eight times as large won't have exactly eight times the capacitance due to factors such as the behavior of the electric field around the edge of the capacitor, inaccuracies that may make the capacitor slightly larger or smaller than desired, or etching variability around the edges. By building larger capacitors out of identical smaller capacitors, the values can match very well, up to ±0.01% according to The Art of Analog Layout. (With laser trimming, matching of ±0.001% is possible, but that is much more accuracy than the MF10 required.) ↩ Most chips from this era have a single layer of polysilicon, so I was surprised to find two layers in this chip. I've seen two layers of polysilicon before, in the MK4116 DRAM chip and AMD's LANCE Ethernet chip. In both cases, the second layer of polysilicon was used for storage devices. ↩ A standard op-amp integrator is an inverting integrator, and the output is negative. However, the MF10 uses the four-switch switched capacitor that inverts the input voltage. The two negatives cancel out, so the MF-10's integrator is a non-inverting integrator. See Introducing the MF10: A Versatile Monolithic Active Filter Building Block for details. ↩ To remove the metal layer, I used Whink rust stain remover (1.5-3.5% HF) to remove the oxide layer and hydrochloric acid to dissolve the metal. I applied Whink for 20 minutes and HCl for 16 minutes in total. I alternated each chemical for about 3 minutes each, applying a few drops at a time. I examined the die under the microscope after each application to gauge the progress. I stopped at this point since the metal was removed and the underlying transistors were visible. Moreover, the silicon became differentially stained, with NMOS transistors significantly darker than PMOS transistors. Some more Whink would probably improve the appearance of the die, but the risk is that the polysilicon might get removed, which would be bad for reverse engineering. In other words, I'd rather stop too early than destroy the features that I want to see. ↩ For reference, the full block diagram of the chip is below, from the datasheet. Block diagram of the MF10 from the Texas Instruments datasheet.  ↩ The filter frequency of the MF10 can be set to either the clock frequency divided by 50 or divided by 100. You might wonder where these ratios come from, since the capacitors on the chip are in 8:1 or 16:1 ratios, not 50:1 or 100:1. The formula for a switched-capacitor integrator is that the filter frequency is the clock frequency divided by 2π times the capacitor ratio. (This can be derived from the op-amp integrator formula and the equivalent resistance of a switched capacitor.) It turns out 2π×8 is 50.27 and 2π×16 is 100.5, providing the 50 and 100 values. Note that these values aren't exactly 50 and 100; they are off by 0.5%. Curiously, the datasheet specifies that the typical frequency error is ±0.2%, significantly smaller. I suspect that the explanation is that the capacitor ratio is not precisely 16:1, due to stray capacitance in the wiring and other factors, and the designers ensured that these factors tweaked the ratio in the desired direction. ↩ I suspect that the ternary input pin was used because the chip didn't have enough physical pins for all the functions they wanted. Note that the two filters are entirely independent, even with separate clocks, except for the 50/100 ratio control and the A/B mode control. I'm sure that these two functions would have independent control pins if the chip had pins available. They could have used a standard 24-pin package for the chip rather than the somewhat unusual 20-pin package, but maybe they had a motivation for avoiding a much larger 24-pin package. ↩

39 minutes ago • 1 votes
Radxa's Q8B has 2x the performance and expansion of the Pi 5

There was a time I'd look at a board like the Radxa Dragon Q8B (at left, above) and be like, "there's no way I'd spend $209 on an SBC with 8 gigs of RAM". But we're in 2026, and seeing the 8 gig Raspberry Pi 5 going for almost the same amount, I figured I'd give it a shot. On paper, the Q8B beats the Pi 5 in pretty much every way. A lot of that is thanks to this Snapdragon 8cx Gen 3 chip, which is the same chip I tested on Microsoft's Windows Dev Kit 2023.

21 hours ago • 1 votes
Three years later

Reflections on October 7th

2 days ago • 1 votes
The Sting

The Sting belongs in the pantheon of films I'm deeply embarrassed to have not watched earlier. Not just because it's a great film — and it is — but because it is so incredibly my shit that I feel retroactively spurned for not having watched it sooner.

2 days ago • 1 votes
It's a Gas!

If everything worked as well as the product called Evapo-Rust, the world would be a much better place. That’s just one of the many lessons learned during my recent — successful! — project to transform my old, nonfunctioning gasoline-powered generator into something much better.

3 days ago • 1 votes
📚 BoredReading

You seem to be enjoying this.

Join free to unlock everything.

Create free account

Already have an account? Sign in