Full Width [alt+shift+f] Shortcuts [alt+shift+k]
Sign Up [alt+shift+s] Log In [alt+shift+l]
1

Dataset Engineering: The Art and Science of Data Preparation

from Alex Strick van Linschoten [alt+shift+b] in AI

Finally back on track and reading the next chapter of Chip Huyen’s book, ‘AI Engineering’. Here are my notes on the chapter. Overview and Core Philosophy “Data will be mostly just toil, tears and sweat.” This is how we start the chapter :) This candid assessment frames dataset engineering as a discipline that requires both technical sophistication and pragmatic persistence. While the chapter’s placement might have been suitable earlier in the book, its position allows it to build effectively on previously established concepts. Data Curation: The Foundation Data curation addresses various use cases including fine-tuning, pre-training, and training from scratch, with specific considerations for chain of thought reasoning and tool use. The process addresses three fundamental aspects: Data Quality: The equivalent of ingredient quality in cooking Data Coverage: Analogous to having the right mix of ingredients Data Quantity: Determining the optimal volume of ingredients Quality Criteria Data quality encompasses multiple dimensions: Relevance to task requirements Consistency in format and structure Sufficient uniqueness Regulatory compliance (especially critical in regulated industries) Coverage Considerations Coverage involves strategic decisions about data proportions: Large language models often utilize significant code data (up to 50%) in training, which appears to enhance logical reasoning capabilities beyond just coding Language distribution can be surprisingly efficient (even 1% representation of a language can enable meaningful capabilities) Training proportions may vary across different stages of the training process Quantity and Optimization A key phenomenon discussed is ossification, where extensive pre-training can effectively freeze model weights, potentially hampering fine-tuning adaptability. This effect is particularly pronounced in smaller models. Key quantity considerations include: Task complexity correlation with data requirements Base model performance...
4th Feb 2025

Stay updated

Get a weekly newsletter with the top 5 articles worth reading every week.

More from Alex Strick van Linschoten

Testing out instrumenting LLM tracing for litellm with Braintrust and Langfuse

I previously tried (and failed) to setup LLM tracing for hinbox using Arize Phoenix and litellm. Since this is sort of a priority for being able to follow along with the Hamel / Shreya evals course with my practical application, I’ll take another stab using a tool with which I’m familiar: Braintrust. Let’s start simple and then if it works the way we want we can set things up for hinbox as well. Simple Braintrust tracing with litellm callbacks Callbacks are listed in the litellm docs as the way to do tracing with Braintrust. So we can do something like this: import litellm litellm.callbacks = ["braintrust"] completion_response = litellm.completion( model="openrouter/google/gemma-3n-e4b-it:free", messages=[ { "content": "What's the capital of China? Just give me the name.", "role": "user", } ], metadata={ # "project_id": "1235-a70e-4571-abcd-234235", "project_name": "hinbox", }, ) print(completion_response.choices[0].message.content) You can pass in a project_id or a project_name and the traces will be routed there. Here’s what it looks like in the Braintrust dashboard: Our first trace logged in Braintrust Note how you can’t see which model was used for the LLM call, nor any cost estimates. The docs mention that you can pass metadata into Braintrust using the metadata property: “braintrust_* - any metadata field starting with braintrust_ will be passed as metadata to the logging request” (link) This seems a bit rudimentary, however. If we take a look at the full tracing documentation on the Braintrust docs we can see that they seem to recommend wrapping the OpenAI client object instead: import os from braintrust import init_logger, traced, wrap_openai from openai import OpenAI logger = init_logger(project="hinbox") client = wrap_openai(OpenAI(api_key=os.environ["OPENAI_API_KEY"])) # @traced automatically logs the input (args) and output (return value) # of this function to a span. To ensure the span is named `answer_question`, # you should name the function `answer_question`. @traced def answer_question(body: str) -> str: prompt = [ {"role": "system", "content": "You are a helpful assistant."}, {"role": "user", "content": body}, ] result = client.chat.completions.create( model="gpt-3.5-turbo", messages=prompt, ) return result.choices[0].message.content def main(): input_text = "What's the capital of China? Just give me the name." result = answer_question(input_text) print(result) if __name__ == "__main__": main() This indeed does label the span as answer_question but it doesn’t do much else. Even the model name isn’t captured here. Instrumenting a series of calls to handle ‘deeply nested code’ (as their docs puts it) even didn’t log the things it was supposed to: import os import random from braintrust import current_span, init_logger, start_span, traced, wrap_openai from openai import OpenAI logger = init_logger(project="hinbox") client = wrap_openai(OpenAI(api_key=os.environ["OPENAI_API_KEY"])) @traced def run_llm(input): model = "gpt-4o" if random.random() > 0.5 else "gpt-4o-mini" result = client.chat.completions.create( model=model, messages=[{"role": "user", "content": input}] ) current_span().log(metadata={"randomModel": model}) return result.choices[0].message.content @traced def some_logic(input): return run_llm("You are a magical wizard. Answer the following question: " + input) def simple_handler(input_text: str): with start_span() as span: output = some_logic(input_text) span.log(input=input_text, output=output, metadata=dict(user_id="test_user")) print(output) if __name__ == "__main__": question = "What's the capital of China? Just give me the name." simple_handler(question) This is adapted from the example they pasted in their docs as their one isn’t even a functional code example on its own. It is seeming increasingly clear that Braintrust isn’t going to be the right choice, at least as long as I want to keep using litellm. I know that Langfuse has a very nice integration with litellm, so I think I’ll pivot over to that now. Basic tracing with Langfuse and litellm Simple tracing is easy: import litellm litellm.callbacks = ["langfuse"] def query_llm(prompt: str): completion_response = litellm.completion( model="openrouter/google/gemma-3n-e4b-it:free", messages=[ { "content": "What's the capital of China? Just give me the name.", "role": "user", } ], ) return completion_response.choices[0].message.content def my_llm_application(): query1 = query_llm("What's the capital of China? Just give me the name.") query2 = query_llm("What's the capital of Japan? Just give me the name.") return (query1, query2) print(my_llm_application()) We specify langfuse for the callback and each llm call is logged as a separate trace + span. Here you can see what this looks like in the dashboard: Basic trace and span in Langfuse dashboard The litellm docs include information on how to specify custom metadata and grouping instructions for Langfuse. Notably, we can specify (as of June 2025, at least!) things like a session_id, tags, a trace_name and/or trace_id as well as custom trace metadata and so on. So we can get most of what we want to specify in the following way: import litellm litellm.callbacks = ["langfuse"] def query_llm(prompt: str, trace_id: str): completion_response = litellm.completion( model="openrouter/google/gemma-3n-e4b-it:free", messages=[ { "content": "What's the capital of China? Just give me the name.", "role": "user", } ], metadata={ "trace_id": trace_id, "trace_name": "my_llm_application", "project": "hinbox", }, ) return completion_response.choices[0].message.content def my_llm_application(): query1 = query_llm( "What's the capital of China? Just give me the name.", "my_llm_application_run_789", ) query2 = query_llm( "What's the capital of Japan? Just give me the name.", "my_llm_application_run_789", ) return (query1, query2) if __name__ == "__main__": print(my_llm_application()) This looks like this in the Langfuse dashboard: Spans grouped into traces This is honestly most of what I’m looking for in terms of my tracing. If I were to use a non-OpenRouter model, moreover, I’d also get full costs in the Langfuse dashboard, e.g.: LLM costs in Langfuse dashboard As such, I can monitor costs from within OpenRouter and have the option to keep track of costs in Langfuse by passing custom metadata should I wish. I’ll make a separate blog where I actually go into how I set up + instrumented hinbox for this kind of tracing while continuing to use litellm.

3rd Jun 2025 • 1 votes
How to think about evals

Today was the first session of Hamel + Shreya’s course, “AI Evals for Engineers and PMs”. The first session was all about mental models for thinking about the topic as a whole, mixed in with some teasers of practical examples and advice. I’ll try to keep up with blogging about what I learn as we go. Most of the actual content will go up online at some point in the future, I’m assuming, so not much point writing up super detailed notes. (There is also a book coming, which I assume will be great, and about which you can learn more here.) So in general I’ll try to be doing the following as I blog along: highlight things I found interesting or inspiring based on the formal ‘lectures’ anything that comes up while doing the practical ‘homework’ (there are some optional exercises assigned to ground everything) contextualise or situate things that come up in my own experience having worked on a few LLM-driven projects Today, fresh out of the first class, I wanted to write about the mental model of the ‘three gulfs’ that they propose, the improvement loop that they suggest is how to measurably improve your applications, and also prompting through the lens of evals. Finally I’ll round off with a bit about what I’ll be exploring this week. The Three Gulfs: Specification, Generalization and Comprehension So there’s this image that they shared in the book chapter preview discussion that came up again during the lesson today: The three gulfs of LLM application development (They’ve shared it already in the YouTube discussion + I see it on Twitter being shared so I think I’m not sharing something I ought not to!) The course is very practically focused, especially so for application developers, so this diagram is in that context. The diagram offers up a way of thinking about LLM application development that pinpoints the places where you might do your work, and it’s also a way of thinking through things systematically, too. I was especially interested in the differentiation between the gulf of specification and the gulf of generalisation, since these can often feel similar, but actually the way to get out of them is actually slightly different. I’ll go into a bit more detail below, but basically with the gulf of specification you might want to be working on your prompts + how specific you are, whereas with the generalisation gulf you might need things like splitting up your tasks or making sure your system is outputting things in a structured way, etc etc. Note also that the world of tooling also doesn’t help you in a specific or targeted way to focus on one aspect of this diagram. Too often the tools try to cover the whole picture and probably also muddy the water by eliding the differences between the different tasks and challenges of each island or the gulf in between. All this is pretty abstract, so let’s go through them one by one. The Gulf of Comprehension in Practice This was seen as sort of the starting point for thinking through LLM Application improvement. At this point your big problem is that you’re trying to understand the data that comes your way from your users. You’re trying to understand the inputs to your application (what your users are typing, assuming that text is the medium of communication / input) and you’re trying to understand what the application or LLM is outputting. The challenge comes because you can’t read every single log or morsel of data. You have to filter things down somehow! If this were something more like traditional ML you’d have statistics to help boil down your data, but mostly we’re talking about unstructured text data so it’s much more unwieldy. This challenge means that people often get stuck at this point. This is where POC applications live, breathe and eventually die. You have enough sense that things are ‘kinda’ working, but you don’t really know what the failure modes are, so you don’t know how to improve it. You’ve tried out one or two things in a halfway systematic way, but really you have no idea what’s working well and what’s not. On Tools vs Process Hamel made the good point that it’s probably not so useful to think about tools too much when thinking at this stage. Generally speaking what’s going on is most often actually a process problem and trying to go straight to ‘what tool do I need’ is probably avoiding the real issue. The Gulf of Specification in Practice This is the place where you are trying to translate intent into precise instructions that the LLM will follow. You’re trying to be explicit and specific in the hope that the LLM will do what you want it to do, and not do the things that you don’t want it to do. The obvious manifestation of this is people writing bad prompts. It might seem that it’s also present when you try to have an LLM solve one problem when it’s either unsuited for that task or the task needs to be broken up and so on, but that’s the sister gulf of generalisation. Here, we’re focused on how to improve the specificity of your prompts. When you split things up and highlight the fact that prompts are something that you’ll need to work on and to improve, it becomes clear that it’s something you wouldn’t want to outsource or to skimp time on. Really the prompt writing is the thing that you (at certain moments, and where it’s identified as the thing needing focus / improvement) want to be working on in partnership with domain experts. For small applications, you might be the same person as the domain expert! For bigger projects, you might be working with domain experts. Just be aware that often the domain expert might not necessarily be detached enough to be able to figure out what needs the focus, or where the weaknesses of a prompt are. That’s what the iterative process / error analysis and everything else that’ll be taught in the course is for (see below and see future posts). Another point Hamel made was about why prompts are actually so important: “you have to express your project taste somewhere”. Given that your application might be fully / mostly driven by LLMs, the prompt is actually a really crucial place to express this taste and as such might be thought of as your ‘moat’. I know just from having experienced a variety of LLM-driven applications, it’s quite easy to tell the ones where the product team gave their prompts and their specification some real love. It’s the difference between POC junk that will die a slow and lonely death and something that delights and solves real user problems. Gulf of Generalisation Shreya didn’t really get into the details around the generalisation gulf in practical terms in this lesson, but I think this one can be a sort of place of comfort for the technically-minded to make refuge in. It’s one where there’s a ton of tools and technologies and techniques to play with, and vendors also live in this space and try to claim that their particular product or special sauce is the thing to help you and so on. The Improvement Loop for LLM Applications We also got a high-level overview of the loop that allows you to iteratively improve an LLM application: The analyse, measure and improve loop; adapted from an image used in the course There’s a lot to unpack in all these different stages, and we didn’t really get into the details in the session today but you can see how this offers a really powerful way of thinking through what it means to iteratively improve an LLM application. Learning how to implement this in a practical way will be the main thing I want to get good at by the end of this course. The process is made up of a bunch of techniques, but in my experience companies or use cases that struggle with improving what they built also lack the scaffold of this loop to orient themselves. Prompting through the lens of evals As we explored above, prompting is sort of the table stakes of improving your LLM application. In order to get good at prompting, it can help to appreciate what they are good at and what they struggle with. So, as Shreya put it, “leverage their strengths and anticipate their weaknesses” (when prompting). At this point Shreya got into some points around what kinds of things went into a good prompt but I think I’ll write a separate blog on that and I don’t want to just regurgitate what we listened to. Today was more of a high-level introduction, and in any case it was much more about the outer-loop process instead of the inner loop (where tooling + specific techniques play more of a role.) A slide from a talk I gave about the inner loop vs the outer loop of GenAI development So it’s great that the course gets into the weeds (esp in the course materials, which include the draft of the book Hamel & Shreya are writing) but I think the really useful thing they’re doing is situating the tactical improvements and techniques within the strategic patterns and workflows that teams and individuals should be doing to work on these LLM applications. At a high level, what are we talking about: how to tease out failure scenarios for these applications and their behaviours conversely, how to understand exactly which domains it does well for Things I want to think about more There was a ton of really rich discussion around prompting in the Discord. I’m interested in exploring more: cross-provider prompting decisions (i.e. how prompting an OpenAI model differs from what you do with a Llama model or whatever) prompts that work with reasoning models vs non-reasoning models the tradeoffs of whether you put your instructions in system prompts vs user instructions In general there’s been a bunch of noise recently about so-called ‘leaked’ system prompts from a bunch of LLM API providers and I’ve mainly been struck by just how detailed they are. I consider myself pretty good at improving and iterating on prompts, but I’ll admit I’m not writing these multi-thousand word tomes. I’d like to explore which scenarios it makes sense to do so, and how to calculate at what point it makes sense from a cost or latency perspective to do so. As I’m sure you can detect, I’m really enthusiastic about the lesson to come and will work in the meanwhile on some of the readings that have been set as well as the homework task of writing a system prompt for a LLM-powered recipe recommendation application!

19th May 2025 • 1 votes
First impressions of the new Gemini Deep Research (with 2.5 Pro)

Google released an updated iteration of their Deep Research tool that uses the new 2.5 Pro model. This was taken from a post originally made on Twitter, so please excuse the terseness. First impressions: a bit too eager to jump into a deep research task even when I just ask a clarifying question quite verbose, just like the OpenAI version. Not sure why both play this up a lot. It looks impressive but in practice I think we need more entry points into this. The ‘Executive Summary’ and other concluding headers are nice touches but I feel maybe there should be some more adherence to user requests for short reports. (I get that as UI it’s maybe weird to think for 10 mins and then spit out a very concise version, but it might actually be more useful.) I continue to be annoyed about how these Gemini DR reports handle footnotes (i.e. as endnotes whacked on at the end of the report). Almost a deal-breaker IMO. It’s almost like GDR tries to show how scholarly and serious it is by giving you these walls of prose (vs OpenAI DR which throws in a lot more bullet points). Not sure one is better than the other but would appreciate a bit more flexibility! The portability of these reports has always been not great. Yes you can export them to Google Docs but markdown (+ other options) would have been much better. In practice, this means that whenever I use GDR the report stays stuck there and I’m far less likely to share it with anyone, whereas the OpenAI DR reports I drop parts/all into a Github Gist etc. These reports have been getting better and better, all things considered. I’ve been following along and using GDR from the early days (even pre-OpenAI DR) and this latest version is the best version of it so far (as you’d hope!) (It’s also a little bit annoying that GDR has removed any way to use the older versions of GDR with Gemini Pro 2.0 and 1.5 etc. Makes it harder to actually compare these things.) Please let’s get an ipad version of the Gemini iPad app soon, too? Feels a bit regressive to have to use GDR on the web interface always. For serious research (as opposed to simply generating a nice report on some area where you don’t know much about already), all these tools remain hamstrung by the quality of the sources. In areas where I am (or very recently used to be) a leading scholar / researcher, the difference between what I’d expect (in terms of taste / discernment for picking out these sources) is especially egregious. Make the models better, yes, but have better filters + retrieval. So yeah, these tools are getting good! Kudos to the teams who are implementing this stuff. Hard to make it perform reproducibly well on so many open-ended uses. But more work to be done! IMO the really great implementations of this ‘deep research’ pattern will all be in-house where you can have control over: source selection (i.e. high-quality inputs only, not just some random things on the internet) how long it spends thinking about a particular area / loop of the research (or decides to backtrack and dig deeper etc) output types / templates / length different modalities of Q&A (sometimes you want reports, other times you want a quick question answered, other times you want visual guides etc etc.) different models for different kinds of tasks possibly you have little sub-research agents / processes which will go off and work on some hypothesis, possibly involving actual datasets / analysis of tabular data etc, something clearly missing from the current versions we have A few other things: GDR’s ‘clarification step’ (which I’ve heard them discuss on podcasts etc) is not as good or useful as the OpenAI DR clarification questions. In practice, because it’s buried under a concealment button that you have to click etc, and where the entire UI seems to be screaming at you to ‘Start Research’, you basically never update or amend the research plan. And when you do, it’s really not clear what’s changed because you don’t get some feedback or diff that your comments were understood; you just get an entire new research plan (again buried under the concealment button) Going forward we’re probably going to want / need ways of navigating the layers to this research. A global overview report will have subsections that (should you wish) can be expanded into their own more detailed or granular reports. This is how research works, after all. Not just endless new reports all trailed one after another pointed in the same direction. The other thing that I think we’re really going to need to work on is research taste. Like the LLMs that power them, GDR and OpenAI DR offer a level of research taste developed to the mean. (I know people are thinking about this since it came up on Dwarkesh’s podcast with the AI 2027 guys, but they were focused on scientific research.) I think there’s not a single answer for this which is, again, why I see the end result as people bringing these things in-house where they get to develop and refine what makes their particular flavour of research unique. (In the human-generated research world this is very much the case, where certain institutions (or even particular authors) are known for how deep they go, or what kinds of sources they prefer, or how they choose to feature or highlight the primary sources they access, and so on.) There are many possible variations of how this manifest, and I hope that we’re headed into a world where all the AI ‘deep researchers’ will be unique and quirky in all the best senses of that word.

8th Apr 2025 • 1 votes
Learnings from a week of building with local LLMs

I took the past week off to work on a little side project. More on that at some point, but at its heart it’s an extension of what I worked on with my translation package tinbox. (The new project uses translated sources to bootstrap a knowledge database.) Building in an environment which has less pressure / deadlines gives you space to experiment, so I both tried out a bunch of new tools and also experimented with different ways of using my tried-and-tested development tools/processes. Along the way, there were a bunch of small insights which occurred to me so I thought I’d write them down. As usual with this blog, I’m mainly writing for my future self but I think there might be parts that are useful for others! Apologies for the somewhat rushed nature of these observations; better I get the blog finished and published than not at all! 🤖 Local Models During this project, I experimented with several local models, which continue to impress me with their evolving capabilities. The recent launch of gemma3 was particularly timely - I found myself regularly using the 27B version, which performed admirably across various tasks. There are three or four models I keep returning to. mistral-small stands out as an exceptional model that’s been relatively recently updated and seems a bit underrated / underappreciated. The original mistral model continues to hold up remarkably well, particularly for structured extraction tasks and general writing needs like summarization. One important realization when working with real-world use cases: benchmarks can be deceptive. While helpful as general indicators, each model has its own strengths and quirks. Many newer models are heavily optimized for structured data extraction, but their performance ultimately depends on whether their training documents align with your specific use case. It’s crucial to test models against your actual requirements rather than relying solely on published benchmarks. For robust results with local models, I’ve found that implementing a “reflection, iterate and improve” pattern significantly enhances performance. When you need a model to summarize or analyze content in a particular format, having a secondary model (or even the same model!) review the output against the original prompt requirements is incredibly valuable. This reviewer model can suggest improvements to better fulfill the original request. Running this loop for 2-5 iterations (depending on complexity) can yield results approaching those of proprietary models like Claude or GPT-4, which might achieve similar quality in a single pass. For local deployments, this iterative improvement pattern is essentially non-negotiable. I also explored vision models, particularly llava and llama-3.2-vision. These were my primary tools for extracting context from images, generating captions, and analyzing visual content. Their effectiveness varies based on content type and language, but they represent impressive capabilities that can run entirely on local systems. A significant portion of my work involved non-English languages, including some relatively rare ones. This is another area where benchmark claims about supporting “hundreds of languages” often don’t align with real-world performance. Models might list impressive language coverage in their specifications, but actual proficiency varies dramatically. It reinforces my earlier point - always verify benchmark claims against your specific use case before committing to a particular model. 💬 Prompting & Instruction Following Working extensively with various models during this project reinforced some fundamental insights about prompting that might seem basic, but prove critical in practical applications. These observations are particularly relevant when working with local models, though they apply to cloud-based systems as well. Context matters significantly more than we might assume. While we’ve grown accustomed to proprietary models like Claude or GPT-4o performing admirably with minimal guidance, local models require more deliberate direction. The more relevant context you can provide (within reasonable token limits), the better your results will be. If you would naturally provide certain background information to a human performing the task, make sure to include it in your prompt to the model as well. Another key insight: every model has its unique characteristics. Techniques that work brilliantly with one model might fall flat with another, especially in the local model ecosystem. They each require slightly different prompting approaches, specific phrasing patterns, and tailored guidance. This necessitates running small experiments to understand how different models respond to various prompting styles. It’s still more art than science, but this experimentation phase is crucial when implementing local models effectively. Perhaps the most valuable lesson I rediscovered is that breaking complex tasks into smaller components yields superior results compared to using a single comprehensive prompt. This is particularly true with local models. When performing extensive data extraction or when dealing with structured data where the extraction targets differ significantly from each other, don’t expect the model to handle everything in one pass – even a human might struggle with such an approach. Instead, break down the task into logical components, create targeted mini-prompts for each aspect, and then recombine the results once all the separate LLM calls are completed. Yes, this approach adds processing time and complexity, but the quality improvement is well worth the trade-off. When accuracy matters more than speed, this decomposition strategy consistently delivers better outcomes. 🧰 Process & Tools My development environment during this project provided plenty of opportunities to evaluate various tools and workflows. As context, I primarily work on a Mac while maintaining access to a separate (local) machine with GPU capabilities for more intensive tasks. This setup allows me to flexibly experiment with both local and cloud-based models. For managing local models, Ollama continues to be my go-to solution for downloading, running, and interfacing with these models. A recent discovery that significantly improved my workflow is Bolt AI, an excellent Mac interface that provides seamless switching between local Ollama models and cloud-based alternatives. If you’re working in a hybrid model environment, Bolt AI is definitely worth exploring. I’ve also recently integrated OpenRouter into my toolkit, which solves the problem of managing countless API keys across different inference providers. OpenRouter not only offers native connections to many cloud providers but also allows you to incorporate your own API keys, streamlining access to a diverse model ecosystem through a unified interface. It also helps with setting spend limits on various models or projects. In terms of development insights, I was impressed by how rapidly front-end development can progress with the assistance of models like Claude 3.7 and OpenAI’s O1-Pro. These models perform exceptionally well when supplemented with documentation (such as an llms.txt file) alongside your prompts. While I can’t speak to their effectiveness with extremely complex applications or massive frontend codebases, they demonstrate remarkable proficiency with small to medium-sized projects. A significant portion of my experimentation involved RepoPrompt, a tool that recently transitioned from free beta to a paid license model. RepoPrompt addresses the challenge of getting your codebase into an LLM-friendly format. Unlike standard CLI tools that simply export code to clipboard or text files, RepoPrompt generates a structured XML representation that, when modified by an LLM and pasted back, creates a reviewable diff of the proposed changes. At least, that’s one of the things it allows you to do! It’s actually a bit more powerful / flexible than that and here’s a video so you can see it in action: RepoPrompt Demo Video While tools like Cursor and Windsurf offer similar functionality, they tend to become less reliable as project complexity increases. RepoPrompt shines when paired with an OpenAI Pro subscription, enabling effective integration of models like O1 Pro and o3-mini-high into your development lifecycle. In my testing, the RepoPrompt + O1 Pro/O3 Mini High combination consistently delivered superior results compared to using Cursor with Claude 3.7 (even with ‘Thinking Mode’ enabled). Despite the occasional pauses while these models process complex problems, the quality improvement justifies the wait. Additionally, I continued working with Claude Code and CodeBuff, both CLI-driven tools focused on code improvement. Of the two, CodeBuff has become my preferred option. Both tools require careful supervision—I typically keep Cursor open to monitor changes in real-time, occasionally needing to revert modifications or redirect the approach. These tools excel when you clearly articulate your objectives and maintain oversight of the implementation process. CodeBuff particularly impresses with larger codebases and demonstrates superior stability overall. An interesting pattern emerged during development: whenever files approached 800-900 lines, it signaled the need to refactor into smaller submodules to maintain LLM comprehension, especially when using agent mode in Cursor. The modular approach significantly improved model performance. I was genuinely surprised by the effectiveness of the RepoPrompt and O1 Pro combination. For smaller, targeted modifications, CodeBuff continues to demonstrate remarkable capability. While I didn’t evaluate these tools in conjunction with local models, I suspect such combinations would require more iterative refinement to achieve comparable results. 🧑‍🔬 Software Engineering Patterns Throughout this experimental project, several software engineering principles proved particularly valuable when working with LLM-assisted development. These patterns aren’t revolutionary, but their importance amplifies in the context of AI-augmented workflows. The principle of simplicity served as a cornerstone approach. Breaking development into the smallest logical next task repeatedly demonstrated its value, especially during the exploratory phases when project architecture was still taking shape. While some engineers might possess the cognitive bandwidth to fully conceptualize complex systems with perfect abstractions from the outset, I’ve found incremental development leads to more robust outcomes. This approach aligns naturally with how most developers actually think through problems and provides clear checkpoints for evaluating progress. Data visibility emerged as another critical factor. When leveraging LLM-assisted coding, comprehensive logging becomes even more essential than in traditional development. Strategically placed log outputs create a diagnostic trail that proves invaluable when troubleshooting unexpected behaviors. This practice creates a feedback loop that strengthens both your understanding of the system and the LLM’s ability to assist effectively. A particularly underappreciated practice I haven’t seen widely discussed is the importance of dead code detection. When working with LLM-assisted development, code cruft tends to accumulate more rapidly than in conventional programming. Tools like deadcode and vulture provide static analysis of Python projects to identify unused functions and variables. Running these tools periodically helps maintain codebase clarity by flagging remnants that might otherwise cause confusion during review. I’m not certain whether newer tools like ruff from Astral include this functionality (particularly for function calls), but the capability is invaluable for maintaining a clean, navigable codebase. Taking time to think offline—away from the keyboard—often yields surprising clarity. This deliberate pause creates space to articulate precisely what you need for the next development increment. When you can express your requirements with precision, the LLM’s output improves proportionally. Ambiguous instructions inevitably produce suboptimal results, whereas clarity fosters efficiency. A final observation worth emphasizing: having experience as an engineer in the pre-LLM era remains tremendously advantageous. When confronting complex workflows involving chained LLM calls with interdependencies and reflection patterns, traditional debugging skills become indispensable. Knowing when to step away from AI assistance and dive into manual debugging with tools like pdb, stepping through code execution and inspecting variables directly, represents a crucial judgment call. LLMs and coding agents often demonstrate a bias toward generating new code rather than methodically analyzing existing problems. Recognizing the moment when direct human intervention becomes more efficient than continually prompting an AI is a skill that comes with experience. Once you’ve manually identified the underlying issue, you can return to the LLM with precisely targeted prompts that yield superior results. 🌐 Appendix 1: FastHTML As a practical addition to my experimentation, I implemented FastHTML for the first time to build a frontend for my knowledge base extraction assistant. The experience was remarkably frictionless, particularly when leveraging their llms.txt file—a markdown-formatted documentation set that integrates seamlessly with your frontend codebase when provided alongside prompts. This approach works exceptionally well with models like O1 Pro or O3 Mini High, creating a development workflow that feels intuitive and responsive. Despite having substantial JavaScript experience from previous roles, I found FastHTML significantly more manageable than complex JavaScript frameworks that dominate the ecosystem today. The reduced cognitive overhead and natural integration with Python-based workflows makes FastHTML a compelling choice for ML practitioners who prefer to minimize context-switching between languages and paradigms. The framework strikes an excellent balance between capability and simplicity that aligns perfectly with rapid prototyping and iterative development cycles common in ML projects. For those building interfaces to ML systems, it’s definitely worth considering as your frontend solution. 📃 Appendix 2: OCR + Translation Another interesting challenge I tackled involved OCR and translation of handwritten documents in non-English languages—a task that proved impossible to accomplish in a single pass with local models, particularly for less common languages. The solution emerged through methodical problem decomposition: Breaking down PDFs into individual page images Segmenting each page into overlapping image chunks (critical for handwriting where text may slant across traditional line boundaries) Applying OCR to extract text in the original source language from each image segment Using translation models to convert the extracted text to English This multi-stage pipeline allowed me to overcome the limitations of local models when confronted with the combined complexity of handwriting recognition and translation. Both gemma3 and llama-3.3 performed admirably within this decomposed workflow, demonstrating that even resource-constrained local deployments can achieve impressive results when problems are thoughtfully restructured. This case exemplifies a core principle of effective ML implementation: when dealing with complex, multi-faceted challenges, breaking them into targeted sub-problems often yields better outcomes than attempting end-to-end solutions—especially when working with constrained computational resources. While this approach may increase processing time, the quality improvement justifies the trade-off for many practical applications.

15th Mar 2025 • 1 votes

More in AI

Agentic Primitives 101

Inside the harnesses that let AI carry a job across hours, days, and conversations.

5 hours ago • 1 votes
Agent governance looks less like moral education and more like institutional economics

how agents turn ruthless on paper first

2 days ago • 1 votes
We are going to kill “unalive”

I talk to a lot of old people, those who were born in the 20th century, and if I ask them what the word “unalive” means, they usually have no idea what I’m talking about, except for some of them who have kids or who study Internet culture. This will, of course, probably seem very weird to most people who were born in or grew up in the 21st century. Just to recap for the olds: “unalive” is the word you use to represent concepts like dying, or death, or killing or being killed, on digital platforms where saying those words accurately will cause the algorithm to punish or censor you. Or, maybe, where the perception is that using those words will result in being censored by the algorithm, and no one is actually willing to find out what happens if you use the forbidden words. This sort of attack on people’s expression started on platforms like TikTok, where nearly all content is distributed through an algorithmic feed, but has since become ubiquitous in nearly all digital media that we see. In fact, these tics are now so prevalent that it’s routine to hear people using this kind of language in everyday life, even though there’s not yet an algorithm to appease in the physical world. I’ve heard people say, out loud, “he unalived himself”, in reference to someone dying by suicide. And all of this has become even more visible in recent days as online conversation has turned to discussion of the horrific lack of accountability around the tragic rape case at Cornell University. Across the Internet, people are routinely referring to the central crime in the case as r*pe or “grape” or even using the 🍇 emoji, without a second thought for what it means that the very word can’t be said online anymore. Or, at least, the assumption is that it can’t be said. To be clear, I am very much in favor of people using content warnings or sensitivity markers for content, and fine with people using abbreviations like “SA” for references to disturbing or triggering topics like sexual assault; we should provide people with as much context and control as possible when choosing what information they want to consume and when. I also know that sometimes, people use lesser terms for stressful subjects like death or assault to create a bit of ironic distance from painful or upsetting topics. But most of the different variations of wording and emojis are coming from trying to appease the platforms, and there’s a heavy cost for those who are worried about being mindful: If someone is using a tool to filter out content, it will no longer be effective because everyone is using misspellings and euphemisms and imagery to get around the algorithm. The spread of censored and mangled syntax is happening because people believe, or have experienced, that platforms will silence them for accurately describing the world in plain language. This shit is terrible, and it has to stop. You Were Not Born With These Constraints One of the things that’s most concerning to me is that an entire generation has grown up not realizing how extreme it is that their very language is being chosen for them by platforms run by people who hate that generation’s ability to express itself, and who hate the things it has to say. From their youngest days, this generation grew up watching people make stupid faces at them for YouTube thumbnails and never had a chance to reflect on the fact that those creators didn’t want to be humiliating themselves by making those expressions — the demands of the algorithms of Big Tech forced them to do that. The rituals of feeding the algorithm are so built into people’s everyday habits that they’re invisible to people who weren’t alive before today’s platforms took over. Every parent of my cohort remembers the first time they heard their toddler finish doing something cute in their living room, and then turn around and say, “please like and subscribe!” afterwards. It’s a ghastly, sickening feeling to confront the fact that our little kids were being brainwashed into thinking that every adorable thing they did should be followed by a prompt to provide data to Google. Over on Instagram, where people originally signed up thinking they were going to see someone’s vacation pictures, or shots of their cousin’s kids, you’re now stuck watching people beg for everyone to reply with cultish phrases in the comments, which will then earn them an obviously AI-generated response in return, all in service of “showing activity” to the algorithm, like it’s an angry god that needs a sacrifice. They’re just not sure exactly what the angry god wants. Your free speech was taken away from you, and the people who did it are the same ones who spent years pretending to care about “free expression”. They contrived examples of lack of free speech on college campuses while squashing protests, and cried crocodile tears about “cancel culture” while getting people fired for political criticism. Now they have no problem with billionaires deciding exactly what words everyone is allowed to say. Larry Ellison is not content with his family owning all of the movies and TV shows — his family has to control what words people are allowed to speak on TikTok, too. Elon Musk isn’t content to merely generate and distribute child sexual abuse material for profit — he wants to silence the messages of the few decent people who are foolish enough to remain on Twitter/X, too. (That’s why I wrote you a guide on how to get your organization off of that cursed platform.) Now that an entire generation has grown up using these Orwellian euphemisms, and all of the Big AI products are trained on the Internet that was created under this regime, do you think today’s AI tools even know that the real, uncensored world exists? If you can’t say “genocide” on any of the major platforms, yet those are the ones all of the Big AI tools used as their training data... well, then the AI tools sure aren’t very likely to know much about genocide, are they? Fuck the Algorithm Our creativity can be constrained by the language we use — our imaginations are limited by what we can think to say. If we’re trained to limit the words we speak just by habit, and those limits are put in place by people whose social, political, cultural and moral goals are the opposite of what we value, then our work is unalive before it is even born. The answer to this is simple: say what you mean. This will take, to some degree, courage. It may even take, I hesitate to say, some sacrifice. When I suggest this course of action to people, they inevitably say, “But it will cost me audience!” or “But what if I lose followers!” or “What if they demonetize me!” Okay, what if they do? What if they do. Are you willing to push on this? To make a point about it? To move to platforms where you can actually say what you mean? Or to remember that you already have platforms where you can say what you mean? On an email newsletter or podcast you can say whatever the hell you want and nobody can stop you. On my blog right here, I can even curse in a headline and it won’t affect anything about how my site operates. (And a reminder: Substack is not an email newsletter, and a Spotify show is not a podcast — they’ll be unaliving your distribution any day now.) If you are a 20th century relic like me, it is incumbent upon you to remind the generation that grew up inside the algorithm that another world is possible, and that we know this because we lived it. We were able to style a MySpace page in any way that we wanted; the code for LiveJournal was entirely open source so there was no part of the algorithm that was unknowable. A blog like the one you’re reading right now could be made by anyone, and put up for pennies, and nobody could stop it from being read by millions of people. (And that last one? It’s still possible.) If you are from this century, forget all that rambling bullshit about ancient history: all that matters is you getting what you deserve, because you’ve been fucked over by the same billionaires who’ve poisoned your planet and infested your world with slop. The best artists around you are invisible to you, and the most important statements by activists that you care about are being silenced. It’s not a conspiracy, it’s a system working as designed. And the proof is as obvious as the fact that the angriest activists you know can’t even talk about systemic abuses or state violence without having to put it in algorithmically approved speech or censoring their captions like they’re going to be read by 5-year-olds. It should make you furious. It is time to kill “unalive”. The response is simple: For every message you put out, start by saying what you mean. Don’t work backwards from what the algorithm wants or what a platform permits. Build a presence on every platform you can, even the ones where you have fewer followers or where you’re harder to find. Tell your audience that your speech being free matters more than corporate convenience. Keep saying it, and keep it positive: independence gets them better art, better information, and more connected communities. Build alliances with other artists, activists and people who share your values, and let them know you’re going to start sharing your work uncensored. Start releasing your work uncensored and see where the platforms push back. (You may be surprised: sometimes you were censoring yourself in anticipation of limits that weren’t even there.) If a platform does try to limit your reach or expression, make a LOT OF NOISE about it. Tell the press, rally your alliance, and spread the word on your other platforms, using the moment to build audience and raise support there. Get others to amplify the parts of your work that don’t violate platform policy, so the controversy drives people to the rest. Find the others pushing back on algorithmic control of expression, and raise and praise their work when they do the same. If we keep accepting the words that are forced upon us by TikTok and Meta and Google and all the rest, while platforms like Twitter/X allow the most hateful and harmful content in the world to be distributed completely unfettered, we’ll only see authoritarianism rise, and the harms against the vulnerable accelerate. But what breaks my heart almost as much is that we’ll see so many brilliant artists and activists and thinkers whose genius will be muted or silenced by mindless, heartless algorithms that capriciously decide who gets to say exactly what words, in what ways. I get angry every time I think about it. The tech tycoons get ever more brazen in what they’re willing to say publicly, boasting about how they’re going to cause the end of the world, or calling for ethnic cleansing, all while putting tighter and tighter reins on the speech and expression of ordinary people. It’s time for “unalive” to die.

2 days ago • 1 votes
What is Codemode

More than a year ago I wrote a few posts here that recommended people not to load custom tools into their context (or MCP servers) but to just use more scripts. Most importantly I wrote that Code Is All You Need and I wrote about that MCP needs code. With Pi 1.0 we now added MCP support via Codemode which in some ways is a long time coming, but then also maybe somewhat surprising to some. So I want to share some updated thoughts on this blog on what this all means. What Are Tools When a harness like Pi provides tools for an LLM to call, it does so by supplying some tool definitions which then translate into some token structure on the server side. Whether a model is encouraged to call a tool is the result of the reinforcement learning process. Something I wrote about before if you want to learn more. One of the reasons we strongly lean towards CLI and bash is because it allows easy composition of calls, and because the model also learns how the file system works when it’s trained. So when it invokes a tool like echo foo > /tmp/test.txt the model also learns that after that tool call, there is now a file called test.txt in /tmp. However bash has one fundamental limitation which is that it can only compose programs that run. And there are some things, which are not programs, but native tools to the LLM and they sort of have to be. The most obvious example here is read or view_image. If a multimodal model needs to read an image, it cannot use cat for that because the harness needs to inject the actual image payload into the protocol of the LLM. Another quite vivid example are sub agents. In order to spawn and orchestrate sub agents, it’s tricky to avoid tools that are provided by the harness. While in theory the agent could provide a CLI tool that talks to the outer harness via environment variables and Unix sockets, it’s a rather crude process. It however has another issue, and that is where the code runs. Brains vs Hands To better understand that, it’s important to think a bit more about where all the bits and pieces run. There really usually are two different systems involved. The first is the brain, the harness: it runs on one machine. It’s trusted. The second is often the same machine, but it’s really where the tools are executing: the hands. In Pi we now call this the execution environment, but you can think of it as the target of all the operations. Crucially what is important for us, is that there is a dividing line between the harness brain and the target environment that runs bash and executes the tools. And splitting this in half has some really important consequences. For a start it means that they are running on different file systems and they have different levels of trust. If you for instance use a sandboxing solution like Gondolin your bash stuff will be sandboxed just fine, but the harness itself will not be. Orchestrating The Harness Which brings us to what Codemode really does: it’s a way for the LLM to express and orchestrate complex operations on the harness side, but not the execution environment side. Codemode runs in the harness, in its own sandbox. In case of Pi it’s running in QuickJS within a WASM runtime with intentional limitations: no network, no file system, no timers, limited RAM. The only way is to call more tools. You could also imagine that Codemode could run Scheme or some other language as well. If you are not familiar with Codemode, it’s basically just a way to issue tool calls from within some language, in our case JavaScript. That allows you to compose those calls without necessarily going through the LLM’s context. Credit for naming goes to our friends at Cloudflare who coined it. For instance if you issue a bash call as a regular tool call in the LLM, then we only throw the trailing 2000 lines into the context and if the agent wants more, it needs to look at the overflow file itself. If however the agent issues that invocation via Codemode, then the Codemode side gets larger outputs sent structurally. Most importantly, because Codemode is JavaScript the agent can express concurrent operations and basic workflows. A common way in which you see agents now use this, is to first probe at 5-10 items from some tool response to see what it looks like, and to then write a Codemode script that processes the next n items. Codemode also allows you to throw state into the transcript! That means that one Codemode invocation can stash away data, that the next call in the session can load again. And remember: this is on the harness host, not the sandbox. In case of Pi, Codemode also allows you to issue calls that naturally do not make any sense in Pi’s traditional interface. For instance if you want to generate images with an image model or you want to classify some text with a one shot classifier model, those Pi APIs are exposed via Codemode, but not via regular tools where they would just waste context. What It Looks Like So now that we talked a bunch about it, it’s probably worth being a bit more explicit about it. Let’s walk ourselves through some invocations of Codemode of recent Pi sessions of mine. Note that none of this code is human written. It’s from real sessions of Pi, just re-indented for your viewing pleasure. The agent starts using Codemode automatically either because it’s a task where the model already naturally picks up that tool, or because a user asked it to. Note that Codemode is by default only enabled in Pi when MCP is enabled, but you can turn it on with "defaultTools": ["+codemode"] in the settings. Just ask Pi to enable it for you. Generating Images Let’s start simple with image generation. Image generation is a feature that Pi supports in the AI SDK core, but it’s not a tool that the agent can use. In the past the only way to use image models has been to write a bespoke extension or to have the agent run node itself and use the internal image APIs. However because we expose quite a few of the internal model APIs within Codemode, it means that the agent can use it: const [painter] = await models.getAvailableOfType("image"); const result = await models.generateImages(painter, { input: [{ type: "text", text: "A cute little puppy sitting on a grassy " + "lawn, soft natural light, photorealistic" }], }); if (result.stopReason !== "stop") return result.errorMessage; for (const block of result.output) { if (block.type === "image") image(block); else text(block.text); } Note that the call to image() sends the image back as image content to the LLM. On the harness side it feeds it directly into both the agent, as well as onto disk as a temporary artifact in case the agent wants to be able to pass that image back to bash. Classifying Things Similar things apply to classifier models such as Jev. They also do not fit well into the workflows of an agent through the typical tools. But rather than making a bespoke tool available, Codemode just allows the agent to reach into the AI SDK and invoke those directly. Here you can see how Jev is used to mass process GitHub issues for a quick sentiment analysis: const jev = await models.getModelOfType("classifier", "typesafe", "jev-latest"); const r = await tools.bash({ command: "gh issue list --state open --limit 100 " + "--json number,title,body,comments", }); const issues = JSON.parse(r.output); const results = await Promise.all(issues.map(async (issue) => { const res = await models.classify(jev, { state: { title: issue.title, body: (issue.body || "").slice(0, 4000), comments: issue.comments.slice(-5).map(c => c.body.slice(0, 800)), }, questions: { sentiment: { type: "choice", instructions: "What is the overall sentiment of the author towards pi?", criteria: { positive: "Appreciative, happy, constructive praise", neutral: "Matter-of-fact report or request without emotion", negative: "Frustrated, annoyed, upset, or angry", }, }, frustration: { type: "score", instructions: "How frustrated is the reporter?", criteria: ["not at all", "mildly", "clearly frustrated", "very angry"], }, kind: { type: "choice", instructions: "What kind of issue is this?", criteria: { bug: "Bug report or regression", feature: "Feature request or enhancement", question: "Question or support request", other: "Docs, discussion, meta, spam", }, }, }, }); if (res.stopReason !== "stop") { return { n: issue.number, title: issue.title, error: res.errorMessage }; } return { n: issue.number, title: issue.title, ...res.answers }; })); store("sentiment_results", results); return results .filter(r => !r.error) .sort((a, b) => b.frustration.score - a.frustration.score) .slice(0, 12) .map(r => `#${r.n} ${r.frustration.score.toFixed(2)} [${r.kind.choice}] ${r.title}`); Note how in that above example we also call store() which dumps the result of that execution into the session transcript. A future invocation of Codemode can thus read back that result if it wants to. The Promise.all here is fine, because Pi limits the total number of concurrent tool executions itself to four and maintains a queue for the rest. A more adventurous example is to use Jev to drive a game engine for debugging purposes: Codemode with Jev for Game Debugging Here it knows about my tankctl command and it built itself quickly a minimal harness around it to drive a game loop to assist a user with debugging a problem. Note how it built a 30 step loop in which each step goes back to both the game engine to get a text dump of what’s going on, and then to Jev to determine what to do next: const jev = await models.getModelOfType("classifier", "typesafe", "jev-latest"); const tank = async (cmd) => (await tools.bash({ command: `tools/tankctl "${cmd}"` })).output; await tank("start --map assets/maps/night_arena.map"); const questions = { action: { type: "choice", instructions: "You control the tank '@' in a top-down tank game. " + "Choose the best next action.", criteria: { attack: "an enemy has line of sight to you and you can fire at it", approach: "no enemy has line of sight; drive toward the nearest enemy", dodge: "an enemy shot is heading at you and will hit soon", powerup: "a powerup is close and no enemy threatens you", }, }, }; function commandFor(choice, st) { const p = st.player; const enemy = st.enemies.filter(e => !e.dead) .sort((a, b) => (b.los - a.los) || (a.dist - b.dist))[0]; if (choice === "attack" && enemy) { return `fire_at tank ${enemy.id}; frames 30 until clear,damage,kill`; } if (choice === "dodge") { // move perpendicular to the closest incoming shot const s = st.projectiles.filter(s => !s.yours) .sort((a, b) => a.eta - b.eta)[0]; const dir = s && Math.abs(s.vel[0]) > Math.abs(s.vel[1]) ? (p.pos[1] > s.pos[1] ? "+down" : "+up") : (p.pos[0] > (s ? s.pos[0] : 0) ? "+right" : "+left"); return `input ${dir}; frames 20 until damage; input stop`; } const powerup = st.powerups.filter(u => u.available) .sort((a, b) => a.dist - b.dist)[0]; if (choice === "powerup" && powerup) { return `goto ${powerup.pos[0]} ${powerup.pos[1]} 180`; } return enemy ? `goto ${enemy.pos[0]} ${enemy.pos[1]} 90` : null; } const log = []; for (let step = 0; step < 30; step++) { const st = JSON.parse(await tank("state")); if (st.state !== "playing") break; const threats = st.projectiles .filter(s => !s.yours && s.miss_dist < 1.5 && s.eta < 1.5) .map(s => `incoming shot dist ${s.dist} eta ${s.eta}s`) .join("\n") || "no incoming shots"; const r = await models.classify(jev, { state: { map: await tank("view 8"), threats, hp: st.player.hp }, questions, }); if (r.stopReason !== "stop") { log.push(`#${step} classifier error: ${r.errorMessage}`); break; } const choice = r.answers.action.choice; const cmd = commandFor(choice, st); if (!cmd) break; log.push(`#${step} hp=${st.player.hp} ${choice} -> ${await tank(cmd)}`); } return log.join("\n"); Calling MCP Servers Lastly, Codemode obviously is great for calling MCP servers. And because we do not actually expose any of the MCP tools to the LLM, the agent first uses provided APIs to issue a tool search within Codemode to discover what it might be able to do with the connected servers. This form of progressive discovery makes the whole MCP business work well enough for a lot of use cases today. Here for instance you can see the agent reach for the Sentry MCP straight away, even without discovering the tools, presumably because it has learned during the RL process already about what the Sentry MCP looks like. But it learns from what we inject into the system prompt, that the Sentry server is available to begin with. It’s not completely guessing here. const orgs = await tools.mcp__sentry__find_organizations({}); const { organizations } = orgs.structuredContent; const results = await Promise.allSettled(organizations.map(org => tools.mcp__sentry__find_projects({ organizationSlug: org.slug, regionUrl: org.regionUrl, }) )); return organizations.map((org, i) => { const r = results[i]; if (r.status !== "fulfilled") return { org: org.slug, error: String(r.reason) }; if (r.value.isError) return { org: org.slug, error: r.value.content }; return { org: org.slug, projects: r.value.structuredContent.projects.map(p => p.slug), }; }); Modern MCP Is A Fight I really don’t want to talk too much about MCP here, but MCP is in fact a protocol that greatly benefits from Codemode. The problem in parts is that MCP in practice often targets harnesses that do not (yet?) use Codemode. But the tide is shifting. In the meantime, a temporary crutch has been to do what Cloudflare did, and do Codemode within the MCP server. But now we have Codemode in Codemode which is pretty bad. It means double JSON escaping, easy for smaller models to get confused by and the inner code cannot call the outer tools. So if you for instance use the Cloudflare MCP servers in Pi, the agent needs to write JavaScript and funnel it through more JavaScript. This is really not optimal, but it’s also understandable that this is happening: const accRes = await tools.mcp__cloudflare__execute({ code: `async () => { const r = await cloudflare.request({ method: "GET", path: "/accounts" }); return r.result.map(a => ({ id: a.id, name: a.name })); }`, }); const accounts = JSON.parse(accRes.content.map(c => c.text).join("")); const out = []; for (const account of accounts) { const r = await tools.mcp__cloudflare__execute({ account_id: account.id, code: `async () => { const r = await cloudflare.request({ method: "GET", path: \`/accounts/\${accountId}/workers/scripts\`, }); return r.result.map(s => ({ id: s.id, modified: s.modified_on })); }`, }); out.push({ account: account.name, workers: r.content.map(c => c.text).join("") }); } return out; MCP Desires So to end things off: how well does Codemode work with MCP today? Well … not amazingly well. That’s because MCP servers are not really targeting harnesses that use Codemode yet (though at this point I think most harnesses support it). For this to work well some recommendations: Structured content: Codemode wants calls to return some nicely formatted JSON. So that needs to come back from the server, and many don’t do that yet. The outputSchema system in MCP is great for that. Consistent results: an interesting failure case is when an MCP server does not return consistent data. For instance because it tries to token optimize things depending on how many items are in the result set. This can cause an initial probe with 5 items to succeed, but then fail when the server returns the maximum batch size. Large binary data: today MCP does not yet support large binary data so quite a few use cases that are really interesting do not work well at all yet. You end up with all kinds of weird workarounds such as pre-signed URLs to allow file uploads then to happen through non MCP channels. Composable tool search: the MCP server might know better than the MCP client which tool is appropriate for a task. But there is no good mechanism today that allows a harness to fan out tool searches across multiple MCP servers. It’s all emergent behavior and it does not scale well to multiple active servers. Future of Codemode So where does this leave us? Is this a reversal of what I wrote a year ago where I encouraged CLIs? I don’t think so. In fact, the MCP ecosystem from my perspective picked up on exactly what we pointed out a year ago works: code. But Codemode goes beyond MCP in that it can act as a capable mechanism within the harness to express more freedom for the agent. There are however also some things that we still need to figure out. For one, durability with Codemode is trickier. We might have to adopt some ideas from durable workflow engines here to snapshot invocations. Or maybe, something like Starlark is a better composition language than JavaScript given its deterministic nature. Images, binary data and just the inability of this pattern to work with smaller models is also something that needs to be fleshed out. So it’s for sure not a perfect solution yet, but it’s quite a useful pattern that I expect us to leverage more.

2 days ago • 1 votes
Do you actually like hard problems?

I spent years becoming the engineer people reached for. Now they still reach for me, but for different reasons.

2 days ago • 1 votes
📚 BoredReading

You seem to be enjoying this.

Join free to unlock everything.

Create free account

Already have an account? Sign in