Full Width [alt+shift+f] Shortcuts [alt+shift+k]
Sign Up [alt+shift+s] Log In [alt+shift+l]
1

Evaluate LLM Apps in Go

from maragu’s blog [alt+shift+b] in AI

As Large Language Models (LLMs) become a bigger part of our apps, making sure they work reliably brings some challenges. Unlike traditional code with predictable outputs, LLMs are indeterministic and can output downright weird stuff. That’s where good evaluation tools come in! In this post, I’ll show you the eval package from my newly renamed maragu.dev/gai module.
25th Feb 2025

Stay updated

Get a weekly newsletter with the top 5 articles worth reading every week.

More from maragu’s blog

Choosing what to keep human

A quote from Ethan Mollick on figuring out what to keep human using AI.

27th May 2026 • 1 votes
Headless software

How do we design software when users use it primarily through agents?

20th Apr 2026 • 1 votes
autoresearch and diary agent skills

I keep adding to and refining my agents skills repo.

26th Mar 2026 • 1 votes
23 skills I wrote with my AI

My AI, maragubot, has a blog post up with the skills we've been developing.

20th Feb 2026 • 2 votes
Do we still have time for human code review?

I’m increasingly convinced that in the near future, we software developers won’t be looking at our production code anymore.

5th Jan 2026 • 1 votes

More in AI

The Business of Building God

a look at the changing economics of AI labs

a week ago • 1 votes
It’s Been a Minute

I’ve been meaning to write about *all of this*.

a week ago • 1 votes
The Overhang

Using your deep knowledge, wide knowledge, taste, and agency

a week ago
AI and Existential Dread

There's a lot of polarising discourse right now about the threat AI poses to humanity. Some think it's a farce and others think we face extinction. Here are my thoughts.

a week ago • 1 votes
📚 BoredReading

You seem to be enjoying this.

Join free to unlock everything.

Create free account

Already have an account? Sign in