More from Eugene Yan
A sandboxed target, inputs that influence task difficulty, tools, and a grader.
Context as infra, taste as config, verification for autonomy, scale via delegation, closing the loop.
Label some data, align LLM-evaluators, and run the eval harness with each change.
Based on what I've learned from role models and mentors in Amazon
More in AI
There's a lot of polarising discourse right now about the threat AI poses to humanity. Some think it's a farce and others think we face extinction. Here are my thoughts.
A middle ground between the cybersecurity and AI safety communities
An overview of the current state of the engineering market and the AI skills that are in demand