How to Build Frontier-Lab Quality Evals with Daniel McKinnon, ex-PM at Meta, Google
How to Build Frontier-Lab Quality Evals with Daniel McKinnon, ex-PM at Meta, GoogleHe wrote evals for Gemini, for Llama, and for Ray-Ban Meta. Today he builds one from a blank spreadsheet, live!
Check out the conversation on Apple, Spotify, and YouTube. Brought to you by:
Today’s EpisodeRasty Turek spent the past year building with coding agents, and he mapped how his process changed over that time. His eval started as QA. Then it became the spec. He reckons he now spends around 90% of his time on evals. And the difference is “indescribable.” Just about every developer and builder has experienced a similar pattern. This is the opportunity for PMs to add great value to the development team. Just as the value of a long PRD was waning, evals sprang up. And PMs are often considered the best owner. So how do you do them well? I’ve now covered evals basics, excel implementation, advanced theory, Braintrust implementation, and Arize implementation. Today, I cover how to practically write a great one to hand off to another team. And I do it with the perfect person. I’ve brought in Daniel McKinnon, who was a PM on the Llama models at Meta. He wrote the viral piece “Show, Don’t Tell” where he said:
So what would a good eval to send him look like? Today, we show you. π I’m doing a free webinar Thursday on getting AI PM interviews. Join me: The next cohort of my LandPMJob program starts August 17th! This is cohort 4, and will be the best yet. Sign up. Today’s Deep DiveThis is the guide to building your first eval before you have a single user.
This post was written entirely by hand: I use AI in my editing and ideation phases, but am here to give you HUMAN content. 1. The cold-start problemEverything I’ve taught you about evals so far has the same first step: Look at your data. Read the traces ➡️ Find the failures ➡️ Categorize them. That’s the error analysis method, and it’s still the most calibrated way to build evals over time. But it has a prerequisite. You need outputs to analyze. On day zero of an AI feature, you have none. This is the gap Daniel’s method fills. He builds the eval before the product exists, out of domain knowledge instead of production data. I’m calling it the cold-start eval, and it’s the one most PMs never learn. When a product is expected to do nearly everything, a normal PRD usually folds because it goes into specifics like how it behaves in particular situations and how the user gets value. So examples become the specification.
Offline vs Online Evals: An ExplainerSay you run a recipe site and you want to generate ice cream recipes. Your offline eval is a hundred prompts covering ice cream flavours, each with a possible correct scoring method. Online is what happens after you ship it and people use it. The bet you’re making is that a good offline score predicts a satisfied online user. It doesn’t always hold, yet it’s still the best predictor available. Sometimes the offline eval tells you the thing you want to build is not possible with today’s models, and you find that out in week one instead of after launch. Then you make changes accordingly. One note before we build, Daniel ran his live demo in Codex. The method has nothing to do with the tool, so I’ve written every step below in a way that could work in Claude Code or ChatGPT Work also. 2. Step-by-step, how to build your eval setYou have a feature idea and no data. So this is how you get to a working eval set in one sitting. The shape of what you’ll build will kinda look like this ↓ Step 1 - Write the problem in one sentenceYou have to understand exactly what the feature will do in one sentence. For example: Extract sender’s name from support e-mails. Step 2 - Validate your domain expertiseYou will not write a good eval for a space you don’t know. So, if you’ve never shipped an AI feature, Daniel’s advice is blunt. Find someone who has and get them to walk you through your first one. And that’s where I come in :D Step 3 - Set up the workspaceOpen Claude or ChatGPT and create a Project. Upload whatever your feature reads, like, sample support tickets, example transcripts, a product spec, your scoring guidelines. Then paste this into the project instructions: You are helping me build an offline eval set for [FEATURE]. Step 4 - Find your floorPick the easiest genuine case you can think of and check the model can do it at all. Daniel picked cystic fibrosis - a disease with a known genetic cause and a canonical gene behind it. While Daniel’s version was pretty niche, your version can look like: Here is the easiest real case for [FEATURE]: [PASTE THE CASE]. Now run it on the model you’ll use. Call this the floor test. It’s where you find out that your easy case sits above the ceiling of the model. If your floor fails, you can make the case easier. Step 5 - Find your ceilingNow the opposite move. Find a case you’re fairly sure nothing can solve. But why bother testing something you expect to fail? Because if you can’t find a case that fails, your eval is already saturated and it will never tell you anything again. Here is a case I believe sits at or beyond the edge of what you can do: Step 6 - Fill in the middleYou have a floor and a ceiling. Everything else is a binary search between them. Here are two cases for [FEATURE]: Run that a few times and keep going until you have a set. How many? Around 100 is typical. And yes, you use AI to build the test for the AI. That’s the only way this takes 90 minutes instead of two weeks. 3. How to score it?Daniel described the scoring step but didn’t write one out on the recording, so I built a version for you all. It’s built on the binary pass/fail principle from the eval guide. You are grading one output. You are not being helpful. You are grading. Calibrate the judge before you trust it. Do this once, at the start. It would take around 30 minutes, but it will make a huge difference between a score and a guess. 4. How to read the numberSay your eval comes back at 50%. Is that shippable? Probably not. So you go one level down and ask which slices you’re good at. Then you put guardrails on the product so it only answers the kinds of questions it gets right 80% of the time. That’s a PM call. Everything below the line goes back to the research or engineering team with a specific ask. You hand them a target instead of a feeling. 5. How evals can help your careerDaniel estimates there are probably fewer than 100 PMs building frontier models worldwide. And, as he said:
Evals, on the other hand, are purely built on judgment. So, this is the place to go deep and differentiate. As you will be one of the few who can. P.S. Many of you are looking for a great AI builder role. Daniel’s hiring. Here’s a handy infographic summarizing what we covered: Where to find Daniel McKinnon
Go Deeper on EvalsTo never miss an episode, subscribe to the podcast on YouTube and follow on Apple & Spotify Last Few ThingsI’m doing a survey on PM salaries. Please fill out this 3-minute survey. I’m giving away 10 copies of Eric Ries “Incorruptible” to participants.
And, as I mentioned earlier, my PM cohort starts next month with a free Webinar Thursday. Invite your friends and earn rewards
If you enjoy Product Growth, share it with your friends and earn rewards when they subscribe.
|







Comments
Post a Comment
VHAVENDA IT SOLUTIONS AND SERVICES WOULD LIKE TO HEAR FROM YOUπ«΅πΌπ«΅πΌπ«΅πΌπ«΅πΌ