Hamel Husain and Shreya Shankar Release Evals Skill for AI Coding Agents
Lenny Rachitsky recommends installing a new evals skill from Hamel Husain and Shreya Shankar that guides AI coding agents in building product-specific AI evals. The linked GitHub repo collects these skills, and his post cites examples of evals improving results at Ramp, Shopify, Harvey and Cursor.
Original post · 1 min read
github.com/ai-evals-course/evals-skills
Lenny Rachitsky @lennysangithub.comGitHub - ai-evals-course/evals-skills: Skills that guide AI coding agents to help you build product-specific AI evals.Skills that guide AI coding agents to help you build product-specific AI evals. - ai-evals-course/evals-skillsEvals have been coming up more and more in my conversations with podcast guests and PM friends.
Nearly half of the 25 awesome PM job openings I shared last week ask for experience writing evals. And leading companies keep sharing what investing in evals bought them:
— @tryramp took its automatic receipt collection from 35% to 83% accuracy.
— @Shopify shipped an AI workflow builder that's 2.2x faster and 68% cheaper than the frontier-model setup it replaced.
— @harvey__ai rebuilt its AI contract reviewer, nearly doubling its internal quality score.
— @cursor_ai tuned its Auto Balance routing, …