Answer extracted from the ACQ2 by Acquired podcast — listen to the full episode below.
Build personal test harnesses called 'evals' — structured folders of prompts with expected and judged results — that you run against every new AI model to identify exactly where its capabilities end and what thresholds it has crossed. This approach gives you a repeatable, systematic method to benchmark each model release against your previous tests and determine whether the technology has genuinely advanced.
Instead of vaguely testing an AI model or relying on marketing claims, create a real testing apparatus. Think of your evals as batch jobs that repeatedly test the same capabilities and judge the results consistently — much like how unit tests transformed software development in the late 1990s and early 2000s.
The premise is simple: write a set of prompts that probe specific capabilities, define what a correct answer looks like, and then run every new model against that same set. This removes guesswork and emotion from evaluation.
As Lütke explains in the podcast, the core insight is that software now exists that no human wrote — and the only way to truly understand it is to interview it systematically. That interview process is your eval framework.
The real value of personal test harnesses is that they help you spot the exact moment a new model crosses into a capability threshold you care about. When Claude or GPT-5 arrives next month, you don't start from scratch — you run it against your existing evals to see what's changed.
This approach also surfaces the edges: where a model is still limited, what it still fails at, and what you still can't reliably ask it to do. That knowledge becomes actionable for product teams building with AI, as detailed in this episode of ACQ2 by Acquired.
"It is a privilege of a lifetime to be part of another platform shift."
Tobi Lütke — Founder and CEO of Shopify. Under his leadership since its founding, Shopify has grown from a single e-commerce platform into a nearly $200 billion company with approximately $10 billion in annual revenue, navigating multiple market cycles and platform transitions to position the company for emerging technologies like AI.
The deeper reason Lütke spends so much time interviewing new models is to understand how to make them work as idealized, nonjudgmental teachers for users. Without structured evals, you can't answer that question reliably. With them, you have a repeatable method for learning what each generation of AI can genuinely do — and what still lies ahead.
At age 15, Sorkin created Sports Page Magazine at Scarsdale High School, which published sports journalism written by high school students and targeted the local community.
In 1995, Sorkin arrived at the New York Times office with a visitor pass, wearing a suit and tie, and performed clerical tasks like copying and filing while learning the business from the ground up.
When launching DealBook during the dot-com bust, the New York Times thought the total addressable market was 30,000 free subscribers. Today, DealBook has grown to serve hundreds of thousands of readers.