3 Takeaways™ The answer lives in this podcast

What was the state of ChatGPT when it first launched, and what changed to make modern AI so much more capable?

When ChatGPT first rolled out, OpenAI had just 200 employees and the initial model — likely GPT-3.5 — produced what Katrina Mulligan describes as junior high school level outputs. Two breakthroughs transformed this: the discovery of scaling laws, showing that multiplying compute and data by 10x yields predictably better performance; and teaching models to use tools and agentic harnesses, enabling them to search the web, access emails, navigate document repositories, and perform computer use.

200 people, a mediocre model — and a discovery that changed the math

The early version of ChatGPT was genuinely underwhelming by today's standards. Mulligan is direct about it: the output quality was roughly what you'd expect from a middle schooler. OpenAI was a small team of 200 people, operating far from the infrastructure scale the company has since reached.

What broke the ceiling was the formalization of scaling laws. Researchers found that increasing compute by 10x and training data by 10x produced predictably better model performance — not incrementally, but reliably and measurably. This wasn't luck. It was a reproducible law, and it gave OpenAI a clear roadmap for capability growth. As Mulligan explains in 3 Takeaways™, this discovery fundamentally reframed what was possible.

Scaling laws, defined: The empirical finding that AI model performance improves in a predictable, measurable way when both compute power and training data are multiplied — typically by a factor of 10x. This discovery gave researchers a reliable mechanism for capability improvement, rather than trial and error.

The second leap: giving the model hands

Raw language capability was only part of the story. The second transformation came from teaching models to use tools — what the field now calls agentic harnesses. Rather than simply generating text, models gained the ability to run searches, read and send emails, navigate document repositories, and operate computers directly.

This shift from passive language generation to active task completion is what makes today's AI qualitatively different from the 2022 launch version. It's a distinction Mulligan draws clearly when discussing AI's trajectory in this episode — the model didn't just get smarter, it got hands.

The practical consequence is real and measurable. Mulligan herself estimates she is at least 30% more effective as a leader and manager than she was a year ago, directly attributing this to how she has matured her use of these agentic tools. That's not a marketing claim — it's a firsthand account from someone operating at the intersection of national security and frontier AI, as she details in 3 Takeaways™.

"I honestly think that I am at least 30% more effective, maybe more than I was a year ago, because of how I've matured my use of these tools. I'm a better leader. I'm a better manager."

Katrina Mulligan — Head of National Security Partnerships, OpenAI.
Mulligan spent years in the most hierarchical corridors of American government — senior roles at the Department of Defense, the National Security Council, and the Department of Justice, including as the number two at the Pentagon overseeing special operations. She then made what she calls the hardest professional transition of her career: moving to OpenAI, the exact organizational opposite of everything she had known. That dual vantage point — inside both Washington's national security apparatus and frontier AI development — makes her perspective on AI capability uniquely grounded. She speaks with authority on 3 Takeaways™ not as an observer, but as a practitioner.

See also

What is recursive self-improvement in AI and how close is it to becoming reality?

Mulligan explains that recursive self-improvement is the idea that AI models would eventually become good enough to conduct the research that improves themselves — a threshold that, if crossed, would dramatically accelerate AI capability in ways that are difficult to predict or control.

What makes the AI revolution fundamentally different from every previous technology revolution in American history?

According to Mulligan, this is the first time in American history that a technology of this consequence is being developed exclusively by the private sector — rather than through government-led programs like GPS or the Human Genome Project.

What does Katrina Mulligan mean when she says the world is looking at a calendar while OpenAI is looking at a watch?

Mulligan uses this metaphor to explain that the unit of time at OpenAI is fundamentally compressed compared to anywhere else: she says internally that a month at OpenAI feels like a quarter anywhere else, meaning three to four months equates to roughly a year of experience elsewhere.

Key takeaways

The full conversation with Katrina Mulligan — including her take on America's AI lead over China and what governments must do now — is available on Listenly.

Listen to the episode on Listenly
Entities: OpenAI, ChatGPT, GPT-3.5, Scaling laws, Agentic AI, Department of Defense, National Security Council, Department of Justice, Pentagon, Three Takeaways, Katrina Mulligan, Lynn Thoman, Listenly