Anthropic’s first technical PM on token maxing, the jagged edge, and living in the future | Dianne Penn

Summary of Anthropic’s first technical PM on token maxing, the jagged edge, and living in the future | Dianne Penn

by Lenny Rachitsky

1h 33mJuly 26, 2026

Overview of Anthropic’s first technical PM on token maxing, the jagged edge, and living in the future | Dianne Penn

This episode is a deep dive into how Anthropic builds frontier AI products, from the company’s scrappy early days to the current era of rapid model leaps, agentic workflows, and tighter safety constraints. Dianne Penn, Anthropic’s head of product for research and labs, explains how the team thinks about product discovery in a world where capabilities emerge unpredictably, why “evals are the new PRDs,” and how products like Claude Code helped unlock the value of models like Opus 4.5. A major theme throughout is that the AI product role is becoming more hands-on, more experimental, and more tied to real model behavior than traditional software PM work.

Key Inflection Points in Anthropic’s Growth

Early startup phase

  • When Dianne joined in 2023, Anthropic was tiny: only a handful of product engineers and one engineer on the API business.
  • The company was still searching for its identity beyond “another chatbot.”
  • A defining early moment was Golden Gate Claude, a quirky interpretability demo that turned a research finding into a public product experience in about 24 hours.

Claude Opus 3

  • Opus 3 was an important internal milestone because it proved Anthropic could train a true frontier model.
  • The work built deep trust across research, inference, fine-tuning, and product teams.
  • It also surfaced a key insight: coding was becoming a major use case, especially long-form code generation.

Claude Opus 4.5 + Claude Code

  • Dianne describes Opus 4.5 as a bigger inflection because the model and the product experience matured together.
  • The combination of a capable model plus a strong “vehicle” like Claude Code created a new kind of adoption loop.
  • Her core thesis: frontier products are needed to make frontier models feel real.

How Anthropic Thinks About Product and Research

“Evals are the new PRDs”

  • The team uses evals as the primary way to define user problems and track improvements.
  • Instead of writing a traditional product spec first, they often:
    • analyze user feedback,
    • reproduce failures,
    • identify the exact model behavior that caused them,
    • and turn that into an eval set.
  • This is especially important because AI issues are often vague at first (“Claude hallucinated”) and need to be translated into actionable model-level problems.

Product work is tightly coupled to model behavior

  • The team looks at whether failures come from:
    • tool use,
    • search/knowledge,
    • alignment,
    • formatting/schema issues,
    • or reasoning quality.
  • Example: one of the earliest big wins was getting Claude to reliably output JSON, which was crucial for agents and tool use.

PRDs still matter, but differently

  • Anthropic still uses PRDs for:
    • larger organizational alignment,
    • model launches,
    • stakeholder coordination,
    • and ambiguous new product areas.
  • But for many research-driven problems, evals are now the more important planning artifact.

Labs: Anthropic’s Engine for Big Bets

What Labs does

  • Labs exists to explore discontinuous, high-upside bets that may not fit the core roadmap.
  • It’s where products like:
    • Claude Code,
    • Skills,
    • Cloud Design,
    • MCP,
    • and other experimental ideas often start.

Why it works

  • Small pods, often with just one engineer initially.
  • Strong culture of autonomy and experimentation.
  • Team members are selected for zero-to-one energy and comfort with ambiguity.
  • The team is willing to let ideas percolate and revisit them across model generations if needed.

The Jagged Edge: Why AI Progress Feels Uneven

  • Dianne emphasizes that AI capability does not improve uniformly.
  • Models can suddenly become much better at one thing while still having rough edges elsewhere.
  • This is why:
    • benchmarking matters,
    • red-teaming matters,
    • and product teams must stay adaptable.

Key implication

  • You can’t assume you know exactly what the next model will unlock.
  • The job is to continuously ask:
    “If Claude 8 exists, what changes in what users do?”

Token Maxing, Experimentation, and “Living in the Future”

  • Dianne agrees that heavy model usage today gives people a preview of how they’ll work in the future.
  • But she reframes “token maxing” as a means, not the end.
  • The real goal is more experimentation and better outcomes.

Her advice

  • Use the models a lot.
  • Don’t just consume them—build with them.
  • The best ideas come from hands-on interaction, not abstract strategy.

A key cultural insight

  • Experimentation is not an individual sport.
  • Anthropic intentionally encourages people to work in public, share prototypes, and discover use cases together.

What Makes a Strong AI Product Manager Now

Traits that matter most

  • First principles thinking
  • Hands-on familiarity with the models
  • Comfort sweating the tokens
  • Curiosity and experimentation
  • Ability to translate messy feedback into actionable evals
  • Ambition paired with practical judgment

What’s changed

  • PMs can’t just “manage” AI products from a distance.
  • To be effective, they need to:
    • use the models directly,
    • ship with them,
    • understand their failure modes,
    • and have a real point of view.

What hasn’t changed

  • User empathy still matters.
  • Good product work still means:
    • understanding user pain,
    • defining the right problem,
    • and making sure the team is moving in the same direction.

Why Claude’s Alignment Makes It Better

  • Dianne argues that Claude’s safety and alignment work does not just constrain the model—it improves it.
  • The goal is to make Claude a thinking partner, not a compliant echo chamber.
  • A good AI should:
    • push back when necessary,
    • help sharpen your thinking,
    • and improve your outcomes.

What Human Judgment Will Still Be Valuable For

Dianne believes humans will remain especially important in areas like:

  • Judgment
  • Proactivity
  • Persistence
  • Subject-matter expertise
  • Deciding what to build
  • Choosing the right level of ambition

She thinks software engineering is already being transformed, while areas like biology and life sciences are still earlier on the curve.

Staying Sharp, Avoiding Burnout, and Raising Kids in the AI Era

Staying sane at Anthropic

  • The work is intense, but the team structure is collaborative and supportive.
  • She credits:
    • strong team culture,
    • “radical ownership,”
    • and people stepping in for each other during launches.

Parenting and the future

  • For her kids, she emphasizes:
    • curiosity,
    • persistence,
    • inner voice,
    • and forming an opinion before outsourcing thinking to AI.
  • One memorable parenting idea she mentions: don’t rush kids to fully automated answers too early—let them struggle a bit.

Notable Recommendations from the Lightning Round

Books

  • How to Raise an Adult
  • Incorruptible by Eric Ries

Show

  • Fallout on Amazon Prime

Favorite product

  • Claude Tag / internal tagging-style workflows for working with Claude in a more agentic way

Final Takeaway

Dianne’s core message is that the AI era rewards people who are:

  • close to the technology,
  • willing to experiment,
  • able to think from first principles,
  • and humble enough to learn from the models and from each other.

Her view of the future is optimistic but demanding:
the best teams will be the ones that combine human judgment with relentless hands-on use of AI.