1:04:05

Anthropic Head of Pretraining on Scaling Laws, Compute, and the Future of AI

From Y Combinator · Published Oct 23, 2025 · Watch on YouTube

TL;DR

The video is a conversation with Nick Joseph, Head of Pretraining at Anthropic, covering the fundamentals of pretraining (next-token prediction on internet-scale data), the empirical dominance of scaling laws over architectural details, and the shift toward balancing pretraining with post-training (RL) as compute continues to scale.

Key insights

  • Scaling laws (power law + constant) allow predictable loss reduction with more compute, but the field's critical feedback loop is: train model → sell product → buy more compute → train better model.
  • Autoregressive (next-token) pretraining won over other objectives (BERT, BART) mainly empirically; its ability to sample open-endedly enables product use, and compute thrown at any objective will like
  • At small scale, hyperparameter tuning matters little compared to throwing more compute; the real risk is hitting a deviation from the expected power law without knowing if it's a fundamental limit or

Want the full analysis - every claim cited to the second it was said?

This page only shows a teaser. Sign up to chat with the complete, cited breakdown of "Anthropic Head of Pretraining on Scaling Laws, Compute, and the Future of AI".