34:48

François Chollet: How We Get To AGI

From Y Combinator · Published Jul 20, 2025 · Watch on YouTube

TL;DR

François Chollet argues that scaling LLM pretraining alone does not produce fluid intelligence—only static memorized skills—as demonstrated by near‑zero accuracy on his ARC benchmark even after 50,000× scaling. The 2024 shift to test‑time adaptation (TTA) enabled human‑level performance on ARC1, but ARC2 reveals current TTA systems are still far below human efficiency.

Key insights

  • The cost of compute has fallen by two orders of magnitude per decade since 1940, enabling deep learning’s success but not fluid intelligence.
  • Scaling pretraining yields predictable benchmark gains (scaling laws) but fails on tasks requiring novel problem‑solving; ARC1 accuracy stayed near zero after 50,000× model scale‑up.
  • Test‑time adaptation (TTA) – models that modify themselves at inference – is the first paradigm to show genuine fluid intelligence, e.g., OpenAI’s O3 model reached human level on ARC1.

Want the full analysis - every claim cited to the second it was said?

This page only shows a teaser. Sign up to chat with the complete, cited breakdown of "François Chollet: How We Get To AGI".