34:06How To Build Generative AI Models Like OpenAI's Sora
From Y Combinator · Published Jul 20, 2024 · Watch on YouTube
TL;DR
OpenAI’s Sora generates long, physically consistent videos by combining transformer and diffusion models with temporal “spacetime patches.” Despite high compute costs, YC batch companies built competing foundation models with only $500k by hacking data (low‑res video, synthetic data, high‑quality small datasets), compute (Azure GPU cluster credits), and expertise (self‑teaching via papers).
Key insights
- Sora’s major advance over prior image models is its ability to correctly spell text in generated scenes.
- Long‑term visual consistency (e.g., same architectural style across a one‑minute clip) distinguishes Sora from earlier frame‑by‑frame generation.
- Under the hood, Sora combines a transformer (usually for text) with a diffusion model (for images) plus a temporal component for frame consistency, using “spacetime patches” as the video equivalent of
Want the full analysis - every claim cited to the second it was said?
This page only shows a teaser. Sign up to chat with the complete, cited breakdown of "How To Build Generative AI Models Like OpenAI's Sora".