30:29

Baidu's AI Lab Director on Advancing Speech Recognition and Simulation

From Y Combinator · Published Jul 22, 2018 · Watch on YouTube

TL;DR

Baidu’s Silicon Valley AI lab, directed by Adam Coates, is mission-oriented toward creating AI technologies that impact at least 100 million people. The lab achieved superhuman speech recognition (Deep Speech) by scaling neural networks with massive data (10,000–20,000+ hours of audio) and moving from hand-engineered modules to end-to-end deep learning.

Key insights

  • Speech recognition became superhuman for Mandarin short queries by scaling up standard deep learning models with far more data and computational investment than prior approaches.
  • Crowdsourcing services that pay people to read books aloud can provide cheap, large-scale audio data for English training sets (~10,000–20,000 hours).
  • Deep learning eliminates the need for hand-engineered intermediate representations (e.g., phonetic decoders) – the network learns to map audio to characters (or phonemes) directly.

Want the full analysis - every claim cited to the second it was said?

This page only shows a teaser. Sign up to chat with the complete, cited breakdown of "Baidu's AI Lab Director on Advancing Speech Recognition and Simulation".