30:29Baidu's AI Lab Director on Advancing Speech Recognition and Simulation
From Y Combinator · Published Jul 22, 2018 · Watch on YouTube
TL;DR
Baidu’s Silicon Valley AI lab, directed by Adam Coates, is mission-oriented toward creating AI technologies that impact at least 100 million people. The lab achieved superhuman speech recognition (Deep Speech) by scaling neural networks with massive data (10,000–20,000+ hours of audio) and moving from hand-engineered modules to end-to-end deep learning.
Key insights
- Speech recognition became superhuman for Mandarin short queries by scaling up standard deep learning models with far more data and computational investment than prior approaches.
- Crowdsourcing services that pay people to read books aloud can provide cheap, large-scale audio data for English training sets (~10,000–20,000 hours).
- Deep learning eliminates the need for hand-engineered intermediate representations (e.g., phonetic decoders) – the network learns to map audio to characters (or phonemes) directly.
Want the full analysis - every claim cited to the second it was said?
This page only shows a teaser. Sign up to chat with the complete, cited breakdown of "Baidu's AI Lab Director on Advancing Speech Recognition and Simulation".