Deep Learning Scientist
Job Summary
We are seeking a Speech Deep Learning Scientist to develop and enhance advanced Speech AI solutions and improve conversational AI experiences for millions of users. This role will focus on speech synthesis model development, speech data processing, model evaluation, and driving innovation in speech technologies
Responsibilities
Train Speech Synthesis Mel Spectrogram and Vocoder models
Measure and benchmark model performance
Maintain TTS model evaluation systems
Analyze model accuracy and bias and recommend improvements and next steps
Improve processes for speech data processing, augmentation, filtering, and TTS training set preparation
Gather knowledge on TTS datasets for training and evaluation
Characterize performance and quality metrics across platforms for various Speech AI components
Collaborate with multiple teams on new product features and enhancements to existing products
Participate in code development and reviews, design document reviews, use case reviews, and test plan reviews
Help innovate, identify problems, recommend solutions, and perform triage in a collaborative team environment
Requirements
5+ years of relevant experience
Excellent programming skills in Python
Strong fundamentals in programming, optimization, and software design
Strong knowledge of machine learning and deep learning techniques, algorithms, and tools, including exposure to Autoregressive Speech Language Models, Audio Diffusion, and Flow Matching models
Knowledge of deep learning applications for speech synthesis, Large Language Models, and speech to speech translation
Hands on experience with speech technologies such as speech synthesis and voice cloning
Experience training speech models
Experience with the PyTorch deep learning framework
Exposure to speech digital signal processing and feature extraction techniques including FFT, MFCC, and Mel Spectrograms
General background with version control and code review tools such as Git, Gerrit, and GitLab
Strong collaborative and interpersonal skills with a proven ability to guide and influence within a dynamic matrix environment Preferred Qualifications
Native or near native fluency in a non English language including Spanish, Mandarin, German, Japanese, Russian, French, UK English, Arabic, Hindi, Korean, Italian, or Portuguese
Experience developing multilingual code switched TTS, voice cloning, and cross lingual voice cloning solutions
Experience developing WFST and neural network based Text Normalization and Inverse Text Normalization solutions
Experience working with G2P systems across multiple languages
Strong personal interest in learning, researching, and creating technologies related to foreign languages, linguistics, phonetics, phonology, and language technology
Comfortable working in a fast paced, highly collaborative, and dynamic environment
Strong C++ programming skills
Familiarity with GPU technologies including CUDA, CuDNN, and TensorRT
Experience deploying machine learning models on data center, cloud, and embedded systems
Tools and Technologies:
Python
PyTorch
Machine Learning
Deep Learning
Speech Synthesis
Voice Cloning
Speech to Speech Translation
Autoregressive Speech Language Models
Audio Diffusion
Flow Matching
FFT
MFCC
Mel Spectrograms
Git
Gerrit
GitLab
C++
CUDA
CuDNN
TensorRT
Pay: $50.00 - $55.00 per hour
Experience:
Deep learning: 5 years (Required)
autoregressive speech language models (LMs): 2 years (Required)
Work Location: Remote