{"schemaVersion":"jobsearcher.job.v1","id":"10a526c61d5c4fc2b0e60fd5","url":"https://jobsearcher.com/jobs/10a526c61d5c4fc2b0e60fd5","canonicalUrl":"https://jobsearcher.com/jobs/10a526c61d5c4fc2b0e60fd5","title":"Deep Learning Scientist","description":"Job Summary\nWe are seeking a Speech Deep Learning Scientist to develop and enhance advanced Speech AI solutions and improve conversational AI experiences for millions of users. This role will focus on speech synthesis model development, speech data processing, model evaluation, and driving innovation in speech technologies\nResponsibilities\nTrain Speech Synthesis Mel Spectrogram and Vocoder models\nMeasure and benchmark model performance\nMaintain TTS model evaluation systems\nAnalyze model accuracy and bias and recommend improvements and next steps\nImprove processes for speech data processing, augmentation, filtering, and TTS training set preparation\nGather knowledge on TTS datasets for training and evaluation\nCharacterize performance and quality metrics across platforms for various Speech AI components\nCollaborate with multiple teams on new product features and enhancements to existing products\nParticipate in code development and reviews, design document reviews, use case reviews, and test plan reviews\nHelp innovate, identify problems, recommend solutions, and perform triage in a collaborative team environment\nRequirements\n5+ years of relevant experience\nExcellent programming skills in Python\nStrong fundamentals in programming, optimization, and software design\nStrong knowledge of machine learning and deep learning techniques, algorithms, and tools, including exposure to Autoregressive Speech Language Models, Audio Diffusion, and Flow Matching models\nKnowledge of deep learning applications for speech synthesis, Large Language Models, and speech to speech translation\nHands on experience with speech technologies such as speech synthesis and voice cloning\nExperience training speech models\nExperience with the PyTorch deep learning framework\nExposure to speech digital signal processing and feature extraction techniques including FFT, MFCC, and Mel Spectrograms\nGeneral background with version control and code review tools such as Git, Gerrit, and GitLab\nStrong collaborative and interpersonal skills with a proven ability to guide and influence within a dynamic matrix environment Preferred Qualifications\nNative or near native fluency in a non English language including Spanish, Mandarin, German, Japanese, Russian, French, UK English, Arabic, Hindi, Korean, Italian, or Portuguese\nExperience developing multilingual code switched TTS, voice cloning, and cross lingual voice cloning solutions\nExperience developing WFST and neural network based Text Normalization and Inverse Text Normalization solutions\nExperience working with G2P systems across multiple languages\nStrong personal interest in learning, researching, and creating technologies related to foreign languages, linguistics, phonetics, phonology, and language technology\nComfortable working in a fast paced, highly collaborative, and dynamic environment\nStrong C++ programming skills\nFamiliarity with GPU technologies including CUDA, CuDNN, and TensorRT\nExperience deploying machine learning models on data center, cloud, and embedded systems\nTools and Technologies:\nPython\nPyTorch\nMachine Learning\nDeep Learning\nSpeech Synthesis\nVoice Cloning\nSpeech to Speech Translation\nAutoregressive Speech Language Models\nAudio Diffusion\nFlow Matching\nFFT\nMFCC\nMel Spectrograms\nGit\nGerrit\nGitLab\nC++\nCUDA\nCuDNN\nTensorRT\nPay: $50.00 - $55.00 per hour\nExperience:\nDeep learning: 5 years (Required)\nautoregressive speech language models (LMs): 2 years (Required)\nWork Location: Remote","company":"Cloud Destinations","rawCompany":"cloud destinations","city":"Remote","state":"OR","isRemote":false,"isActive":false,"createdAt":"2026-08-05T00:14:29.928Z","occupations":[{"code":"29-1127.00","title":"Speech-Language Pathologists","slug":"speech-language-pathologists"},{"code":"15-2051.00","title":"Data Scientists","slug":"data-scientists"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"}],"industries":[{"code":"541990","title":"All Other Professional, Scientific, and Technical Services","slug":"all-other-professional-scientific-and-technical-services"},{"code":"541930","title":"Translation and Interpretation Services","slug":"translation-and-interpretation-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Deep Learning Scientist","description":"Job Summary\nWe are seeking a Speech Deep Learning Scientist to develop and enhance advanced Speech AI solutions and improve conversational AI experiences for millions of users. This role will focus on speech synthesis model development, speech data processing, model evaluation, and driving innovation in speech technologies\nResponsibilities\nTrain Speech Synthesis Mel Spectrogram and Vocoder models\nMeasure and benchmark model performance\nMaintain TTS model evaluation systems\nAnalyze model accuracy and bias and recommend improvements and next steps\nImprove processes for speech data processing, augmentation, filtering, and TTS training set preparation\nGather knowledge on TTS datasets for training and evaluation\nCharacterize performance and quality metrics across platforms for various Speech AI components\nCollaborate with multiple teams on new product features and enhancements to existing products\nParticipate in code development and reviews, design document reviews, use case reviews, and test plan reviews\nHelp innovate, identify problems, recommend solutions, and perform triage in a collaborative team environment\nRequirements\n5+ years of relevant experience\nExcellent programming skills in Python\nStrong fundamentals in programming, optimization, and software design\nStrong knowledge of machine learning and deep learning techniques, algorithms, and tools, including exposure to Autoregressive Speech Language Models, Audio Diffusion, and Flow Matching models\nKnowledge of deep learning applications for speech synthesis, Large Language Models, and speech to speech translation\nHands on experience with speech technologies such as speech synthesis and voice cloning\nExperience training speech models\nExperience with the PyTorch deep learning framework\nExposure to speech digital signal processing and feature extraction techniques including FFT, MFCC, and Mel Spectrograms\nGeneral background with version control and code review tools such as Git, Gerrit, and GitLab\nStrong collaborative and interpersonal skills with a proven ability to guide and influence within a dynamic matrix environment Preferred Qualifications\nNative or near native fluency in a non English language including Spanish, Mandarin, German, Japanese, Russian, French, UK English, Arabic, Hindi, Korean, Italian, or Portuguese\nExperience developing multilingual code switched TTS, voice cloning, and cross lingual voice cloning solutions\nExperience developing WFST and neural network based Text Normalization and Inverse Text Normalization solutions\nExperience working with G2P systems across multiple languages\nStrong personal interest in learning, researching, and creating technologies related to foreign languages, linguistics, phonetics, phonology, and language technology\nComfortable working in a fast paced, highly collaborative, and dynamic environment\nStrong C++ programming skills\nFamiliarity with GPU technologies including CUDA, CuDNN, and TensorRT\nExperience deploying machine learning models on data center, cloud, and embedded systems\nTools and Technologies:\nPython\nPyTorch\nMachine Learning\nDeep Learning\nSpeech Synthesis\nVoice Cloning\nSpeech to Speech Translation\nAutoregressive Speech Language Models\nAudio Diffusion\nFlow Matching\nFFT\nMFCC\nMel Spectrograms\nGit\nGerrit\nGitLab\nC++\nCUDA\nCuDNN\nTensorRT\nPay: $50.00 - $55.00 per hour\nExperience:\nDeep learning: 5 years (Required)\nautoregressive speech language models (LMs): 2 years (Required)\nWork Location: Remote","datePosted":"2026-08-05T00:14:29.928Z","dateModified":"2026-08-05T00:14:29.928Z","hiringOrganization":{"@type":"Organization","name":"Cloud Destinations","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Remote","addressRegion":"OR","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"10a526c61d5c4fc2b0e60fd5"},"url":"https://jobsearcher.com/jobs/10a526c61d5c4fc2b0e60fd5"}}