{"schemaVersion":"jobsearcher.job.v1","id":"2c46f320c0d370d1994b3dca","url":"https://jobsearcher.com/jobs/2c46f320c0d370d1994b3dca","canonicalUrl":"https://jobsearcher.com/jobs/2c46f320c0d370d1994b3dca","title":"Audio Inference Engineer, Model Efficiency","description":"Overview\nAs a member of Cohere's Model Efficiency team, you will advance audio model serving performance—focusing on latency, throughput, and quality for real-time, streaming workloads. You collaborate with training and serving infra teams to ensure seamless deployment of audio models. You will solve bottlenecks and implement innovative optimizations that scale audio inference. This role offers the chance to influence core ML systems at a fast-growing AI company shaping the future of audio applications.\n\nCompensation / Benefitsopen and inclusive culturehealth and dental benefits6 weeks of vacationparental leave top-upremote-friendly with office stipendsco-working stipend\nResponsibilitiesDevelop and optimize high-performance audio or ML inference systemsImprove latency, throughput, and quality for audio streaming workloadsDiagnose bottlenecks and implement practical, creative solutionsCollaborate with training and serving infrastructure teams for end-to-end integrationSupport real-time and streaming audio inference workflows\nKey requirementsSignificant experience developing high-performance audio or machine learning inference systemsProficiency with C++ and PythonHands-on experience with deep learning models for audio, speech, or language applicationsBias for action and a results-oriented mindsetbias for actionresults-orientedcross-functional collaborationGPU programminglow-level system optimizationmodel parallelization techniques over multiple GPUs","company":"Cohere","rawCompany":"cohere","city":"White Plains","state":"NY","isRemote":false,"isActive":false,"createdAt":"2026-09-15T04:58:51.562Z","occupations":[{"code":"29-1181.00","title":"Audiologists","slug":"audiologists"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"}],"industries":[{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Audio Inference Engineer, Model Efficiency","description":"Overview\nAs a member of Cohere's Model Efficiency team, you will advance audio model serving performance—focusing on latency, throughput, and quality for real-time, streaming workloads. You collaborate with training and serving infra teams to ensure seamless deployment of audio models. You will solve bottlenecks and implement innovative optimizations that scale audio inference. This role offers the chance to influence core ML systems at a fast-growing AI company shaping the future of audio applications.\n\nCompensation / Benefitsopen and inclusive culturehealth and dental benefits6 weeks of vacationparental leave top-upremote-friendly with office stipendsco-working stipend\nResponsibilitiesDevelop and optimize high-performance audio or ML inference systemsImprove latency, throughput, and quality for audio streaming workloadsDiagnose bottlenecks and implement practical, creative solutionsCollaborate with training and serving infrastructure teams for end-to-end integrationSupport real-time and streaming audio inference workflows\nKey requirementsSignificant experience developing high-performance audio or machine learning inference systemsProficiency with C++ and PythonHands-on experience with deep learning models for audio, speech, or language applicationsBias for action and a results-oriented mindsetbias for actionresults-orientedcross-functional collaborationGPU programminglow-level system optimizationmodel parallelization techniques over multiple GPUs","datePosted":"2026-09-15T04:58:51.562Z","dateModified":"2026-09-15T04:58:51.562Z","hiringOrganization":{"@type":"Organization","name":"Cohere","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"White Plains","addressRegion":"NY","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"2c46f320c0d370d1994b3dca"},"url":"https://jobsearcher.com/jobs/2c46f320c0d370d1994b3dca"}}