JOBSEARCHER

Audio Inference Engineer, Model Efficiency

CohereBrooklyn, NYL4 MidSeptember 15th, 2026
Overview As a member of Cohere's Model Efficiency team, you will advance audio model serving performance—focusing on latency, throughput, and quality for real-time, streaming workloads. You collaborate with training and serving infra teams to ensure seamless deployment of audio models. You will solve bottlenecks and implement innovative optimizations that scale audio inference. This role offers the chance to influence core ML systems at a fast-growing AI company shaping the future of audio applications. Compensation / Benefitsopen and inclusive culturehealth and dental benefits6 weeks of vacationparental leave top-upremote-friendly with office stipendsco-working stipend ResponsibilitiesDevelop and optimize high-performance audio or ML inference systemsImprove latency, throughput, and quality for audio streaming workloadsDiagnose bottlenecks and implement practical, creative solutionsCollaborate with training and serving infrastructure teams for end-to-end integrationSupport real-time and streaming audio inference workflows Key requirementsSignificant experience developing high-performance audio or machine learning inference systemsProficiency with C++ and PythonHands-on experience with deep learning models for audio, speech, or language applicationsBias for action and a results-oriented mindsetbias for actionresults-orientedcross-functional collaborationGPU programminglow-level system optimizationmodel parallelization techniques over multiple GPUs