JOBSEARCHER

Senior Technical Program Manager, Machine Learning Infrastructure

Overview As Technical Program Manager for Machine Learning Infrastructure, you will oversee the portfolio spanning inference, serving, and endpoints to scale Cohere’s AI infrastructure for a growing user base. You’ll coordinate cross-functionally with Modeling and customer-facing teams, identify pain points, and implement processes that boost development velocity while meeting user needs. You’ll foster a culture of continuous improvement, drive prioritized delivery, and partner with senior leaders to solve strategic problems that impact the company. This role offers exposure to cutting-edge AI tech and a dynamic, fast-paced environment with meaningful impact. Compensation / Benefitsweekly lunch stipendhealth and dental benefitsRRSP matching, 401K, Pension6 weeks of paid vacationparential leave top-up (up to 6 months)home office stipend ResponsibilitiesManage the ML infrastructure program portfolio (inference, serving, efficiency, endpoints) to scale for internal and external usersLead end-to-end program coordination with cross-functional partners (Modeling, customer-facing teams)Identify bottlenecks and establish processes that enable development while serving user needsPromote continuous improvement and incident management practices with timely root-cause analysis and fixesPrioritize overlapping projects to align with top company prioritiesCollaborate with stakeholders to set timelines, deliverables, budgets, and scopeDeliver clear updates to engineering, leadership, and non-technical audiencesAct as a tactical and strategic partner to senior program leads, addressing both execution and high-level strategy Key requirements5+ years of Technical Program Management experience in Machine Learning Infrastructure (inference, serving, efficiency, endpoints)Deep technical knowledge of ML infrastructure design and engineering best practicesComfort in fast-paced, low-structure environments and willingness to do hands-on work if neededExperience advocating for infrastructure and modeling teams to build scalable, high-performance ML systemspragmatic and execution-orientedorganized and cross-functional collaborationadaptable in fast-changing environmentsML infrastructure design and implementationmodel inference and servingsystem efficiency improvements