JOBSEARCHER

Staff Technical Program Manager, Site Reliability Engineering

MongoDBBrooklyn, NYL7 ManagerSeptember 15th, 2026
Overview As a TPM for SRE, you partner with SRE leaders to scale MongoDB’s cloud platform, driving program execution and reliability practices. You coordinate across US and EMEA teams to deliver smoother launches, clearer roadmaps, and stronger reliability metrics. You’ll shape scalable processes and reduce hero culture by empowering teams to operate independently. This role offers remote work from the East Coast and the chance to impact platform reliability at scale. Compensation / Benefitsremote work optionsequity and employee stock purchase programflexible paid time off20 weeks fully-paid gender-neutral parental leavefertility and adoption assistancehealth benefits including mental health support ResponsibilitiesDefine and drive program planning and execution with SRE engineers and leaders; manage dependencies and track work in Jira to ensure on-time deliveryLead production reliability efforts, including change management, launch readiness, and defining operational SLOs/SLIsCoordinate cross-functional efforts with Security, Compliance, Cloud platform, and other engineering teams; drive incident response and follow-throughDesign lightweight frameworks and processes to enable scalable, reliable delivery and reduce reliance on individual heroes Key requirements8+ years in technical program management, engineering management, or similar role with software engineering teamsProven track record leading large-scale, cross-team platform initiatives through ambiguity and changeStrong knowledge of production change management, SDLC, and reliability metrics (SLOs/SLIs)Skilled at shaping roadmaps and managing dependenciesAbility to query and interpret metrics, logs, or data sources to inform decisions and communicate riskExcellent communicator—clear, concise, calm—across engineers, partners, and executivesLow-ego, highly collaborative, ownership mindset for end-to-end problemscollaborationclear communicationownership of hard problemsKubernetes (Nice to Have)cloud networking (Nice to Have)observability stacks (metrics, logs, tracing, alerting)