Senior Site Reliability Engineer
Overview
In this role you will strengthen and scale Vanguard’s operational resiliency. You’ll work with product and platform teams to identify and remediate reliability risks, design standards and tooling, and lead blameless post-incident reviews. You’ll implement durable solutions, monitor production health, and evangelize resilience practices across the organization. The position offers meaningful impact on client outcomes through proactive reliability improvements and collaboration with cross-functional teams.
ResponsibilitiesEvaluate applications, platforms, and vendors for resiliency and riskDesign and implement enterprise reliability standards and processesLead blameless post-incident reviews for high-severity eventsPartner with product/platform teams to remediate reliability risksDevelop and evangelize new standards, tools, and frameworksTroubleshoot production issues and implement durable preventive solutionsParticipate in on-call rotation to support stabilityEvaluate and onboard resiliency toolingEngage in reliability communities of practice and share learningsContribute to strategic initiatives advancing operational maturity and resiliency posture
Key requirementsObservability platforms experience (Splunk, Honeycomb, CloudWatch, Dynatrace, AppDynamics)Strong understanding of SLIs, SLOs, and SLAs with dashboarding/reportingMonitoring and alerting design, anomaly detection, predictive alerting, synthetic monitoringAutomation and resilience practices (Python automation, RPA platforms e.g., Blue Prism, UiPath), chaos engineering, failure analysis (FMEA)Collaborative cross-functional partnershipProactive problem solvingCommunications and evangelism of standardsObservability platformsReliability metrics (SLIs/SLOs/SLAs)Monitoring & alerting