Sr Staff DevOps Engineer
Overview
In this role you will strengthen the reliability and scale of the Vantor Hub platform. You will work with cross-functional teams to build, test, and release software securely across multiple AWS accounts, balancing speed with quality. You will lead reliability initiatives, shaping production readiness and automation. This is a mission-driven opportunity to help customers see the world differently while advancing platform maturity.
Compensation / Benefits401(k) with company matchmental health resourcesstudent loan repayment assistanceadoption reimbursementpet insuranceincentive eligibility
ResponsibilitiesLead reliability engineering efforts for the Vantor Hub platform and related infrastructure servicesDesign, implement, and maintain scalable CI/CD pipelines for software and infrastructure deliveryBuild, operate, and improve cloud infrastructure automation using Terraform, CloudFormation, Kubernetes, Docker, and AWS-native servicesTroubleshoot complex infrastructure, deployment, networking, performance, reliability, and production issuesImprove service availability, scalability, maintainability, reliability, security, and overall operational readinessDesign and operate highly available, resilient, observable, and secure infrastructure across commercial and government AWS environmentsPartner with engineering teams to improve deployment safety, observability, production readiness, and operational supportParticipate in a team on-call rotation (approx. one week every twelve weeks) for production reliability and incident responseLeverage AI development tools to accelerate software design, testing, and documentation while maintaining reliability and securityContribute to shared engineering standards for responsible AI-assisted development, including validation, documentation, and review patternsUse AI-assisted tooling to speed up automation, documentation, runbooks, tests, and incident analysis with peer review and security validation
Key requirementsBachelor's Degree in Software Engineering or Computer Science (or related field) or equivalent experience8+ years in DevOps, Platform Engineering, Site Reliability Engineering, or cloud operations with production ownershipExperience owning, operating, and improving production systems in cloud environmentsStrong SRE practices: SLIs/SLOs, error budgets, incident response, post-incident reviews, reliability metrics, automationStrong AWS proficiency, including multi-account production workloadsStrong Docker and Kubernetes experienceStrong IaC experience with Terraform or CloudFormationScripting ability in Python, Bash, or PowerShellSolid networking understandingEffective communication and collaboration with team membersU.S. citizen and willingness to obtain a U.S. Government security clearancestrong communicationcollaborationability to mentor othersAWS multi-account environmentsDockerKubernetes