Sr. Hardware Reliability Engineer, Infrastructure Reliability & Quality
Overview
In this role you drive reliability risk identification, assessment and mitigation for AWS data center infrastructure. You will lead root cause analysis of critical failures and push continuous improvements to maximize datacenter availability. You’ll collaborate with internal teams and external partners to guide product specifications, risk plans and vendor selections. This position rewards ownership, actionable execution, and clear communication in an open, collaborative environment. A strong impact will be shaping reliability for mission-critical infrastructure at scale.
Compensation / Benefitshealth insurance401(k) matchingRSUs and sign-on paymentspaid time offparential leaveflexible working culture
ResponsibilitiesDrive Design for Reliability (DFR) for new product designsQualification of third-party critical infrastructure equipment for AWS data centersOversee factory and site testing of third-party equipment across categories (Liquid Cooling, generator, chiller, air handler, etc.)Lead root cause analysis of field failures and validate remediation actionsProvide reliability data-driven recommendations on maintenance and replacementsOffer feedback to sourcing for vendor performanceAnalyze internal reliability data to drive high reliability at low costSupport DFMEAs as needed and develop end-of-life strategies for critical equipment
Key requirements8+ years of Reliability Engineering in high-reliability environments3+ years in accelerated life testing, stress analysis and finite element analysisExperience with external design/manufacturing supply chain partnersStrong problem-solving, communication and vendor management skillsWillingness to travel domestically and internationallyExperience with reliability risk identification and assessment from component to system levelKnowledge of statistical techniques and models for data analysisstrong communication skillsvendor managementproblem solvingreliability engineering methods and toolsDFR and DFMEAreliability risk assessment from component to system level