JOBSEARCHER

Site Reliability Engineer

Job Title: Site Reliability EngineerLocation: Remote (US, EST hours required)Interview Process: 2x InterviewsClient OverviewThis is a global materials science leader with a 170 plus year history, operating dozens of manufacturing and R&D sites worldwide. This role sits within a research and development group supporting advanced computing infrastructure behind ongoing materials science innovation.Top 3 SkillsKubernetes cluster operations and management, including provisioning, upgrades, and troubleshooting across on premises and cloud environmentsRancher for Kubernetes platform managementLinux systems administration, including performance tuning, scripting, and networkingWhat You'll DoMaintain and enhance Kubernetes platforms across on premises and cloud environmentsSupport provisioning, upgrades, troubleshooting, and lifecycle management of Kubernetes clusters managed through RancherProvide deep Linux systems administration support, including performance tuning, troubleshooting, and automationDevelop and maintain infrastructure as code solutions to standardize and automate platform deploymentSupport and improve GitOps workflows using ArgoCD to manage cluster and application configurationCollaborate with developers, scientists, and infrastructure teams to deliver reliable platform servicesIdentify opportunities to improve platform resilience, observability, security, and maintainabilityWhat We Need From You5 plus years of professional experience in site reliability engineering, platform engineering, DevOps, or systems engineeringHands on experience operating Kubernetes platforms in production environments, both on premises and cloud basedExperience with Rancher for Kubernetes cluster managementStrong Linux systems administration skills, including troubleshooting, scripting, and system performance analysisExperience implementing infrastructure as code solutions for platform provisioning and lifecycle managementPreferred/BonusBachelor's degree in Computer Science, Software Engineering, Information Technology, or related fieldExperience with Cluster API (CAPI)Experience with hybrid infrastructure spanning on premises and public cloud (AWS, Azure, GCP)Familiarity with Kubernetes observability, logging, monitoring, and alerting toolingExperience supporting scientific research, high performance computing, or computational science environmentsExperience with Agile teams (Scrum, Kanban)Elevait Solutions was founded by veterans who believe that how you treat people is the only thing that actually matters in this industry. We're a team that stays in your corner before, during, and after placement. We show up for the communities we work in, we tell you the truth, and we work hard to make sure every placement is a good fit for both sides. If that sounds like the kind of team you want behind you, we'd like to talk.