Infrastructure/Cloud DevOps - SRE
Infrastructure/Cloud DevOps - SREW2 ContractPay Rate: $55 - $65 per hourLocation: Cupertino, CA - Remote RoleJob Summary:We are looking for a highly motivated DevOps / Site Reliability Engineer to support large-scale Kubernetes-based infrastructure and platform operations. This role is focused on building, automating, and operating highly reliable systems that power critical engineering platforms and services.Duties and Responsibilities:Design, build, automate, and support scalable Kubernetes-based platforms and servicesOperate and troubleshoot production environments running at scaleDevelop automation and tooling to improve operational efficiency and reliabilityMonitor platform health, performance, and availability using observability toolingTroubleshoot infrastructure, application, and networking issues across distributed systemsWork closely with engineering teams to improve deployment, reliability, and scalability practicesParticipate in operational support, incident response, and root cause analysisImprove CI/CD workflows and deployment automationDrive operational excellence through documentation, automation, and process improvementsTake ownership of projects and independently drive deliverables to completionRequirements and Qualifications:Strong hands-on experience with Kubernetes platforms such as:EKSGKEAKS or similarExperience running and supporting applications on Kubernetes at scaleStrong understanding of containerized infrastructure and distributed systemsExperience with monitoring and observability tools, preferably:GrafanaPrometheusExperience with CI/CD pipelines and deployment automationExperience with Splunk logging, log analysis, and troubleshootingStrong scripting and automation experience using Python and/or GolangExperience troubleshooting production systems under pressureStrong communication and collaboration skillsSelf-starter mentality with strong ownership and accountabilityPreferred QualificationsExperience operating Ray clusters/servicesStrong networking and troubleshooting experienceExperience with cloud infrastructure and platform servicesExperience with Infrastructure as Code and automation frameworksExperience supporting high-scale production systemsFamiliarity with SRE principles and operational best practicesDesired Skills and ExperienceKubernetes, Amazon EKS, Google Kubernetes Engine (GKE), Azure Kubernetes Service (AKS), DevOps, Site Reliability Engineering (SRE), Cloud Infrastructure, Containerized Infrastructure, Distributed Systems, Platform Engineering, Production Operations, CI/CD, Deployment Automation, Infrastructure as Code (IaC), Python, Golang, Grafana, Prometheus, Splunk, Observability, Log Analysis, Incident Response, Root Cause Analysis, Network Troubleshooting, Ray Clusters, Systems Automation, Performance Monitoring, Scalability, High Availability, Operational Excellence, Technical Documentation, Cross-Functional Collaboration, Project OwnershipBayside Solutions, Inc. is not able to sponsor any candidates at this time. Additionally, candidates for this position must qualify as a W2 candidate.Bayside Solutions, Inc. may collect your personal information during the position application process. Please reference Bayside Solutions, Inc.'s CCPA Privacy Policy at www.baysidesolutions.com.