Senior Site Reliability Engineer
We're working with a leading technology solutions provider specializing in delivering innovative services across various industries, from IT consulting to digital transformation.The Role Design, implement, and support Kubernetes on baremetal and hypervisor platforms in a private cloud environment. Architect, design, and standardize enterprise compute and hypervisor environments using Infrastructure-as-Code and GitOps practices. Lead the architecture and design of enterprise compute and hypervisor platform solutions across hardware, OS, virtualization, cloud orchestration, and container orchestration layers. Define standards and automation frameworks for bare metal provisioning and lifecycle management. Design and implement Bare Metal as a Service (BMaaS) capabilities for scalable infrastructure consumption. Perform deep troubleshooting across storage, Kubernetes, hypervisors, networking, and Linux systems.What You'll Need 6+ years of experience in infrastructure engineering, platform engineering, or DevOps with a strong focus on Compute system design. Proven experience designing and automating bare metal compute environments at scale, including PXE boot and automated server imaging. Practical experience using Redfish APIs for hardware provisioning and lifecycle operations. Deep expertise with Ubuntu Linux in enterprise environments and strong hands-on experience with KVM hypervisors (e.g., Suse Harvester, OpenStack). Experience designing and deploying production-grade Kubernetes clusters. Proficiency with Infrastructure as Code tools (e.g., Terraform, Ansible) and strong scripting skills (Python, Bash).What's On Offer The opportunity to work on cutting-edge compute platform technologies, including Kubernetes on bare metal and hypervisors. A role that offers significant impact on enterprise compute and hypervisor environments. Collaboration with experienced platform and SRE teams to build secure, performant, and multi-tenant services.Apply via Haystack today!