JOBSEARCHER

Infrastructure Engineer

Site Reliability Engineer (SRE) – Microsoft Hyper-V & Private CloudHyper- V SRE ( Senior Infrastructure Engineer — Hyper-V)Mandatory Skills for Hyper-V SRE Production SupportMicrosoft Hyper-V Administration (Deployment, Troubleshooting, Optimization)Hyper-V Failover Clustering & High AvailabilityWindows Server 2016/2019/2022 Administration & OS patchingStorage Spaces Direct (S2D), CSV, SAN/NAS & StorageSite Reliability Engineering (SRE) Principles, SLI/SLO, ReliabilityPowerShell Scripting & AutomationDisaster Recovery, Backup, Hyper-V Replica & Business ContinuitySCVMM (System Center Virtual Machine Manager) Desired Skills for Hyper-V SRE Production SupportAzure Monitor / SCOM / Prometheus / Grafana / SplunkVeeam / Altaro Backup SolutionsVDI InfrastructureInfrastructure as Code (IaC)Hybrid Cloud / Private CloudITIL (Incident, Change, Problem Management)OverviewWe are seeking an experienced Site Reliability Engineer (SRE) – Microsoft Hyper-V & Private Cloud to operate highly available private cloud and Virtual Desktop Infrastructure (VDI) platforms based on Microsoft Hyper-V. This role combines deep Hyper-V expertise with modern Site Reliability Engineering practices to improve platform reliability, scalability, performance, automation, and operational excellence.The ideal candidate will be responsible for ensuring platform reliability through infrastructure automation, OS upgrades, proactive monitoring, incident response, capacity planning, and continuous service improvement while collaborating with infrastructure, security, and application teams.Key ResponsibilitiesOperate enterprise-scale private cloud infrastructure built on Microsoft Hyper-V.Optimize, and support highly available VDI environments on Hyper-V.Improve platform reliability, availability, scalability, and resiliency by applying SRE principles and engineering best practices.Disaster recovery, backup, patch management, and business continuity strategies.Define and maintain Service Level Indicators (SLIs), Service Level Objectives (SLOs), and operational metrics for critical infrastructure services.Automate infrastructure provisioning, configuration management, and operational workflows using PowerShell and Infrastructure as Code (IaC) principles wherever applicable.Manage Hyper-V Failover Clusters, host lifecycle, storage, networking, and capacity to ensure high availability and business continuity.Develop proactive monitoring, alerting, logging, and observability capabilities to detect and prevent service degradation.Lead incident response for infrastructure-related outages, perform root cause analysis (RCA), and implement preventive actions through post-incident reviews.Perform capacity planning, performance tuning, and resource optimization across Hyper-V clusters and VDI platforms.Support infrastructure migration initiatives including P2V, V2V, workload modernization, and private cloud transformations.Collaborate closely with Security, Networking, Platform Engineering, and Application teams to improve platform reliability and operational efficiency.Develop and maintain technical documentation, architecture diagrams, operational runbooks, automation scripts, and standard operating procedures.Mentor junior engineers and promote SRE culture, automation, and operational best practices across the team.Required Technical SkillsSite Reliability EngineeringStrong understanding of Site Reliability Engineering principles and operational excellence.Experience with infrastructure reliability, service availability, resiliency, and performance optimization.