{"schemaVersion":"jobsearcher.job.v1","id":"8aed87f33a13bf44b53679bd","url":"https://jobsearcher.com/jobs/8aed87f33a13bf44b53679bd","canonicalUrl":"https://jobsearcher.com/jobs/8aed87f33a13bf44b53679bd","title":"Manager, Site Reliability Engineering","description":"Paradigm is a software company transforming the way that the residential, construction & building product industries operate across the globe. We are looking for a Manager, Site Reliability Engineering to be part of revolutionizing these industries.\nWe're looking for a hands-on SRE leader to build and develop a high-performing team that oversees reliability across our Azure-based platform. You'll promote modern SRE practices, drive down incident response times, and shape a culture where automation replaces toil and every incident becomes a learning opportunity.\nThis role combines technical depth with people leadership. You'll design reliability frameworks, lead incident response, coach engineers, and partner with product teams to embed reliability into everything we build. Working closely with the Senior Director of SRE & Cloud Operations, you'll transform reactive operations into proactive, data-driven service management with increasing use of AI and automation to get there faster.\nWhat You Will Do:\nLead and grow a team of site reliability engineers. Provide guidance, mentorship, and career development.\nContribute to and mature SRE practices across production services: SLOs, SLIs, error budgets, toil reduction, and blameless post-mortems that turn incidents into lasting improvements.\nOversee the incident management lifecycle end-to-end including detection, response, resolution, post-incident review, and systemic improvement.\nDesign on-call rotations, runbooks, and escalation procedures that balance service reliability with engineer well-being and sustainable work practices.\nDrive measurable reductions in MTTR and MTTD through improved observability, intelligent automation, and predictive monitoring.\nBuild automation to eliminate manual operational work including provisioning, deployment, scaling, self-healing, and reporting.\nImplement chaos engineering practices to validate system resilience and surface weaknesses before they cause outages.\nPartner with engineering and product teams to embed reliability requirements into the development lifecycle, from design through deployment.\nCollaborate with the observability team to ensure comprehensive instrumentation, smart alerting, and actionable dashboards across all critical services.\nMeasure, report, and advocate for reliability improvements with both technical and executive stakeholders using data to drive investment decisions.\nWhat You Need to Succeed:\nBachelor’s degree in Engineering, or a related field or equivalent experience.\n7+ years in site reliability engineering, DevOps, or infrastructure engineering, with at least 1 year in people management (or demonstrated tech lead experience with direct influence over team processes and career growth).\nHands-on experience running production systems on Azure (including proficiency with key services such as AKS, App Services, Service Bus, Event Grid, and Azure Monitor) or comparable cloud platforms.\nProven track record implementing SRE practices with measurable reliability improvements and familiarity with modern observability platforms (Datadog, Prometheus/Grafana, or equivalent). AI-enhanced observability experience is preferred.\nExperience leading incident response for high-severity production issues and running effective post-mortems.\nStrong background in automation, infrastructure as code (Terraform, Bicep, or similar), and systematically eliminating manual operational work.\nExperience with Kubernetes container orchestration with production-grade operational experience.\nAbility to automate workflows and build scripts using Python, Bash, PowerShell, or Go.\nExperience with AI coding assistants and CI/CD systems (GitHub Actions, Azure DevOps, ArgoCD) with automation capabilities is preferred.\nKnowledge of distributed systems patterns is preferred.\nExposure to AIOps platforms or using LLMs for operational automation is preferred.\nStrong communication with the ability to make complex technical issues clear for both engineers and executives.\nData-driven approach. You use metrics and telemetry to guide decisions, not gut feel.\nYou are collaborative cross-functionally and build trust and alignment naturally.\nReady to Join? Apply now at myparadigm.com/careers/\n#Paradigm","company":"Paradigm","rawCompany":"paradigm","city":"Middleton","state":"WI","isRemote":false,"isActive":false,"createdAt":"2026-04-12T18:07:38.025Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"11-3021.00","title":"Computer and Information Systems Managers","slug":"computer-and-information-systems-managers"},{"code":"11-9041.00","title":"Architectural and Engineering Managers","slug":"architectural-and-engineering-managers"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Manager, Site Reliability Engineering","description":"Paradigm is a software company transforming the way that the residential, construction & building product industries operate across the globe. We are looking for a Manager, Site Reliability Engineering to be part of revolutionizing these industries.\nWe're looking for a hands-on SRE leader to build and develop a high-performing team that oversees reliability across our Azure-based platform. You'll promote modern SRE practices, drive down incident response times, and shape a culture where automation replaces toil and every incident becomes a learning opportunity.\nThis role combines technical depth with people leadership. You'll design reliability frameworks, lead incident response, coach engineers, and partner with product teams to embed reliability into everything we build. Working closely with the Senior Director of SRE & Cloud Operations, you'll transform reactive operations into proactive, data-driven service management with increasing use of AI and automation to get there faster.\nWhat You Will Do:\nLead and grow a team of site reliability engineers. Provide guidance, mentorship, and career development.\nContribute to and mature SRE practices across production services: SLOs, SLIs, error budgets, toil reduction, and blameless post-mortems that turn incidents into lasting improvements.\nOversee the incident management lifecycle end-to-end including detection, response, resolution, post-incident review, and systemic improvement.\nDesign on-call rotations, runbooks, and escalation procedures that balance service reliability with engineer well-being and sustainable work practices.\nDrive measurable reductions in MTTR and MTTD through improved observability, intelligent automation, and predictive monitoring.\nBuild automation to eliminate manual operational work including provisioning, deployment, scaling, self-healing, and reporting.\nImplement chaos engineering practices to validate system resilience and surface weaknesses before they cause outages.\nPartner with engineering and product teams to embed reliability requirements into the development lifecycle, from design through deployment.\nCollaborate with the observability team to ensure comprehensive instrumentation, smart alerting, and actionable dashboards across all critical services.\nMeasure, report, and advocate for reliability improvements with both technical and executive stakeholders using data to drive investment decisions.\nWhat You Need to Succeed:\nBachelor’s degree in Engineering, or a related field or equivalent experience.\n7+ years in site reliability engineering, DevOps, or infrastructure engineering, with at least 1 year in people management (or demonstrated tech lead experience with direct influence over team processes and career growth).\nHands-on experience running production systems on Azure (including proficiency with key services such as AKS, App Services, Service Bus, Event Grid, and Azure Monitor) or comparable cloud platforms.\nProven track record implementing SRE practices with measurable reliability improvements and familiarity with modern observability platforms (Datadog, Prometheus/Grafana, or equivalent). AI-enhanced observability experience is preferred.\nExperience leading incident response for high-severity production issues and running effective post-mortems.\nStrong background in automation, infrastructure as code (Terraform, Bicep, or similar), and systematically eliminating manual operational work.\nExperience with Kubernetes container orchestration with production-grade operational experience.\nAbility to automate workflows and build scripts using Python, Bash, PowerShell, or Go.\nExperience with AI coding assistants and CI/CD systems (GitHub Actions, Azure DevOps, ArgoCD) with automation capabilities is preferred.\nKnowledge of distributed systems patterns is preferred.\nExposure to AIOps platforms or using LLMs for operational automation is preferred.\nStrong communication with the ability to make complex technical issues clear for both engineers and executives.\nData-driven approach. You use metrics and telemetry to guide decisions, not gut feel.\nYou are collaborative cross-functionally and build trust and alignment naturally.\nReady to Join? Apply now at myparadigm.com/careers/\n#Paradigm","datePosted":"2026-04-12T18:07:38.025Z","dateModified":"2026-04-12T18:07:38.025Z","hiringOrganization":{"@type":"Organization","name":"Paradigm","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Middleton","addressRegion":"WI","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"8aed87f33a13bf44b53679bd"},"url":"https://jobsearcher.com/jobs/8aed87f33a13bf44b53679bd"}}