JOBSEARCHER

Information Technology - DevOps Engineer (Cloud Engineer)

Job Title:Information Technology - DevOps Engineer (Cloud Engineer)Location: Chicago, ILType: W2 ContractRequired EducationBachelor s degree in Computer Science, Information Technology, Engineering, or equivalent work experienceRequired Experience7+ years in Backup Engineering, Infrastructure Engineering, or Site Reliability Engineering5+ years designing enterprise backup solutions3+ years supporting cyber recovery architecturesExperience implementing SRE principles within enterprise infrastructure environmentsStrong understanding of distributed systems and high availability architecturesCohesityDell PowerProtect Data ManagerDell Data DomainDell Cyber RecoveryRubrikCommvaultVeritas NetBackupVeeamAir-gapped vaultsImmutable backupsClean RoomsIsolated Recovery Environments (IRE)Recovery orchestrationCyber resilience testingRansomware recoveryRecovery validationMicrosoft AzureAWSGoogle Cloud PlatformCloud-native backupCross-region recoveryHybrid cloud resiliencyVMwareHyper-VKubernetesOpenShiftLinuxWindows ServerActive DirectoryEnterprise storage platformsAnsibleTerraformPythonPowerShellBashGitHubGitHub ActionsCI/CD pipelinesDynatraceGrafanaPrometheusSplunkELK StackServiceNowZero Trust architectureNIST Cybersecurity FrameworkCIS ControlsEncryption and key managementIdentity and Access Management (IAM)Multi-factor authentication (MFA)Secure recovery processesPreferred QualificationsExperience in financial services or another highly regulated industryExperience supporting GSIB cyber resiliency programsKnowledge of regulatory expectations from agencies such as the Federal Reserve, OCC, or FFIECExperience with chaos engineering and resilience testingFamiliarity with SRE tooling and reliability metricsExperience implementing AI-assisted operations (AIOps) and predictive analyticsStrong systems thinking and engineering mindsetExcellent troubleshooting and root cause analysis skillsAbility to lead cross-functional technical recovery effortsStrong communication and executive presentation skillsProven ability to influence engineering standards and drive operational excellenceCommitment to continuous improvement through automation and reliability engineeringDescriptionJob Description Project OverviewSeeking a highly technical Senior Site Reliability Engineer (SRE) with deep expertise in enterprise backup engineering, cyber recovery, and platform resiliencyResponsible for engineering highly available, secure, and automated recovery capabilities that protect against operational failures, ransomware, and other cyber threatsCombines traditional SRE principles (automation, observability, reliability engineering, and resilience) with experience designing and operating enterprise backup platforms, immutable storage, air-gapped cyber vaults, isolated recovery environments (IREs), and recovery orchestrationPartners closely with Infrastructure, Cyber Security, Cloud Engineering, Application Development, and Disaster Recovery teams to ensure critical services remain recoverable, resilient, and continuously validatedEngineer and maintain highly available, resilient enterprise platforms using SRE principlesDefine and measure Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for backup and recovery servicesDevelop automation to reduce operational toil and improve reliabilityPerform root cause analysis (RCA) and implement permanent corrective actionsContinuously improve platform reliability, scalability, performance, and recoverabilityEstablish proactive monitoring, alerting, and observability for backup and cyber recovery platformsParticipate in incident response and major incident recovery activitiesDesign, implement, and administer enterprise backup and recovery solutions across on-premises, cloud, and SaaS platformsEngineer immutable backup architectures that support ransomware resilienceDesign backup strategies for virtual environments, physical servers, databases, Kubernetes/OpenShift, cloud-native workloads, NAS/Object Storage, and enterprise applicationsOptimize backup performance, retention, replication, encryption, and recovery objectivesImplement policy-based backup automation and lifecycle managementEnsure compliance with enterprise RPO and RTO requirementsDesign and implement enterprise cyber recovery solutions including air-gapped recovery vaults, clean rooms, Isolated Recovery Environments (IRE), and immutable storage architecturesDevelop secure recovery workflows following cyberattack scenariosEngineer automated malware scanning and recovery validation processesDesign and test recovery orchestration for severe-but-plausible cyber eventsSupport recovery point validation and promotion into production recovery environmentsCollaborate with Cyber Security teams on ransomware resilience strategiesDevelop Infrastructure as Code (IaC) and Recovery as Code automationBuild automated recovery runbooks using Ansible, Terraform, PowerShell, Python, and GitHub ActionsAutomate recovery validation, reporting, and compliance evidence generationEliminate manual recovery processes wherever possibleImplement monitoring for backup success rates, replication health, recovery readiness, storage utilization, cyber vault health, and infrastructure dependenciesBuild dashboards for executive and operational visibilityIntegrate with enterprise observability platforms (Dynatrace, Grafana, Splunk, Prometheus)Plan and execute cyber recovery exercises, clean room validation, air-gap recovery testing, full isolated recovery environment exercises, Bare Metal Recovery (BMR) testing, and Disaster Recovery testingValidate application recoverability against defined RTO/RPO objectivesProduce executive reporting on recovery readiness and testing outcomes