JOBSEARCHER

Cloud Engineer

Cloud EngineerEducation: Bachelor's DegreeYears of Experience: 5 yearsLocation: Telework (within the United States)Background Investigation: Tier 2 / Moderate Risk Background Investigation required Description: The Cloud Engineer executes day-to-day infrastructure operations, deployment, and maintenance for a complex, mission-critical cloud platform hosted in AWS GovCloud supporting 300+ applications and services across production, staging, and sandbox environments. This role manages EKS clusters, CI/CD pipeline operations, secrets and certificate management, backup and recovery automation, S3 lifecycle management, failover automation, and configuration drift remediation, ensuring all changes are integrated into approved DevSecOps workflows and Infrastructure as Code practices. Responsibilities:Deploy and maintain EKS clusters, VPC infrastructure, IAM configurations, and networking across all platform environments using approved Terraform IaC templatesManage day-to-day operations of 3,000+ CI/CD pipelines supporting 130+ monthly production releases across 288 production servicesImplement and operate centralized secrets management ensuring automated rotation for 100% of production secrets where automation is supported and manual rotation at least every 90 days for non-automated secretsManage Public Key Infrastructure operations including TLS certificate issuance, renewal, rotation, storage, and integrity ensuring zero disruption to production services and 100% renewal no later than 30 days prior to expirationConfigure and maintain AWS Backup and VA-approved Kubernetes-native backup tooling for cross-region snapshot replication with independent per-region encryptionManage S3 Cross-Region Replication configurations, lifecycle policies, and Intelligent-Tiering for mission-critical image dataExecute failover automation procedures: Route 53 health-check configuration, DNS failover activation, and secondary-region infrastructure managementPerform automated configuration drift detection at least weekly and remediate identified drift within three business daysEnsure all environment changes are fully validated in non-production environments prior to deployment to productionConduct post-deployment production verification testing within two hours of each production changePrepare and maintain documented backout and rollback plans for 100% of production changesSupport application teams with manual deployments and other work through ticketed and CCB-managed processesManage node lifecycle including detection, cordon, drain, and replacement of failed nodes within 30 minutesEnsure autoscaling policies prevent resource saturation exceeding 80% for more than five minutesMonitor EBS volume utilization ensuring volumes do not exceed 80% for more than one hourEnsure 100% of platform components are integrated with centralized logging and monitoring solutionsSupport phased software upgrades progressing through sandbox, staging, and production with documented validation procedures and rollback plansParticipate in daily Change Control Board meetings and on-call rotationParticipate in Agile and SAFe ceremonies including sprint planning, backlog refinement, demos, and retrospectives Qualifications:Demonstrated experience in cloud infrastructure engineering and operations for Federal or enterprise environmentsHands-on experience with AWS Elastic Kubernetes Service (EKS), Kubernetes cluster operations, and containerized application deploymentProficiency with Infrastructure as Code tooling, specifically Terraform, for provisioning and managing cloud infrastructureExperience with CI/CD pipeline operations and DevSecOps workflows including build, test, deploy, and promotion processesHands-on experience with secrets management tools (AWS Secrets Manager, HashiCorp Vault, or Kubernetes secrets)Experience with PKI operations including TLS certificate management using AWS Certificate Manager, HashiCorp Vault, or equivalentExperience with AWS Backup, Velero, or equivalent backup and recovery tooling for Kubernetes environmentsKnowledge of S3 operations including Cross-Region Replication, lifecycle policies, versioning, and encryptionExperience with Route 53 DNS management, health checks, and failover configurationsFamiliarity with configuration drift detection and remediation practicesKnowledge of AWS managed services (RDS, ElastiCache, DocumentDB, Kinesis, MSK, Athena)Experience with Helm charts and Kubernetes deployment manifests for application deploymentExperience working in Agile or SAFe environmentsAWS Solutions Architect Associate or AWS Certified DevOps EngineerCertified Kubernetes Administrator (CKA) or equivalent certificationBachelor's Degree from an accredited academic institution with a minimum of 5 years of experience in cloud infrastructure engineering and operations for Federal or enterprise IT environments Preferred Qualifications:Department of Veterans Affairs (VA) experienceActive Personal Identity Verification (PIV) cardExperience operating in AWS GovCloud or equivalent FedRAMP-authorized cloud environmentsExperience with FIPS 140-3 encryption requirements and Federal security compliance standardsExperience managing platforms supporting 200+ applications across multiple environments