Cloud Engineer
Cloud Engineer
Education: Bachelor's Degree
Years of Experience: 5 years
Location: Telework (within the United States)
Background Investigation: Tier 2 / Moderate Risk Background Investigation required
Description: The Cloud Engineer executes day-to-day infrastructure operations, deployment, and maintenance for a complex, mission-critical cloud platform hosted in AWS GovCloud supporting 300+ applications and services across production, staging, and sandbox environments. This role manages EKS clusters, CI/CD pipeline operations, secrets and certificate management, backup and recovery automation, S3 lifecycle management, failover automation, and configuration drift remediation, ensuring all changes are integrated into approved DevSecOps workflows and Infrastructure as Code practices.
Responsibilities:
Deploy and maintain EKS clusters, VPC infrastructure, IAM configurations, and networking across all platform environments using approved Terraform IaC templates
Manage day-to-day operations of 3,000+ CI/CD pipelines supporting 130+ monthly production releases across 288 production services
Implement and operate centralized secrets management ensuring automated rotation for 100% of production secrets where automation is supported and manual rotation at least every 90 days for non-automated secrets
Manage Public Key Infrastructure operations including TLS certificate issuance, renewal, rotation, storage, and integrity ensuring zero disruption to production services and 100% renewal no later than 30 days prior to expiration
Configure and maintain AWS Backup and VA-approved Kubernetes-native backup tooling for cross-region snapshot replication with independent per-region encryption
Manage S3 Cross-Region Replication configurations, lifecycle policies, and Intelligent-Tiering for mission-critical image data
Execute failover automation procedures: Route 53 health-check configuration, DNS failover activation, and secondary-region infrastructure management
Perform automated configuration drift detection at least weekly and remediate identified drift within three business days
Ensure all environment changes are fully validated in non-production environments prior to deployment to production
Conduct post-deployment production verification testing within two hours of each production change
Prepare and maintain documented backout and rollback plans for 100% of production changes
Support application teams with manual deployments and other work through ticketed and CCB-managed processes
Manage node lifecycle including detection, cordon, drain, and replacement of failed nodes within 30 minutes
Ensure autoscaling policies prevent resource saturation exceeding 80% for more than five minutes
Monitor EBS volume utilization ensuring volumes do not exceed 80% for more than one hour
Ensure 100% of platform components are integrated with centralized logging and monitoring solutions
Support phased software upgrades progressing through sandbox, staging, and production with documented validation procedures and rollback plans
Participate in daily Change Control Board meetings and on-call rotation
Participate in Agile and SAFe ceremonies including sprint planning, backlog refinement, demos, and retrospectives
Qualifications:
Demonstrated experience in cloud infrastructure engineering and operations for Federal or enterprise environments
Hands-on experience with AWS Elastic Kubernetes Service (EKS), Kubernetes cluster operations, and containerized application deployment
Proficiency with Infrastructure as Code tooling, specifically Terraform, for provisioning and managing cloud infrastructure
Experience with CI/CD pipeline operations and DevSecOps workflows including build, test, deploy, and promotion processes
Hands-on experience with secrets management tools (AWS Secrets Manager, HashiCorp Vault, or Kubernetes secrets)
Experience with PKI operations including TLS certificate management using AWS Certificate Manager, HashiCorp Vault, or equivalent
Experience with AWS Backup, Velero, or equivalent backup and recovery tooling for Kubernetes environments
Knowledge of S3 operations including Cross-Region Replication, lifecycle policies, versioning, and encryption
Experience with Route 53 DNS management, health checks, and failover configurations
Familiarity with configuration drift detection and remediation practices
Knowledge of AWS managed services (RDS, ElastiCache, DocumentDB, Kinesis, MSK, Athena)
Experience with Helm charts and Kubernetes deployment manifests for application deployment
Experience working in Agile or SAFe environments
AWS Solutions Architect Associate or AWS Certified DevOps Engineer
Certified Kubernetes Administrator (CKA) or equivalent certification
Bachelor's Degree from an accredited academic institution with a minimum of 5 years of experience in cloud infrastructure engineering and operations for Federal or enterprise IT environments
Preferred Qualifications:
Department of Veterans Affairs (VA) experience
Active Personal Identity Verification (PIV) card
Experience operating in AWS GovCloud or equivalent FedRAMP-authorized cloud environments
Experience with FIPS 140-3 encryption requirements and Federal security compliance standards
Experience managing platforms supporting 200+ applications across multiple environments
We are committed to fostering an inclusive and professional work environment. As an equal opportunity employer, we consider all qualified applicants for employment without regard to legally protected characteristics. We offer comprehensive healthcare benefits, which will be discussed in detail prior to your first interview.