AWS DevOps Engineer
Remote RolePosition Title: Senior AWS DevOps/Platform EngineerLocation: Denver CO/ RemoteDuration: 12+ Months ContractPosition Type- W2 OnlyExp Level- 10+YearsReq Skills- Platform Engineer, AWS, EC2, S3, IAM, VPC, Glue, Athena, EMR, Secrets Manager, CloudWatch, (infrastructure-as-code, Terraform preferred, CloudFormation), (AI agent deployments using GitLab CI/CD, Docker, Artifactory), AWS architectures, VPC peering, PrivateLink, transit gateway configurations, Docker, CI/CD pipelines (GitLab CI preferred), (Kafka, SNS/SQS, EventBridge), (CloudWatch, Prometheus, Grafana, or similar), (DNS, CIDR, NAT, VPC endpoints, firewalls), (REST APIs, Cloud SDKs, AWS SDK / boto3), (IIA Data Lake, S3 storage, Glue data catalog, Athena query engine, EMR compute clusters) JOB SUMMARYPlatform Engineer designs, builds, and maintains the AWS infrastructure that underpins the IIA Data Lake, agent runtime environments, CI/CD pipelines, and graph database systems, and develops the application code, automation, and internal tooling for those systems.This role ensures production environments are stable, scalable, and secure while enabling data science and agentic AI workloads to operate reliably at scale.MAJOR DUTIES AND RESPONSIBILITIESResponsibilities span the following areas. Individual focus areas will be determined based on team needs and candidate strengths:Infrastructure and Data LakeDesign and manage AWS infrastructure for the IIA Data Lake including S3 storage, Glue data catalog, Athena query engine, and EMR compute clusters.Manage cross-account connectivity, VPC networking, security groups, and IAM roles/policies to enable secure data flow between IIA, upstream data providers, and downstream consumers.Build and maintain infrastructure for AI agent runtime environments, including compute resources for LangGraph agents deployed via LangSmith Deployments.Design and implement event-driven infrastructure including triggers, message queues, and pub/sub frameworks that orchestrate agent execution and data pipeline workflowsSupport deployment and operation of AWS Neptune for the network topology graph (digital twin), including capacity planning, schema design support, and performance tuning.Implement and manage infrastructure-as-code (Terraform, CloudFormation) for repeatable, auditable environment provisioning.Manage IAM access key rotations, secrets management (AWS Secrets Manager, Delinea), and security compliance for on-premises and cloud integrations (e.g., Splunk Edge Processor).CI/CD and Agent DeploymentsBuild and maintain CI/CD pipelines for AI agent deployments using GitLab CI/CD, Docker, and Artifactory.Manage container lifecycle for agents deployed via LangSmith Deployments, including image builds, versioning, and rollback procedures.Automate deployment workflows to enable rapid, reliable promotion of agents from development through production.Coordinate with SpecGPT platform team on AI Gateway integration, cross-account deployment, and connectivity requirements.Application Development and ToolingSupport and maintain existing ETL pipelines written in Scala/Spark, including troubleshooting, enhancements, and onboarding new data sources.Develop and maintain application code: scripts, CLIs, small services, and automation utilities (primarily Python) that support data ingestion, deployment, environment provisioning, and operational workflows.Write integration code and glue services that connect IIA systems with upstream data providers, downstream consumers, and external platforms.Production OperationsEnsure production environment stability through monitoring, alerting, and incident response. Maintain SLAs for data pipeline availability and agent uptime.Implement production monitoring and alerting for deployed agents (health checks, error rates, latency, resource utilization).Coordinate with upstream data teams and platform teams (SpecGPT, Splunk, Public Cloud) on connectivity, firewall requests, and integration requirements.Support data engineering team with infrastructure needs for new data source onboarding (storage provisioning, access controls, pipeline compute) and coordinate on scheduler maintenance and development (Airflow, event-driven triggers).Perform other duties as required.REQUIRED QUALIFICATIONSSkills/Abilities and KnowledgeStrong communication skills with ability to explain infrastructure decisions to non-infrastructure stakeholdersExpert-level experience with AWS services: EC2, S3, IAM, VPC, Glue, Athena, EMR, Secrets Manager, CloudWatchStrong experience with infrastructure-as-code (Terraform preferred, CloudFormation acceptable)Experience managing cross-account AWS architectures, VPC peering, PrivateLink, and transit gateway configurationsExperience with IAM policy design, least-privilege access patterns, and service account managementExperience with containerization (Docker) and container orchestrationExperience with CI/CD pipelines (GitLab CI preferred)Experience with event-driven architectures and messaging systems (e.g., Kafka, SNS/SQS, EventBridge)Proficiency with Linux-based operating systems and shell scriptingExperience with monitoring and alerting tools (CloudWatch, Prometheus, Grafana, or similar)Understanding of networking fundamentals: DNS, CIDR, NAT, VPC endpoints, firewalls, security groupsDemonstrated ability to work across teams and coordinate with external platform owners on connectivity and access requirementsProficiency in Python (or a comparable general-purpose language) for building automation, tooling, and applicationsSolid software engineering fundamentals: Git-based workflows, code review, modular and reusable design, dependency management, and writing maintainable, documented codeExperience writing automated tests (unit/integration) for application and infrastructure code, and integrating those tests into CI/CDAbility to write integration code against REST APIs and cloud SDKs (e.g., AWS SDK / boto3)PREFERRED QUALIFICATIONSSkills/Abilities and KnowledgeExperience with graph databases (AWS Neptune, Neo4j) including deployment, scaling, and operational managementExperience with Apache Kafka or similar streaming platformsExperience with Apache Spark (Scala preferred) for distributed and streaming data processing such as Spark Streaming or structured streaming.Experience with Airflow or similar workflow orchestration platformsExperience in the telecommunications industry or other large-scale network operations environmentsFamiliarity with AI/ML infrastructure requirements (model serving, GPU/CPU compute, artifact management via MLflow or similar)Experience with Splunk integration, particularly Edge Processor and MCP connectivityAWS certifications (Solutions Architect, DevOps Engineer, or similar)Experience developing and operating small services or APIs (e.g., FastAPI/Flask) in a production environment