AI / Cloud Infrastructure Engineer
Job SummaryExperience: 5+ YearsWe are seeking an experienced AI / Cloud Infrastructure Engineer to design, deploy, automate, and maintain scalable cloud infrastructure supporting artificial intelligence, machine learning, and generative AI workloads. This role combines cloud engineering, DevOps, Kubernetes, infrastructure automation, GPU computing, and MLOps.The ideal candidate will have hands-on experience with AWS, Azure, or Google Cloud and a strong understanding of infrastructure requirements for model training, inference, data processing, and production AI applications.Key ResponsibilitiesDesign and maintain secure, scalable, highly available cloud infrastructure for AI and machine learning workloads.Deploy and manage compute, storage, networking, GPU, and container infrastructure across cloud environments.Build and administer Kubernetes clusters for model training, inference, and AI application deployment.Automate infrastructure provisioning using Terraform, CloudFormation, Bicep, or Pulumi.Develop CI/CD pipelines for AI applications, machine learning models, containers, and infrastructure changes.Configure and optimize GPU-enabled environments for model training and large-scale inference.Build reliable infrastructure for large language models, generative AI applications, RAG pipelines, and vector databases.Deploy and support model-serving solutions for real-time and batch inference.Implement autoscaling, workload scheduling, resource quotas, load balancing, and disaster-recovery capabilities.Monitor cloud infrastructure, model endpoints, latency, resource utilization, and service availability.Optimize cloud and GPU costs through rightsizing, autoscaling, reserved capacity, and workload scheduling.Implement IAM, network security, encryption, secrets management, vulnerability scanning, and audit logging.Support ML platforms such as MLflow, Kubeflow, Amazon SageMaker, Azure Machine Learning, or Vertex AI.Develop operational dashboards, alerts, runbooks, and incident-response procedures.Troubleshoot performance, networking, container, deployment, and infrastructure-related production issues.Collaborate with AI engineers, data scientists, software engineers, DevOps teams, and security teams.Required QualificationsBachelor’s or Master’s degree in Computer Science, Information Technology, Engineering, or a related discipline.5+ years of experience in cloud infrastructure, DevOps, Site Reliability Engineering, platform engineering, or MLOps.Hands-on experience with AWS, Microsoft Azure, or Google Cloud Platform.Strong experience with Kubernetes, Docker, and Helm.Experience provisioning cloud infrastructure through Infrastructure as Code.Proficiency in Python, Bash, PowerShell, or a comparable scripting language.Experience with CI/CD tools such as GitHub Actions, GitLab CI, Jenkins, or Azure DevOps.Strong knowledge of cloud networking, including VPCs/VNets, subnets, DNS, routing, firewalls, load balancers, and private endpoints.Experience managing Linux-based production environments.Knowledge of observability and monitoring tools such as Prometheus, Grafana, OpenTelemetry, Datadog, CloudWatch, or Azure Monitor.Understanding of machine learning lifecycle, model deployment, inference, and data pipelines.Strong knowledge of IAM, secrets management, encryption, security policies, and cloud governance.Experience supporting highly available and business-critical production systems.Preferred QualificationsExperience managing NVIDIA GPU infrastructure and CUDA-based workloads.Experience with Amazon EKS, Azure Kubernetes Service, or Google Kubernetes Engine.Knowledge of distributed training and inference environments.Experience with NVIDIA Triton, KServe, Ray Serve, vLLM, TorchServe, or TensorFlow Serving.Familiarity with generative AI, large language models, embeddings, RAG, and vector search.Experience with vector databases such as Pinecone, Weaviate, Milvus, Qdrant, OpenSearch, or pgvector.Knowledge of MLflow, Kubeflow, SageMaker, Azure ML, Vertex AI, or Databricks.Experience with service mesh technologies such as Istio or Linkerd.Familiarity with Kafka, Spark, Airflow, Snowflake, and enterprise data platforms.Experience implementing FinOps practices for cloud and AI infrastructure.Cloud, Kubernetes, DevOps, or security certifications are preferred.Core Technical SkillsCloud: AWS, Azure, or Google CloudContainers: Docker, Kubernetes, HelmInfrastructure as Code: Terraform, CloudFormation, Bicep, or PulumiCI/CD: GitHub Actions, GitLab CI, Jenkins, or Azure DevOpsProgramming: Python, Bash, PowerShell, SQLAI Infrastructure: GPU computing, model serving, distributed workloads, LLM infrastructureMLOps: MLflow, Kubeflow, SageMaker, Azure ML, Vertex AIMonitoring: Prometheus, Grafana, OpenTelemetry, Datadog, CloudWatchSecurity: IAM, encryption, secrets management, network security, vulnerability managementSuccess in This RoleThe successful candidate will deliver a stable, secure, and cost-efficient cloud foundation that enables AI teams to train, deploy, scale, and monitor production AI applications reliably.