{"schemaVersion":"jobsearcher.job.v1","id":"e52c56bb7fe24b50cf2fa52b","url":"https://jobsearcher.com/jobs/e52c56bb7fe24b50cf2fa52b","canonicalUrl":"https://jobsearcher.com/jobs/e52c56bb7fe24b50cf2fa52b","title":"Staff/Principal DevOps Engineer, AI Inference","description":"Your Impact at LILA\nThe Staff/Principal DevOps Engineer - AI Inference will drive the design, implementation, and optimization of infrastructure purpose-built for serving machine learning models at scale. This role bridges platform engineering, site reliability, and ML infrastructure, building the systems that power low-latency, high-throughput inference across GPU clusters and cloud accelerators. You will collaborate with ML engineers, research scientists, and software engineers to build inference platforms that serve models reliably to production users while maximizing compute efficiency.\nWhat You'll Be Building\nGPU/accelerator infrastructure on Kubernetes: scheduling, resource isolation, multi-tenant GPU sharing, device plugins, and topology-aware placement for inference workloads\nModel serving platforms using frameworks such as vLLM, Triton Inference Server, TGI, or custom serving stacks with optimized batching, caching, and request routing\nIntelligent request routing and load balancing across heterogeneous accelerator fleets (NVIDIA GPUs, AWS Inferentia/Trainium) to maximize utilization and minimize latency\nAutoscaling systems that dynamically match inference compute supply with demand across production, research, and experimental workloads\nProduction-grade deployment pipelines for ML models: canary rollouts, A/B testing, model versioning, and safe rollback across multi-region deployments\nInfrastructure-as-code with Terraform and Helm for GPU-accelerated EKS clusters, including node pools, spot/on-demand strategies, and accelerator-specific networking\nObservability and performance optimization: GPU utilization monitoring, inference latency profiling, token throughput dashboards, and SLO/SLI tracking for model endpoints\nCI/CD pipelines for model artifacts: container image builds with CUDA/driver dependencies, model registry integration, and automated inference benchmarking in CI\nAWS cloud infrastructure for ML: EKS with GPU node groups, EC2 accelerated instances (P4/P5, Inf2, Trn1), S3 model storage, EFA/high-bandwidth networking, and IAM least privilege\nCost optimization and capacity planning: right-sizing accelerator instances, spot instance strategies for inference, and fleet-wide efficiency reporting\nWhat You'll Need to Succeed\nExpertise in DevOps, SRE, or Platform Engineering with significant experience operating GPU/accelerator infrastructure at scale\nDeep experience with Kubernetes for ML workloads: GPU scheduling, resource quotas, node affinity, and accelerator device management\nStrong proficiency deploying to AWS using infrastructure-as-code (Terraform, Helm) with hands-on experience managing GPU-based compute (EKS, EC2 P-series/Inf/Trn instances)\nExperience with model serving infrastructure: inference servers, request batching, KV-cache optimization, or LLM serving frameworks\nStrong understanding of networking for distributed inference: high-bandwidth interconnects, NCCL, VPC/PrivateLink, and load balancing at L4/L7\nStrong proficiency in Python for automation, tooling, and integration with ML frameworks\nBonus Points For\nExperience with LLM inference optimization: continuous batching, speculative decoding, quantization (GPTQ, AWQ, FP8), tensor parallelism, and pipeline parallelism\nHands-on experience with multiple accelerator families (NVIDIA A100/H100, AWS Inferentia2, Trainium, AMD MI300X) and maintaining hardware-agnostic serving infrastructure\nMulti-region deployment experience with geographic routing and failover for latency-sensitive inference endpoints\nProficiency in Rust or Go for performance-critical infrastructure components\nSRE practices for ML systems: chaos engineering on GPU workloads, incident management, capacity modeling for bursty inference traffic\nExperience with model registries, artifact versioning, and ML supply chain security\nObservability platform expertise: building custom metrics for token-level throughput, time-to-first-token, and per-request GPU memory profiling\nPrior startup/high-growth experience balancing velocity with reliability in rapidly scaling AI systems\nAbout LILA\nLila Sciences is building Scientific Superintelligence™ to solve humankind's greatest challenges. We believe science is the most inspiring frontier for AI. Rather than hard-coding expert knowledge into tools, LILA builds systems that can learn for themselves.\nLILA combines advanced AI models with proprietary AI Science Factory™ instruments into an operating system for science that executes the entire scientific method autonomously, accelerating discovery at unprecedented speed, scale, and impact across medicine, materials, and energy. Learn more at www.lila.ai.\nGuided by our core values of truth, trust, curiosity, grit, and velocity, we move with startup speed while tackling problems of historic importance. If this sounds like an environment you'd love to work in, even if you don't meet every qualification listed above, we encourage you to apply.\nWe're All In\nLila Sciences is committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity or Veteran status.\nInformation you provide during your application process will be handled in accordance with our Candidate Privacy Policy.\nA Note to Agencies\nLila Sciences does not accept unsolicited resumes from any source other than candidates. The submission of unsolicited resumes by recruitment or staffing agencies to Lila Sciences or its employees is strictly prohibited unless contacted directly by Lila Science's internal Talent Acquisition team. Any resume submitted by an agency in the absence of a signed agreement will automatically become the property of Lila Sciences, and Lila Sciences will not owe any referral or other fees with respect thereto.","company":"Lila Sciences","rawCompany":"lila sciences","city":"Somerville","state":"MA","isRemote":false,"isActive":false,"createdAt":"2026-07-30T10:54:08.408Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1221.00","title":"Computer and Information Research Scientists","slug":"computer-and-information-research-scientists"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Staff/Principal DevOps Engineer, AI Inference","description":"Your Impact at LILA\nThe Staff/Principal DevOps Engineer - AI Inference will drive the design, implementation, and optimization of infrastructure purpose-built for serving machine learning models at scale. This role bridges platform engineering, site reliability, and ML infrastructure, building the systems that power low-latency, high-throughput inference across GPU clusters and cloud accelerators. You will collaborate with ML engineers, research scientists, and software engineers to build inference platforms that serve models reliably to production users while maximizing compute efficiency.\nWhat You'll Be Building\nGPU/accelerator infrastructure on Kubernetes: scheduling, resource isolation, multi-tenant GPU sharing, device plugins, and topology-aware placement for inference workloads\nModel serving platforms using frameworks such as vLLM, Triton Inference Server, TGI, or custom serving stacks with optimized batching, caching, and request routing\nIntelligent request routing and load balancing across heterogeneous accelerator fleets (NVIDIA GPUs, AWS Inferentia/Trainium) to maximize utilization and minimize latency\nAutoscaling systems that dynamically match inference compute supply with demand across production, research, and experimental workloads\nProduction-grade deployment pipelines for ML models: canary rollouts, A/B testing, model versioning, and safe rollback across multi-region deployments\nInfrastructure-as-code with Terraform and Helm for GPU-accelerated EKS clusters, including node pools, spot/on-demand strategies, and accelerator-specific networking\nObservability and performance optimization: GPU utilization monitoring, inference latency profiling, token throughput dashboards, and SLO/SLI tracking for model endpoints\nCI/CD pipelines for model artifacts: container image builds with CUDA/driver dependencies, model registry integration, and automated inference benchmarking in CI\nAWS cloud infrastructure for ML: EKS with GPU node groups, EC2 accelerated instances (P4/P5, Inf2, Trn1), S3 model storage, EFA/high-bandwidth networking, and IAM least privilege\nCost optimization and capacity planning: right-sizing accelerator instances, spot instance strategies for inference, and fleet-wide efficiency reporting\nWhat You'll Need to Succeed\nExpertise in DevOps, SRE, or Platform Engineering with significant experience operating GPU/accelerator infrastructure at scale\nDeep experience with Kubernetes for ML workloads: GPU scheduling, resource quotas, node affinity, and accelerator device management\nStrong proficiency deploying to AWS using infrastructure-as-code (Terraform, Helm) with hands-on experience managing GPU-based compute (EKS, EC2 P-series/Inf/Trn instances)\nExperience with model serving infrastructure: inference servers, request batching, KV-cache optimization, or LLM serving frameworks\nStrong understanding of networking for distributed inference: high-bandwidth interconnects, NCCL, VPC/PrivateLink, and load balancing at L4/L7\nStrong proficiency in Python for automation, tooling, and integration with ML frameworks\nBonus Points For\nExperience with LLM inference optimization: continuous batching, speculative decoding, quantization (GPTQ, AWQ, FP8), tensor parallelism, and pipeline parallelism\nHands-on experience with multiple accelerator families (NVIDIA A100/H100, AWS Inferentia2, Trainium, AMD MI300X) and maintaining hardware-agnostic serving infrastructure\nMulti-region deployment experience with geographic routing and failover for latency-sensitive inference endpoints\nProficiency in Rust or Go for performance-critical infrastructure components\nSRE practices for ML systems: chaos engineering on GPU workloads, incident management, capacity modeling for bursty inference traffic\nExperience with model registries, artifact versioning, and ML supply chain security\nObservability platform expertise: building custom metrics for token-level throughput, time-to-first-token, and per-request GPU memory profiling\nPrior startup/high-growth experience balancing velocity with reliability in rapidly scaling AI systems\nAbout LILA\nLila Sciences is building Scientific Superintelligence™ to solve humankind's greatest challenges. We believe science is the most inspiring frontier for AI. Rather than hard-coding expert knowledge into tools, LILA builds systems that can learn for themselves.\nLILA combines advanced AI models with proprietary AI Science Factory™ instruments into an operating system for science that executes the entire scientific method autonomously, accelerating discovery at unprecedented speed, scale, and impact across medicine, materials, and energy. Learn more at www.lila.ai.\nGuided by our core values of truth, trust, curiosity, grit, and velocity, we move with startup speed while tackling problems of historic importance. If this sounds like an environment you'd love to work in, even if you don't meet every qualification listed above, we encourage you to apply.\nWe're All In\nLila Sciences is committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity or Veteran status.\nInformation you provide during your application process will be handled in accordance with our Candidate Privacy Policy.\nA Note to Agencies\nLila Sciences does not accept unsolicited resumes from any source other than candidates. The submission of unsolicited resumes by recruitment or staffing agencies to Lila Sciences or its employees is strictly prohibited unless contacted directly by Lila Science's internal Talent Acquisition team. Any resume submitted by an agency in the absence of a signed agreement will automatically become the property of Lila Sciences, and Lila Sciences will not owe any referral or other fees with respect thereto.","datePosted":"2026-07-30T10:54:08.408Z","dateModified":"2026-07-30T10:54:08.408Z","hiringOrganization":{"@type":"Organization","name":"Lila Sciences","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Somerville","addressRegion":"MA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"e52c56bb7fe24b50cf2fa52b"},"url":"https://jobsearcher.com/jobs/e52c56bb7fe24b50cf2fa52b"}}