{"schemaVersion":"jobsearcher.job.v1","id":"54407547803e2069725ac7e2","url":"https://jobsearcher.com/jobs/54407547803e2069725ac7e2","canonicalUrl":"https://jobsearcher.com/jobs/54407547803e2069725ac7e2","title":"LLM DevOps / Inference Engineer","description":"Must-Have RequirementsHands-on experience launching and managing inference containersStrong experience building and maintaining AWS infrastructure as code, including networking, compute, orchestration, secrets, observability, and CI/CDExperience working in a startup or fast-paced environmentStrong AI/LLM and model infrastructure experienceExcellent communication skillsUS Citizens will be given first preferenceActive and complete LinkedIn profileHealthcare experience is a plus, but not requiredRole OverviewWe are looking for an LLM DevOps / Inference Engineer to support a healthcare-focused AI benchmark and evaluation platform.The engineer will own the infrastructure supporting the benchmark harness, self-hosted model serving stack, and overall platform reliability. This role is ideal for someone comfortable provisioning GPU infrastructure while also tuning and optimizing model inference.Key ResponsibilitiesBuild and maintain AWS infrastructure using Terraform, including networking, compute, orchestration, secrets, observability, and CI/CDDeploy and manage self-hosted LLM inference workloads on GPU infrastructureImplement batching, quantization, autoscaling, and multi-GPU strategies for efficient inferenceDevelop standardized processes for onboarding new models quicklyBuild provider abstraction layers for API-based models such as OpenAI, Anthropic, and Gemini, as well as self-hosted modelsImplement rate limiting, retries, backoff, quota management, request/response logging, and cost attributionEnsure benchmark runs are reproducible and cost-efficient through pinned model/container versions and optimized cloud capacityBuild observability dashboards covering throughput, latency, token usage, GPU-hour costs, failures, and per-model performanceSupport burst workloads while preventing orphaned or unnecessary cloud resourcesContribute to Trusted Execution Environment (TEE) architecture and evaluate AWS Nitro Enclaves and comparable confidential-computing solutionsRequired SkillsProduction-scale AWS infrastructure experienceStrong knowledge of EKS/ECS, EC2 GPU instances (G5/G6, P4d/P5), VPC, IAM, KMS, Secrets Manager, ECR, CloudWatch, and Service QuotasStrong Terraform / Infrastructure as Code experienceHands-on experience with LLM model serving and inference optimizationPractical knowledge of batching strategies, KV cache behavior, quantization, and multi-GPU shardingExperience managing GPU workloads with EKS/ECS, node autoscaling, and CUDA-based container pipelinesStrong reliability and cost-engineering experience, including SLOs, monitoring, alerting, and cloud cost optimizationDesirable SkillsExperience with AWS SageMaker and BedrockStrong Python development skillsKnowledge of confidential computing, including Enclaves, remote attestation, and sealed key releaseUnderstanding of healthcare compliance, including HIPAA-eligible services, BAA requirements, audit logging, and PHI access controlsExperience with security hardening, image scanning, and network egress controls","company":"Xaxis Solutions","rawCompany":"xaxis solutions","city":"Denver","state":"CO","isRemote":false,"isActive":false,"createdAt":"2026-08-14T14:59:38.273Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1244.00","title":"Network and Computer Systems Administrators","slug":"network-and-computer-systems-administrators"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"LLM DevOps / Inference Engineer","description":"Must-Have RequirementsHands-on experience launching and managing inference containersStrong experience building and maintaining AWS infrastructure as code, including networking, compute, orchestration, secrets, observability, and CI/CDExperience working in a startup or fast-paced environmentStrong AI/LLM and model infrastructure experienceExcellent communication skillsUS Citizens will be given first preferenceActive and complete LinkedIn profileHealthcare experience is a plus, but not requiredRole OverviewWe are looking for an LLM DevOps / Inference Engineer to support a healthcare-focused AI benchmark and evaluation platform.The engineer will own the infrastructure supporting the benchmark harness, self-hosted model serving stack, and overall platform reliability. This role is ideal for someone comfortable provisioning GPU infrastructure while also tuning and optimizing model inference.Key ResponsibilitiesBuild and maintain AWS infrastructure using Terraform, including networking, compute, orchestration, secrets, observability, and CI/CDDeploy and manage self-hosted LLM inference workloads on GPU infrastructureImplement batching, quantization, autoscaling, and multi-GPU strategies for efficient inferenceDevelop standardized processes for onboarding new models quicklyBuild provider abstraction layers for API-based models such as OpenAI, Anthropic, and Gemini, as well as self-hosted modelsImplement rate limiting, retries, backoff, quota management, request/response logging, and cost attributionEnsure benchmark runs are reproducible and cost-efficient through pinned model/container versions and optimized cloud capacityBuild observability dashboards covering throughput, latency, token usage, GPU-hour costs, failures, and per-model performanceSupport burst workloads while preventing orphaned or unnecessary cloud resourcesContribute to Trusted Execution Environment (TEE) architecture and evaluate AWS Nitro Enclaves and comparable confidential-computing solutionsRequired SkillsProduction-scale AWS infrastructure experienceStrong knowledge of EKS/ECS, EC2 GPU instances (G5/G6, P4d/P5), VPC, IAM, KMS, Secrets Manager, ECR, CloudWatch, and Service QuotasStrong Terraform / Infrastructure as Code experienceHands-on experience with LLM model serving and inference optimizationPractical knowledge of batching strategies, KV cache behavior, quantization, and multi-GPU shardingExperience managing GPU workloads with EKS/ECS, node autoscaling, and CUDA-based container pipelinesStrong reliability and cost-engineering experience, including SLOs, monitoring, alerting, and cloud cost optimizationDesirable SkillsExperience with AWS SageMaker and BedrockStrong Python development skillsKnowledge of confidential computing, including Enclaves, remote attestation, and sealed key releaseUnderstanding of healthcare compliance, including HIPAA-eligible services, BAA requirements, audit logging, and PHI access controlsExperience with security hardening, image scanning, and network egress controls","datePosted":"2026-08-14T14:59:38.273Z","dateModified":"2026-08-14T14:59:38.273Z","hiringOrganization":{"@type":"Organization","name":"Xaxis Solutions","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Denver","addressRegion":"CO","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"54407547803e2069725ac7e2"},"url":"https://jobsearcher.com/jobs/54407547803e2069725ac7e2"}}