{"schemaVersion":"jobsearcher.job.v1","id":"1a1a4fbea62b1443fdba0962","url":"https://jobsearcher.com/jobs/1a1a4fbea62b1443fdba0962","canonicalUrl":"https://jobsearcher.com/jobs/1a1a4fbea62b1443fdba0962","title":"ML Platform Engineer - GPU Infrastructure","description":"Job Title: ML Platform Engineer - GPU Infrastructure\r\nSupport team by designing, implementing, and maintaining the automation and ML workload enablement layer of the GPU cluster platform. This role focuses on optimizing GPU compute environments for AI/ML training and Isaac Sim simulation workloads, integrating GPU jobs into CI/CD pipelines, standardizing runtime environments, and supporting reliable storage and artifact management.\r\nRequired Experience\r\n3+ years of experience in ML Platform Engineering, DevOps, Infrastructure Engineering, or related field\r\nBachelor's or Master's degree in Systems Engineering, Computer Science, Computer Engineering, or related discipline\r\nResponsibilities\r\nSupport GPU cluster platforms for AI/ML and simulation workloads\r\nOptimize GPU compute environments for ML training and Isaac Sim execution\r\nIntegrate GPU workload execution into CI/CD pipelines\r\nStandardize runtime environments using containers and automation tools\r\nManage storage, artifacts, and workload outputs\r\nTroubleshoot and improve platform reliability, scalability, and performance\r\nCollaborate with ML, infrastructure, and engineering teams\r\nRequired Skills\r\nExperience with Linux, Kubernetes, Docker, and GPU infrastructure\r\nKnowledge of CI/CD tools and automation scripting (Python/Bash)\r\nExperience supporting AI/ML workloads and distributed systems\r\nFamiliarity with NVIDIA GPU technologies and containerized environments\r\nStrong troubleshooting and performance optimization skills\r\nPreferred Skills\r\nExperience with Isaac Sim or simulation workloads\r\nExposure to cloud platforms (AWS, Azure, or GCP)\r\nKnowledge of monitoring and observability tools such as Grafana or Prometheus\r\nJ-18808-Ljbffr","company":"Optimal Cae","rawCompany":"optimal cae","city":"Brooklyn","state":"NY","isRemote":false,"isActive":false,"createdAt":"2026-06-25T01:12:02.537Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1244.00","title":"Network and Computer Systems Administrators","slug":"network-and-computer-systems-administrators"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"ML Platform Engineer - GPU Infrastructure","description":"Job Title: ML Platform Engineer - GPU Infrastructure\r\nSupport team by designing, implementing, and maintaining the automation and ML workload enablement layer of the GPU cluster platform. This role focuses on optimizing GPU compute environments for AI/ML training and Isaac Sim simulation workloads, integrating GPU jobs into CI/CD pipelines, standardizing runtime environments, and supporting reliable storage and artifact management.\r\nRequired Experience\r\n3+ years of experience in ML Platform Engineering, DevOps, Infrastructure Engineering, or related field\r\nBachelor's or Master's degree in Systems Engineering, Computer Science, Computer Engineering, or related discipline\r\nResponsibilities\r\nSupport GPU cluster platforms for AI/ML and simulation workloads\r\nOptimize GPU compute environments for ML training and Isaac Sim execution\r\nIntegrate GPU workload execution into CI/CD pipelines\r\nStandardize runtime environments using containers and automation tools\r\nManage storage, artifacts, and workload outputs\r\nTroubleshoot and improve platform reliability, scalability, and performance\r\nCollaborate with ML, infrastructure, and engineering teams\r\nRequired Skills\r\nExperience with Linux, Kubernetes, Docker, and GPU infrastructure\r\nKnowledge of CI/CD tools and automation scripting (Python/Bash)\r\nExperience supporting AI/ML workloads and distributed systems\r\nFamiliarity with NVIDIA GPU technologies and containerized environments\r\nStrong troubleshooting and performance optimization skills\r\nPreferred Skills\r\nExperience with Isaac Sim or simulation workloads\r\nExposure to cloud platforms (AWS, Azure, or GCP)\r\nKnowledge of monitoring and observability tools such as Grafana or Prometheus\r\nJ-18808-Ljbffr","datePosted":"2026-06-25T01:12:02.537Z","dateModified":"2026-06-25T01:12:02.537Z","hiringOrganization":{"@type":"Organization","name":"Optimal Cae","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Brooklyn","addressRegion":"NY","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"1a1a4fbea62b1443fdba0962"},"url":"https://jobsearcher.com/jobs/1a1a4fbea62b1443fdba0962"}}