{"schemaVersion":"jobsearcher.job.v1","id":"73f3bcbc75b1ac0cb690b193","url":"https://jobsearcher.com/jobs/73f3bcbc75b1ac0cb690b193","canonicalUrl":"https://jobsearcher.com/jobs/73f3bcbc75b1ac0cb690b193","title":"Infrastructure Engineer (GPU & Compute)","description":"Who We Are\nLightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems—designed to take ideas from research to production with less friction.\nThrough our merger with Voltage Park, a neocloud and AI Factory, Lightning AI combines developer-first software with cost-efficient, large-scale compute. Teams get the tools they need for experimentation, training, and production inference, with security, observability, and control built in.\nWe serve solo researchers, startups, and large enterprises. Lightning AI operates globally with offices in New York City, San Francisco, Seattle, and London, and is backed by Coatue, Index Ventures, Bain Capital Ventures, and Firstminute.\n\nWhat We're Looking For\nLightning AI is seeking a GPU & Compute Infrastructure Engineer to join our Infrastructure Engineering team.\nIn this role, you will own image management, system diagnostics, and validation across large-scale bare-metal compute infrastructure, with a particular focus on GPU-enabled systems. You will work at the intersection of hardware, systems, and software—developing automation, improving reliability, and enabling efficient cluster bring-up for AI/ML and HPC workloads.\nYou will play a key role in owning and evolving our image pipeline, running validation environments and test clusters, and supporting both system-level and GPU hardware qualification. This role is critical to ensuring that our infrastructure is consistent, performant, and ready to support demanding AI workloads from day one.\nWe're flexible on location for this team. This role can work hybrid out of one of our US-based hubs (Seattle, NYC, or SF) or fully remote within the U.S., with occasional company and team offsites. We are not able to provide visa sponsorship for this position at this time.\nWhat You'll Do\nSystems, Image & Validation Infrastructure\nOwn and evolve systems for image management, deployment, and validation across bare-metal infrastructure\nRun and maintain test clusters used for system validation, diagnostics, and bring-up\nValidate firmware, drivers, and OS images across compute and GPU-enabled systems\nSupport hardware qualification efforts for next-generation platforms\nGPU Diagnostics & Performance\nOwn GPU diagnostics and validation workflows across large-scale infrastructure\nDiagnose and resolve complex issues across GPUs, drivers, OS, and hardware layers\nAnalyze system and GPU performance using tools such as NVIDIA DCGM\nIdentify failure patterns and drive improvements in system stability and validation coverage\nAutomation & Tooling\nBuild and maintain automation for provisioning, validation, and system bring-up\nDevelop Python-based tools and workflows to improve efficiency and reduce manual operational overhead\nImprove the reliability, repeatability, and scalability of image pipelines and validation systems\nSystems & Operations\nManage and operate Linux-based systems in production and validation environments\nManage virtualization technology\nSupport bare-metal provisioning workflows, including PXE and image-based systems\nInterface with hardware management systems (e.g., IPMI, Redfish) for monitoring and debugging\nCross-Functional Collaboration\nPartner with Infrastructure, Hardware, and Data Center teams on system bring-up and validation\nCollaborate with platform and ML teams to ensure systems meet workload requirements\nContribute to best practices for provisioning, diagnostics, and lifecycle management of infrastructure\nWhat You'll Need\nRequired Qualifications\n5+ years of experience in infrastructure engineering, systems engineering, or related roles\nStrong Linux systems experience in production environments\nHands-on experience with GPU-enabled systems and tools such as NVIDIA DCGM\nFamiliarity with bare-metal provisioning and system bring-up workflows\nProficiency in Python or similar scripting/programming languages for automation\nAbility to debug complex issues across hardware, OS, GPUs, and system software\nIdeal Experience\nExperience with high-performance interconnects (e.g., InfiniBand, NVLink)\nExperience with PXE boot environments, LiveCD systems, or image-based provisioning workflows\nExperience with hardware management interfaces such as iDRAC, IPMI, or Redfish\nData center operations experience, including working with physical hardware\nExperience supporting AI/ML or HPC workloads at scale\nExperience with GPU validation frameworks or large-scale hardware qualification processes\n\nCompensation\nBenefits and Perks\nWe offer a comprehensive and competitive benefits package designed to support our employees' health, well-being, and long-term success. Benefits may vary by location, team, and role.\nBenefits include:\nComprehensive medical, dental and vision coverage (U.S.); Private medical and dental insurance (U.K.)\nRetirement and financial wellness support (U.S.); Pension contribution (U.K.)\nGenerous paid time off, plus holidays\nPaid parental leave\nProfessional development support\nWellness and work-from-home stipends\nFlexible work environment\n\nAt Lightning AI, we are committed to fostering an inclusive and diverse workplace. We believe that diverse teams drive innovation and create better products. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity, national origin, age, disability, veteran status, or any other protected characteristic. We are dedicated to building a culture where everyone can thrive and contribute to their fullest potential.","company":"Lightningai","rawCompany":"lightningai","city":"Millbrae","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-08-05T11:50:20.558Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"11-3021.00","title":"Computer and Information Systems Managers","slug":"computer-and-information-systems-managers"},{"code":"15-1244.00","title":"Network and Computer Systems Administrators","slug":"network-and-computer-systems-administrators"}],"industries":[{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Infrastructure Engineer (GPU & Compute)","description":"Who We Are\nLightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems—designed to take ideas from research to production with less friction.\nThrough our merger with Voltage Park, a neocloud and AI Factory, Lightning AI combines developer-first software with cost-efficient, large-scale compute. Teams get the tools they need for experimentation, training, and production inference, with security, observability, and control built in.\nWe serve solo researchers, startups, and large enterprises. Lightning AI operates globally with offices in New York City, San Francisco, Seattle, and London, and is backed by Coatue, Index Ventures, Bain Capital Ventures, and Firstminute.\n\nWhat We're Looking For\nLightning AI is seeking a GPU & Compute Infrastructure Engineer to join our Infrastructure Engineering team.\nIn this role, you will own image management, system diagnostics, and validation across large-scale bare-metal compute infrastructure, with a particular focus on GPU-enabled systems. You will work at the intersection of hardware, systems, and software—developing automation, improving reliability, and enabling efficient cluster bring-up for AI/ML and HPC workloads.\nYou will play a key role in owning and evolving our image pipeline, running validation environments and test clusters, and supporting both system-level and GPU hardware qualification. This role is critical to ensuring that our infrastructure is consistent, performant, and ready to support demanding AI workloads from day one.\nWe're flexible on location for this team. This role can work hybrid out of one of our US-based hubs (Seattle, NYC, or SF) or fully remote within the U.S., with occasional company and team offsites. We are not able to provide visa sponsorship for this position at this time.\nWhat You'll Do\nSystems, Image & Validation Infrastructure\nOwn and evolve systems for image management, deployment, and validation across bare-metal infrastructure\nRun and maintain test clusters used for system validation, diagnostics, and bring-up\nValidate firmware, drivers, and OS images across compute and GPU-enabled systems\nSupport hardware qualification efforts for next-generation platforms\nGPU Diagnostics & Performance\nOwn GPU diagnostics and validation workflows across large-scale infrastructure\nDiagnose and resolve complex issues across GPUs, drivers, OS, and hardware layers\nAnalyze system and GPU performance using tools such as NVIDIA DCGM\nIdentify failure patterns and drive improvements in system stability and validation coverage\nAutomation & Tooling\nBuild and maintain automation for provisioning, validation, and system bring-up\nDevelop Python-based tools and workflows to improve efficiency and reduce manual operational overhead\nImprove the reliability, repeatability, and scalability of image pipelines and validation systems\nSystems & Operations\nManage and operate Linux-based systems in production and validation environments\nManage virtualization technology\nSupport bare-metal provisioning workflows, including PXE and image-based systems\nInterface with hardware management systems (e.g., IPMI, Redfish) for monitoring and debugging\nCross-Functional Collaboration\nPartner with Infrastructure, Hardware, and Data Center teams on system bring-up and validation\nCollaborate with platform and ML teams to ensure systems meet workload requirements\nContribute to best practices for provisioning, diagnostics, and lifecycle management of infrastructure\nWhat You'll Need\nRequired Qualifications\n5+ years of experience in infrastructure engineering, systems engineering, or related roles\nStrong Linux systems experience in production environments\nHands-on experience with GPU-enabled systems and tools such as NVIDIA DCGM\nFamiliarity with bare-metal provisioning and system bring-up workflows\nProficiency in Python or similar scripting/programming languages for automation\nAbility to debug complex issues across hardware, OS, GPUs, and system software\nIdeal Experience\nExperience with high-performance interconnects (e.g., InfiniBand, NVLink)\nExperience with PXE boot environments, LiveCD systems, or image-based provisioning workflows\nExperience with hardware management interfaces such as iDRAC, IPMI, or Redfish\nData center operations experience, including working with physical hardware\nExperience supporting AI/ML or HPC workloads at scale\nExperience with GPU validation frameworks or large-scale hardware qualification processes\n\nCompensation\nBenefits and Perks\nWe offer a comprehensive and competitive benefits package designed to support our employees' health, well-being, and long-term success. Benefits may vary by location, team, and role.\nBenefits include:\nComprehensive medical, dental and vision coverage (U.S.); Private medical and dental insurance (U.K.)\nRetirement and financial wellness support (U.S.); Pension contribution (U.K.)\nGenerous paid time off, plus holidays\nPaid parental leave\nProfessional development support\nWellness and work-from-home stipends\nFlexible work environment\n\nAt Lightning AI, we are committed to fostering an inclusive and diverse workplace. We believe that diverse teams drive innovation and create better products. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity, national origin, age, disability, veteran status, or any other protected characteristic. We are dedicated to building a culture where everyone can thrive and contribute to their fullest potential.","datePosted":"2026-08-05T11:50:20.558Z","dateModified":"2026-08-05T11:50:20.558Z","hiringOrganization":{"@type":"Organization","name":"Lightningai","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Millbrae","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"73f3bcbc75b1ac0cb690b193"},"url":"https://jobsearcher.com/jobs/73f3bcbc75b1ac0cb690b193"}}