{"schemaVersion":"jobsearcher.job.v1","id":"2619362eea6caab4cf6e48cd","url":"https://jobsearcher.com/jobs/2619362eea6caab4cf6e48cd","canonicalUrl":"https://jobsearcher.com/jobs/2619362eea6caab4cf6e48cd","title":"AI Platform Support Engineer (US)","description":"Who We Are\nLightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems—designed to take ideas from research to production with less friction.\nThrough our merger with Voltage Park, a neocloud and AI Factory, Lightning AI combines developer-first software with cost-efficient, large-scale compute. Teams get the tools they need for experimentation, training, and production inference, with security, observability, and control built in.\nWe serve solo researchers, startups, and large enterprises. Lightning AI operates globally with offices in New York City, San Francisco, Seattle, and London, and is backed by Coatue, Index Ventures, Bain Capital Ventures, and Firstminute.\nThe Way We Work\nThe people who thrive here are builders who move fast, communicate openly, take ownership, and continuously improve themselves, their teams, and our company. Here's what that looks like in practice:\nMove with Urgency: We move quickly, make thoughtful decisions, and keep momentum. We value action over perfection and learn by shipping.\nTake Ownership: We own outcomes, not just our individual work. We make decisions that move the company forward and follow through.\nCommunicate Openly: We communicate directly, seek to understand, and create clarity for others. Honest conversations help us move faster together.\nBuild Great Teams: We lead by example, empower others, and create healthy teams where people can do their best work.\nRaise the Bar: We're always improving ourselves. We learn from feedback, consistently challenge ourselves to grow, and focus on the work that matters most.\nThink Long-Term: We design for what's next. We create scalable systems, simplify complexity, and use AI and automation to amplify our impact.\nWhat We're Looking For\nLightning AI is looking to hire an AI Platform Support Engineer to join our US Customer Experience team, supporting ML engineers running large-scale training and inference workloads across cloud infrastructure, Kubernetes, and GPU platforms in production environments.\nThis role sits at the intersection of ML systems, cloud infrastructure, Kubernetes, and customers. You'll support engineers training models, deploying inference systems, and scaling GPU workloads in production.You are not a ticket router or traditional support engineer. You are a technical partner to ML teams - helping diagnose failures, improve reliability, and guide customers through complex distributed systems problems.\nThe problems range from Kubernetes scheduling and GPU orchestration to distributed PyTorch failures, inference latency, networking bottlenecks, storage performance, and platform reliability. You'll gain exposure to a wide variety of real world AI workloads across industries and help shape the infrastructure powering the next generation of ML applications.\nThis role is hybrid out of our Seattle, San Francisco, or New York office hubs, with an in-office requirement of at least 2 days per week and occasional team and company offsites. The role follows a Monday–Friday schedule, with working hours from 8:00 AM to 5:00 PM PST. We are not able to provide visa sponsorship for this role at this time.\n\nWhat You'll Do\nWork Directly With ML Engineers\nPartner directly with customer engineering teams running training and inference workloads in production\nHelp customers diagnose and resolve complex distributed systems and ML infrastructure issues\nAct as a technical advisor during high impact incidents and platform degradation events\nTranslate infrastructure level issues into actionable guidance for ML engineers\nBuild credibility with customers through strong technical reasoning and clear communication\nDebug ML Infrastructure & Distributed Workloads\nInvestigate failures involving distributed training, Kubernetes orchestration, GPU allocation, networking, and storage systems\nTroubleshoot PyTorch, CUDA, NCCL, and inference serving related issues\nAnalyze logs, metrics, traces, and system behavior to isolate root causes\nDebug containerized workloads running across Kubernetes and bare metal GPU environments\nSupport customers scaling workloads across multi node GPU systems\nDiagnose performance bottlenecks involving compute, memory, networking, or storage\nImprove Reliability & Platform Operations\nIdentify recurring patterns across customer issues and drive long term reliability improvements\nContribute to post incident reviews and operational improvements\nBuild internal tooling, automation, documentation, and runbooks\nPartner closely with infrastructure, networking, and platform engineering teams\nHelp improve observability, operational visibility, and troubleshooting workflows\nImprove the customer experience through better processes and technical guidance\nWhat This Role Is Not\nTo set clear expectations:\nThis is not a traditional help desk or ticket routing support role\nThis is not purely customer success or account management\nThis is not a backend engineering role\nThis is not a passive escalation position\nThis role is for engineers who enjoy solving difficult technical problems while working closely with other engineers.\n\nWhat You'll Need\nRequired Qualifications\nInfrastructure & Systems\nStrong software engineering and systems troubleshooting background\nExperience with Kubernetes and containerized environments\nLinux systems knowledge, including networking, storage, process management, and performance tuning\nExperience with cloud infrastructure and distributed systems\nExperience with observability and debugging tools such as Prometheus, Grafana, or OpenTelemetry\nML Infrastructure Experience\nHands on experience operating machine learning workloads in production or research environments\nExperience with distributed ML systems and tooling such as PyTorch, CUDA, or NCCL\nFamiliarity with GPU infrastructure and orchestration\nExperience troubleshooting performance, reliability, or scaling issues in ML infrastructure\nUnderstanding of the operational challenges involved in running ML systems at scale\nCollaboration\nStrong communication skills and ability to work directly with highly technical customers and engineering teams\nComfortable operating in fast moving, highly ambiguous environments\nEnjoys solving complex technical problems collaboratively\nIdeal Experience\nExperience with large scale model training or distributed inference systems\nFamiliarity with Ray, Kubeflow, Slurm, or similar distributed scheduling platforms\nExperience with InfiniBand, RDMA, or high-performance networking\nExperience operating bare metal infrastructure\nFamiliarity with storage systems commonly used in ML environments\nExperience working at an AI infrastructure, cloud, MLOps, or developer tooling company\nContributions to platform engineering, developer infrastructure, or operational tooling projects\nExperience writing automation, tooling, or scripts in Python or similar languages\nBenefits and Perks\nWe offer a comprehensive and competitive benefits package designed to support our employees' health, well-being, and long-term success:\nComprehensive Health Coverage: Medical, dental, and vision coverage for employees and eligible dependents.\nMeaningful Equity: RSUs that give employees a stake in the company's long-term success.\nRetirement Savings: 401(k) matching (U.S.) and pension contributions (U.K.).\nFlexible Time Off: Unlimited PTO, company holidays, and floating holidays to support work-life balance.\nCompany-Wide Winter Break: Two weeks of company closure each winter to disconnect and recharge.\nPaid Parental & Family Leave: Paid leave to support you and your family through life's important moments.\nProfessional Development: Annual learning and development allowance to support your professional growth.\nWellness Benefits: Wellness and work-from-home stipends to support your physical and mental well-being.\nSabbatical Program: Four weeks of paid sabbatical leave after four years of service.\nFlexible Work: Flexible schedules and a hybrid work model for our office-based teams.\nIn-Office Meals: Complimentary meals at our office hubs.\nBenefits may vary by location, team, and role.\n\nAt Lightning AI, we are committed to fostering an inclusive and diverse workplace. We believe that diverse teams drive innovation and create better products. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity, national origin, age, disability, veteran status, or any other protected characteristic. We are dedicated to building a culture where everyone can thrive and contribute to their fullest potential.","company":"Lightningai","rawCompany":"lightningai","city":"Millbrae","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-08-04T17:32:31.300Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1232.00","title":"Computer User Support Specialists","slug":"computer-user-support-specialists"},{"code":"15-1221.00","title":"Computer and Information Research Scientists","slug":"computer-and-information-research-scientists"}],"industries":[{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"AI Platform Support Engineer (US)","description":"Who We Are\nLightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems—designed to take ideas from research to production with less friction.\nThrough our merger with Voltage Park, a neocloud and AI Factory, Lightning AI combines developer-first software with cost-efficient, large-scale compute. Teams get the tools they need for experimentation, training, and production inference, with security, observability, and control built in.\nWe serve solo researchers, startups, and large enterprises. Lightning AI operates globally with offices in New York City, San Francisco, Seattle, and London, and is backed by Coatue, Index Ventures, Bain Capital Ventures, and Firstminute.\nThe Way We Work\nThe people who thrive here are builders who move fast, communicate openly, take ownership, and continuously improve themselves, their teams, and our company. Here's what that looks like in practice:\nMove with Urgency: We move quickly, make thoughtful decisions, and keep momentum. We value action over perfection and learn by shipping.\nTake Ownership: We own outcomes, not just our individual work. We make decisions that move the company forward and follow through.\nCommunicate Openly: We communicate directly, seek to understand, and create clarity for others. Honest conversations help us move faster together.\nBuild Great Teams: We lead by example, empower others, and create healthy teams where people can do their best work.\nRaise the Bar: We're always improving ourselves. We learn from feedback, consistently challenge ourselves to grow, and focus on the work that matters most.\nThink Long-Term: We design for what's next. We create scalable systems, simplify complexity, and use AI and automation to amplify our impact.\nWhat We're Looking For\nLightning AI is looking to hire an AI Platform Support Engineer to join our US Customer Experience team, supporting ML engineers running large-scale training and inference workloads across cloud infrastructure, Kubernetes, and GPU platforms in production environments.\nThis role sits at the intersection of ML systems, cloud infrastructure, Kubernetes, and customers. You'll support engineers training models, deploying inference systems, and scaling GPU workloads in production.You are not a ticket router or traditional support engineer. You are a technical partner to ML teams - helping diagnose failures, improve reliability, and guide customers through complex distributed systems problems.\nThe problems range from Kubernetes scheduling and GPU orchestration to distributed PyTorch failures, inference latency, networking bottlenecks, storage performance, and platform reliability. You'll gain exposure to a wide variety of real world AI workloads across industries and help shape the infrastructure powering the next generation of ML applications.\nThis role is hybrid out of our Seattle, San Francisco, or New York office hubs, with an in-office requirement of at least 2 days per week and occasional team and company offsites. The role follows a Monday–Friday schedule, with working hours from 8:00 AM to 5:00 PM PST. We are not able to provide visa sponsorship for this role at this time.\n\nWhat You'll Do\nWork Directly With ML Engineers\nPartner directly with customer engineering teams running training and inference workloads in production\nHelp customers diagnose and resolve complex distributed systems and ML infrastructure issues\nAct as a technical advisor during high impact incidents and platform degradation events\nTranslate infrastructure level issues into actionable guidance for ML engineers\nBuild credibility with customers through strong technical reasoning and clear communication\nDebug ML Infrastructure & Distributed Workloads\nInvestigate failures involving distributed training, Kubernetes orchestration, GPU allocation, networking, and storage systems\nTroubleshoot PyTorch, CUDA, NCCL, and inference serving related issues\nAnalyze logs, metrics, traces, and system behavior to isolate root causes\nDebug containerized workloads running across Kubernetes and bare metal GPU environments\nSupport customers scaling workloads across multi node GPU systems\nDiagnose performance bottlenecks involving compute, memory, networking, or storage\nImprove Reliability & Platform Operations\nIdentify recurring patterns across customer issues and drive long term reliability improvements\nContribute to post incident reviews and operational improvements\nBuild internal tooling, automation, documentation, and runbooks\nPartner closely with infrastructure, networking, and platform engineering teams\nHelp improve observability, operational visibility, and troubleshooting workflows\nImprove the customer experience through better processes and technical guidance\nWhat This Role Is Not\nTo set clear expectations:\nThis is not a traditional help desk or ticket routing support role\nThis is not purely customer success or account management\nThis is not a backend engineering role\nThis is not a passive escalation position\nThis role is for engineers who enjoy solving difficult technical problems while working closely with other engineers.\n\nWhat You'll Need\nRequired Qualifications\nInfrastructure & Systems\nStrong software engineering and systems troubleshooting background\nExperience with Kubernetes and containerized environments\nLinux systems knowledge, including networking, storage, process management, and performance tuning\nExperience with cloud infrastructure and distributed systems\nExperience with observability and debugging tools such as Prometheus, Grafana, or OpenTelemetry\nML Infrastructure Experience\nHands on experience operating machine learning workloads in production or research environments\nExperience with distributed ML systems and tooling such as PyTorch, CUDA, or NCCL\nFamiliarity with GPU infrastructure and orchestration\nExperience troubleshooting performance, reliability, or scaling issues in ML infrastructure\nUnderstanding of the operational challenges involved in running ML systems at scale\nCollaboration\nStrong communication skills and ability to work directly with highly technical customers and engineering teams\nComfortable operating in fast moving, highly ambiguous environments\nEnjoys solving complex technical problems collaboratively\nIdeal Experience\nExperience with large scale model training or distributed inference systems\nFamiliarity with Ray, Kubeflow, Slurm, or similar distributed scheduling platforms\nExperience with InfiniBand, RDMA, or high-performance networking\nExperience operating bare metal infrastructure\nFamiliarity with storage systems commonly used in ML environments\nExperience working at an AI infrastructure, cloud, MLOps, or developer tooling company\nContributions to platform engineering, developer infrastructure, or operational tooling projects\nExperience writing automation, tooling, or scripts in Python or similar languages\nBenefits and Perks\nWe offer a comprehensive and competitive benefits package designed to support our employees' health, well-being, and long-term success:\nComprehensive Health Coverage: Medical, dental, and vision coverage for employees and eligible dependents.\nMeaningful Equity: RSUs that give employees a stake in the company's long-term success.\nRetirement Savings: 401(k) matching (U.S.) and pension contributions (U.K.).\nFlexible Time Off: Unlimited PTO, company holidays, and floating holidays to support work-life balance.\nCompany-Wide Winter Break: Two weeks of company closure each winter to disconnect and recharge.\nPaid Parental & Family Leave: Paid leave to support you and your family through life's important moments.\nProfessional Development: Annual learning and development allowance to support your professional growth.\nWellness Benefits: Wellness and work-from-home stipends to support your physical and mental well-being.\nSabbatical Program: Four weeks of paid sabbatical leave after four years of service.\nFlexible Work: Flexible schedules and a hybrid work model for our office-based teams.\nIn-Office Meals: Complimentary meals at our office hubs.\nBenefits may vary by location, team, and role.\n\nAt Lightning AI, we are committed to fostering an inclusive and diverse workplace. We believe that diverse teams drive innovation and create better products. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity, national origin, age, disability, veteran status, or any other protected characteristic. We are dedicated to building a culture where everyone can thrive and contribute to their fullest potential.","datePosted":"2026-08-04T17:32:31.300Z","dateModified":"2026-08-04T17:32:31.300Z","hiringOrganization":{"@type":"Organization","name":"Lightningai","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Millbrae","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"2619362eea6caab4cf6e48cd"},"url":"https://jobsearcher.com/jobs/2619362eea6caab4cf6e48cd"}}