{"schemaVersion":"jobsearcher.job.v1","id":"6ddd2dd313437046c16c27bb","url":"https://jobsearcher.com/jobs/6ddd2dd313437046c16c27bb","canonicalUrl":"https://jobsearcher.com/jobs/6ddd2dd313437046c16c27bb","title":"Platform Support Engineer (APAC)","description":"Who We Are\nLightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems—designed to take ideas from research to production with less friction.\nThrough our merger with Voltage Park, a neocloud and AI Factory, Lightning AI combines developer-first software with cost-efficient, large-scale compute. Teams get the tools they need for experimentation, training, and production inference, with security, observability, and control built in.\nWe serve solo researchers, startups, and large enterprises. Lightning AI operates globally with offices in New York City, San Francisco, Seattle, and London, and is backed by Coatue, Index Ventures, Bain Capital Ventures, and Firstminute.\nOur Values\nMove Fast: We act with speed and precision, breaking down big challenges into achievable steps.\nFocus: We complete one goal at a time with care, collaborating as a team to deliver features with precision.\nBalance: Sustained performance comes from rest and recovery. We ensure a healthy work-life balance to keep you at your best.\nCraftsmanship: Innovation through excellence. Every detail matters, and we take pride in mastering our craft.\nMinimal: Simplicity drives our innovation. We eliminate complexity through discipline and focus on what truly matters.\nWhat We’re Looking For\nLightning AI is looking to hire a Platform Support Engineer to join our APAC Customer Experience team, supporting ML engineers running large-scale training and inference workloads across cloud infrastructure, Kubernetes, and GPU platforms in production environments.\nThis role is not a ticket router or traditional support engineer. You are a technical partner to ML teams - helping diagnose failures, improve reliability, and guide customers through complex distributed systems problems.The problems range from Kubernetes scheduling and GPU orchestration to distributed PyTorch failures, inference latency, networking bottlenecks, storage performance, and platform reliability. You’ll gain exposure to a wide variety of real world AI workloads across industries and help shape the infrastructure powering the next generation of ML applications.\nThis role is remote and open to candidates based in either the Philippines or Singapore. The role follows a Thursday–Sunday schedule, with working hours from 7:00 AM to 5:00 PM local time (UTC+8).\n\nWhat You'll Do\nWork Directly With ML Engineers\nPartner directly with customer engineering teams running training and inference workloads in production\nHelp customers diagnose and resolve complex distributed systems and ML infrastructure issues\nAct as a technical advisor during high impact incidents and platform degradation events\nTranslate infrastructure level issues into actionable guidance for ML engineers\nBuild credibility with customers through strong technical reasoning and clear communication\nDebug ML Infrastructure & Distributed Workloads\nInvestigate failures involving distributed training, Kubernetes orchestration, GPU allocation, networking, and storage systems\nTroubleshoot PyTorch, CUDA, NCCL, and inference serving related issues\nAnalyze logs, metrics, traces, and system behavior to isolate root causes\nDebug containerized workloads running across Kubernetes and bare metal GPU environments\nSupport customers scaling workloads across multi node GPU systems\nDiagnose performance bottlenecks involving compute, memory, networking, or storage\nImprove Reliability & Platform Operations\nIdentify recurring patterns across customer issues and drive long term reliability improvements\nContribute to post incident reviews and operational improvements\nBuild internal tooling, automation, documentation, and runbooks\nPartner closely with infrastructure, networking, and platform engineering teams\nHelp improve observability, operational visibility, and troubleshooting workflows\nImprove the customer experience through better processes and technical guidance\nWhat This Role Is Not\nTo set clear expectations:\nThis is not a traditional help desk or ticket routing support role\nThis is not purely customer success or account management\nThis is not a backend engineering role\nThis is not a passive escalation position\nThis role is for engineers who enjoy solving difficult technical problems while working closely with other engineers.\n\nWhat You’ll Need\nRequired Qualifications\nInfrastructure & Systems\nStrong software engineering and systems troubleshooting background\nExperience with Kubernetes and containerized environments\nLinux systems knowledge, including networking, storage, process management, and performance tuning\nExperience with cloud infrastructure and distributed systems\nExperience with observability and debugging tools such as Prometheus, Grafana, or OpenTelemetry\nML Infrastructure Experience\nHands on experience operating machine learning workloads in production or research environments\nExperience with distributed ML systems and tooling such as PyTorch, CUDA, or NCCL\nFamiliarity with GPU infrastructure and orchestration\nExperience troubleshooting performance, reliability, or scaling issues in ML infrastructure\nUnderstanding of the operational challenges involved in running ML systems at scale\nCollaboration\nStrong communication skills and ability to work directly with highly technical customers and engineering teams\nComfortable operating in fast moving, highly ambiguous environments\nEnjoys solving complex technical problems collaboratively\nNice-to-Haves\nExperience with large scale model training or distributed inference systems\nFamiliarity with Ray, Kubeflow, Slurm, or similar distributed scheduling platforms\nExperience with InfiniBand, RDMA, or high-performance networking\nExperience operating bare metal infrastructure\nFamiliarity with storage systems commonly used in ML environments\nExperience working at an AI infrastructure, cloud, MLOps, or developer tooling company\nContributions to platform engineering, developer infrastructure, or operational tooling projects\nExperience writing automation, tooling, or scripts in Python or similar languages\nBenefits and Perks\nWe offer a comprehensive and competitive benefits package designed to support our employees’ health, well-being, and long-term success. Benefits may vary by location, team, and role.\nBenefits include:\nComprehensive medical, dental and vision coverage (U.S.); Private medical and dental insurance (U.K.)\nRetirement and financial wellness support (U.S.); Pension contribution (U.K.)\nGenerous paid time off, plus holidays\nPaid parental leave\nProfessional development support\nWellness and work-from-home stipends\nFlexible work environment\n\nAt Lightning AI, we are committed to fostering an inclusive and diverse workplace. We believe that diverse teams drive innovation and create better products. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity, national origin, age, disability, veteran status, or any other protected characteristic. We are dedicated to building a culture where everyone can thrive and contribute to their fullest potential.","company":"Lightningai","rawCompany":"lightningai","city":"Remote","state":"OR","isRemote":false,"isActive":false,"createdAt":"2026-07-19T14:46:27.678Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1244.00","title":"Network and Computer Systems Administrators","slug":"network-and-computer-systems-administrators"}],"industries":[{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Platform Support Engineer (APAC)","description":"Who We Are\nLightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems—designed to take ideas from research to production with less friction.\nThrough our merger with Voltage Park, a neocloud and AI Factory, Lightning AI combines developer-first software with cost-efficient, large-scale compute. Teams get the tools they need for experimentation, training, and production inference, with security, observability, and control built in.\nWe serve solo researchers, startups, and large enterprises. Lightning AI operates globally with offices in New York City, San Francisco, Seattle, and London, and is backed by Coatue, Index Ventures, Bain Capital Ventures, and Firstminute.\nOur Values\nMove Fast: We act with speed and precision, breaking down big challenges into achievable steps.\nFocus: We complete one goal at a time with care, collaborating as a team to deliver features with precision.\nBalance: Sustained performance comes from rest and recovery. We ensure a healthy work-life balance to keep you at your best.\nCraftsmanship: Innovation through excellence. Every detail matters, and we take pride in mastering our craft.\nMinimal: Simplicity drives our innovation. We eliminate complexity through discipline and focus on what truly matters.\nWhat We’re Looking For\nLightning AI is looking to hire a Platform Support Engineer to join our APAC Customer Experience team, supporting ML engineers running large-scale training and inference workloads across cloud infrastructure, Kubernetes, and GPU platforms in production environments.\nThis role is not a ticket router or traditional support engineer. You are a technical partner to ML teams - helping diagnose failures, improve reliability, and guide customers through complex distributed systems problems.The problems range from Kubernetes scheduling and GPU orchestration to distributed PyTorch failures, inference latency, networking bottlenecks, storage performance, and platform reliability. You’ll gain exposure to a wide variety of real world AI workloads across industries and help shape the infrastructure powering the next generation of ML applications.\nThis role is remote and open to candidates based in either the Philippines or Singapore. The role follows a Thursday–Sunday schedule, with working hours from 7:00 AM to 5:00 PM local time (UTC+8).\n\nWhat You'll Do\nWork Directly With ML Engineers\nPartner directly with customer engineering teams running training and inference workloads in production\nHelp customers diagnose and resolve complex distributed systems and ML infrastructure issues\nAct as a technical advisor during high impact incidents and platform degradation events\nTranslate infrastructure level issues into actionable guidance for ML engineers\nBuild credibility with customers through strong technical reasoning and clear communication\nDebug ML Infrastructure & Distributed Workloads\nInvestigate failures involving distributed training, Kubernetes orchestration, GPU allocation, networking, and storage systems\nTroubleshoot PyTorch, CUDA, NCCL, and inference serving related issues\nAnalyze logs, metrics, traces, and system behavior to isolate root causes\nDebug containerized workloads running across Kubernetes and bare metal GPU environments\nSupport customers scaling workloads across multi node GPU systems\nDiagnose performance bottlenecks involving compute, memory, networking, or storage\nImprove Reliability & Platform Operations\nIdentify recurring patterns across customer issues and drive long term reliability improvements\nContribute to post incident reviews and operational improvements\nBuild internal tooling, automation, documentation, and runbooks\nPartner closely with infrastructure, networking, and platform engineering teams\nHelp improve observability, operational visibility, and troubleshooting workflows\nImprove the customer experience through better processes and technical guidance\nWhat This Role Is Not\nTo set clear expectations:\nThis is not a traditional help desk or ticket routing support role\nThis is not purely customer success or account management\nThis is not a backend engineering role\nThis is not a passive escalation position\nThis role is for engineers who enjoy solving difficult technical problems while working closely with other engineers.\n\nWhat You’ll Need\nRequired Qualifications\nInfrastructure & Systems\nStrong software engineering and systems troubleshooting background\nExperience with Kubernetes and containerized environments\nLinux systems knowledge, including networking, storage, process management, and performance tuning\nExperience with cloud infrastructure and distributed systems\nExperience with observability and debugging tools such as Prometheus, Grafana, or OpenTelemetry\nML Infrastructure Experience\nHands on experience operating machine learning workloads in production or research environments\nExperience with distributed ML systems and tooling such as PyTorch, CUDA, or NCCL\nFamiliarity with GPU infrastructure and orchestration\nExperience troubleshooting performance, reliability, or scaling issues in ML infrastructure\nUnderstanding of the operational challenges involved in running ML systems at scale\nCollaboration\nStrong communication skills and ability to work directly with highly technical customers and engineering teams\nComfortable operating in fast moving, highly ambiguous environments\nEnjoys solving complex technical problems collaboratively\nNice-to-Haves\nExperience with large scale model training or distributed inference systems\nFamiliarity with Ray, Kubeflow, Slurm, or similar distributed scheduling platforms\nExperience with InfiniBand, RDMA, or high-performance networking\nExperience operating bare metal infrastructure\nFamiliarity with storage systems commonly used in ML environments\nExperience working at an AI infrastructure, cloud, MLOps, or developer tooling company\nContributions to platform engineering, developer infrastructure, or operational tooling projects\nExperience writing automation, tooling, or scripts in Python or similar languages\nBenefits and Perks\nWe offer a comprehensive and competitive benefits package designed to support our employees’ health, well-being, and long-term success. Benefits may vary by location, team, and role.\nBenefits include:\nComprehensive medical, dental and vision coverage (U.S.); Private medical and dental insurance (U.K.)\nRetirement and financial wellness support (U.S.); Pension contribution (U.K.)\nGenerous paid time off, plus holidays\nPaid parental leave\nProfessional development support\nWellness and work-from-home stipends\nFlexible work environment\n\nAt Lightning AI, we are committed to fostering an inclusive and diverse workplace. We believe that diverse teams drive innovation and create better products. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity, national origin, age, disability, veteran status, or any other protected characteristic. We are dedicated to building a culture where everyone can thrive and contribute to their fullest potential.","datePosted":"2026-07-19T14:46:27.678Z","dateModified":"2026-07-19T14:46:27.678Z","hiringOrganization":{"@type":"Organization","name":"Lightningai","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Remote","addressRegion":"OR","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"6ddd2dd313437046c16c27bb"},"url":"https://jobsearcher.com/jobs/6ddd2dd313437046c16c27bb"}}