{"schemaVersion":"jobsearcher.job.v1","id":"6a3cbb4458bf1e3f2bd51a88","url":"https://jobsearcher.com/jobs/6a3cbb4458bf1e3f2bd51a88","canonicalUrl":"https://jobsearcher.com/jobs/6a3cbb4458bf1e3f2bd51a88","title":"Infrastructure Engineer (Storage)","description":"Who We Are\nLightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems—designed to take ideas from research to production with less friction.\nThrough our merger with Voltage Park, a neocloud and AI Factory, Lightning AI combines developer-first software with cost-efficient, large-scale compute. Teams get the tools they need for experimentation, training, and production inference, with security, observability, and control built in.\nWe serve solo researchers, startups, and large enterprises. Lightning AI operates globally with offices in New York City, San Francisco, Seattle, and London, and is backed by Coatue, Index Ventures, Bain Capital Ventures, and Firstminute.\nThe Way We Work\nThe people who thrive here are builders who move fast, communicate openly, take ownership, and continuously improve themselves, their teams, and our company. Here's what that looks like in practice:\nMove with Urgency: We move quickly, make thoughtful decisions, and keep momentum. We value action over perfection and learn by shipping.\nTake Ownership: We own outcomes, not just our individual work. We make decisions that move the company forward and follow through.\nCommunicate Openly: We communicate directly, seek to understand, and create clarity for others. Honest conversations help us move faster together.\nBuild Great Teams: We lead by example, empower others, and create healthy teams where people can do their best work.\nRaise the Bar: We're always improving ourselves. We learn from feedback, consistently challenge ourselves to grow, and focus on the work that matters most.\nThink Long-Term: We design for what's next. We create scalable systems, simplify complexity, and use AI and automation to amplify our impact.\nWhat We're Looking For\nLightning AI is seeking a Storage Infrastructure Engineer to join our Infrastructure Engineering team.\nIn this role, you will focus on building and operating the storage systems that power large-scale AI/ML training, inference, and HPC workloads. You will work at the intersection of software, hardware, and operations—developing automation, improving reliability, and scaling distributed storage systems across our bare-metal infrastructure.\nYou will help own the data plane of our storage infrastructure, supporting high-throughput, low-latency data access for some of the most demanding AI workloads. You'll play a key role in managing and evolving our storage stack (including VAST and S3-compatible systems like Ceph), ensuring performance, reliability, and efficiency at scale.\nThis role is based in one of our hubs (NYC, SF, Seattle, or London), with a minimum of 2 in-office days per week and occasional team and company offsites. We are not able to provide visa sponsorship for this position at this time.\nWhat You'll Do\nStorage Systems & Infrastructure\nOperate and scale distributed storage systems, including VAST and S3-compatible object storage (e.g., Ceph)\nImprove performance, reliability, and efficiency of storage systems supporting large-scale AI/ML workloads\nTroubleshoot complex storage and data path issues across hardware and software layers\nOptimize storage performance to support high-throughput, low-latency AI training and inference workloads\nAutomation & Tooling\nBuild and maintain automation for provisioning, managing, and monitoring storage infrastructure\nDevelop Python-based tools and workflows to reduce manual operational overhead\nImprove lifecycle management of storage clusters, from deployment through maintenance and scaling\nSystems & Operations\nManage and operate Linux-based systems in production, including bare-metal environments\nPartner with infrastructure and data center teams on hardware bring-up, upgrades, and issue resolution\nSupport capacity planning, utilization tracking, and forecasting for storage systems\nLeverage monitoring and telemetry to diagnose issues and improve system performance and reliability\nCross-Functional Collaboration\nWork closely with Infrastructure Engineering, Network Engineering, and Platform teams to integrate storage into the broader platform\nContribute to design discussions around new infrastructure deployments and scaling strategies\nHelp define best practices for operating storage systems in high-performance computing environments\nWhat You'll Need\nRequired Qualifications\n5+ years of experience in infrastructure engineering, systems engineering, or related roles\nHands-on experience operating distributed storage systems (e.g., VAST, Ceph, or similar)\nStrong Linux systems experience in production environments\nProficiency in Python or similar scripting/programming languages for automation\nExperience working with bare-metal infrastructure and hardware-oriented systems\nAbility to debug complex issues across system boundaries (storage, OS, hardware, networking)\nExperience with storage networking protocols (e.g., NFS or similar)\nExperience with capacity planning, monitoring, and performance tuning\nIdeal Experience\nExperience with VAST storage systems in production environments\nExperience operating S3-compatible object storage at scale\nData center operations experience, including working with physical hardware\nFamiliarity with AI/ML or HPC workloads and their storage requirements\nBackground in high-performance or low-latency distributed systems\nFamiliarity with high-performance data transfer technologies (e.g., RDMA, GPU Direct Storage)\nExperience supporting GPU-based workloads or large-scale compute clusters\n\nCompensation\nBenefits and Perks\nWe offer a comprehensive and competitive benefits package designed to support our employees' health, well-being, and long-term success:\nComprehensive Health Coverage: Medical, dental, and vision coverage for employees and eligible dependents.\nMeaningful Equity: RSUs that give employees a stake in the company's long-term success.\nRetirement Savings: 401(k) matching (U.S.) and pension contributions (U.K.).\nFlexible Time Off: Unlimited PTO, company holidays, and floating holidays to support work-life balance.\nCompany-Wide Winter Break: Two weeks of company closure each winter to disconnect and recharge.\nPaid Parental & Family Leave: Paid leave to support you and your family through life's important moments.\nProfessional Development: Annual learning and development allowance to support your professional growth.\nWellness Benefits: Wellness and work-from-home stipends to support your physical and mental well-being.\nSabbatical Program: Four weeks of paid sabbatical leave after four years of service.\nFlexible Work: Flexible schedules and a hybrid work model for our office-based teams.\nIn-Office Meals: Complimentary meals at our office hubs.\nBenefits may vary by location, team, and role.\n\nAt Lightning AI, we are committed to fostering an inclusive and diverse workplace. We believe that diverse teams drive innovation and create better products. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity, national origin, age, disability, veteran status, or any other protected characteristic. We are dedicated to building a culture where everyone can thrive and contribute to their fullest potential.","company":"Lightningai","rawCompany":"lightningai","city":"Millbrae","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-08-04T16:17:58.501Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"11-3021.00","title":"Computer and Information Systems Managers","slug":"computer-and-information-systems-managers"},{"code":"15-1221.00","title":"Computer and Information Research Scientists","slug":"computer-and-information-research-scientists"}],"industries":[{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"334112","title":"Computer Storage Device Manufacturing","slug":"computer-storage-device-manufacturing"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Infrastructure Engineer (Storage)","description":"Who We Are\nLightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems—designed to take ideas from research to production with less friction.\nThrough our merger with Voltage Park, a neocloud and AI Factory, Lightning AI combines developer-first software with cost-efficient, large-scale compute. Teams get the tools they need for experimentation, training, and production inference, with security, observability, and control built in.\nWe serve solo researchers, startups, and large enterprises. Lightning AI operates globally with offices in New York City, San Francisco, Seattle, and London, and is backed by Coatue, Index Ventures, Bain Capital Ventures, and Firstminute.\nThe Way We Work\nThe people who thrive here are builders who move fast, communicate openly, take ownership, and continuously improve themselves, their teams, and our company. Here's what that looks like in practice:\nMove with Urgency: We move quickly, make thoughtful decisions, and keep momentum. We value action over perfection and learn by shipping.\nTake Ownership: We own outcomes, not just our individual work. We make decisions that move the company forward and follow through.\nCommunicate Openly: We communicate directly, seek to understand, and create clarity for others. Honest conversations help us move faster together.\nBuild Great Teams: We lead by example, empower others, and create healthy teams where people can do their best work.\nRaise the Bar: We're always improving ourselves. We learn from feedback, consistently challenge ourselves to grow, and focus on the work that matters most.\nThink Long-Term: We design for what's next. We create scalable systems, simplify complexity, and use AI and automation to amplify our impact.\nWhat We're Looking For\nLightning AI is seeking a Storage Infrastructure Engineer to join our Infrastructure Engineering team.\nIn this role, you will focus on building and operating the storage systems that power large-scale AI/ML training, inference, and HPC workloads. You will work at the intersection of software, hardware, and operations—developing automation, improving reliability, and scaling distributed storage systems across our bare-metal infrastructure.\nYou will help own the data plane of our storage infrastructure, supporting high-throughput, low-latency data access for some of the most demanding AI workloads. You'll play a key role in managing and evolving our storage stack (including VAST and S3-compatible systems like Ceph), ensuring performance, reliability, and efficiency at scale.\nThis role is based in one of our hubs (NYC, SF, Seattle, or London), with a minimum of 2 in-office days per week and occasional team and company offsites. We are not able to provide visa sponsorship for this position at this time.\nWhat You'll Do\nStorage Systems & Infrastructure\nOperate and scale distributed storage systems, including VAST and S3-compatible object storage (e.g., Ceph)\nImprove performance, reliability, and efficiency of storage systems supporting large-scale AI/ML workloads\nTroubleshoot complex storage and data path issues across hardware and software layers\nOptimize storage performance to support high-throughput, low-latency AI training and inference workloads\nAutomation & Tooling\nBuild and maintain automation for provisioning, managing, and monitoring storage infrastructure\nDevelop Python-based tools and workflows to reduce manual operational overhead\nImprove lifecycle management of storage clusters, from deployment through maintenance and scaling\nSystems & Operations\nManage and operate Linux-based systems in production, including bare-metal environments\nPartner with infrastructure and data center teams on hardware bring-up, upgrades, and issue resolution\nSupport capacity planning, utilization tracking, and forecasting for storage systems\nLeverage monitoring and telemetry to diagnose issues and improve system performance and reliability\nCross-Functional Collaboration\nWork closely with Infrastructure Engineering, Network Engineering, and Platform teams to integrate storage into the broader platform\nContribute to design discussions around new infrastructure deployments and scaling strategies\nHelp define best practices for operating storage systems in high-performance computing environments\nWhat You'll Need\nRequired Qualifications\n5+ years of experience in infrastructure engineering, systems engineering, or related roles\nHands-on experience operating distributed storage systems (e.g., VAST, Ceph, or similar)\nStrong Linux systems experience in production environments\nProficiency in Python or similar scripting/programming languages for automation\nExperience working with bare-metal infrastructure and hardware-oriented systems\nAbility to debug complex issues across system boundaries (storage, OS, hardware, networking)\nExperience with storage networking protocols (e.g., NFS or similar)\nExperience with capacity planning, monitoring, and performance tuning\nIdeal Experience\nExperience with VAST storage systems in production environments\nExperience operating S3-compatible object storage at scale\nData center operations experience, including working with physical hardware\nFamiliarity with AI/ML or HPC workloads and their storage requirements\nBackground in high-performance or low-latency distributed systems\nFamiliarity with high-performance data transfer technologies (e.g., RDMA, GPU Direct Storage)\nExperience supporting GPU-based workloads or large-scale compute clusters\n\nCompensation\nBenefits and Perks\nWe offer a comprehensive and competitive benefits package designed to support our employees' health, well-being, and long-term success:\nComprehensive Health Coverage: Medical, dental, and vision coverage for employees and eligible dependents.\nMeaningful Equity: RSUs that give employees a stake in the company's long-term success.\nRetirement Savings: 401(k) matching (U.S.) and pension contributions (U.K.).\nFlexible Time Off: Unlimited PTO, company holidays, and floating holidays to support work-life balance.\nCompany-Wide Winter Break: Two weeks of company closure each winter to disconnect and recharge.\nPaid Parental & Family Leave: Paid leave to support you and your family through life's important moments.\nProfessional Development: Annual learning and development allowance to support your professional growth.\nWellness Benefits: Wellness and work-from-home stipends to support your physical and mental well-being.\nSabbatical Program: Four weeks of paid sabbatical leave after four years of service.\nFlexible Work: Flexible schedules and a hybrid work model for our office-based teams.\nIn-Office Meals: Complimentary meals at our office hubs.\nBenefits may vary by location, team, and role.\n\nAt Lightning AI, we are committed to fostering an inclusive and diverse workplace. We believe that diverse teams drive innovation and create better products. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity, national origin, age, disability, veteran status, or any other protected characteristic. We are dedicated to building a culture where everyone can thrive and contribute to their fullest potential.","datePosted":"2026-08-04T16:17:58.501Z","dateModified":"2026-08-04T16:17:58.501Z","hiringOrganization":{"@type":"Organization","name":"Lightningai","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Millbrae","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"6a3cbb4458bf1e3f2bd51a88"},"url":"https://jobsearcher.com/jobs/6a3cbb4458bf1e3f2bd51a88"}}