{"schemaVersion":"jobsearcher.job.v1","id":"eaccc6f4571cc2cb52fb8b7b","url":"https://jobsearcher.com/jobs/eaccc6f4571cc2cb52fb8b7b","canonicalUrl":"https://jobsearcher.com/jobs/eaccc6f4571cc2cb52fb8b7b","title":"Infrastructure Engineer (Observability)","description":"Who We Are\nLightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems—designed to take ideas from research to production with less friction.\nThrough our merger with Voltage Park, a neocloud and AI Factory, Lightning AI combines developer-first software with cost-efficient, large-scale compute. Teams get the tools they need for experimentation, training, and production inference, with security, observability, and control built in.\nWe serve solo researchers, startups, and large enterprises. Lightning AI operates globally with offices in New York City, San Francisco, Seattle, and London, and is backed by Coatue, Index Ventures, Bain Capital Ventures, and Firstminute.\nThe Way We Work\nThe people who thrive here are builders who move fast, communicate openly, take ownership, and continuously improve themselves, their teams, and our company. Here's what that looks like in practice:\nMove with Urgency: We move quickly, make thoughtful decisions, and keep momentum. We value action over perfection and learn by shipping.\nTake Ownership: We own outcomes, not just our individual work. We make decisions that move the company forward and follow through.\nCommunicate Openly: We communicate directly, seek to understand, and create clarity for others. Honest conversations help us move faster together.\nBuild Great Teams: We lead by example, empower others, and create healthy teams where people can do their best work.\nRaise the Bar: We're always improving ourselves. We learn from feedback, consistently challenge ourselves to grow, and focus on the work that matters most.\nThink Long-Term: We design for what's next. We create scalable systems, simplify complexity, and use AI and automation to amplify our impact.\nWhat We're Looking For\nLightning AI is seeking an Observability Infrastructure Engineer to join our Infrastructure Engineering team.\nIn this role, you will own and evolve observability systems across large-scale, GPU-enabled bare-metal infrastructure. You'll operate at the intersection of infrastructure, data, and product, building platforms for metrics, logs, traces, and alerting that power both internal operations and customer-facing visibility.\nYou will play a key role in productizing observability, enabling scalable, multi-tenant monitoring experiences while keeping pace with rapid infrastructure buildouts. This includes designing telemetry pipelines, improving signal quality, and delivering actionable insights that ensure reliability and transparency across our platform.\nThis role is based in one of our hubs (NYC, SF, Seattle, or London), with a minimum of 2 in-office days per week and occasional team and company offsites. We are not able to provide visa sponsorship for this position at this time.\nWhat You'll Do\nObservability Platform & Productization\nOwn and evolve a scalable observability platform spanning metrics, logs, traces, and events\nDrive the productization of observability capabilities for both internal teams and external customers\nDesign multi-tenant observability systems with scoped access, RBAC, and customer-facing visibility\nContinuously improve observability systems to keep pace with rapid infrastructure buildouts\nTelemetry & Data Pipelines\nDesign and operate telemetry pipelines ingesting data from GPUs, CPUs, networking (Ethernet & InfiniBand), containers, APIs, and BMC/Redfish\nBuild systems to correlate signals across infrastructure layers to enable faster debugging and root cause analysis\nImplement streaming and real-time data pipelines using tools such as Kafka, OTEL, Promtail, or similar\nAlerting, Reliability & Insights\nDesign and implement noise-resistant alerting systems to improve signal quality and reduce operational load\nCreate dashboards and alerting for InfraOps, Engineering, and Customer Success teams\nBuild automated insights and enable proactive detection, forecasting, and system health visibility at scale\nSystems & Infrastructure Engineering\nContribute to broader infrastructure engineering projects beyond observability\nPartner with infrastructure and platform teams to embed observability into core systems and workflows\nSupport large-scale, distributed systems across compute, networking, and storage environments\nCross-Functional Collaboration\nWork closely with customer-facing teams to deliver external observability experiences\nCollaborate with engineering, operations, and support teams to improve system transparency and reliability\nHelp define best practices for observability across the organization\nWhat You'll Need\nRequired Qualifications\n5+ years of experience in infrastructure engineering, SRE, or observability-focused roles\nStrong experience with monitoring systems such as Prometheus, Grafana, ELK, or VictoriaMetrics\nExperience building and operating observability platforms at scale\nProficiency in Python, Go, or bash for automation and data integration\nFamiliarity with containerized environments and Kubernetes observability\nExperience with streaming telemetry pipelines (Kafka, OTEL, Promtail, or equivalent)\nExperience with multi-tenant monitoring architectures\nStrong written and verbal communication skills\nIdeal Experience\nExperience with GPU observability, particularly NVIDIA DCGM\nExperience monitoring large-scale GPU or HPC clusters\nFamiliarity with InfiniBand fabric observability\nExperience building customer-facing or productized infrastructure systems\nExperience with correlation engines, RCA workflows, or predictive alerting systems\nBroad exposure to infrastructure domains including networking, storage, and provisioning\nCompensation\nBenefits and Perks\nWe offer a comprehensive and competitive benefits package designed to support our employees' health, well-being, and long-term success:\nComprehensive Health Coverage: Medical, dental, and vision coverage for employees and eligible dependents.\nMeaningful Equity: RSUs that give employees a stake in the company's long-term success.\nRetirement Savings: 401(k) matching (U.S.) and pension contributions (U.K.).\nFlexible Time Off: Unlimited PTO, company holidays, and floating holidays to support work-life balance.\nCompany-Wide Winter Break: Two weeks of company closure each winter to disconnect and recharge.\nPaid Parental & Family Leave: Paid leave to support you and your family through life's important moments.\nProfessional Development: Annual learning and development allowance to support your professional growth.\nWellness Benefits: Wellness and work-from-home stipends to support your physical and mental well-being.\nSabbatical Program: Four weeks of paid sabbatical leave after four years of service.\nFlexible Work: Flexible schedules and a hybrid work model for our office-based teams.\nIn-Office Meals: Complimentary meals at our office hubs.\nBenefits may vary by location, team, and role.\n\nAt Lightning AI, we are committed to fostering an inclusive and diverse workplace. We believe that diverse teams drive innovation and create better products. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity, national origin, age, disability, veteran status, or any other protected characteristic. We are dedicated to building a culture where everyone can thrive and contribute to their fullest potential.","company":"Lightningai","rawCompany":"lightningai","city":"Millbrae","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-08-03T22:52:42.880Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1244.00","title":"Network and Computer Systems Administrators","slug":"network-and-computer-systems-administrators"},{"code":"11-3021.00","title":"Computer and Information Systems Managers","slug":"computer-and-information-systems-managers"}],"industries":[{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Infrastructure Engineer (Observability)","description":"Who We Are\nLightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems—designed to take ideas from research to production with less friction.\nThrough our merger with Voltage Park, a neocloud and AI Factory, Lightning AI combines developer-first software with cost-efficient, large-scale compute. Teams get the tools they need for experimentation, training, and production inference, with security, observability, and control built in.\nWe serve solo researchers, startups, and large enterprises. Lightning AI operates globally with offices in New York City, San Francisco, Seattle, and London, and is backed by Coatue, Index Ventures, Bain Capital Ventures, and Firstminute.\nThe Way We Work\nThe people who thrive here are builders who move fast, communicate openly, take ownership, and continuously improve themselves, their teams, and our company. Here's what that looks like in practice:\nMove with Urgency: We move quickly, make thoughtful decisions, and keep momentum. We value action over perfection and learn by shipping.\nTake Ownership: We own outcomes, not just our individual work. We make decisions that move the company forward and follow through.\nCommunicate Openly: We communicate directly, seek to understand, and create clarity for others. Honest conversations help us move faster together.\nBuild Great Teams: We lead by example, empower others, and create healthy teams where people can do their best work.\nRaise the Bar: We're always improving ourselves. We learn from feedback, consistently challenge ourselves to grow, and focus on the work that matters most.\nThink Long-Term: We design for what's next. We create scalable systems, simplify complexity, and use AI and automation to amplify our impact.\nWhat We're Looking For\nLightning AI is seeking an Observability Infrastructure Engineer to join our Infrastructure Engineering team.\nIn this role, you will own and evolve observability systems across large-scale, GPU-enabled bare-metal infrastructure. You'll operate at the intersection of infrastructure, data, and product, building platforms for metrics, logs, traces, and alerting that power both internal operations and customer-facing visibility.\nYou will play a key role in productizing observability, enabling scalable, multi-tenant monitoring experiences while keeping pace with rapid infrastructure buildouts. This includes designing telemetry pipelines, improving signal quality, and delivering actionable insights that ensure reliability and transparency across our platform.\nThis role is based in one of our hubs (NYC, SF, Seattle, or London), with a minimum of 2 in-office days per week and occasional team and company offsites. We are not able to provide visa sponsorship for this position at this time.\nWhat You'll Do\nObservability Platform & Productization\nOwn and evolve a scalable observability platform spanning metrics, logs, traces, and events\nDrive the productization of observability capabilities for both internal teams and external customers\nDesign multi-tenant observability systems with scoped access, RBAC, and customer-facing visibility\nContinuously improve observability systems to keep pace with rapid infrastructure buildouts\nTelemetry & Data Pipelines\nDesign and operate telemetry pipelines ingesting data from GPUs, CPUs, networking (Ethernet & InfiniBand), containers, APIs, and BMC/Redfish\nBuild systems to correlate signals across infrastructure layers to enable faster debugging and root cause analysis\nImplement streaming and real-time data pipelines using tools such as Kafka, OTEL, Promtail, or similar\nAlerting, Reliability & Insights\nDesign and implement noise-resistant alerting systems to improve signal quality and reduce operational load\nCreate dashboards and alerting for InfraOps, Engineering, and Customer Success teams\nBuild automated insights and enable proactive detection, forecasting, and system health visibility at scale\nSystems & Infrastructure Engineering\nContribute to broader infrastructure engineering projects beyond observability\nPartner with infrastructure and platform teams to embed observability into core systems and workflows\nSupport large-scale, distributed systems across compute, networking, and storage environments\nCross-Functional Collaboration\nWork closely with customer-facing teams to deliver external observability experiences\nCollaborate with engineering, operations, and support teams to improve system transparency and reliability\nHelp define best practices for observability across the organization\nWhat You'll Need\nRequired Qualifications\n5+ years of experience in infrastructure engineering, SRE, or observability-focused roles\nStrong experience with monitoring systems such as Prometheus, Grafana, ELK, or VictoriaMetrics\nExperience building and operating observability platforms at scale\nProficiency in Python, Go, or bash for automation and data integration\nFamiliarity with containerized environments and Kubernetes observability\nExperience with streaming telemetry pipelines (Kafka, OTEL, Promtail, or equivalent)\nExperience with multi-tenant monitoring architectures\nStrong written and verbal communication skills\nIdeal Experience\nExperience with GPU observability, particularly NVIDIA DCGM\nExperience monitoring large-scale GPU or HPC clusters\nFamiliarity with InfiniBand fabric observability\nExperience building customer-facing or productized infrastructure systems\nExperience with correlation engines, RCA workflows, or predictive alerting systems\nBroad exposure to infrastructure domains including networking, storage, and provisioning\nCompensation\nBenefits and Perks\nWe offer a comprehensive and competitive benefits package designed to support our employees' health, well-being, and long-term success:\nComprehensive Health Coverage: Medical, dental, and vision coverage for employees and eligible dependents.\nMeaningful Equity: RSUs that give employees a stake in the company's long-term success.\nRetirement Savings: 401(k) matching (U.S.) and pension contributions (U.K.).\nFlexible Time Off: Unlimited PTO, company holidays, and floating holidays to support work-life balance.\nCompany-Wide Winter Break: Two weeks of company closure each winter to disconnect and recharge.\nPaid Parental & Family Leave: Paid leave to support you and your family through life's important moments.\nProfessional Development: Annual learning and development allowance to support your professional growth.\nWellness Benefits: Wellness and work-from-home stipends to support your physical and mental well-being.\nSabbatical Program: Four weeks of paid sabbatical leave after four years of service.\nFlexible Work: Flexible schedules and a hybrid work model for our office-based teams.\nIn-Office Meals: Complimentary meals at our office hubs.\nBenefits may vary by location, team, and role.\n\nAt Lightning AI, we are committed to fostering an inclusive and diverse workplace. We believe that diverse teams drive innovation and create better products. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity, national origin, age, disability, veteran status, or any other protected characteristic. We are dedicated to building a culture where everyone can thrive and contribute to their fullest potential.","datePosted":"2026-08-03T22:52:42.880Z","dateModified":"2026-08-03T22:52:42.880Z","hiringOrganization":{"@type":"Organization","name":"Lightningai","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Millbrae","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"eaccc6f4571cc2cb52fb8b7b"},"url":"https://jobsearcher.com/jobs/eaccc6f4571cc2cb52fb8b7b"}}