JOBSEARCHER

Senior Systems Software Engineer, Observability and Telemetry Platform

NVIDIARoanoke, VAL6 LeadSeptember 15th, 2026
Overview In this Senior role you will design, build, and sustain a large-scale Observability and Telemetry platform that ensures high reliability and uptime for GPU cloud services. You will work across design, development, and operations to improve performance, latency, and capacity while enabling developers to safely evolve the system. You will collaborate with cross-functional teams to automate and optimize production systems, and participate in blameless postmortems and proactive reliability improvements. This role offers impact through shaping critical infrastructure that underpins NVIDIA’s cloud services and developer experience. Compensation / Benefitsequitybenefitsremote work option ResponsibilitiesDesign, implement and support operational and reliability aspects of a large-scale Observability & Telemetry collection platform with focus on performance at scale, real-time monitoring, logging and alertingEngage in and improve the full lifecycle of services—from inception and design through deployment, operation and refinementSupport services pre-launch via system design consulting, tooling, platforms, capacity management and launch reviewsMaintain live services by measuring availability, latency and overall system healthScale systems sustainably through automation and reliability-driven changesPractice sustainable incident response and blameless postmortemsParticipate in an on-call rotation to support production systems Key requirementsBS degree in Computer Science or related field, or equivalent experience5+ years of experience with infrastructure automation and distributed systems design5+ years delivering foundational infrastructure and observability platformsExperience in Python, Go, Perl or RubyProficiency in Linux, networking and containersExperience with Kubernetes, OpenStack and Docker in private/public cloud environmentsExperience running observability tools like Grafana, OpenTelemetry, Prometheusstrong communicationownership and drivesystematic problem-solvinginfrastructure automationdistributed systems designobservability platforms
Senior Systems Software Engineer, Observability and Telemetry Platform at NVIDIA | JobSearcher