Senior Systems Software Engineer, Observability and Telemetry Platform
Overview
In this Senior role you will design, build, and sustain a large-scale Observability and Telemetry platform that ensures high reliability and uptime for GPU cloud services. You will work across design, development, and operations to improve performance, latency, and capacity while enabling developers to safely evolve the system. You will collaborate with cross-functional teams to automate and optimize production systems, and participate in blameless postmortems and proactive reliability improvements. This role offers impact through shaping critical infrastructure that underpins NVIDIA’s cloud services and developer experience.
Compensation / Benefitsequitybenefitsremote work option
ResponsibilitiesDesign, implement and support operational and reliability aspects of a large-scale Observability & Telemetry collection platform with focus on performance at scale, real-time monitoring, logging and alertingEngage in and improve the full lifecycle of services—from inception and design through deployment, operation and refinementSupport services pre-launch via system design consulting, tooling, platforms, capacity management and launch reviewsMaintain live services by measuring availability, latency and overall system healthScale systems sustainably through automation and reliability-driven changesPractice sustainable incident response and blameless postmortemsParticipate in an on-call rotation to support production systems
Key requirementsBS degree in Computer Science or related field, or equivalent experience5+ years of experience with infrastructure automation and distributed systems design5+ years delivering foundational infrastructure and observability platformsExperience in Python, Go, Perl or RubyProficiency in Linux, networking and containersExperience with Kubernetes, OpenStack and Docker in private/public cloud environmentsExperience running observability tools like Grafana, OpenTelemetry, Prometheusstrong communicationownership and drivesystematic problem-solvinginfrastructure automationdistributed systems designobservability platforms