JOBSEARCHER

Senior Software Engineer - Distributed Systems Engineer, EDA Infrastructure

NVIDIATacoma, WAL6 LeadSeptember 15th, 2026
Overview In this role you will architect and scale infrastructure that supports EDA workloads across GPU and CPU compute fleets. You will work with cross-functional teams to automate provisioning, configuration, and lifecycle management of large-scale systems. You’ll build monitoring, remediation, and recovery workflows to improve reliability and utilization. The position emphasizes collaboration, ownership, and delivering resilient platforms for critical chip-design workloads. Compensation / Benefitsequitybenefits ResponsibilitiesDesign and build automation platforms for provisioning, configuring, operating, and lifecycle managing large GPU/CPU infrastructureDevelop monitoring, health-management, and remediation systems to improve reliability and utilizationAutomate hardware deployment, OS configuration, firmware/software updates, cluster enrollment, and recovery workflowsBuild services and workflows that integrate with workload schedulers, infrastructure management, and observability platformsAnalyze hardware diagnostics, OS signals, scheduler data, and telemetry to identify failures and restore servicesCollaborate with EDA, infrastructure, networking, storage, and hardware teams to deliver scalable solutionsParticipate in incident response, root-cause analysis, capacity planning, and continuous improvement of production services Key requirements5+ years of software or infrastructure engineering for large-scale production systemsBS in Computer Science, Engineering, Physics, Mathematics, or related field, or equivalent experienceStrong programming in Go or Python with solid data structures, algorithms, testing, and designExperience designing automation for distributed systems and large Linux compute fleetsUnderstanding of performance, security, reliability, fault tolerance, state management, data consistencyExperience with infrastructure automation, deployment, observability, and operational recoveryStrong communication and cross-team collaboration skillsSystematic problem solving with ownership and reducing operational toilstrong communicationcross-team collaborationownership and accountabilityGoPythondistributed systems