JOBSEARCHER

Sr. Systems Design Engineer - Data Center GPU

AMDChicago, ILL6 LeadSeptember 15th, 2026
Overview As part of AMD’s Data Center GPU team, you will drive post-silicon validation and system-level debugging for SOC IP blocks. You’ll collaborate across design, firmware, software, and kernel teams to validate features, triage failures, and optimize performance. The role emphasizes scalable, automated validation and deep hardware-software understanding to deliver high-quality GPU technology. This high-visibility position offers opportunities to shape IP behavior and fix root causes across multi-die designs while advancing AMD’s mission in HPC and AI. ResponsibilitiesExecute post-silicon validation for SOC-level IP blocks, including test plan development, execution, coverage tracking, and issue reportingDebug hardware and system-level issues during bring-up, validation, and production across IP blocks like DMA engines, interrupt controllers, and data path logicPerform post-silicon debug analysis using scan dump tools and debug reports to investigate hangs and errorsTriager failures by correlating debug data across multiple IP domains with guidance from senior engineersCoordinate test execution with multiple teams to ensure features are validated and optimized on timeCollaborate with design, firmware, driver, software runtime, and kernel teams to understand IP behavior, error propagation, and power management dependenciesDevelop and maintain validation content for error handling, fault injection, and power management scenariosEngage in hardware/software modeling and debugging frameworks to reproduce and root-cause silicon failuresParticipate in cross-team triage and debug efforts, progressively taking ownership of specific IP areas Key requirementsPost-silicon validation or hardware debug experienceProgramming in C/C++, Python, or PerlPost-silicon debug techniques and methodologiesBoard/platform-level debugging experience (bring-up, sequencing, analysis, optimization)Knowledge of SoC architectures, including multi-die or chiplet designsUnderstanding of cache, memory subsystems, DMA/IO architectures, and interconnectsExperience with software runtime environments, kernel drivers, or OS interfaces affecting hardware IPsExposure to RAS concepts (error detection, poison propagation, machine check logging, watchdogs)Familiarity with scan dump analysis or JTAG-based post-silicon debugging toolsStrong analytical, problem-solving, and detail orientationSelf-starter with ability to ramp up on new IP domains via docs and hands-on debuggingStrong analytical thinkingDetail-orientedSelf-motivated and proactiveC/C++Python (or Perl)Post-silicon validation