Systems Validation Engineer – Data Center GPU
Overview
As a Systems Validation Engineer in AMD’s Datacenter and GPU Accelerated Processing division, you drive SOC and system-level validation for data center GPU products across pre-silicon to post-silicon phases. You will own validation strategy, execute tests, and collaborate with cross-functional teams to ensure high-quality AI and HPC products. This role offers impact through shaping validation methodologies, improving coverage, and enabling scalable, reliable platforms. You will work in a hybrid Markham, ON environment helping AMD advance data center technology.
ResponsibilitiesPlan and execute validation content with domain experts across accelerators and racksOwn validation planning and execution through emulation, bring-up, post-silicon enablement, regression, and sustaining phasesLead triage, debug, root-cause analysis, and issue closure across hardware, firmware, platform, diagnostics, and system-level interactionsPartner with architecture, design, firmware, diagnostics, platform, automation, customer engineering, and program teams to align scope, risks, dependencies, and exit criteriaBuild and improve automation, validation content, debug-analysis scripts, and regression checks to improve coverage, efficiency, and repeatabilityReview RTL, specifications, firmware interfaces, register maps, and validation collateral to identify coverage gaps and improve validation effectivenessSupport lab-based and system-level debug using platform tools, firmware logs, register-level analysis, and standard measurement equipmentSupport customer and sustaining issues by providing SOC and system test technical expertise, debug guidance, and issue-resolution supportMentor engineers and help improve team validation methodology, debug practices, and product quality
Key requirementsStrong experience in semiconductor validation, silicon bring-up, SoC validation, firmware validation, or hardware/software integrationHands-on experience with system and rack level validation in data center environmentsExperience developing validation plans, test content, automation, and debug methodologies for complex SoC or GPU productsStrong debug skills across hardware, firmware, software, diagnostics, and platform interactionsExperience with emulation, simulation, or FPGA/prototyping environmentsProficiency in Python and C/C++ for automation, test infrastructure, or debug-analysis toolsFamiliarity with Linux, lab bring-up, register-level debug, firmware logs, and platform debug toolsAbility to execute independently, manage priorities, influence cross-functional teams, and drive technical issues to closurestrong communication with technical and program stakeholdersmentoring and leadershipindependent planning and prioritizationSoC validationsystem/rack level validationdebug and triage across hardware/firmware/software