{"schemaVersion":"jobsearcher.job.v1","id":"1c819662fb219143b685e52a","url":"https://jobsearcher.com/jobs/1c819662fb219143b685e52a","canonicalUrl":"https://jobsearcher.com/jobs/1c819662fb219143b685e52a","title":"Sr. Data Center GPU Validation and Debug Engineer","description":"Overview:\n\nADVANCE YOUR CAREER. ADVANCE THE WORLD.\n\nAt AMD, we believe technology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future.\n\nWhether you’re designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger — technology that moves the world forward. Join us and, together, we’ll advance your career.\n\nResponsibilities:\n\nTHE ROLE\n\nWe are seeking an experienced engineer to validate and debug data center GPU products across the complete software and hardware stack, spanning workload behavior and performance qualification as well as platform health and root-cause isolation. The role demands strong technical judgment, disciplined failure isolation, and the ability to trace unexpected behavior from application to hardware, working across firmware, kernel, compiler, library, and framework teams to drive issues to resolution.\n\nKEY RESPONSIBILITIES:\n\nDebug GPU failures including hangs, crashes, memory faults, firmware errors, and performance regressions across the full software and hardware stack.\nReproduce failures and reduce them to minimal, actionable test cases.\nTriage system-level interactions across GPU, CPU, memory, networking, power, and platform topology.\nValidate AI, HPC, and communication workloads across GPU platforms, firmware, and software releases.\nCharacterize performance and identify compute, memory, and communication bottlenecks.\nAnalyze single- and multi-GPU behavior across workload configurations, topology, NUMA, power, and thermal conditions.\nEstablish benchmarks, baselines, and regression-detection methods; correlate microbenchmark results with real workload behavior.\nBuild automated diagnostics, stress tests, performance suites, and regression infrastructure.\nSupport customer-representative validation and escalation reproduction.\n\nREQUIRED QUALIFICATIONS:\n\nStrong Linux and systems-engineering experience, with C/C++, Python, and shell automation.\nExperience with data center GPUs, accelerators, or comparable high-performance systems.\nUnderstanding of GPU architecture, memory hierarchy, parallel execution, communication, and synchronization.\nExperience debugging across software, firmware, kernel, and hardware boundaries.\nAbility to analyze logs, traces, hardware telemetry, and performance counters.\n\nPREFERRED EXPERIENCE:\n\nStrong knowledge of Linux kernel drivers, PCIe, firmware interaction, and memory management.\nFamiliarity with GPU programming models and runtimes such as ROCm/HIP, RCCL, or CUDA equivalents.\nExperience with AI training or inference frameworks such as PyTorch, vLLM, or SGLang.\nExperience benchmarking AI inference, training, or HPC workloads, including roofline analysis and model-level profiling.\nExperience with profiling tools, statistical analysis, and regression detection.\nFamiliarity with multi-GPU topology, PCIe/fabric interconnects, NUMA, GPU firmware, RAS, and rack-scale platforms.\nExperience with system bring-up, qualification, or production data center operations.\nFamiliarity with CI systems, automated validation infrastructure, and fleet-scale testing.\n\n ACADEMIC CREDENTIALS:\n\nBachelor’s or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent\n\nThis role is not eligible for visa sponsorship.\n\n#LI-G11\n\n#LI-HYBRID\n\nQualifications:\n\nBenefits offered are described: AMD benefits at a glance.\n\nAMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.\n\nAMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.\n\nThis posting is for an existing vacancy.","company":"Advanced Micro Devices","rawCompany":"advanced micro devices","city":"Austin","state":"TX","isRemote":false,"isActive":true,"createdAt":"2026-09-14T10:25:13.747Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"17-2061.00","title":"Computer Hardware Engineers","slug":"computer-hardware-engineers"},{"code":"17-2112.02","title":"Validation Engineers","slug":"validation-engineers"}],"industries":[{"code":"334111","title":"Electronic Computer Manufacturing","slug":"electronic-computer-manufacturing"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Sr. Data Center GPU Validation and Debug Engineer","description":"Overview:\n\nADVANCE YOUR CAREER. ADVANCE THE WORLD.\n\nAt AMD, we believe technology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future.\n\nWhether you’re designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger — technology that moves the world forward. Join us and, together, we’ll advance your career.\n\nResponsibilities:\n\nTHE ROLE\n\nWe are seeking an experienced engineer to validate and debug data center GPU products across the complete software and hardware stack, spanning workload behavior and performance qualification as well as platform health and root-cause isolation. The role demands strong technical judgment, disciplined failure isolation, and the ability to trace unexpected behavior from application to hardware, working across firmware, kernel, compiler, library, and framework teams to drive issues to resolution.\n\nKEY RESPONSIBILITIES:\n\nDebug GPU failures including hangs, crashes, memory faults, firmware errors, and performance regressions across the full software and hardware stack.\nReproduce failures and reduce them to minimal, actionable test cases.\nTriage system-level interactions across GPU, CPU, memory, networking, power, and platform topology.\nValidate AI, HPC, and communication workloads across GPU platforms, firmware, and software releases.\nCharacterize performance and identify compute, memory, and communication bottlenecks.\nAnalyze single- and multi-GPU behavior across workload configurations, topology, NUMA, power, and thermal conditions.\nEstablish benchmarks, baselines, and regression-detection methods; correlate microbenchmark results with real workload behavior.\nBuild automated diagnostics, stress tests, performance suites, and regression infrastructure.\nSupport customer-representative validation and escalation reproduction.\n\nREQUIRED QUALIFICATIONS:\n\nStrong Linux and systems-engineering experience, with C/C++, Python, and shell automation.\nExperience with data center GPUs, accelerators, or comparable high-performance systems.\nUnderstanding of GPU architecture, memory hierarchy, parallel execution, communication, and synchronization.\nExperience debugging across software, firmware, kernel, and hardware boundaries.\nAbility to analyze logs, traces, hardware telemetry, and performance counters.\n\nPREFERRED EXPERIENCE:\n\nStrong knowledge of Linux kernel drivers, PCIe, firmware interaction, and memory management.\nFamiliarity with GPU programming models and runtimes such as ROCm/HIP, RCCL, or CUDA equivalents.\nExperience with AI training or inference frameworks such as PyTorch, vLLM, or SGLang.\nExperience benchmarking AI inference, training, or HPC workloads, including roofline analysis and model-level profiling.\nExperience with profiling tools, statistical analysis, and regression detection.\nFamiliarity with multi-GPU topology, PCIe/fabric interconnects, NUMA, GPU firmware, RAS, and rack-scale platforms.\nExperience with system bring-up, qualification, or production data center operations.\nFamiliarity with CI systems, automated validation infrastructure, and fleet-scale testing.\n\n ACADEMIC CREDENTIALS:\n\nBachelor’s or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent\n\nThis role is not eligible for visa sponsorship.\n\n#LI-G11\n\n#LI-HYBRID\n\nQualifications:\n\nBenefits offered are described: AMD benefits at a glance.\n\nAMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.\n\nAMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.\n\nThis posting is for an existing vacancy.","datePosted":"2026-09-14T10:25:13.747Z","dateModified":"2026-09-14T10:25:13.747Z","hiringOrganization":{"@type":"Organization","name":"Advanced Micro Devices","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Austin","addressRegion":"TX","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"1c819662fb219143b685e52a"},"url":"https://jobsearcher.com/jobs/1c819662fb219143b685e52a"}}