{"schemaVersion":"jobsearcher.job.v1","id":"e62f9f88bb8e34058f3fac66","url":"https://jobsearcher.com/jobs/e62f9f88bb8e34058f3fac66","canonicalUrl":"https://jobsearcher.com/jobs/e62f9f88bb8e34058f3fac66","title":"AI Cluster Validation Engineer","description":"Overview:\nWHAT YOU DO AT AMD CHANGES EVERYTHING\nAt AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career.\nResponsibilities:\nTHE ROLE:\nWe are looking for a Senior Engineer to drive validation of next-generation AI cluster solutions. In this role, you will be at the forefront of optimizing GPU cluster, working across the full system stack to ensure our solutions meet the demanding requirements of large-scale AI workloads. The primary focus of this role is on the RDMA fabric at the heart of these systems. Understanding data flows between GPUs, NICs, and the cluster network to optimize performance at scale.\n\nThe ideal candidate brings strong expertise in GPU architectures, parallel computing, and hands-on experience validating complex, high-performance systems.\n\nThis is a high-impact engineering role for someone who thrives at the intersection of hardware, networking, and systems software and who wants to shape the performance and reliability of AI infrastructure at scale.\n\nTHE PERSON:\nYou are a seasoned engineer who thrives on hands-on problem-solving. Equally comfortable shaping long-term strategy and diving deep into complex hardware, firmware, and driver issues. You communicate clearly across teams, take full ownership of your work, and bring the drive and work ethic to see tough problems through to resolution.\nOur team is built on a culture of technical innovation and continuous career development, where your impact will be felt across performance, automation, and development.\nKEY RESPONSIBILITIES:\nScalability Testing: Evaluate GPU cluster scalability through rigorous testing across diverse workloads, cluster sizes, configurations, and networking technologies including RoCE.\nBenchmarking & Performance Profiling: Develop and execute comprehensive benchmarking strategies to establish performance baselines, identify bottlenecks, and generate actionable insights for improvement.\nPerformance Tuning: Implement optimization strategies across protocol enhancements, load balancing, and parallel processing to drive measurable gains in RDMA throughput, latency, and collective communications.\nNIC & Cluster Optimization: Collaborate with hardware and software teams to enhance GPU cluster performance end-to-end, with a focus on the NIC-to-network data path.\nCross-functional Collaboration: Partner closely with hardware engineers, software developers, and system architects to integrate performance improvements into cluster architecture.\nDocumentation: Produce clear, detailed documentation of performance analysis, tuning methodologies, and outcomes for both internal teams and senior stakeholders.\nContinuous Learning: Stay current with advances in GPU architectures, parallel processing, and emerging networking technologies to inform ongoing improvement efforts.\nPREFERRED EXPERIENCE:\nProven experience in optimizing the performance of GPU clusters.\nRDMA network configuration, troubleshooting and performance tuning.\nStrong understanding of GPU architectures, parallel computing concepts, and network protocols.\nProficiency in scripting languages (e.g., Python, Bash) for automation and performance analysis.\nExperience with system level performance analysis tools and methodologies for GPU clusters.\nAnalytical mindset with excellent problem-solving and debug skills.\nExcellent communication and collaboration skills for effective teamwork.\nACADEMIC CREDENTIALS:\nBachelors or Masters degree in electrical or computer engineering preferred\n\nLOCATION: Austin, TX, Santa Clara CA, Seattle WA, Secaucus NJ\n\nThis role is not eligible for visa sponsorship.\n\n#LI-SC3\nQualifications:\nBenefits offered are described: AMD benefits at a glance.\n\nAMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.\n\nAMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.\n\nThis posting is for an existing vacancy.","company":"Advanced Micro Devices","rawCompany":"advanced micro devices","city":"Austin","state":"TX","isRemote":false,"isActive":false,"createdAt":"2026-08-03T19:35:17.127Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"17-2112.02","title":"Validation Engineers","slug":"validation-engineers"},{"code":"17-2061.00","title":"Computer Hardware Engineers","slug":"computer-hardware-engineers"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"334111","title":"Electronic Computer Manufacturing","slug":"electronic-computer-manufacturing"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"AI Cluster Validation Engineer","description":"Overview:\nWHAT YOU DO AT AMD CHANGES EVERYTHING\nAt AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career.\nResponsibilities:\nTHE ROLE:\nWe are looking for a Senior Engineer to drive validation of next-generation AI cluster solutions. In this role, you will be at the forefront of optimizing GPU cluster, working across the full system stack to ensure our solutions meet the demanding requirements of large-scale AI workloads. The primary focus of this role is on the RDMA fabric at the heart of these systems. Understanding data flows between GPUs, NICs, and the cluster network to optimize performance at scale.\n\nThe ideal candidate brings strong expertise in GPU architectures, parallel computing, and hands-on experience validating complex, high-performance systems.\n\nThis is a high-impact engineering role for someone who thrives at the intersection of hardware, networking, and systems software and who wants to shape the performance and reliability of AI infrastructure at scale.\n\nTHE PERSON:\nYou are a seasoned engineer who thrives on hands-on problem-solving. Equally comfortable shaping long-term strategy and diving deep into complex hardware, firmware, and driver issues. You communicate clearly across teams, take full ownership of your work, and bring the drive and work ethic to see tough problems through to resolution.\nOur team is built on a culture of technical innovation and continuous career development, where your impact will be felt across performance, automation, and development.\nKEY RESPONSIBILITIES:\nScalability Testing: Evaluate GPU cluster scalability through rigorous testing across diverse workloads, cluster sizes, configurations, and networking technologies including RoCE.\nBenchmarking & Performance Profiling: Develop and execute comprehensive benchmarking strategies to establish performance baselines, identify bottlenecks, and generate actionable insights for improvement.\nPerformance Tuning: Implement optimization strategies across protocol enhancements, load balancing, and parallel processing to drive measurable gains in RDMA throughput, latency, and collective communications.\nNIC & Cluster Optimization: Collaborate with hardware and software teams to enhance GPU cluster performance end-to-end, with a focus on the NIC-to-network data path.\nCross-functional Collaboration: Partner closely with hardware engineers, software developers, and system architects to integrate performance improvements into cluster architecture.\nDocumentation: Produce clear, detailed documentation of performance analysis, tuning methodologies, and outcomes for both internal teams and senior stakeholders.\nContinuous Learning: Stay current with advances in GPU architectures, parallel processing, and emerging networking technologies to inform ongoing improvement efforts.\nPREFERRED EXPERIENCE:\nProven experience in optimizing the performance of GPU clusters.\nRDMA network configuration, troubleshooting and performance tuning.\nStrong understanding of GPU architectures, parallel computing concepts, and network protocols.\nProficiency in scripting languages (e.g., Python, Bash) for automation and performance analysis.\nExperience with system level performance analysis tools and methodologies for GPU clusters.\nAnalytical mindset with excellent problem-solving and debug skills.\nExcellent communication and collaboration skills for effective teamwork.\nACADEMIC CREDENTIALS:\nBachelors or Masters degree in electrical or computer engineering preferred\n\nLOCATION: Austin, TX, Santa Clara CA, Seattle WA, Secaucus NJ\n\nThis role is not eligible for visa sponsorship.\n\n#LI-SC3\nQualifications:\nBenefits offered are described: AMD benefits at a glance.\n\nAMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.\n\nAMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.\n\nThis posting is for an existing vacancy.","datePosted":"2026-08-03T19:35:17.127Z","dateModified":"2026-08-03T19:35:17.127Z","hiringOrganization":{"@type":"Organization","name":"Advanced Micro Devices","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Austin","addressRegion":"TX","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"e62f9f88bb8e34058f3fac66"},"url":"https://jobsearcher.com/jobs/e62f9f88bb8e34058f3fac66"}}