{"schemaVersion":"jobsearcher.job.v1","id":"512a5688caf0c8209dfeb9d0","url":"https://jobsearcher.com/jobs/512a5688caf0c8209dfeb9d0","canonicalUrl":"https://jobsearcher.com/jobs/512a5688caf0c8209dfeb9d0","title":"Failure Analysis Engineer","description":"Overview:\nWHAT YOU DO AT AMD CHANGES EVERYTHING\nAt AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career.\nResponsibilities:\nTHE ROLE\nJoin a highly visible engineering team responsible for bringing up, debugging, and improving next-generation server and rack-scale platforms deployed in data center environments. As a Senior Failure Analysis Engineer, you will be among the first engineers to investigate, troubleshoot, and resolve complex system-level issues on new hardware platforms, partnering closely with design, firmware, validation, manufacturing, quality, and infrastructure teams. This role offers the opportunity to work on cutting-edge technologies while driving root cause analysis, platform reliability, and manufacturing readiness across the product lifecycle. You will play a key role in identifying and resolving complex hardware and firmware interactions, supporting platform bring-up activities, improving debug processes, and helping scale organizational knowledge through technical documentation and AI-assisted engineering workflows. This is an ideal opportunity for an engineer who enjoys solving difficult technical problems, working across multiple disciplines, and making a measurable impact on the success of next-generation data center products.\n\nTHE PERSON\nThe ideal candidate is naturally curious, analytical, and motivated by solving complex technical challenges. They thrive in environments where not every answer is immediately available and are comfortable investigating issues across multiple engineering domains to identify root causes and drive resolution. They demonstrate strong communication skills, effectively manage stakeholder expectations during ongoing investigations, and can clearly communicate technical findings to both engineering and cross-functional teams.\nSuccessful candidates are collaborative, adaptable, and persistent problem-solvers who enjoy learning new technologies, working in highly dynamic environments, and continuously improving how engineering teams operate. They embrace new tools and technologies, including AI-assisted workflows, to improve efficiency, accelerate troubleshooting, and scale technical knowledge across the organization.\n\nKEY RESPONSIBILITIES\nPerform system-level and rack-level failure analysis across server and data center platforms, driving issues from initial symptom identification through root cause determination and resolution.\nDebug complex hardware, firmware, and platform-related issues involving CPUs, GPUs, memory, PCIe, networking, power delivery, thermal systems, and firmware interactions.\nAnalyze system logs, platform telemetry, BIOS, BMC, and other diagnostic data to identify failure patterns, isolate root causes, and validate corrective actions.\nReproduce failures in lab environments and utilize appropriate diagnostic tools and methodologies to validate fixes and improve overall platform reliability.\nPartner closely with design, firmware, validation, manufacturing, quality, and infrastructure teams to investigate issues, drive corrective actions, and improve system performance.\nSupport manufacturing and ODM partners by providing technical guidance, failure triage, structured debug processes, and escalation support for complex platform issues.\nDevelop and maintain technical documentation, troubleshooting guides, SOPs, debug methodologies, and knowledge-sharing resources to improve organizational effectiveness.\nDrive root cause analysis activities and contribute to continuous improvements in reliability, serviceability, manufacturability, and operational readiness.\nLeverage AI-assisted tools, automation, and knowledge systems to improve troubleshooting efficiency, accelerate investigations, and scale engineering best practices.\nServe as a technical resource for cross-functional teams by communicating findings, managing stakeholder expectations, and providing clear recommendations based on data-driven analysis.\nContribute expertise in areas such as networking, signal integrity, system architecture, or platform reliability to support the successful deployment of next-generation data center technologies.\n\nPREFERRED EXPERIENCE\nExperience performing system-level, server-level, or rack-level troubleshooting and failure analysis.\nBackground in hardware debugging, root cause analysis, and complex issue resolution.\nFamiliarity with Linux-based environments, platform logs, and diagnostic workflows.\nExperience supporting server, data center, networking, storage, or enterprise hardware platforms.\nExposure to BIOS, BMC, IPMI, firmware interactions, or platform management technologies.\nUnderstanding of networking technologies, signal integrity concepts, power delivery, or high-speed interfaces.\nExperience collaborating with manufacturing, ODM, validation, quality, or design teams.\nAbility to develop technical documentation, SOPs, troubleshooting guides, and knowledge-sharing resources.\nFamiliarity with lab debugging equipment and system diagnostic tools.\nExperience leveraging AI tools, AI agents, automation, or knowledge systems to improve debugging efficiency and technical problem-solving.\nWillingness to occasionally travel in support of manufacturing, validation, or deployment activities.\n\nACADEMIC CREDENTIALS\nBachelor's or master's degree preferred in Electrical Engineering, Computer Engineering, Systems Engineering, Computer Science, or a related technical field.\n\nLOCATION: Secaucus, NJ (Onsite)\n\nTHIS ROLE IS NOT ELEGIBLE FOR VISA SUPPORT\n\n#LI-CS1\nQualifications:\nBenefits offered are described: AMD benefits at a glance.\n\nAMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.\n\nAMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.\n\nThis posting is for an existing vacancy.","company":"Advanced Micro Devices","rawCompany":"advanced micro devices","city":"Secaucus","state":"NJ","isRemote":false,"isActive":false,"createdAt":"2026-07-30T11:33:15.697Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1211.00","title":"Computer Systems Analysts","slug":"computer-systems-analysts"},{"code":"17-2061.00","title":"Computer Hardware Engineers","slug":"computer-hardware-engineers"}],"industries":[{"code":"334111","title":"Electronic Computer Manufacturing","slug":"electronic-computer-manufacturing"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"},{"code":"541513","title":"Computer Facilities Management Services","slug":"computer-facilities-management-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Failure Analysis Engineer","description":"Overview:\nWHAT YOU DO AT AMD CHANGES EVERYTHING\nAt AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career.\nResponsibilities:\nTHE ROLE\nJoin a highly visible engineering team responsible for bringing up, debugging, and improving next-generation server and rack-scale platforms deployed in data center environments. As a Senior Failure Analysis Engineer, you will be among the first engineers to investigate, troubleshoot, and resolve complex system-level issues on new hardware platforms, partnering closely with design, firmware, validation, manufacturing, quality, and infrastructure teams. This role offers the opportunity to work on cutting-edge technologies while driving root cause analysis, platform reliability, and manufacturing readiness across the product lifecycle. You will play a key role in identifying and resolving complex hardware and firmware interactions, supporting platform bring-up activities, improving debug processes, and helping scale organizational knowledge through technical documentation and AI-assisted engineering workflows. This is an ideal opportunity for an engineer who enjoys solving difficult technical problems, working across multiple disciplines, and making a measurable impact on the success of next-generation data center products.\n\nTHE PERSON\nThe ideal candidate is naturally curious, analytical, and motivated by solving complex technical challenges. They thrive in environments where not every answer is immediately available and are comfortable investigating issues across multiple engineering domains to identify root causes and drive resolution. They demonstrate strong communication skills, effectively manage stakeholder expectations during ongoing investigations, and can clearly communicate technical findings to both engineering and cross-functional teams.\nSuccessful candidates are collaborative, adaptable, and persistent problem-solvers who enjoy learning new technologies, working in highly dynamic environments, and continuously improving how engineering teams operate. They embrace new tools and technologies, including AI-assisted workflows, to improve efficiency, accelerate troubleshooting, and scale technical knowledge across the organization.\n\nKEY RESPONSIBILITIES\nPerform system-level and rack-level failure analysis across server and data center platforms, driving issues from initial symptom identification through root cause determination and resolution.\nDebug complex hardware, firmware, and platform-related issues involving CPUs, GPUs, memory, PCIe, networking, power delivery, thermal systems, and firmware interactions.\nAnalyze system logs, platform telemetry, BIOS, BMC, and other diagnostic data to identify failure patterns, isolate root causes, and validate corrective actions.\nReproduce failures in lab environments and utilize appropriate diagnostic tools and methodologies to validate fixes and improve overall platform reliability.\nPartner closely with design, firmware, validation, manufacturing, quality, and infrastructure teams to investigate issues, drive corrective actions, and improve system performance.\nSupport manufacturing and ODM partners by providing technical guidance, failure triage, structured debug processes, and escalation support for complex platform issues.\nDevelop and maintain technical documentation, troubleshooting guides, SOPs, debug methodologies, and knowledge-sharing resources to improve organizational effectiveness.\nDrive root cause analysis activities and contribute to continuous improvements in reliability, serviceability, manufacturability, and operational readiness.\nLeverage AI-assisted tools, automation, and knowledge systems to improve troubleshooting efficiency, accelerate investigations, and scale engineering best practices.\nServe as a technical resource for cross-functional teams by communicating findings, managing stakeholder expectations, and providing clear recommendations based on data-driven analysis.\nContribute expertise in areas such as networking, signal integrity, system architecture, or platform reliability to support the successful deployment of next-generation data center technologies.\n\nPREFERRED EXPERIENCE\nExperience performing system-level, server-level, or rack-level troubleshooting and failure analysis.\nBackground in hardware debugging, root cause analysis, and complex issue resolution.\nFamiliarity with Linux-based environments, platform logs, and diagnostic workflows.\nExperience supporting server, data center, networking, storage, or enterprise hardware platforms.\nExposure to BIOS, BMC, IPMI, firmware interactions, or platform management technologies.\nUnderstanding of networking technologies, signal integrity concepts, power delivery, or high-speed interfaces.\nExperience collaborating with manufacturing, ODM, validation, quality, or design teams.\nAbility to develop technical documentation, SOPs, troubleshooting guides, and knowledge-sharing resources.\nFamiliarity with lab debugging equipment and system diagnostic tools.\nExperience leveraging AI tools, AI agents, automation, or knowledge systems to improve debugging efficiency and technical problem-solving.\nWillingness to occasionally travel in support of manufacturing, validation, or deployment activities.\n\nACADEMIC CREDENTIALS\nBachelor's or master's degree preferred in Electrical Engineering, Computer Engineering, Systems Engineering, Computer Science, or a related technical field.\n\nLOCATION: Secaucus, NJ (Onsite)\n\nTHIS ROLE IS NOT ELEGIBLE FOR VISA SUPPORT\n\n#LI-CS1\nQualifications:\nBenefits offered are described: AMD benefits at a glance.\n\nAMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.\n\nAMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.\n\nThis posting is for an existing vacancy.","datePosted":"2026-07-30T11:33:15.697Z","dateModified":"2026-07-30T11:33:15.697Z","hiringOrganization":{"@type":"Organization","name":"Advanced Micro Devices","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Secaucus","addressRegion":"NJ","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"512a5688caf0c8209dfeb9d0"},"url":"https://jobsearcher.com/jobs/512a5688caf0c8209dfeb9d0"}}