Senior Datacenter Platform/Debug Engineer
Join AMD's Datacenter Platform Engineering GroupAt AMD, we believe technology has the power to solve the world's most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future.Whether you're designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger — technology that moves the world forward. Join us and, together, we'll advance your career.Join AMD's Datacenter Platform Engineering Group (DPEG) and help support the deployment, availability, and operational success of next-generation AI and HPC infrastructure. As a Platform Systems Engineer, you will work on cutting-edge GPU and server platforms, partnering with hardware, firmware, software, validation, and datacenter engineering teams to troubleshoot complex system issues and ensure reliable operation of large-scale compute environments.This role offers the opportunity to work directly with advanced datacenter technologies, participate in system bring-up and deployment activities, and become a key contributor in resolving critical platform-level issues. Ideal candidates enjoy solving challenging technical problems, collaborating across multiple engineering disciplines, and making a direct impact on the success of AMD's datacenter infrastructure.The ideal candidate is a hands-on systems engineer who enjoys deep technical troubleshooting and thrives in fast-paced datacenter environments. They possess strong analytical skills, can quickly isolate and resolve complex issues, and are comfortable working across hardware, firmware, and software layers of a system.Successful candidates will demonstrate:Strong troubleshooting and root cause analysis skillsExcellent communication and collaboration abilitiesA proactive and self-driven approach to problem solvingAbility to mentor and guide junior engineersStrong documentation and organizational skillsComfort operating in highly technical and mission-critical environmentsA passion for learning new technologies and solving complex engineering challengesKey responsibilities include:Support datacenter deployments and help maintain the availability and uptime of large-scale compute systems.Perform system-level debugging and triage across hardware, firmware, software, and operating system layers.Investigate and resolve complex platform issues impacting GPU and server infrastructure.Support system bring-up, initialization, validation, and operational readiness activities.Utilize industry-standard debug tools and diagnostic methods to identify root causes.Provide technical leadership and guidance to junior engineers during troubleshooting activities.Document debug methodologies, troubleshooting procedures, and best practices.Collaborate with cross-functional engineering teams to drive issue resolution and continuous improvement.Preferred experience includes:System-level hardware, firmware, and software debugging experienceDatacenter, server, HPC, or AI infrastructure environmentsRoot cause analysis and triage of complex platform issuesGPU, PCIe, memory, retimer, networking, and system architecture knowledgeRAS (Reliability, Availability, Serviceability) concepts and methodologiesLinux operating system experiencePython, Bash, or similar scripting experienceHands-on experience with industry-standard debug tools and diagnosticsServer bring-up, system initialization, and validation activitiesTechnical leadership, mentoring, and cross-functional collaborationAcademic credentials:Bachelor's or Master's degree preferred in Computer Engineering, Electrical Engineering, Computer Science, or a related technical disciplineLocation: Rockdale, Texas | 100% OnsiteThis role is not eligible for visa support.