DCGPU Platform System Manager
Dcgpu Platform System ManagerAt AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you'll discover the real differentiator is our culture. We push the limits of innovation to solve the world's most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career.The data center platform engineering group (DPEG) manager that will be the primary on-site leader for a data center in north Austin, Texas. This individual will need to be on-site daily driving day-to-day activities that will include the deployment and availability of highly complex AMD instinct platforms that will be the backbone of the work required to release state-of-the-art technologies. This manager will have personnel responsibilities and be required to manage the work required to run a data center with well over 1000 systems.Experienced, self-motivated individual that has previously managed a data center preferably in the GPU system space w/500+ systems/platform. Person should be able to communicate updates on the state of the fleet, ensure the team is working on deployment/availability of the fleet, as well be able to solve technical issues that arise.Key responsibilities:On-site management of a data center site with a vast array of differing platforms/systemPersonnel and work assigned management of data center engineers and techniciansClear / concise communication in open daily meetings on "high attention" given to system availabilityKnowledge in requirements for root causing issues and understanding/mapping debug methodologiesProvide leadership input/recommendations for improvements / help drive organizational initiativesPreferred experience:Experience managing GPU data center employees, systems, day-to-day activitiesKnowledge in solving issues around capacity planning, power, thermal, networking, clusteringExecutive-level focused communicationManagement experience in GPU data centers that have a vast array of different systems/platformsCo-work with external stakeholders/vendors and ability to openly communicate/drive issuesCollaborate with internal stakeholders on root causing issues, driving issue meetingsHands-on experience with day-to-day issues that arise in a data center (capacity, network, power, thermal)Leadership and communication skills that require presentations in executive forumAbility to clearly articulate the work by the team (ins/outs), needs, and recruit necessary skillsAcademic credentials:Bachelors degree in engineeringThis role is not eligible for visa sponsorship.